Software Engineer, On-Device Machine Learning

Google

New York (NY)

On-site

USD 151,000 - 252,000

Full time

2 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Google is seeking a software engineer to work on LiteRT, an on-device AI framework designed to maximize performance and efficiency across mobile, desktop, and embedded devices. You will contribute to state-of-the-art ML acceleration and deployment strategies, collaborating with multiple product teams.

The role focuses on building cross‑platform infrastructure for Google products, enabling on‑device AI with scalable performance and device flexibility.

Qualifications

  • Bachelor’s degree or equivalent practical experience is required.
  • Five years of software development experience in one or more programming languages.
  • Three years of experience with ML infrastructure (deployment, evaluation, data processing, debugging).
  • Two years of experience developing compiler technology.
  • Experience in mobile development is required.

Responsibilities

  • Collaborate with peers and stakeholders through design and code reviews to ensure best practices.
  • Design and implement solutions in one or more ML areas, leverage ML infrastructure, and demonstrate expertise.
  • Develop LiteRT, Google's on-device AI framework for first- and third-party use with hardware acceleration.
  • Enable on-device deployment of key models across accelerators on Android, Chrome, iOS, desktop, and more.
  • Improve performance of on-device model inference via optimizations in the runtime and kernels.

Skills

Software development
ML infrastructure
Mobile development
Compiler development
Programming languages

Education

Bachelor’s degree or equivalent

Tools

PyTorch
JAX
TensorFlow
TensorFlow Lite
ExecuTorch
Core ML
SNPE/QNN

Job description

Note: By applying to this position you will have an opportunity to share your preferred working location from the following: Sunnyvale, CA, USA; New York, NY, USA.

Minimum qualifications:
  • Bachelor’s degree or equivalent practical experience.
  • 5 years of experience with software development in one or more programming languages.
  • 3 years of experience with ML infrastructure (e.g., model deployment, model evaluation, optimization, data processing, debugging).
  • 2 years of experience developing compilers.
  • Experience in mobile development.
Preferred qualifications:
  • Experience leading and delivering successful ML projects focused on on-device deployment (Android, iOS, web browsers, or embedded devices).
  • Experience in ML frameworks (e.g., PyTorch, JAX, TensorFlow).
  • Experience with on-device ML software development kits (SDKs)/tooling (e.g., TensorFlow Lite, ExecuTorch, Core ML, SNPE/QNN).
  • Understanding of Generative AI model architectures and their optimization for on-device execution.
  • Excellent communication and collaboration skills.
  • Passion for innovation and a strong desire to push the boundaries of what's possible with on-device ML.
About The Job

Google's software engineers develop the next-generation technologies that change how billions of users connect, explore, and interact with information and one another. Our products need to handle information at massive scale, and extend well beyond web search. We're looking for engineers who bring fresh ideas from all areas, including information retrieval, distributed computing, large-scale system design, networking and data storage, security, artificial intelligence, natural language processing, UI design and mobile; the list goes on and is growing every day. As a software engineer, you will work on a specific project critical to Google’s needs with opportunities to switch teams and projects as you and our fast-paced business grow and evolve. We need our engineers to be versatile, display leadership qualities and be enthusiastic to take on new problems across the full-stack as we continue to push technology forward.

LiteRT is Google’s next generation on-device AI framework, succeeding TensorFlow Lite (TFLite). It is designed to maximize the performance, efficiency, and portability of ML models on a wide array of edge devices, from mobile phones to embedded systems. LiteRT significantly upgrades GPU acceleration and introduces native NPU acceleration, while maintaining and enhancing the robust CPU performance inherited from TFLite.

LiteRT enables developers and Google products to deploy AI across mobile, web, desktop, and embedded. Our team focuses on building cross‑platform infrastructure aligned with Google's business needs, serving top Google products (Android, Chrome, Photos, Meet, Youtube, etc), third‑party developers, and specialized Pixel solutions. Our goal is to provide on‑device AI infrastructure with exceptional performance, enabling framework and device flexibility at scale.

Google Cloud accelerates every organization’s ability to digitally transform its business and industry. We deliver enterprise‑grade solutions that leverage Google’s cutting‑edge technology, and tools that help developers build more sustainably. Customers in more than 200 countries and territories turn to Google Cloud as their trusted partner to enable growth and solve their most critical business problems.

Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

US: $174000 - $252000 (USD) + 15% bonus target + equity + benefits

Responsibilities

Learn more about benefits at Google .

  • Collaborate with peers and stakeholders through design and code reviews to ensure best practices amongst available technologies (e.g., style guidelines, checking code in, accuracy, testability, and efficiency).
  • Design and implement solutions in one or more specialized ML areas, leverage ML infrastructure, and demonstrate expertise in a chosen field.
  • Develop LiteRT, Google's on-device AI framework for first- and third-party, enabling state-of-the-art (SOTA) hardware acceleration and use cases on edge platforms.
  • Enable on-device deployment of key models, such as Gemini Nano and Gemma, across various accelerators (GPU/Pixel TPU/NPUs/CPU) on Android, Chrome, iOS, desktop, and more.
  • Improve performance of on‑device model inference via optimizations in the model symbol, on‑device runtime and kernel implementation.

Google is proud to be an equal opportunity workplace and is an affirmative action employer. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status. We also consider qualified applicants regardless of criminal histories, consistent with legal requirements. See also Google's EEO Policy and EEO is the Law. If you have a disability or special need that requires accommodation, please let us know by completing our Accommodations for Applicants form .

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, On-Device Machine Learning
Software Engineer, On-Device Machine Learning

Google • Sunnyvale (CA)

On-site
USD 174,000 - 252,000
Software Engineer, On-Device Machine Learning
Software Engineer, On-Device Machine Learning

Google Inc. • Sunnyvale (CA)

On-site
USD 174,000 - 253,000
Equity
Benefits
Senior Staff Software Engineer, AI Edge
Senior Staff Software Engineer, AI Edge

Google • Sunnyvale (CA)

On-site
USD 262,000 - 365,000
Senior Staff Software Engineer, AI Edge
Senior Staff Software Engineer, AI Edge

Google • Sunnyvale (CA)

On-site
USD 262,000 - 365,000
Staff Software Engineer, AI Engines 3P TPU Inference
Staff Software Engineer, AI Engines 3P TPU Inference

Google Inc. • Mountain View (CA)

On-site
USD 207,000 - 300,000
Staff Software Engineer, TPU Performance
Staff Software Engineer, TPU Performance

Google • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Staff AI Engines Engineer - TPU Inference & ML Runtime
Staff AI Engines Engineer - TPU Inference & ML Runtime

Google Inc. • Mountain View (CA)

On-site
Staff Software Engineer, AI Engines 3P TPU Inference
Staff Software Engineer, AI Engines 3P TPU Inference

Socket.dev • Mountain View (CA)

On-site
USD 207,000 - 300,000
Staff Software Engineer, TPU Performance
Staff Software Engineer, TPU Performance

Google • Kirkland (WA)

On-site
USD 207,000 - 300,000
Health insurance
Dental, Vision, Life, Disability
401(k) with company match
+5
Staff Software Engineer, ML Systems Co-Design
Staff Software Engineer, ML Systems Co-Design

Google Inc. • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Health, dental, vision, life, and(dis)
401(k) with company match
Paid Time Off: 20 days/year
+4