Staff Software Engineer, TPU, Performance

Google

New York (NY)

On-site

USD 207,000 - 300,000

Full time

8 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Google is seeking a Staff Software Engineer, TPU, Performance to advance ML performance on TPU systems and contribute to Gemini and OSS ML models. You will work on optimizing model architectures, tuning compilers and runtime for large-scale ML workloads, and collaborating with researchers to harness TPUs at scale.

Role involves strong ownership, stakeholder influence, and hands-on work with ML design, ML infrastructure, and performance tuning across GPU/TPU architectures.

Qualifications

  • Bachelor’s degree or equivalent practical experience.
  • 8+ years of software development experience, with 5+ years in ML/AI domains.
  • 5+ years designing and deploying ML models and ML infrastructure.

Responsibilities

  • Identify and maintain ML training and serving benchmarks representative of Google production.
  • Drive performance improvements for ML workloads on TPUs and GPUs.
  • Collaborate with product teams and researchers to onboard new ML models on TPU hardware and optimize for large-scale training.
  • Analyze performance metrics to identify bottlenecks and implement scalable solutions.
  • Engage with cross-functional teams to push core ML infrastructure and compiler/runtime optimizations.

Skills

Software dev
ML experience
Speech/Audio
Reinforcement learning
ML infrastructure
GPU acceleration

Education

Bachelor's degree or equivalent practical experience
Master’s degree or PhD in Engineering/CS

Tools

TensorFlow/TPU
OpenXLA
MLIR
CUDA

Job description

Staff Software Engineer, TPU, Performance

Google - Sunnyvale, CA, USA; New York, NY, USA

Advanced

Experience owning outcomes and decision making, solving ambiguous problems and influencing stakeholders; deep expertise in domain.

In most instances, this position requires in-person interviews as part of the hiring process. Note: By applying to this position you will have an opportunity to share your preferred working location from the following: Sunnyvale, CA, USA; New York, NY, USA.

Minimum qualifications:
  • Bachelor’s degree or equivalent practical experience.
  • 8 years of experience in software development.
  • 5 years of experience with one or more of the following: speech/audio (e.g., technology duplicating and responding to the human voice), reinforcement learning (e.g., sequential decision making), ML infrastructure, or specialization in another ML field.
  • 5 years of experience with ML design and ML infrastructure (e.g., model deployment, model evaluation, data processing, debugging, fine tuning).
  • 5 years of experience testing, and launching software products, and 3 years of experience with software design and architecture.
Preferred qualifications:
  • Master’s degree or PhD in Engineering, Computer Science, or a related technical field.
  • 8 years of experience with data structures and algorithms.
  • Experience with machine learning, compiler optimization, code generation, and runtime systems for GPU architectures (OpenXLA, MLIR, Triton, etc).
  • Experience in tailoring algorithms and ML models to exploit ML accelerator architecture strengths and minimize weaknesses.
  • Experience in low-level GPU programming (CUDA, OpenCL, etc.) and performance tuning techniques.
  • Understanding of modern Graphics Processing Unit (GPU), TPU or other ML accelerator architectures, memory hierarchies, and performance bottlenecks.
About the job

Google's software engineers develop the next-generation technologies that change how billions of users connect, explore, and interact with information and one another. Our products need to handle information at massive scale, and extend well beyond web search. We're looking for engineers who bring fresh ideas from all areas, including information retrieval, distributed computing, large-scale system design, networking and data storage, security, artificial intelligence, natural language processing, UI design and mobile; the list goes on and is growing every day. As a software engineer, you will work on a specific project critical to Google’s needs with opportunities to switch teams and projects as you and our fast-paced business grow and evolve. We need our engineers to be versatile, display leadership qualities and be enthusiastic to take on new problems across the full-stack as we continue to push technology forward.

Google’s Core Machine Learning (ML) organization is seeking software engineers to join the team known for pioneering work with Tensor Processing Units (TPUs). In this role, you will work on Gemini, as well as industry leading open-source models, to understand model architecture and optimize the performance of these Machine Learning (ML) models on TPU systems for both Just After eXecution (JAX) and PyTorch platforms. You will improve the performance of ever-evolving ML workloads, achieving results. These fundamental efforts will influence next-generation (next-gen) TPU architectures via partnerships, ensuring performance for Gemini and Open-Source Software (OSS) Machine Learning (ML) models. The Core team builds the technical foundation behind Google’s flagship products. We are owners and advocates for the underlying design elements, developer platforms, product components, and infrastructure at Google. These are the essential building blocks for excellent, safe, and coherent experiences for our users and drive the pace of innovation for every developer. We look across Google’s products to build central solutions, break down technical barriers and strengthen existing systems. As the Core team, we have a mandate and a unique opportunity to impact important technical decisions across the company. Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

US: $207000 - $300000 (USD) + 20% bonus target + equity + benefits

Learn more about benefits at Google.

Responsibilities
  • Identify and maintain ML training and serving benchmarks that are representative to Google production and broader ML industry.
  • Achieve performance for customer launches, and in case of third-party/Open-Source Software (3P/OSS) models, for engaged benchmark submissions ML commons, InferenceMAX, etc.).
  • Use the benchmarks to identify performance opportunities and drive out-of-the-box performance toward improving the compiler, runtime, etc., in collaboration with those teams.
  • Engage with Google Product teams and researchers to solve their performance problems (e.g., onboard new ML models and products on Google new TPU hardware, enabling larger models (giant models) to train efficiently on a very large-scale (i.e., thousands of TPUs).
  • Analyze performance and efficiency metrics to identify bottlenecks, design, and implement solutions at Google fleet-wide scale.

Information collected and processed as part of your Google Careers profile, and any job applications you choose to submit is subject to Google's Applicant and Candidate Privacy Policy.

Google is proud to be an equal opportunity and affirmative action employer. We are committed to building a workforce that is representative of the users we serve, creating a culture of belonging, and providing an equal employment opportunity regardless of race, creed, color, religion, gender, sexual orientation, gender identity/expression, national origin, disability, age, genetic information, veteran status, marital status, pregnancy or related condition (including breastfeeding), expecting or parents‑to-be, criminal histories consistent with legal requirements, or any other basis protected by law. See also Google's EEO Policy, Know your rights: workplace discrimination is illegal, Belonging at Google, and How we hire.

If you have a need that requires accommodation, please let us know by completing our Accommodations for Applicants form.

Google is a global company and, in order to facilitate efficient collaboration and communication globally, English proficiency is a requirement for all roles unless stated otherwise in the job posting.

To all recruitment agencies: Google does not accept agency resumes. Please do not forward resumes to our jobs alias, Google employees, or any other organization location. Google is not responsible for any fees related to unsolicited resumes.

Equity is granted exclusively and discretionarily by Alphabet Inc. on the basis of an agreement concluded between you and Alphabet Inc. Alphabet Inc. is your sole contractual partner with respect to equity grants. GSU grants are not guaranteed, are discretionary, are subject to approval by the Alphabet Inc. board of directors or its delegate, the terms of the relevant Alphabet Inc. stock plan, and your grant agreement. They have no impact on statutory payments. Current or past grants do not confer an acquired right.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Software Engineer, TPU, Performance
Staff Software Engineer, TPU, Performance

Google • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Equity
Bonus target 20%
Company benefits
Staff Software Engineer, ML Systems Co-Design
Staff Software Engineer, ML Systems Co-Design

Google Inc. • Mountain View (CA)

On-site
USD 207,000 - 300,000
Software Engineer III, TPU Performance, Hardware and Software Codesign
Software Engineer III, TPU Performance, Hardware and Software Codesign

Google Inc. • Sunnyvale (CA)

On-site
USD 147,000 - 210,000
Senior Software Engineer, Fleet-level ML Performance
Senior Software Engineer, Fleet-level ML Performance

Google • United States

On-site
USD 174,000 - 252,000
Staff Software Engineer, ML Systems Co-Design
Staff Software Engineer, ML Systems Co-Design

Google Inc. • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Health, dental, vision, life, and(dis)
401(k) with company match
Paid Time Off: 20 days/year
+4
Staff Software Engineer, ML Systems Co-Design
Staff Software Engineer, ML Systems Co-Design

Google • Town of Montana (WI)

On-site
USD 207,000 - 300,000
Health insurance
401(k) with company match
Paid time off
+4
Senior Performance Co-Design Engineer, TPU
Senior Performance Co-Design Engineer, TPU

Google • Sunnyvale (CA)

On-site
USD 174,000 - 252,000
Senior Staff Software Engineer, Machine Learning, ML Training
Senior Staff Software Engineer, Machine Learning, ML Training

Google • Town of Montana (WI)

On-site
USD 262,000 - 364,000
Health/dental/vision
401(k) match
Paid time off
+4
Senior Staff Performance Codesign Engineer, TPU
Senior Staff Performance Codesign Engineer, TPU

Google • Town of Montana (WI)

On-site
USD 240,000 - 333,000
Equity
Benefits
Tech Lead Manager, ML Accelerator Fleet Efficiency
Tech Lead Manager, ML Accelerator Fleet Efficiency

Google • Sunnyvale (CA)

On-site
USD 207,000 - 300,000