Staff Software Engineer, ML Performance, GPU

Google LLC

Sunnyvale (CA)

On-site

USD 186,000 - 228,000

Full time

4 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Google is seeking a Staff Software Engineer for ML Performance on GPU to drive high-performance ML workloads and next-gen GPU architectures. You will work across model deployment, evaluation, and data processing, ensuring peak efficiency on large-scale systems.

The role requires hands-on CUDA/Triton/CUTLASS experience, deep knowledge of ML workloads, and collaboration with cross-functional teams to push performance forward.

Qualifications

  • Bachelor’s degree or equivalent practical experience.
  • 8 years of software development experience.
  • 5 years of ML design/infrastructure experience.
  • GPU architectures and performance bottlenecks expertise.
  • Low-level GPU programming (CUDA, Triton, CUTLASS) experience.
  • Experience with deployment of LLMs on AI accelerators.

Responsibilities

  • Identify and maintain LLM training and serving benchmarks; use them to identify performance opportunities and guide optimizations across GPUs.
  • Partner with product teams to onboard, optimize, and scale ML models on GPU hardware.
  • Conduct architecture-level simulations and roofline analyses to guide system designs.
  • Analyze fleet-wide performance metrics to drive scalable optimizations across Google's infrastructure.
  • Research and implement model/data efficiency techniques and profiling mechanisms.

Skills

ML design infra
GPU architectures
Low-level GPU programming
LLM deployment
Performance engineering

Education

Bachelor’s degree or equivalent practical experience
Master’s or PhD in Engineering/CS or related field

Tools

CUDA
Triton
CUTLASS
OpenXLA
XLA
Roofline analysis tools

Job description

Staff Software Engineer, ML Performance, GPU

Share Staff Software Engineer, ML Performance, GPU

corporate_fare Google place Sunnyvale, CA, USA

X In most instances, this position requires in-person interviews as part of the hiring process.

  • Bachelor’s degree or equivalent practical experience.
  • 8 years of experience in software development.
  • 5 years of experience with ML design and ML infrastructure (e.g., model deployment, model evaluation, data processing, debugging, fine tuning).
  • Experience with modern GPU architectures, memory hierarchies, and performance bottlenecks.
  • Experience with low-level GPU programming (CUDA, Triton, CUTLASS, etc.) and performance engineering techniques.
  • Experience with modern LLMs and their deployment on AI accelerators.
Preferred qualifications:
  • Master’s degree or PhD in Engineering, Computer Science, or a related technical field.
  • 8 years of experience with data structures and algorithms.
  • 3 years of experience in a technical leadership role leading project teams and setting technical direction.
  • 3 years of experience working in a complex, matrixed organization involving cross-functional, or cross-business projects.
  • Experience in hardware-aware algorithm design and compiler stacks (e.g., OpenXLA), tailoring large-scale ML models and distributed systems for peak performance across accelerator hardware.
About the job

Google's software engineers develop the next-generation technologies that change how billions of users connect, explore, and interact with information and one another. Our products need to handle information at massive scale, and extend well beyond web search. We're looking for engineers who bring fresh ideas from all areas, including information retrieval, distributed computing, large-scale system design, networking and data storage, security, artificial intelligence, natural language processing, UI design and mobile; the list goes on and is growing every day. As a software engineer, you will work on a specific project critical to Google’s needs with opportunities to switch teams and projects as you and our fast-paced business grow and evolve. We need our engineers to be versatile, display leadership qualities and be enthusiastic to take on new problems across the full-stack as we continue to push technology forward.

While known for pioneering work with TPUs, GPUs are an equally vital and rapidly expanding frontier within Google's ML infrastructure. GPUs are indispensable to Google’s ever-evolving landscape for strategic, pragmatic, and performance-driven reasons — ensuring top performance for our ML models, adapting to ML workloads, achieving results, and influencing next-gen GPU architectures via partnerships.
Core ML's GPU Performance team is responsible for optimizing, modeling, and evaluating GPU systems for comparative analysis and benchmarking for internal and external ML workloads. Our team’s focus on performance analysis and optimization identifies opportunities in Google production and research ML workloads and lands optimizations to entire fleet. We evaluate current and future ML workloads and runs performance/total cost of ownership simulations to collect roofline estimates and guide decision-making for the hardware teams.

Behind everything our users see online is the architecture built by the Technical Infrastructure team to keep it running. From developing and maintaining our data centers to building the next generation of Google platforms, we make Google's product portfolio possible. We're proud to be our engineers' engineers and love voiding warranties by taking things apart so we can rebuild them. We keep our networks up and running, ensuring our users have the best and fastest experience possible.

Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

About the job

Google's software engineers develop the next-generation technologies that change how billions of users connect, explore, and interact with information and one another. Our products need to handle information at massive scale, and extend well beyond web search. We're looking for engineers who bring fresh ideas from all areas, including information retrieval, distributed computing, large-scale system design, networking and data storage, security, artificial intelligence, natural language processing, UI design and mobile; the list goes on and is growing every day. As a software engineer, you will work on a specific project critical to Google’s needs with opportunities to switch teams and projects as you and our fast-paced business grow and evolve. We need our engineers to be versatile, display leadership qualities and be enthusiastic to take on new problems across the full-stack as we continue to push technology forward.

While known for pioneering work with TPUs, GPUs are an equally vital and rapidly expanding frontier within Google's ML infrastructure. GPUs are indispensable to Google’s ever-evolving landscape for strategic, pragmatic, and performance-driven reasons — ensuring top performance for our ML models, adapting to ML workloads, achieving results, and influencing next-gen GPU architectures via partnerships.
Core ML's GPU Performance team is responsible for optimizing, modeling, and evaluating GPU systems for comparative analysis and benchmarking for internal and external ML workloads. Our team’s focus on performance analysis and optimization identifies opportunities in Google production and research ML workloads and lands optimizations to entire fleet. We evaluate current and future ML workloads and runs performance/total cost of ownership simulations to collect roofline estimates and guide decision-making for the hardware teams.

Behind everything our users see online is the architecture built by the Technical Infrastructure team to keep it running. From developing and maintaining our data centers to building the next generation of Google platforms, we make Google's product portfolio possible. We're proud to be our engineers' engineers and love voiding warranties by taking things apart so we can rebuild them. We keep our networks up and running, ensuring our users have the best and fastest experience possible.

Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

US: $207000 - $300000 (USD) + 20% bonus target + equity + benefits

Learn more about benefits at Google .

  • Identify and maintain LLM training and serving benchmarks; use them to identify performance opportunities, drive XLA:GPU/Triton performance and guide XLA releases.
  • Partner with product teams (e.g., Google DeepMind) to onboard, optimize, and scale LLMs and machine learning models on GPU hardware.
  • Conduct architecture-level simulations, performance benchmarking, and roofline analyses using tools like TRT-LLM, vLLM, and SGLang to guide system designs.
  • Analyze fleet-wide performance and efficiency metrics to identify bottlenecks and engineer scalable optimizations across Google's infrastructure.
  • Research and implement model/data efficiency techniques, tooling, and profiling mechanisms to improve workload performance and training efficiency.

Google is proud to be an equal opportunity and affirmative action employer. We are committed to building a workforce that is representative of the users we serve, creating a culture of belonging, and providing an equal employment opportunity regardless of race, creed, color, religion, gender, sexual orientation, gender identity/expression, national origin, disability, age, genetic information, veteran status, marital status, pregnancy or related condition (including breastfeeding), expecting or parents-to-be, criminal histories consistent with legal requirements, or any other basis protected by law. See also Google's EEO Policy , Know your rights: workplace discrimination is illegal , Belonging at Google , and How we hire .

Google is a global company and, in order to facilitate efficient collaboration and communication globally, English proficiency is a requirement for all roles unless stated otherwise in the job posting.

To all recruitment agencies: Google does not accept agency resumes. Please do not forward resumes to our jobs alias, Google employees, or any other organization location. Google is not responsible for any fees related to unsolicited resumes.

Equity is granted exclusively and discretionarily by Alphabet Inc. on the basis of an agreement concluded between you and Alphabet Inc. Alphabet Inc. is your sole contractual partner with respect to equity grants. GSU grants are not guaranteed, are discretionary, are subject to approval by the Alphabet Inc. board of directors or its delegate, the terms of the relevant Alphabet Inc. stock plan, and your grant agreement. They have no impact on statutory payments. Current or past grants do not confer an acquired right.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Software Engineer, TPU, Performance
Staff Software Engineer, TPU, Performance

Google • New York (NY)

On-site
USD 207,000 - 300,000
Staff Software Engineer, TPU, Performance
Staff Software Engineer, TPU, Performance

Google • Ionia (NY)

On-site
USD 207,000 - 300,000
Staff Software Engineer, ML Systems Co-Design
Staff Software Engineer, ML Systems Co-Design

Google Inc. • Mountain View (CA)

On-site
USD 207,000 - 300,000
Staff Software Engineer, ML Systems Co-Design
Staff Software Engineer, ML Systems Co-Design

Google Inc. • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Health, dental, vision, life, and(dis)
401(k) with company match
Paid Time Off: 20 days/year
+4
Software Engineering Manager, GPU Reliability
Software Engineering Manager, GPU Reliability

Google Inc. • Seattle (WA)

On-site
USD 207,000 - 300,000
Health insurance
401(k) match
Paid time off
+4
Software Engineer III, TPU Performance, Hardware and Software Codesign
Software Engineer III, TPU Performance, Hardware and Software Codesign

Google Inc. • Sunnyvale (CA)

On-site
USD 147,000 - 210,000
Senior Staff Software Engineer, GPU System Software
Senior Staff Software Engineer, GPU System Software

Google • Sunnyvale (CA)

On-site
USD 262,000 - 364,000
Equity grants
Bonus target
Staff Software Engineer, ML Systems Co-Design
Staff Software Engineer, ML Systems Co-Design

Google • Town of Montana (WI)

On-site
USD 207,000 - 300,000
Health insurance
401(k) with company match
Paid time off
+4
Staff Software Engineer, ML Frameworks
Staff Software Engineer, ML Frameworks

Google Inc. • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Senior Staff Software Engineer, AI/ML, Google Cloud
Senior Staff Software Engineer, AI/ML, Google Cloud

Google • Seattle (WA)

On-site
USD 262,000 - 364,000
Health insurance
Dental insurance
Vision insurance
+8