Staff ML Systems Engineer - GPU & Performance

Google

Sunnyvale (CA)

On-site

USD 207,000 - 300,000

Full time

5 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Equity
Performance bonus

Job summary

Google is seeking a senior software engineer for the Core ML GPU Performance team in Sunnyvale to optimize and model GPU systems for large-scale ML workloads. You will help drive performance improvements across production and research pipelines, collaborating with cross-functional teams to scale LLMs on accelerator hardware.

You will work on architecture-level optimizations, benchmarking, and tooling to improve efficiency and throughput, shaping the future of ML infrastructure at Google.

Qualifications

  • Bachelor's degree or equivalent practical experience.
  • 8 years of experience in software development.
  • 5 years of experience with ML design and ML infrastructure (e.g., model deployment, model evaluation, data processing, debugging, fine tuning).
  • Experience with modern GPU architectures, memory hierarchies, and performance bottlenecks.
  • Experience with low-level GPU programming (CUDA, Triton, CUTLASS, etc.) and performance engineering techniques.
  • Experience with modern LLMs and their deployment on AI accelerators.

Responsibilities

  • Identify and optimize LLM training and serving benchmarks; drive XLA:GPU/Triton performance and guide XLA releases.
  • Partner with product teams to onboard, optimize, and scale LLMs on GPU hardware.
  • Conduct architecture-level simulations, performance benchmarking, and roofline analyses to guide system designs.
  • Analyze fleet-wide performance and efficiency metrics to identify bottlenecks and engineer scalable optimizations.
  • Research and implement model/data efficiency techniques and profiling mechanisms to improve workload performance.

Skills

Software development
ML design & infrastructure
GPU architectures
Low-level GPU programming
LLM deployment

Education

Bachelor's degree or equivalent practical experience
Master's degree or PhD (preferred)

Tools

OpenXLA

Job description

Google is seeking a senior software engineer for the Core ML GPU Performance team in Sunnyvale to optimize and model GPU systems for large-scale ML workloads. You will help drive performance improvements across production and research pipelines, collaborating with cross-functional teams to scale LLMs on accelerator hardware.

You will work on architecture-level optimizations, benchmarking, and tooling to improve efficiency and throughput, shaping the future of ML infrastructure at Google.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff ML Performance Engineer - GPU Optimization
Staff ML Performance Engineer - GPU Optimization

Google LLC • Sunnyvale (CA)

On-site
USD 186,000 - 228,000
Performance SOC Architect - GPU/ML Profiling & Optimization
Performance SOC Architect - GPU/ML Profiling & Optimization

Google LLC • Sunnyvale (CA)

On-site
USD 174,000 - 252,000
Staff Software Engineer, ML Performance, GPU
Staff Software Engineer, ML Performance, GPU

Google • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Equity
Performance bonus
ML Performance Engineering Manager (TPU & Optimization)
ML Performance Engineering Manager (TPU & Optimization)

Google • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Staff ML Systems Engineer — TPU Performance & Gemini
Staff ML Systems Engineer — TPU Performance & Gemini

AI Chopping Block, Inc. • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Engineering Manager, ML Performance
Engineering Manager, ML Performance

Google • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Staff ML Compiler Performance Engineer – Co-Design Lead
Staff ML Compiler Performance Engineer – Co-Design Lead

Google LLC • Mountain View (CA)

On-site
USD 207,000 - 300,000
Senior AI/ML Software Architect — GPU Infrastructure
Senior AI/ML Software Architect — GPU Infrastructure

Google • Seattle (WA)

On-site
USD 262,000 - 364,000
Health insurance
Dental insurance
Vision insurance
+8
Staff ML Systems Architect - Co-Design & TPU Performance
Staff ML Systems Architect - Co-Design & TPU Performance

Google • Mountain View (CA)

On-site
USD 207,000 - 300,000
Staff Software Engineer, TPU Performance & ML Infrastructure
Staff Software Engineer, TPU Performance & ML Infrastructure

Google • New York (NY)

On-site
USD 207,000 - 300,000