Machine Learning Engineer, GPU Kernel and Runtime

Waymo

Mountain View (CA)

On-site

USD 213,000 - 263,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Annual bonus
Equity incentive
Company benefits

Job summary

Waymo is an autonomous driving technology company hiring for a role within the ML Infrastructure team. The position focuses on optimizing ML workloads, CUDA kernel development, and NVIDIA stack integration across perception, behavior, and planning models.

Candidates should have 5+ years in system performance or ML compilers, with strong C++/CUDA skills and experience with NVIDIA GPUs. Waymo offers a competitive compensation package and equity incentives.

Qualifications

  • B.S. or M.S. in CS, EE, Deep Learning or a related field.
  • 5+ years of industry experience on system performance, hardware-level GPU optimization, or ML compilers.
  • Strong C++ and CUDA programming skills.
  • Extensive experience in NVIDIA GPU Kernel development to accelerate deep learning models.
  • Proven debugging and optimization experience on the XLA:GPU compiler, as well as the NVIDIA runtime stack.
  • Passion for developing and optimizing ML software stacks for modern ML accelerator architectures (framework, runtime library, ML compiler, efficient deep learning etc.)

Responsibilities

  • Collaborate with ML practitioners on models for perception, behavior prediction, and planning, to understand their models and accelerate them onboard through custom NVIDIA GPU kernel development.
  • Deep dive into the NVIDIA ML software and runtime stack, from custom CUDA ops to the XLA:GPU compiler and low-level libraries. Analyze numeric behaviors, debug complex compiler issues, and ensure inference results are stable and consistent. Develop tools/system software for optimal resource usage, hardware efficiency, and platform reliability in an ML serving system.
  • Analyze ML workload performance at the hardware level; apply manual and AI agent-assisted techniques and develop highly optimized, custom CUDA/Triton operator libraries tailored to Waymo’s specific architectures.
  • Build tools to benchmark, profile GPU execution, and productize deep learning models for a streamlined and robust onboard and offboard deployment.

Skills

C++
CUDA
GPU kernel
XLA:GPU
NVIDIA runtime
ML software stacks

Education

B.S./M.S. in CS/EE/Deep Learning
M.Sc or PhD in Computer Science/Mathematics

Tools

Nsight Compute
cuda-gdb
Triton
MLIR
CuTe DSL
cuTile

Job description

Waymo is an autonomous driving technology company with the mission to be the world’s most trusted driver. Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on building the Waymo Driver—The World’s Most Experienced Driver™—to improve access to mobility while saving thousands of lives now lost to traffic crashes. The Waymo Driver powers Waymo’s fully autonomous ride-hail service and can also be applied to a range of vehicle platforms and product use cases. The Waymo Driver has provided over ten million rider-only trips, enabled by its experience autonomously driving over 100 million miles on public roads and tens of billions in simulation across 15+ U.S. states.

The Waymo ML Infrastructure team accelerates Waymo’s mission, by building the best ecosystem for sustainably innovating and shipping ML powered intelligence. Research, Production, and the Hardware teams are our primary stakeholders and our work powers the development of the state of the art models in the areas of Perception and Trajectory planning that are core to our autonomous driving software. We enable our partners by offering the best in class solutions for the entire model development lifecycle. These solutions include understanding the model business goals and platform hardware characteristics, and codesign the models for the hardwares. These solutions are developed in close collaboration with teams at different modeling teams. Scale and efficiency are core tenets our infra follows.

You Will
  • Collaborate with ML practitioners on models for perception, behavior prediction, and planning, to understand their models and accelerate them onboard through custom NVIDIA GPU kernel development.
  • Deep dive into the NVIDIA ML software and runtime stack, from custom CUDA ops to the XLA:GPU compiler and low-level libraries. Analyze numeric behaviors, debug complex compiler issues, and ensure inference results are stable and consistent. Develop tools/system software for optimal resource usage, hardware efficiency, and platform reliability in an ML serving system.
  • Analyze ML workload performance at the hardware level; apply manual and AI agent-assisted techniques and develop highly optimized, custom CUDA/Triton operator libraries tailored to Waymo’s specific architectures.
  • Build tools to benchmark, profile GPU execution, and productize deep learning models for a streamlined and robust onboard and offboard deployment.
You Have
  • B.S. or M.S. in CS, EE, Deep Learning or a related field
  • 5+ years of industry experience on system performance, hardware-level GPU optimization, or ML compilers
  • Strong C++ and CUDA programming skills
  • Extensive experience in NVIDIA GPU Kernel development to accelerate deep learning models
  • Proven debugging and optimization experience on the XLA:GPU compiler, as well as the NVIDIA runtime stack
  • Passion for developing and optimizing ML software stacks for modern ML accelerator architectures (framework, runtime library, ML compiler, efficient deep learning etc.)
We Prefer
  • M.Sc or PhD in Computer Science, Mathematics or a related field.
  • Strong Python programming skills.
  • Experience with advanced NVIDIA profiling (e.g., Nsight Compute) and debugging (e.g. cuda-gdb) tools.
  • Solid experience with designing, training and debugging deep learning models to achieve the highest scores/accuracies.
  • In-depth knowledge of ML frameworks, ML compilers, and IRs (Triton, HLO, MLIR, CuTe DSL, cuTile) or modern ML system architectures.
  • Role-Related Knowledge
Skill Areas
  • C++ Coding
  • CUDA profiling & debugging
  • ML runtime optimization
  • Custom GPU kernel development

Waymo employees are also eligible to participate in Waymo’s discretionary annual bonus program, equity incentive plan, and generous Company benefits program, subject to eligibility requirements.

The expected base salary range for this full-time position across US locations is listed below. Actual starting pay will be based on job-related factors, including exact work location, experience, relevant training and education, and skill level.

Salary Range
  • $213,000—$263,000 USD
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Engineer, Runtime & Optimization
Machine Learning Engineer, Runtime & Optimization

Waymo • Mountain View (CA)

On-site
USD 204,000 - 259,000
Discretionary annual bonus program
Equity incentive plan
Generous company benefits program
Senior Machine Learning Engineer, Runtime and Serving
Senior Machine Learning Engineer, Runtime and Serving

Waymo • Mountain View (CA)

On-site
USD 213,000 - 263,000
Annual bonus program
Equity incentive plan
Generous company benefits program
Staff ML Engineer, Generative Model Performance & Efficiency
Staff ML Engineer, Generative Model Performance & Efficiency

Waymo • United States

On-site
USD 251,000 - 310,000
Staff ML Engineer, Generative Model Performance & Efficiency
Staff ML Engineer, Generative Model Performance & Efficiency

Waymo • Mountain View (NY)

On-site
USD 251,000 - 310,000
Discretionary annual bonus program
Equity incentive plan
Generous company benefits
Staff ML Engineer, Generative Model Performance & Efficiency
Staff ML Engineer, Generative Model Performance & Efficiency

Waymo • Mountain View (CA)

On-site
USD 251,000 - 310,000
Annual bonus program
Equity incentive plan
Generous Company benefits program
Software Engineer, GPU
Software Engineer, GPU

Waymo • Mountain View (CA)

On-site
USD 204,000 - 259,000
Discretionary annual bonus program
Equity incentive plan
Generous benefits program
Technical Lead Manager, Machine Learning Runtime & Serving Waymo Mountain View, California
Technical Lead Manager, Machine Learning Runtime & Serving Waymo Mountain View, California

Neura Market • Mountain View (CA)

Hybrid
USD 251,000 - 310,000
Bonus program
Equity incentive plan
Benefits program
Software Engineer, GPU
Software Engineer, GPU

Waymo • New York (NY)

Hybrid
USD 204,000 - 259,000
Discretionary annual bonus
Equity incentive plan
Benefits for employees
Machine Learning Engineer, Simulation Realism
Machine Learning Engineer, Simulation Realism

Waymo • Mountain View (CA)

Hybrid
USD 213,000 - 263,000
Annual bonus program
Equity incentive plan
Generous company benefits
Machine Learning Engineer, Driver Understanding and Evaluation
Machine Learning Engineer, Driver Understanding and Evaluation

Waymo • Mountain View (WY)

Hybrid
USD 164,000 - 222,000
Annual bonus
Equity incentives
Benefits package