Senior GPU Kernel & ML Compiler Engineer

Waymo

Mountain View (CA)

On-site

USD 213,000 - 263,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Annual bonus
Equity incentive
Company benefits

Job summary

Waymo is an autonomous driving technology company hiring for a role within the ML Infrastructure team. The position focuses on optimizing ML workloads, CUDA kernel development, and NVIDIA stack integration across perception, behavior, and planning models.

Candidates should have 5+ years in system performance or ML compilers, with strong C++/CUDA skills and experience with NVIDIA GPUs. Waymo offers a competitive compensation package and equity incentives.

Qualifications

  • B.S. or M.S. in CS, EE, Deep Learning or a related field.
  • 5+ years of industry experience on system performance, hardware-level GPU optimization, or ML compilers.
  • Strong C++ and CUDA programming skills.
  • Extensive experience in NVIDIA GPU Kernel development to accelerate deep learning models.
  • Proven debugging and optimization experience on the XLA:GPU compiler, as well as the NVIDIA runtime stack.
  • Passion for developing and optimizing ML software stacks for modern ML accelerator architectures (framework, runtime library, ML compiler, efficient deep learning etc.)

Responsibilities

  • Collaborate with ML practitioners on models for perception, behavior prediction, and planning, to understand their models and accelerate them onboard through custom NVIDIA GPU kernel development.
  • Deep dive into the NVIDIA ML software and runtime stack, from custom CUDA ops to the XLA:GPU compiler and low-level libraries. Analyze numeric behaviors, debug complex compiler issues, and ensure inference results are stable and consistent. Develop tools/system software for optimal resource usage, hardware efficiency, and platform reliability in an ML serving system.
  • Analyze ML workload performance at the hardware level; apply manual and AI agent-assisted techniques and develop highly optimized, custom CUDA/Triton operator libraries tailored to Waymo’s specific architectures.
  • Build tools to benchmark, profile GPU execution, and productize deep learning models for a streamlined and robust onboard and offboard deployment.

Skills

C++
CUDA
GPU kernel
XLA:GPU
NVIDIA runtime
ML software stacks

Education

B.S./M.S. in CS/EE/Deep Learning
M.Sc or PhD in Computer Science/Mathematics

Tools

Nsight Compute
cuda-gdb
Triton
MLIR
CuTe DSL
cuTile

Job description

Waymo is an autonomous driving technology company hiring for a role within the ML Infrastructure team. The position focuses on optimizing ML workloads, CUDA kernel development, and NVIDIA stack integration across perception, behavior, and planning models.

Candidates should have 5+ years in system performance or ML compilers, with strong C++/CUDA skills and experience with NVIDIA GPUs. Waymo offers a competitive compensation package and equity incentives.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Engineer, GPU Kernel and Runtime
Machine Learning Engineer, GPU Kernel and Runtime

Waymo • Mountain View (CA)

On-site
USD 213,000 - 263,000
Annual bonus
Equity incentive
Company benefits
Senior ML Compiler Engineer - Hybrid, AI Acceleration
Senior ML Compiler Engineer - Hybrid, AI Acceleration

Waymo • New York (NY)

Hybrid
USD 213,000 - 263,000
Senior ML Compiler Engineer - High-Performance AI, Hybrid
Senior ML Compiler Engineer - High-Performance AI, Hybrid

Waymo • Mountain View (CA)

Hybrid
USD 213,000 - 263,000
Discretionary annual bonus
Equity incentive plan
Generous benefits program
Senior ML Compiler Engineer, Compute
Senior ML Compiler Engineer, Compute

Waymo • New York (NY)

Hybrid
USD 213,000 - 263,000
Senior ML Compiler Engineer, Compute
Senior ML Compiler Engineer, Compute

Waymo • Mountain View (CA)

Hybrid
USD 213,000 - 263,000
Discretionary annual bonus
Equity incentive plan
Generous benefits program
GPU Software Engineer - High-Perf Compute & Profiling
GPU Software Engineer - High-Perf Compute & Profiling

Waymo • New York (NY)

Hybrid
USD 204,000 - 259,000
Discretionary annual bonus
Equity incentive plan
Benefits for employees
Senior ML ASIC Design Engineer
Senior ML ASIC Design Engineer

Waymo • New York (NY)

Hybrid
USD 175,000 - 215,000
Discretionary annual bonus
Equity incentive plan
Company benefits package
Staff ML Engineer - Simulation & Efficient Inference
Staff ML Engineer - Simulation & Efficient Inference

Waymo • United States

On-site
USD 251,000 - 310,000
Software Engineer, GPU
Software Engineer, GPU

Waymo • Mountain View (CA)

On-site
USD 204,000 - 259,000
Discretionary annual bonus program
Equity incentive plan
Generous benefits program
Software Engineer, GPU
Software Engineer, GPU

Waymo • New York (NY)

Hybrid
USD 204,000 - 259,000
Discretionary annual bonus
Equity incentive plan
Benefits for employees