Senior AI Training Infrastructure Engineer

Kodiak

Mountain View (CA)

On-site

USD 190,000 - 260,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Equity
Annual bonus
Medical/Dental/Vision
Flexible PTO
Office perks: lunch/kitchen
Headspace/Calm services
401(k)
Commuter benefits

Job summary

Kodiak Robotics, Inc. in Mountain View, CA is seeking ML systems engineers to design and optimize high‑throughput data loading and distributed training infrastructure for autonomous vehicle AI workloads. You will enable multi‑node GPU clusters, mixed precision, and scalable dataset pipelines.

The role requires strong Python and PyTorch skills, plus systems-level experience in C++/CUDA/Triton. Competitive compensation and comprehensive benefits are offered.

Qualifications

  • Experience with distributed training frameworks and techniques (PyTorch DDP/FSDP, DeepSpeed, Megatron, NCCL).
  • Experience building high-performance data pipelines for large-scale training (streaming formats, sharding, storage/network-aware loading).
  • Deep understanding of GPU performance: mixed precision, memory hierarchy, kernel fusion, profiling tools (Nsight, PyTorch Profiler).
  • Strong Python skills and PyTorch internals; systems-level experience (C++/CUDA/Triton) a plus.
  • Passion for scalable AI infrastructure for fast, reliable training.

Responsibilities

  • Design high-throughput data loading and streaming systems for multimodal sensor data, including dataset formats, sharding, and prefetching pipelines that keep GPUs saturated.
  • Build and optimize distributed training infrastructure across multi-node GPU clusters, applying various parallelism strategies.
  • Maximize utilization of accelerators (BF16/FP8), memory optimization, fused kernels, and overlap of compute and communication.
  • Profile end-to-end training pipelines to identify bottlenecks across storage, network, CPU preprocessing, and GPU compute.
  • Develop scalable dataset construction pipelines converting raw driving logs into training-ready formats.
  • Partner with ML teams to scale architectures from prototype to full-cluster training runs.

Skills

Distributed training
PyTorch DDP/FSDP
DeepSpeed
Megatron
NCCL
Data pipelines
Python
C++/CUDA/Triton

Education

BS/MS/PhD in CS or related field

Tools

WebDataset
MosaicML Streaming/MDS
NVLink
InfiniBand

Job description

Kodiak Robotics, Inc. in Mountain View, CA is seeking ML systems engineers to design and optimize high‑throughput data loading and distributed training infrastructure for autonomous vehicle AI workloads. You will enable multi‑node GPU clusters, mixed precision, and scalable dataset pipelines.

The role requires strong Python and PyTorch skills, plus systems-level experience in C++/CUDA/Triton. Competitive compensation and comprehensive benefits are offered.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Training Infrastructure Engineer
Senior AI Training Infrastructure Engineer

Omaze • Mountain View (CA)

On-site
USD 190,000 - 260,000
Competitive compensation & equity
Excellent health plans
Generous PTO & holidays
+1
Senior ML Engineer - Deploy & Optimize Onboard Autonomy
Senior ML Engineer - Deploy & Optimize Onboard Autonomy

Kodiak • San Francisco (CA)

On-site
USD 200,000 - 265,000
Competitive compensation
Health + dental + vision
Flexible PTO
+2
Senior AI Infrastructure Engineer (Model Training)
Senior AI Infrastructure Engineer (Model Training)

Kodiak Robotics • Mountain View (CA)

On-site
USD 190,000 - 260,000
Competitive compensation & equity
Excellent health plans
Generous PTO & holidays
+1
Staff ML Engineer for Autonomous Driving Deployment
Staff ML Engineer for Autonomous Driving Deployment

Kodiak • Mountain View (CA)

On-site
USD 200,000 - 265,000
Competitive compensation
Medical/Dental/Vision
Flexible PTO
+4
Senior AI Infrastructure Engineer - Model Training
Senior AI Infrastructure Engineer - Model Training

Kodiak • Mountain View (CA)

On-site
USD 190,000 - 260,000
Equity
Annual bonus
Medical/Dental/Vision
+5
Lead Architect, Multimodal Transformers
Lead Architect, Multimodal Transformers

Omaze • Mountain View (CA)

On-site
USD 200,000 - 230,000
Competitive compensation
Equity bonus
Excellent health plan
+5
Senior Multimodal AI Engineer — Real-Time Transformer
Senior Multimodal AI Engineer — Real-Time Transformer

Kodiak • Mountain View (CA)

On-site
USD 200,000 - 260,000
Competitive compensation package including equity and bonuses
Medical, Dental, and Vision plans
Flexible PTO and generous parental leave
+2
Senior DL Infra Engineer: Scalable GPU AI Training
Senior DL Infra Engineer: Scalable GPU AI Training

NVIDIA • California (MO)

On-site
USD 224,000 - 431,000
Equity
Benefits
Senior ML Infrastructure Engineer - GPU Training & MLOps
Senior ML Infrastructure Engineer - GPU Training & MLOps

Atoms • San Francisco (CA)

On-site
USD 224,000 - 280,000
Medical, Dental, Vision, Disability, and Life Insurance
Flexible Spending Account / Health Savings Account Options
401(k)
+2
Lead ML Training Systems Engineer — Scalable Multimodal AI
Lead ML Training Systems Engineer — Scalable Multimodal AI

Rhoda AI • Mountain View (CA)

On-site
USD 140,000 - 180,000