Senior AI Infrastructure Engineer - Model Training

Kodiak

Mountain View (CA)

On-site

USD 190,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation
Equity
Medical/Dental/Vision
Flexible PTO
Mountain View office
Lunch provided
EV charging
Wellbeing benefits
401(k)

Job summary

Kodiak Robotics, Inc. is seeking engineers to accelerate model training for autonomous driving data pipelines in Mountain View, CA. The role focuses on building high‑throughput data loading, scalable distributed training, and GPU‑level optimizations to saturate modern accelerators.

You will work with PyTorch, DeepSpeed, and related tools to scale multimodal sensor data across multi‑node GPU clusters, ensuring efficient, reliable training performance for large autonomous driving models.

Qualifications

  • BS, MS, or PhD in Computer Science or a related field, and at least 2–3 years of industry experience in ML systems or infrastructure.
  • Experience with distributed training frameworks and techniques (PyTorch DDP/FSDP, DeepSpeed, Megatron, NCCL) and a strong grasp of parallelism trade-offs.
  • Experience building high‑performance data pipelines for large‑scale training, including streaming dataset formats (WebDataset, MosaicML Streaming/MDS, or similar), sharding, and storage/network‑aware loading.

Responsibilities

  • Design high-throughput data loading and streaming systems for multimodal sensor data (camera, LiDAR, radar), including dataset formats, sharding strategies, and prefetching pipelines that keep GPUs saturated
  • Build and optimize distributed training infrastructure across multi‑node GPU clusters, applying data, tensor, pipeline, and fully sharded (FSDP/ZeRO) parallelism to models that don't fit on a single device
  • Maximize utilization of modern accelerators such as NVIDIA B200s through mixed‑precision training (BF16/FP8), fused kernels, memory optimization, and communication/computation overlap
  • Profile end‑to‑end training pipelines to find and eliminate bottlenecks across storage, network, CPU preprocessing, and GPU compute
  • Develop scalable dataset construction pipelines that convert petabytes of raw driving logs into training‑ready, streamable formats
  • Partner with ML teams to scale new architectures from prototype to full‑cluster training runs efficiently and reliably

Skills

Distributed training
Data pipelines
GPU optimization
Python/PyTorch
C++/CUDA/Triton

Education

BS/MS/PhD in Computer Science or related field

Tools

WebDataset
MosaicML Streaming
NCCL
NVLink
InfiniBand

Job description

Mountain View, CA

Kodiak Robotics, Inc. was founded in 2018 and has become a leader in autonomous ground transportation committed to a safer and more efficient future for all. The company has developed an artificial intelligence (AI) powered technology stack purpose-built for commercial trucking and the public sector. The company delivers freight daily for its customers across the southern United States using its autonomous technology. In 2024, Kodiak became the first known company to publicly announce delivering a driverless semi‑truck to a customer. Kodiak is also leveraging its commercial self‑driving software to develop, test and deploy autonomous capabilities for the U.S. Department of Defense.

Kodiak's AI is only as good as the speed at which we can train it. Every improvement to our models – from GigaFusionNet to large‑scale world models – depends on infrastructure that turns thousands of hours of multimodal driving data into training throughput. We are looking for engineers who make model training fast: streaming massive camera, LiDAR, and radar datasets without stalling a single GPU, sharding data and models efficiently across nodes, and extracting every FLOP from the latest hardware. If you measure your impact in tokens per second and GPU utilization, this role is for you.

In this role, you will:
  • Design high‑throughput data loading and streaming systems for multimodal sensor data (camera, LiDAR, radar), including dataset formats, sharding strategies, and prefetching pipelines that keep GPUs saturated
  • Build and optimize distributed training infrastructure across multi‑node GPU clusters, applying data, tensor, pipeline, and fully sharded (FSDP/ZeRO) parallelism to models that don't fit on a single device
  • Maximize utilization of modern accelerators such as NVIDIA B200s through mixed‑precision training (BF16/FP8), fused kernels, memory optimization, and communication/computation overlap
  • Profile end‑to‑end training pipelines to find and eliminate bottlenecks across storage, network, CPU preprocessing, and GPU compute
  • Develop scalable dataset construction pipelines that convert petabytes of raw driving logs into training‑ready, streamable formats
  • Partner with ML teams to scale new architectures from prototype to full‑cluster training runs efficiently and reliably
What you’ll bring:
  • BS, MS, or PhD in Computer Science or a related field, and at least 2–3 years of industry experience in ML systems or infrastructure
  • Hands‑on experience with distributed training frameworks and techniques (PyTorch DDP/FSDP, DeepSpeed, Megatron, NCCL) and a strong grasp of parallelism trade‑offs
  • Experience building high‑performance data pipelines for large‑scale training, including streaming dataset formats (WebDataset, MosaicML Streaming/MDS, or similar), sharding, and storage/network‑aware loading
  • Deep understanding of GPU performance: mixed precision, memory hierarchy, kernel fusion, profiling tools (Nsight, PyTorch Profiler), and interconnects (NVLink, InfiniBand)
  • Strong Python skills and proficiency in PyTorch internals; systems‑level experience (C++/CUDA/Triton) a plus
  • Passion for building the infrastructure that lets AI for the physical world train faster, scale further, and improve continuously
What we offer:
  • Competitive compensation package including equity and annual bonuses
  • Excellent Medical, Dental, and Vision plans through Kaiser Permanente, Cigna, and MetLife (including a medical plan with infertility benefits)
  • MetLife Legal Services, Identity & Fraud Protection, Hospital Indemnity Insurance, Accident Insurance, and Critical Illness Insurance
  • Flexible PTO, 10 paid holidays, and generous parental leave policies
  • Our office is centrally located in Mountain View, CA
  • Office perks: dog‑friendly, free catered lunch, a fully stocked kitchen, and free EV charging
  • Wellbeing Benefits – Headspace through Cigna, Calm through Kaiser, One Medical, Gympass, Spring Health through Cigna, Rula (mental health navigation)
  • Fidelity 401(k)
  • Commuter, FSA, Dependent Care FSA, HSA

The pay range listed below reflects the base salary in our SF/Silicon Valley location, across several internal levels. Actual starting pay will be based on job‑related factors including: work location, experience, relevant training, education, skill level and performance during interview. Total compensation at Kodiak includes base pay, equity, bonus and a competitive benefits package.

$190,000 – $260,000 USD

At Kodiak, we strive to build a diverse community working towards our common company goals in a safe and collaborative environment where harassment of any kind is strictly prohibited. Kodiak is committed to equal opportunity employment regardless of race, ethnicity, religion, gender identity, sexual orientation, age, disability, or veteran status, or any other basis protected by applicable law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infrastructure Engineer (Model Training)
Senior AI Infrastructure Engineer (Model Training)

Omaze • Mountain View (CA)

On-site
USD 190,000 - 260,000
Competitive compensation & equity
Excellent health plans
Generous PTO & holidays
+1
Staff Machine Learning Engineer - Deployment
Staff Machine Learning Engineer - Deployment

Kodiak • San Francisco (CA)

On-site
USD 200,000 - 265,000
Competitive compensation package
Medical, Dental, and Vision plans
Flexible PTO
+2
Staff Machine Learning Engineer - Data
Staff Machine Learning Engineer - Data

Kodiak • Mountain View (CA)

On-site
USD 200,000 - 265,000
Competitive compensation package
Medical, Dental and Vision plans
Flexible PTO and paid holidays
+2
Senior Applied AI Engineer - Multimodal Transformers
Senior Applied AI Engineer - Multimodal Transformers

Kodiak Robotics • Mountain View (CA)

On-site
USD 200,000 - 260,000
Competitive compensation including equity and annual bonuses
Medical, Dental, and Vision plans
Flexible PTO and paid holidays
+3
Senior Applied AI Engineer - Multimodal Transformers
Senior Applied AI Engineer - Multimodal Transformers

Kodiak • Mountain View (CA)

On-site
USD 200,000 - 260,000
Competitive compensation package including equity and bonuses
Medical, Dental, and Vision plans
Flexible PTO and generous parental leave
+2
Senior Software Engineer, Multimodal Transformers
Senior Software Engineer, Multimodal Transformers

Omaze • Mountain View (CA)

On-site
USD 200,000 - 230,000
Equity
Annual bonuses
Medical coverage
+10
Technical Lead, Multimodal Transformers
Technical Lead, Multimodal Transformers

Socket.dev • Mountain View (CA)

On-site
USD 230,000 - 300,000
Competitive compensation
Equity
Medical benefits
+1
Staff Machine Learning Engineer - Data
Staff Machine Learning Engineer - Data

Kodiak • Mountain View (WY)

On-site
USD 200,000 - 265,000
Equity and annual bonuses
Medical, Dental, Vision plans
Flexible PTO
+5
World Model Research Scientist- Physical AI
World Model Research Scientist- Physical AI

Kodiak • Mountain View (CA)

On-site
USD 180,000 - 240,000
Competitive compensation package
Medical, Dental, and Vision plans
Flexible PTO
+3
Senior Onboard Infrastructure Software Engineer
Senior Onboard Infrastructure Software Engineer

Kodiak • Mountain View (WY)

On-site
USD 150,000 - 250,000
Competitive compensation
Health insurance plans
Flexible PTO
+2