Senior ML Training Infrastructure Engineer

Dyna Robotics, Inc

Redwood City, Northern (CA, KY)

Hybrid

USD 180,000 - 260,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Dyna Robotics in Redwood City, CA seeks an ML Training Infrastructure Engineer to architect and build the end-to-end training infrastructure for our multi-cloud GPU fleet. You will ensure every GPU is busy, every run is reproducible, and researchers can run experiments with a single command.

You will own data pipelines, scheduling, and inference pipelines using TensorRT, Triton, and DeepSpeed. The role demands deep HPC/ML infra expertise and a strong ownership mindset.

Qualifications

  • 7+ years of engineering experience in HPC or ML infrastructure.
  • Proficiency with PyTorch and distributed training (DeepSpeed, Accelerate).
  • Cloud GPU management (GCP/AWS) and Kubernetes expertise.
  • Low-level distributed systems knowledge including NCCL, memory management, race conditions.
  • Ownership mindset to design, build, and operate end-to-end systems.

Responsibilities

  • Scale distributed training for large-scale GPU clusters with optimization techniques.
  • Build a research codebase and job scheduler using Kubernetes/SLURM.
  • Design high-throughput pipelines for terabytes of multimodal data.
  • Create low-latency production inference pipelines with TensorRT and Triton.
  • Profile GPU utilization, I/O bottlenecks, and memory to maximize compute efficiency.

Skills

7+ years experience
PyTorch
DeepSpeed
Accelerate
Kubernetes
GCP/AWS
NCCL
Distributed systems
Ownership mindset
SLURM
TensorRT
Triton

Job description

Dyna Robotics in Redwood City, CA seeks an ML Training Infrastructure Engineer to architect and build the end-to-end training infrastructure for our multi-cloud GPU fleet. You will ensure every GPU is busy, every run is reproducible, and researchers can run experiments with a single command.

You will own data pipelines, scheduling, and inference pipelines using TensorRT, Triton, and DeepSpeed. The role demands deep HPC/ML infra expertise and a strong ownership mindset.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Infrastructure Engineer, Training Redwood City, CA Fulltime
ML Infrastructure Engineer, Training Redwood City, CA Fulltime

Dyna Robotics, Inc • Redwood City (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Senior ML Infrastructure Engineer — GPU & Robotics
Senior ML Infrastructure Engineer — GPU & Robotics

Nimble • San Francisco (CA)

On-site
USD 210,000 - 300,000
Unlimited Flexible Time Off
Health Insurance
Paid Parental Leave
+4
Senior ML Training Infrastructure Architect
Senior ML Training Infrastructure Architect

General Motors • Sunnyvale (CA), Northern (KY)

Hybrid
USD 170,000 - 241,000
Medical
Dental
Vision
+1
Senior ML Systems Engineer - Training Infra
Senior ML Systems Engineer - Training Infra

OpenAI • California (MO)

On-site
USD 295,000 - 380,000
Relocation assistance
Senior ML Infrastructure Engineer – Training & Inference
Senior ML Infrastructure Engineer – Training & Inference

Nimble • San Francisco (CA), Northern (KY)

Hybrid
USD 210,000 - 300,000
Unlimited Flexible Time Off
Health Insurance
Equity
+3
Senior ML Infrastructure Engineer - GPU Training & MLOps
Senior ML Infrastructure Engineer - GPU Training & MLOps

Atoms • San Francisco (CA)

On-site
USD 224,000 - 280,000
Medical, Dental, Vision, Disability, and Life Insurance
Flexible Spending Account / Health Savings Account Options
401(k)
+2
Senior ML Training Systems Engineer - Distributed CUDA
Senior ML Training Systems Engineer - Distributed CUDA

Genesis AI • San Francisco (CA)

On-site
USD 180,000 - 260,000
Real-Time ML Infrastructure Engineer for Robotics
Real-Time ML Infrastructure Engineer for Robotics

General Robotics • Redmond (WA)

On-site
USD 155,000 - 200,000
Medical benefits
401K
Senior Training Infra Engineer - 800+ GPU Scale
Senior Training Infra Engineer - 800+ GPU Scale

Figureai • San Jose (CA)

On-site
USD 150,000 - 350,000
Software Engineer: ML Infra
Software Engineer: ML Infra

Generalist • Somerville (MA), San Mateo (CA)

On-site
USD 120,000 - 160,000