ML Systems Engineer: GPU Training & Infra

Dyna Robotics

Redwood City (CA)

On-site

USD 150,000 - 210,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Dyna Robotics is seeking a Research Engineer to architect and build end-to-end training infrastructure for a multi-cloud GPU fleet. You will ensure GPUs are utilized efficiently, runs are reproducible, and researchers can iterate with a single command.

You will own training infrastructure end-to-end, working on scalable systems for large multimodal models and real-time robot data processing, with emphasis on performance, reliability, and deployable inference.

Qualifications

  • 7+ years of engineering experience in HPC or ML infrastructure.
  • Deep experience with PyTorch and distributed training frameworks (DeepSpeed, Accelerate).
  • Hands-on cloud GPU management (GCP/AWS) and container orchestration (Kubernetes).
  • Strong understanding of distributed systems, NCCL inter-node communication, memory management.

Responsibilities

  • Scale distributed training infrastructure for large-scale GPU clusters, including sharding, checkpointing, and memory optimization.
  • Build a research codebase and job scheduling system to enable fast iteration and reliable retries.
  • Create high-throughput data pipelines for terabytes of multimodal robot data, ensuring GPUs are never starved.
  • Develop low-latency production inference pipelines with quantization and model compilation (TensorRT, Triton).
  • Profile GPU utilization and memory to optimize performance across the compute fleet.

Job description

Dyna Robotics is seeking a Research Engineer to architect and build end-to-end training infrastructure for a multi-cloud GPU fleet. You will ensure GPUs are utilized efficiently, runs are reproducible, and researchers can iterate with a single command.

You will own training infrastructure end-to-end, working on scalable systems for large multimodal models and real-time robot data processing, with emphasis on performance, reliability, and deployable inference.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Systems Engineer - Distributed AI & GPU
ML Systems Engineer - Distributed AI & GPU

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits package
ML Infra Engineer: Scale GPU Training & Data Pipelines
ML Infra Engineer: Scale GPU Training & Data Pipelines

Humble Robotics • United States

On-site
USD 150,000 - 230,000
ML Infra Engineer: GPU Fleet & Inference Orchestrator
ML Infra Engineer: GPU Fleet & Inference Orchestrator

Generalist • San Francisco (CA)

On-site
USD 120,000 - 160,000
Distributed ML Training Engineer - Scale GPUs, Unlimited PTO
Distributed ML Training Engineer - Scale GPUs, Unlimited PTO

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Unlimited PTO
Parental leave
+1
ML Infra Engineer — GPU Clusters & Distributed Systems
ML Infra Engineer — GPU Clusters & Distributed Systems

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 250,000
Industry-leading compensation and/or:?
Unlimited PTO
Top-tier medical, dental, and vision
+1
Lead ML Training Systems Engineer - Multimodal, Large-Scale
Lead ML Training Systems Engineer - Multimodal, Large-Scale

Rhoda AI • Palo Alto (CA)

On-site
USD 210,000 - 320,000
ML Platform Engineer — Infra for Research on GPU Fleets
ML Platform Engineer — Infra for Research on GPU Fleets

cursor • New York (NY), San Francisco (CA)

On-site
USD 120,000 - 180,000
ML Infra & GPU Fleet Engineer
ML Infra & GPU Fleet Engineer

Generalist • Somerville (MA), San Mateo (CA)

On-site
USD 120,000 - 160,000
Remote GPU Performance Engineer for Large-Scale Models
Remote GPU Performance Engineer for Large-Scale Models

United States Digital Space LLC • United States

Remote
USD 140,000 - 210,000
Five weeks paid leave
Comprehensive healthcare (vision +</p>
ML Platform Engineer – Training Orchestration
ML Platform Engineer – Training Orchestration

Rhoda AI • Mountain View (CA)

On-site
USD 100,000 - 140,000