Robotics ML Training Engineer: End-to-End Scale

Weave Robotics

San Francisco (CA)

On-site

USD 180,000 - 250,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Weave Robotics is seeking a senior ML systems engineer to own end-to-end training stack development, from raw fleet uploads to GPU-accelerated model training. You will optimize data pipelines, orchestrate experiments, and push production-grade infrastructure for terabyte-scale robot data in real homes and businesses.

The role emphasizes deep PyTorch/JAX expertise, multi-node distributed training, and performance tuning of GPUs, memory, and I/O.

Qualifications

  • Experience building ML training stacks for large-scale data
  • Ability to read loss curves and identify data bugs and instability
  • Hands-on experience with distributed training across multi-node clusters

Responsibilities

  • Build training stack end to end: distributed training, data loading, checkpointing, run orchestration, experiment tracking.
  • Large-scale data handling: ingests and process terabytes of multimodal robot data (video, proprioception, sensor streams).
  • Optimize research productivity by growing a scalable, reproducible codebase for experiments.
  • Sampling and curation decisions that affect model behavior.
  • Make runs fast: profile, fix bottlenecks, ensure reproducibility from data loading to GPU kernels.
  • Transition research prototypes into production infrastructure used by the team.

Skills

ML system expertise
Distributed training
Performance engineering
Loss curve interpretation

Tools

PyTorch
JAX
Nsight Systems
NCCL
CUDA
Kubernetes
SLURM
GCP
AWS

Job description

Weave Robotics is seeking a senior ML systems engineer to own end-to-end training stack development, from raw fleet uploads to GPU-accelerated model training. You will optimize data pipelines, orchestrate experiments, and push production-grade infrastructure for terabyte-scale robot data in real homes and businesses.

The role emphasizes deep PyTorch/JAX expertise, multi-node distributed training, and performance tuning of GPUs, memory, and I/O.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Robotics ML Data Engineer Scale Terabyte Pipelines
Robotics ML Data Engineer Scale Terabyte Pipelines

Weave Robotics, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 200,000
Robotics ML Engineer: Scale Real-World Fleet Learning
Robotics ML Engineer: Scale Real-World Fleet Learning

Watney Robotics • San Francisco (CA)

On-site
USD 150,000 - 190,000
ML Research Engineer, Training
ML Research Engineer, Training

Weave Robotics • San Francisco (CA)

On-site
USD 180,000 - 250,000
Robotics ML Data Engineer for Real-World Robots
Robotics ML Data Engineer for Real-World Robots

Weave Robotics • San Francisco (CA)

On-site
USD 140,000 - 190,000
Staff ML Infra Engineer: Scalable Training & Pipelines
Staff ML Infra Engineer: Scalable Training & Pipelines

Watney Robotics Inc • San Francisco (CA)

On-site
USD 140,000 - 210,000
ML Research Engineer, Data San Francisco, CA · On-site
ML Research Engineer, Data San Francisco, CA · On-site

Weave Robotics, Inc. • San Francisco (CA), Northern (KY)

On-site
USD 140,000 - 200,000
ML Research Engineer, Data
ML Research Engineer, Data

Weave Robotics • San Francisco (CA)

On-site
USD 140,000 - 190,000
Robotics ML Research Scientist - Real-World Learning
Robotics ML Research Scientist - Real-World Learning

Weave Robotics, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 230,000
Lead ML Training Systems Engineer - Multimodal, Large-Scale
Lead ML Training Systems Engineer - Multimodal, Large-Scale

Rhoda AI • Palo Alto (CA)

On-site
USD 210,000 - 320,000
Robotics ML Training Infrastructure Engineer
Robotics ML Training Infrastructure Engineer

Eka Robotics • Boston (MA)

On-site
USD 140,000 - 190,000