ML Research Engineer, Training

Weave Robotics

San Francisco (CA)

On-site

USD 180,000 - 250,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Weave Robotics is seeking a senior ML systems engineer to own end-to-end training stack development, from raw fleet uploads to GPU-accelerated model training. You will optimize data pipelines, orchestrate experiments, and push production-grade infrastructure for terabyte-scale robot data in real homes and businesses.

The role emphasizes deep PyTorch/JAX expertise, multi-node distributed training, and performance tuning of GPUs, memory, and I/O.

Qualifications

  • Experience building ML training stacks for large-scale data
  • Ability to read loss curves and identify data bugs and instability
  • Hands-on experience with distributed training across multi-node clusters

Responsibilities

  • Build training stack end to end: distributed training, data loading, checkpointing, run orchestration, experiment tracking.
  • Large-scale data handling: ingests and process terabytes of multimodal robot data (video, proprioception, sensor streams).
  • Optimize research productivity by growing a scalable, reproducible codebase for experiments.
  • Sampling and curation decisions that affect model behavior.
  • Make runs fast: profile, fix bottlenecks, ensure reproducibility from data loading to GPU kernels.
  • Transition research prototypes into production infrastructure used by the team.

Skills

ML system expertise
Distributed training
Performance engineering
Loss curve interpretation

Tools

PyTorch
JAX
Nsight Systems
NCCL
CUDA
Kubernetes
SLURM
GCP
AWS

Job description

Join Us, and Ship Robots

Weave was founded to build the robots we’d want to have in our own home. We believe the next generation of robotics will transform everyday life by enabling people to do more and to reclaim time to spend on what’s important.

We also believe robots are in a sense like any other product: to matter, they have to ship. Our robots are already operating in real homes and businesses, giving us the opportunity to rapidly improve from real-world experience. With a growing team, strong customer demand, and capital for expansion, we’re entering an exciting stage of growth—and we’re looking for people with exceptional talent and standards to help bring home robotics to millions of households.

The Role

Most robot learning research is graded on evals that don't survive contact with the field. Ours is graded by robots doing useful work in real homes and businesses, every day. We're one of the first companies with a deployed fleet generating real-world robot data at terabyte scale. The pipeline and training stack you build are what turns that data into capability.

Model quality is set as much by training as by architecture: what data gets in, how it's sampled, whether the run is stable, or whether a silent bug ate the gradient three days ago. You'll own that layer from raw fleet uploads to the batch that hits the GPU. When the stack is right, ideas become models in training in days, and deployed in weeks.

Responsibilities
  • Build training stack end to end: distributed training, data loading, checkpointing, run orchestration, experiment tracking.
  • Large-scale data handling: Develop high-throughput data ingestion, transformation, and storage systems capable of processing terabytes of multimodal robot data, including video, proprioception, and sensor streams.
  • Optimize research productivity: Grow the codebase that makes experiments reproducible, scalable, and easy to launch, monitor, retry/recover, and debug.
  • Sampling and curation: Drive sampling and curation decisions that show up in model behavior.
  • Make runs fast and honest: profile and fix throughput bottlenecks, chase down loss spikes and silent data bugs, keep results reproducible enough to trust: from data loading to GPU kernels.
  • From research to production: Turn research prototypes into infrastructure the whole team trains on.
What You'll Bring
  • ML system expertise: Deep PyTorch or JAX experience, including multi-node distributed training (FSDP, DDP, or equivalent) on real workloads.
  • Performance engineering: Experience profiling and optimizing GPU utilization, data pipelines, I/O bottlenecks, memory usage, and distributed training performance, including CUDA-level profiling tools (e.g. Nsight Systems) and NCCL tuning.
  • Training run judgement: You can read a loss curve, tell instability from a data bug and know when to kill a run.
Nice to Have
  • Robot learning exposure: you’ve trained policies (VLAs, world models, RL) and can tell a data problem from a model problem.
  • Cluster and cloud infrastructure experience: Kubernetes, SLURM, GCP/AWS.
  • Large-scale post-training experience: SFT, reward modeling, RL fine-tuning.
  • On-robot inference optimization experience: TensorRT, quantization, distillation.
  • CUDA or Triton kernel work.
  • Experience with video-heavy datasets: transcoding, chunking, and the storage/compute tradeoffs of training on video at scale.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Research Engineer, Data San Francisco, CA · On-site
ML Research Engineer, Data San Francisco, CA · On-site

Weave Robotics, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 200,000
ML Research Scientist, Robot Learning
ML Research Scientist, Robot Learning

Weave Robotics • San Francisco (CA)

On-site
USD 140,000 - 210,000
ML Research Engineer, Data
ML Research Engineer, Data

Weave Robotics • San Francisco (CA)

On-site
USD 140,000 - 190,000
ML Research Scientist, Robot Learning
ML Research Scientist, Robot Learning

Weave Robotics, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 230,000
Founding ML Engineer
Founding ML Engineer

a16z-speedrun • San Francisco (CA)

On-site
USD 150,000 - 230,000
Research Member of Technical Staff- Training Systems
Research Member of Technical Staff- Training Systems

Rhoda AI • Mountain View (CA)

On-site
USD 140,000 - 180,000
Research Member of Technical Staff- Training Systems
Research Member of Technical Staff- Training Systems

Rhoda AI • Mountain View (CA)

On-site
USD 150,000 - 200,000
Research Member of Technical Staff - Training Platform
Research Member of Technical Staff - Training Platform

Rhoda AI • Mountain View (CA)

On-site
USD 100,000 - 140,000
Research Member of Technical Staff- Training Systems
Research Member of Technical Staff- Training Systems

Rhoda AI • Palo Alto (CA)

On-site
USD 210,000 - 320,000
Research Member of Technical Staff- Robot Learning Systems & Reliability
Research Member of Technical Staff- Robot Learning Systems & Reliability

Rhoda • Mountain View (CA)

On-site
USD 180,000 - 240,000