About the Role
This is a founding training infrastructure role at an early-stage robotics AI company pretraining a large-scale foundation model on 15PB+ of video and expert demonstration data. You will own the compute the research team runs on, building GPU and data clusters largely from scratch, and your architectural decisions will directly set the pace of model iteration as runs scale to hundreds of GPUs.
What You'll Do
- Own distributed training end-to-end: parallelism strategy, multi-node performance, scaling efficiency, and GPU capacity from cloud providers.
- Build the full data path from storage to GPU, including high-throughput loaders, pre-encoding pipelines, and sampling infrastructure.
- Design fault tolerance for long-running training jobs: checkpointing, checkpoint durability, and automatic failure detection and restart.
- Build an evaluation harness that runs automatically against every checkpoint, including probes, calibration checks, and research-facing dashboards.
- Own observability and reproducibility: experiment tracking, alerting, pinned environments, and profiling across compute, networking, and storage.
- Translate research requirements into production-grade infrastructure that keeps experimentation fast and large runs routine.
What We're Looking For
- 3+ years of hands-on experience building or operating distributed training infrastructure for large-scale model pretraining.
- Demonstrated expertise across the full training infrastructure stack: distributed training, storage-to-GPU data paths, fault tolerance, evaluation systems, and observability.
- Experience building or operating distributed training systems across multi-GPU or multi-node setups at 100+ GPU scale.
- Hands-on experience with large-scale video or world models, such as vision-language-action models, image-to-video models, or robot action policies.
- Production experience with PyTorch or JAX, and strong Python fundamentals.
- Prior GPU infrastructure optimization work on pretraining runs.
- Experience in early-stage startup environments or sole ownership of training infrastructure systems is a strong plus.
- Publications at top-tier ML conferences (ICML, ICLR, NeurIPS) are a plus.
- Must be eligible to work in the United States without visa sponsorship.
Compensation & Benefits
Salary range: $200,000 to $375,000 USD annually. No visa sponsorship is available.
Location
On-site in San Francisco, California, United States.