Founding ML Infra Engineer — Large-Scale Video Models

Wintermeyer Ventures

San Francisco (CA)

On-site

USD 200,000 - 375,000

Full time

9 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Wintermeyer Ventures is seeking a founding training infrastructure engineer for an early-stage robotics AI company in San Francisco. You will own the compute used by the research team, building GPU and data clusters largely from scratch and shaping the pace of model iteration as runs scale to hundreds of GPUs.

You will own distributed training end-to-end, from parallelism strategy to data path and fault-tolerance, delivering production-grade infra for fast experimentation and reliable research.

Qualifications

  • 3+ years building or operating distributed training infra for large-scale pretraining.
  • Expertise across training infra stack: distributed training, storage-to-GPU paths, fault tolerance, evaluation systems, observability.
  • Experience with multi-GPU or multi-node setups at 100+ GPUs.
  • Hands-on with large-scale video or world models (vision-language-action, image-to-video, or robot policies).
  • Production experience with PyTorch or JAX and strong Python fundamentals.
  • Startup environment experience or sole ownership of training infra is a plus.
  • Eligible to work in the United States without visa sponsorship.

Responsibilities

  • Own distributed training end-to-end: parallelism strategy, multi-node performance, scaling efficiency, and GPU capacity.
  • Build the full data path from storage to GPU, including high-throughput loaders and sampling infrastructure.
  • Design fault tolerance for long-running training jobs: checkpointing and automatic restart.
  • Build an evaluation harness that runs automatically against checkpoints.
  • Own observability and reproducibility: experiment tracking, environment pinning, profiling.
  • Translate research requirements into production-grade infra for fast experimentation.

Skills

Distributed training infra
Python fundamentals
GPU optimization
Observability
Experiment tracking

Tools

PyTorch
JAX

Job description

Wintermeyer Ventures is seeking a founding training infrastructure engineer for an early-stage robotics AI company in San Francisco. You will own the compute used by the research team, building GPU and data clusters largely from scratch and shaping the pace of model iteration as runs scale to hundreds of GPUs.

You will own distributed training end-to-end, from parallelism strategy to data path and fault-tolerance, delivering production-grade infra for fast experimentation and reliable research.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff: Training Infrastructure
Member of Technical Staff: Training Infrastructure

Wintermeyer Ventures • San Francisco (CA)

On-site
USD 200,000 - 375,000
ML Infra Engineer: Scale & Optimize Large-Scale Training
ML Infra Engineer: Scale & Optimize Large-Scale Training

Physical Intelligence • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior ML Platform Engineer — Scale Research ML Infra
Senior ML Platform Engineer — Scale Research ML Infra

techire ai • San Francisco (CA)

On-site
USD 270,000 - 330,000
Stock options
Founding ML Systems Engineer – High-Performance AI Infra
Founding ML Systems Engineer – High-Performance AI Infra

Strativ Group • Palo Alto (CA)

On-site
USD 450,000 - 550,000
Founding equity
Direct exposure to founders
Competitive compensation
Artificial Intelligence Engineer
Artificial Intelligence Engineer

Acceler8 Talent • California (MO)

Hybrid
USD 180,000 - 260,000
Senior ML Infra Engineer for Large-Scale Mid-Training & RL
Senior ML Infra Engineer for Large-Scale Mid-Training & RL

Peano AI • Palo Alto (CA)

On-site
USD 200,000 - 260,000
Software Engineer: ML Infra
Software Engineer: ML Infra

Generalist • Somerville (MA), San Mateo (CA)

On-site
USD 120,000 - 160,000
Founding ML Infra Engineer — Equity & Open-Source AI
Founding ML Infra Engineer — Equity & Open-Source AI

Peano AI • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Infrastructure Engineer for Large-Scale ML Training
Infrastructure Engineer for Large-Scale ML Training

Thinking Machines Lab Inc. • San Francisco (CA)

On-site
USD 300,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Distributed Training Infra Engineer for Large Models
Distributed Training Infra Engineer for Large Models

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 240,000