Member of Technical Staff: Training Infrastructure

Wintermeyer Ventures

San Francisco (CA)

On-site

USD 200,000 - 375,000

Full time

4 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Wintermeyer Ventures is seeking a founding training infrastructure engineer for an early-stage robotics AI company in San Francisco. You will own the compute used by the research team, building GPU and data clusters largely from scratch and shaping the pace of model iteration as runs scale to hundreds of GPUs.

You will own distributed training end-to-end, from parallelism strategy to data path and fault-tolerance, delivering production-grade infra for fast experimentation and reliable research.

Qualifications

  • 3+ years building or operating distributed training infra for large-scale pretraining.
  • Expertise across training infra stack: distributed training, storage-to-GPU paths, fault tolerance, evaluation systems, observability.
  • Experience with multi-GPU or multi-node setups at 100+ GPUs.
  • Hands-on with large-scale video or world models (vision-language-action, image-to-video, or robot policies).
  • Production experience with PyTorch or JAX and strong Python fundamentals.
  • Startup environment experience or sole ownership of training infra is a plus.
  • Eligible to work in the United States without visa sponsorship.

Responsibilities

  • Own distributed training end-to-end: parallelism strategy, multi-node performance, scaling efficiency, and GPU capacity.
  • Build the full data path from storage to GPU, including high-throughput loaders and sampling infrastructure.
  • Design fault tolerance for long-running training jobs: checkpointing and automatic restart.
  • Build an evaluation harness that runs automatically against checkpoints.
  • Own observability and reproducibility: experiment tracking, environment pinning, profiling.
  • Translate research requirements into production-grade infra for fast experimentation.

Skills

Distributed training infra
Python fundamentals
GPU optimization
Observability
Experiment tracking

Tools

PyTorch
JAX

Job description

About the Role

This is a founding training infrastructure role at an early-stage robotics AI company pretraining a large-scale foundation model on 15PB+ of video and expert demonstration data. You will own the compute the research team runs on, building GPU and data clusters largely from scratch, and your architectural decisions will directly set the pace of model iteration as runs scale to hundreds of GPUs.

What You'll Do
  • Own distributed training end-to-end: parallelism strategy, multi-node performance, scaling efficiency, and GPU capacity from cloud providers.
  • Build the full data path from storage to GPU, including high-throughput loaders, pre-encoding pipelines, and sampling infrastructure.
  • Design fault tolerance for long-running training jobs: checkpointing, checkpoint durability, and automatic failure detection and restart.
  • Build an evaluation harness that runs automatically against every checkpoint, including probes, calibration checks, and research-facing dashboards.
  • Own observability and reproducibility: experiment tracking, alerting, pinned environments, and profiling across compute, networking, and storage.
  • Translate research requirements into production-grade infrastructure that keeps experimentation fast and large runs routine.
What We're Looking For
  • 3+ years of hands-on experience building or operating distributed training infrastructure for large-scale model pretraining.
  • Demonstrated expertise across the full training infrastructure stack: distributed training, storage-to-GPU data paths, fault tolerance, evaluation systems, and observability.
  • Experience building or operating distributed training systems across multi-GPU or multi-node setups at 100+ GPU scale.
  • Hands-on experience with large-scale video or world models, such as vision-language-action models, image-to-video models, or robot action policies.
  • Production experience with PyTorch or JAX, and strong Python fundamentals.
  • Prior GPU infrastructure optimization work on pretraining runs.
  • Experience in early-stage startup environments or sole ownership of training infrastructure systems is a strong plus.
  • Publications at top-tier ML conferences (ICML, ICLR, NeurIPS) are a plus.
  • Must be eligible to work in the United States without visa sponsorship.
Compensation & Benefits

Salary range: $200,000 to $375,000 USD annually. No visa sponsorship is available.

Location

On-site in San Francisco, California, United States.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Engineer, Infrastructure, Training Systems
Research Engineer, Infrastructure, Training Systems

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

On-site
USD 350,000 - 475,000
Health benefits
Unlimited PTO
Parental leave
+1
Member of Technical Staff, Post-Training & Applied Research
Member of Technical Staff, Post-Training & Applied Research

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 275,000 - 315,000
Relocation assistance
Infrastructure, Large-scale Training San Jose
Infrastructure, Large-scale Training San Jose

Hark, Inc. • San Jose (CA)

On-site
USD 180,000 - 450,000
Research Engineer - Distributed Training
Research Engineer - Distributed Training

Prime Intellect • San Francisco (CA), Northern (KY)

On-site
USD 150,000 - 350,000
Senior Research Engineer, Model Training / Pretraining
Senior Research Engineer, Model Training / Pretraining

Intelix.AI • Seattle (WA)

On-site
USD 147,000 - 220,000
Member of Technical Staff, AI Compute & Data Infrastructure Vinci
Member of Technical Staff, AI Compute & Data Infrastructure Vinci

CDFAM - Computational Design Symposium • Palo Alto (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Member of Technical Staff, AI Compute & Data Infrastructure
Member of Technical Staff, AI Compute & Data Infrastructure

Vinci • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff — Training Infrastructure
Member of Technical Staff — Training Infrastructure

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 240,000
Research Engineer - RL Infrastructure
Research Engineer - RL Infrastructure

Prime Intellect • San Francisco (CA), Northern (KY)

On-site
USD 150,000 - 350,000
Visa sponsorship
Relocation assistance
Remote work option
Research Member of Technical Staff- Training Systems
Research Member of Technical Staff- Training Systems

Rhoda AI • Palo Alto (CA)

On-site
USD 210,000 - 320,000