Distributed ML Infra Engineer - Large-Scale Pre-Training

Lever, Inc.

Sunnyvale (CA)

On-site

USD 150,000 - 450,000

Full time

5 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Medical, dental, and vision benefits
Bonus
401K Plan
Generous paid time off
Parental Leave
Employee Assistance Program
Life insurance and disability

Job summary

The Institute of Foundation Models is seeking a distributed ML infrastructure engineer to extend and scale our training systems. You’ll work with world‑class researchers to extend distributed training frameworks (DeepSpeed, FSDP, FairScale, Horovod) and build robust multi‑node launch scripts.

You will own experiment tracking, metrics logging, and job monitoring for external visibility, and aim to improve reliability and performance of large‑scale pre‑training pipelines.

Qualifications

  • 5+ years of experience in ML systems, infra, or distributed training.
  • Experience modifying distributed ML frameworks (e.g., DeepSpeed, FSDP, FairScale, Horovod).
  • Strong software engineering fundamentals (Python, systems design, testing).
  • Proven multi-node experience (Slurm, Kubernetes, Ray) and debugging skills (NCCL/GLOO).
  • Ability to implement algorithms across GPUs/nodes based on mathematical specs.
  • Experience working on an ML platform/ infrastructure, and/or distributed inference optimization team.
  • Experience with large-scale machine learning workloads and pre-training.

Responsibilities

  • Distributed Framework Ownership – Extend or modify training frameworks (e.g., DeepSpeed, FSDP) to support new use cases and architectures.
  • Optimizer Implementation – Translate mathematical optimizer specs into distributed implementations.
  • Launch Config & Debugging – Create and debug multi-node launch scripts with flexible batch sizes, parallelism strategies, and hardware targets.
  • Metrics & Monitoring – Build systems for experiment tracking, job monitoring, and logging usable by collaborators and researchers.
  • Infra Engineering – Write production-quality code and tests for ML infra in PyTorch or JAX; ensure reliability and maintainability at scale.
  • Data Loading & Checkpointing – Build and maintain data-loading and checkpoint/restart workflows, restoring model, optimizer, RNG, and data progress after interruptions.

Skills

ML systems
Distributed training
Python
Slurm
Kubernetes
Ray
NCCL/GLOO

Tools

DeepSpeed
FSDP
FairScale
Horovod

Job description

The Institute of Foundation Models is seeking a distributed ML infrastructure engineer to extend and scale our training systems. You’ll work with world‑class researchers to extend distributed training frameworks (DeepSpeed, FSDP, FairScale, Horovod) and build robust multi‑node launch scripts.

You will own experiment tracking, metrics logging, and job monitoring for external visibility, and aim to improve reliability and performance of large‑scale pre‑training pipelines.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior ML Infra Engineer - Large-Scale Training & Pipelines
Senior ML Infra Engineer - Large-Scale Training & Pipelines

Kindredventures • San Francisco (CA)

On-site
USD 160,000 - 220,000
Senior ML Systems Engineer – Distributed Training
Senior ML Systems Engineer – Distributed Training

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Equity
Health benefits
Remote-friendly US culture
+1
RL Infrastructure Engineer: Scale End-to-End ML
RL Infrastructure Engineer: Scale End-to-End ML

Lever, Inc. • Sunnyvale (CA)

On-site
USD 150,000 - 450,000
Medical benefits
Dental benefits
Vision benefits
+7
ML Infra Engineer: Scale & Optimize Large-Scale Training
ML Infra Engineer: Scale & Optimize Large-Scale Training

Physical Intelligence • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff - ML Infra
Member of Technical Staff - ML Infra

Kindredventures • San Francisco (CA)

On-site
USD 160,000 - 220,000
Distributed Training Infra Engineer for Large Models
Distributed Training Infra Engineer for Large Models

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 240,000
Machine Learning Engineer — Pre-training (LLM)
Machine Learning Engineer — Pre-training (LLM)

Lever, Inc. • Sunnyvale (CA)

On-site
USD 150,000 - 450,000
Medical, dental, and vision benefits
Bonus
401K Plan
+4
Senior ML Infra Engineer for Large-Scale Mid-Training & RL
Senior ML Infra Engineer for Large-Scale Mid-Training & RL

Peano AI • Palo Alto (CA)

On-site
USD 200,000 - 260,000
Distributed ML Systems Engineer: Scalable Training
Distributed ML Systems Engineer: Scalable Training

Susquehanna International Group, LLP • Lower Merion Township

On-site
USD 120,000 - 170,000
ML Infra Engineer: Scale & Optimize Large-Scale Training
ML Infra Engineer: Scale & Optimize Large-Scale Training

Goliath Partners Inc. • San Francisco (CA)

On-site
USD 150,000 - 210,000