Machine Learning Infrastructure Engineer

Institute of Foundation Models

Sunnyvale (CA)

On-site

USD 150,000 - 450,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Comprehensive medical, dental, and vision
401(k) program
Generous PTO
Paid parental leave
Gym access
Complimentary lunch

Job summary

A leading research lab in Sunnyvale is seeking a distributed ML infrastructure engineer to extend and scale training systems. The ideal candidate must have over 5 years of experience in ML systems with strong expertise in distributed training frameworks like DeepSpeed and FSDP. This role offers a competitive salary ranging from $150,000 to $450,000 annually along with comprehensive benefits and amenities.

Qualifications

  • 5+ years of experience in ML systems, infra, or distributed training.
  • Strong software engineering fundamentals and proven multi-node experience.
  • Ability to implement algorithms across GPUs/nodes based on mathematical specs.

Responsibilities

  • Extend or modify training frameworks to support new use cases.
  • Translate mathematical optimizer specs into distributed implementations.
  • Build systems for experiment tracking and job monitoring.

Skills

Distributed training frameworks (e.g., DeepSpeed, FSDP, FairScale, Horovod)
Python
Slurm
Kubernetes
Ray
NCCL/GLOO debugging skills
Mixed-precision training
CUDA
Triton kernel

Job description

About the Institute of Foundation Models

We are a dedicated research lab for building, understanding, using, and risk-managing foundation models. Our mandate is to advance research, nurture the next generation of AI builders, and drive transformative contributions to a knowledge-driven economy.

As part of our team, you’ll have the opportunity to work on the core of cutting‑edge foundation model training, alongside world‑class researchers, data scientists, and engineers, tackling the most fundamental and impactful challenges in AI development. You will participate in the development of groundbreaking AI solutions that have the potential to reshape entire industries. Strategic and innovative problem‑solving skills will be instrumental in establishing MBZUAI as a global hub for high‑performance computing in deep learning, driving impactful discoveries that inspire the next generation of AI pioneers.

The Role

We're looking for a distributed ML infrastructure engineer to help extend and scale our training systems. You’ll work side‑by‑side with world‑class researchers and engineers to:

  • Extend distributed training frameworks (e.g., DeepSpeed, FSDP, FairScale, Horovod)
  • Implement distributed optimizers from mathematical specs
  • Build robust config + launch systems across multi‑node, multi‑GPU clusters
  • Own experiment tracking, metrics logging, and job monitoring for external visibility
  • Improve training system reliability, maintainability, and performance
  • While much of the work will support large‑scale pre‑training, pre‑training experience is not required. Strong infrastructure and systems experience is what we value most.
Key Responsibilities
  • Distributed Framework Ownership – Extend or modify training frameworks (e.g., DeepSpeed, FSDP) to support new use cases and architectures.
  • Optimizer Implementation – Translate mathematical optimizer specs into distributed implementations.
  • Launch Config & Debugging – Create and debug multi‑node launch scripts with flexible batch sizes, parallelism strategies, and hardware targets.
  • Metrics & Monitoring – Build systems for experiment tracking, job monitoring, and logging usable by collaborators and researchers.
  • Infra Engineering – Write production‑quality code and tests for ML infra in PyTorch or JAX; ensure reliability and maintainability at scale.
Qualifications
Must-Haves:
  • 5+ years of experience in ML systems, infra, or distributed training
  • Experience modifying distributed ML frameworks (e.g., DeepSpeed, FSDP, FairScale, Horovod)
  • Strong software engineering fundamentals (Python, systems design, testing)
  • Proven multi‑node experience (e.g., Slurm, Kubernetes, Ray) and debugging skills (e.g., NCCL/GLOO)
  • Ability to implement algorithms across GPUs/nodes based on mathematical specs
  • Experience working on an ML platform/ infrastructure, and/or distributed inference optimization team
  • Experience with large‑scale machine learning workloads (strong ML fundamentals)
Nice-to-Haves:
  • Exposure to mixed‑precision training (e.g., bf16, fp8) with accuracy validation
  • Familiarity with performance profiling, kernel fusion, or memory optimization
  • Open‑source contributions or published research (MLSys, ICML, NeurIPS)
  • CUDA or Triton kernel experience
  • Experience with large‑scale pre‑training
  • Experience building custom training pipelines at scale and modifying them for custom needs
  • Deep familiarity with training infrastructure and performance tuning

$150,000 - $450,000 a year

Benefits
  • Comprehensive medical, dental, and vision
  • 401(k) program
  • Generous PTO, sick leave, and holidays
  • Paid parental leave and family‑friendly benefits
  • On‑site amenities and perks: Complimentary lunch, gym access, and a short walk to the Sunnyvale Caltrain station
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Scientist - Distributed Machine Learning
Research Scientist - Distributed Machine Learning

Institute of Foundation Models • Sunnyvale (CA)

On-site
USD 300,000 - 600,000
Comprehensive medical, dental, and vision benefits
401K Plan
Generous paid time off
+3
Distributed Machine Learning Engineer
Distributed Machine Learning Engineer

Institute of Foundation Models • Sunnyvale (CA)

On-site
USD 150,000 - 450,000
Comprehensive medical, dental, and vision benefits
Bonus
401K Plan
+4
Machine Learning Engineer
Machine Learning Engineer

Institute of Foundation Models • Sunnyvale (CA)

On-site
USD 150,000 - 450,000
Comprehensive medical, dental, and vision benefits
Generous paid time off
Paid Parental Leave
Machine Learning Infrastructure Engineer
Machine Learning Infrastructure Engineer

David Joseph & Company • San Francisco (CA)

On-site
USD 200,000 - 400,000
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training

Preference Model • Seattle (WA)

On-site
USD 180,000 - 300,000
Health insurance
Vision insurance
Dental insurance
+3
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training

Preference Model • San Francisco (CA)

On-site
USD 180,000 - 300,000
Cash and equity
Ownership & autonomy
Visa sponsorship
+5
Machine Learning Engineer – World Model
Machine Learning Engineer – World Model

Institute of Foundation Models • Sunnyvale (CA)

On-site
USD 150,000 - 450,000
Comprehensive medical, dental, and vision benefits
Bonus
401K Plan
+4
ML Infrastructure Engineer
ML Infrastructure Engineer

Lattice, Inc. • San Francisco (CA)

Hybrid
USD 200,000 - 280,000
Competitive salary
Premium health, dental, and vision insurance
Unlimited PTO
+2
Research Engineer, Infrastructure, Training Systems
Research Engineer, Infrastructure, Training Systems

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Unlimited PTO
Parental leave
+1
Software Engineer - ML Infrastructure
Software Engineer - ML Infrastructure

Epsilon Health • San Francisco (CA)

On-site
USD 150,000 - 250,000