Machine Learning Engineer — Pre-training (LLM)

Lever, Inc.

Sunnyvale (CA)

In loco

USD 150.000 - 450.000

Tempo pieno

4 giorni fa
Candidati tra i primi
Generatore di candidature

Ricevi una risposta da questo datore di lavoro — un curriculum e una lettera di presentazione personalizzati, che corrispondono esattamente a ciò che sta cercando.

Supera i filtri ATS

Vantaggi offerti da questo lavoro

Medical, dental, and vision benefits
Bonus
401K Plan
Generous paid time off
Parental Leave
Employee Assistance Program
Life insurance and disability

Descrizione del lavoro

The Institute of Foundation Models is seeking a distributed ML infrastructure engineer to extend and scale our training systems. You’ll work with world‑class researchers to extend distributed training frameworks (DeepSpeed, FSDP, FairScale, Horovod) and build robust multi‑node launch scripts.

You will own experiment tracking, metrics logging, and job monitoring for external visibility, and aim to improve reliability and performance of large‑scale pre‑training pipelines.

Competenze

  • 5+ years of experience in ML systems, infra, or distributed training.
  • Experience modifying distributed ML frameworks (e.g., DeepSpeed, FSDP, FairScale, Horovod).
  • Strong software engineering fundamentals (Python, systems design, testing).
  • Proven multi-node experience (Slurm, Kubernetes, Ray) and debugging skills (NCCL/GLOO).
  • Ability to implement algorithms across GPUs/nodes based on mathematical specs.
  • Experience working on an ML platform/ infrastructure, and/or distributed inference optimization team.
  • Experience with large-scale machine learning workloads and pre-training.

Mansioni

  • Distributed Framework Ownership – Extend or modify training frameworks (e.g., DeepSpeed, FSDP) to support new use cases and architectures.
  • Optimizer Implementation – Translate mathematical optimizer specs into distributed implementations.
  • Launch Config & Debugging – Create and debug multi-node launch scripts with flexible batch sizes, parallelism strategies, and hardware targets.
  • Metrics & Monitoring – Build systems for experiment tracking, job monitoring, and logging usable by collaborators and researchers.
  • Infra Engineering – Write production-quality code and tests for ML infra in PyTorch or JAX; ensure reliability and maintainability at scale.
  • Data Loading & Checkpointing – Build and maintain data-loading and checkpoint/restart workflows, restoring model, optimizer, RNG, and data progress after interruptions.

Conoscenze

ML systems
Distributed training
Python
Slurm
Kubernetes
Ray
NCCL/GLOO

Strumenti

DeepSpeed
FSDP
FairScale
Horovod

Descrizione del lavoro

About the Institute of Foundation Models


We are a dedicated research lab for building, understanding, using, and risk-managing foundation models. Our mandate is to advance research, nurture the next generation of AI builders, and drive transformative contributions to a knowledge-driven economy.


As part of our team, you’ll have the opportunity to work on the core of cutting-edge foundation model training, alongside world-class researchers, data scientists, and engineers, tackling the most fundamental and impactful challenges in AI development. You will participate in the development of groundbreaking AI solutions that have the potential to reshape entire industries. Strategic and innovative problem-solving skills will be instrumental in establishing MBZUAI as a global hub forhigh-performance computing in deep learning, driving impactful discoveries that inspire the next generation of AIpioneers.


The Role

We’re looking for a distributed ML infrastructure engineer to help extend and scale our training systems. You’ll work side-by-side with world-class researchers and engineers to:



  • Extend distributed training frameworks (e.g., DeepSpeed, FSDP, FairScale, Horovod)

  • Implement distributed optimizers from mathematical specs

  • Build robust config + launch systems across multi-node, multi-GPU clusters

  • Own experiment tracking, metrics logging, and job monitoring for external visibility

  • Improve training system reliability, maintainability, and performance


Much of the work will support large-scale pre-training, and pre-training experience is required. Strong infrastructure and systems experience is what we value most.


Key Responsibilities


  • Distributed Framework Ownership – Extend or modify training frameworks (e.g., DeepSpeed, FSDP) to support new use cases and architectures.

  • Optimizer Implementation – Translate mathematical optimizer specs into distributed implementations.

  • Launch Config & Debugging – Create and debug multi-node launch scripts with flexible batch sizes, parallelism strategies, and hardware targets. Select and validate data, tensor, pipeline, expert, and context parallelism strategies as appropriate for the model and cluster.

  • Metrics & Monitoring – Build systems for experiment tracking, job monitoring, and logging usable by collaborators and researchers.

  • Infra Engineering – Write production-quality code and tests for ML infra in PyTorch or JAX; ensure reliability and maintainability at scale.

  • Data Loading & Checkpointing – Build and maintain data-loading and checkpoint/restart workflows, restoring model, optimizer, RNG, and data progress after interruptions.


Qualifications

Must-Haves:


  • 5+ years of experience in ML systems, infra, or distributed training

  • Experience modifying distributed ML frameworks (e.g., DeepSpeed, FSDP, FairScale, Horovod)

  • Strong software engineering fundamentals (Python, systems design, testing)

  • Proven multi-node experience (e.g., Slurm, Kubernetes, Ray) and debugging skills (e.g., NCCL/GLOO)

  • Ability to implement algorithms across GPUs/nodes based on mathematical specs

  • Experience working on an ML platform/ infrastructure, and/or distributed inference optimization team

  • Experience with large-scale machine learning workloads (strong ML fundamentals)

  • Experience with large-scale pre-training


Nice-to-Haves:


  • Exposure to mixed-precision training (e.g., bf16, fp8) with accuracy validation

  • Familiarity with performance profiling, kernel fusion, or memory optimization

  • Open-source contributions or published research (MLSys, ICML, NeurIPS)

  • CUDA or Triton kernel experience

  • Experience building custom training pipelines at scale and modifying them for custom needs

  • Deep familiarity with training infrastructure and performance tuning


$150,000 - $450,000 a year


Salary Range


The posted salary range represents the company’s good faith estimate of the compensation for this position upon hire. The actual compensation offered may vary within this range depending on individual qualifications, including but not limited to relevant skills, experience, education, certifications, geographic location, and specific business needs.


The posted salary range represents the company’s good faith estimate of the compensation for this position upon hire. The actual compensation offered may vary within this range depending on individual qualifications, including but not limited to relevant skills, experience, education, certifications, geographic location, and specific business needs.


Benefits Include


  • Comprehensive medical, dental, and vision benefits

  • Bonus

  • 401K Plan

  • Generous paid time off, sick leave and holidays

  • Paid Parental Leave

  • Employee Assistance Program

  • Life insurance and disability

Ottieni la revisione del curriculum gratis e riservata.

o trascina qui il file.

Similar jobs

Offerte di lavoro simili che vale la pena confrontare

Machine Learning Engineer — Reinforcement Learning
Machine Learning Engineer — Reinforcement Learning

Lever, Inc. • Sunnyvale (CA)

In loco
USD 150.000 - 450.000
Medical benefits
Dental benefits
Vision benefits
+7
Machine Learning Engineer
Machine Learning Engineer

Institute of Foundation Models • Sunnyvale (CA)

In loco
USD 140.000 - 180.000
Medical benefits
Bonus
401K Plan
+3
ML Infrastructure Engineer
ML Infrastructure Engineer

Lattice, Inc. • San Francisco (CA)

In loco
USD 200.000 - 280.000
Competitive salary
Premium health, dental, and vision insurance
Unlimited PTO
+2
Software Engineer - ML Infrastructure
Software Engineer - ML Infrastructure

Epsilon Health • San Francisco (CA)

In loco
USD 150.000 - 250.000
Senior Engineering Manager, ML Platform
Senior Engineering Manager, ML Platform

Boston Dynamics, Inc. • Waltham (MA)

In loco
USD 198.000 - 300.000
Medical
Dental
Vision
+3
Machine Learning Engineer — GPU Kernel
Machine Learning Engineer — GPU Kernel

Lever, Inc. • Sunnyvale (CA)

In loco
USD 150.000 - 450.000
Comprehensive medical, dental, and视觉 V
Bonus
401K Plan
+4
Software Engineer - ML Infrastructure
Software Engineer - ML Infrastructure

Epsilon • San Francisco (CA), Northern (KY)

Ibrido
USD 180.000 - 280.000
Senior ML Systems Engineer, Frameworks & Tooling
Senior ML Systems Engineer, Frameworks & Tooling

Cohere • San Francisco (CA)

In loco
USD 150.000 - 180.000
Inclusive culture
Weekly lunch stipend
Health and dental benefits
+4
Member of Technical Staff - ML Infrastructure Engineer, Post-training
Member of Technical Staff - ML Infrastructure Engineer, Post-training

Preference Model • San Francisco (CA)

In loco
USD 200.000 - 350.000
Health, vision, dental benefits
401K match
Lunch provided onsite
+2
Member of Technical Staff
Member of Technical Staff

Harrison Clarke • San Francisco (CA)

In loco
USD 180.000 - 280.000