Distributed ML Research Scientist — Frontier-Scale Training

Ifm Us

Sunnyvale (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health benefits
401K Plan
Paid time off
Parental leave
EAP

Job summary

Institute of Foundation Models in Sunnyvale is seeking a senior ML engineer to lead frontier-scale pre-training and systems work. You will set up distributed training across multi-node GPU clusters and implement production-grade kernels in CUDA/Triton.

You will prototype optimizers and attention methods, drive mixed-precision paths, and mentor interns and junior engineers while collaborating with researchers and engineers on end-to-end AI systems.

Qualifications

  • 5+ years combined industry or hands-on research experience with large-scale deep-learning training.
  • Led at least one large-scale transformer pre-training run.
  • Proficient in PyTorch or JAX/Flax and DeepSpeed/FSDP/Megatron-LM.
  • Experience with distributed training at scale (100+ GPUs).
  • Strong software engineering on large ML codebases; clear design docs.

Responsibilities

  • Set up DeepSpeed / FSDP / Megatron-LM across multi-node GPU clusters.
  • Create robust launch scripts, checkpoints, and job monitoring (NCCL/GLOO/GPU).
  • Prototype optimizers or attention methods in PyTorch/JAX/NumPy and convert to CUDA/Triton kernels.
  • Lead mixed-precision training (bf16, fp8, 4-bit) and analyze accuracy vs speed.
  • Apply kernel fusion, memory optimization, and performance tuning for throughput.
  • Build logging, metrics, and experiment-tracking tools for rapid iteration.
  • Design ablation studies and statistical tests; mentor interns and junior engineers.

Skills

5+ years experience
Transformer pre-training
PyTorch
JAX/Flax
DeepSpeed / FSDP / Megatron-LM / MOSA

Tools

Slurm
K8s
Ray
NCCL
CUDA
Triton

Job description

Institute of Foundation Models in Sunnyvale is seeking a senior ML engineer to lead frontier-scale pre-training and systems work. You will set up distributed training across multi-node GPU clusters and implement production-grade kernels in CUDA/Triton.

You will prototype optimizers and attention methods, drive mixed-precision paths, and mentor interns and junior engineers while collaborating with researchers and engineers on end-to-end AI systems.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Distributed ML Training Engineer
Senior Distributed ML Training Engineer

Visa Hunt • San Francisco (CA), Northern (KY)

Hybrid
USD 220,000 - 310,000
Stock options
Health/dental/vision insurance
Meals provided in office
+1
Senior ML Training Systems Engineer - Distributed CUDA
Senior ML Training Systems Engineer - Distributed CUDA

Genesis AI • San Francisco (CA)

On-site
USD 180,000 - 260,000
Distributed ML Engineer for High-Performance AI Training
Distributed ML Engineer for High-Performance AI Training

Ifm Us • Sunnyvale (CA)

On-site
USD 140,000 - 210,000
Medical benefits
Dental benefits
Vision benefits
+7
Distributed ML Training Systems Engineer
Distributed ML Training Systems Engineer

Mind Robotics • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Senior ML Training Systems Engineer - Distributed GPU Infra
Senior ML Training Systems Engineer - Distributed GPU Infra

Baseten • San Francisco (CA)

On-site
USD 150,000 - 200,000
Competitive compensation, including equity
100% coverage of medical, dental, and vision insurance
Generous PTO policy
+2
Senior Foundation Model Systems Engineer
Senior Foundation Model Systems Engineer

Jobtailor • California (MO)

On-site
USD 180,000 - 260,000
AI Training Systems Architect (Distributed)
AI Training Systems Architect (Distributed)

Unconventional AI • Palo Alto (CA)

On-site
USD 210,000 - 320,000
Health benefits
401k matching
Unlimited PTO
+1
Remote Senior Training Infrastructure Engineer—Multi-GPU AI
Remote Senior Training Infrastructure Engineer—Multi-GPU AI

Luma AI • San Francisco (CA)

Hybrid
USD 187,000 - 395,000
Pre-Training ML Engineer—Foundation Models
Pre-Training ML Engineer—Foundation Models

Trading Interview • New York (NY)

On-site
USD 300,000 - 350,000
Medical insurance
Dental insurance
Vision insurance
+3
Staff ML Systems Engineer — Distributed Training at Scale
Staff ML Systems Engineer — Distributed Training at Scale

RadixArk • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Comprehensive benefits
Flexible work arrangements