Turn this role into an interview — a resume and cover letter built around what this employer wants.
Get past ATS filters
Benefits offered by this job
Comprehensive medical, dental, and vision
401(k) program
Generous PTO
Paid parental leave
Gym access
Complimentary lunch
Job summary
A leading research lab in Sunnyvale is seeking a distributed ML infrastructure engineer to extend and scale training systems. The ideal candidate must have over 5 years of experience in ML systems with strong expertise in distributed training frameworks like DeepSpeed and FSDP. This role offers a competitive salary ranging from $150,000 to $450,000 annually along with comprehensive benefits and amenities.
Qualifications
5+ years of experience in ML systems, infra, or distributed training.
Strong software engineering fundamentals and proven multi-node experience.
Ability to implement algorithms across GPUs/nodes based on mathematical specs.
Responsibilities
Extend or modify training frameworks to support new use cases.
Translate mathematical optimizer specs into distributed implementations.
Build systems for experiment tracking and job monitoring.
Skills
Distributed training frameworks (e.g., DeepSpeed, FSDP, FairScale, Horovod)
Python
Slurm
Kubernetes
Ray
NCCL/GLOO debugging skills
Mixed-precision training
CUDA
Triton kernel
Job description
A leading research lab in Sunnyvale is seeking a distributed ML infrastructure engineer to extend and scale training systems. The ideal candidate must have over 5 years of experience in ML systems with strong expertise in distributed training frameworks like DeepSpeed and FSDP. This role offers a competitive salary ranging from $150,000 to $450,000 annually along with comprehensive benefits and amenities.