Senior Distributed Systems Engineer MoE at Scale

Ifm Us

Sunnyvale (CA)

On-site

USD 190,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Bonus
401K Plan
Generous paid time off

Job summary

The Institute of Foundation Models (IFM) is seeking a deeply technical engineer to co‑design and optimize the communication stack for large‑scale distributed training, including hybrid parallelism and MoE workloads.

This role focuses on performance engineering, distributed debugging, and cross‑layer optimization across thousands of GPUs, with a strong emphasis on fault‑tolerant execution and topology‑aware design.

Qualifications

  • Experience optimizing distributed training at large GPU scales.
  • Hands‑on with RDMA, InfiniBand, RoCE, and GPUDirect RDMA.
  • Familiarity with NCCL and UCX internals.
  • Strong systems programming ability in C/C++, Rust, or Go.

Responsibilities

  • Design and optimize communication for large‑scale distributed training (thousand+ GPUs).
  • Drive high‑performance hierarchical collectives for MoE workloads.
  • Co‑design runtime orchestration with topology awareness.
  • Debug and profile training performance bottlenecks in communication.

Skills

NCCL/UCX internals
GPUDirect RDMA
RDMA/InfiniBand/RoCE
C/C++/Rust/Go
PyTorch familiarity
distributed debugging
performance engineering

Education

Master’s degree or Bachelor’s + 1 year experience

Tools

NCCL
UCX
InfiniBand

Job description

The Institute of Foundation Models (IFM) is seeking a deeply technical engineer to co‑design and optimize the communication stack for large‑scale distributed training, including hybrid parallelism and MoE workloads.

This role focuses on performance engineering, distributed debugging, and cross‑layer optimization across thousands of GPUs, with a strong emphasis on fault‑tolerant execution and topology‑aware design.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Distributed Systems Engineer
Senior Distributed Systems Engineer

Ifm Us • Sunnyvale (CA)

On-site
USD 190,000 - 260,000
Bonus
401K Plan
Generous paid time off
Staff Engineer – Foundation Model Serving & APIs
Staff Engineer – Foundation Model Serving & APIs

RoShay Services • San Francisco (CA)

On-site
USD 180,000 - 240,000
ML Infra Engineer: Scale Distributed Training & Research
ML Infra Engineer: Scale Distributed Training & Research

Doist • San Francisco (CA)

On-site
USD 180,000 - 250,000
Distributed ML Training Engineer - Scale GPUs, Unlimited PTO
Distributed ML Training Engineer - Scale GPUs, Unlimited PTO

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Unlimited PTO
Parental leave
+1
Senior Distributed ML Training Engineer
Senior Distributed ML Training Engineer

Visa Hunt • San Francisco (CA), Northern (KY)

Hybrid
USD 220,000 - 310,000
Stock options
Health/dental/vision insurance
Meals provided in office
+1
Senior Foundation Model Systems Engineer
Senior Foundation Model Systems Engineer

Jobtailor • California (MO)

On-site
USD 180,000 - 260,000
Distributed Training Infra Engineer for Large Models
Distributed Training Infra Engineer for Large Models

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 240,000
Remote Inference Optimization Engineer
Remote Inference Optimization Engineer

Modular Mailing Systems, Inc. • Los Altos (CA)

Hybrid
USD 198,000 - 286,000
Premier insurance plans
5% 401k matching
Flexible paid time off
+2
Distributed Training Engineer for Multimodal Models
Distributed Training Engineer for Multimodal Models

Luma • Redwood City (CA)

On-site
USD 210,000 - 260,000
Training Infra Engineer for Scalable LLM Systems
Training Infra Engineer for Scalable LLM Systems

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000