Senior Distributed Systems Engineer — High-Perf GPU

Lever, Inc.

Sunnyvale (CA)

On-site

USD 180,000 - 250,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

The Institute of Foundation Models designs ultra-scale GPU systems to train foundation models and optimize cross-layer performance and scalability. We seek a deeply technical engineer to co-design and optimize the communication stack for large-scale distributed training, including MoE workloads, hybrid parallelism, and fault-tolerant execution across thousands of GPUs.

This role emphasizes performance engineering, distributed debugging, and production-grade optimizations at the boundary of

Qualifications

  • Experience optimizing distributed training at 1,000+ GPU scale.
  • Hands-on with RDMA, InfiniBand, RoCE, and GPUDirect RDMA.
  • Deep familiarity with NCCL and/or UCX internals.
  • Systems programming in C/C++, Rust, or Go.
  • Experience with PyTorch in large-scale training.
  • Ability to profile and fix communication bottlenecks.

Responsibilities

  • Design and optimize expert-parallel and hybrid-parallel communication patterns.
  • Drive high-performance hierarchical collectives for MoE workloads.
  • Co-design runtime orchestration with communication topology awareness.
  • Reduce tail latency and improve determinism across thousands of GPUs.
  • Architect fault-tolerant distributed execution under real-world cluster failures.

Skills

RDMA/InfiniBand optimization
NCCL/UCX internals
C/C++ / Rust / Go systems programming
PyTorch large-scale training
Distributed debugging
Troubleshooting communication bottlene

Tools

NCCL
UCX
GPUDirect RDMA
InfiniBand
RoCE

Job description

The Institute of Foundation Models designs ultra-scale GPU systems to train foundation models and optimize cross-layer performance and scalability. We seek a deeply technical engineer to co-design and optimize the communication stack for large-scale distributed training, including MoE workloads, hybrid parallelism, and fault-tolerant execution across thousands of GPUs.

This role emphasizes performance engineering, distributed debugging, and production-grade optimizations at the boundary of

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Distributed Systems Engineer
Senior Distributed Systems Engineer

Lever, Inc. • Sunnyvale (CA)

On-site
USD 180,000 - 250,000
Staff Engineer, Distributed GPU Clusters
Staff Engineer, Distributed GPU Clusters

Kindredventures • San Francisco (CA)

On-site
USD 140,000 - 230,000
Remote GPU Performance Engineer for Large-Scale Models
Remote GPU Performance Engineer for Large-Scale Models

United States Digital Space LLC • United States

Remote
USD 140,000 - 210,000
Five weeks paid leave
Comprehensive healthcare (vision +</p>
Distributed Training Engineer — High-Perf GPU Scale
Distributed Training Engineer — High-Perf GPU Scale

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Relocation assistance
Visa sponsorship
Senior ML Training Systems Engineer - Distributed GPU Infra
Senior ML Training Systems Engineer - Distributed GPU Infra

Baseten • San Francisco (CA)

On-site
USD 150,000 - 200,000
Competitive compensation, including equity
100% coverage of medical, dental, and vision insurance
Generous PTO policy
+2
Senior AI Training Performance Engineer (GPU & Scale)
Senior AI Training Performance Engineer (GPU & Scale)

figure.ai • San Jose (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Staff Software Engineer, GPU Infra & ML Systems
Staff Software Engineer, GPU Infra & ML Systems

Reflection AI Ltd • New York (NY)

On-site
USD 180,000 - 240,000
Top-tier compensation
Stock options
Health & wellness
+4
Staff Engineer: Foundation Model API & GPU Inference
Staff Engineer: Foundation Model API & GPU Inference

Databricks Inc. • San Francisco (CA)

On-site
USD 192,000 - 260,000
Comprehensive benefits
Diversity and inclusion initiatives
Senior GPU Cluster Engineer for AI Infrastructure
Senior GPU Cluster Engineer for AI Infrastructure

Sciforium • San Francisco (CA)

On-site
USD 150,000 - 220,000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2
GPU Systems Engineer
GPU Systems Engineer

Career Techniques • New York (NY)

On-site
USD 200,000 - 300,000