Senior Distributed Systems Engineer

Lever, Inc.

Sunnyvale (CA)

On-site

USD 180,000 - 250,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

The Institute of Foundation Models designs ultra-scale GPU systems to train foundation models and optimize cross-layer performance and scalability. We seek a deeply technical engineer to co-design and optimize the communication stack for large-scale distributed training, including MoE workloads, hybrid parallelism, and fault-tolerant execution across thousands of GPUs.

This role emphasizes performance engineering, distributed debugging, and production-grade optimizations at the boundary of

Qualifications

  • Experience optimizing distributed training at 1,000+ GPU scale.
  • Hands-on with RDMA, InfiniBand, RoCE, and GPUDirect RDMA.
  • Deep familiarity with NCCL and/or UCX internals.
  • Systems programming in C/C++, Rust, or Go.
  • Experience with PyTorch in large-scale training.
  • Ability to profile and fix communication bottlenecks.

Responsibilities

  • Design and optimize expert-parallel and hybrid-parallel communication patterns.
  • Drive high-performance hierarchical collectives for MoE workloads.
  • Co-design runtime orchestration with communication topology awareness.
  • Reduce tail latency and improve determinism across thousands of GPUs.
  • Architect fault-tolerant distributed execution under real-world cluster failures.

Skills

RDMA/InfiniBand optimization
NCCL/UCX internals
C/C++ / Rust / Go systems programming
PyTorch large-scale training
Distributed debugging
Troubleshooting communication bottlene

Tools

NCCL
UCX
GPUDirect RDMA
InfiniBand
RoCE

Job description

About the Institute of Foundation Models

The Institute of Foundation Models (IFM) designs and operates ultra-scale GPU supercomputing systems to train next-generation foundation models. We believe performance, fault tolerance, and scalability are co-designed across model architecture, communication systems, runtime, and hardware topology.

This role sits at the core of that effort — driving communication performance, distributed reliability, and cross-layer optimization for large-scale training workloads.

The Mission

We are looking for a deeply technical engineer to co-design and optimize the communication stack for large-scale distributed training, including hybrid parallelism and Mixture-of-Experts (MoE) workloads.

This is not a network operations role. This is a systems-level engineering position focused on performance engineering, distributed debugging, and communication-runtime co-design.

  • Design and optimize expert-parallel and hybrid-parallel communication patterns
  • Drive high-performance hierarchical collectives for MoE workloads
  • Co-design runtime orchestration with communication topology awareness
  • Reduce tail latency and improve determinism across thousands of GPUs
  • Architect fault-tolerant distributed execution under real‑world cluster failures
Core Technical Scope
  • Communication-compute overlap and topology-aware collective optimization
  • Deep debugging of NCCL, RDMA, and custom communication layers
  • Hybrid expert parallel strategies in modern large-scale MoE systems
  • Elastic and resilient distributed job orchestration concepts
  • Congestion analysis and routing optimization across InfiniBand/RoCE fabrics
  • Microbenchmarking and performance modeling for communication-heavy workloads
Expected Technical Depth
  • Hybrid expert parallel communication for Mixture-of-Experts training
  • Scaling behavior under network pressure
  • Distributed orchestration for elastic, large-scale training
  • Fault detection and recovery in distributed GPU workloads
  • Cross-layer bottlenecks: GPU NIC PCIe NVSwitch Fabric Scheduler
Required Background
  • Experience optimizing distributed training at 1,000+ GPU scale (or equivalent depth)
  • Hands‑on expertise with RDMA, InfiniBand, RoCE, and GPUDirect RDMA
  • Deep familiarity with NCCL and/or UCX internals
  • Strong systems programming ability (C/C++, Rust, or Go)
  • Strong familiarity with modern model training frameworks such as PyTorch
  • Ability to troubleshoot and profile training performance issues related to communication bottlenecks
  • Ability to translate research ideas into production‑grade optimizations
  • Experience debugging distributed hangs, desynchronization, and performance regressions
What We Mean by "Hardcore"
  • You can explain why an communication degrades at scale and how to fix it
  • You have improved real cluster throughput via communication redesign
  • You can trace a distributed hang across ranks and identify the root cause
  • You are comfortable working at the boundary between hardware and runtime
Application Requirements
  • Include a link to your GitHub (required)
  • Provide links to relevant distributed systems, HPC, or large-scale training projects
  • Include a list of publications and/or public technical reports (if applicable)
  • Describe the hardest distributed debugging problem you solved
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Distributed Systems Engineer — High-Perf GPU
Senior Distributed Systems Engineer — High-Perf GPU

Lever, Inc. • Sunnyvale (CA)

On-site
USD 180,000 - 250,000
Cluster Design
Cluster Design

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 230,000
Large-Scale GPU Cluster Engineering Lead (GPU · Cluster · Orchestration)
Large-Scale GPU Cluster Engineering Lead (GPU · Cluster · Orchestration)

NJF Global Holdings Ltd • New York (NY)

On-site
USD 150,000 - 200,000
Principal Infrastructure Engineer, AI Cluster Performance & Validation
Principal Infrastructure Engineer, AI Cluster Performance & Validation

Nscale • New York (NY), San Francisco (CA), Seattle (WA)

On-site
USD 180,000 - 240,000
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Member of Technical Staff - ML Performance
Member of Technical Staff - ML Performance

Veeda Innovation • Northern (KY)

On-site
USD 150,000 - 230,000
GPU Systems Engineer
GPU Systems Engineer

Career Techniques • New York (NY)

On-site
USD 200,000 - 300,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Hyperbolic • San Francisco (CA)

On-site
USD 180,000 - 260,000
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda Innovation • California (MO)

Hybrid
USD 180,000 - 240,000
Member of Technical Staff — Compute Cluster
Member of Technical Staff — Compute Cluster

Kindredventures • San Francisco (CA)

On-site
USD 140,000 - 230,000