Staff Engineer, AI Cloud Orchestration

Lambda

United States

Remote

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Lambda, The Superintelligence Cloud, is seeking a Staff Engineer to advance our managed Kubernetes platform with bare-metal, GPU-aware orchestration for AI workloads. You will shape infrastructure across compute, storage, networking, and security, collaborating with teams to deliver reliable scale.

Ideal candidates have leadership experience in distributed systems, Kubernetes, and GPU-accelerated computing, and can bridge open-source ecosystems with internal platforms to empower AI training and

Qualifications

  • Experience leading infra projects or platforms at scale.
  • Strong background in distributed systems and cloud infrastructure.
  • Proven ability to design and deliver reliable, high-performance systems for AI workloads.

Responsibilities

  • Drive technical vision for Lambda's Managed Kubernetes bare-metal platform, focusing on scalability and high availability.
  • Integrate and extend NVIDIA open-source ecosystem: GPU Operator, NCCL, DCGM, and related projects for topology-aware scheduling.
  • Design GPU-aware orchestration systems across compute, networking, and storage layers.
  • Lead development of services powering managed Kubernetes and higher-level platform services for inference and AIOps.
  • Collaborate with Network team on CNI integrations, InfiniBand/RDMA, and GPUDirect.
  • Partner with Storage teams to define storage architecture needs for AI workloads.
  • Build foundation for Slurm on Kubernetes to run HPC workloads alongside Kubernetes workloads.
  • Shape higher-level services for model serving infrastructure and autoscaling based on inference load.

Skills

Distributed systems
Kubernetes
GPU-accelerated computing
Networking for AI workloads
Cloud infrastructure
Leadership

Tools

GPU Operator
CNI (Cilium, Multus)
RDMA/InfiniBand
Slurm on Kubernetes
NCCL
Topograph

Job description

Lambda, The Superintelligence Cloud, is seeking a Staff Engineer to advance our managed Kubernetes platform with bare-metal, GPU-aware orchestration for AI workloads. You will shape infrastructure across compute, storage, networking, and security, collaborating with teams to deliver reliable scale.

Ideal candidates have leadership experience in distributed systems, Kubernetes, and GPU-accelerated computing, and can bridge open-source ecosystems with internal platforms to empower AI training and

Get your free, confidential resume review.
or drag and drop your file here.