Lead Large-Scale GPU Cluster Engineer for AI Research

Linuxcareers

San Francisco (CA)

On-site

USD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Linuxcareers in San Francisco is building AI research infrastructure. You will design, deploy, and operate large-scale GPU clusters powering training, evaluation, and serving for the research team.

The role emphasizes extending orchestration with Kubernetes/Slurm, building unified interfaces, and ensuring reliability with observability. You'll work with researchers to optimize performance and placement.

Qualifications

  • Experience operating large-scale GPU clusters.
  • Proficient with Kubernetes, Slurm and Docker.
  • Strong Linux, networking, storage, and IaC skills.
  • Familiar with cloud ML services (GCP/AWS/Azure).
  • CUDA/NCCL experience and profiling for distributed workloads.
  • Ability to deliver end-to-end from requirements through autonomous execution.

Responsibilities

  • Design, deploy, and operate large distributed GPU clusters end-to-end, including provisioning, imaging, upgrades, and capacity planning
  • Extend scheduling and orchestration systems such as Kubernetes and Slurm for topology-aware placement, preemption, quotas, and multi-tenancy across training and inference workloads
  • Build software that abstracts cluster management and presents a unified, self-serve interface to researchers and engineers
  • Own cluster storage and artifact paths for checkpoints and logs, with clear retention and lineage
  • Monitor and continuously improve reliability and error recovery; build observability to catch failures before researchers do
  • Partner with researchers to unblock large-scale runs and advise on performance and placement trade-offs

Skills

GPU clusters
Kubernetes
Slurm
Docker
Linux systems
Networking
Storage
Infrastructure as code
Cloud platforms
CUDA/NCCL
Observability
End-to-end ownership

Tools

Kubernetes
Slurm
Docker

Job description

Linuxcareers in San Francisco is building AI research infrastructure. You will design, deploy, and operate large-scale GPU clusters powering training, evaluation, and serving for the research team.

The role emphasizes extending orchestration with Kubernetes/Slurm, building unified interfaces, and ensuring reliability with observability. You'll work with researchers to optimize performance and placement.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead AI Infrastructure Engineer: GPU Clusters & Reliability
Lead AI Infrastructure Engineer: GPU Clusters & Reliability

Luma AI • San Francisco (CA)

On-site
USD 300,000 - 420,000
Senior AI GPU Cluster Architect
Senior AI GPU Cluster Architect

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior AI Infrastructure Engineer — Scale GPU Clusters
Senior AI Infrastructure Engineer — Scale GPU Clusters

AI Breaking Wire • San Francisco (CA)

On-site
USD 280,000 - 400,000
Equity
Medical, dental, and vision benefits
Unlimited PTO
+2
AI Infra & Cluster Engineer — Scale GPU/CPU Orchestration
AI Infra & Cluster Engineer — Scale GPU/CPU Orchestration

Linuxcareers • San Francisco (CA)

On-site
USD 120,000 - 160,000
GPU Systems Engineer for AI Training Clusters
GPU Systems Engineer for AI Training Clusters

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Senior AI Infrastructure Engineer — Scalable GPU Clusters
Senior AI Infrastructure Engineer — Scalable GPU Clusters

NVIDIA AI • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Member of Technical Staff — Compute Cluster
Member of Technical Staff — Compute Cluster

Linuxcareers • San Francisco (CA)

On-site
USD 120,000 - 180,000
GPU Infra Solutions Architect for Large-Scale AI Clusters
GPU Infra Solutions Architect for Large-Scale AI Clusters

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Lead AI Infrastructure Architect for Large-Scale GPU Clusters
Lead AI Infrastructure Architect for Large-Scale GPU Clusters

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity and benefits
Senior AI Infrastructure Engineer | Scale GPU Clusters
Senior AI Infrastructure Engineer | Scale GPU Clusters

Fuel Talent LLC • Seattle (WA)

Hybrid
USD 126,000 - 189,000