Lead Large-Scale GPU Cluster Engineer for AI Research

Linuxcareers

San Francisco (CA)

On-site

USD 120,000 - 180,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Linuxcareers in San Francisco is building AI research infrastructure. You will design, deploy, and operate large-scale GPU clusters powering training, evaluation, and serving for the research team.

The role emphasizes extending orchestration with Kubernetes/Slurm, building unified interfaces, and ensuring reliability with observability. You'll work with researchers to optimize performance and placement.

Qualifications

  • Experience operating large-scale GPU clusters.
  • Proficient with Kubernetes, Slurm and Docker.
  • Strong Linux, networking, storage, and IaC skills.
  • Familiar with cloud ML services (GCP/AWS/Azure).
  • CUDA/NCCL experience and profiling for distributed workloads.
  • Ability to deliver end-to-end from requirements through autonomous execution.

Responsibilities

  • Design, deploy, and operate large distributed GPU clusters end-to-end, including provisioning, imaging, upgrades, and capacity planning
  • Extend scheduling and orchestration systems such as Kubernetes and Slurm for topology-aware placement, preemption, quotas, and multi-tenancy across training and inference workloads
  • Build software that abstracts cluster management and presents a unified, self-serve interface to researchers and engineers
  • Own cluster storage and artifact paths for checkpoints and logs, with clear retention and lineage
  • Monitor and continuously improve reliability and error recovery; build observability to catch failures before researchers do
  • Partner with researchers to unblock large-scale runs and advise on performance and placement trade-offs

Skills

GPU clusters
Kubernetes
Slurm
Docker
Linux systems
Networking
Storage
Infrastructure as code
Cloud platforms
CUDA/NCCL
Observability
End-to-end ownership

Tools

Kubernetes
Slurm
Docker

Job description

Linuxcareers in San Francisco is building AI research infrastructure. You will design, deploy, and operate large-scale GPU clusters powering training, evaluation, and serving for the research team.

The role emphasizes extending orchestration with Kubernetes/Slurm, building unified interfaces, and ensuring reliability with observability. You'll work with researchers to optimize performance and placement.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead AI Infrastructure Engineer: GPU Clusters & Reliability
Lead AI Infrastructure Engineer: GPU Clusters & Reliability

Luma AI • San Francisco (CA)

On-site
USD 300,000 - 420,000
AI Infra & Cluster Engineer — Scale GPU/CPU Orchestration
AI Infra & Cluster Engineer — Scale GPU/CPU Orchestration

Linuxcareers • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior GPU Infrastructure Engineer — HPC & Clusters
Senior GPU Infrastructure Engineer — HPC & Clusters

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
Member of Technical Staff — Compute Cluster
Member of Technical Staff — Compute Cluster

Linuxcareers • San Francisco (CA)

On-site
USD 120,000 - 180,000
Lead AI Infrastructure Architect for Large-Scale GPU Clusters
Lead AI Infrastructure Architect for Large-Scale GPU Clusters

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 356,500
Equity and benefits
Staff Infra Engineer – Large-Scale GPU Training
Staff Infra Engineer – Large-Scale GPU Training

Hark • San Jose (CA)

On-site
USD 180,000 - 450,000
Senior AI Infrastructure Engineer | Scale GPU Clusters
Senior AI Infrastructure Engineer | Scale GPU Clusters

Fuel Talent LLC • Seattle (WA)

Hybrid
USD 126,000 - 189,000
ML Infra Engineer — GPU Clusters & Distributed Systems
ML Infra Engineer — GPU Clusters & Distributed Systems

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 250,000
Industry-leading compensation and/or:?
Unlimited PTO
Top-tier medical, dental, and vision
+1
Senior AI Infrastructure Lead: GPU Clusters & LLMs
Senior AI Infrastructure Lead: GPU Clusters & LLMs

Cadence Design Systems • San Jose (CA)

On-site
USD 137,000 - 254,000
Lead HPC/GPU Cluster Architect | Automate & Scale
Lead HPC/GPU Cluster Architect | Automate & Scale

The San Francisco Compute Company • Boston (MA)

Hybrid
USD 140,000 - 200,000
Generous equity grant
Visa Sponsorships
Retirement matching
+5