Lead AI Infrastructure Engineer: GPU Clusters & Reliability

Luma AI

San Francisco (CA)

On-site

USD 300,000 - 420,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Luma AI in San Francisco is seeking a leader to define reliability for a frontier AI infrastructure. You will architect and operate large GPU environments, pushing the limits of training and inference while partnering with research and product to scale systems and improve availability.

The role demands deep Linux/distributed systems expertise, strong Kubernetes mastery, and a track record of delivering scalable, production‑grade infrastructure.

Qualifications

  • Deep expertise in Linux and distributed systems.
  • Experience operating GPU/accelerator clusters in real production environments.
  • Strong fluency in Kubernetes and modern open‑source infrastructure.
  • Comfortable debugging across hardware → kernel → runtime → orchestration.
  • You understand how systems behave under contention and at scale.
  • You write code and build automation.
  • You think in bottlenecks, failure modes, and tradeoffs.
  • Engineers trust your judgment, especially when things break.

Responsibilities

  • Architect and operate large, heterogeneous GPU environments under extreme demand.
  • Improve utilization and performance where small gains materially change company outcomes.
  • Resolve failures that span hardware, OS, runtimes, and orchestration.
  • Eliminate entire classes of instability.
  • Build mechanisms that make heroics unnecessary.

Skills

Linux
Distributed systems
Kubernetes
GPU clusters
Open‑source infra
Debugging hardware to orchestration
Performance & bottlenecks
Automation / SRE scripting
Leadership / mentoring

Tools

Kubernetes
Linux

Job description

Luma AI in San Francisco is seeking a leader to define reliability for a frontier AI infrastructure. You will architect and operate large GPU environments, pushing the limits of training and inference while partnering with research and product to scale systems and improve availability.

The role demands deep Linux/distributed systems expertise, strong Kubernetes mastery, and a track record of delivering scalable, production‑grade infrastructure.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Large-Scale GPU Cluster Engineer for AI Research
Lead Large-Scale GPU Cluster Engineer for AI Research

Linuxcareers • San Francisco (CA)

On-site
USD 120,000 - 180,000
Senior AI Infrastructure Lead: GPU Clusters & LLMs
Senior AI Infrastructure Lead: GPU Clusters & LLMs

Cadence Design Systems • San Jose (CA)

On-site
USD 137,000 - 254,000
ML Infra Engineer — GPU Clusters & Distributed Systems
ML Infra Engineer — GPU Clusters & Distributed Systems

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 250,000
Industry-leading compensation and/or:?
Unlimited PTO
Top-tier medical, dental, and vision
+1
ML Inference Systems Engineer — Kubernetes & GPU Scale
ML Inference Systems Engineer — Kubernetes & GPU Scale

Luma AI • United States

Remote
USD 140,000 - 190,000
Staff AI Infrastructure Engineer
Staff AI Infrastructure Engineer

Luma AI • San Francisco (CA)

On-site
USD 300,000 - 420,000
Senior AI Infrastructure Engineer - GPU & Kubernetes
Senior AI Infrastructure Engineer - GPU & Kubernetes

HCL Technologies Limited • California (MO)

On-site
USD 120,000 - 180,000
401(k) retirement plan
Paid time off (PTO)
Paid holidays
+1
Senior GPU Infrastructure Engineer — HPC & Clusters
Senior GPU Infrastructure Engineer — HPC & Clusters

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
Senior Distributed AI Training Architect
Senior Distributed AI Training Architect

Luma AI • United States

Remote
USD 180,000 - 280,000
Remote Senior Training Infrastructure Engineer—Multi-GPU AI
Remote Senior Training Infrastructure Engineer—Multi-GPU AI

Luma AI • San Francisco (CA)

Hybrid
USD 187,500 - 395,000
GPU Cluster Infra Lead — Tech Strategy & Team Growth
GPU Cluster Infra Lead — Tech Strategy & Team Growth

Far Ai • United States

Remote
USD 180,000 - 250,000