Lead AI Infrastructure Engineer: GPU Clusters & Reliability

Luma AI

San Francisco (CA)

On-site

USD 300,000 - 420,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Luma AI in San Francisco is seeking a leader to define reliability for a frontier AI infrastructure. You will architect and operate large GPU environments, pushing the limits of training and inference while partnering with research and product to scale systems and improve availability.

The role demands deep Linux/distributed systems expertise, strong Kubernetes mastery, and a track record of delivering scalable, production‑grade infrastructure.

Qualifications

  • Deep expertise in Linux and distributed systems.
  • Experience operating GPU/accelerator clusters in real production environments.
  • Strong fluency in Kubernetes and modern open‑source infrastructure.
  • Comfortable debugging across hardware → kernel → runtime → orchestration.
  • You understand how systems behave under contention and at scale.
  • You write code and build automation.
  • You think in bottlenecks, failure modes, and tradeoffs.
  • Engineers trust your judgment, especially when things break.

Responsibilities

  • Architect and operate large, heterogeneous GPU environments under extreme demand.
  • Improve utilization and performance where small gains materially change company outcomes.
  • Resolve failures that span hardware, OS, runtimes, and orchestration.
  • Eliminate entire classes of instability.
  • Build mechanisms that make heroics unnecessary.

Skills

Linux
Distributed systems
Kubernetes
GPU clusters
Open‑source infra
Debugging hardware to orchestration
Performance & bottlenecks
Automation / SRE scripting
Leadership / mentoring

Tools

Kubernetes
Linux

Job description

Luma AI in San Francisco is seeking a leader to define reliability for a frontier AI infrastructure. You will architect and operate large GPU environments, pushing the limits of training and inference while partnering with research and product to scale systems and improve availability.

The role demands deep Linux/distributed systems expertise, strong Kubernetes mastery, and a track record of delivering scalable, production‑grade infrastructure.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff AI Infra Engineer: GPU Fleet Reliability Leader
Staff AI Infra Engineer: GPU Fleet Reliability Leader

Luma AI • United States

Remote
USD 210,000 - 320,000
Senior GPU Infra Reliability Engineer - Remote
Senior GPU Infra Reliability Engineer - Remote

Luma AI • United States

Remote
USD 180,000 - 240,000
Lead Large-Scale GPU Cluster Engineer for AI Research
Lead Large-Scale GPU Cluster Engineer for AI Research

Linuxcareers • San Francisco (CA)

On-site
USD 120,000 - 180,000
GPU Cluster Infra Engineer - Reliability & Automation
GPU Cluster Infra Engineer - Reliability & Automation

Doist • San Francisco (CA)

On-site
USD 150,000 - 210,000
Equity
Health benefits
401k matching
+2
Staff AI Infrastructure Engineer
Staff AI Infrastructure Engineer

Luma AI • San Francisco (CA)

On-site
USD 300,000 - 420,000
Lead AI Systems Engineer: GPU Clusters & AI Ops
Lead AI Systems Engineer: GPU Clusters & AI Ops

Semiconductor Engineering • San Jose (CA)

On-site
USD 140,000 - 220,000
Senior AI Infrastructure Engineer — Scale GPU Clusters
Senior AI Infrastructure Engineer — Scale GPU Clusters

AI Breaking Wire • San Francisco (CA)

On-site
USD 280,000 - 400,000
Equity
Medical, dental, and vision benefits
Unlimited PTO
+2
Senior AI Cloud SRE — HPC & GPU Reliability Lead
Senior AI Cloud SRE — HPC & GPU Reliability Lead

Lambda • San Francisco (CA)

On-site
USD 180,000 - 230,000
Health, dental, and vision
401k with company match
Flexible paid time off
+2
Remote Senior Training Infrastructure Engineer—Multi-GPU AI
Remote Senior Training Infrastructure Engineer—Multi-GPU AI

Luma AI • San Francisco (CA)

Hybrid
USD 187,000 - 395,000
Senior Distributed AI Training Architect
Senior Distributed AI Training Architect

Luma AI • United States

Remote
USD 180,000 - 280,000