Staff AI Infrastructure Engineer - Frontier GPU Scale

lumalabs-ai

San Francisco (CA)

On-site

USD 230,000 - 360,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Luma AI is seeking a senior infrastructure engineering leader to define reliability for a new generation of AI infrastructure. You will guide GPU clusters, runtimes, and orchestration while partnering with research and product teams to scale models and workloads.

You will set technical direction, hire and develop engineers, and shape platform strategy to ensure performance and reliability at scale across a fast-moving research environment.

Qualifications

  • Deep expertise in Linux and distributed systems.
  • Experience operating GPU/accelerator clusters in production.
  • Strong fluency in Kubernetes and modern open-source infra.

Responsibilities

  • Architect and operate large, heterogeneous GPU environments under extreme demand.
  • Improve utilization and performance where small gains materially change outcomes.
  • Resolve failures spanning hardware, OS, runtimes, and orchestration.
  • Eliminate entire classes of instability and build scalable reliability.
  • Lead and mentor engineers to raise reliability standards across the company.

Skills

Linux
Distributed systems
Kubernetes
GPU clusters
Automation
Debugging
Performance optimization
Code writing

Job description

Luma AI is seeking a senior infrastructure engineering leader to define reliability for a new generation of AI infrastructure. You will guide GPU clusters, runtimes, and orchestration while partnering with research and product teams to scale models and workloads.

You will set technical direction, hire and develop engineers, and shape platform strategy to ensure performance and reliability at scale across a fast-moving research environment.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead AI Infrastructure Engineer: GPU Clusters & Reliability
Lead AI Infrastructure Engineer: GPU Clusters & Reliability

Luma AI • San Francisco (CA)

On-site
USD 300,000 - 420,000
Staff AI Infrastructure Engineer
Staff AI Infrastructure Engineer

Luma AI • San Francisco (CA)

On-site
USD 300,000 - 420,000
Staff AI Infrastructure Engineer
Staff AI Infrastructure Engineer

lumalabs-ai • San Francisco (CA)

On-site
USD 230,000 - 360,000
Senior SRE — AI GPU Infra Architect (Multi-Cloud)
Senior SRE — AI GPU Infra Architect (Multi-Cloud)

lumalabs-ai • San Francisco (CA)

On-site
USD 170,000 - 290,000
Staff AI Systems Architect
Staff AI Systems Architect

Luma • Redwood City (CA)

On-site
USD 260,000 - 380,000
Senior AI Infrastructure Lead: GPU Clusters & LLMs
Senior AI Infrastructure Lead: GPU Clusters & LLMs

Cadence Design Systems • San Jose (CA)

On-site
USD 137,000 - 254,000
Senior GPU Cluster Infra Engineer | Remote
Senior GPU Cluster Infra Engineer | Remote

AISafety • Berkeley (CA)

Hybrid
USD 120,000 - 180,000
Health Insurance
401(k) match
PTO 25 days per year
+3
Senior Site Reliability Engineer
Senior Site Reliability Engineer

lumalabs-ai • San Francisco (CA)

On-site
USD 170,000 - 290,000
Remote Senior Training Infrastructure Engineer—Multi-GPU AI
Remote Senior Training Infrastructure Engineer—Multi-GPU AI

Luma AI • San Francisco (CA)

Hybrid
USD 187,000 - 395,000
Distributed AI Training Engineer (GPU Clusters)
Distributed AI Training Engineer (GPU Clusters)

lumalabs-ai • San Francisco (CA)

On-site
USD 188,000 - 395,000