Staff AI Infra Engineer: GPU Fleet Reliability Leader

Luma AI

United States

Remote

USD 210,000 - 320,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Luma AI is seeking a Staff AI Infrastructure Engineer to own the reliability of our large GPU fleet, spanning scheduling, efficiency, and resilience. This role involves close-to-the-metal work with kernels, containers, schedulers, networking, and storage, under high-demand conditions.

You will lead and grow a team, partner with research, and shape how infrastructure evolves as models scale, ensuring latency remains predictable while performance improves.

Qualifications

  • Deep expertise in Linux and distributed systems.
  • Experience operating GPU or accelerator clusters in real production environments.
  • Strong fluency in Kubernetes and modern open-source infrastructure.

Responsibilities

  • Architect and operate large, heterogeneous GPU environments under extreme demand.
  • Resolve failures spanning hardware, OS, runtimes, and orchestration, and eliminate instability.
  • Define how infrastructure and workloads evolve as cluster size and concurrency grow.

Skills

Linux
Distributed systems
GPU clusters
Kubernetes
Open-source infra
Automation
Production ownership

Tools

Docker

Job description

Luma AI is seeking a Staff AI Infrastructure Engineer to own the reliability of our large GPU fleet, spanning scheduling, efficiency, and resilience. This role involves close-to-the-metal work with kernels, containers, schedulers, networking, and storage, under high-demand conditions.

You will lead and grow a team, partner with research, and shape how infrastructure evolves as models scale, ensuring latency remains predictable while performance improves.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead AI Infrastructure Engineer: GPU Clusters & Reliability
Lead AI Infrastructure Engineer: GPU Clusters & Reliability

Luma AI • San Francisco (CA)

On-site
USD 300,000 - 420,000
Senior GPU Infra Reliability Engineer - Remote
Senior GPU Infra Reliability Engineer - Remote

Luma AI • United States

Remote
USD 180,000 - 240,000
Autonomous AI Infrastructure Engineer: GPU Fleet Mastery
Autonomous AI Infrastructure Engineer: GPU Fleet Mastery

Together • San Francisco (CA)

On-site
USD 190,000 - 270,000
Health insurance
Startup equity
Competitive benefits
Staff AI Infrastructure Engineer
Staff AI Infrastructure Engineer

Luma AI • San Francisco (CA)

On-site
USD 300,000 - 420,000
Engineering Manager, AI Cloud GPU Fleet
Engineering Manager, AI Cloud GPU Fleet

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Engineering Manager, GPU Fleet & Infra (Hybrid)
Engineering Manager, GPU Fleet & Infra (Hybrid)

Socket.dev • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Health, dental, and vision coverage
Wellness stipend
401k plan with company match (USA)
Engineering Manager - AI Cloud Infrastructure & GPU Fleet
Engineering Manager - AI Cloud Infrastructure & GPU Fleet

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Health, dental, and vision coverage
401k with company match
Flexible paid time off
Hands-On Tech Lead for Large-Scale Inference
Hands-On Tech Lead for Large-Scale Inference

Luma AI • United States

Remote
USD 180,000 - 320,000
AI Infra Engineer: GPU Fleet Automation
AI Infra Engineer: GPU Fleet Automation

Fal.ai Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 250,000
Relocation assistance
Health, dental, and vision insurance (
Regular team events and offsites
GPU Fleet Infra Engineer — Scale, Automation & Kubernetes
GPU Fleet Infra Engineer — Scale, Automation & Kubernetes

OpenAI • New York (NY)

Hybrid
USD 180,000 - 240,000
Relocation assistance
Hybrid work model