Head of Global Compute Capacity & Platform Strategy

lumalabs-ai

San Francisco (CA)

On-site

USD 250,000 - 450,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

LumaLabs AI in San Francisco is seeking a senior Compute leader to own the company’s global compute footprint, bridging capacity strategy, capital allocation, and systems architecture. You will design scaling roadmaps from silicon up, ensuring uninterrupted runway for research and robotics teams to ship frontier world models.

As a member of the executive team, you will turn capital into capability, overseeing the platform, vendor partnerships, and the multi-year roadmap to deliver scalable

Qualifications

  • 10+ years of engineering leadership in large-scale distributed systems.
  • Deep fluency in high-performance cluster topology and distributed training economics.
  • Direct operational oversight of 10k+ accelerator environments.

Responsibilities

  • Architect multi-year compute strategy and capacity roadmap.
  • Direct the platform org including infra, datacenters, and operations teams.
  • Maximize MFU and optimize fleet utilization for flagship training runs.
  • Negotiate and manage large capital deployments with Finance.
  • Unify global capacity to support world-model training and real-time inference.
  • Interface with NVIDIA, AMD, hyperscalers and frontier silicon vendors.

Skills

Compute strategy leadership
Capacity planning
Vendor partnerships
High-performance distributed systems
Budget & capital allocation

Tools

InfiniBand/RoCE
GPU topology
On-prem vs cloud strategy

Job description

The Role

Compute is the ultimate physical and financial prerequisite for the robotics foundation models we are building. This role owns Luma’s global compute footprint end-to-end—bridging macro capacity strategy, multi-million dollar capital allocation, and top-tier systems architecture. You will design our scaling roadmap from the silicon up, ensuring our research and robotics teams have the uninterrupted runway they need to ship frontier world models. As a member of the executive team, you will be the single person responsible for turning capital into capability.

What You'll Do
  • Architect Multi-Year Compute Strategy: Lead capacity planning, global vendor and cloud partnerships, on-prem vs. cloud mix, and accelerator supply chain roadmaps (H/B-series GPUs, custom silicon evaluation).
  • Direct the Platform Org: Provide strategic leadership to our infrastructure, distributed systems, and datacenter operations teams—scaling the organization to support next-generation compute demands.
  • Maximize Fleet Utilization: Oversee the architectural efficiency of our cluster configurations to deliver >50% Model Flops Utilization (MFU) on flagship training runs.
  • Command a Megawatt Budget: Negotiate, secure, and operate our largest-scale capital deployments for compute infrastructure, partnering directly with Finance to optimize unit economics and risk management.
  • Unify Global Capacity: Champion the platform strategy that enables world-model training, heavy simulation rollouts, and real-time on-robot inference to seamlessly share a single, elastic fleet.
  • Act as Principal Executive Interface: Serve as the primary commercial and strategic bridge to NVIDIA, AMD, hyperscalers, and frontier silicon vendors.
Qualifications
  • 10+ years of engineering leadership experience in large-scale distributed systems, infrastructure, or technical supply chain, with a proven track record of leading compute platform strategy at a frontier AI lab, hyperscaler, or major autonomy program.
  • Deep technical & commercial fluency in high-performance cluster topology, high-speed interconnects (InfiniBand/RoCE), large-scale data systems, and the economics of distributed training architectures.
  • Direct operational oversight of 10k+ accelerator environments in high-performance production settings.
Preferred qualifications
  • Scale Credentials: Experience orchestrating capital or infrastructure for training runs at the >100B-parameter or >100k-GPU-day scale.
  • Robotics/Autonomy Context: Familiarity with the unique capacity and latency demands of edge-to-cloud inference and real-time autonomous systems.
Compensation

The base pay range for this role is $250,000 – $450,000 per year.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
Global Compute Architect & Platform Strategy Leader
Global Compute Architect & Platform Strategy Leader

lumalabs-ai • San Francisco (CA)

On-site
USD 250,000 - 450,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • United States

On-site
USD 120,000 - 150,000
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Member of Technical Staff - Datacenter Operations
Member of Technical Staff - Datacenter Operations

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Member of Technical Staff - Datacenter Operations
Member of Technical Staff - Datacenter Operations

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 300,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Primeintellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

lumalabs-ai • San Francisco (CA)

On-site
USD 170,000 - 290,000
Member of Technical Staff - Datacenter Operations
Member of Technical Staff - Datacenter Operations

Primeintellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Member of Technical Staff - Datacenter Operations
Member of Technical Staff - Datacenter Operations

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000