Hands-On Tech Lead for Large-Scale Inference

Luma AI

United States

Remote

USD 180,000 - 320,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Luma AI seeks a hands-on Tech Lead Manager to own the inference serving stack, including routing, scheduling, and fleet-wide orchestration across thousands of GPUs and multiple clouds. The role demands staying hands-on while hiring, growing the team, and steering technical direction.

Expected to lead the team through the first 90 days by diagnosing the stack, shipping a meaningful platform improvement, and scaling reliability and cost efficiency for production multi-cluster deployments.

Qualifications

  • 8+ years in large-scale distributed systems or ML infrastructure.

Responsibilities

  • Own and troubleshoot the serving stack latency, reliability, and cost.
  • Lead, grow, and coach inference engineering team including on-call and postmortems.
  • Set technical roadmap for serving engines, routing, scheduling, and autoscaling.
  • Own platform SLOs and budget, monitor GPU utilization and cost per generation.
  • Collaborate with research to ship architectures to production and integrate with online eval loops.

Skills

Distributed systems
ML infrastructure
LLM serving
Python
PyTorch
Kubernetes
Scheduling
GPU clusters
Cost optimization
Incident response

Tools

vLLM
SGLang
TensorRT-LLM
KV-cache management

Job description

Luma AI seeks a hands-on Tech Lead Manager to own the inference serving stack, including routing, scheduling, and fleet-wide orchestration across thousands of GPUs and multiple clouds. The role demands staying hands-on while hiring, growing the team, and steering technical direction.

Expected to lead the team through the first 90 days by diagnosing the stack, shipping a meaningful platform improvement, and scaling reliability and cost efficiency for production multi-cluster deployments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff AI Infra Engineer: GPU Fleet Reliability Leader
Staff AI Infra Engineer: GPU Fleet Reliability Leader

Luma AI • United States

Remote
USD 210,000 - 320,000
Lead AI Infrastructure Engineer: GPU Clusters & Reliability
Lead AI Infrastructure Engineer: GPU Clusters & Reliability

Luma AI • San Francisco (CA)

On-site
USD 300,000 - 420,000
Software Engineer, Inference
Software Engineer, Inference

Luma AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Lead, AI Compute Infra for Scalable LLM Inference
Lead, AI Compute Infra for Scalable LLM Inference

ByteDance • Seattle (WA)

On-site
USD 232,560 - 427,500
Inference Systems Engineer for Scalable Multimodal AI
Inference Systems Engineer for Scalable Multimodal AI

Luma AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Engineering Manager: Inference Infrastructure Leader
Engineering Manager: Inference Infrastructure Leader

EngineersOfAI • New York (NY), Northern (KY)

Hybrid
USD 230,000 - 360,000
Senior AI Inference DevOps Engineer
Senior AI Inference DevOps Engineer

Lila Sciences • Cambridge (MA)

On-site
USD 192,000 - 272,000
Equity in company equity
Medical, dental, vision coverage
Generous PTO and holidays
Senior GPU Infra Reliability Engineer - Remote
Senior GPU Infra Reliability Engineer - Remote

Luma AI • United States

Remote
USD 180,000 - 240,000
Tech Lead, AI Inference & High-Performance Systems
Tech Lead, AI Inference & High-Performance Systems

Visa Hunt • United States

On-site
USD 180,000 - 240,000
Medical, Dental, Vision
401(K)
Flexible Time Off
Inference Optimization Lead
Inference Optimization Lead

Up Top • United States

Hybrid
USD 180,000 - 320,000