Hands-on Tech Lead: Inference Platform & Scaling

lumalabs-ai

San Francisco (CA)

On-site

USD 230,000 - 350,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Luma is hiring a Tech Lead Manager to guide the inference platform team that governs the serving stack across GPUs and clusters. The role blends hands-on engineering with leadership, focusing on performance, reliability, and cost efficiency to enable rapid production of large models.

You will set technical direction, own roadmap, and mentor engineers while collaborating with research, product, and infrastructure to push new model architectures into production.

Qualifications

  • 8+ years of engineering experience in large-scale distributed systems or ML infrastructure.
  • Experience running inference platforms at scale across multiple clusters or clouds.
  • Technical leadership experience with hands-on work and growth mindset.
  • Deep expertise in LLM and foundation-model serving engines (vLLM, SGLang, TensorRT-LLM).
  • Strong skills in serving-performance tools: batching, KV-cache, quantization, decoding.
  • Proficiency in Python and PyTorch; Kubernetes at scale.
  • Experience with queues, scheduling, and fleet management at scale.

Responsibilities

  • Hands-on design, build, and debugging in the serving stack while leading the team.
  • Hire, coach, and grow inference engineers; build operational culture.
  • Set roadmap for serving platform: routing, autoscaling, observability, deployment.
  • Own platform SLOs, latency, availability, GPU utilization, and cost per generation.
  • Collaborate with research to ship models to production and integrate RL loops.

Skills

Distributed systems
ML infrastructure
Leadership
Hands-on coding
Python
PyTorch
Kubernetes
System design

Tools

TensorRT-LLM
vLLM
SGLang
Kubernetes

Job description

Luma is hiring a Tech Lead Manager to guide the inference platform team that governs the serving stack across GPUs and clusters. The role blends hands-on engineering with leadership, focusing on performance, reliability, and cost efficiency to enable rapid production of large models.

You will set technical direction, own roadmap, and mentor engineers while collaborating with research, product, and infrastructure to push new model architectures into production.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Inference Systems Engineer — Kubernetes & GPU Scale
ML Inference Systems Engineer — Kubernetes & GPU Scale

Luma AI • United States

Remote
USD 140,000 - 190,000
Tech Lead Manager, Inference
Tech Lead Manager, Inference

lumalabs-ai • San Francisco (CA)

On-site
USD 230,000 - 350,000
Software Engineer, Inference
Software Engineer, Inference

Luma AI • United States

Remote
USD 140,000 - 190,000
Software Engineer, Inference
Software Engineer, Inference

Luma AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Inference Systems Engineer for Scalable Multimodal AI
Inference Systems Engineer for Scalable Multimodal AI

Luma AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Lead AI Infrastructure Engineer: GPU Clusters & Reliability
Lead AI Infrastructure Engineer: GPU Clusters & Reliability

Luma AI • San Francisco (CA)

On-site
USD 300,000 - 420,000
Engineering Manager: Inference Infrastructure Leader
Engineering Manager: Inference Infrastructure Leader

EngineersOfAI • New York (NY), Northern (KY)

Hybrid
USD 230,000 - 360,000
Technical Lead: Inference Benchmarking & ML Infra
Technical Lead: Inference Benchmarking & ML Infra

NVIDIA • United States

On-site
USD 224,000 - 356,500
Equity and benefits
Comprehensive benefits package
Competitive salaries
Tech Lead Manager for Scalable LLM Training Platform
Tech Lead Manager for Scalable LLM Training Platform

United States Digital Space LLC • San Francisco (CA), New York (NY)

On-site
USD 290,000 - 363,000
ML Platform Engineer: Scale AI & Inference
ML Platform Engineer: Scale AI & Inference

Apply • San Francisco (CA)

Hybrid
USD 245,000 - 345,000
Flexible Time Off
Health Insurance
Work From Home Allowance
+2