Inference Systems Engineer for Scalable Multimodal AI

Luma AI

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Luma AI in San Francisco is building cutting-edge multimodal AI systems. We seek a seasoned Platform Engineer to ship new model architectures into our inference engine and optimize deployments across clusters.

You will collaborate with research, engineering and infra, develop scalable scheduling, CI/CD pipelines, and tooling to measure and ensure uptime for inference workloads across thousands of GPUs. Proficiency in Python, Linux, Docker, Kubernetes, and model deployment frameworks is required;

Qualifications

  • Strong Python and system architecture skills.
  • Experience deploying models with PyTorch, Huggingface, vLLM, SGLang, or tensorRT-LLM.
  • Experience with Linux, Docker, and Kubernetes.
  • Familiarity with queues, scheduling, and fleet management at scale.
  • Bonus: RDMA networking, large-scale ML systems (>100 GPUs), FFmpeg.

Responsibilities

  • Ship new model architectures by integrating them into our inference engine.
  • Collaborate across research, engineering and infrastructure to optimize deployments.
  • Build internal tooling to measure, profile, and track inference jobs.
  • Automate, test and maintain inference services for maximum uptime and reliability.
  • Optimize deployment workflows to scale across thousands of machines.
  • Manage inference workloads across clusters and hardware providers.
  • Build scheduling systems to efficiently leverage GPU resources while meeting SLOs.
  • Develop CI/CD pipelines for processing/optimizing model checkpoints and platform components.

Skills

Python
System architecture
Model deployment
PyTorch
Huggingface
vLLM
SGLang
tensorRT-LLM
Queues & scheduling
Linux
Docker
Kubernetes
RDMA
High scale ML
FFmpeg

Job description

Luma AI in San Francisco is building cutting-edge multimodal AI systems. We seek a seasoned Platform Engineer to ship new model architectures into our inference engine and optimize deployments across clusters.

You will collaborate with research, engineering and infra, develop scalable scheduling, CI/CD pipelines, and tooling to measure and ensure uptime for inference workloads across thousands of GPUs. Proficiency in Python, Linux, Docker, Kubernetes, and model deployment frameworks is required;

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Inference Systems Engineer — Kubernetes & GPU Scale
ML Inference Systems Engineer — Kubernetes & GPU Scale

Luma AI • United States

Remote
USD 140,000 - 190,000
Software Engineer, Inference
Software Engineer, Inference

Luma AI • United States

Remote
USD 140,000 - 190,000
Software Engineer, Inference
Software Engineer, Inference

Luma AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Multimodal AI Evaluation Engineer
Senior Multimodal AI Evaluation Engineer

Luma AI • San Francisco (CA), New York (NY)

On-site
USD 170,000 - 210,000
Inference Infra Architect for Multimodal ML Systems
Inference Infra Architect for Multimodal ML Systems

Elorian AI • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
Hands-on Tech Lead: Inference Platform & Scaling
Hands-on Tech Lead: Inference Platform & Scaling

lumalabs-ai • San Francisco (CA)

On-site
USD 230,000 - 350,000
AI/ML Inference & Vision Infrastructure Engineer
AI/ML Inference & Vision Infrastructure Engineer

Front Door Defense • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 210,000
Restaurant d'entreprise
Indemnités de stage/alternance
Lead AI Infrastructure Engineer: GPU Clusters & Reliability
Lead AI Infrastructure Engineer: GPU Clusters & Reliability

Luma AI • San Francisco (CA)

On-site
USD 300,000 - 420,000
Distributed RL Systems Engineer — Scale Training & Inference
Distributed RL Systems Engineer — Scale Training & Inference

Luma AI • United States

Remote
USD 180,000 - 240,000
Senior Distributed AI Training Architect
Senior Distributed AI Training Architect

Luma AI • United States

Remote
USD 180,000 - 280,000