Inference Systems Engineer for Scalable Multimodal AI

Luma AI

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Luma AI in San Francisco is building cutting-edge multimodal AI systems. We seek a seasoned Platform Engineer to ship new model architectures into our inference engine and optimize deployments across clusters.

You will collaborate with research, engineering and infra, develop scalable scheduling, CI/CD pipelines, and tooling to measure and ensure uptime for inference workloads across thousands of GPUs. Proficiency in Python, Linux, Docker, Kubernetes, and model deployment frameworks is required;

Qualifications

  • Strong Python and system architecture skills.
  • Experience deploying models with PyTorch, Huggingface, vLLM, SGLang, or tensorRT-LLM.
  • Experience with Linux, Docker, and Kubernetes.
  • Familiarity with queues, scheduling, and fleet management at scale.
  • Bonus: RDMA networking, large-scale ML systems (>100 GPUs), FFmpeg.

Responsibilities

  • Ship new model architectures by integrating them into our inference engine.
  • Collaborate across research, engineering and infrastructure to optimize deployments.
  • Build internal tooling to measure, profile, and track inference jobs.
  • Automate, test and maintain inference services for maximum uptime and reliability.
  • Optimize deployment workflows to scale across thousands of machines.
  • Manage inference workloads across clusters and hardware providers.
  • Build scheduling systems to efficiently leverage GPU resources while meeting SLOs.
  • Develop CI/CD pipelines for processing/optimizing model checkpoints and platform components.

Skills

Python
System architecture
Model deployment
PyTorch
Huggingface
vLLM
SGLang
tensorRT-LLM
Queues & scheduling
Linux
Docker
Kubernetes
RDMA
High scale ML
FFmpeg

Job description

Luma AI in San Francisco is building cutting-edge multimodal AI systems. We seek a seasoned Platform Engineer to ship new model architectures into our inference engine and optimize deployments across clusters.

You will collaborate with research, engineering and infra, develop scalable scheduling, CI/CD pipelines, and tooling to measure and ensure uptime for inference workloads across thousands of GPUs. Proficiency in Python, Linux, Docker, Kubernetes, and model deployment frameworks is required;

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Inference
Software Engineer, Inference

Luma AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Multimodal AI Evaluation Engineer
Senior Multimodal AI Evaluation Engineer

Luma AI • San Francisco (CA), New York (NY)

On-site
USD 170,000 - 210,000
Inference Infra Architect for Multimodal ML Systems
Inference Infra Architect for Multimodal ML Systems

Elorian AI • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
ML Systems Engineer - Scalable Training & Inference
ML Systems Engineer - Scalable Training & Inference

Scale AI, Inc. • New York (NY)

On-site
USD 189,000 - 237,000
Equity
Benefits
Commuter stipend
ML Research Data Infrastructure Engineer
ML Research Data Infrastructure Engineer

Luma • Redwood City (CA)

On-site
USD 170,000 - 360,000
Lead AI Infrastructure Engineer: GPU Clusters & Reliability
Lead AI Infrastructure Engineer: GPU Clusters & Reliability

Luma AI • San Francisco (CA)

On-site
USD 300,000 - 420,000
Distributed RL Systems Engineer — Scale Training & Inference
Distributed RL Systems Engineer — Scale Training & Inference

Luma AI • United States

Remote
USD 180,000 - 240,000
Distributed LLM Inference & Optimization Engineer
Distributed LLM Inference & Optimization Engineer

Together AI • San Francisco (CA)

On-site
USD 160,000 - 230,000
Startup equity
Health insurance
Competitive benefits
Senior Distributed AI Training Architect
Senior Distributed AI Training Architect

Luma AI • United States

Remote
USD 180,000 - 280,000
Research Scientist / Engineer – Training Infrastructure
Research Scientist / Engineer – Training Infrastructure

Luma AI • San Francisco (CA)

Hybrid
USD 187,000 - 395,000