ML Inference Systems Engineer: Scale GPU Deployments

Luma

United Kingdom

On-site

GBP 90,000 - 150,000

Full time

28 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Luma is building unified general intelligence that can generate, understand, and operate in the physical world. This role owns large-scale inference systems, integrating new architectures into the inference engine, and scaling deployments across thousands of machines.

You will work with research, engineering and infrastructure to optimize model efficiency, build tooling, automate inference services, and manage GPU resources with Kubernetes and CI/CD for model checkpoints.

Qualifications

  • Strong Python and system-architecture skills.
  • Experience deploying models with PyTorch, Hugging Face, vLLM, SGLang, TensorRT-LLM, or similar.
  • Experience with queues, scheduling, traffic control, and fleet management at scale.
  • Experience with Linux, Docker, and Kubernetes, and with orchestration, deployment, and scheduling.
  • Familiarity with Redis and S3-compatible storage.

Responsibilities

  • Ship new model architectures by integrating them into the inference engine.
  • Collaborate across research, engineering, and infrastructure to optimize model efficiency and deployments.
  • Build internal tooling to measure, profile, and track the lifetime of inference jobs and workflows.
  • Automate, test, and maintain inference services for maximum uptime and reliability.
  • Manage and optimize inference workloads across clusters and hardware providers, and scale deployments across thousands of machines.
  • Build scheduling systems that use expensive GPU resources optimally while meeting SLOs, and maintain CI/CD for model checkpoints and SDKs.

Skills

Python
System architecture
PyTorch
Hugging Face
vLLM
SGLang
TensorRT-LLM
Scheduling at scale
Linux
Docker
Kubernetes
Redis
S3 storage

Tools

Docker
Kubernetes
Redis
S3 storage

Job description

Luma is building unified general intelligence that can generate, understand, and operate in the physical world. This role owns large-scale inference systems, integrating new architectures into the inference engine, and scaling deployments across thousands of machines.

You will work with research, engineering and infrastructure to optimize model efficiency, build tooling, automate inference services, and manage GPU resources with Kubernetes and CI/CD for model checkpoints.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Inference Systems Engineer (GPU/Cluster Scaling)
Inference Systems Engineer (GPU/Cluster Scaling)

AItoolnavio • Greater London

Hybrid
GBP 100,000 - 180,000
Senior Systems Engineer - Large-Scale AI Inference
Senior Systems Engineer - Large-Scale AI Inference

Speedrun Talent Network • Greater London

Hybrid
GBP 90,000 - 140,000
Software Engineer, Inference
Software Engineer, Inference

Speedrun Talent Network • Greater London

Hybrid
GBP 90,000 - 140,000
Software Engineer, Inference
Software Engineer, Inference

AItoolnavio • Greater London

Hybrid
GBP 100,000 - 180,000
Senior Distributed Training Engineer — Large-Scale GPU Systems
Senior Distributed Training Engineer — Large-Scale GPU Systems

Speedrun Talent Network • Greater London

Hybrid
GBP 120,000 - 180,000
RL Infrastructure Engineer - Scale & Post-Training Systems
RL Infrastructure Engineer - Scale & Post-Training Systems

AItoolnavio • Greater London

Hybrid
GBP 120,000 - 180,000
Lead RL Infra Engineer: Scale Training & Environments
Lead RL Infra Engineer: Scale Training & Environments

Luma • Greater London

On-site
GBP 147,000 - 299,000
ML/AI Engineer: Scalable ML Ops & GPU Inference
ML/AI Engineer: Scalable ML Ops & GPU Inference

Lloyds Bank plc • Manchester

Hybrid
GBP 73,000 - 81,000
Hybrid Working
Job Share
Pension contribution up to 15%
+5
RL Systems Engineer: Scalable Post-Training Infrastructure
RL Systems Engineer: Scalable Post-Training Infrastructure

Luma • United Kingdom

On-site
GBP 150,000 - 190,000
ML/AI Engineer — Kubernetes, CI/CD & GPU Inference
ML/AI Engineer — Kubernetes, CI/CD & GPU Inference

Lloyds Banking Group • Manchester

Hybrid
GBP 73,000 - 81,000
Pension up to 15%
Annual bonus
Share schemes
+4