Inference Systems Engineer (GPU/Cluster Scaling)

AItoolnavio

Greater London

Hybrid

GBP 100,000 - 180,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Luma is seeking a systems engineer to own large-scale model serving and inference pipelines. You will integrate new architectures into the inference engine, scale deployments across thousands of machines, and keep GPU fleets busy while meeting internal SLOs.

You’ll work on scheduling, fleet management, and reliable deployment pipelines across clusters, with a strong emphasis on efficiency and uptime. This is the systems side of ML, not pure modeling.

Qualifications

  • Experience deploying models at scale.
  • Strong Python and system design skills.
  • Familiarity with Linux, Docker and Kubernetes.
  • Experience with distributed systems and model serving toolchains.

Responsibilities

  • Ship new model architectures by integrating them into the inference engine.
  • Collaborate across research, engineering, and infrastructure to optimize model efficiency and deployments.
  • Build internal tooling to measure, profile, and track the lifetime of inference jobs and workflows.
  • Automate, test, and maintain inference services for maximum uptime and reliability.
  • Manage and optimize inference workloads across clusters and hardware providers, and scale deployments across thousands of machines.
  • Build scheduling systems that use expensive GPU resources optimally while meeting SLOs, and maintain CI/CD for model checkpoints and SDKs.

Skills

Python
System architecture
Model deployment
PyTorch
Hugging Face
vLLM
SGLang
TensorRT-LLM
Queues
Scheduling
Fleet management
Linux
Docker
Kubernetes
Redis
S3 storage

Tools

Docker
Kubernetes
Redis
S3 storage

Job description

Luma is seeking a systems engineer to own large-scale model serving and inference pipelines. You will integrate new architectures into the inference engine, scale deployments across thousands of machines, and keep GPU fleets busy while meeting internal SLOs.

You’ll work on scheduling, fleet management, and reliable deployment pipelines across clusters, with a strong emphasis on efficiency and uptime. This is the systems side of ML, not pure modeling.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Inference Systems Engineer: Scale GPU Deployments
ML Inference Systems Engineer: Scale GPU Deployments

Luma • United Kingdom

On-site
GBP 90,000 - 150,000
Senior Software Engineer, Large-Scale Model Inference
Senior Software Engineer, Large-Scale Model Inference

Luma • Greater London

On-site
GBP 147,000 - 261,000
Senior Systems Engineer - Large-Scale AI Inference
Senior Systems Engineer - Large-Scale AI Inference

Speedrun Talent Network • Greater London

Hybrid
GBP 90,000 - 140,000
Software Engineer, Inference
Software Engineer, Inference

AItoolnavio • Greater London

Hybrid
GBP 100,000 - 180,000
Software Engineer, Inference
Software Engineer, Inference

Speedrun Talent Network • Greater London

Hybrid
GBP 90,000 - 140,000
Software Engineer, Inference
Software Engineer, Inference

Luma • Greater London

On-site
GBP 147,000 - 261,000
Distributed ML Infrastructure Engineer for Large-Scale Models
Distributed ML Infrastructure Engineer for Large-Scale Models

Lumaai • Greater London

On-site
GBP 120,000 - 180,000
Senior Distributed Training Engineer — Large-Scale GPU Systems
Senior Distributed Training Engineer — Large-Scale GPU Systems

Speedrun Talent Network • Greater London

Hybrid
GBP 120,000 - 180,000
Senior Distributed ML Systems Engineer
Senior Distributed ML Systems Engineer

Luma • Greater London

On-site
GBP 147,000 - 299,000
RL Infrastructure Engineer - Scale & Post-Training Systems
RL Infrastructure Engineer - Scale & Post-Training Systems

AItoolnavio • Greater London

Hybrid
GBP 120,000 - 180,000