Senior Software Engineer, Large-Scale Model Inference

Luma

Greater London

On-site

GBP 147,000 - 261,000

Full time

17 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Luma seeks a Senior Systems Engineer to own how its models are served, integrating new architectures into the inference engine and scaling deployments across thousands of machines.

This is large-scale inference systems work, including scheduling, fleet management, deployment pipelines, and reliability across clusters and hardware providers. You’ll work across research, engineering, and infrastructure to optimize efficiency.

Qualifications

  • Strong Python and system-architecture skills.
  • Experience deploying models with PyTorch, Hugging Face, vLLM, SGLang, or TensorRT-LLM.
  • Experience with queues, scheduling, traffic control, and fleet management at scale.
  • Experience with Linux, Docker, and Kubernetes, and with orchestration, deployment, and scheduling.
  • Familiarity with Redis and S3-compatible storage.

Responsibilities

  • Ship new model architectures by integrating them into the inference engine.
  • Collaborate across research, engineering, and infrastructure to optimize model efficiency and deployments.
  • Build internal tooling to measure, profile, and track the lifetime of inference jobs and workflows.
  • Automate, test, and maintain inference services for maximum uptime and reliability.
  • Manage and optimize inference workloads across clusters and hardware providers, and scale deployments across thousands of machines.
  • Build scheduling systems that use expensive GPU resources optimally while meeting SLOs, and maintain CI/CD for model checkpoints and SDKs.

Skills

Python
System architecture
Model serving at scale
distributed systems

Tools

Docker
Kubernetes
Redis
S3-compatible storage

Job description

Luma seeks a Senior Systems Engineer to own how its models are served, integrating new architectures into the inference engine and scaling deployments across thousands of machines.

This is large-scale inference systems work, including scheduling, fleet management, deployment pipelines, and reliability across clusters and hardware providers. You’ll work across research, engineering, and infrastructure to optimize efficiency.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Systems Engineer - Large-Scale AI Inference
Senior Systems Engineer - Large-Scale AI Inference

Speedrun Talent Network • Greater London

Hybrid
GBP 90,000 - 140,000
Inference Systems Engineer (GPU/Cluster Scaling)
Inference Systems Engineer (GPU/Cluster Scaling)

AItoolnavio • Greater London

Hybrid
GBP 100,000 - 180,000
ML Inference Systems Engineer: Scale GPU Deployments
ML Inference Systems Engineer: Scale GPU Deployments

Luma • United Kingdom

On-site
GBP 90,000 - 150,000
Software Engineer, Inference
Software Engineer, Inference

AItoolnavio • Greater London

Hybrid
GBP 100,000 - 180,000
Software Engineer, Inference
Software Engineer, Inference

Speedrun Talent Network • Greater London

Hybrid
GBP 90,000 - 140,000
Software Engineer, Inference
Software Engineer, Inference

Luma • Greater London

On-site
GBP 147,000 - 261,000
Distributed ML Infrastructure Engineer for Large-Scale Models
Distributed ML Infrastructure Engineer for Large-Scale Models

Lumaai • Greater London

On-site
GBP 120,000 - 180,000
Senior Distributed Training Engineer — Large-Scale GPU Systems
Senior Distributed Training Engineer — Large-Scale GPU Systems

Speedrun Talent Network • Greater London

Hybrid
GBP 120,000 - 180,000
Senior Distributed ML Systems Engineer
Senior Distributed ML Systems Engineer

Luma • Greater London

On-site
GBP 147,000 - 299,000
RL Infrastructure Engineer - Scale & Post-Training Systems
RL Infrastructure Engineer - Scale & Post-Training Systems

AItoolnavio • Greater London

Hybrid
GBP 120,000 - 180,000