ML Research Infra Engineer — GPU, Reproducible Experiments

Axiōma Search

Greater London

Hybrid

GBP 110,000 - 140,000

Full time

2 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Stealth AI Lab in London (Hybrid) is seeking a Member of Technical Staff – Infrastructure to own and scale GPU-driven training platforms. You'll build and operate GPU clusters, manage distributed job scheduling, networking, and storage for ML workloads, and improve observability and failure recovery to empower researchers.

Collaborate with researchers to design reliable pipelines for datasets and model weights, and ensure reproducible experiments across large-scale environments.

Qualifications

  • Strong Linux and distributed systems fundamentals.
  • Experience with Docker, Kubernetes and/or Slurm.
  • Strong Python or Go.
  • Good understanding of concurrency, networking and failure recovery.
  • Familiarity with GPU networking and NCCL.
  • Familiarity with Ray, Terraform, Prometheus, Grafana or OpenTelemetry.

Responsibilities

  • Build and operate GPU clusters for model training and inference.
  • Manage job scheduling, networking and storage for distributed ML workloads.
  • Build systems that can run large numbers of training environments concurrently.
  • Improve GPU utilisation, startup times and overall platform reliability.
  • Build reliable pipelines for datasets, model weights and checkpoints.
  • Improve monitoring, debugging and failure recovery.
  • Give researchers simple, reproducible ways to run experiments.

Job description

Stealth AI Lab in London (Hybrid) is seeking a Member of Technical Staff – Infrastructure to own and scale GPU-driven training platforms. You'll build and operate GPU clusters, manage distributed job scheduling, networking, and storage for ML workloads, and improve observability and failure recovery to empower researchers.

Collaborate with researchers to design reliable pipelines for datasets and model weights, and ensure reproducible experiments across large-scale environments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Infrastructure Engineer
Research Infrastructure Engineer

Axiōma Search • Greater London

Hybrid
GBP 110,000 - 140,000
Senior ML Infra Engineer: Scale GPU Clusters & Pipelines
Senior ML Infra Engineer: Scale GPU Clusters & Pipelines

Ellison Institute, LLC • Oxford

Hybrid
GBP 90,000 - 140,000
Competitive salary
25 days annual leave + 8 bank holidays
3 additional days between Christmas &.
+11
Senior Compute Infrastructure Engineer – AI Training & LLMs
Senior Compute Infrastructure Engineer – AI Training & LLMs

Inherentlabs • Greater London

On-site
GBP 70,000 - 90,000
Good lunch and dinner
Collaborative work culture
No bureaucracy
ML Performance Engineer – Scale GPU/CPU Workloads
ML Performance Engineer – Scale GPU/CPU Workloads

Barlowe LLP • Greater London

On-site
GBP 90,000 - 150,000
Lunch provided
35 days’ annual leave
9% company pension contributions
+4
Lead ML Infra Engineer for Scalable Physics AI
Lead ML Infra Engineer for Scalable Physics AI

PhysicsX • City Of London

On-site
GBP 80,000 - 120,000
Senior ML Infra Architect for Large-Scale AI Simulations
Senior ML Infra Architect for Large-Scale AI Simulations

Hamilton Barnes Associates Limited • United Kingdom

Hybrid
GBP 90,000 - 130,000
Significant stock option packages
Remote-first working setup
Fully paid travel and accommodation
+1
Machine Learning & Cloud Infra Engineer
Machine Learning & Cloud Infra Engineer

SpAItial AI • Greater London

On-site
GBP 60,000 - 85,000
ML Research Engineer: Generative Models & Scalable Training
ML Research Engineer: Generative Models & Scalable Training

Harnham • London

On-site
GBP 40,000 - 50,000
Staff AI Systems Engineer - Pre-Training Infra
Staff AI Systems Engineer - Pre-Training Infra

Reflection • Greater London

On-site
GBP 70,000 - 100,000
Senior ML Compute Architect: Scalable GPU Pipelines
Senior ML Compute Architect: Scalable GPU Pipelines

Ellison Institute of Technology • Oxford

On-site
GBP 90,000 - 130,000
Travel allowance
Pension 7.5%
Private Medical Insurance
+2