Research Infrastructure Engineer

Axiōma Search

Greater London

Hybrid

GBP 110,000 - 140,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Stealth AI Lab in London (Hybrid) is seeking a Member of Technical Staff – Infrastructure to own and scale GPU-driven training platforms. You'll build and operate GPU clusters, manage distributed job scheduling, networking, and storage for ML workloads, and improve observability and failure recovery to empower researchers.

Collaborate with researchers to design reliable pipelines for datasets and model weights, and ensure reproducible experiments across large-scale environments.

Qualifications

  • Strong Linux and distributed systems fundamentals.
  • Experience with Docker, Kubernetes and/or Slurm.
  • Strong Python or Go.
  • Good understanding of concurrency, networking and failure recovery.
  • Familiarity with GPU networking and NCCL.
  • Familiarity with Ray, Terraform, Prometheus, Grafana or OpenTelemetry.

Responsibilities

  • Build and operate GPU clusters for model training and inference.
  • Manage job scheduling, networking and storage for distributed ML workloads.
  • Build systems that can run large numbers of training environments concurrently.
  • Improve GPU utilisation, startup times and overall platform reliability.
  • Build reliable pipelines for datasets, model weights and checkpoints.
  • Improve monitoring, debugging and failure recovery.
  • Give researchers simple, reproducible ways to run experiments.

Job description

Member of Technical Staff – Infrastructure

Stealth AI Lab | Paris or London (Hybrid)

About

Training models at scale creates a lot of infrastructure problems. Researchers need reliable access to GPUs, experiments need to be reproducible, and failures need to be easy to understand. You'll own the platform that makes that possible.

The company is building AI systems that learn how to carry out complex work inside large organisations. They recreate real-world workflows as interactive training environments, then use those environments to train models through practice and feedback — so the models get better at completing long, multi-step tasks reliably, rather than simply generating answers.

You'll work across GPU infrastructure, distributed execution, storage and observability. The aim is straightforward: make compute productive and give researchers simple tools to run, inspect and debug their work.

What you'll do

  • Build and operate GPU clusters for model training and inference
  • Manage job scheduling, networking and storage for distributed ML workloads
  • Build systems that can run large numbers of training environments concurrently
  • Improve GPU utilisation, startup times and overall platform reliability
  • Build reliable pipelines for datasets, model weights and checkpoints
  • Improve monitoring, debugging and failure recovery
  • Give researchers simple, reproducible ways to run experiments

What you'll need

  • Strong Linux and distributed systems fundamentals
  • Experience with Docker, Kubernetes and/or Slurm
  • Strong Python or Go
  • Good understanding of concurrency, networking and failure recovery
  • Familiarity with GPU networking and NCCL
  • Familiarity with Ray, Terraform, Prometheus, Grafana or OpenTelemetry

Shortlisted candidates will be contacted within 48 hours.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Research Infra Engineer — GPU, Reproducible Experiments
ML Research Infra Engineer — GPU, Reproducible Experiments

Axiōma Search • Greater London

Hybrid
GBP 110,000 - 140,000
Training / AI Infrastructure Engineering & Research London
Training / AI Infrastructure Engineering & Research London

Genesis • Greater London

Hybrid
GBP 90,000 - 130,000
Senior ML Infrastructure Engineer (Research Initiatives) - Systems Integrator
Senior ML Infrastructure Engineer (Research Initiatives) - Systems Integrator

Hamilton Barnes Associates Limited • United Kingdom

Hybrid
GBP 90,000 - 130,000
Significant stock option packages
Remote-first working setup
Fully paid travel and accommodation
+1
Research Engineer (Post-Training)
Research Engineer (Post-Training)

Axiōma Search • Greater London

Hybrid
GBP 90,000 - 140,000
Staff Software Engineer, Inference / Compute Infrastructure Engineering
Staff Software Engineer, Inference / Compute Infrastructure Engineering

Together AI • Greater London

Hybrid
GBP 100,000 - 160,000
Member of Technical Staff (Infrastructure Engineer, Compute Infrastructure)
Member of Technical Staff (Infrastructure Engineer, Compute Infrastructure)

Inherentlabs • Greater London

On-site
GBP 70,000 - 90,000
Good lunch and dinner
Collaborative work culture
No bureaucracy
Software Engineer, GPU Infrastructure- ChatGPT Engineering
Software Engineer, GPU Infrastructure- ChatGPT Engineering

OpenAI • Greater London

On-site
GBP 120,000 - 190,000
Machine Learning & Cloud Infra Engineer
Machine Learning & Cloud Infra Engineer

SpAItial AI • Greater London

On-site
GBP 60,000 - 85,000
Senior Platform Engineer (Product Initiatives) - Systems Integrator
Senior Platform Engineer (Product Initiatives) - Systems Integrator

Hamilton Barnes Associates Limited • United Kingdom

Remote
GBP 120,000 - 190,000
High-Upside Equity
Flexible remote setup
Work-Life Balance
+1
Member of Technical Staff (Infrastructure Engineer, Training and Inference Systems)
Member of Technical Staff (Infrastructure Engineer, Training and Inference Systems)

Inherentlabs • Greater London

On-site
GBP 70,000 - 90,000