ML Infrastructure Engineer - Distributed Systems & HPC

Stanford Black Limited

Greater London

On-site

GBP 100,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation
Bonus structure
Autonomy from day one
Close collaboration with researchers

Job summary

Stanford Black Limited, London, is seeking an ML infrastructure engineer to design and optimize large-scale training and inference systems. You will work with researchers to scale ML workloads across a GPU estate, spanning software, hardware, networking, and compilers.

The role offers significant autonomy, collaboration with researchers, and a focus on performance across distributed systems. Strong Python/C++ engineering and experience with PyTorch/TensorFlow are expected, with opportunities to

Qualifications

  • Degree in CS/Math/Physics/Engineering or related field.
  • Strong software engineering in Python and C++.
  • Experience with PyTorch, JAX or TensorFlow.

Responsibilities

  • Design and optimise large-scale training and inference systems for ML workloads.
  • Improve throughput, latency, GPU utilisation and efficiency across distributed environments.
  • Build infrastructure and tooling to accelerate experimentation and model development.
  • Partner with researchers to productionise novel ML approaches.
  • Drive performance improvements across software, hardware and networking layers.
  • Influence the technical direction of ML infrastructure.

Skills

Python
C++
PyTorch/TensorFlow
Distributed Systems
Performance Optimisation
Research Engineering

Education

Degree in CS/Math/Physics/Engineering

Tools

CUDA
GPGPU Tools
Kubernetes

Job description

Stanford Black Limited, London, is seeking an ML infrastructure engineer to design and optimize large-scale training and inference systems. You will work with researchers to scale ML workloads across a GPU estate, spanning software, hardware, networking, and compilers.

The role offers significant autonomy, collaboration with researchers, and a focus on performance across distributed systems. Strong Python/C++ engineering and experience with PyTorch/TensorFlow are expected, with opportunities to

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Performance Engineer: Scale GPU/CPU ML Workloads
ML Performance Engineer: Scale GPU/CPU ML Workloads

G-Research • Greater London

Hybrid
GBP 90,000 - 150,000
Competitive pay
Lunch provided
Annual leave 35d
+5
ML Performance Engineer: Large-Scale GPU/CPU Optimization
ML Performance Engineer: Large-Scale GPU/CPU Optimization

gresearch • Greater London

On-site
GBP 90,000 - 130,000
Lunch provided
35 days annual leave
9% pension contributions
+3
ML Infrastructure Engineering Manager – Distributed Systems
ML Infrastructure Engineering Manager – Distributed Systems

eFinancialCareers • Greater London

On-site
GBP 120,000 - 180,000
ML Performance Engineer – Scale GPU/CPU Workloads
ML Performance Engineer – Scale GPU/CPU Workloads

Barlowe LLP • Greater London

On-site
GBP 90,000 - 150,000
Lunch provided
35 days’ annual leave
9% company pension contributions
+4
Senior AI Infrastructure Engineer - Scale Multi-GPU Training
Senior AI Infrastructure Engineer - Scale Multi-GPU Training

LinuxRecruit • Greater London

On-site
GBP 90,000 - 120,000
Competitive salary
Equity in early-stage startup
London lab
Machine Learning Specialist
Machine Learning Specialist

Stanford Black Limited • Greater London

On-site
GBP 100,000 - 150,000
Competitive compensation
Bonus structure
Autonomy from day one
+1
ML Infrastructure Engineer: Scalable LLM Serving Platform
ML Infrastructure Engineer: Scalable LLM Serving Platform

Neura Market • Greater London

On-site
GBP 110,000 - 165,000
Senior AI Infra & DS Engineer — London
Senior AI Infra & DS Engineer — London

LinuxRecruit • Greater London

On-site
GBP 100,000 - 140,000
Machine Learning Performance Engineer
Machine Learning Performance Engineer

G-Research • Greater London

Hybrid
GBP 90,000 - 150,000
Competitive pay
Lunch provided
Annual leave 35d
+5
Remote Performance Engineer: ML Training & Kernels
Remote Performance Engineer: ML Training & Kernels

Cohere • Greater London

On-site
GBP 75,000 - 95,000
Co-working benefit
Daily lunch program
Regular community and social events