Machine Learning Specialist

Stanford Black Limited

Greater London

On-site

GBP 100,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation
Bonus structure
Autonomy from day one
Close collaboration with researchers

Job summary

Stanford Black Limited, London, is seeking an ML infrastructure engineer to design and optimize large-scale training and inference systems. You will work with researchers to scale ML workloads across a GPU estate, spanning software, hardware, networking, and compilers.

The role offers significant autonomy, collaboration with researchers, and a focus on performance across distributed systems. Strong Python/C++ engineering and experience with PyTorch/TensorFlow are expected, with opportunities to

Qualifications

  • Degree in CS/Math/Physics/Engineering or related field.
  • Strong software engineering in Python and C++.
  • Experience with PyTorch, JAX or TensorFlow.

Responsibilities

  • Design and optimise large-scale training and inference systems for ML workloads.
  • Improve throughput, latency, GPU utilisation and efficiency across distributed environments.
  • Build infrastructure and tooling to accelerate experimentation and model development.
  • Partner with researchers to productionise novel ML approaches.
  • Drive performance improvements across software, hardware and networking layers.
  • Influence the technical direction of ML infrastructure.

Skills

Python
C++
PyTorch/TensorFlow
Distributed Systems
Performance Optimisation
Research Engineering

Education

Degree in CS/Math/Physics/Engineering

Tools

CUDA
GPGPU Tools
Kubernetes

Job description

  • We're partnering with a highly quantitative research organisation building some of the most advanced machine learning systems in industry.
  • Engineers in this team operate at the intersection of machine learning, distributed systems, and high-performance computing, helping scale modern AI workloads across a large GPU estate. The work spans distributed training, inference optimisation, compute infrastructure, systems design, and performance engineering.
  • You'll work directly with researchers to take cutting-edge ML ideas from prototype to production, solving problems that span software, hardware, networking, compilers, and large-scale distributed systems.
  • This is an opportunity to tackle technical challenges rarely seen outside leading AI labs and top-tier quantitative research firms.

Responsibilities

  • Design and optimise large-scale training and inference systems for modern ML workloads.
  • Improve throughput, latency, GPU utilisation and training efficiency across distributed environments.
  • Build infrastructure and tooling that accelerates experimentation and model development.
  • Partner with researchers to productionise novel ML approaches.
  • Drive performance improvements across software, hardware and networking layers.
  • Influence the technical direction of critical ML infrastructure used across the organisation.

What We're Looking For

  • Strong experience in Machine Learning Engineering, Research Engineering, ML Infrastructure, Distributed Systems or Performance Engineering.
  • Excellent software engineering skills in Python and/or C++.
  • Experience working with modern ML frameworks such as PyTorch, JAX or TensorFlow.
  • Experience training, deploying or optimising large-scale machine learning models.
  • Strong understanding of distributed systems, parallel computing and performance optimisation.
  • Degree in Computer Science, Mathematics, Physics, Engineering or a related quantitative discipline, or equivalent industry experience.

Particularly Relevant Experience

  • Large-scale distributed training (DeepSpeed, FSDP, Megatron, Ray, DDP or similar).
  • GPU programming and optimisation (CUDA, Triton, NCCL, XLA, PTX).
  • Multi-GPU or multi-node training environments.
  • HPC, Kubernetes, Slurm or large-scale compute infrastructure.
  • Foundation models, LLMs, recommendation systems or large-scale deep learning.
  • Compiler technologies, kernel optimisation, inference optimisation or systems-level ML performance work.

Why Join?

  • Work on some of the largest and most computationally intensive ML workloads in industry.
  • Solve challenging problems across distributed systems, GPU computing, machine learning infrastructure and performance optimisation.
  • Collaborate closely with exceptional researchers, engineers and quantitative scientists.
  • Significant autonomy and ownership from day one.
  • Deep investment in compute infrastructure and engineering excellence.
  • Competitive compensation and bonus structure.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Performance Engineer
Machine Learning Performance Engineer

G-Research • Greater London

Hybrid
GBP 90,000 - 150,000
Competitive pay
Lunch provided
Annual leave 35d
+5
Machine Learning Performance Engineer
Machine Learning Performance Engineer

gresearch • Greater London

On-site
GBP 90,000 - 130,000
Lunch provided
35 days annual leave
9% pension contributions
+3
Senior ML Infrastructure Engineer (Research Initiatives) - Systems Integrator
Senior ML Infrastructure Engineer (Research Initiatives) - Systems Integrator

Hamilton Barnes Associates Limited • United Kingdom

Hybrid
GBP 90,000 - 130,000
Significant stock option packages
Remote-first working setup
Fully paid travel and accommodation
+1
Machine Learning Performance Engineer
Machine Learning Performance Engineer

Barlowe LLP • Greater London

On-site
GBP 90,000 - 150,000
Lunch provided
35 days’ annual leave
9% company pension contributions
+4
Research Engineer, Machine Learning (AI for Science)
Research Engineer, Machine Learning (AI for Science)

Generative • Greater London

On-site
GBP 90,000 - 150,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Dimensionalconsultingltd • Manchester

Hybrid
GBP 70,000 - 90,000
Competitive salary with annual performance bonus
£2,000 annual conference and learning budget
Private healthcare and dental cover
+4
Machine Learning Engineer - Scaling
Machine Learning Engineer - Scaling

BioTalent • Greater London

On-site
GBP 60,000 - 80,000
Machine Learning Performance Engineer
Machine Learning Performance Engineer

Quant Blueprint LLC • Greater London

On-site
GBP 50,000 - 70,000
Machine Learning Performance Engineer
Machine Learning Performance Engineer

Janestreet • Greater London

On-site
GBP 111,000 - 140,000
Machine Learning Engineer
Machine Learning Engineer

G-Research • Greater London

On-site
GBP 85,000 - 130,000
Highly competitive compensation
Annual discretionary bonus
Lunch provided via Just Eat
+6