ML Performance Engineer: Scale GPU/CPU ML Workloads

G-Research

Greater London

Hybrid

GBP 90,000 - 150,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive pay
Lunch provided
Annual leave 35d
Pension 9%
Informal dress code
Healthcare & life assurance
Cycle-to-work
Monthly events

Job summary

G-Research in London seeks an exceptional ML Performance Engineer to optimise large-scale workloads across GPU and CPU infrastructure. You will profile and tune training and inference jobs, develop reference implementations, and collaborate with research and platform teams to evolve the compute stack.

The role requires strong Python, CUDA, and PyTorch knowledge, plus experience with HPC schedulers and Kubernetes.

Qualifications

  • Proven track record profiling, benchmarking and optimising distributed workloads.
  • Experience with Python.
  • Knowledge of CUDA.
  • Experience with HPC schedulers and Kubernetes-based workload orchestration.
  • Strong understanding of PyTorch and deep learning frameworks.
  • Strong background in data structures, algorithms and parallel programming.
  • Deep understanding of Linux OS fundamentals.
  • Familiarity with profiling/monitoring tools (nsys, ncu, eBPF).
  • Strong communication skills across teams.

Responsibilities

  • Collaborate with researchers and engineers to understand compute challenges and design optimised solutions.
  • Profile, benchmark and tune large-scale training and inference workloads on distributed CPU, GPU, and memory-intensive jobs.
  • Develop reference implementations, libraries and tools to improve job efficiency and reliability.
  • Collaborate with systems, architecture and platform teams to evolve the compute stack.
  • Influence long-term platform and infrastructure decisions.

Skills

Profiling workloads
Benchmarking
Optimising distributed workloads
Python
CUDA
HPC schedulers
Kubernetes
PyTorch
Data structures & algorithms
Parallel programming
Linux fundamentals
Profiling tools (nsys, ncu, eBPF)
Communication

Education

Bachelors/Masters/PhD in CS

Tools

Kubernetes

Job description

G-Research in London seeks an exceptional ML Performance Engineer to optimise large-scale workloads across GPU and CPU infrastructure. You will profile and tune training and inference jobs, develop reference implementations, and collaborate with research and platform teams to evolve the compute stack.

The role requires strong Python, CUDA, and PyTorch knowledge, plus experience with HPC schedulers and Kubernetes.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Performance Engineer – Scale GPU/CPU Workloads
ML Performance Engineer – Scale GPU/CPU Workloads

Barlowe LLP • Greater London

On-site
GBP 90,000 - 150,000
Lunch provided
35 days’ annual leave
9% company pension contributions
+4
ML Performance Engineer: Large-Scale GPU/CPU Optimization
ML Performance Engineer: Large-Scale GPU/CPU Optimization

gresearch • Greater London

On-site
GBP 90,000 - 130,000
Lunch provided
35 days annual leave
9% pension contributions
+3
Machine Learning Performance Engineer
Machine Learning Performance Engineer

G-Research • Greater London

Hybrid
GBP 90,000 - 150,000
Competitive pay
Lunch provided
Annual leave 35d
+5
Machine Learning Performance Engineer
Machine Learning Performance Engineer

gresearch • Greater London

On-site
GBP 90,000 - 130,000
Lunch provided
35 days annual leave
9% pension contributions
+3
Machine Learning Performance Engineer
Machine Learning Performance Engineer

Barlowe LLP • Greater London

On-site
GBP 90,000 - 150,000
Lunch provided
35 days’ annual leave
9% company pension contributions
+4
ML Infrastructure Engineer - Distributed Systems & HPC
ML Infrastructure Engineer - Distributed Systems & HPC

Stanford Black Limited • Greater London

On-site
GBP 100,000 - 150,000
Competitive compensation
Bonus structure
Autonomy from day one
+1
Senior AI Infrastructure Engineer - Scale Multi-GPU Training
Senior AI Infrastructure Engineer - Scale Multi-GPU Training

LinuxRecruit • Greater London

On-site
GBP 90,000 - 120,000
Competitive salary
Equity in early-stage startup
London lab
Remote Performance Engineer: ML Training & Kernels
Remote Performance Engineer: ML Training & Kernels

Cohere • Greater London

On-site
GBP 75,000 - 95,000
Co-working benefit
Daily lunch program
Regular community and social events
Machine Learning Specialist
Machine Learning Specialist

Stanford Black Limited • Greater London

On-site
GBP 100,000 - 150,000
Competitive compensation
Bonus structure
Autonomy from day one
+1
ML Performance Engineer: GPU & Systems Optimisation
ML Performance Engineer: GPU & Systems Optimisation

Trading Interview • Greater London

Hybrid
GBP 120,000 - 180,000