Lead AI Training Infrastructure Engineer

Genesis

Greater London

Hybrid

GBP 90,000 - 130,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Genesis seeks a senior ML infrastructure engineer to design and optimize distributed training systems on multi-node GPU clusters. You will push wall-clock time to convergence by profiling, refining data pipelines, and low-level kernels in CUDA/CuDNN/Triton.

You will build scalable PyTorch-based workflows, balance CPU/GPU workloads, and develop monitoring tools to diagnose performance regressions at scale.

Qualifications

  • Deep experience in distributed systems, ML infrastructure, or high-performance computing (8+ years).
  • Production-grade expertise in Python.
  • Low-level performance mastery: CUDA/cuDNN/Triton, CPU-GPU interactions, data movement, and kernel optimization.
  • Scaling at the frontier: experience with PyTorch and training jobs using data, context, pipeline, and model parallelism.
  • System-level mindset with a track record of tuning hardware-software interactions for maximum utilization.

Responsibilities

  • Drive down wall-clock time to convergence by profiling and eliminating bottlenecks across the foundation model training stack, from data pipelines to GPU kernels.
  • Design, build, and optimize distributed training systems (PyTorch) for multi-node GPU clusters, ensuring scalability, robustness, and high utilization
  • Implement efficient low-level code (CUDA, cuDNN, Triton, custom kernels) and integrate it seamlessly into high-level training frameworks
  • Optimize workloads for hardware efficiency: CPU/GPU compute balance, memory management, data throughput, and networking
  • Develop monitoring and debugging tools for large-scale runs, enabling rapid diagnosis of performance regressions and failures

Skills

Distributed systems
Python
CUDA
cuDNN
Triton
CPU-GPU interactions
Kernel optimization
PyTorch
Performance tuning

Tools

PyTorch

Job description

Genesis seeks a senior ML infrastructure engineer to design and optimize distributed training systems on multi-node GPU clusters. You will push wall-clock time to convergence by profiling, refining data pipelines, and low-level kernels in CUDA/CuDNN/Triton.

You will build scalable PyTorch-based workflows, balance CPU/GPU workloads, and develop monitoring tools to diagnose performance regressions at scale.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Training / AI Infrastructure Engineering & Research London
Training / AI Infrastructure Engineering & Research London

Genesis • Greater London

Hybrid
GBP 90,000 - 130,000
Senior Data Infra Engineer for Large-Scale AI Pipelines
Senior Data Infra Engineer for Large-Scale AI Pipelines

Genesis • Greater London

Hybrid
GBP 90,000 - 130,000
Senior ML Infra Engineer: Scale GPU Clusters & Pipelines
Senior ML Infra Engineer: Scale GPU Clusters & Pipelines

Ellison Institute, LLC • Oxford

Hybrid
GBP 90,000 - 140,000
Competitive salary
25 days annual leave + 8 bank holidays
3 additional days between Christmas &.
+11
AI Research Platform Engineer
AI Research Platform Engineer

Genesis • Greater London

Hybrid
GBP 90,000 - 120,000
GPU Performance Engineer: Scale Inference & Training
GPU Performance Engineer: Scale Inference & Training

Anthropic • York and North Yorkshire

On-site
GBP 90,000 - 140,000
Comprehensive health insurance
Fertility benefits
22 weeks parental leave
+1
ML Performance Engineer – Scale GPU/CPU Workloads
ML Performance Engineer – Scale GPU/CPU Workloads

Barlowe LLP • Greater London

On-site
GBP 90,000 - 150,000
Lunch provided
35 days’ annual leave
9% company pension contributions
+4
Lead GPU Infrastructure Architect for Scalable AI Clusters
Lead GPU Infrastructure Architect for Scalable AI Clusters

Hamilton Barnes Associates Limited • Greater London

On-site
GBP 110,000 - 150,000
Senior AI Compute Kernels Engineer — High-Performance
Senior AI Compute Kernels Engineer — High-Performance

Graphcore • West of England

Hybrid
GBP 110,000 - 170,000
Work-life balance
Private Medical Insurance
Pension
+3
Senior Compute Infrastructure Engineer – AI Training & LLMs
Senior Compute Infrastructure Engineer – AI Training & LLMs

Inherentlabs • Greater London

On-site
GBP 70,000 - 90,000
Good lunch and dinner
Collaborative work culture
No bureaucracy
AI Infrastructure Lead: GPUs, HPC & Distributed Systems
AI Infrastructure Lead: GPUs, HPC & Distributed Systems

JD.com International Limited • Greater London

On-site
GBP 90,000 - 130,000