Training / AI Infrastructure

Genesis AI

Greater London

Hybrid

GBP 120,000 - 170,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Genesis AI seeks a senior ML infrastructure engineer to optimize distributed training pipelines and systems on multi-node GPU clusters in London. You will profile bottlenecks, implement low-level kernels, and integrate CUDA/cuDNN/Triton into high-level frameworks to maximize throughput and utilization.

The role requires deep expertise in distributed systems, Python, and PyTorch with a track record of hardware-aware optimizations across CPU/GPU boundaries and data movement.

Qualifications

  • 8+ years of experience in distributed systems, ML infrastructure, or HPC.
  • Production-grade Python experience.
  • Low-level performance mastery with CUDA/cuDNN/Triton and kernel optimization.
  • Experience scaling PyTorch training jobs with data, context, pipeline, and model parallelism.
  • System-level mindset tuning hardware–software interactions for max utilization.

Responsibilities

  • Profile and remove bottlenecks to reduce wall-clock convergence time.
  • Design and optimize distributed training systems for multi-node GPU clusters.
  • Implement low-level code (CUDA/cuDNN/Triton) and integrate into training frameworks.
  • Optimize workloads for CPU/GPU balance, memory, data throughput, and networking.
  • Develop monitoring and debugging tools for large-scale runs.

Skills

Distributed systems
Python
Performance optimization
PyTorch
Hardware-tuning

Tools

CUDA
cuDNN
Triton

Job description

What You'll Do
  • Drive down wall-clock time to convergence by profiling and eliminating bottlenecks across the foundation model training stack stack, from data pipelines to GPU kernels

  • Design, build, and optimize distributed training systems (PyTorch) for multi-node GPU clusters, ensuring scalability, robustness, and high utilization

  • Implement efficient low-level code (CUDA, cuDNN, Triton, custom kernels) and integrate it seamlessly into high-level training frameworks

  • Optimize workloads for hardware efficiency: CPU/GPU compute balance, memory management, data throughput, and networking

  • Develop monitoring and debugging tools for large-scale runs, enabling rapid diagnosis of performance regressions and failures

What You'll Bring
  • Deep experience in distributed systems, ML infrastructure, or high-performance computing (8+ years)

  • Production-grade expertise in Python

  • Low-level performance mastery: CUDA/cuDNN/Triton, CPU-GPU interactions, data movement, and kernel optimization

  • Scaling at the frontier: experience with PyTorch and training jobs using data, context, pipeline, and model parallelism

  • System-level mindset with a track record of tuning hardware-software interactions for maximum utilization

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Training / AI Infrastructure Engineering & Research London
Training / AI Infrastructure Engineering & Research London

Genesis • Greater London

Hybrid
GBP 90,000 - 130,000
Inference
Inference

Genesis AI • Greater London

On-site
GBP 110,000 - 150,000
Member of Technical Staff, ML Performance
Member of Technical Staff, ML Performance

Odyssey • Greater London

On-site
GBP 70,000 - 90,000
Performance Engineer (GPU)
Performance Engineer (GPU)

Anthropic • York and North Yorkshire

On-site
GBP 90,000 - 140,000
Comprehensive health insurance
Fertility benefits
22 weeks parental leave
+1
Principal ML Performance Engineer (GPU Optimization)
Principal ML Performance Engineer (GPU Optimization)

Sponsor Finder • Boston

On-site
GBP 120,000 - 180,000
Lead AI Training Infrastructure Engineer
Lead AI Training Infrastructure Engineer

Genesis • Greater London

Hybrid
GBP 90,000 - 130,000
Senior AI Training Infra Engineer - Distributed PyTorch
Senior AI Training Infra Engineer - Distributed PyTorch

Genesis AI • Greater London

Hybrid
GBP 120,000 - 170,000
Founding AI Infrastructure Engineer
Founding AI Infrastructure Engineer

Imperial College London • Greater London

On-site
GBP 90,000 - 130,000
Equity
Founding-team influence
Small team culture
Data Infrastructure
Data Infrastructure

Genesis • Greater London

On-site
GBP 110,000 - 150,000
Research Infrastructure Engineer
Research Infrastructure Engineer

Axiōma Search • Greater London

Hybrid
GBP 110,000 - 140,000