Senior AI Training Infra Engineer - Distributed PyTorch

Genesis AI

Greater London

Hybrid

GBP 120,000 - 170,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Genesis AI seeks a senior ML infrastructure engineer to optimize distributed training pipelines and systems on multi-node GPU clusters in London. You will profile bottlenecks, implement low-level kernels, and integrate CUDA/cuDNN/Triton into high-level frameworks to maximize throughput and utilization.

The role requires deep expertise in distributed systems, Python, and PyTorch with a track record of hardware-aware optimizations across CPU/GPU boundaries and data movement.

Qualifications

  • 8+ years of experience in distributed systems, ML infrastructure, or HPC.
  • Production-grade Python experience.
  • Low-level performance mastery with CUDA/cuDNN/Triton and kernel optimization.
  • Experience scaling PyTorch training jobs with data, context, pipeline, and model parallelism.
  • System-level mindset tuning hardware–software interactions for max utilization.

Responsibilities

  • Profile and remove bottlenecks to reduce wall-clock convergence time.
  • Design and optimize distributed training systems for multi-node GPU clusters.
  • Implement low-level code (CUDA/cuDNN/Triton) and integrate into training frameworks.
  • Optimize workloads for CPU/GPU balance, memory, data throughput, and networking.
  • Develop monitoring and debugging tools for large-scale runs.

Skills

Distributed systems
Python
Performance optimization
PyTorch
Hardware-tuning

Tools

CUDA
cuDNN
Triton

Job description

Genesis AI seeks a senior ML infrastructure engineer to optimize distributed training pipelines and systems on multi-node GPU clusters in London. You will profile bottlenecks, implement low-level kernels, and integrate CUDA/cuDNN/Triton into high-level frameworks to maximize throughput and utilization.

The role requires deep expertise in distributed systems, Python, and PyTorch with a track record of hardware-aware optimizations across CPU/GPU boundaries and data movement.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead AI Training Infrastructure Engineer
Lead AI Training Infrastructure Engineer

Genesis • Greater London

Hybrid
GBP 90,000 - 130,000
Lead AI Training Infrastructure Engineer
Lead AI Training Infrastructure Engineer

Genesis • Greater London

Hybrid
GBP 90,000 - 130,000
Training / AI Infrastructure Engineering & Research London
Training / AI Infrastructure Engineering & Research London

Genesis • Greater London

Hybrid
GBP 90,000 - 130,000
Senior Data Infra Engineer for Large-Scale AI Pipelines
Senior Data Infra Engineer for Large-Scale AI Pipelines

Genesis • Greater London

Hybrid
GBP 90,000 - 130,000
Training / AI Infrastructure
Training / AI Infrastructure

Genesis AI • Greater London

Hybrid
GBP 120,000 - 170,000
Training / AI Infrastructure
Training / AI Infrastructure

Genesis AI • Greater London

Hybrid
GBP 120,000 - 170,000
Staff AI Systems Engineer - Pre-Training Infra
Staff AI Systems Engineer - Pre-Training Infra

Reflection • Greater London

On-site
GBP 70,000 - 100,000
Top-tier compensation
Comprehensive health insurance
Fully paid parental leave
+2
Senior Inference Architect — Low-Latency, On-Device & GPU
Senior Inference Architect — Low-Latency, On-Device & GPU

Genesis AI • Greater London

On-site
GBP 110,000 - 150,000
Senior Data Infrastructure Engineer - Petabyte-Scale AI Pipelines
Senior Data Infrastructure Engineer - Petabyte-Scale AI Pipelines

Genesis AI • Greater London

Hybrid
GBP 120,000 - 180,000
Senior AI Infra Engineer - ML Platform & GPU Systems
Senior AI Infra Engineer - ML Platform & GPU Systems

Salient Group • Greater London

Hybrid
GBP 120,000 - 180,000
Equity grant