Training / AI Infrastructure Engineering & Research London

Genesis

Greater London

Hybrid

GBP 90,000 - 130,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Genesis seeks a senior ML infrastructure engineer to design and optimize distributed training systems on multi-node GPU clusters. You will push wall-clock time to convergence by profiling, refining data pipelines, and low-level kernels in CUDA/CuDNN/Triton.

You will build scalable PyTorch-based workflows, balance CPU/GPU workloads, and develop monitoring tools to diagnose performance regressions at scale.

Qualifications

  • Deep experience in distributed systems, ML infrastructure, or high-performance computing (8+ years).
  • Production-grade expertise in Python.
  • Low-level performance mastery: CUDA/cuDNN/Triton, CPU-GPU interactions, data movement, and kernel optimization.
  • Scaling at the frontier: experience with PyTorch and training jobs using data, context, pipeline, and model parallelism.
  • System-level mindset with a track record of tuning hardware-software interactions for maximum utilization.

Responsibilities

  • Drive down wall-clock time to convergence by profiling and eliminating bottlenecks across the foundation model training stack, from data pipelines to GPU kernels.
  • Design, build, and optimize distributed training systems (PyTorch) for multi-node GPU clusters, ensuring scalability, robustness, and high utilization
  • Implement efficient low-level code (CUDA, cuDNN, Triton, custom kernels) and integrate it seamlessly into high-level training frameworks
  • Optimize workloads for hardware efficiency: CPU/GPU compute balance, memory management, data throughput, and networking
  • Develop monitoring and debugging tools for large-scale runs, enabling rapid diagnosis of performance regressions and failures

Skills

Distributed systems
Python
CUDA
cuDNN
Triton
CPU-GPU interactions
Kernel optimization
PyTorch
Performance tuning

Tools

PyTorch

Job description

What You'll Do
  • Drive down wall-clock time to convergence by profiling and eliminating bottlenecks across the foundation model training stack stack, from data pipelines to GPU kernels

  • Design, build, and optimize distributed training systems (PyTorch) for multi-node GPU clusters, ensuring scalability, robustness, and high utilization

  • Implement efficient low-level code (CUDA, cuDNN, Triton, custom kernels) and integrate it seamlessly into high-level training frameworks

  • Optimize workloads for hardware efficiency: CPU/GPU compute balance, memory management, data throughput, and networking

  • Develop monitoring and debugging tools for large-scale runs, enabling rapid diagnosis of performance regressions and failures

What You'll Bring
  • Deep experience in distributed systems, ML infrastructure, or high-performance computing (8+ years)

  • Production-grade expertise in Python

  • Low-level performance mastery: CUDA/cuDNN/Triton, CPU-GPU interactions, data movement, and kernel optimization

  • Scaling at the frontier: experience with PyTorch and training jobs using data, context, pipeline, and model parallelism

  • System-level mindset with a track record of tuning hardware-software interactions for maximum utilization

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Performance Engineer (GPU)
Performance Engineer (GPU)

Anthropic • York and North Yorkshire

On-site
GBP 90,000 - 140,000
Comprehensive health insurance
Fertility benefits
22 weeks parental leave
+1
Member of Technical Staff, ML Performance
Member of Technical Staff, ML Performance

Odyssey • Greater London

On-site
GBP 70,000 - 90,000
Lead AI Training Infrastructure Engineer
Lead AI Training Infrastructure Engineer

Genesis • Greater London

Hybrid
GBP 90,000 - 130,000
Machine Learning Performance Engineer
Machine Learning Performance Engineer

Quant Blueprint LLC • Greater London

On-site
GBP 50,000 - 70,000
Research Infrastructure Engineer
Research Infrastructure Engineer

Axiōma Search • Greater London

Hybrid
GBP 110,000 - 140,000
Senior ML Infrastructure Engineer (Research Initiatives) - Systems Integrator
Senior ML Infrastructure Engineer (Research Initiatives) - Systems Integrator

Hamilton Barnes Associates Limited • United Kingdom

Hybrid
GBP 90,000 - 130,000
Significant stock option packages
Remote-first working setup
Fully paid travel and accommodation
+1
Machine Learning Performance Engineer
Machine Learning Performance Engineer

Trading Interview • Greater London

Hybrid
GBP 120,000 - 180,000
Performance Engineer
Performance Engineer

Anthropic • York and North Yorkshire

On-site
GBP 110,000 - 150,000
Health insurance
Fertility benefits
Parental leave 22 weeks
+12
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

CVFine by Instrovate Technologies • Greater London

On-site
GBP 70,000 - 90,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Arcus Search • Greater London

On-site
GBP 80,000 - 120,000