Senior AI Infrastructure Engineer - Scale Multi-GPU Training

LinuxRecruit

Greater London

On-site

GBP 90,000 - 120,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salary
Equity in early-stage startup
London lab

Job summary

LinuxRecruit is seeking an experienced AI Infrastructure or MLOps Engineer to join our London lab. You will architect and optimize distributed training across multiple GPUs and machines in AWS, eliminate bottlenecks in the data path, and manage cluster orchestration with Slurm and Kubernetes.

The role requires deep PyTorch expertise, familiarity with transformer models, and experience deploying production AI systems.

Qualifications

  • Proven track record building production AI systems.
  • Strong PyTorch experience and understanding of transformers.
  • Experience with distributed training and high-performance compute.

Responsibilities

  • Architect and optimize distributed training across multiple GPUs and machines.
  • Manage cluster orchestration using Slurm and Kubernetes; expand to specialised GPU providers.
  • Collaborate across teams in a fast-moving startup to shape technology decisions.

Skills

Distributed training design
Performance optimization
Collaboration & ownership
Problem solving under pressure

Tools

PyTorch
Slurm
Kubernetes
AWS

Job description

LinuxRecruit is seeking an experienced AI Infrastructure or MLOps Engineer to join our London lab. You will architect and optimize distributed training across multiple GPUs and machines in AWS, eliminate bottlenecks in the data path, and manage cluster orchestration with Slurm and Kubernetes.

The role requires deep PyTorch expertise, familiarity with transformer models, and experience deploying production AI systems.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infra & DS Engineer — London
Senior AI Infra & DS Engineer — London

LinuxRecruit • Greater London

On-site
GBP 100,000 - 140,000
AI infrastructure engineer
AI infrastructure engineer

LinuxRecruit • Greater London

On-site
GBP 90,000 - 120,000
Competitive salary
Equity in early-stage startup
London lab
Lead AI Infrastructure & Distributed Systems Engineer
Lead AI Infrastructure & Distributed Systems Engineer

LinuxRecruit • Greater London

On-site
GBP 100,000 - 140,000
MLOps Engineer: Scalable AI Infra & Deployment
MLOps Engineer: Scalable AI Infra & Deployment

Talenzon group • Greater London

On-site
GBP 70,000 - 90,000
ML Performance Engineer: Scale GPU/CPU ML Workloads
ML Performance Engineer: Scale GPU/CPU ML Workloads

G-Research • Greater London

Hybrid
GBP 90,000 - 150,000
Competitive pay
Lunch provided
Annual leave 35d
+5
Scale-Focused AI Infra & MLOps Engineer
Scale-Focused AI Infra & MLOps Engineer

EngineersOfAI • Greater London

On-site
GBP 90,000 - 120,000
MLOps & Infrastructure Engineer for AI HPC
MLOps & Infrastructure Engineer for AI HPC

EngineersOfAI • United Kingdom

On-site
GBP 40,000 - 70,000
Onsite MLOps Platform Engineer — AI Infra
Onsite MLOps Platform Engineer — AI Infra

DGH Recruitment • City Of London

On-site
GBP 85,000 - 115,000
MLOps Engineer – AI Infrastructure & Deployment
MLOps Engineer – AI Infrastructure & Deployment

Talenzon group • Greater London

On-site
GBP 70,000 - 90,000
AI Engineer - Hybrid, Generative AI & MLOps
AI Engineer - Hybrid, Generative AI & MLOps

Harnham - Data and Analytics Recruitment • Greater London

Hybrid
GBP 90,000 - 120,000
Hybrid working model
Professional development
Career progression