Senior ML Systems Engineer: Training Frameworks & HPC

Visa Hunt

Greater London

Hybrid

GBP 110,000 - 170,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Weekly lunch stipend
Health and dental benefits
RRSP matching / Pension
Parental Leave top-up
Enrichment benefits (arts, fitness, or
Paid vacation 6 weeks
Travel budget to other offices
Home office stipend

Job summary

Cohere is seeking a senior engineer to help build, maintain and evolve the training framework powering frontier-scale language models. This role sits at the intersection of large-scale training, distributed systems, and HPC infrastructure.

You will design and maintain the core components that enable fast, reliable, and scalable model training — and build the tooling that connects research ideas to thousands of GPUs.

Qualifications

  • Experience in large-scale distributed training or HPC systems.
  • Familiarity with multi-node cluster orchestration (Slurm, Ray, Kubernetes).
  • Ability to debug performance issues across CUDA/NCCL, networking, IO and data pipelines.

Responsibilities

  • Build and own the training framework for large-scale LLM training.
  • Design distributed training abstractions (data/tensor/pipeline parallelism, memory management, checkpointing).
  • Improve training throughput and stability on multi-node clusters.
  • Develop tooling for monitoring, logging, debugging, and developer ergonomics.
  • Collaborate with infra teams to ensure clusters and hardware configurations support high-performance training.
  • Investigate and resolve performance bottlenecks across the ML systems stack.
  • Build robust systems for reproducible, debuggable large-scale runs.

Skills

Distributed training
JAX internals
Kubernetes/Ray
CUDA debugging
Containerization
Performance tuning
Collaboration

Tools

Slurm
Ray
Kubernetes
Docker
Apptainer

Job description

Cohere is seeking a senior engineer to help build, maintain and evolve the training framework powering frontier-scale language models. This role sits at the intersection of large-scale training, distributed systems, and HPC infrastructure.

You will design and maintain the core components that enable fast, reliable, and scalable model training — and build the tooling that connects research ideas to thousands of GPUs.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote Performance Engineer: ML Training & Kernels
Remote Performance Engineer: ML Training & Kernels

Cohere • Greater London

On-site
GBP 75,000 - 95,000
Co-working benefit
Daily lunch program
Regular community and social events
Staff Engineer – Training Infra & ML Systems
Staff Engineer – Training Infra & ML Systems

Cohere • Greater London

On-site
GBP 70,000 - 90,000
Open and inclusive culture
Weekly lunch stipend of $75/£75
Full health and dental benefits
+2
Senior ML Systems Engineer, Frameworks & Tooling
Senior ML Systems Engineer, Frameworks & Tooling

Visa Hunt • Greater London

Hybrid
GBP 110,000 - 170,000
Weekly lunch stipend
Health and dental benefits
RRSP matching / Pension
+5
ML Infrastructure Engineer - Distributed Systems & HPC
ML Infrastructure Engineer - Distributed Systems & HPC

Stanford Black Limited • Greater London

On-site
GBP 100,000 - 150,000
Competitive compensation
Bonus structure
Autonomy from day one
+1
Senior ML Systems Engineer - Simulations
Senior ML Systems Engineer - Simulations

Oriole • Greater London

On-site
GBP 90,000 - 140,000
ML Systems Performance Engineer
ML Systems Performance Engineer

Quant Blueprint LLC • Greater London

On-site
GBP 50,000 - 70,000
ML Performance Engineer: Scale GPU/CPU ML Workloads
ML Performance Engineer: Scale GPU/CPU ML Workloads

G-Research • Greater London

Hybrid
GBP 90,000 - 150,000
Competitive pay
Lunch provided
Annual leave 35d
+5
Machine Learning Specialist
Machine Learning Specialist

Stanford Black Limited • Greater London

On-site
GBP 100,000 - 150,000
Competitive compensation
Bonus structure
Autonomy from day one
+1
Senior Foundation Model Engineer: CUDA&Distributed Systems
Senior Foundation Model Engineer: CUDA&Distributed Systems

PulseRise Technologies • Greater London

On-site
GBP 120,000 - 180,000
Principal ML Infra Engineer – Large-Scale Physics Models
Principal ML Infra Engineer – Large-Scale Physics Models

Linuxconfig • Greater London

Hybrid
GBP 90,000 - 140,000
Equity options
Enhanced pension contribution
Private medical insurance
+1