Senior ML Systems Engineer - Large-Scale Training & HPC

Cohere

Greater London

Remote

GBP 90,000 - 150,000

Full time

12 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Equity / stock options
Six weeks' paid vacation
Remote-friendly culture
Wellness allowance
Education benefit

Job summary

Cohere is hiring a senior engineer to build and evolve the training framework powering frontier-scale language models. You will own core components enabling fast, reliable multi-node training and create tooling that accelerates research onto thousands of GPUs.

You’ll work across ML systems, collaborate with infra teams, and own end-to-end training stack components. Expect high impact on performance, scalability, and research velocity in a remote-capable environment.

Qualifications

  • Experience with multi-node cluster orchestration (Slurm, Ray, Kubernetes, or similar).
  • Experience with training LLMs or other large transformer architectures.
  • Deep familiarity with JAX internals, distributed training libraries, or custom kernels/fused ops.
  • Strong engineering experience in large-scale distributed training or HPC systems.
  • Comfort debugging performance issues across CUDA/NCCL, networking, IO, and data pipelines.
  • Excellent judgment around trade-offs: performance vs complexity, velocity vs maintainability.
  • Experience with data pipeline optimization, sharded datasets, or caching strategies.
  • Contributions to ML frameworks (PyTorch, JAX, DeepSpeed, Megatron, xFormers, etc.).
  • Experience working with containerized environments (Docker, Singularity/Apptainer).
  • Background in performance engineering, profiling, or low-level systems.

Responsibilities

  • Build, maintain and evolve the training framework for frontier-scale language models.
  • Design distributed training abstractions (data/tensor/pipeline parallelism, FSDP/ZeRO strategies).
  • Improve training throughput and stability on multi-node clusters.
  • Develop tooling for monitoring, logging, debugging, and developer ergonomics.
  • Collaborate with infra to ensure cluster, containers, and hardware support high-performance training.
  • Investigate and resolve ML systems bottlenecks across the stack.
  • Build robust, reproducible, debuggable large-scale runs.
  • Own critical components of the training stack and drive infrastructure evolution.
  • Sample projects include high-performance data loading, caching, profiling, metrics, and regression testing.

Skills

Multi-node orchestration
LLM training experience
Distributed training
CUDA/NCCL debugging
Performance profiling
JAX internals
Containerized environments
PyTorch/JAX/DeepSpeed experience
Data pipeline optimization
Sharded datasets / caching
Fused kernels
Team collaboration

Tools

Slurm
Ray
Kubernetes
Docker
Singularity/Apptainer

Job description

Cohere is hiring a senior engineer to build and evolve the training framework powering frontier-scale language models. You will own core components enabling fast, reliable multi-node training and create tooling that accelerates research onto thousands of GPUs.

You’ll work across ML systems, collaborate with infra teams, and own end-to-end training stack components. Expect high impact on performance, scalability, and research velocity in a remote-capable environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Systems Engineer: Training Frameworks & HPC
Senior ML Systems Engineer: Training Frameworks & HPC

Visa Hunt • Greater London

Hybrid
GBP 110,000 - 170,000
Weekly lunch stipend
Health and dental benefits
RRSP matching / Pension
+5
Remote Performance Engineer: ML Training & Kernels
Remote Performance Engineer: ML Training & Kernels

Cohere • Greater London

On-site
GBP 75,000 - 95,000
Co-working benefit
Daily lunch program
Regular community and social events
Senior Machine Learning Systems Engineer (Frameworks & Tooling)
Senior Machine Learning Systems Engineer (Frameworks & Tooling)

Cohere • Greater London

Remote
GBP 90,000 - 150,000
Equity / stock options
Six weeks' paid vacation
Remote-friendly culture
+2
Staff Engineer – Training Infra & ML Systems
Staff Engineer – Training Infra & ML Systems

Cohere • Greater London

On-site
GBP 70,000 - 90,000
Open and inclusive culture
Weekly lunch stipend of $75/£75
Full health and dental benefits
+2
Senior ML Systems Engineer, Frameworks & Tooling
Senior ML Systems Engineer, Frameworks & Tooling

Visa Hunt • Greater London

Hybrid
GBP 110,000 - 170,000
Weekly lunch stipend
Health and dental benefits
RRSP matching / Pension
+5
Senior ML Engineer — Foundation Models in Production
Senior ML Engineer — Foundation Models in Production

Pulserise Technologies Ltd • Greater London

On-site
GBP 120,000 - 180,000
Principal ML Infra Engineer – Large-Scale Physics Models
Principal ML Infra Engineer – Large-Scale Physics Models

Linuxconfig • Greater London

Hybrid
GBP 90,000 - 140,000
Equity options
Enhanced pension contribution
Private medical insurance
+1
Senior ML Platform Engineer - Scalable AI Deployment
Senior ML Platform Engineer - Scalable AI Deployment

Scale AI, Inc. • Greater London

On-site
GBP 100,000 - 150,000
Remote Staff ML Efficiency Engineer — Scale & Optimize
Remote Staff ML Efficiency Engineer — Scale & Optimize

Reddit, Inc. • Greater London

On-site
GBP 110,000 - 160,000
Global Benefits
Family Planning
Mental Health Support
+4
Applied ML Staff Lead for Enterprise AI Solutions
Applied ML Staff Lead for Enterprise AI Solutions

Visa Hunt • Greater London

Hybrid
GBP 140,000 - 190,000
Weekly lunch stipend
Health and dental benefits
RRSP matching
+6