Senior Machine Learning Systems Engineer (Frameworks & Tooling)

Cohere

Greater London

Remote

GBP 90,000 - 150,000

Full time

12 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Equity / stock options
Six weeks' paid vacation
Remote-friendly culture
Wellness allowance
Education benefit

Job summary

Cohere is hiring a senior engineer to build and evolve the training framework powering frontier-scale language models. You will own core components enabling fast, reliable multi-node training and create tooling that accelerates research onto thousands of GPUs.

You’ll work across ML systems, collaborate with infra teams, and own end-to-end training stack components. Expect high impact on performance, scalability, and research velocity in a remote-capable environment.

Qualifications

  • Experience with multi-node cluster orchestration (Slurm, Ray, Kubernetes, or similar).
  • Experience with training LLMs or other large transformer architectures.
  • Deep familiarity with JAX internals, distributed training libraries, or custom kernels/fused ops.
  • Strong engineering experience in large-scale distributed training or HPC systems.
  • Comfort debugging performance issues across CUDA/NCCL, networking, IO, and data pipelines.
  • Excellent judgment around trade-offs: performance vs complexity, velocity vs maintainability.
  • Experience with data pipeline optimization, sharded datasets, or caching strategies.
  • Contributions to ML frameworks (PyTorch, JAX, DeepSpeed, Megatron, xFormers, etc.).
  • Experience working with containerized environments (Docker, Singularity/Apptainer).
  • Background in performance engineering, profiling, or low-level systems.

Responsibilities

  • Build, maintain and evolve the training framework for frontier-scale language models.
  • Design distributed training abstractions (data/tensor/pipeline parallelism, FSDP/ZeRO strategies).
  • Improve training throughput and stability on multi-node clusters.
  • Develop tooling for monitoring, logging, debugging, and developer ergonomics.
  • Collaborate with infra to ensure cluster, containers, and hardware support high-performance training.
  • Investigate and resolve ML systems bottlenecks across the stack.
  • Build robust, reproducible, debuggable large-scale runs.
  • Own critical components of the training stack and drive infrastructure evolution.
  • Sample projects include high-performance data loading, caching, profiling, metrics, and regression testing.

Skills

Multi-node orchestration
LLM training experience
Distributed training
CUDA/NCCL debugging
Performance profiling
JAX internals
Containerized environments
PyTorch/JAX/DeepSpeed experience
Data pipeline optimization
Sharded datasets / caching
Fused kernels
Team collaboration

Tools

Slurm
Ray
Kubernetes
Docker
Singularity/Apptainer

Job description

  • We're looking for a senior engineer to help build, maintain and evolve the training framework that powers our frontier-scale language models. This role sits at the intersection of large-scale training, distributed systems, and HPC infrastructure
  • You will design and maintain the core components that enable fast, reliable, and scalable model training - and build the tooling that connects research ideas to thousands of GPUs
  • If you enjoy working across the full stack of ML systems, this role gives you the opportunity and autonomy to have massive impact
  • Build and own the training framework responsible for large-scale LLM training
  • Design distributed training abstractions (data/tensor/pipeline parallelism, FSDP/ZeRO strategies, memory management, checkpointing)
  • Improve training throughput and stability on multi-node clusters (e.g., GB200/300, AMD, H200/100)
  • Develop and maintain tooling for monitoring, logging, debugging, and developer ergonomics
  • Collaborate closely with infra teams to ensure our cluster, container environments, and hardware configurations support high-performance training
  • Investigate and resolve performance bottlenecks across the ML systems stack
  • Build robust systems that ensure reproducible, debuggable, large-scale runs
  • You'll work on some of the most challenging and consequential ML systems problems today
  • You'll collaborate with a world-class team working fast and at scale
  • You'll have end-to-end ownership over critical components of the training stack
  • You'll shape the next generation of infrastructure for frontier-scale models
  • You'll build tools and systems that directly accelerate research and model quality
  • Sample Projects:
  • Build a high-performance data loading and caching pipeline
  • Implement performance profiling across the ML systems stack
  • Develop internal metrics and monitoring for training runs
  • Build reproducibility and regression testing infrastructure
  • Develop a performant fault-tolerant distributed checkpointing system
Benefits
  • Six weeks' paid vacation
  • Equity / stock options
  • RRSP, 401(k), and Pension Scheme contributions
  • Coverage for 100% of your insurance premiums across health, dental, vision, and travel
  • Additional coverage for accessing mental health providers/services
  • Six months of fully paid parental leave, including adoption and surrogacy
  • Financial support for egg freezing and IVF in Canada and the UK
  • A monthly fitness and wellness allowance
  • Globally dispersed company that supports a remote work culture
  • A $2,000 annual education benefit for professional development
  • A weekly stipend for meals when working remotely and catered lunch when working from one of our global offices
  • A monthly arts and culture allowance
  • A monthly quality time allowance
  • A track record of building tools that increase developer velocity for ML teamsExperience with multi-node cluster orchestration (Slurm, Ray, Kubernetes, or similar)Bonus: paper at top-tier venues (such as NeurIPS, ICML, ICLR, AIStats, MLSys, JAX, AAAI, Nature, COLING, ACL, EMNLP)Experience with training LLMs or other large transformer architecturesDeep familiarity with JAX internals, distributed training libraries, or custom kernels/fused opsStrong engineering experience in large-scale distributed training or HPC systemsComfort debugging performance issues across CUDA/NCCL, networking, IO, and data pipelinesExcellent judgment around trade-offs: performance vs complexity, research velocity vs maintainabilityExperience with data pipeline optimization, sharded datasets, or caching strategiesContributions to ML frameworks (PyTorch, JAX, DeepSpeed, Megatron, xFormers, etc.)Experience working with containerized environments (Docker, Singularity/Apptainer)Background in performance engineering, profiling, or low-level systemsIf some of the above doesn't line up perfectly with your experience, we still encourage you to apply!Strong collaboration skills - you'll work closely with infra, research, and deployment teamsFamiliarity with evaluation and serving frameworks (vLLM, TensorRT-LLM, custom KV caches)
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Systems Engineer, Frameworks & Tooling
Senior ML Systems Engineer, Frameworks & Tooling

Visa Hunt • Greater London

Hybrid
GBP 110,000 - 170,000
Weekly lunch stipend
Health and dental benefits
RRSP matching / Pension
+5
Principal Machine Learning Engineer – Production Systems
Principal Machine Learning Engineer – Production Systems

SoftInWay Inc • Bristol

On-site
GBP 70,000 - 100,000
Principal Machine Learning Infrastructure Engineer London, United Kingdom
Principal Machine Learning Infrastructure Engineer London, United Kingdom

PhysicsX Ltd • Greater London

On-site
GBP 80,000 - 100,000
Equity options
10% employer pension contribution
Free office lunches
+6
Principal Machine Learning Infrastructure Engineer
Principal Machine Learning Infrastructure Engineer

PhysicsX • City Of London

On-site
GBP 80,000 - 120,000
Equity options
10% employer pension contribution
Free office lunches
+2
Senior ML Infrastructure Engineer (Research Initiatives) - Systems Integrator
Senior ML Infrastructure Engineer (Research Initiatives) - Systems Integrator

Hamilton Barnes Associates Limited • United Kingdom

Hybrid
GBP 90,000 - 130,000
Significant stock option packages
Remote-first working setup
Fully paid travel and accommodation
+1
Senior ML Systems Engineer: Training Frameworks & HPC
Senior ML Systems Engineer: Training Frameworks & HPC

Visa Hunt • Greater London

Hybrid
GBP 110,000 - 170,000
Weekly lunch stipend
Health and dental benefits
RRSP matching / Pension
+5
Member of Technical Staff, ML Performance
Member of Technical Staff, ML Performance

Odyssey • Greater London

On-site
GBP 70,000 - 90,000
Senior ML Systems Engineer - Large-Scale Training & HPC
Senior ML Systems Engineer - Large-Scale Training & HPC

Cohere • Greater London

Remote
GBP 90,000 - 150,000
Equity / stock options
Six weeks' paid vacation
Remote-friendly culture
+2
Member of Technical Staff (AI Infrastructure Engineer)
Member of Technical Staff (AI Infrastructure Engineer)

Perplexity • Greater London

On-site
GBP 65,000 - 90,000
Machine Learning & Cloud Infra Engineer
Machine Learning & Cloud Infra Engineer

SpAItial AI • Greater London

On-site
GBP 60,000 - 85,000