Senior ML Systems Engineer: Training Frameworks & Tooling

Cohere

United States

Remote

USD 150,000 - 230,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Cohere is seeking a senior engineer to help build, maintain, and evolve the training framework that powers frontier-scale language models. This role sits at the intersection of large-scale training, distributed systems, and HPC infrastructure.

You will design and maintain core components that enable fast, reliable model training, and build tooling that connects research ideas to thousands of GPUs. You’ll work across the full ML systems stack with autonomy and impact.

Qualifications

  • Strong engineering experience in large-scale distributed training or HPC systems.
  • Familiarity with JAX internals, distributed training libraries, or custom kernels/fused ops.
  • Experience with multi-node cluster orchestration (Slurm, Ray, Kubernetes, or similar).
  • Comfort debugging performance issues across CUDA/NCCL, networking, IO, and data pipelines.
  • Experience working with containerized environments (Docker, Singularity/Apptainer).
  • A track record of building tools that increase developer velocity for ML teams.
  • Excellent judgment around trade-offs: performance vs complexity, research velocity vs maintainability.

Responsibilities

  • Build and own the training framework responsible for large-scale LLM training.
  • Design distributed training abstractions (data/tensor/pipeline parallelism, FSDP/ZeRO strategies, memory management, checkpointing).
  • Improve training throughput and stability on multi-node clusters (e.g., GB200/300, AMD, H200/100).
  • Develop and maintain tooling for monitoring, logging, debugging, and developer ergonomics.
  • Collaborate closely with infra teams to ensure our cluster, container environments, and hardware configurations support high-performance training.
  • Investigate and resolve performance bottlenecks across the ML systems stack.
  • Build robust systems that ensure reproducible, debuggable, large-scale runs.

Skills

Distributed training
JAX internals
Multi-node orchestration
CUDA/NCCL debugging
Container tooling
ML tooling for teams

Education

Bachelor's degree in CS or related field

Tools

JAX
Slurm
Ray
Kubernetes
Docker

Job description

Cohere is seeking a senior engineer to help build, maintain, and evolve the training framework that powers frontier-scale language models. This role sits at the intersection of large-scale training, distributed systems, and HPC infrastructure.

You will design and maintain core components that enable fast, reliable model training, and build tooling that connects research ideas to thousands of GPUs. You’ll work across the full ML systems stack with autonomy and impact.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior ML Training Frameworks & Tools Engineer
Senior ML Training Frameworks & Tools Engineer

Cohere • New York (NY)

Remote
USD 180,000 - 280,000
Health & dental benefits
Parental leave top-up
6 weeks vacation
Senior ML Systems Engineer, Frameworks & Tooling
Senior ML Systems Engineer, Frameworks & Tooling

Cohere • United States

Remote
USD 150,000 - 230,000
Staff Engineer, Training Infrastructure & ML Pipelines
Staff Engineer, Training Infrastructure & ML Pipelines

Cohere • United States

Remote
USD 180,000 - 260,000
Lunch stipend
Health & dental benefits
401K / Pension
+2
Staff Engineer, Production AI & Scalable LLMs
Staff Engineer, Production AI & Scalable LLMs

cohere • New York (NY)

On-site
USD 150,000 - 230,000
Lunch stipend
Health benefits
RRSP/401K
+4
Senior ML Systems Engineer, Frameworks & Tooling
Senior ML Systems Engineer, Frameworks & Tooling

Cohere • San Francisco (CA)

On-site
USD 150,000 - 180,000
Inclusive culture
Weekly lunch stipend
Health and dental benefits
+4
Senior ML Systems Engineer: Scalable Training Frameworks
Senior ML Systems Engineer: Scalable Training Frameworks

Cohere • San Francisco (CA)

Hybrid
USD 150,000 - 180,000
Inclusive culture
Weekly lunch stipend
Health and dental benefits
+4
Senior ML Systems Engineer – Distributed Training
Senior ML Systems Engineer – Distributed Training

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Equity
Health benefits
Remote-friendly US culture
+1
Staff Training Infra Engineer for Scalable ML Pipelines
Staff Training Infra Engineer for Scalable ML Pipelines

WeHireYou • Paris (IN)

Hybrid
USD 102,000 - 147,000
Lunch stipend
Health & dental benefits
Parental leave
+4
Senior ML Systems Engineer, Frameworks & Tooling
Senior ML Systems Engineer, Frameworks & Tooling

Cohere • New York (NY)

Remote
USD 180,000 - 280,000
Health & dental benefits
Parental leave top-up
6 weeks vacation
Senior ML Infrastructure Engineer — Frontier RL & LLM Training
Senior ML Infrastructure Engineer — Frontier RL & LLM Training

Preference Model • San Francisco (CA)

On-site
USD 200,000 - 350,000
Health, vision, dental benefits
401K match
Lunch provided onsite
+2