Senior ML Systems Engineer: Scalable Training Frameworks

Cohere

San Francisco (CA)

Hybrid

USD 150,000 - 180,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Inclusive culture
Weekly lunch stipend
Health and dental benefits
100% Parental Leave top-up
Personal enrichment benefits
Remote-flexible work
6 weeks of vacation

Job summary

A leading AI research firm located in San Francisco is seeking a Senior ML Systems Engineer to build and maintain the training framework for large-scale language models. The role involves designing distributed training solutions and improving training throughput across multi-node clusters. The ideal candidate will have strong engineering experience in distributed training, familiarity with JAX, and excellent collaboration skills. This position promises significant ownership over critical components and engagement with cutting-edge AI technologies while offering a flexible work environment.

Qualifications

  • Strong engineering experience in large-scale distributed training or HPC systems.
  • Experience with multi-node cluster orchestration.
  • Comfort debugging performance issues across the ML systems stack.

Responsibilities

  • Build and own the training framework for large-scale LLM training.
  • Design distributed training abstractions.
  • Improve training throughput on multi-node clusters.

Skills

Large-scale distributed training
HPC systems
Performance debugging
Containerized environments
Collaboration skills

Tools

JAX
CUDA/NCCL
Docker
Kubernetes

Job description

A leading AI research firm located in San Francisco is seeking a Senior ML Systems Engineer to build and maintain the training framework for large-scale language models. The role involves designing distributed training solutions and improving training throughput across multi-node clusters. The ideal candidate will have strong engineering experience in distributed training, familiarity with JAX, and excellent collaboration skills. This position promises significant ownership over critical components and engagement with cutting-edge AI technologies while offering a flexible work environment.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff ML Systems Engineer — Distributed Training at Scale
Staff ML Systems Engineer — Distributed Training at Scale

RadixArk • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Comprehensive benefits
Flexible work arrangements
Tech Lead for Distributed ML Systems & Training Platform
Tech Lead for Distributed ML Systems & Training Platform

Scale AI • New York (NY), San Francisco (CA)

On-site
USD 275,000 - 350,000
Health, dental, and vision coverage
Retirement benefits
Learning and development stipend
+2
Senior ML Training Systems Engineer - Distributed GPU Infra
Senior ML Training Systems Engineer - Distributed GPU Infra

Baseten • San Francisco (CA)

On-site
USD 150,000 - 200,000
Competitive compensation, including equity
100% coverage of medical, dental, and vision insurance
Generous PTO policy
+2
ML Systems Engineer: Distributed LLM Training & Inference
ML Systems Engineer: Distributed LLM Training & Inference

Scale AI • Seattle (WA), New York (NY), San Francisco (CA)

On-site
USD 200,800 - 251,000
Comprehensive health coverage
Equity-based compensation
Retirement benefits
+3
Staff ML Engineer - Build Scalable Multimodal AI Platform
Staff ML Engineer - Build Scalable Multimodal AI Platform

kadence • San Francisco (CA)

On-site
USD 180,000 - 250,000
Research Software Engineer — Scalable RL & Distributed Training
Research Software Engineer — Scalable RL & Distributed Training

Reflection AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Top-tier compensation
Comprehensive health insurance
Fully paid parental leave
+2
Senior ML Training Frameworks & Tools Engineer
Senior ML Training Frameworks & Tools Engineer

Cohere • New York (NY)

Remote
USD 180,000 - 280,000
Health & dental benefits
Parental leave top-up
6 weeks vacation
RL Research Engineer - Scalable, Safe AI Systems
RL Research Engineer - Scalable, Safe AI Systems

Anthropic • San Francisco (CA)

Hybrid
USD 500,000 - 850,000
Competitive compensation
Equity donation matching
Generous vacation and parental leave
+2
Senior ML Engineer — Scalable AI & Personalization (Remote)
Senior ML Engineer — Scalable AI & Personalization (Remote)

Quizlet • San Francisco (CA)

On-site
USD 194,834 - 241,913
Inclusive culture
Cutting-edge technology
Strong company mission
Tech Lead Manager for Scalable LLM Training Platform
Tech Lead Manager for Scalable LLM Training Platform

Scale AI • San Francisco (CA)

On-site
USD 120,000 - 160,000