Senior Machine Learning Engineer, Core Systems

AI Breaking Wire

San Francisco, Northern (CA, KY)

Hybrid

USD 180,000 - 280,000

Full time

6 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Equity
Health benefits
Remote-friendly US culture
Learning stipend

Job summary

Anthropic is seeking a Senior ML Engineer to join the Core Systems team to optimize training and inference infrastructure for the Claude family of foundation models. You will design high-performance distributed pipelines and work with researchers to remove scaling bottlenecks.

Ideal candidates have 5+ years of ML infrastructure experience, expert C++/Python skills, and hands-on work with AWS/GCP and Kubernetes.

Qualifications

  • 5+ years of software engineering experience with a focus on large-scale ML infrastructure.
  • Expert-level C++ and Python programming.
  • Deep familiarity with distributed training paradigms (FSDP, Tensor Parallelism, Pipeline Parallelism).

Responsibilities

  • Architect and implement high-performance distributed training and inference pipelines for massive transformer models.
  • Collaborate with researchers to unblock scaling bottlenecks and improve hardware utilization.
  • Develop low-latency inference runtimes and model optimization techniques (quantization, pruning, kernel fusion).
  • Monitor and triage large-scale cluster performance issues across heterogeneous GPU clusters.

Skills

C++
Python
Distributed ML
Kubernetes
AWS / GCP

Tools

PyTorch
TensorFlow

Job description

About the Role

Anthropic is dedicated to building reliable, interpretable, and steerable AI systems. We are seeking a Senior ML Engineer to join our Core Systems team to optimize the training and inference infrastructure for our Claude family of foundation models.

Responsibilities
  • Architect and implement high-performance distributed training and inference pipelines for massive transformer models.
  • Work hand-in-hand with research scientists to unblock scaling bottlenecks and improve hardware utilization.
  • Develop low-latency inference runtimes and model optimization techniques (quantization, pruning, kernel fusion).
  • Monitor and triage large-scale cluster performance issues across heterogeneous GPU clusters.
Requirements
  • 5+ years of software engineering experience with a strong focus on large-scale machine learning infrastructure.
  • Expert-level C++ and Python programming skills.
  • Deep familiarity with distributed training paradigms (FSDP, Tensor Parallelism, Pipeline Parallelism).
  • Experience working with cloud providers (AWS, GCP) and orchestration tools like Kubernetes.
Benefits
  • Competitive salary and generous equity package.
  • Premium medical, dental, and vision benefits.
  • Flexible working hours and remote-friendly culture within the US.
  • Annual learning and development stipend.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Systems Engineer – Distributed Training
Senior ML Systems Engineer – Distributed Training

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Equity
Health benefits
Remote-friendly US culture
+1
Machine Learning Engineer
Machine Learning Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 150,000 - 190,000
Member of Technical Staff, Applied AI
Member of Technical Staff, Applied AI

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Health insurance
Dental insurance
Vision insurance
+1
Senior Machine Learning Engineer, Infrastructure
Senior Machine Learning Engineer, Infrastructure

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
Senior Machine Learning Engineer, Alignment and Safety
Senior Machine Learning Engineer, Alignment and Safety

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Equity compensation
Competitive salary
Remote and hybrid options
+1
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training

Preference Model • San Francisco (CA)

On-site
USD 180,000 - 300,000
Cash and equity
Ownership & autonomy
Visa sponsorship
+5
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training

Preference Model • Seattle (WA)

On-site
USD 180,000 - 300,000
Health insurance
Vision insurance
Dental insurance
+3
Staff Software Engineer, Scalable AI Inference Systems
Staff Software Engineer, Scalable AI Inference Systems

Anthropic • San Francisco (CA)

On-site
USD 320,000 - 485,000
ML Infrastructure Engineer
ML Infrastructure Engineer

Lattice, Inc. • San Francisco (CA)

Hybrid
USD 200,000 - 280,000
Competitive salary
Premium health, dental, and vision insurance
Unlimited PTO
+2
Member of Technical Staff, ML Engineer
Member of Technical Staff, ML Engineer

Physical Superintelligence • Boston (MA)

Hybrid
USD 140,000 - 210,000