Production ML Engineer - Scale & Reliability

Neura Market

San Francisco, Northern (CA, KY)

Hybrid

USD 350,000 - 850,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Anthropic is seeking a Research Engineer for its ML Performance and Scaling team in San Francisco. You will ensure production pretrained models train reliably and efficiently at scale, bridging research and engineering across the full training stack.

You will design experiments, optimize performance, and respond to incidents during model launches while collaborating with teams in SF and London. This role combines deep technical work with operational excellence and impact on safe, beneficial AI

Qualifications

  • Hands-on experience training large language models.
  • Experience with JAX, TPU, PyTorch, or large-scale distributed systems.
  • Willingness to be on-call during model launches and incidents.
  • Strong collaboration skills across time zones and teams.

Responsibilities

  • Own production pretraining pipeline components (model ops, performance, observability, reliability).
  • Debug across full stack from hardware to training dynamics and eval infra.
  • Design experiments to improve training efficiency and uptime.
  • Respond to on-call incidents during launches and coordinate solutions.
  • Build/maintain production logging, dashboards, and evaluation tools.
  • Add capabilities to training codebase (long context, novel architectures).
  • Collaborate with SF and London teams and with other ML groups.
  • Document systems and debugging approaches for institutional knowledge.

Skills

Large language models
Distributed systems
Research engineering
On-call readiness
Cross-team collaboration

Education

Bachelor's degree in a relevant field

Tools

JAX
TPU
PyTorch
Open-source ML frameworks

Job description

Anthropic is seeking a Research Engineer for its ML Performance and Scaling team in San Francisco. You will ensure production pretrained models train reliably and efficiently at scale, bridging research and engineering across the full training stack.

You will design experiments, optimize performance, and respond to incidents during model launches while collaborating with teams in SF and London. This role combines deep technical work with operational excellence and impact on safe, beneficial AI

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Scale-Focused Research Engineer, Pretraining
Scale-Focused Research Engineer, Pretraining

SignalAI • San Francisco (CA)

On-site
USD 350,000 - 850,000
Production ML Engineer — Scale & Train LLMs
Production ML Engineer — Scale & Train LLMs

Anthropic • San Francisco (CA)

On-site
USD 350,000 - 850,000
Generous vacation and parental leave
Flexible working hours
Office space for collaboration
Applied ML Engineer: Scale & Deploy Large Models
Applied ML Engineer: Scale & Deploy Large Models

AI Breaking Wire • San Francisco (CA)

Hybrid
USD 250,000 - 380,000
Equity
Medical, dental, and vision
Unlimited PTO
+2
Production ML Ops Engineer: Scale AI Deployment
Production ML Ops Engineer: Scale AI Deployment

Ent • San Francisco (CA)

Hybrid
USD 150,000 - 230,000
Equity
Medical, dental, and vision coverage
Flexible PTO
+3
Production ML Engineer - Systems & Reliability
Production ML Engineer - Systems & Reliability

Sprinter Health • San Francisco (CA)

Hybrid
USD 150,000 - 200,000
Medical, dental, and vision plans 100%
Flexible PTO
401(k) with match
+2
ML Infrastructure Engineer — Scale Research to Production
ML Infrastructure Engineer — Scale Research to Production

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 500,000
Office in San Francisco
Research Engineer, Large-Scale ML Training (Equity)
Research Engineer, Large-Scale ML Training (Equity)

Togetherai • San Francisco (CA)

On-site
USD 200,000 - 290,000
Competitive compensation
Startup equity
Health insurance
+1
Research Engineer, Pretraining Scaling
Research Engineer, Pretraining Scaling

Anthropic • San Francisco (CA)

On-site
USD 350,000 - 850,000
ML Infra Engineer: Scale Research Pipelines & Tools
ML Infra Engineer: Scale Research Pipelines & Tools

Anthropic • San Francisco (CA)

On-site
USD 350,000 - 500,000
Research Engineer: Scalable Pretraining
Research Engineer: Scalable Pretraining

Anthropic • San Francisco (CA)

On-site
USD 350,000 - 850,000