Research Engineer, Production-Scale Pretraining

Anthropic

California (MO)

On-site

USD 350,000 - 850,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation
Equity donation matching
Generous vacation & parental leave
Flexible working hours
Office space

Job summary

Anthropic is seeking a Research Engineer for its ML Performance and Scaling team in San Francisco. You will help train production pretrained models, optimize performance, and ensure reliable, scalable operations across the full training stack.

You will work at the intersection of research and engineering, tackling incidents during launches and designing experiments to improve efficiency, uptime, and model quality.

Qualifications

  • Hands-on experience training large language models with modern ML stacks.
  • Experience with production ML systems, observability, and deployment pipelines.
  • Ability to work across research and engineering and communicate clearly.

Responsibilities

  • Own critical aspects of the production pretraining pipeline, including model operations, performance optimization, observability, and reliability.
  • Debug and resolve complex issues across the full stack—from hardware errors and networking to training dynamics and evaluation infrastructure.
  • Design and run experiments to improve training efficiency, reduce step time, increase uptime, and enhance model performance.
  • Respond to on-call incidents during model launches, diagnosing problems quickly and coordinating solutions across teams.
  • Build and maintain production logging, monitoring dashboards, and evaluation infrastructure.
  • Add new capabilities to the training codebase, such as long context support or novel architectures.
  • Collaborate across SF and London teams, as well as with Tokens, Architectures, and Systems teams.
  • Document systems, debugging approaches, and lessons learned.

Skills

LLM training
JAX/TPU
PyTorch
Distributed systems

Education

Bachelor’s degree in a relevant field

Tools

JAX
PyTorch
TPU
Distributed training

Job description

Anthropic is seeking a Research Engineer for its ML Performance and Scaling team in San Francisco. You will help train production pretrained models, optimize performance, and ensure reliable, scalable operations across the full training stack.

You will work at the intersection of research and engineering, tackling incidents during launches and designing experiments to improve efficiency, uptime, and model quality.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer: Scalable Pretraining
Research Engineer: Scalable Pretraining

Anthropic • San Francisco (CA)

On-site
USD 350,000 - 850,000
Production ML Engineer — Scale & Train LLMs
Production ML Engineer — Scale & Train LLMs

Anthropic • San Francisco (CA)

On-site
USD 350,000 - 850,000
Generous vacation and parental leave
Flexible working hours
Office space for collaboration
Research Engineer: AI Scientist Infrastructure
Research Engineer: AI Scientist Infrastructure

Anthropic • California (MO)

On-site
USD 350,000 - 850,000
Research Engineer, Pretraining Scaling
Research Engineer, Pretraining Scaling

Anthropic • California (MO)

On-site
USD 350,000 - 850,000
Competitive compensation
Equity donation matching
Generous vacation & parental leave
+2
Research Engineer, Pretraining Scaling
Research Engineer, Pretraining Scaling

Anthropic • San Francisco (CA)

On-site
USD 350,000 - 850,000
Tech Lead, Distributed Pre-Training Evals & Scale
Tech Lead, Distributed Pre-Training Evals & Scale

Anthropic • San Francisco (CA)

Hybrid
USD 500,000 - 850,000
Equity donation matching
Generous vacation and parental leave
Flexible working hours
+1
Distributed ML Training Performance Engineer
Distributed ML Training Performance Engineer

OpenAI • California (MO)

Hybrid
USD 170,000 - 260,000
Relocation assistance
Research Data Platform Engineer
Research Data Platform Engineer

Anthropic Limited • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 405,000
Research Engineer — Product-Driven ML & Deployment
Research Engineer — Product-Driven ML & Deployment

SupportFinity™ • San Francisco (CA)

Hybrid
USD 295,000 - 555,000
Hybrid work model
Relocation assistance
Research Engineer: AI Scientist Infra & Pipelines
Research Engineer: AI Scientist Infra & Pipelines

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 850,000
Flexible hours
Generous vacation
Parental leave
+2