Tech Lead, Distributed Pre-Training Evals & Scale

Anthropic

San Francisco (CA)

Hybrid

USD 500,000 - 850,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity donation matching
Generous vacation and parental leave
Flexible working hours
Office in San Francisco

Job summary

Anthropic is seeking an experienced tech lead for the Evals Infrastructure team in San Francisco to own the distributed systems that measure model performance at scale. You will lead a team, collaborate with researchers, and ensure fast, reliable evals with reproducible results.

The role blends hands-on engineering with people leadership in a hybrid office setting, shaping how we evaluate and ship safe AI at scale.

Qualifications

  • Bachelor’s degree or equivalent in a related field.
  • 1+ year of managing engineers or tech-lead-with-reports experience.
  • Strong Python and Rust skills.
  • Experience with large-scale distributed systems.

Responsibilities

  • Lead the team building the distributed systems that schedule, orchestrate, and execute evals for our frontier model training.
  • Own eval throughput and cost: compute allocation across suites, queueing against constrained accelerator pools, caching and reuse of eval work.
  • Build and scale the harnesses researchers use to define, run, and iterate on evals.
  • Make eval results trustworthy — determinism, reproducibility, and honest uncertainty quantification on reported metrics.
  • Ensure eval signal reaches dashboards and reviews where launch decisions are made.
  • Contribute directly as an engineer while managing and growing the team, prioritizing its work, and coaching your reports.

Skills

Python
Rust
Distributed systems
Leadership
Cloud

Education

Bachelor's degree

Job description

Anthropic is seeking an experienced tech lead for the Evals Infrastructure team in San Francisco to own the distributed systems that measure model performance at scale. You will lead a team, collaborate with researchers, and ensure fast, reliable evals with reproducible results.

The role blends hands-on engineering with people leadership in a hybrid office setting, shaping how we evaluate and ship safe AI at scale.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Eval Infrastructure Tech Lead - Scale Safe AI Metrics
Eval Infrastructure Tech Lead - Scale Safe AI Metrics

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 500,000 - 850,000
Equity option
Vacation and leave
Flexible hours
+1
Eval Infra Tech Lead — Scalable Evals Orchestrator
Eval Infra Tech Lead — Scalable Evals Orchestrator

anthropic • San Francisco (CA)

Hybrid
USD 500,000 - 850,000
Equity donation matching
Generous vacation
Parental leave
Eval Infrastructure Tech Lead — Scalable AI Evaluation
Eval Infrastructure Tech Lead — Scalable AI Evaluation

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 500,000 - 850,000
Pre-training Distributed Systems Tech Lead / Manager
Pre-training Distributed Systems Tech Lead / Manager

Anthropic • San Francisco (CA)

Hybrid
USD 500,000 - 850,000
Equity donation matching
Generous vacation and parental leave
Flexible working hours
+1
Production ML Engineer - Scale & Reliability
Production ML Engineer - Scale & Reliability

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 850,000
Evals Infrastructure Tech Lead / Manager
Evals Infrastructure Tech Lead / Manager

anthropic • San Francisco (CA)

On-site
USD 500,000 - 850,000
Equity donation matching
Generous vacation
Parental leave
Evals Infrastructure Tech Lead / Manager Anthropic San Francisco, CA
Evals Infrastructure Tech Lead / Manager Anthropic San Francisco, CA

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 500,000 - 850,000
Equity option
Vacation and leave
Flexible hours
+1
Research Engineer - AI Scientist Infra & Scale
Research Engineer - AI Scientist Infra & Scale

SignalAI • San Francisco (CA)

On-site
USD 350,000 - 850,000
Competitive compensation
Equity donation matching
Generous vacation and parental leave
+2
Eval Engineer for AI Model Evaluations
Eval Engineer for AI Model Evaluations

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 500,000 - 850,000
AI Research Infrastructure Engineer
AI Research Infrastructure Engineer

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 850,000