Staff+ Software Engineer, ML Sampling Path

EngineersOfAI

San Francisco, Northern (CA, KY)

Hybrid

USD 320,000 - 485,000

Full time

9 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Visa sponsorship

Job summary

Anthropic is seeking a seasoned software engineer for the Safeguards ML Sampling Path team in California. You will design and operate backend services powering Claude’s safety systems, ensuring low latency and high reliability across the token generation path.

You will own SLOs, respond to incidents, and implement safe deployment strategies while collaborating with inference and research teams to optimize per-token performance and cost as models grow.

Qualifications

  • 8+ years of industry software engineering experience.
  • Familiarity with LLM inference systems and transformer-based models (not required, but a plus).
  • Designed, built, and operated high QPS systems in production with incident response and postmortem follow-through.

Responsibilities

  • Design, build, and operate backend systems that process every token on the generation path for Claude requests, including the streaming contract with the API and inference engines.
  • Own latency and reliability end to end: define and maintain SLOs and error budgets, lead incident response and postmortem follow-through.
  • Ship changes to the hot path rapidly but safely with canaries, gradual rollouts, and latency gating; drive per-token performance and cost control.
  • Set technical direction for the sampling path; lead design reviews, balance latency, reliability, and cost with teams, mentor engineers.

Skills

High QPS systems
Distributed systems
Latency optimization
Incident response

Education

Bachelor's degree in computer science or related field

Job description

About Anthropic

Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.

About the role:

The Safeguards ML Sampling Path team builds and operates the production services that power Claude's safety systems. These services sit on the token generation path across every platform Claude runs on: every request must pass through them, and each millisecond of added latency is wait time for our users. You’ll keep p99 latency flat as traffic grows, build for robustness as dependencies time out or partially fail, and ship changes safely to a system that cannot go down.

What you'll do:
  • Design, build, and operate the backend systems that process every token on the generation path for Claude requests, including the streaming contract with the API and inference engines.
  • Own latency and reliability end to end: define and maintain SLOs and error budgets for added latency, time-to-first-token, and availability, and lead incident response and postmortem follow-through.
  • Ship changes to the hot path rapidly but safely — canaried and gradual rollouts, error budget and latency gating, fast rollbacks — and drive per-token performance: chase tail latency and keep cost flat as traffic, models, and checks per request grow.
  • Set technical direction for the sampling path: lead design reviews, make latency, reliability, and cost trade-off calls with the inference and research teams, mentor engineers, and raise the operational bar for the wider Safeguards organization.
You may be a good fit if you:
  • Have designed, built, and operated high QPS systems at global scale, and were accountable for them in production: incident response, outages, and postmortem-driven remediation.
  • Have a strong foundation in distributed systems: replication, consistency tradeoffs, failure modes, and SLO management under load.
  • Design systems for graceful degradation: you plan for a slow dependency, a dropped stream, or a half-rolled-out deploy before it happens, and build so the system degrades predictably instead of failing.
  • Have successfully shipped broad or all-encompassing changes to mission critical systems (e.g., database migrations, interface changes, rewrites).
Strong candidates may also have:
  • 8+ years of industry software engineering experience.
  • Familiarity with LLM inference systems and transformer-based models (not required, but a plus).

The annual compensation range for this role is listed below.

For sales roles, the range provided is the role’s On Target Earnings ("OTE") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role.

Annual Salary:

$320,000 — $485,000 USD

Logistics

Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience

Required field of study:A field relevant to the role as demonstrated through coursework, training, or professional experience

Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position

Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices.

Visa sponsorship:We do sponsor visas! However, we aren't able to successfully sponsor visas for every rol

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff+ Software Engineer, ML Sampling Path
Staff+ Software Engineer, ML Sampling Path

Socket.dev • San Francisco (CA)

Hybrid
USD 320,000 - 485,000
Staff+ Site Reliability Engineer, Safeguards ML Infra
Staff+ Site Reliability Engineer, Safeguards ML Infra

Anthropic • Seattle (WA), New York (NY), San Francisco (CA)

On-site
USD 405,000 - 485,000
Staff+ Software Engineer, ML Inference Path
Staff+ Software Engineer, ML Inference Path

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 272,000 - 368,000
Staff+ Software Engineer, ML Inference Path
Staff+ Software Engineer, ML Inference Path

Socket.dev • San Francisco (CA)

Hybrid
USD 320,000 - 485,000
Equity donation matching
Generous vacation
Parental leave
+2
Staff+ Software Engineer, Safeguards ML Infrastructure
Staff+ Software Engineer, Safeguards ML Infrastructure

Doist • San Francisco (CA)

Hybrid
USD 320,000 - 485,000
Staff+ Site Reliability Engineer, Safeguards ML Infra
Staff+ Site Reliability Engineer, Safeguards ML Infra

Socket.dev • San Francisco (CA), Northern (KY)

Hybrid
USD 291,000 - 485,000
Staff Software Engineer (AI Reliability)
Staff Software Engineer (AI Reliability)

Anthropic • San Francisco (CA)

On-site
USD 325,000 - 485,000
Competitive compensation
Generous vacation
Flexible working hours
Technical Program Manager, Safeguards (Infrastructure & Evals)
Technical Program Manager, Safeguards (Infrastructure & Evals)

Anthropic • Seattle (WA)

On-site
USD 290,000 - 365,000
Staff+ Software Engineer, Claude App Infrastructure
Staff+ Software Engineer, Claude App Infrastructure

Anthropic • New York (NY)

Hybrid
USD 320,000 - 485,000
Staff + Sr. Software Engineer, Cloud Inference
Staff + Sr. Software Engineer, Cloud Inference

Anthropic • San Francisco (CA)

Hybrid
USD 300,000 - 485,000