Staff + Senior Software Engineer (Scaling)

Anthropic

Seattle (WA)

On-site

USD 180,000 - 210,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Health insurance
Parental leave
Flexible PTO
Relocation assistance
Daily meals
Education stipend

Job summary

Anthropic in Seattle seeks a senior software engineer to design, build, and maintain distributed inference systems that serve Claude at scale. You will tackle routing, load balancing, and fleet orchestration across cloud providers, accelerators, and Kubernetes, with a focus on performance and reliability across millions of users.

You should have extensive experience with large-scale distributed systems, Python or Rust, and ML systems at scale; familiarity with AWS/GCP/Azure, model serving

Qualifications

  • Significant software engineering experience with distributed systems.

Responsibilities

  • Design, build, and maintain the distributed systems that serve Claude to millions of users worldwide.
  • Develop intelligent request routing, load balancing, and traffic management across thousands of accelerators and cloud providers.
  • Maximize compute efficiency and optimize cost across the fleet by autoscaling and orchestrating workloads across production, research, and experimental environments.
  • Build and operate production-grade deployment pipelines for releasing new models to users.
  • Provide high-performance inference infrastructure that enables researchers to develop next-generation models.
  • Integrate new AI accelerator platforms and support inference for new model architectures.
  • Analyze observability data to tune performance based on real-world production workloads.
  • Manage multi-region deployments and geographic routing for global customers.

Skills

Distributed systems
Kubernetes
Cloud infrastructure
Python
Rust
LLM inference optimization
Batching
Caching strategies
Load balancing
Traffic management
Observability

Tools

Kubernetes
AWS
GCP
Azure

Job description

  • Our Inference team is responsible for building and scaling the critical systems that serve Claude to millions of users worldwide
  • We bring Claude to life by serving our models via the industry’s largest compute-agnostic inference deployments
  • We are responsible for the entire stack from intelligent request routing to fleet-wide orchestration across diverse AI accelerators
  • The team has a dual mandate: maximizing compute efficiency to reliably serve our explosive customer growth, while enabling breakthrough research by giving our scientists the high-performance inference infrastructure they need to develop next-generation models
  • We tackle complex, distributed systems challenges across multiple accelerator families and emerging AI hardware running in multiple cloud platforms
  • Inference systems are highly performance sensitive distributed systems
  • Inference serves hundreds of thousands of customers every day, and the size & span of the inference fleet requires sophisticated routing, scaling, and networking systems
  • Design, build, and maintain the distributed systems that serve Claude to millions of users worldwide
  • Develop resilient, flexible systems that adapt in real time to real world events
  • Develop intelligent request routing, load balancing, and traffic management systems across thousands of accelerators and multiple cloud providers
  • Maximize compute efficiency and optimize cost across the fleet by autoscaling and orchestrating production, research, and experimental workloads across multiple cloud providers
  • Build and operate production-grade deployment pipelines for releasing new models to users
  • Provide high-performance inference infrastructure that enables researchers to develop next-generation models
  • Integrate new AI accelerator platforms and support inference for new model architectures
  • Representative projects
  • Designing intelligent routing algorithms that optimize request distribution across many accelerators in different environments
  • Autoscaling our compute fleet to dynamically match supply with demand across production, research, and experimental workloads
  • Building production-grade deployment pipelines for releasing new models to millions of users reliably
  • Contributing to new inference features
  • Supporting inference for new model architectures
  • Analyzing observability data to tune performance based on real-world production workloads
  • Managing multi-region deployments and geographic routing for global customers
Benefits
  • Comprehensive health, dental, and vision insurance for you and your dependents
  • Inclusive fertility benefits via Carrot Fertility
  • 22 weeks of paid parental leave
  • Flexible paid time off and absence policies
  • Mental health support for you and your dependents
  • Competitive salary and equity packages
  • Optional equity donation matching at a 1:1 ratio, up to 25% of your equity grant
  • Retirement plans with competitive matching
  • Life and income protection plans
  • $500/month flexible wellness and time saver stipend
  • Commuter benefits
  • Annual education stipend
  • Home office stipends
  • Relocation support for those moving for Anthropic
  • Daily meals and snacks in the office

Care about the societal impacts of your workThrive in environments where technical excellence directly drives both business results and research breakthroughsSignificant software engineering experience, particularly with distributed systemsDesire to learn more about machine learning systems and infrastructureWillingness to pick up slack, even if it goes outside your job descriptionResults-oriented, with a bias towards flexibility and impactWe encourage you to apply even if you do not believe you meet every single qualificationExperience with high-performance, large-scale distributed systemsExperience with Kubernetes and cloud infrastructure (AWS, GCP, Azure)Familiarity with LLM inference optimization, batching, and caching strategiesExperience with load balancing, request routing, or traffic management systemsProficiency in Python or RustExperience implementing and deploying machine learning systems at scale

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff + Senior Software Engineer, Inference Deployment
Staff + Senior Software Engineer, Inference Deployment

Anthropic • United States

Hybrid
USD 320,000 - 485,000
Equity donation matching
Generous vacation
Parental leave
+2
Staff + Sr. Software Engineer, Cloud Inference
Staff + Sr. Software Engineer, Cloud Inference

Anthropic • San Francisco (CA)

Hybrid
USD 300,000 - 485,000
Staff + Senior Software Engineer, Inference
Staff + Senior Software Engineer, Inference

Anthropic • New York (NY)

On-site
USD 320,000 - 485,000
Engineering Manager, Inference Infrastructure
Engineering Manager, Inference Infrastructure

Anthropic • New York (NY)

Hybrid
USD 405,000 - 625,000
Competitive compensation
Equity donation
Vacation and parental leave
+2
Engineering Manager, Inference Infrastructure
Engineering Manager, Inference Infrastructure

Anthropic • Seattle (WA)

Hybrid
USD 405,000 - 625,000
Competitive compensation
Equity matching
Parental leave
+2
Staff + Senior Software Engineer, Inference San Francisco, CA | New York City, NY | Seattle, WA
Staff + Senior Software Engineer, Inference San Francisco, CA | New York City, NY | Seattle, WA

Anthropic • New York (NY)

Hybrid
USD 320,000 - 485,000
Staff + Sr. Software Engineer, Scaling
Staff + Sr. Software Engineer, Scaling

Anthropic • New York (NY)

Hybrid
USD 320,000 - 485,000
Competitive compensation
Equity donation matching (optional)
Generous vacation
+3
Inference Performance Engineer
Inference Performance Engineer

Adaption Labs • San Francisco (CA)

On-site
USD 180,000 - 260,000
Flexible work
Adaption Passport
Lunch stipend
+1
Staff Applied AI Inference Engineer
Staff Applied AI Inference Engineer

Crusoe Energy Systems • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health benefits
401(k) match
Paid time off
+1
Staff + Sr. Software Engineer, Cloud Inference Launch Engineering
Staff + Sr. Software Engineer, Cloud Inference Launch Engineering

Anthropic • San Francisco (CA)

On-site
USD 320,000 - 485,000