Staff + Senior Software Engineer, Inference

Anthropic

New York (NY)

Hybrid

USD 320,000 - 485,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Anthropic is looking for an experienced engineer to join the Inference Team in New York. The role involves designing and maintaining systems that serve Claude to millions of users worldwide, emphasizing distributed systems and machine learning infrastructure.

Candidates should have a Bachelor's degree and significant software engineering experience, particularly with distributed systems. The position offers a hybrid work model allowing flexibility while ensuring collaboration and productivity.

Qualifications

  • Significant software engineering experience, especially with distributed systems.
  • Enjoy pair programming and learning about machine learning systems.
  • Thrives in technical excellence-driven environments.

Responsibilities

  • Design, build, and maintain distributed systems that serve Claude.
  • Develop intelligent request routing and load balancing systems.
  • Maximize compute efficiency across the fleet.

Skills

Software engineering
Distributed systems
Machine learning
Kubernetes
Python
Rust

Education

Bachelor’s degree or equivalent

Tools

AWS
GCP
Azure

Job description

About The Role

The Inference Team is responsible for building and maintaining critical systems that serve Claude to millions of users worldwide. We deliver our models via compute‑agnostic inference deployments, handling the entire stack from request routing to fleet‑wide orchestration across diverse AI accelerators.

Our dual mandate is to maximize compute efficiency to serve growing customer demand while enabling breakthrough research by providing scientific teams with high‑performance inference infrastructure.

Key Responsibilities
  • Design, build, and maintain distributed systems that serve Claude to millions of users worldwide.
  • Develop intelligent request routing, load balancing, and traffic management systems across thousands of accelerators.
  • Maximize compute efficiency across the fleet by autoscaling and orchestrating production, research, and experimental workloads.
  • Build and operate production‑grade deployment pipelines for releasing new models to users.
  • Provide high‑performance inference infrastructure that enables researchers to develop next‑generation models.
  • Integrate new AI accelerator platforms and support inference for new model architectures.
  • Use observability data to tune and improve performance based on real‑world production workloads.
Representative Projects
  • Design intelligent routing algorithms that optimize request distribution across thousands of accelerators.
  • Autoscale compute fleet to dynamically match supply with demand across production, research, and experimental workloads.
  • Build production‑grade deployment pipelines for releasing new models to millions of users.
  • Integrate new AI accelerator platforms to maintain hardware‑agnostic competitive advantage.
  • Contribute to new inference features such as structured sampling and prompt caching.
  • Support inference for new model architectures.
  • Analyze observability data to tune performance based on real‑world production workloads.
  • Manage multi‑region deployments and geographic routing for global customers.
Minimum Qualifications
  • Significant software engineering experience, especially with distributed systems.
  • Results‑oriented, with a bias toward flexibility and impact.
  • Willingness to pick up slack, even if it falls outside your job description.
  • Enjoy pair programming.
  • Desire to learn more about machine learning systems and infrastructure.
  • Thriving in environments where technical excellence drives business results and research breakthroughs.
  • Care about the societal impacts of your work.
Preferred Qualifications
  • Experience with high‑performance, large‑scale distributed systems.
  • Experience implementing and deploying machine learning systems at scale.
  • Experience with load balancing, request routing, or traffic management systems.
  • Familiarity with LLM inference optimization, batching, and caching strategies.
  • Experience with Kubernetes and cloud infrastructure (AWS, GCP, Azure).
  • Proficiency in Python or Rust.
Annual Salary

$320,000 – $485,000 USD

Logistics

Minimum education: Bachelor’s degree or equivalent combination of education, training, and/or experience.

Required field of study: Field relevant to the role based on coursework, training, or professional experience.

Minimum years of experience: Correlated with internal job level requirements.

Location: Hybrid model, requiring staff to be in one of our offices at least 25% of the time, with some roles requiring more office presence.

Visa sponsorship: We sponsor visas and will make every reasonable effort to secure a visa when an offer is made.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff + Senior Software Engineer, Inference San Francisco, CA | New York City, NY | Seattle, WA
Staff + Senior Software Engineer, Inference San Francisco, CA | New York City, NY | Seattle, WA

Anthropic • New York (NY)

Hybrid
USD 320,000 - 485,000
Staff + Sr. Software Engineer, Cloud Inference Launch Engineering
Staff + Sr. Software Engineer, Cloud Inference Launch Engineering

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 320,000 - 485,000
Staff + Senior Software Engineer, Inference Deployment
Staff + Senior Software Engineer, Inference Deployment

United States Digital Space LLC • San Francisco (CA)

On-site
USD 320,000 - 485,000
Visa sponsorship
Hybrid work policy where staff are in the office at least 25% of the time
Staff + Sr. Software Engineer, Cloud Inference
Staff + Sr. Software Engineer, Cloud Inference

Anthropic • San Francisco (CA)

Hybrid
USD 300,000 - 485,000
Research Engineer, Infrastructure, Inference
Research Engineer, Infrastructure, Inference

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, vision benefits
Unlimited PTO
Parental leave
+1
Staff+ Software Engineer, Inference Runtime
Staff+ Software Engineer, Inference Runtime

Anthropic • Seattle (WA)

Hybrid
USD 405,000 - 485,000
Staff+ Software Engineer (Inference Runtime)
Staff+ Software Engineer (Inference Runtime)

jobr.pro • San Francisco (CA)

Hybrid
USD 405,000 - 485,000
Competitive compensation
Generous vacation
Flexible working hours
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
Research Engineer, Infrastructure, Inference
Research Engineer, Infrastructure, Inference

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Staff+ Software Engineer, Inference Runtime
Staff+ Software Engineer, Inference Runtime

Menlo Ventures • New York (NY)

Hybrid
USD 405,000 - 485,000