About The Role
The Inference Team is responsible for building and maintaining critical systems that serve Claude to millions of users worldwide. We deliver our models via compute‑agnostic inference deployments, handling the entire stack from request routing to fleet‑wide orchestration across diverse AI accelerators.
Our dual mandate is to maximize compute efficiency to serve growing customer demand while enabling breakthrough research by providing scientific teams with high‑performance inference infrastructure.
Key Responsibilities
- Design, build, and maintain distributed systems that serve Claude to millions of users worldwide.
- Develop intelligent request routing, load balancing, and traffic management systems across thousands of accelerators.
- Maximize compute efficiency across the fleet by autoscaling and orchestrating production, research, and experimental workloads.
- Build and operate production‑grade deployment pipelines for releasing new models to users.
- Provide high‑performance inference infrastructure that enables researchers to develop next‑generation models.
- Integrate new AI accelerator platforms and support inference for new model architectures.
- Use observability data to tune and improve performance based on real‑world production workloads.
Representative Projects
- Design intelligent routing algorithms that optimize request distribution across thousands of accelerators.
- Autoscale compute fleet to dynamically match supply with demand across production, research, and experimental workloads.
- Build production‑grade deployment pipelines for releasing new models to millions of users.
- Integrate new AI accelerator platforms to maintain hardware‑agnostic competitive advantage.
- Contribute to new inference features such as structured sampling and prompt caching.
- Support inference for new model architectures.
- Analyze observability data to tune performance based on real‑world production workloads.
- Manage multi‑region deployments and geographic routing for global customers.
Minimum Qualifications
- Significant software engineering experience, especially with distributed systems.
- Results‑oriented, with a bias toward flexibility and impact.
- Willingness to pick up slack, even if it falls outside your job description.
- Enjoy pair programming.
- Desire to learn more about machine learning systems and infrastructure.
- Thriving in environments where technical excellence drives business results and research breakthroughs.
- Care about the societal impacts of your work.
Preferred Qualifications
- Experience with high‑performance, large‑scale distributed systems.
- Experience implementing and deploying machine learning systems at scale.
- Experience with load balancing, request routing, or traffic management systems.
- Familiarity with LLM inference optimization, batching, and caching strategies.
- Experience with Kubernetes and cloud infrastructure (AWS, GCP, Azure).
- Proficiency in Python or Rust.
Annual Salary
$320,000 – $485,000 USD
Logistics
Minimum education: Bachelor’s degree or equivalent combination of education, training, and/or experience.
Required field of study: Field relevant to the role based on coursework, training, or professional experience.
Minimum years of experience: Correlated with internal job level requirements.
Location: Hybrid model, requiring staff to be in one of our offices at least 25% of the time, with some roles requiring more office presence.
Visa sponsorship: We sponsor visas and will make every reasonable effort to secure a visa when an offer is made.