Staff Cloud Inference Launch Engineer for Scalable AI

Anthropic

San Francisco (CA)

On-site

USD 320,000 - 485,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Anthropic is seeking an experienced Cloud Inference Engineer to scale Claude across major cloud platforms and drive end‑to‑end inference deployment. You will own API integrations, routing, inference execution, capacity management, and day‑to‑day operations for production workloads.

You will work on the model and inference launch pipeline, validating new launches, performance improvements, safeguards, and cross‑platform correctness to ensure reliable delivery to customers.

Qualifications

  • Strong interest in LLM serving; prior inference or ML experience is not required.
  • Significant software engineering experience with large-scale distributed systems serving millions of users.
  • Proven ability to build automation or test infrastructure that measurably improved release velocity or reliability.
  • Experience building or operating services on at least one major cloud platform (AWS, GCP, or Azure), with exposure to Kubernetes, Infrastructure as Code, or container orchestration.
  • Thrive in cross-functional collaboration with both internal teams and external partners.

Responsibilities

  • Be on the critical path for frontier model launches, bringing up inference for new model architectures and shipping them to cloud platforms in lockstep with our first‑party platform.
  • Work with the core inference team to bring new inference features (e.g. structured sampling, prompt caching, and more) to cloud platforms, owning the platform‑specific integration that gets them to production.
  • Identify and dive deep on the gaps that make inference behave differently across first‑party and CSPs — config drift, observability, deployment patterns, hard cross‑platform bugs — and fix them at the source rather than building platform‑specific workarounds.
  • Design, build, and own the CI/CD infrastructure for the inference server and load balancer across cloud platforms, with shadow traffic, performance baselines (throughput and latency), and correctness checks that catch regressions before production
  • Drive down merge‑to‑production cycle time by making validation faster, more parallel, and cost‑effective enough to run on the same constrained accelerator pool that serves customers, without trading away reliability
  • Analyze observability data across providers to identify performance bottlenecks, cost anomalies, and regressions, and drive remediation based on real‑world production workloads

Skills

Distributed systems
Software engineering
Automation
Cloud platforms
Cross-functional collaboration
Autonomy
Learning new tech

Education

Bachelor's degree

Tools

Kubernetes
Infrastructure as Code
Container orchestration
Python
Rust

Job description

Anthropic is seeking an experienced Cloud Inference Engineer to scale Claude across major cloud platforms and drive end‑to‑end inference deployment. You will own API integrations, routing, inference execution, capacity management, and day‑to‑day operations for production workloads.

You will work on the model and inference launch pipeline, validating new launches, performance improvements, safeguards, and cross‑platform correctness to ensure reliable delivery to customers.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cloud Inference Launch Engineer — Staff/Senior
Cloud Inference Launch Engineer — Staff/Senior

3M HEALTHCARE • Seattle (WA), San Francisco (CA)

Hybrid
USD 320,000 - 485,000
Cloud Inference Launch Engineer for Frontier Models
Cloud Inference Launch Engineer for Frontier Models

Anthropic • San Francisco (CA)

Hybrid
USD 320,000 - 485,000
Staff Senior Software Engineer, Inference Deployment
Staff Senior Software Engineer, Inference Deployment

Anthropic • United States

Hybrid
USD 320,000 - 485,000
Equity donation matching
Generous vacation
Parental leave
+2
Staff+ Cloud Infrastructure Engineer for Scalable AI
Staff+ Cloud Infrastructure Engineer for Scalable AI

Anthropic • San Francisco (CA)

Hybrid
USD 320,000 - 485,000
Staff + Sr. Software Engineer, Cloud Inference
Staff + Sr. Software Engineer, Cloud Inference

Anthropic • San Francisco (CA)

Hybrid
USD 300,000 - 485,000
Staff Platform Engineer: Cloud Distribution & Enterprise
Staff Platform Engineer: Cloud Distribution & Enterprise

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 405,000 - 485,000
Equity donation matching (optional)
Generous vacation and parental leave
Flexible working hours
Lead Node Infra Engineer — AI Compute & Cloud
Lead Node Infra Engineer — AI Compute & Cloud

SignalAI • New York (NY)

Hybrid
USD 405,000 - 485,000
Staff + Sr. Software Engineer, Cloud Inference Launch Engineering
Staff + Sr. Software Engineer, Cloud Inference Launch Engineering

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 320,000 - 485,000
Senior Staff Engineer, AI Inference Systems
Senior Staff Engineer, AI Inference Systems

Anthropic • New York (NY)

Hybrid
USD 320,000 - 485,000
Senior Staff Software Engineer — AI Inference Systems
Senior Staff Software Engineer — AI Inference Systems

Anthropic • New York (NY)

Hybrid
USD 320,000 - 485,000