Staff+ Software Engineer, Capacity Engineering

Anthropic

San Francisco, New York, Seattle (CA, NY, WA)

Hybrid

USD 320,000 - 485,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Anthropic seeks a Capacity Engineer to help manage one of the largest AI infrastructure fleets. You will build production systems that ingest telemetry, provide observability, and measure utilization of compute resources across Kubernetes clusters and cloud providers.

You will own planning, capacity allocation, and efficiency programs, rebuild data pipelines in BigQuery, and partner with researchers, finance, and leadership to forecast demand and control costs while maintaining strong SLOs.

Qualifications

  • Production systems in a dev-ops environment.
  • Proficiency in Python and SQL with well-tested code.
  • Hands-on experience with at least one major cloud provider (AWS, GCP, or Azure).
  • Experience with Prometheus, PromQL, Grafana.

Responsibilities

  • Build data pipelines ingesting occupancy and utilization telemetry from Kubernetes clusters.
  • Develop real-time tooling for fleet health, capacity planning and alerting.
  • Measure and improve utilization of training, inference and evaluation workloads.
  • Reconcile cloud provider billing exports with internal telemetry for spend attribution.

Skills

DevOps
Python
SQL
Cloud platforms
Observability

Education

Bachelor's degree

Tools

Prometheus
Grafana
BigQuery

Job description

Capacity Engineer

Anthropic’s Capacity Engineering team manages one of the largest and fastest‑growing infrastructure fleets, spanning multiple accelerator families, CPU families and cloud providers. The engineer builds production systems that ingest telemetry, provide observability, and measure utilization of compute resources.

Key Responsibilities
  • Data Platform: Build data pipelines that ingest occupancy and utilization telemetry from Kubernetes clusters, normalize billing and usage across cloud providers, and serve the resulting tables in BigQuery for analysis by researchers, finance, and leadership. Ensure correctness, completeness and low latency of the data.
  • Planning: Create real‑time tooling for fleet health, capacity planning and alerting. Operate Kubernetes‑native infrastructure at scale and coordinate cross‑team scheduling efforts to reduce fragmentation and improve resource usage.
  • Efficiency: Measure and improve the utilization of training, inference and evaluation workloads. Develop benchmarking infrastructure, establish baseline metrics, and work with system‑owning teams to close gaps.
  • Attribution & Forecasting: Reconcile cloud provider billing exports with internal telemetry, attribute spend to workloads and teams, and produce defensible compute plans that survive finance review.
  • Own the planning and allocation stack used by leadership and teams for capacity allocation, adopting cross‑region and cross‑provider placement guardrails, queueing, and occupancy KPIs.
  • Drive efficiency programs such as rightsizing, unused capacity recovery, and job‑level utilization improvements.
  • Develop and maintain the BigQuery data platform, ensuring SLOs for completeness, latency and detection of data gaps.
Required Experience and Skills
  • Strong track record building and operating production systems in a dev‑ops environment.
  • Proficiency in Python and SQL; code must be idiomatic, well‑tested and maintainable.
  • Hands‑on experience with at least one major cloud provider (AWS, GCP, or Azure) and its operations.
  • Experience with observability tooling (Prometheus, PromQL, Grafana) and building monitoring that teams rely on.
  • Ability to gather requirements and work across organizational boundaries in ambiguous environments.
Preferred Qualifications
  • Capacity planning, resource management or cost attribution experience at a hyperscaler or large‑scale ML environment.
  • Product engineering experience focused on developer experience and internal data products.
  • Experience with scheduling, packing efficiency or profiling‑driven optimization of distributed workloads.
  • Multi‑cloud data ingestion expertise, including billing export normalization and reservation APIs.
  • Knowledge of total cost of ownership and forecasting, including decomposing infrastructure growth drivers.
  • Familiarity with accelerator infrastructure and GPU/TPU utilization metrics.
  • Experience building internal data products with self‑service access, schema contracts and discoverability.
  • Storage efficiency, retention, and lifecycle program expertise at scale.
Compensation

Annual Salary: $320,000 – $485,000 USD

Qualifications and Application Notes

Minimum education: Bachelor’s degree or equivalent combination of education, training, and/or experience. Field of study pertinent to the role as demonstrated through coursework, training, or professional experience.

Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position.

Location-based hybrid policy: Staff are expected to be in an Anthropic office at least 25% of the time.

Visa Sponsorship: Anthropic sponsors visas for qualified candidates. We will make every reasonable effort to obtain the appropriate visa.

Legal and Equal‑Opportunity Statement

Anthropic is an Equal Opportunity Employer and does not discriminate on the basis of race, color, religion, sex, national origin, disability, veteran status, gender identity, sexual orientation, or any other protected characteristic.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Technical Program Manager, Compute San Francisco, CA | New York City, NY | Seattle, WA
Technical Program Manager, Compute San Francisco, CA | New York City, NY | Seattle, WA

Anthropic • San Francisco (CA)

Hybrid
USD 290,000 - 365,000
Infrastructure Capacity Planner, Demand Planning
Infrastructure Capacity Planner, Demand Planning

Anthropic • San Francisco (CA)

On-site
USD 320,000 - 405,000
Senior Engineering Manager, Capacity Engineering
Senior Engineering Manager, Capacity Engineering

Anthropic • San Francisco (CA), New York (NY), Seattle (WA)

On-site
USD 210,000 - 270,000
AI Infrastructure Operations, Demand Planning
AI Infrastructure Operations, Demand Planning

Anthropic Limited • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 405,000
Product Manager, Compute Platform
Product Manager, Compute Platform

Visa Hunt • San Francisco (CA)

Hybrid
USD 305,000 - 385,000
AI Infrastructure Operations, Demand Planning
AI Infrastructure Operations, Demand Planning

Anthropic • New York (NY)

Hybrid
USD 320,000 - 405,000
Competitive compensation
Equity donation matching (optional)
Generous vacation
+2
Product Manager, Compute Platform
Product Manager, Compute Platform

Neura Market • San Francisco (CA)

Hybrid
USD 305,000 - 385,000
Competitive compensation
Optional equity donation matching
Generous vacation and parental leave
+2
Staff+ Software Engineer, Kubernetes Platform
Staff+ Software Engineer, Kubernetes Platform

Anthropic • New York (NY), Seattle (WA), San Francisco (CA)

On-site
USD 320,000 - 405,000
Equity donation matching
Generous vacation
Flexible working hours
+1
Staff Software Engineer, Node Infra San Francisco, CA | New York City, NY | Seattle, WA
Staff Software Engineer, Node Infra San Francisco, CA | New York City, NY | Seattle, WA

Anthropic • San Francisco (CA)

Hybrid
USD 405,000 - 485,000
Competitive compensation
Generous vacation
Flexible working hours
Staff Engineer, Datacenter Server Lifecycle San Francisco, CA | New York City, NY | Seattle, WA
Staff Engineer, Datacenter Server Lifecycle San Francisco, CA | New York City, NY | Seattle, WA

Anthropic • San Francisco (CA)

Hybrid
USD 320,000 - 405,000