Staff Engineer, Datacenter Server Lifecycle

Anthropic

New York (NY)

Hybrid

USD 320,000 - 405,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Anthropic seeks a Staff Engineer for the Datacenter Server Lifecycle team in New York, NY, to manage server lifecycle from provisioning to decommissioning while ensuring security standards. Candidates should have at least 8 years of datacenter operations experience and expertise in managing sensitive AI workloads. The role offers a base salary between $320,000 and $405,000 USD, requiring a bachelor's degree. Expect at least 25% office attendance under a hybrid policy, along with potential travel to datacenter sites.

Qualifications

  • Hands-on experience with server hardware, including troubleshooting and deployment.
  • End-to-end understanding of hardware lifecycle management.
  • Proficiency in at least one programming language (Python, Rust, Go, Java).

Responsibilities

  • Lead automation build-out for datacenters with thousands of servers.
  • Define server lifecycle strategy from provisioning to decommissioning.
  • Work with Infrastructure Security on trusted compute standards.

Skills

Server hardware experience
Lifecycle management understanding
Programming proficiency
Cloud infrastructure knowledge
Communication skills
Problem-solving abilities
Willingness to travel

Education

Bachelor's degree or equivalent

Tools

Kubernetes
AWS
GCP

Job description

Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.

About The Role

Anthropic is expanding beyond cloud infrastructure, and this role sits at the heart of that effort. As a Staff Engineer on the Datacenter Server Lifecycle team, you will own the end-to-end operational journey of every machine in our facility — from initial provisioning and deployment, across its working life, through maintenance and refresh, and all the way to decommissioning. This is greenfield work: you will help define the processes, tooling, and operational standards that govern how we run and retire hardware at scale.

A distinguishing aspect of this role is its deep intersection with security. The machines in our datacenter handle some of the most sensitive workloads in AI — training frontier models and serving millions of users interacting with Claude. Ensuring that every machine in the fleet is trusted, attested, and operating with a verified chain of integrity from the hardware up is a core part of the job, not an afterthought. You will partner closely with our Infrastructure Security team to define and enforce trusted compute standards across the lifecycle, from secure provisioning through end‑of‑life handling.

Key Responsibilities
  • Lead the build‑out of automation to support datacenters containing tens of thousands of servers.
  • Define and own the end‑to‑end server lifecycle strategy — from provisioning and deployment through operation, maintenance, refresh, and decommissioning — and maintain automation and operational procedures for common lifecycle events (e.g., hardware failures, firmware upgrades, fleet rotations).
  • Partner closely with Infrastructure Security to design and enforce trusted compute standards across the server lifecycle.
  • Work closely with our Networking team to ensure end‑to‑end connectivity across all sites.
  • Build and maintain tooling to track machine health, configuration, and operational status across the full datacenter fleet.
Minimum Qualifications
  • Hands‑on experience with server hardware, including rack deployment, cabling, troubleshooting, and understanding failure modes at scale.
  • End‑to‑end understanding of hardware lifecycle management: asset tracking, provisioning workflows, maintenance scheduling, and decommissioning practices.
  • Proficiency in at least one programming language (e.g., Python, Rust, Go, or Java).
  • Working knowledge of modern cloud infrastructure, including Kubernetes, Infrastructure as Code, AWS, and GCP.
  • Ability to communicate clearly and build consensus with a wide range of stakeholders.
  • Comfort navigating ambiguity and making progress on complex, cross‑functional problems.
  • Willingness to travel occasionally to datacenter sites across North America.
Preferred Qualifications
  • 8+ years of experience in datacenter operations, hardware infrastructure management, or a closely related discipline.
  • Hands‑on experience with GPU or AI accelerator hardware (e.g., NVIDIA A100/H100, AMD MI300, Google TPUs, or AWS Trainium) and an understanding of their operational demands.
  • Familiarity with modern provisioning tooling such as coreboot, LinuxBoot, or u‑root.
  • Experience building or contributing to datacenter automation or fleet management platforms.
  • Experience building and deploying server operating system distributions across thousands of hosts.
  • Background in large‑scale capacity planning and hardware refresh strategy, ideally at a hyperscaler or large cloud provider.
  • Experience with trusted compute and hardware security concepts such as secure boot, TPM, hardware attestation, and firmware verification — or a strong desire to develop deep expertise in this area.
Annual Salary

$320,000—$405,000 USD

Logistics

Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience.

Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience.

Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position.

Location‑based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices.

Visa sponsorship: We sponsor visas. If an offer is made, we will make every reasonable effort to obtain a visa for the candidate.

We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed. We strive to include a range of diverse perspectives on our team.

Guidance on Candidates' AI Usage: Learn about our policy for using AI in our application process.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Engineer, Datacenter Server Lifecycle
Staff Engineer, Datacenter Server Lifecycle

Anthropic • San Francisco (CA)

On-site
USD 320,000 - 405,000
Staff+ Infrastructure Engineer, Cluster Infrastructure
Staff+ Infrastructure Engineer, Cluster Infrastructure

Anthropic • Seattle (WA)

On-site
USD 320,000 - 405,000
Competitive compensation and benefits
Generous vacation and parental leave
Flexible working hours
+1
Data Center Operations Lead - Partner Site Operations
Data Center Operations Lead - Partner Site Operations

Anthropic • San Francisco (CA)

On-site
USD 320,000 - 405,000
Staff+ Software Engineer, Developer Productivity
Staff+ Software Engineer, Developer Productivity

Anthropic • Seattle (WA)

On-site
USD 405,000 - 625,000
Staff+ Software Engineer, Developer Productivity
Staff+ Software Engineer, Developer Productivity

Anthropic • San Francisco (CA)

On-site
USD 405,000 - 625,000
Staff+ Software Engineer, Claude App Infrastructure
Staff+ Software Engineer, Claude App Infrastructure

Anthropic • New York (NY)

On-site
USD 320,000 - 485,000
Technical Program Manager, Data Center Infrastructure
Technical Program Manager, Data Center Infrastructure

anthropic • New York (NY), Seattle (WA), San Francisco (CA)

On-site
USD 365,000 - 435,000
Competitive compensation
Generous vacation and parental leave
Flexible working hours
Data Center Engineer, Reliability & Infrastructure Management – Compute Supply
Data Center Engineer, Reliability & Infrastructure Management – Compute Supply

Anthropic • San Francisco (CA)

Hybrid
USD 320,000 - 405,000
Data Center Engineer, Reliability & Infrastructure Management – Compute Supply
Data Center Engineer, Reliability & Infrastructure Management – Compute Supply

Anthropic • New York (NY)

On-site
USD 320,000 - 405,000
Competitive compensation
Equity donation matching
Generous vacation
+3
Staff+ Software Engineer, Databases
Staff+ Software Engineer, Databases

Anthropic • Seattle (WA)

On-site
USD 320,000 - 485,000