Data Center Operations Lead - Partner Site Operations

Anthropic

United States

Hybrid

USD 320,000 - 405,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Anthropic in the United States seeks a Data Center Operations leader who owns site availability, deployment milestones, and incident response across partner-operated sites. You define standards, governance rhythms, and drive performance with vendors while shaping fleet-wide playbooks.

This role requires 8+ years in data center ops, strong vendor management, and hands-on infra knowledge; experience with on-call rotations, SLAs, and coordinating multi-vendor facilities is essential.

Qualifications

  • 8+ years in data center operations with production availability accountability.
  • Experience managing vendors, MSPs, or contract workforces to measurable outcomes.
  • Hands-on depth in server, network and rack infrastructure to verify vendor claims.
  • Experience standing up operations at new sites or data halls.
  • Experience as incident command or lead-responder and communicating under ambiguity.
  • Understanding of SLA design, contracts, and EHS programs.

Responsibilities

  • Own site availability, deployment milestones, and repair turnaround with objective data.
  • Set daily/weekly vendor priorities and lead operating cadence and reviews.
  • Define deployment, break‑fix, change management, security, and EHS procedures.
  • Track vendor performance against SLAs and staffing commitments; drive corrective actions.
  • Participate in incident escalation and lead vendor response during incidents.
  • Translate engineering requirements into vendor direction and report site constraints to leadership.

Skills

Data center ops
Vendor management
Incident command
On-call readiness

Education

Bachelor's degree

Job description

About Anthropic

Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.



About the role

Anthropic's Data Center Operations (DCO) team ensures compute fleet availability through hardware and IT operations. At our partner-operated sites, this role manages the interface between Anthropic and the strategic site operations partner performing day‑to‑day data hall work. As the site lead, you own site outcomes for your assigned sites including: deployment velocity, availability, and incident response. Rather than managing operations staff directly, you provide tactical direction, set priorities, and define the standards for the vendor's on‑site teams, paired with performance oversight and ongoing operational assessment to ensure all operational commitments are met. You will define the operational processes, quality gates, and governance rhythms for partner-operated sites. Expect to build the playbook as much as you run it, not just at a site level, but defining and developing program improvements fleet‑wide.



What you’ll own


  • Operational outcomes. Own site availability, deployment milestones, and repair turnaround, verified with independent data rather than vendor self‑reporting.

  • Vendor direction. Set daily and weekly priorities and lead the operating cadence, including standups and business reviews.

  • Process definition. Author and improve procedures for deployment, break‑fix, change management, security, and EHS compliance. Analyze operational trends and standardize lessons across the program.

  • Performance management. Track vendor performance against SLAs and staffing commitments, driving corrective actions when necessary.

  • Incident response and on‑call. Participate in the incident escalation on‑call rotation. When designated Anthropic Incident Commander for a site‑specific incident, direct vendor response, own communications, and close out post‑incident actions.

  • Internal interface. Translate engineering requirements into vendor direction and communicate site constraints and risks to leadership.



Representative work


  • Leading weekly operations reviews and scorecards with vendor site leads.

  • Directing deployment surges to meet first‑compute‑online milestones.

  • Analyzing failure patterns to identify root causes and driving fixes with owners.

  • Creating break‑fix ownership matrices and training vendor teams.

  • Serving as Incident Commander for facility events and producing post‑mortems.

  • Establishing operational readiness for new data halls, including spares and security.

  • Identifying process gaps and codifying improvements as program standards.



You may be a good fit if


  • Have 8+ years of experience in data center operations (hardware, IT infrastructure, or critical facilities) as a manager, technical lead or related role, including accountability for production availability.

  • Have managed vendors, MSPs, or contract workforces to measurable outcomes: SOWs, SLAs, operational reviews, and corrective action.

  • Carry hands‑on technical depth in server, network, and rack‑level infrastructure, enough to independently verify vendor claims and audit quality.

  • Have built or substantially improved operational processes, not just run them.

  • Have served in an incident command or lead‑responder role and communicate clearly under ambiguity.

  • Can support non‑standard hours, including an on‑call rotation and availability during deployment surges and maintenance windows.

  • Bachelor's degree in relevant domain or equivalent practical experience.

  • Strong candidates may also have Experience with third‑party colocation providers or partner‑operated sites, delivering IT operations outcomes inside a facility someone else runs.

  • Experience standing up operations at a new site or data hall, from commissioning handoff through first deployment.

  • Experience with GPU/accelerator or high‑density liquid‑cooled infrastructure.

  • Familiarity with multi‑vendor sites where facilities and IT operations are performed by different partners.

  • Experience leading projects from initiation to completion across teams you didn’t own.

  • Background in incident management frameworks, contract/SLA design, or EHS programs.



Logistics


  • Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience

  • Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience

  • Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position

  • Location‑based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices.

  • Visa sponsorship: We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and every candidate. But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this.



Research shows that people who identify as being from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy, so we urge you not to exclude yourself prematurely and to submit an application if you're interested in this work. We think AI systems like the ones we're building have enormous social and ethical implications. We think this makes representation even more important, and we strive to include a range of diverse perspectives on our team.



The annual compensation range for this role is listed below.


Annual Salary: $320,000 - $405,000 USD



How we're different

We believe that the highest‑impact AI research will be big science. At Anthropic we work as a single cohesive team on just a few large‑scale research efforts. And we value impact — advancing our long‑term goals of steerable, trustworthy AI — rather than work on smaller and more specific puzzles. We see AI research as an empirical science, which has as much in common with physics and biology as with traditional efforts in computer science. We're an extremely collaborative group, and we host frequent research discussions to ensure that we are pursuing the highest‑impact work at any given time. As such, we greatly value communication skills.



The easiest way to understand our research directions is to read our recent. This research continues many of the directions our team worked on prior to Anthropic, including: GPT‑3, Circuit‑Based Interpretability, Multimodal Neurons, Scaling Laws, AI & Compute, Concrete Problems in AI Safety, and Learning from Human Preferences.



Come work with us!



Anthropic is a public benefit corporation headquartered in SanFrancisco. We offer competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a lovely office space in which to collaborate with colleagues.



Guidance on Candidates' AI Usage: Learn about our policy for using AI in our application process.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Center Operations Lead - Partner Site Operations
Data Center Operations Lead - Partner Site Operations

Menlo Ventures • San Francisco (CA)

Hybrid
USD 320,000 - 405,000
Data Center Operations Lead - Partner Site Operations
Data Center Operations Lead - Partner Site Operations

Anthropic • San Francisco (CA)

Hybrid
USD 320,000 - 405,000
Equity donation matching
Generous vacation and parental leave
Flexible working hours
+1
Lead, Data Center Security Delivery (Construction to Operations)
Lead, Data Center Security Delivery (Construction to Operations)

Anthropic • United States

Hybrid
USD 290,000 - 365,000
Hardware Lab Manager
Hardware Lab Manager

Anthropic • New York (NY)

On-site
USD 320,000 - 405,000
Competitive pay
Generous vacation
Parental leave
+3
Lead, Data Center Security Delivery (Construction to Operations)
Lead, Data Center Security Delivery (Construction to Operations)

Anthropic • San Francisco (CA)

On-site
USD 219,000 - 365,000
Software Engineer, Research Infrastructure
Software Engineer, Research Infrastructure

Anthropic • New York (NY)

Hybrid
USD 405,000 - 625,000
Competitive compensation
Equity donation matching
Generous vacation and parental leave
+2
Staff+ Software Engineer, Storage + Transfer
Staff+ Software Engineer, Storage + Transfer

Anthropic • New York (NY)

Hybrid
USD 180,000 - 260,000
Equity donation matching
Generous vacation and parental leave
Flexible working hours
+1
Staff+ Software Engineer, Infrastructure (Distributed Systems)
Staff+ Software Engineer, Infrastructure (Distributed Systems)

Anthropic • New York (NY)

Hybrid
USD 320,000 - 485,000
Competitive compensation
Equity donation matching (optional)
Generous vacation and parental leave
+2
Data Center Supply Planning Lead
Data Center Supply Planning Lead

Anthropic • United States

Hybrid
USD 320,000 - 405,000
Technical Program Manager (Infrastructure)
Technical Program Manager (Infrastructure)

Anthropic • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive compensation
Generous vacation
Flexible working hours
+1