Stand out for this role — generate a tailored resume and cover letter in about a minute.
Anthropic is seeking a Repairs Lead to own and scale the end-to-end hardware repair program across our data centers. You will define SLAs, manage backlogs, and partner with OEMs/ODMs to maintain high availability of servers, GPUs, networks, and optics.
You will develop scalable repair processes, train site teams, and drive data-driven improvements to reduce downtime and optimize spare parts, returns, and workflow across multiple sites.
Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.
As a Repairs Lead within the Data Center Infrastructure organization, you will define and manage the end-to-end hardware repair program across Anthropic's growing fleet of data centers. You will be accountable for repair turnaround time and the compute returned to service across every site, covering server, GPU/accelerator, network, and optics break-fix, RMA and reverse logistics with OEMs and ODMs, and the spares and repair inventory that keeps repair SLAs achievable. As a subject matter expert in hardware operations, you will develop scalable repair processes and quality targets, set the standards that site operations partners and repair vendors execute against, and turn failure trends into fixes which are driven upstream with engineering and equipment partners. If you are experienced in at-scale datacenter hardware operations, are passionate about HPC data centers, enjoy working in complex, fast-paced environments, we welcome you to apply.