Engineering Manager, Accelerator Platform

Anthropic

San Francisco (CA)

Hybrid

USD 405,000 - 485,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity donation matching
Generous vacation and parental leave
Flexible working hours
Collaborative office space

Job summary

Anthropic is looking for an Engineering Manager for the Accelerator Platform team in San Francisco. In this role, you will lead engineering efforts for hardware integrations, manage the team, and define the technical direction for new accelerator platforms.

The ideal candidate will have significant experience in infrastructure management, a strong understanding of ML infrastructure, and a Bachelor's degree in a related field. This role requires excellent cross-functional collaboration and strategic planning for hardware roadmaps.

Qualifications

  • 3+ years in engineering management experience.
  • Deep understanding of distributed systems and hardware/software co-design.
  • Experience managing diverse compute infrastructure at scale.

Responsibilities

  • Lead the Accelerator Platform team and manage engineering talent.
  • Oversee bring-up lifecycle for new accelerator platforms.
  • Ensure integration with Anthropic's inference serving stack.

Skills

Infrastructure management
Technical fluency in systems programming
Experience with heterogeneous compute infrastructure
Cross-functional relationship building
Strategic thinking on hardware roadmaps

Education

Bachelor's degree in a related field

Tools

Kubernetes
ML accelerator architectures

Job description

About Anthropic

Anthropic's mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.

About the Role

Every time someone talks to Claude—through the API, claude.ai, our cloud partners, or any of our expanding surfaces—the request lands on an AI accelerator. Not one kind, many kinds: TPUs, Trainium chips, GPUs. Each arrives with its own software stack, performance characteristics, failure modes, and operational quirks. Someone has to take raw silicon and turn it into a platform that the rest of Anthropic can build on without thinking about which chip is underneath. That's us.

The Accelerator Platform team owns the bringup and normalization of new hardware platforms for Anthropic's first party inference fleet. We sit between the low-level systems teams and the serving infrastructure that runs production inference—bridging the gap so that every new accelerator generation ships as a first‑class production platform. It's deeply technical work at the intersection of hardware enablement, distributed systems, and ML infrastructure, and it is directly on the critical path for Anthropic's compute strategy.

We're hiring an Engineering Manager to build and lead this team. You'll inherit a small nucleus of experienced engineers and grow it into a standalone platform organization. You'll set technical direction, hire a strong team, and partner closely with hardware vendors, cloud providers, and teams across Inference to bring new accelerator generations online quickly and reliably.

Responsibilities
  • Build and lead the Accelerator Platform team—hiring, developing, and retaining engineers who thrive at the hardware/software boundary
  • Own the end‑to‑end bring‑up lifecycle for new accelerator platforms (multiple generations of Trainium, TPUs, and GPUs), from initial silicon availability through production‑ready inference
  • Define and drive the platform normalization layer—ensuring new hardware integrates cleanly with Anthropic's inference serving stack to provide a consistent abstraction
  • Partner with cloud providers (AWS, GCP, Microsoft Azure) and chip vendors on hardware roadmaps, capacity planning, and platform‑specific technical challenges
  • Collaborate closely with teams across Inference and Infrastructure to ensure new platforms meet production reliability and latency requirements from day one
  • Contribute to Anthropic's multi‑cloud compute strategy—helping the organization maintain optionality across accelerator families and avoid lock‑in to any single vendor
  • Manage the team's priorities across competing demands: new platform bring‑up, ongoing production support for existing platforms, and longer‑term investments in tooling and automation.
You may be a good fit if you
  • Have significant experience managing infrastructure or platform engineering teams (3+ years in engineering management)
  • Have deep technical fluency in systems programming, distributed systems, or hardware/software co‑design—you need to understand the stack deeply enough to make sound technical and hiring decisions
  • Have experience bringing up or operating heterogeneous compute infrastructure at scale—whether that's GPU clusters, TPU pods, custom ASICs, or FPGA deployments.
  • Are comfortable with ambiguity and can build structure where none exists. This team is being carved out as a new entity; you'll be defining its charter, processes, and culture from scratch
  • Think strategically about hardware roadmaps and can translate vendor capabilities into engineering plans
  • Build strong cross‑functional relationships—this role requires tight collaboration with hardware vendors, cloud partners, and half a dozen internal teams
  • Care deeply about both technical excellence and the people doing the work.
Strong candidates may also
  • Have direct experience with ML accelerator architectures (GPU/CUDA, TPU/XLA, Trainium/Neuron, or similar)
  • Have worked on ML inference serving infrastructure at scale (1000+ accelerators)
  • Have experience with Kubernetes‑based ML workload orchestration
  • Understand ML‑specific networking (RDMA, InfiniBand, NVLink, ICI) and how interconnect topology affects serving performance
  • Have experience managing vendor relationships and influencing hardware/software roadmaps
  • Have led teams through rapid growth phases (hiring 5+ engineers in a short timeframe).
Compensation

Annual Salary: $405,000 - $485,000 USD

Logistics

Education requirements: We require at least a Bachelor's degree in a related field or equivalent experience.

Location‑based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices.

Visa sponsorship: We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and every candidate. But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this.

Benefits

We offer competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a lovely office space in which to collaborate with colleagues.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff+ Software Engineer (Inference Runtime)
Staff+ Software Engineer (Inference Runtime)

jobr.pro • San Francisco (CA)

Hybrid
USD 405,000 - 485,000
Competitive compensation
Generous vacation
Flexible working hours
Staff+ Software Engineer, Inference Runtime
Staff+ Software Engineer, Inference Runtime

Anthropic • San Francisco (CA)

Hybrid
USD 405,000 - 485,000
Competitive compensation
Flexible working hours
Generous vacation and parental leave
Engineering Manager — Accelerator Platform Lead
Engineering Manager — Accelerator Platform Lead

Anthropic • San Francisco (CA)

On-site
Staff+ Software Engineer, Inference Runtime
Staff+ Software Engineer, Inference Runtime

Anthropic • New York (NY)

Hybrid
USD 405,000 - 485,000
Generous vacation
Parental leave
Flexible working hours
+1
Staff+ Software Engineer, Inference Runtime
Staff+ Software Engineer, Inference Runtime

Menlo Ventures • New York (NY)

Hybrid
USD 405,000 - 485,000
Staff+ Software Engineer, Inference Runtime
Staff+ Software Engineer, Inference Runtime

Anthropic • Seattle (WA)

Hybrid
USD 405,000 - 485,000
Staff+ Software Engineer, Inference Runtime Remote-Friendly (Travel-Required) | San Francisco, [...]
Staff+ Software Engineer, Inference Runtime Remote-Friendly (Travel-Required) | San Francisco, [...]

Anthropic • New York (NY)

Hybrid
USD 405,000 - 485,000
Staff Software Engineer, Node Infra San Francisco, CA | New York City, NY | Seattle, WA
Staff Software Engineer, Node Infra San Francisco, CA | New York City, NY | Seattle, WA

Anthropic • San Francisco (CA)

Hybrid
USD 405,000 - 485,000
Competitive compensation
Generous vacation
Flexible working hours
Engineering Manager, GPU (ML Accelerator)
Engineering Manager, GPU (ML Accelerator)

Anthropic • New York (NY)

On-site
USD 425,000 - 560,000
Competitive compensation
Flexible working hours
Generous vacation and parental leave
+2
Technical Program Manager, Hardware Systems
Technical Program Manager, Hardware Systems

Anthropic • San Francisco (CA)

Hybrid
USD 365,000 - 435,000