Lead Node Infra Engineer — AI Compute & Cloud

SignalAI

New York (NY)

Hybrid

USD 405,000 - 485,000

Full time

12 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Anthropic is seeking a Node Infra Lead to own the technical strategy and lifecycle for accelerator compute across multiple clouds. You will steer ingestion, bring-up, health checks, and automated repair for a large AI compute fleet powering Claude.

You’ll collaborate with research, inference, and product teams to shape long-term compute, data, and infrastructure strategy, while driving incident response and operational excellence across the team.

Qualifications

  • Minimum: deep expertise in distributed systems and cloud platforms.
  • Proficient in at least one systems language (e.g., Rust, Go, Python).
  • Hands-on experience with ML accelerators (GPUs/TPUs/Trainium).
  • Proven ability to lead multi-quarter initiatives across teams.
  • Strong stakeholder alignment and communication.

Responsibilities

  • Own the technical strategy and roadmap for node lifecycle management - ingestion, bring-up, health checking, and automated repair.
  • Drive cross-team initiatives to build and scale AI clusters across multiple clouds and accelerator families.
  • Design and operate the systems that detect, isolate, and remediate unhealthy hardware automatically, driving up fleet MTBI and minimizing stranded capacity.
  • Define infrastructure architecture, ensuring the hardest problems get solved.
  • Work closely with cloud providers and internal research/inference/product teams to shape long-term compute, data, and infrastructure strategy.
  • Establish and evolve operational excellence practices (incident response, postmortem culture, on-call).
  • Support the growth of engineers around you through technical mentorship and coaching.

Skills

Distributed systems
Cloud platforms
Kubernetes
Terraform
Systems programming
Leadership
Stakeholder alignment

Education

Bachelor's degree or equivalent

Tools

Karpenter
Mesos

Job description

Anthropic is seeking a Node Infra Lead to own the technical strategy and lifecycle for accelerator compute across multiple clouds. You will steer ingestion, bring-up, health checks, and automated repair for a large AI compute fleet powering Claude.

You’ll collaborate with research, inference, and product teams to shape long-term compute, data, and infrastructure strategy, while driving incident response and operational excellence across the team.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Cloud Inference Launch Engineer for Scalable AI
Staff Cloud Inference Launch Engineer for Scalable AI

Anthropic • San Francisco (CA)

On-site
USD 320,000 - 485,000
Cloud Inference Launch Engineer — Staff/Senior
Cloud Inference Launch Engineer — Staff/Senior

3M HEALTHCARE • Seattle (WA), San Francisco (CA)

Hybrid
USD 320,000 - 485,000
Staff Node Infra Engineer - Scalable AI Clusters
Staff Node Infra Engineer - Scalable AI Clusters

Anthropic • San Francisco (CA)

On-site
USD 320,000 - 405,000
Staff+ Cloud Infrastructure Engineer for Scalable AI
Staff+ Cloud Infrastructure Engineer for Scalable AI

Anthropic • San Francisco (CA)

Hybrid
USD 320,000 - 485,000
Senior AI Infrastructure Engineer
Senior AI Infrastructure Engineer

Anthropic • New York (NY)

Hybrid
USD 180,000 - 280,000
Health insurance
Fertility benefits
Parental leave 22 weeks
+7
Senior Node Infra Engineer — Scale Clusters Flexible
Senior Node Infra Engineer — Scale Clusters Flexible

Anthropic • San Francisco (CA)

Hybrid
USD 405,000 - 485,000
Competitive compensation
Generous vacation
Flexible working hours
Staff Senior Software Engineer, Inference Deployment
Staff Senior Software Engineer, Inference Deployment

Anthropic • United States

Hybrid
USD 320,000 - 485,000
Equity donation matching
Generous vacation
Parental leave
+2
Lead AI Networking Architect for High-Performance Infra
Lead AI Networking Architect for High-Performance Infra

nscaleoperationsukltd • Seattle (WA)

On-site
USD 180,000 - 260,000
Staff Infrastructure Engineer: Scale Compute & Secure Clusters
Staff Infrastructure Engineer: Scale Compute & Secure Clusters

SignalAI • New York (NY)

On-site
USD 405,000 - 485,000
Staff Engineer, Datacenter Server Lifecycle
Staff Engineer, Datacenter Server Lifecycle

Anthropic • San Francisco (CA)

Hybrid
USD 320,000 - 405,000