Staff AI Reliability Engineer – Scale, Resilience & Observability

Mat Vin

Greater London

Ibrido

GBP 325.000 - 390.000

Tempo pieno

14 giorni+
Generatore di candidature

Distinguiti per questa posizione — genera un curriculum e una lettera di presentazione personalizzati in circa un minuto.

Supera i filtri ATS

Descrizione del lavoro

Anthropic is seeking a Staff Software Engineer in London to strengthen AI reliability engineering across serving paths, from SDK to network and infrastructure. You will lead incident response, design robust monitoring, and work with cross-team partners to ensure Claude’s reliability and performance.

The role emphasizes building scalable, high-availability systems with global reach and a strong emphasis on safety commitments and observability tooling.

Competenze

  • Bachelor’s degree or equivalent required.
  • Experience with large-scale distributed systems and reliability.
  • Experience operating large-scale model serving infrastructure.

Mansioni

  • Develop service level objectives for large language model serving systems, balancing availability and latency with development velocity.
  • Design and implement monitoring and observability systems across the token path.
  • Assist in the design and implementation of high-availability serving infrastructure across multiple regions and cloud providers.
  • Lead incident response for critical AI services, ensuring rapid recovery, thorough incident reviews, and systematic improvements.
  • Support the reliability of safeguard model serving and safety commitments.

Conoscenze

Distributed systems
Reliability engineering
Incident response
Cross-team collaboration
Communication skills

Formazione

Bachelor’s degree or equivalent

Strumenti

Model serving infra
GPUs/TPUs hardware
RDMA/InfiniBand
ML observability tools

Descrizione del lavoro

Anthropic is seeking a Staff Software Engineer in London to strengthen AI reliability engineering across serving paths, from SDK to network and infrastructure. You will lead incident response, design robust monitoring, and work with cross-team partners to ensure Claude’s reliability and performance.

The role emphasizes building scalable, high-availability systems with global reach and a strong emphasis on safety commitments and observability tooling.

Ottieni la revisione del curriculum gratis e riservata.

o trascina qui il file.

Similar jobs

Offerte di lavoro simili che vale la pena confrontare

Staff Engineer, AI Safeguards Infrastructure
Staff Engineer, AI Safeguards Infrastructure

Anthropic • Greater London

Ibrido
GBP 325.000 - 395.000
Equity donation matching
Generous vacation
Parental leave
+2
Staff AI Inference Systems Engineer
Staff AI Inference Systems Engineer

Mat Vin • Greater London

Ibrido
GBP 325.000 - 390.000
Competitive compensation
Equity donation matching
Generous vacation and parental leave
+2
London-based Staff Engineer, Observability & Profiling
London-based Staff Engineer, Observability & Profiling

Anthropic Limited • Greater London

Ibrido
GBP 120.000 - 160.000
Staff Software Engineer, AI Safety & Safeguards
Staff Software Engineer, AI Safety & Safeguards

Anthropic • York and North Yorkshire

In loco
GBP 110.000 - 160.000
Comprehensive health insurance
Fertility benefits via Carrot Fertilty
Paid parental leave 22 weeks
+1
Staff Software Engineer, AI Reliability Engineering
Staff Software Engineer, AI Reliability Engineering

Mat Vin • Greater London

In loco
GBP 325.000 - 390.000
Staff Infrastructure Engineer: Scalable AI Compute
Staff Infrastructure Engineer: Scalable AI Compute

Anthropic • Greater London

Ibrido
GBP 90.000 - 120.000
Staff Software Engineer - Scale Distributed Systems & AI
Staff Software Engineer - Scale Distributed Systems & AI

Gigs • Greater London

Ibrido
GBP 107.000 - 128.000
Stock options
Home office stipend
Learning and development budget
+1
Platform Systems Engineer: Scalable AI Infra & Observability
Platform Systems Engineer: Scalable AI Infra & Observability

OpenAI • Greater London

In loco
GBP 60.000 - 80.000
Staff Observability Engineer: Scalable Profiling & Telemetry
Staff Observability Engineer: Scalable Profiling & Telemetry

AI Chopping Block, Inc. • Greater London

Ibrido
GBP 325.000 - 390.000
Competitive compensation
Equity donation matching
Generous vacation
+3
Production Engineering Lead - AI-Driven Reliability
Production Engineering Lead - AI-Driven Reliability

Meta • City of Westminster

In loco
GBP 110.000 - 150.000