Staff Software Engineer AI Reliability Engineering

Humanloop

Greater London

On-site

GBP 57,000 - 73,000

Full time

4 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Competitive compensation and benefits
Equity donation matching (optional)
Generous vacation and parental leave
Flexible working hours
Hybrid policy: in-office presence 25%+
Visa sponsorship where possible

Job summary

Anthropic in the UK seeks an experienced reliability-focused engineer to own large-scale AI serving infrastructure. You will define SLAs, build robust monitoring, and drive fast incident response across multiple regions. This role emphasizes cross-team collaboration, ownership, and impact in a highly technical environment.

You should have a strong background in distributed systems, production engineering, or SRE, with a focus on reliability in AI workloads and large-scale model serving.

Qualifications

  • Bachelor's degree or equivalent required.
  • Experience with distributed systems and reliability concepts.
  • Strong communication and collaboration skills; ownership mindset.

Responsibilities

  • Define and monitor service level objectives for AI serving systems.
  • Design and implement monitoring and observability across token paths.
  • Help build high-availability serving infrastructure across regions/clouds.
  • Lead incident response for critical AI services and drive improvements.
  • Support reliability of safeguard model serving and safety commitments.

Skills

Distributed systems
Reliability engineering
SRE
Communication
Collaboration
Ownership
Open-source contributions
Incident response

Education

Bachelor's degree
Field relevant to role

Tools

AI observability tools
Chaos engineering tooling
Observability frameworks
Large-scale model serving infrastructure
GPU/ML accelerators

Job description

Salary: £57,000 - 73,000 per year

Requirements
  • We are looking for strong distributed systems, infrastructure, or reliability backgrounds, including reliability-minded software engineers and SREs.
  • We value candidates who are curious and brave, and who are comfortable jumping into unfamiliar systems during an incident to help drive resolution.
  • We look for people who think holistically about how systems compose and where the seams are.
  • We need someone who can build lasting relationships across teams and work as a trusted teammate.
  • We value people who care about users and feel ownership over outcomes, even for systems they do not own.
  • Excellent communication and collaboration skills are important, as you will partner across the company.
  • We value diverse experience across product stacks, databases, distributed systems, and related technical domains.
  • A bachelors degree or an equivalent combination of education, training, and/or experience is required.
  • A field relevant to the role through coursework, training, or professional experience is required.
  • Years of experience should align with the internal job level requirements for the position.
  • Strong candidates may also have experience as an SRE, Production Engineer, or in similar reliability-focused roles on large-scale systems.
  • Experience operating large-scale model serving or training infrastructure, including environments with more than 1,000 GPUs, is a plus.
  • Experience with ML hardware accelerators such as GPUs, TPUs, or Trainium is a plus.
  • Knowledge of ML-specific networking optimizations such as RDMA and InfiniBand is a plus.
  • Expertise in AI-specific observability tools and frameworks is a plus.
  • Experience with chaos engineering and systematic resilience testing is a plus.
  • Contributions to open-source infrastructure or ML tooling are a plus.
Responsibilities
  • We develop appropriate Service Level Objectives for large language model serving systems, balancing availability and latency with development velocity.
  • We design and implement monitoring and observability systems across the token path.
  • We assist in the design and implementation of high-availability serving infrastructure across multiple regions and cloud providers.
  • We lead incident response for critical AI services, ensuring rapid recovery, thorough incident reviews, and systematic improvements.
  • We support the reliability of safeguard model serving, which is critical for both site reliability and our safety commitments.
Technologies
  • AI
  • API
  • Cloud
  • Hardware
  • InfiniBand
  • Support
  • Model Serving
  • Network
  • RDMA
Benefits
  • Competitive compensation and benefits
  • Optional equity donation matching
  • Generous vacation and parental leave
  • Flexible working hours
  • A lovely office space in which to collaborate with colleagues
  • A hybrid policy with in-office presence expected at least 25% of the time
  • We also sponsor visas where possible

We encourage applicants from underrepresented groups to apply.

More

We are Anthropic, a public benefit corporation headquartered in San Francisco, and our mission is to create reliable, interpretable, and steerable AI systems that are safe and beneficial for our users and for society. Our AIRE (AI Reliability Engineering) team works across Anthropic to improve reliability across our most critical serving paths, from the SDK through our network, API layers, serving infrastructure, and accelerators and back. We work in a highly collaborative environment with dynamic, cross-cutting exposure to the systems that matter most, and we value communication, teamwork, and impact. We offer competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, a lovely office space in which to collaborate with colleagues, and a hybrid policy with in-office presence expected at least 25% of the time. We also sponsor visas where possible.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer, AI Reliability Engineering London, UK
Senior Software Engineer, AI Reliability Engineering London, UK

Anthropic • Greater London

Hybrid
GBP 120,000 - 180,000
Competitive compensation
Flexible working hours
Generous vacation and parental leave
Staff Software Engineer, Infrastructure (Distributed Systems)
Staff Software Engineer, Infrastructure (Distributed Systems)

Menlo Ventures • Greater London

Hybrid
GBP 325,000 - 390,000
Staff+ Software Engineer, Safeguards Infrastructure
Staff+ Software Engineer, Safeguards Infrastructure

Anthropic • Greater London

Hybrid
GBP 325,000 - 395,000
Equity donation matching
Generous vacation
Parental leave
+2
Staff Software Security Engineer
Staff Software Security Engineer

Jackalope Digital LLC • Greater London

Hybrid
GBP 255,000 - 325,000
Staff Software Engineer, Infrastructure (Distributed Systems)
Staff Software Engineer, Infrastructure (Distributed Systems)

AI Chopping Block • Greater London

Hybrid
GBP 325,000 - 390,000
Competitive compensation
Equity donation matching
Generous vacation
+2
Staff Software Engineer, Infrastructure (Distributed Systems)
Staff Software Engineer, Infrastructure (Distributed Systems)

AI Startups UK • Greater London

Hybrid
GBP 325,000 - 390,000
Competitive compensation
Equity donation matching (optional)
Generous vacation and parental leave
+1
Staff Software Engineer, Observability & Profiling London, UK
Staff Software Engineer, Observability & Profiling London, UK

Anthropic Limited • Greater London

Hybrid
GBP 120,000 - 160,000
Staff Software Security Engineer
Staff Software Security Engineer

Anthropic • City Of London

Hybrid
GBP 270,000 - 325,000
Staff Software Engineer, Observability & Profiling
Staff Software Engineer, Observability & Profiling

AI Startups UK • Greater London

Hybrid
GBP 325,000 - 390,000
Office space
Flexible working hours
Generous vacation
+2
Staff Software Engineer, Observability & Profiling
Staff Software Engineer, Observability & Profiling

AI Chopping Block • Greater London

Hybrid
GBP 325,000 - 390,000
Competitive compensation
Equity donation matching
Generous vacation
+3