Senior AI Reliability Engineer – Model Serving

Humanloop

Greater London

On-site

GBP 57,000 - 73,000

Full time

4 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Competitive compensation and benefits
Equity donation matching (optional)
Generous vacation and parental leave
Flexible working hours
Hybrid policy: in-office presence 25%+
Visa sponsorship where possible

Job summary

Anthropic in the UK seeks an experienced reliability-focused engineer to own large-scale AI serving infrastructure. You will define SLAs, build robust monitoring, and drive fast incident response across multiple regions. This role emphasizes cross-team collaboration, ownership, and impact in a highly technical environment.

You should have a strong background in distributed systems, production engineering, or SRE, with a focus on reliability in AI workloads and large-scale model serving.

Qualifications

  • Bachelor's degree or equivalent required.
  • Experience with distributed systems and reliability concepts.
  • Strong communication and collaboration skills; ownership mindset.

Responsibilities

  • Define and monitor service level objectives for AI serving systems.
  • Design and implement monitoring and observability across token paths.
  • Help build high-availability serving infrastructure across regions/clouds.
  • Lead incident response for critical AI services and drive improvements.
  • Support reliability of safeguard model serving and safety commitments.

Skills

Distributed systems
Reliability engineering
SRE
Communication
Collaboration
Ownership
Open-source contributions
Incident response

Education

Bachelor's degree
Field relevant to role

Tools

AI observability tools
Chaos engineering tooling
Observability frameworks
Large-scale model serving infrastructure
GPU/ML accelerators

Job description

Anthropic in the UK seeks an experienced reliability-focused engineer to own large-scale AI serving infrastructure. You will define SLAs, build robust monitoring, and drive fast incident response across multiple regions. This role emphasizes cross-team collaboration, ownership, and impact in a highly technical environment.

You should have a strong background in distributed systems, production engineering, or SRE, with a focus on reliability in AI workloads and large-scale model serving.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Software Engineer AI Reliability Engineering
Staff Software Engineer AI Reliability Engineering

Humanloop • Greater London

On-site
GBP 57,000 - 73,000
Competitive compensation and benefits
Equity donation matching (optional)
Generous vacation and parental leave
+3
Senior AI Reliability Engineer: LLM Serving & Resilience
Senior AI Reliability Engineer: LLM Serving & Resilience

Anthropic • Greater London

Hybrid
GBP 120,000 - 180,000
Competitive compensation
Flexible working hours
Generous vacation and parental leave
Production Engineering Lead - AI-Driven Reliability
Production Engineering Lead - AI-Driven Reliability

Meta • City of Westminster

On-site
GBP 110,000 - 150,000
Senior Software Engineer, AI Reliability Engineering London, UK
Senior Software Engineer, AI Reliability Engineering London, UK

Anthropic • Greater London

Hybrid
GBP 120,000 - 180,000
Competitive compensation
Flexible working hours
Generous vacation and parental leave
Infra Engineer: AI-Powered Reliability & Scale
Infra Engineer: AI-Powered Reliability & Scale

Neura Market • Greater London

Hybrid
GBP 89,160 - 133,740
Generous PTO
Health & dental insurance
Parental leave
+7
Senior AI Platform SRE Engineer
Senior AI Platform SRE Engineer

Amach • Greater London

Hybrid
GBP 70,000 - 90,000
Flexible working environment
Learning and development opportunities
Options for career advancement
+1
Backend API Engineer for Scalable AI Model Serving
Backend API Engineer for Scalable AI Model Serving

Xai • Greater London

On-site
GBP 107,000 - 262,000
Equity
Healthcare coverage
Dental coverage
+4
Senior AI Engineer - Cloud-Native ML & API Microservices
Senior AI Engineer - Cloud-Native ML & API Microservices

Infused Solutions Ltd • Greater London

Hybrid
GBP 65,000 - 75,000
Senior AI Engineer - Hybrid, Client-Facing & Platform
Senior AI Engineer - Hybrid, Client-Facing & Platform

develop • Greater London

Hybrid
GBP 42,000 - 70,000
Share options
Discretionary bonus
Senior AI Infrastructure Engineer – AWS, Terraform, Multi-Region
Senior AI Infrastructure Engineer – AWS, Terraform, Multi-Region

Visa Hunt • Greater London

Hybrid
GBP 90,000 - 140,000