Platform Reliability Engineer for AI Infrastructure

Callosum Technologies Ltd.

Greater London

Hybrid

GBP 90,000 - 130,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Equity
Private healthcare
Visa sponsorship
Relocation assistance
London office

Job summary

Callosum Technologies Ltd. in London is seeking an experienced Platform Reliability Engineer to own the production health of our platform, defining SLOs, observability, incident response and capacity planning across heterogeneous compute backends.

You will collaborate with platform, hardware and orchestration teams to expose backends reliably, build a blameless culture, and drive reliability improvements at scale.

Qualifications

  • Strong SRE or production-engineering background running customer-facing systems at scale.
  • Fluency with modern operational tooling: observability stacks, container orchestration, infrastructure-as-code, CI/CD.
  • Experience owning incident response and driving reliability improvements.
  • Compliance execution experience.

Responsibilities

  • Define service-level objectives, monitoring, alerting and observability for a production platform.
  • Lead on-call and incident response with runbooks, escalation, blameless postmortems, and follow-through.
  • Perform capacity planning and operate across heterogeneous compute backends.

Skills

SRE background
Production engineering
Incident response
Reliability improvements
Capacity planning
Open-source contributions
AI-native / Early-stage experience

Tools

Observability stacks
Container orchestration (Kubernetes)
Infrastructure as code (Terraform)
CI/CD pipelines
Compliance tooling

Job description

Callosum Technologies Ltd. in London is seeking an experienced Platform Reliability Engineer to own the production health of our platform, defining SLOs, observability, incident response and capacity planning across heterogeneous compute backends.

You will collaborate with platform, hardware and orchestration teams to expose backends reliably, build a blameless culture, and drive reliability improvements at scale.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE for AI Platform Reliability & Observability
SRE for AI Platform Reliability & Observability

Callosum • Greater London

On-site
GBP 90,000 - 150,000
Equity & Ownership
Private healthcare
Visa sponsorship & relocation
+1
Platform Reliability Director
Platform Reliability Director

LinuxRecruit • Greater London

On-site
GBP 120,000 - 180,000
Senior Cloud SRE for AI Platform — Reliability & Scale
Senior Cloud SRE for AI Platform — Reliability & Scale

Mistral • Greater London

On-site
GBP 90,000 - 140,000
Healthcare coverage
Parental leave
Retirement plans
+3
Platform Reliability Engineer – Cloud Infrastructure
Platform Reliability Engineer – Cloud Infrastructure

Talenzon group • Greater London

On-site
GBP 60,000 - 80,000
Senior Platform Engineer & SRE for Cloud Reliability
Senior Platform Engineer & SRE for Cloud Reliability

Myn • Greater London

Hybrid
GBP 90,000 - 120,000
Platform Operations Lead – Java & AWS Reliability
Platform Operations Lead – Java & AWS Reliability

develop • Greater London

Hybrid
GBP 72,000 - 88,000
Hybrid working (3 days per week in the
Annual performance bonus
Career progression opportunities
Founding Applied AI SRE Engineer for Platform Reliability
Founding Applied AI SRE Engineer for Platform Reliability

Mistral • Greater London

On-site
GBP 90,000 - 140,000
Healthcare
Relocation
Wellness program
+1
Lead Site Reliability Engineer – Resilience & Automation
Lead Site Reliability Engineer – Resilience & Automation

Addition • England

On-site
GBP 90,000 - 120,000
Annual bonus
Pension up to 8.5%
26 days leave + holidays
+5
Senior Platform Engineer: Equity, Scale & Reliability
Senior Platform Engineer: Equity, Scale & Reliability

Attio • United Kingdom

On-site
GBP 105,000 - 125,000
Equity in early-stage company
25 days holiday
Apple hardware
+4
Senior Platform Engineer / SRE
Senior Platform Engineer / SRE

Myn • Greater London

Hybrid
GBP 90,000 - 120,000