Senior AI Platform SRE: Reliability & Automation

CloudFactory

Reading

On-site

GBP 70,000 - 110,000

Full time

8 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

CloudFactory is seeking a Site Reliability Engineer to keep production systems reliable, scalable, and secure. You will work with engineers and operators to fuse engineering, operation, and security for platform and service excellence.

The role emphasizes building golden paths, developer tooling, and end-to-end software delivery governance to support ML/LLM workloads and cloud deployments. This is a chance to grow in a mission-driven, globally connected environment.

Qualifications

  • 5+ years in infrastructure engineering, DevOps, or SRE.
  • Experience with Kubernetes in large-scale production.
  • Proficiency in Python or Go for automation.
  • Experience with AI/ML workloads on Kubernetes is a plus.
  • Strong collaboration across product, backend, and frontend teams.

Responsibilities

  • Ensure reliability of platform including ML/LLM workloads and infrastructure.
  • Implement observability and tracing for ML models and services.
  • Contribute to company-wide technical direction and golden paths.
  • Develop reusable tooling and automation to accelerate delivery.
  • Package common open-source tools (Grafana, Istio, CloudNative stack, ML tooling).
  • Embed security, compliance, and cost governance into platform design.

Skills

Kubernetes
Python
Go
Terraform
Helm
CloudFormation
AI/ML tooling
Observability

Tools

Terraform
CloudFormation
Helm

Job description

CloudFactory is seeking a Site Reliability Engineer to keep production systems reliable, scalable, and secure. You will work with engineers and operators to fuse engineering, operation, and security for platform and service excellence.

The role emphasizes building golden paths, developer tooling, and end-to-end software delivery governance to support ML/LLM workloads and cloud deployments. This is a chance to grow in a mission-driven, globally connected environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

CloudFactory • Reading

On-site
GBP 70,000 - 110,000
Senior Cloud SRE: Scale AI Platform & Reliability
Senior Cloud SRE: Scale AI Platform & Reliability

Mistral AI • Greater London

On-site
GBP 75,000 - 110,000
Healthcare coverage
Relocation support
Retirement plans
+3
Senior SRE — Build Reliable, Scalable Platforms
Senior SRE — Build Reliable, Scalable Platforms

Deepl-Se • Greater London

On-site
GBP 90,000 - 130,000
Senior Cloud SRE for AI Platform — Reliability & Scale
Senior Cloud SRE for AI Platform — Reliability & Scale

Mistral • Greater London

On-site
GBP 90,000 - 140,000
Healthcare coverage
Parental leave
Retirement plans
+3
Founding Cloud SRE — AI/ML Platform & GPU Compute
Founding Cloud SRE — AI/ML Platform & GPU Compute

Icehouseventures • Greater London

Hybrid
GBP 70,000 - 90,000
SRE Associate: Build Reliable Cloud Platforms
SRE Associate: Build Reliable Cloud Platforms

WeAreTechWomen • Birmingham

On-site
GBP 65,000 - 90,000
Senior Cloud SRE: Build Reliable, Scalable Platforms
Senior Cloud SRE: Build Reliable, Scalable Platforms

Nice • Greater London

Hybrid
GBP 70,000 - 110,000
NICE-FLEX hybrid model
SRE for AI-Driven Financial Infrastructure
SRE for AI-Driven Financial Infrastructure

United States Digital Space LLC • Greater London

On-site
GBP 90,000 - 130,000
Daily catered lunches
Modern office environment
Tech talks and knowledge sharing
Senior Platform Engineer & SRE for Cloud Reliability
Senior Platform Engineer & SRE for Cloud Reliability

Myn • Greater London

Hybrid
GBP 90,000 - 120,000
Cloud SRE: Build Reliable, Scalable Platforms
Cloud SRE: Build Reliable, Scalable Platforms

Lloyds Banking Group • Manchester

Hybrid
GBP 70,000 - 110,000
Pension up to 15%
Annual bonus
Share schemes
+3