Senior Site Reliability Engineer

CloudFactory

Reading

On-site

GBP 70,000 - 110,000

Full time

8 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

CloudFactory is seeking a Site Reliability Engineer to keep production systems reliable, scalable, and secure. You will work with engineers and operators to fuse engineering, operation, and security for platform and service excellence.

The role emphasizes building golden paths, developer tooling, and end-to-end software delivery governance to support ML/LLM workloads and cloud deployments. This is a chance to grow in a mission-driven, globally connected environment.

Qualifications

  • 5+ years in infrastructure engineering, DevOps, or SRE.
  • Experience with Kubernetes in large-scale production.
  • Proficiency in Python or Go for automation.
  • Experience with AI/ML workloads on Kubernetes is a plus.
  • Strong collaboration across product, backend, and frontend teams.

Responsibilities

  • Ensure reliability of platform including ML/LLM workloads and infrastructure.
  • Implement observability and tracing for ML models and services.
  • Contribute to company-wide technical direction and golden paths.
  • Develop reusable tooling and automation to accelerate delivery.
  • Package common open-source tools (Grafana, Istio, CloudNative stack, ML tooling).
  • Embed security, compliance, and cost governance into platform design.

Skills

Kubernetes
Python
Go
Terraform
Helm
CloudFormation
AI/ML tooling
Observability

Tools

Terraform
CloudFormation
Helm

Job description

At CloudFactory, we are a mission-driven team passionate about unlocking the potential of AI to transform the world. By combining advanced technology with a global network of talented people, we make unusable data usable, driving real-world impact at scale.

More than just a workplace, we’re a global community founded on strong relationships and the belief that meaningful work transforms lives. Our commitment to earning, learning, and serving fuels everything we do as we strive to connect one million people to meaningful work and build leaders worth following.

Our Culture

At CloudFactory, we believe in building a workplace where everyone feels empowered, valued, and inspired to bring their authentic selves to work. We are:

  • Mission-Driven: We focus on creating economic and social impact.
  • People-Centric: We care deeply about our team’s growth, well‑being, and sense of belonging.
  • Innovative: We embrace change and find better ways to do things together.
  • Globally Connected: We foster collaboration between diverse cultures and perspectives.

If you’re passionate about innovation, collaboration, and making a real impact, we’d love to have you on board!

Role Summary

As a Site Reliability Engineer, you will play a key role in keeping all production systems running smoothly. You will work closely with other engineers and operators to fuse engineering principles, operational knowledge, security, and automation to work towards platform/service production excellence from an angle of infrastructure, reliability, and security.

The SRE team owns the foundation of AI Platform’s Core platform - the services and infrastructure that let us deploy to a multitude of public cloud providers and that powers many ML and LLM powered features. We give every other engineering team a reliable base to build on, and we own the software delivery lifecycle end to end: the tooling, patterns, and automation that reduce friction for the whole org.

This is an exciting opportunity to grow professionally while contributing to a mission-driven organization.

Responsibilities

What you’ll own

  • Reliability of platform(includes ML and LLM workloads) - model serving and inference infrastructure (GPU-backed endpoints, autoscaling, latency and cost tradeoffs), with SLOs, on-call, and incident response that cover models, not just services
  • Observability(includes ML models) - drift and performance monitoring for ML, plus LLM-specific tracing, evals, and guardrails, wired into the same metrics and logging stacks we run everywhere else
  • Company‑wide technical direction: shaping the roadmap and building golden paths that raise the baseline for every team
  • Developer tooling and automation that compounds - reusable GitHub Actions, GitOps workflows, Terraform modules - so every engineer ships faster
  • Reusable components packaging common open‑source tools (Grafana, Istio, CloudNative stack, and ML tooling such as model registries and feature stores) for teams to deploy in any environment
  • Secure-by-default infrastructure - baking security, compliance audits, cost governance, and audit trails into the platform in close partnership with our lead/backend/staff engineers.
Who you are (must-haves)
  • 5+ years in infrastructure engineering, DevOps, or SRE, operating large-scale, high-availability production systems using Kubernetes
  • Production Operational experience - a live cluster under real load, not a lab. Fluent with Helm, and Terraform or Cloudformation, on at least one major cloud (AWS preferred).
  • Good proficiency in Python or Go or general scripting for automation and tooling(automation with higher language preferred)
  • AI is already in your daily loop - Agentic tooling (Claude Code, Codex, Droid, internal skills) is part of how you ship and not what you are experimenting with. We believe AI tools can be great with human judgement and we want the SRE team to bring the next wave day to day operations.
  • First-principles reasoning - Reasoning from constraints and failure modes naming the tradeoff in business terms (reliability vs. velocity, cost vs. blast radius, standardisation vs. one-off)
  • At least one infrastructure build you owned end to end - with the outcome metric attached (deploy time, MTTR, cost, adoption, availability).
  • Cross-functional strength. Track record working with product, backend/frontend teams to pull through collective initiative.
ML & AI platform (strongly preferred)
  • Running ML workloads on Kubernetes - GPU scheduling, capacity, and cost management
  • Model serving and inference at production scale (eg KServe, RayServe, Triton, vLLM, or similar) with real latency and cost constraints(preferred RayServe)
  • MLOps pipeline tooling - training pipelines, model registries, feature stores, and lineage (Kubeflow, MLflow, Feast, Weights & Biases, or equivalents)
  • LLMOps in production - inference serving, prompt/version management, and LLM observability (tracing, evals, drift, guardrails, cost per request)
  • Governing ML/LLM workloads as platform capabilities: data-residency and PII controls, and audit trails
Any other General requirements
  • Global Collaboration: Ability to work across global teams and different cultures across various time zones with strong communication skills.
  • Problem Solving: Ability to break down complex problems into simple, actionable solutions.
  • Ownership & Drive: Tendency to go above and beyond to meet deadlines, manage own deliverables, and assist team members.
  • Availability: Willingness to support processes for 24x7 operational support.

At CloudFactory, we believe that work should be more than just a job—it should be a platform for growth, impact, and community. Here, you’ll earn with purpose, learn every day, and serve a mission that truly matters. If you're looking for a career where you can develop professionally, contribute meaningfully, and be part of a global movement, we’d love to have you on this journey!

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Platform SRE: Reliability & Automation
Senior AI Platform SRE: Reliability & Automation

CloudFactory • Reading

On-site
GBP 70,000 - 110,000
Forward Deployed Engineer
Forward Deployed Engineer

CloudFactory • Reading

On-site
GBP 60,000 - 90,000
Relocation assistance
Frequent travel to client sites
Forward Deployed Engineer
Forward Deployed Engineer

CloudFactory Limited • Reading

Hybrid
GBP 60,000 - 90,000
Relocation assistance
Site Reliability Engineer, Mistral Cloud
Site Reliability Engineer, Mistral Cloud

Mistral AI • Greater London

On-site
GBP 75,000 - 110,000
Healthcare coverage
Relocation support
Retirement plans
+3
Principal Site Reliability Engineer, Infrastructure Observability
Principal Site Reliability Engineer, Infrastructure Observability

United States Digital Space LLC • Greater London

Hybrid
GBP 120,000 - 170,000
Hybrid work up to 3 days per week
Site Reliability Engineer
Site Reliability Engineer

United States Digital Space LLC • Greater London

On-site
GBP 90,000 - 130,000
Daily catered lunches
Modern office environment
Tech talks and knowledge sharing
Senior Cloud Engineer, AI Platform SRE
Senior Cloud Engineer, AI Platform SRE

CreateFuture • City of Edinburgh, Manchester, Greater London, Leeds

On-site
GBP 90,000 - 130,000
35 days leave
Private medical insurance
Enhanced parental and adoption leave
+1
Site Reliability - Member of Technical Staff
Site Reliability - Member of Technical Staff

Callosum • Greater London

On-site
GBP 90,000 - 150,000
Equity & Ownership
Private healthcare
Visa sponsorship & relocation
+1
Senior Cloud Engineer, AI Platform SRE
Senior Cloud Engineer, AI Platform SRE

CreateFuture • Greater London

Hybrid
GBP 90,000 - 130,000
35 days leave including bank holidays
Private medical insurance
Enhanced parental and adoption leave
+1
Senior Site Reliability Engineer - Selby Jennings
Senior Site Reliability Engineer - Selby Jennings

eFinancialCareers • Greater London

On-site
GBP 90,000 - 130,000