Senior Cloud Engineer, AI Platform SRE

CreateFuture

Greater London

On-site

GBP 90,000 - 130,000

Full time

13 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

35 days leave including bank holidays
Private medical insurance
Enhanced parental and adoption leave
Paid learning and development

Job summary

CreateFuture is an AI-native consulting partner helping organizations build reliable AI platforms. You will design, build, and maintain the Kubernetes infrastructure for AI workloads, and shape CI/CD pipelines bespoke to AI/agent workflows.

Expect to define SLOs, manage on-call, and build observability dashboards to surface issues early. The role emphasizes Python automation, Terraform, and cloud experience (AWS/GCP) with strong collaboration across teams and a culture focused on craft and

Qualifications

  • 4+ years in DevOps/SRE or platform engineering with production ownership of Kubernetes-based systems.
  • Hands-on experience operating AI/ML systems in production is a strong plus.
  • Proficiency in Python for automation with Terraform and at least one major cloud provider (AWS or GCP).

Responsibilities

  • Design, build, and maintain Kubernetes infrastructure supporting AI workloads with IaC (Terraform).
  • Design and maintain CI/CD pipelines for AI/agent workflows with automated eval and regression checks.
  • Define and track SLOs/SLAs and bring SRE discipline to incident response and postmortems.
  • Participate in on-call rotations and maintain runbooks.
  • Build observability for AI concerns with dashboards and alerts.
  • Collaborate with Inference Control Plane and other teams to ensure operability from day one.

Skills

Kubernetes
Python
Terraform
CI/CD
Observability

Tools

GitHub Actions
GitLab CI
Jenkins
Argo CD
Datadog
Prometheus
Grafana

Job description

Working at CreateFutureCreateFuture is an AI-native consulting partner where people do work that matters and are supported to do it well. We work alongside organisations such as PayPal, adidas, NatWest, FanDuel and Money Saving Expert, building digital products and services that make a difference while always putting people first.

We’re a team of creators. We write code, shape delivery, build go-to-market strategies, develop AI solutions and create the practices that support our people. We work side by side with our clients, challenging what’s not working and helping them to build the future. Our commitment to craft, quality, and culture has helped us scale to over 600 people in just a few years.

  • 35 days leave (including bank holidays).
  • Private medical insurance.
  • Enhanced parental and adoption leave.
  • 40 hours of paid learning and development.

Join us on our journey. Let’s create tomorrow, together, today.

About the role and team:

You'll be bringing SRE discipline to how AI platforms are run in production. You'll help build and operate our Kubernetes-based platform for AI workloads, support CI/CD pipelines purpose-built for AI and agentic systems, and bring the SLOs, on-call, and incident response rigour that this space has historically lacked. Success looks like an AI platform that runs reliably and scales predictably.

What you'll be doing:
  • Design, build, and maintain the Kubernetes infrastructure that supports AI workloads, including model serving, agent orchestration, and batch inference, with infrastructure as code in Terraform.
  • Design and maintain CI/CD pipelines tailored to AI and agentic workflows, including model deployment, agent/tool updates, and prompt or configuration rollouts, with automated eval and regression checks before release.
  • Define and track SLOs/SLAs for AI platform services, and bring SRE rigour to incident response, root cause analysis, and postmortems.
  • Participate in on-call rotations and maintain clear, usable runbooks.
  • Build observability for AI-specific concerns — latency, token usage, cost per request, model/agent error rates, and drift — with dashboards and alerting that surface issues before they reach users.
  • Partner with the Inference Control Plane, Evals/Observability, and Context & Knowledge Platform teams so new agents, tools, and knowledge systems are built with operability in mind from day one.
We'd love to talk to you if you:
  • Have 4+ years' experience in DevOps, SRE, or platform engineering, including production ownership of Kubernetes-based systems.
  • Have hands‑on experience operating AI or ML systems in production — model serving, LLM inference, or MLOps pipelines — a strong plus.
  • Are strong in Python for automation and operational tooling, with production experience in Terraform and at least one major cloud provider (AWS or GCP).
  • Have built and maintained CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins, Argo CD) and are comfortable with observability stacks (Datadog, Prometheus, Grafana).
  • Have calm, rigorous incident management instincts, and ideally some familiarity with LLM/agent ecosystems (model APIs, vector databases, MCP, orchestration frameworks).
What we’ll offer you:

We trust people to do their best work. That means flexibility over rigid rules, impact over activity, and real investment in your growth both professionally and personally. You’ll be part of a supportive, and friendly culture, surrounded by smart, curious people who care deeply about what they do.
We offer flexible working, including hybrid and remote options. Our office hubs are located in Edinburgh, Leeds, Manchester, London and Bulgaria, with occasional travel to client sites or CreateFuture offices when needed.

We trust you to manage your time balancing collaboration with client time and focused work. What matters is the impact you have, not how busy you look.

Our hiring process

We try to keep our hiring process clear, fair and respectful of your time. We aim to get back to everyone who applies and we will be upfront about where you are in the process.

It usually looks like this:

  • Call with our Talent Acquisition Team
  • Role specific capability interview

Depending on the role, we might also ask you to do a short presentation, a practical or technical task or have a values focused conversation. We will explain what is involved before anything happens.

Inclusion at CreateFuture

We believe diverse teams build better workplaces and better products. We want CreateFuture to be a place where people feel able to be themselves and do their best work.

If you need any adjustments or support during the application process, just. We will do what we can to help.

We look forward to your application!

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Cloud Engineer, AI Platform SRE
Senior Cloud Engineer, AI Platform SRE

CreateFuture • City of Edinburgh, Manchester, Greater London, Leeds

Hybrid
GBP 90,000 - 130,000
35 days leave
Private medical insurance
Enhanced parental and adoption leave
+1
Senior Cloud Engineer
Senior Cloud Engineer

CreateFuture • Leeds

Hybrid
GBP 90,000 - 130,000
35 days leave
Private medical insurance
Enhanced parental leave
+2
Senior Cloud Engineer
Senior Cloud Engineer

CreateFuture • City of Edinburgh

Hybrid
GBP 90,000 - 125,000
35 days leave incl. bank holidays
Private medical insurance
Enhanced parental & adoption leave
+2
Senior Cloud Engineer
Senior Cloud Engineer

CreateFuture • Manchester

Hybrid
GBP 85,000 - 110,000
35 days leave
Private medical insurance
Enhanced parental leave
+1
Lead Cloud Engineer
Lead Cloud Engineer

CreateFuture • City of Edinburgh

Hybrid
GBP 90,000 - 130,000
35 days leave
Private medical insurance
Enhanced parental leave
+1
Senior AI Platform Engineer
Senior AI Platform Engineer

CreateFuture • Manchester, City of Edinburgh, Leeds, Greater London

Hybrid
GBP 90,000 - 130,000
35 days leave
Private medical insurance
Enhanced parental leave
+1
AI Platform Engineer
AI Platform Engineer

CreateFuture • City of Edinburgh, Manchester, Greater London, Leeds

Hybrid
GBP 70,000 - 110,000
35 days leave
Private medical insurance
Enhanced parental and adoption leave
+1
Lead Cloud Engineer
Lead Cloud Engineer

CreateFuture • City of Edinburgh, Manchester, Greater London, Leeds

Hybrid
GBP 110,000 - 150,000
Flexible working
Hybrid/Remote options
Office hubs in Edinburgh, Leeds, Mach,
+1
Senior AI Engineer (Amazon Bedrock)
Senior AI Engineer (Amazon Bedrock)

CreateFuture • City of Edinburgh

Hybrid
GBP 90,000 - 130,000
35 days leave
Private medical insurance
Enhanced parental and adoption leave
+1
Principal AI Platform Engineer
Principal AI Platform Engineer

CreateFuture • City of Edinburgh, Manchester, Greater London, Leeds

Hybrid
GBP 80,000 - 100,000
Flexible working
Hybrid/Remote
Office hubs