AI Platform SRE - Infra, Security & Observability

duvo.ai

United Kingdom

Remote

GBP 70,000 - 120,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Unlimited AI budget
Autonomy to do your best work
Real AI product with real customers
Ownership and candor culture
Equity opportunity

Job summary

duvo.ai is seeking an experienced Infrastructure/SRE engineer to own the reliability, security, and infrastructure of our AI operations platform. You will oversee sandbox infrastructure, capacity, and observability, partnering with product engineers to build a scalable, secure, and observable system that runs AI agents for enterprise customers.

You will join a growing SRE team and help shape the reliability culture from the ground up, including Terraform modules, OpenTelemetry pipelines, and

Qualifications

  • Design and operate scalable systems with strong reliability.
  • Security-minded handling of enterprise data and tenant isolation.
  • Observability and incident response to minimize customer impact.
  • Automation of infrastructure and CI/CD pipelines; IaC emphasis.

Responsibilities

  • Own platform reliability, infrastructure, observability, and incident response.
  • Lead reliability initiatives from proposal to production; measure results.
  • Make informed trade-offs between reliability and speed of delivery.

Skills

Distributed systems
Security mindset
Observability
Infrastructure as code
Automation
Ownership

Tools

GCP
Kubernetes
Terraform
Docker
OpenTelemetry
Postgres
Redis

Job description

duvo.ai is seeking an experienced Infrastructure/SRE engineer to own the reliability, security, and infrastructure of our AI operations platform. You will oversee sandbox infrastructure, capacity, and observability, partnering with product engineers to build a scalable, secure, and observable system that runs AI agents for enterprise customers.

You will join a growing SRE team and help shape the reliability culture from the ground up, including Terraform modules, OpenTelemetry pipelines, and

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Platform Tech Lead - Observability & SRE
AI Platform Tech Lead - Observability & SRE

Cisco Systems Inc • City of Westminster

On-site
GBP 110,000 - 165,000
Remote AI Platform Engineer — Production Systems & Orchestration
Remote AI Platform Engineer — Production Systems & Orchestration

duvo.ai • United Kingdom

Remote
GBP 90,000 - 130,000
Unlimited AI budget
Autonomy to do your best work
Real AI product with real customers
+1
Senior AI Platform SRE: Reliability & Automation
Senior AI Platform SRE: Reliability & Automation

CloudFactory • Reading

On-site
GBP 70,000 - 110,000
AI Platform Cloud SRE Lead — Remote & Flexible
AI Platform Cloud SRE Lead — Remote & Flexible

CreateFuture • City of Edinburgh, Manchester, Greater London, Leeds

Hybrid
GBP 90,000 - 130,000
35 days leave
Private medical insurance
Enhanced parental and adoption leave
+1
Senior Platform & SRE Leader — Remote
Senior Platform & SRE Leader — Remote

aitrainer • United Kingdom

Remote
GBP 120,000 - 180,000
SRE for AI Platform Reliability & Observability
SRE for AI Platform Reliability & Observability

Callosum • Greater London

On-site
GBP 90,000 - 150,000
Equity & Ownership
Private healthcare
Visa sponsorship & relocation
+1
GenAI Platform SRE Lead: Reliability, Automation & Scale
GenAI Platform SRE Lead: Reliability, Automation & Scale

Aviva plc • United Kingdom

Hybrid
GBP 65,000 - 74,000
Bonus opportunity
Generous pension up to 14%
29 days holiday + bank holidays (buy/s
+5
AI Platform & SRE Leader — Scale & Govern AI
AI Platform & SRE Leader — Scale & Govern AI

Capgemini • Greater London

Hybrid
GBP 90,000 - 130,000
Senior Platform Engineer — Secure, Scalable Infra for Enterprise AI
Senior Platform Engineer — Secure, Scalable Infra for Enterprise AI

Dex • Greater London

On-site
GBP 90,000 - 130,000
Site Reliability Engineer
Site Reliability Engineer

Atarus • Greater London

On-site
GBP 90,000 - 150,000