Lead Site Reliability Engineer London

Zego

Greater London

Hybrid

GBP 90,000 - 130,000

Full time

4 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Salary
Annual performance bonus
Share options
Private medical insurance
Pension
Generous holiday
Travel to office/learning budget

Job summary

Zego is hiring a Site Reliability Engineer to build and own the SRE function across our AI-powered insurance platform. You will embed reliability, observability and operational excellence as core engineering concerns, and drive automation that scales with a 120+ engineering team.

You will lead SLI/SLO definitions, incident response, and runbooks, with a focus on reducing toil and enabling teams to ship confidently. Hybrid work with offices in London or Halifax.

Qualifications

  • Deep SRE experience with SLI/SLO definition and error budgets.
  • Incident response ownership to blameless retrospectives.
  • Strong Python software engineering.
  • Observability depth with modern tools (OpenTelemetry, Datadog).
  • Experience at AWS scale and Kubernetes platforms.
  • Automation and runbooks for reliable systems.

Responsibilities

  • Define and land reliability standards (SLIs/SLOs, error budgets, production readiness).
  • Automate standards into CI and tooling; provide safe defaults and guardrails.
  • Build SRE capability on the AI platform, extend MCP servers and agents; enable machine-readable runbooks.
  • Own observability practice, instrumentation, signal quality and cost.
  • Lead end-to-end incident response and post-incident retrospectives for improvements.
  • Measure and reduce toil via tooling and automation.

Skills

SRE engineering
SLI/SLO definition
Incident response
Observability
Python
AWS
Kubernetes
OpenTelemetry
Datadog
IaC (Infrastructure as Code)
AI tooling
MCP servers

Tools

Datadog
OpenTelemetry
AWS
Kubernetes
MCP servers
ArgoCD

Job description

About Zego

At Zego, we're on a mission to do the good thing, not the insurance thing.

Insurance hasn't changed much in over a century. The way we live, work and travel has. We're building the real-time, AI-driven infrastructure that powers innovative insurance, so good drivers get cover that works the way they actually drive.

We're not just updating insurance; We're leading the AI evolution in insurance

For us, AI isn't a line on a roadmap or a buzzword on a slide. It's our operating reality, and it's how we build products that back drivers instead of the old insurance playbook.

We don't do things slowly, and we don't do bureaucracy. We back high-performance builders who want ownership, early responsibility and the chance to do the most career-defining work of their lives. You'll get the space to try things, the tools to move fast, and the room to see your ideas reach millions of drivers.

Do not take our word for it. Read what Zegons say about us on Glassdoor.

If you're ready to build the future of insurance, we're hiring.

Overview of the Role

You will build the Site Reliability Engineering function at Zego, embedding reliability, observability and operational excellence as core engineering concerns - AI is the primary lever for doing that at scale, not a bolt on.

  • You will be our dedicated SRE, working alongside Systems Engineering and embedded with Product and Engineering across roughly 120 engineers. Teams own their services and their own on-call. You own the framework, the instrumentation and the AI tooling that makes that ownership work.
  • You will make teams good at running what they build: alert quality over alert volume, runbooks an agent can execute rather than prose that describes, and diagnostics good enough that an engineer or an agent reaches resolution without an SRE in the room.
  • You will be the primary advocate for reliability with Product and Engineering, making the case with data rather than assertion, so platform health is prioritised alongside delivery commitments.

Key Responsibilities

  • Define and land the reliability standard: SLIs, SLOs, error budgets and production readiness criteria that teams apply to their own services, with adoption measured rather than assumed.
  • Encode standards into automation rather than enforcing them by hand. Guardrails in CI, agents that check reliability and observability posture on pull requests, and safe defaults in shared infrastructure so the reliable path is the easy path.
  • Build SRE capability on Zego's AI platform, extending our MCP servers and agents, and making the estate legible to them through machine readable runbooks, structured telemetry and diagnostics an agent can act on.
  • Own observability as a practice, including instrumentation design, signal quality, cardinality and cost.
  • Raise the standard of incident response end-to-end, from detection and triage through to retrospectives that produce change. Put AI to work where it pays off most, stripping toil out of incidents so responders can focus on judgement
  • Measure toil, publish it, and remove it through tooling that others can run without you.

What you will need to be successful in the Role

We are looking for an engineer who lives SRE and DevOps culture, and treats AI tooling as a default part of how the work gets done, not a side experiment. You will engage and empower teams through decisions grounded in data, and you will be as comfortable building with AI as consuming it: extending MCP servers, writing agents, and automating the operational work that would otherwise fill your week. We treat reliability at Zego as a platform capability, so we care more about what you can make repeatable for others than what you can fix yourself.

What you’ll bring to the Team

  • Deep SRE experience: SLI and SLO definition, error budget management, and ownership of incident response through to blameless retrospectives.
  • Fluency with AI as an engineering tool rather than a chat window: building agents and automation, working with MCP servers or equivalent, and judgement about where an agent can act and where a human stays in the loop.
  • Strong software engineering in Python, writing tested, maintainable code that other engineers can run and extend.
  • Observability depth with OpenTelemetry and Datadog.
  • AWS at scale, and Kubernetes based platforms managed through Infrastructure as Code.

Desirable

  • Istio, Crossplane and ArgoCD. We run all three and will teach the specifics to the right person.
  • Running ML or AI workloads in production, such as model serving or feature pipelines. We do this at Zego and you would help support it.
  • Having been a sole delivery owner somewhere, and knowing what that does and does not mean.
The Zego ways of working

Teams work better with time to collaborate and space to get things done. We call it Zego Hybrid: some of us are in our office (central London or Halifax) weekly, others monthly or quarterly. It's about finding the balance between face time and focus that produces great work and a healthy life around it.

We also make a point of getting everyone in the same room. Teams come together every quarter, and once a year the whole company does, properly. It's a serious investment and consistently one of the best parts of the year.

Here is what the last one looked like.

Benefits
  • Market-competitive salary, benchmarked against your function and reviewed every year.
  • Annual performance bonus, linked to company performance and your contribution.
  • Share options, a real stake in the company and a share in the future you build
  • Private medical insurance.
  • Pension, generous holiday, and £1,000 a year to spend on getting to the office or learning something new.
  • Cutting-edge systems and tools, so you always have what you need to drive your impact.

Zego, leading the AI evolution in insurance.

Equal opportunities

We're an equal opportunity employer and we value diversity at our company. We do not discriminate on the basis of race, religion, national origin, gender, sexual orientation, age, marital status, or disability status.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer
Lead Site Reliability Engineer

Zego • Greater London

Hybrid
GBP 90,000 - 130,000
Market-competitive salary
Annual performance bonus
Share options
+2
Senior AI Product Manager - Platform
Senior AI Product Manager - Platform

Zego • Greater London

Hybrid
GBP 110,000 - 160,000
Market-competitive salary
Annual performance bonus
Share options
+5
Senior AI Product Manager – Growth London
Senior AI Product Manager – Growth London

Zego • Greater London

Hybrid
GBP 90,000 - 150,000
Salary benchmarking
Annual bonus
Share options
+4
Lead Backend Engineer London
Lead Backend Engineer London

Zego • Greater London

Hybrid
GBP 90,000 - 125,000
Salary benchmarked against function
Annual performance bonus
Share options
+3
Data Engineer, Portugal London
Data Engineer, Portugal London

Zego • Greater London

Hybrid
GBP 35,000 - 50,000
Market-competitive salary
Annual performance bonus
Share options
+3
Senior Software Engineer London
Senior Software Engineer London

Zego • Greater London

Hybrid
GBP 70,000 - 110,000
Market-competitive salary
Annual performance bonus
Share options
+3
Lead Data Engineer, Portugal
Lead Data Engineer, Portugal

Zego • Greater London

Hybrid
GBP 95,000 - 140,000
Market salary
Annual bonus
Share options
+3
Senior Software Engineer
Senior Software Engineer

Zego • Greater London

On-site
GBP 90,000 - 120,000
Salary
Annual bonus
Share options
+4
Lead Data Engineer London
Lead Data Engineer London

Zego • Greater London

Hybrid
GBP 90,000 - 120,000
Annual performance bonus
Share options
Private medical insurance
+2
Special Projects Manager London
Special Projects Manager London

Zego • Greater London

Hybrid
GBP 150,000 - 190,000
Market-competitive salary
Annual performance bonus
Share options
+4