Lead Site Reliability Engineer

Zego

Greater London

Hybrid

GBP 90,000 - 130,000

Full time

4 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Market-competitive salary
Annual performance bonus
Share options
Private medical insurance
Pension and learning budget

Job summary

Zego is hiring a Site Reliability Engineer to build and own the reliability framework across its AI-powered insurance platform. You will define SLIs/SLOs, automate guardrails, and lead incident response to reduce toil while enabling rapid scale.

The role blends traditional SRE with AI tooling, observability, and platform engineering across a large engineering org. You will work in a hybrid model with offices in London or Halifax, collaborating with ~120 engineers and product teams to raise

Qualifications

  • Deep SRE experience defining SLIs and SLOs and leading incident retrospectives.
  • Ability to drive reliability through data and automation.
  • Experience integrating AI tooling into engineering workflows.
  • Strong Python development and maintainable code practices.

Responsibilities

  • Define and land reliability standards (SLIs, SLOs, error budgets).
  • Encode standards into automation and safe defaults for reliability.
  • Build SRE capability on Zego's AI platform and runbooks for ops.
  • Own observability with instrumentation design and signal quality.
  • Improve incident response from detection to retrospectives with AI.
  • Measure and reduce toil via tooling that others can reuse.

Skills

SLI/SLO definition
Error budget management
Incident response ownership
Python programming

Tools

OpenTelemetry
Datadog
AWS
Kubernetes
Infrastructure as Code

Job description

About Zego

At Zego, we're on a mission to do the good thing, not the insurance thing.

Insurance hasn't changed much in over a century. The way we live, work and travel has. We're building the real-time, AI-driven infrastructure that powers innovative insurance, so good drivers get cover that works the way they actually drive.

We're not just updating insurance; We're leading the AI evolution in insurance

For us, AI isn't a line on a roadmap or a buzzword on a slide. It's our operating reality, and it's how we build products that back drivers instead of the old insurance playbook.

We don't do things slowly, and we don't do bureaucracy. We back high-performance builders who want ownership, early responsibility and the chance to do the most career-defining work of their lives. You'll get the space to try things, the tools to move fast, and the room to see your ideas reach millions of drivers.

Do not take our word for it. Read what Zegons say about us on Glassdoor.

If you're ready to build the future of insurance, we\'re hiring.

Overview of the Role

You will build the Site Reliability Engineering function at Zego, embedding reliability, observability and operational excellence as core engineering concerns - AI is the primary lever for doing that at scale, not a bolt on.

  • You will be our dedicated SRE, working alongside Systems Engineering and embedded with Product and Engineering across roughly 120 engineers. Teams own their services and their own on-call. You own the framework, the instrumentation and the AI tooling that makes that ownership work.

  • You will make teams good at running what they build: alert quality over alert volume, runbooks an agent can execute rather than prose that describes, and diagnostics good enough that an engineer or an agent reaches resolution without an SRE in the room.

  • You will be the primary advocate for reliability with Product and Engineering, making the case with data rather than assertion, so platform health is prioritised alongside delivery commitments.

Key Responsibilities

  • Define and land the reliability standard: SLIs, SLOs, error budgets and production readiness criteria that teams apply to their own services, with adoption measured rather than assumed.

  • Encode standards into automation rather than enforcing them by hand. Guardrails in CI, agents that check reliability and observability posture on pull requests, and safe defaults in shared infrastructure so the reliable path is the easy path.

  • Build SRE capability on Zego\'s AI platform, extending our MCP servers and agents, and making the estate legible to them through machine readable runbooks, structured telemetry and diagnostics an agent can act on.

  • Own observability as a practice, including instrumentation design, signal quality, cardinality and cost.

  • Raise the standard of incident response end-to-end, from detection and triage through to retrospectives that produce change. Put AI to work where it pays off most, stripping toil out of incidents so responders can focus on judgement

  • Measure toil, publish it, and remove it through tooling that others can run without you.

What you will need to be successful in the Role

We are looking for an engineer who lives SRE and DevOps culture, and treats AI tooling as a default part of how the work gets done, not a side experiment. You will engage and empower teams through decisions grounded in data, and you will be as comfortable building with AI as consuming it: extending MCP servers, writing agents, and automating the operational work that would otherwise fill your week. We treat reliability at Zego as a platform capability, so we care more about what you can make repeatable for others than what you can fix yourself.

What you’ll bring to the Team

  • Deep SRE experience: SLI and SLO definition, error budget management, and ownership of incident response through to blameless retrospectives.

  • Fluency with AI as an engineering tool rather than a chat window: building agents and automation, working with MCP servers or equivalent, and judgement about where an agent can act and where a human stays in the loop.

  • Strong software engineering in Python, writing tested, maintainable code that other engineers can run and extend.

  • Observability depth with OpenTelemetry and Datadog.

  • AWS at scale, and Kubernetes based platforms managed through Infrastructure as Code.

Desirable

  • Istio, Crossplane and ArgoCD. We run all three and will teach the specifics to the right person.

  • Running ML or AI workloads in production, such as model serving or feature pipelines. We do this at Zego and you would help support it.

  • Having been a sole delivery owner somewhere, and knowing what that does and does not mean.

The Zego ways of working

Teams work better with time to collaborate and space to get things done. We call it Zego Hybrid: some of us are in our office (central London or Halifax) weekly, others monthly or quarterly. It\'s about finding the balance between face time and focus that produces great work and a healthy life around it.

We also make a point of getting everyone in the same room. Teams come together every quarter, and once a year the whole company does, properly. It\'s a serious investment and consistently one of the best parts of the year.

Here is what the last one looked like.

Benefits

Market-competitive salary, benchmarked against your function and reviewed every year.

Annual performance bonus, linked to company performance and your contribution.

Share options, a real stake in the company and a share in the future you build

Private medical insurance.

Pension, generous holiday, and £1,000 a year to spend on getting to the office or learning something new.

Cutting-edge systems and tools, so you always have what you need to drive your impact.

Ready to do the most impactful work of your career? We want to hear from you.

Zego, leading the AI evolution in insurance.

Equal opportunities

We\'re an equal opportunity employer and we value diversity at our company. We do not discriminate on the basis of race, religion, national origin, gender, sexual orientation, age, marital status, or disability status.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer London
Lead Site Reliability Engineer London

Zego • Greater London

Hybrid
GBP 90,000 - 130,000
Salary
Annual performance bonus
Share options
+4
Senior Software Engineer London
Senior Software Engineer London

Zego • Greater London

Hybrid
GBP 70,000 - 110,000
Market-competitive salary
Annual performance bonus
Share options
+3
Special Projects Manager London
Special Projects Manager London

Zego • Greater London

Hybrid
GBP 150,000 - 190,000
Market-competitive salary
Annual performance bonus
Share options
+4
Software Engineer London
Software Engineer London

Zego • Greater London

Hybrid
GBP 45,000 - 65,000
Market-competitive salary
Annual performance bonus
Share options
+3
Engineering Manager London
Engineering Manager London

Zego • Greater London

Hybrid
GBP 90,000 - 140,000
Market-competitive salaire
Annual performance bonus
Share options
+4
Data Engineer London
Data Engineer London

Zego • Greater London

Hybrid
GBP 75,000 - 110,000
Private medical insurance
Pension
Annual performance bonus
+2
Senior AI Product Manager – Growth
Senior AI Product Manager – Growth

Zego • Greater London

Hybrid
GBP 90,000 - 150,000
Market salary
Performance bonus
Share options
+4
Lead Backend Engineer London
Lead Backend Engineer London

Zego • Greater London

Hybrid
GBP 90,000 - 125,000
Salary benchmarked against function
Annual performance bonus
Share options
+3
Senior AI Product Manager - Platform
Senior AI Product Manager - Platform

Zego • Greater London

Hybrid
GBP 110,000 - 160,000
Market-competitive salary
Annual performance bonus
Share options
+5
Lead Data Engineer
Lead Data Engineer

Zego • Greater London

Hybrid
GBP 90,000 - 130,000
Private medical insurance
Pension
Share options
+3