Site Reliability Engineer

Camwebdir

United Kingdom

Remote

GBP 90,000 - 120,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

23 days holiday
Birthday day off
Private medical insurance
Life insurance
Pension scheme
Enhanced family leave
Employee assistance program
Cycle to work

Job summary

Darktrace is seeking a Site Reliability Engineer (SRE) to lead reliability within a key domain, shaping platform reliability strategy across engineering groups. You’ll act as the go-to SME, define standards, and drive best practices in collaboration with Platform Engineering and DevSecOps.

The role focuses on a core area—observability, performance, data infrastructure reliability, security-focused SRE, or network reliability—while advancing resilience and scalable systems across the stack.

Qualifications

  • Proven experience in Site Reliability Engineering, DevOps, or infrastructure engineering.
  • Deep expertise in at least one area: Observability, Performance engineering, Data infrastructure reliability, security-focused SRE, or network reliability.
  • Strong programming skills (Go, Python, or similar).
  • Experience with cloud platforms (AWS, GCP, Azure) and Kubernetes.
  • Strong communication skills, with the ability to explain complex technical concepts clearly.
  • Self-driven with the ability to identify and prioritise high-impact work independently.

Responsibilities

  • Domain Expertise & Strategy: act as SME in your reliability domain and set standards.
  • Engineering & Delivery: design solutions for complex reliability challenges and build tooling.
  • Collaboration & Platform Integration: embed your domain in internal platforms and work with DevSecOps.
  • Incident & Operational Excellence: contribute to incident response and runbooks.

Skills

SRE expertise
Observability
Performance engineering
Data reliability
Security‑focused SRE
Network reliability
Go/Python
Cloud platforms
Kubernetes
Communication skills
Self-driven

Job description

Darktrace is a global leader in AI for cybersecurity that keeps organizations ahead of the changing threat landscape every day. Founded in 2013, Darktrace provides the essential cybersecurity platform protecting nearly 10,000 organizations from unknown threats using its proprietary AI.

The Darktrace Active AI Security Platform™ delivers a proactive approach to cyber resilience to secure the business across the entire digital estate – from network to cloud to email. Breakthrough innovations from our R&D teams have resulted in over 200 patent applications filed. Darktrace’s platform and services are supported by over 2,400 employees around the world. To learn more, visit http://www.darktrace.com.

Job Description
About the Role

We’re looking for a Site Reliability Engineer (SRE) to bring deep expertise in a key reliability domain and help shape the future of our platform reliability strategy.

SRE sits at the heart of our operational trifecta alongside Platform Engineering and DevSecOps. In this role, you’ll act as the go‑to authority in your area of specialism, working across teams to embed best practices, solve complex reliability challenges, and improve system resilience at scale.

Unlike a generalist SRE, this role focuses on a core domain of expertise—such as observability, performance engineering, data infrastructure reliability, security‑focused SRE, or network reliability—while influencing reliability standards across the wider engineering organisation.

Key Responsibilities
Domain Expertise & Strategy
  • Act as the subject matter expert in your chosen reliability domain
  • Define and implement standards, frameworks, and best practices across SRE, Platform Engineering, and DevSecOps
  • Stay current with industry trends and bring innovative ideas into the organisation
Engineering & Delivery
  • Design and implement solutions to complex, cross‑cutting reliability challenges
  • Build tooling, automation, and frameworks to improve system resilience and scalability
  • Lead deep‑diving investigations into systemic issues and drive long‑term fixes
Collaboration & Platform Integration
  • Partner with Platform Engineering to ensure your domain is embedded within the internal developer platform
  • Collaborate with DevSecOps to integrate security, compliance, and resilience practices
  • Contribute to cross‑team initiatives that improve reliability across the stack
Incident & Operational Excellence
  • Play a key role in incident response, particularly within your specialism
  • Contribute to on‑call rotations and continuous improvement of operational processes
  • Develop runbooks, documentation, and training materials to support teams
What You’ll Bring
Essential
  • Proven experience in Site Reliability Engineering, DevOps, or infrastructure engineering
  • Deep expertise in at least one of the following areas:
    • Observability & monitoring (metrics, logging, distributed tracing)
    • Performance engineering & capacity planning
    • Data infrastructure reliability (databases, streaming, pipelines)
    • Security‑focused SRE (hardening, compliance automation, secrets management)
    • Network reliability & traffic management
  • Strong programming skills (e.g. Go, Python, or similar)
  • Experience with cloud platforms (AWS, GCP, Azure) and Kubernetes
  • Strong communication skills, with the ability to explain complex technical concepts clearly
  • Self‑driven with the ability to identify and prioritise high‑impact work independently
Desirable
  • Experience building internal developer platforms or tooling
  • Contributions to open‑source, technical blogs, or public speaking
  • Experience working in regulated environments
  • Familiarity with SLO frameworks and error budget management
  • Relevant certifications in your specialist domain
Success Measures
  • Improved reliability and performance within your domain of specialism
  • Adoption of best practices across SRE, Platform Engineering, and DevSecOps
  • Reduction in incidents and faster resolution times
  • Scalable, well‑integrated solutions within the internal platform
  • Strong collaboration across teams and measurable improvements in operational maturity
Why Join Us?
  • Shape reliability strategy in a modern, cloud‑native engineering environment
  • Work on complex, high‑impact systems at scale
  • Collaborate with expert teams across Platform Engineering and DevSecOps
  • Take ownership of a domain and drive meaningful, organisation‑wide impact
Benefits:
  • 23 days’ holiday + all public holidays, rising to 25 days after 2 years of service,
  • Additional day off for your birthday,
  • Private medical insurance which covers you, your cohabiting partner and children,
  • Life insurance of 4 times your base salary,
  • Salary sacrifice pension scheme,
  • Enhanced family leave,
  • Confidential Employee Assistance Program,
  • Cycle to work scheme.

Darktrace is an Equal Opportunity Employer. We consider all qualified applicants for employment without regard to race, color, religion, sex (including pregnancy, childbirth, and related medical conditions), sexual orientation, gender identity or expression, national origin, age, disability, genetic information, marital status, veteran or military status, or any other characteristic protected by applicable federal, state, or local law.

Darktrace is committed to providing reasonable accommodations to qualified individuals with disabilities in accordance with applicable laws. If you require a reasonable accommodation to participate in the application or interview process, please contact your Talent Partner.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Darktrace Ltd • Cambridge

On-site
GBP 60,000 - 80,000
23 days holiday plus public holidays
Private medical insurance
Life insurance of 4 times base salary
+2
Specialist Software Engineer
Specialist Software Engineer

Darktrace • Cambridge

On-site
GBP 90,000 - 130,000
Technical Customer Support Engineer
Technical Customer Support Engineer

Darktrace • Cambridge

On-site
GBP 32,000 - 46,000
Private medical insurance
Life insurance
Salary sacrifice pension scheme
+2
RESPOND Support Engineer
RESPOND Support Engineer

darktrace • Cambridge

On-site
GBP 26,000 - 38,000
23 days holiday
Birthday leave
Private medical insurance
+5
Site Reliability Engineer: Observability & Platform Resilience
Site Reliability Engineer: Observability & Platform Resilience

Darktrace Ltd • Cambridge

On-site
GBP 60,000 - 80,000
23 days holiday plus public holidays
Private medical insurance
Life insurance of 4 times base salary
+2
Senior Software Engineer (Python)
Senior Software Engineer (Python)

Darktrace • Cambridge

Hybrid
GBP 90,000 - 130,000
Senior Software Engineer (Python)
Senior Software Engineer (Python)

Darktrace • United Kingdom

On-site
GBP 90,000 - 120,000
SRE Lead: Observability & Platform Reliability
SRE Lead: Observability & Platform Reliability

Camwebdir • United Kingdom

Remote
GBP 90,000 - 120,000
23 days holiday
Birthday day off
Private medical insurance
+5
Senior DevOps Engineer (Azure)
Senior DevOps Engineer (Azure)

Darktrace • Greater London

Hybrid
GBP 90,000 - 130,000
Private medical insurance
Life insurance
Cycle to work scheme
+3
Senior Solutions Engineer - Enterprise Accounts
Senior Solutions Engineer - Enterprise Accounts

darktrace • Greater London

Hybrid
GBP 90,000 - 130,000