Senior SRE Engineer - Databases & Reliability

Grafana

United Kingdom

Remote

GBP 104,000 - 125,000

Full time

4 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Equity and bonus where applicable
Remote-first company culture

Job summary

Grafana Labs is seeking a Staff Software Engineer—SRE to join our remote-first team. You will own production reliability for high-SLA, multi-tenant environments across AWS, GCP, and Azure, and help define per-tenant SLOs with a focus on reducing SLO burn.

You will lead incident response, drive automation, and mentor engineers while collaborating with product squads to ensure scalable, observable, and reliable systems across Grafana Cloud offerings.

Qualifications

  • 8+ years of engineering experience, including 4+ years in SRE or production engineering.
  • Strong Kubernetes experience in AWS, GCP, or Azure; familiarity with Helm, Terraform, or Jsonnet.
  • Operate multi-tenant production systems and design/implement SLOs.
  • Technical leadership including leading projects and mentoring engineers.
  • Experience with Go, Python or Java.
  • Experience with Linux internals; knowledge of networking, cloud storage, and scaling.
  • Excellent problem-solving and troubleshooting skills; performance, scaling, and failure modes.
  • Experience in blame-free incident response and post-incident reviews.
  • Autonomous work style and close collaboration with product engineering teams.
  • Intellectual curiosity, transparency, bias toward action, and collaboration.

Responsibilities

  • Own production reliability for high-SLA and complex customer environments.
  • Define and evolve per-tenant SLOs and reliability models; reduce SLO burn.
  • Review SLOs and reduce budget burn through monitoring, automation, self-healing, and autoscaling.
  • Design and implement automation and solutions to improve reliability and scalable growth; improve observability and alert quality.
  • Develop fault-tolerant design patterns and promote reliability across the service lifecycle.
  • Serve as primary escalation point and participate in on-call and incident response; post-incident reviews.
  • Lead customer-impacting incident response and post-incident reviews.
  • Collaborate with engineering leaders to influence product strategy and scalable designs.
  • Contribute to design docs and code reviews; collaborate across teams.
  • Teach SRE practices and promote reliability early in feature development.

Skills

SRE
Kubernetes
Go
Python
Java
Linux
Incident response
Cloud
Networking
Troubleshooting

Tools

Helm
Terraform
Jsonnet
Grafana

Job description

Grafana Labs is seeking a Staff Software Engineer—SRE to join our remote-first team. You will own production reliability for high-SLA, multi-tenant environments across AWS, GCP, and Azure, and help define per-tenant SLOs with a focus on reducing SLO burn.

You will lead incident response, drive automation, and mentor engineers while collaborating with product squads to ensure scalable, observable, and reliable systems across Grafana Cloud offerings.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff SRE: Cloud Databases (Multi-Tenant, Remote)
Staff SRE: Cloud Databases (Multi-Tenant, Remote)

Grafana Labs • Greater London

On-site
GBP 103,958 - 124,750
100% Remote work
30 days annual leave
In-person onboarding
Remote Senior Backend Engineer: AI-Driven Data Platform
Remote Senior Backend Engineer: AI-Driven Data Platform

Grafana • Pathhead

Remote
GBP 91,000 - 115,000
Remote-first culture
Global team
30 days annual leave
+1
Senior Backend Engineer - Databases - Analytics | UK | Remote
Senior Backend Engineer - Databases - Analytics | UK | Remote

Grafanalabs • United Kingdom

Remote
GBP 91,000 - 115,000
Equity
Bonus
30 days annual leave
+1
Senior Backend Engineer - Databases - Analytics | UK | Remote
Senior Backend Engineer - Databases - Analytics | UK | Remote

Grafana • Pathhead

Remote
GBP 91,000 - 115,000
Remote-first culture
Global team
30 days annual leave
+1
Remote Staff Backend Engineer, Alerts & Cloud
Remote Staff Backend Engineer, Alerts & Cloud

Grafana Labs • Greater London

On-site
GBP 103,958 - 124,750
RSUs
Remote work
Global culture
+1
Remote Engineering Manager – Observability Platform
Remote Engineering Manager – Observability Platform

Grafanalabs • United Kingdom

Remote
GBP 103,000 - 129,000
RSUs
Remote-first culture
30 days annual leave
Senior SRE: Cloud Reliability & Observability Lead
Senior SRE: Cloud Reliability & Observability Lead

Omilia Natural Language Solutions Ua Ltd • United Kingdom

On-site
GBP 85,000 - 120,000
Fixed compensation
Long-term employment with the working
Development in professional growth (ca
+1
Senior SRE
Senior SRE

Pulse Recruit • Greater London

On-site
GBP 65,000 - 85,000
Backend Platform Engineer - Remote (UK)
Backend Platform Engineer - Remote (UK)

Coinscapture • United Kingdom

Remote
GBP 72,000 - 90,000
RSUs
Remote-first culture
30 days annual leave with shutdowns
Senior SRE: AI Cloud Reliability & GPU Scale
Senior SRE: AI Cloud Reliability & GPU Scale

Nebius Group • Greater London

On-site
GBP 90,000 - 150,000
Competitive pay
Career growth
Flexibility
+3