Staff Site Reliability Engineer

Obsidian Security

Cheltenham

On-site

GBP 124,000 - 141,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation
Equity and 401k
Comprehensive healthcare
Flexible paid time off
Paid holiday time off
12 weeks of new parent or family leave

Job summary

Obsidian Security in Cheltenham seeks a Staff Site Reliability Engineer to lead the reliability strategy for a complex multi-tenant SaaS platform. You will partner with DevOps and Platform Engineering, shape reliability practices, and ensure proactive issue detection. Candidates need over 5 years in SRE or production engineering and expertise in AWS or GCP. The position offers a competitive salary between £124,000 and £141,000 GBP.

Qualifications

  • 5+ years in SRE, Production Engineering, or related roles.
  • 3+ years operating at a senior or technical leadership level.
  • Deep expertise in AWS and/or GCP.
  • Proven experience designing reliability systems for SaaS platforms.

Responsibilities

  • Define and lead long-term reliability strategy across services.
  • Embed reliability, standardize SLI/SLOs across teams.
  • Build intelligent detection systems and enable self-service observability.
  • Define and evolve a tiered incident communication strategy.

Skills

SRE
Production Engineering
AWS
GCP
Kubernetes
Observability stacks
CI/CD systems

Job description

Founded in 2017, Obsidian Security was created to close a critical gap: securing the SaaS applications where modern business happens—platforms like Microsoft 365, Salesforce, and hundreds more. Backed by top investors including Greylock, Norwest Venture Partners, and IVP, we’ve built a complete SaaS security platform to reduce risk, detect and respond to threats, and prevent breaches at the source. Our team includes leaders who helped define the categories of endpoint and identity security at CrowdStrike, Okta, Cylance, and Carbon Black.

Now, we’re transforming how SaaS is secured—in the era of agentic AI. Today, Obsidian is trusted by global enterprises like Snowflake, T‑Mobile, and Pure Storage. We protect more than 200 organizations across North America, Europe, the Middle East, Southeast Asia, Australia, and New Zealand—including many of the world’s largest Fortune 1000 and Global 2000 companies.

With strong global momentum, a growing partner ecosystem including SentinelOne, Databricks, and Google Cloud, and a major fundraise on the horizon, we’re scaling quickly toward long‑term growth and IPO readiness. Join us as we define the future of SaaS security!

Staff Site Reliability Engineer

As a Staff SRE at Obsidian, you will define and drive the company‑wide reliability vision for a complex, multi‑tenant SaaS platform serving enterprise and financial customers. You will operate as a strategic partner to DevOps and Platform Engineering leadership, shaping a unified reliability strategy that scales across the organization.

Your core mandate: ensure Obsidian detects, diagnoses, and communicates system issues before customers are impacted—consistently and predictably. This is a hands‑on technical role that involves architecting and leading the implementation of systems that handle real‑world complexity, including upstream SaaS dependencies, sparse and noisy signals, and mission‑critical enterprise workloads.

Key Responsibilities
  • Reliability Strategy & Architecture – Define and lead long‑term reliability strategy across services. Establish end‑to‑end system visibility frameworks and guide architecture for observability, detection, and resilience.
  • Cross‑Org Leadership – Partner across teams to embed reliability, standardize SLI/SLOs, and serve as a technical escalation expert.
  • Detection & Observability – Build intelligent detection systems (anomaly detection, connector health models) and enable self‑service observability.
  • Incident Management – Define and evolve a tiered incident communication strategy, improve response practices, and lead postmortems to strengthen reliability and customer trust.
  • Execution – Contribute hands‑on to system design, monitoring, and debugging across distributed systems and data pipelines.
Required Qualifications
  • 5+ years in SRE, Production Engineering, or related roles
  • 3+ years operating at a senior or technical leadership level (Staff or equivalent scope)
  • Deep expertise in:
    • AWS and/or GCP
    • Kubernetes and Helm
    • Observability stacks (Prometheus, Grafana, or equivalent)
    • CI/CD systems (GitLab CI/CD, ArgoCD, etc.)
  • Proven experience designing and scaling reliability systems for multi‑tenant SaaS platforms
  • Strong debugging and systems thinking across distributed microservices and legacy systems
  • Demonstrated ability to lead initiatives that improve incident detection, response, and system resilience
  • Hands‑on engineering approach with a track record of building—not just configuring—reliability systems
Preferred Qualifications
  • Experience in B2B SaaS serving enterprise or financial customers
  • Familiarity with third‑party SaaS connector architectures and ingestion patterns
  • Experience building anomaly detection or intelligent alerting systems
  • Experience designing customer‑facing status pages and incident communication frameworks
Why This Role
  • Drive org‑wide reliability strategy
  • Own and build new detection & observability systems
  • Tackle complex distributed systems challenges
  • Safeguard critical infrastructure for financial customers
What Success Looks Like
  • Issues caught and resolved before customer impact
  • Reliability is measurable and continuously improving
  • Teams self‑serve observability with scalable tools
  • Clear, proactive incident communication builds trust
  • Reliability becomes a competitive advantage
Employee Benefits

Our competitive benefits packages are designed to support our employees' well‑being, both at work and at home. Our US‑based employees enjoy competitive compensation with equity and 401k, comprehensive healthcare with dental and vision coverage, flexible paid time off and paid holiday time off, 12 weeks of new parent or family leave, and personal and professional development resources. For more details on our US benefits, or for information on our international benefits, please see here.

Pay Transparency

Please note that the base pay range is a guideline and for candidates who receive an offer, the base pay will vary based on factors such as work location, as well as the knowledge, skills and experience of the candidate. In addition to a competitive base salary, this position is eligible for equity awards and may be eligible for sales commission or incentive compensation based on the role or function within the company.

Base Salary Range: £124,000 GBP – £141,000 GBP

At Obsidian, we are proud to be an equal‑opportunity employer. We value diversity and hire for talent, passion, and compassion. In compliance with federal law, all persons hired will be required to submit satisfactory proof of identity and legal authorization. If you have a need that requires accommodation, please contact accommodations@obsidiansecurity.com. Information collected and processed as part of any job applications you choose to submit is subject to Obsidian’s Applicant Privacy Policy.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Site Reliability Engineer
Staff Site Reliability Engineer

Obsidian Security • Manchester

On-site
GBP 124,000 - 141,000
Equity awards
Comprehensive healthcare
Flexible paid time off
+1
Site Reliability Engineer
Site Reliability Engineer

Obsidian Security • Salford

On-site
GBP 85,000 - 103,000
Competitive compensation with equity
Comprehensive healthcare
Flexible paid time off
+2
Senior Software Engineer - Data Pipelines
Senior Software Engineer - Data Pipelines

Obsidian Security • Cheltenham

Hybrid
GBP 91,000 - 106,000
Competitive compensation
Comprehensive healthcare
Flexible paid time off
+1
Enterprise Account Executive - UK
Enterprise Account Executive - UK

Obsidian Security • Greater London

Hybrid
GBP 60,000 - 120,000
Equity & 401k
Healthcare including dental/vision
Flexible PTO
+2
Director, Sales (EMEA)
Director, Sales (EMEA)

Obsidian Security • United Kingdom

On-site
GBP 100,000 - 150,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

LSEG • Nottingham

On-site
GBP 70,000 - 90,000
Healthcare
Retirement planning
Paid volunteering days
+1
Sr Channel Account Manager (EMEA)
Sr Channel Account Manager (EMEA)

Obsidian Security • United Kingdom

On-site
GBP 132,000 - 154,000
Competitive compensation with equity
Comprehensive healthcare including dental and vision
Flexible paid time off
+2
SRE Architect (68019)
SRE Architect (68019)

Hitachi Digital Services • Greater London

On-site
GBP 90,000 - 150,000
Principal Site Reliability Engineer, Infrastructure Observability
Principal Site Reliability Engineer, Infrastructure Observability

United States Digital Space LLC • Greater London

Hybrid
GBP 120,000 - 170,000
Hybrid work up to 3 days per week
SRE Architect (68019) (DEAI DS) Cloud & Data Engineering United Kingdom
SRE Architect (68019) (DEAI DS) Cloud & Data Engineering United Kingdom

Hitachids • Greater London

On-site
GBP 90,000 - 140,000