Senior Software Engineer II — Reliability & Observability

United States Digital Space LLC

Greater London

Hybrid

GBP 90,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Flexible remote model
Professional development stipend
Health and parental leave

Job summary

the company is hiring a Senior Software Engineer II to join the OPX team within the Developer Experience organization. You will design and build automated reliability and self-healing systems at scale, delivering platform tooling that engineers across the company adopt for their services.

You will own incident management tooling, evolve observability infrastructure with SLOs and real-time signals, and contribute to AI-driven automation that reduces toil and speeds delivery.

Qualifications

  • 8+ years of experience designing and building products in a software engineering team.
  • Bachelor's Degree in Computer Science/Engineering or equivalent practical experience
  • 3+ years of experience on infrastructure and/or platform engineering focused teams
  • Expertise in Observability and reliability, operational metrics and data analysis
  • Proven track record architecting monitoring frameworks, SLO platforms, and automated response workflows (Datadog or equivalent)
  • Proven experience working on large-scale enterprise software applications
  • Experience in Developer Experience (DevEx) & Internal Portals
  • Familiarity with cloud platforms (AWS, GCP or similar)
  • Experience implementing AI-driven automation across the SDLC
  • Experience writing high-quality code (Go, Python or equivalent)
  • Experience mentoring engineers in a technical lead capacity
  • Proactive growth mindset

Responsibilities

  • Design and build automated reliability and self-healing systems that protect production at scale
  • Own and improve incident management tooling and on-call health
  • Develop and evolve observability infrastructure with monitoring, SLOs, and performance regression detection
  • Contribute to AI-driven operational tooling for autonomous remediation
  • Drive incident prevention by identifying systemic patterns and reducing toil
  • Partner with product engineering teams to diagnose reliability gaps
  • Define and champion operational excellence best practices across engineering
  • Embed the company's cultural principles across global teams

Skills

Mentoring engineers
Team collaboration
Growth mindset
Go/Python coding

Education

Bachelor's Degree in Computer Science/Engineering

Tools

Datadog
New Relic
Grafana
PagerDuty
Terraform

Job description

the company is hiring a Senior Software Engineer II to join the OPX team within the Developer Experience organization. You will design and build automated reliability and self-healing systems at scale, delivering platform tooling that engineers across the company adopt for their services.

You will own incident management tooling, evolve observability infrastructure with SLOs and real-time signals, and contribute to AI-driven automation that reduces toil and speeds delivery.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer - Cloud Observability & Automation
Senior Site Reliability Engineer - Cloud Observability & Automation

Omilia • Greater London

On-site
GBP 90,000 - 120,000
Fixed compensation
Long-term vacation
Professional growth
+3
Senior Reliability Engineer – Platform & Observability
Senior Reliability Engineer – Platform & Observability

One Identity, LLC • United Kingdom

Remote
GBP 90,000 - 140,000
Senior Software Engineer — Production Reliability & Observability
Senior Software Engineer — Production Reliability & Observability

One Identity • United Kingdom

On-site
GBP 90,000 - 130,000
Senior Platform SRE: Build Reliability Platform & Observability
Senior Platform SRE: Build Reliability Platform & Observability

United States Digital Space LLC • Greater London

Hybrid
GBP 110,000 - 170,000
Competitive salary
Flexible benefits package
Private medical cover for you and your
+4
Senior Platform SRE: Scale Reliability & Observability
Senior Platform SRE: Scale Reliability & Observability

IG Group • City Of London

Hybrid
GBP 110,000 - 140,000
Competitive salary
Flexible benefits (12%)
Private medical cover
+3
Reliability Engineer for Observability Platform
Reliability Engineer for Observability Platform

Coralogix • Greater London

On-site
GBP 90,000 - 130,000
Senior Software Engineer II — Secure, Scalable Cloud Apps
Senior Software Engineer II — Secure, Scalable Cloud Apps

RELX • Greater London

On-site
GBP 90,000 - 120,000
Performance & Observability Engineer — AI-Driven Reliability
Performance & Observability Engineer — AI-Driven Reliability

Herbert Smith Freehills Kramer • City Of London

On-site
Senior Observability Engineer — Scale Reliability & Telemetry
Senior Observability Engineer — Scale Reliability & Telemetry

IG Group • City Of London

On-site
GBP 110,000 - 150,000
Competitive salary
Private medical cover
Flexible benefits package
+1
Senior SRE: Cloud Reliability & Observability
Senior SRE: Cloud Reliability & Observability

Doist • United Kingdom

On-site
GBP 90,000 - 120,000