SRE-NOC Engineer: Master Incident Response & Reliability

Greenhouse Software, Inc.

United Kingdom

Remote

GBP 60,000 - 96,000

Full time

43 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

NiCE in the United Kingdom is seeking an experienced SRE – NOC to join our team. You will balance traditional NOC responsibilities with reliability engineering, focusing on 24/7 service reliability, incident response, and automation.

You’ll own runbooks, design alerting, build dashboards with Grafana, and work with cross‑functional teams to reduce toil. The role suits those who engineer solutions rather than only respond to alerts, with a focus on SLOs/SLIs and scalable platforms.

Qualifications

  • Experience with incident management and production support.
  • Familiarity with cloud infrastructure (AWS preferred).
  • Monitoring or alerting platforms.
  • Scripting or programming in Python, Bash, Go, or similar.
  • Understanding of networking fundamentals (DNS, TCP/IP, load balancing).

Responsibilities

  • Act as primary or escalation responder in a 24x7 on-call rotation.
  • Lead or support Major Incident (MI) response, including triage, mitigation and resolution.
  • Coordinate across Engineering, Infrastructure, Security, and Product teams.
  • Execute and improve runbooks, playbooks and escalation paths.
  • Drive blameless post-incident reviews (PIRs) and track corrective actions.
  • Own service health monitoring across infrastructure, applications and dependencies.
  • Design and maintain alerting strategies aligned with SLIs/SLOs.
  • Build dashboards using Grafana and other monitoring tools.
  • Automate repetitive operational tasks to reduce manual toil.
  • Improve MTTD and MTTR through tooling and processes.
  • Develop scripts/tools to support NOC/SRE workflows.
  • Implement self-healing and auto-remediation where possible.
  • Collaborate with engineering to improve system design for reliability.
  • Support and troubleshoot capacity planning and availability reviews.

Skills

Incident management
24x7 NOC experience
Python
Bash
Go
Networking fundamentals
Strong communication

Tools

AWS
Grafana

Job description

NiCE in the United Kingdom is seeking an experienced SRE – NOC to join our team. You will balance traditional NOC responsibilities with reliability engineering, focusing on 24/7 service reliability, incident response, and automation.

You’ll own runbooks, design alerting, build dashboards with Grafana, and work with cross‑functional teams to reduce toil. The role suits those who engineer solutions rather than only respond to alerts, with a focus on SLOs/SLIs and scalable platforms.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote SRE-NOC Engineer: Automate, Observe, Resolve
Remote SRE-NOC Engineer: Automate, Observe, Resolve

AI Chopping Block, Inc. • United Kingdom

Remote
GBP 65,000 - 105,000
NOC Engineer / SRE
NOC Engineer / SRE

Greenhouse Software, Inc. • United Kingdom

On-site
GBP 60,000 - 96,000
NOC Engineer / SRE
NOC Engineer / SRE

AI Chopping Block, Inc. • United Kingdom

Remote
GBP 65,000 - 105,000
SRE & Reliability Lead — AI-Ops & Observability
SRE & Reliability Lead — AI-Ops & Observability

LexisNexis Risk Solutions • Carshalton

On-site
GBP 90,000 - 130,000
Sr. Network Site Reliability Engineer (SREs)
Sr. Network Site Reliability Engineer (SREs)

Technopride Ltd • Greater London

On-site
GBP 70,000 - 90,000
Remote UK SRE: Observability & Reliability Lead
Remote UK SRE: Observability & Reliability Lead

Orex Nova, Inc. • United Kingdom

Remote
GBP 75,000 - 110,000
Fully remote UK
On-call compensation
Home-office stipend
SRE
SRE

Technopride Ltd • Hove

On-site
GBP 60,000 - 80,000
Senior SRE Engineer: Automation & Observability
Senior SRE Engineer: Automation & Observability

UK Health Security Agency • Liverpool

On-site
GBP 60,000 - 90,000
SRE & Operations Leader - Reliability at Scale
SRE & Operations Leader - Reliability at Scale

LexisNexis Risk Solutions • United Kingdom

Remote
GBP 90,000 - 130,000
Senior Network Engineer
Senior Network Engineer

IT WORLD LIMITED • England

On-site
GBP 70,000 - 78,000