Site Reliability Engineer

Incite-Insight.co.uk

West of England

On-site

GBP 70,000 - 95,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Incite-Insight.co.uk is seeking an experienced Site Reliability/Platform Engineer to automate and modernise its infrastructure for large-scale environments.

You'll build Python-based automation around incident management, runbooks, and routine tasks, and integrate monitoring and ITSM platforms via APIs, to reduce toil and speed up incident response. This role offers significant autonomy and an opportunity to shape the SRE capabilities within the organisation.

Qualifications

  • What you'll bring includes commercial SRE/Platform Eng/DevOps experience and strong Python automation.
  • Hands-on knowledge of monitoring and alerting frameworks and ITSM integrations.

Responsibilities

  • Build Python-based automation for incident handling and operational tasks.
  • Integrate monitoring, ITSM and infrastructure components via APIs.
  • Develop internal tools, dashboards and CLI utilities to speed issue resolution.
  • Automate existing manual processes to reduce toil and improve reliability.

Skills

SRE / Platform Eng
DevOps
Python automation
Monitoring (Prometheus Grafana)
API integration
Incident management
On-call experience

Tools

Prometheus
Grafana
OpenTelemetry
ServiceNow
Jira Service Management
APIs

Job description

Site Reliability / Platform Engineer

We are recruiting for a growing technology infrastructure business that is building a new operational capability to support large-scale, high-performance computing environments.
This is an excellent opportunity for an experienced Site Reliability Engineer or Platform Engineer who enjoys automating things rather than repeatedly fixing them manually.
The role sits at the intersection of infrastructure, operations and software engineering. You will use Python and modern automation techniques to improve reliability, reduce manual workload and make incident response faster and more effective.

What you'll be doing

You'll build Python-based automation around incident management, operational runbooks and routine infrastructure tasks.
You'll integrate monitoring, infrastructure and ITSM platforms through APIs, helping improve the quality of alerts through better correlation, enrichment, suppression and deduplication.
You'll also develop internal tools, command-line utilities, dashboards and potentially ChatOps capabilities that allow operational teams to resolve issues more quickly.
A major part of the role will be taking existing operational processes and asking:
"Why are we still doing this manually?"
You'll then design a safe, controlled and auditable way of automating it.

What we're looking for

You should have good commercial experience in Site Reliability Engineering, Platform Engineering, DevOps or production infrastructure operations, together with strong hands-on Python automation skills.
You'll also need experience with:

  • Monitoring and observability tools such as Prometheus, Grafana or similar
  • Production incident management and/or on-call environments
  • Automating operational runbooks and repetitive infrastructure processes
  • APIs and systems integration
  • Version-controlled automation and operational tooling
Experience with any of the following would be particularly useful:

ServiceNow, Halo, Jira Service Management, OpenTelemetry, distributed tracing, Slack/Teams automation, datacentre or colocation environments, GPU infrastructure, DCIM, IPAM, virtualisation platforms or LLM-assisted operational automation.

This is not an AI/ML development position. We're looking for someone who understands production infrastructure and can use software engineering and automation to make that infrastructure more reliable.
You'll be joining a growing organisation where you'll have considerable autonomy and the opportunity to help shape the SRE and operational automation capability rather than simply inherit an established environment.

Salary: TBC
Location / hybrid working: TBC

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

ReVybe IT Recruitment Limited • Greater London

Hybrid
GBP 51,000 - 85,000
Bonus
Benefits
Site Reliability Engineer
Site Reliability Engineer

Falconsmartit • Hove

Hybrid
GBP 90,000 - 130,000
Senior Site Reliability Engineer - Selby Jennings
Senior Site Reliability Engineer - Selby Jennings

eFinancialCareers • Greater London

On-site
GBP 90,000 - 130,000
Senior Platform Engineer / SRE
Senior Platform Engineer / SRE

Myn • Greater London

Hybrid
GBP 90,000 - 120,000
SRE Platform Engineer: Python Automation & Incident Response
SRE Platform Engineer: Python Automation & Incident Response

Incite-Insight.co.uk • West of England

On-site
GBP 70,000 - 95,000
Site Reliability Engineer (SRE) / Platform Engineer
Site Reliability Engineer (SRE) / Platform Engineer

Adecco • City Of London

Hybrid
Hybrid work arrangement
Competitive day rate
London-based contract
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Xpertise Recruitment • West Drayton

On-site
GBP 60,000 - 80,000
Site Reliability Engineer - Negotiable
Site Reliability Engineer - Negotiable

Alchemy • Reading

Hybrid
GBP 60,000 - 80,000
Competitive salary
Healthcare benefits
Site Reliability Engineer
Site Reliability Engineer

慨正橡扯 • Manchester

Hybrid
GBP 60,000 - 80,000
SRE
SRE

Technopride Ltd • Hove

Hybrid
GBP 60,000 - 80,000