Senior Site Reliability Engineer II

Talent Octopusventures

England

On-site

GBP 79,000 - 132,000

Full time

4 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Health benefits
401(k) with match
Wellbeing platform
Disability & life insurance
Family benefits
HSA & commuter accounts
Volunteer leave

Job summary

CamWebDir is seeking a Site Reliability Engineer to improve the reliability, availability, and performance of our production systems. You will provide technical input into project plans, schedules, methodologies, and operational readiness activities across multiple environments.

You will lead incident responses and postmortems, perform root‑cause analysis, coordinate with engineering, security, and operations, and drive automation, Kubernetes, monitoring, and system modernization to enhance

Qualifications

  • 5+ years of experience in SRE/related field.
  • Bachelor’s degree in Engineering, Computer Science, IT, or equivalent.
  • Experience supporting highly available production systems.
  • Experience leading incident reviews, postmortems, RCA, and remediation planning.
  • Experience across infrastructure, application, security, and operations teams.
  • Strong problem‑solving, analytical, organizational, and communication skills.
  • Ability to manage multiple priorities and deliver in a fast‑paced environment.

Responsibilities

  • Lead and participate in incident response, postmortems, RCA, and gap assessments.
  • Identify reliability, availability, performance, security, and operational risks.
  • Develop, prioritize, and track corrective and preventive actions.
  • Collaborate with engineering, security, support, and business stakeholders to resolve incidents and gaps.
  • Respond to system alerts and operational exceptions within enterprise systems and products.
  • Provide technical input into project plans, schedules, and readiness activities.
  • Support triage, planning, execution, documentation, and closure of changes and service requests.
  • Lead or contribute to projects involving cloud, Kubernetes, automation, monitoring, and modernization.
  • Improve production quality by creating new capabilities and remediating weaknesses.

Skills

SRE experience
Production systems
Incident reviews
Postmortems
Root-cause analysis
Remediation planning
Infra & App teams
Problem solving
Analytical
Communication
Prioritization
Execution

Education

Bachelor's degree in Engineering, Computer Science, Information Technology

Tools

Kubernetes
Linux
Windows
IaC
Python
Shell
PowerShell
Monitoring

Job description

About the Role:

The SRE role is responsible for improving the reliability, availability, performance, and operational quality of production systems. This role provides technical input into project plans, schedules, methodologies, and operational strategies across multiple system environments.

Job Functions
  • Lead and participate in incident response, postmortems, root-cause analysis, and gap assessments.
  • Identify reliability, availability, performance, security, and operational risks across production environments.
  • Develop, prioritize, and track corrective and preventive actions through completion.
  • Follow up with engineering, development, security, support, and business stakeholders to ensure timely resolution of incidents and identified gaps.
  • Respond to system-management alerts and operational exceptions within assigned enterprise systems and product offerings.
  • Provide technical input into project plans, schedules, implementation methodologies, and operational readiness activities.
  • Support the triage, planning, execution, documentation, and closure of changes, service requests, and operational tasks.
  • Lead or contribute to Operations Team projects involving cloud, on-premises infrastructure, security, Kubernetes, automation, monitoring, and system modernization.
  • Improve production quality and availability by creating new operational capabilities and remediating weaknesses in existing systems and processes.
Qualifications
  • 5+ years of experience in Site Reliability Engineering, Systems Engineering, DevOps, Infrastructure Engineering, or a related field.
  • Bachelor’s degree in Engineering, Computer Science, Information Technology, or equivalent professional experience.
  • Demonstrated experience supporting highly available production systems.
  • Experience leading incident reviews, postmortems, root-cause analysis, and remediation planning.
  • Experience working across infrastructure, application, security, and operations teams.
  • Strong problem-solving, analytical, organizational, and communication skills.
  • Ability to manage multiple priorities and drive work to completion in a fast-paced operational environment.
Technical Skills
  • Strong experience in Site Reliability Engineering (SRE), production operations, and IT service management processes including incident, problem, change, and service request management.
  • Hands‑on expertise with cloud and on‑premises infrastructure, Kubernetes, containerized workloads, virtualization, and distributed systems.
  • Advanced knowledge of Linux/UNIX and Windows environments, storage and file systems, including installation, configuration, troubleshooting, lifecycle management, backup, disaster recovery, and business continuity.
  • Experience with monitoring, alerting, logging, observability, and performance analysis, including the ability to analyze system diagnostics, logs, traces, resource utilization, and operational metrics.
  • Strong automation and infrastructure engineering skills, including Infrastructure as Code (IaC), configuration management, scripting (Python, Shell, PowerShell), system provisioning, deployments, remediation, and security risk mitigation.
Accountabilities
  • Monitor assigned environments, respond to alerts and incidents, diagnose system and performance issues, and coordinate escalation and recovery.
  • Track remediation activities and stakeholder commitments through completion to improve production quality, reliability, and availability.
  • Design and maintain automation, scripts, integrations, runbooks, and workflows for provisioning, health checks, deployments, remediation, and routine operations. Install, configure, troubleshoot, and support hardware, software, storage, network, cloud, Kubernetes, and other infrastructure services.
  • Establish logging, monitoring, alerting, metrics, and tracing standards; improve alert quality by reducing noise and ensuring alerts are actionable. Build dashboards and visualizations that communicate system health, availability, performance, capacity, service‑level objectives, and incident trends.
  • Develop and maintain recovery procedures and participate in disaster‑recovery, resilience, and business‑continuity exercises. Partner with development, operations, security, support teams, vendors, and stakeholders to coordinate work, resolve issues, and meet delivery commitments.
  • Lead or contribute to Operations Team projects from planning and implementation through documentation, transition to support, and closure.
  • Plan, risk‑assess, obtain approval for, implement, document, and close changes, service requests, and operational tasks.
  • Review and improve technical procedures, scripts, automation, and operational documentation while providing guidance to less‑experienced team members.
Working for you

We know that your wellbeing and happiness are key to a long and successful career. These are some of the benefits we are delighted to offer:

  • Health Benefits: Comprehensive, multi‑carrier program for medical, dental and vision benefits
  • Retirement Benefits: 401(k) with match and an Employee Share Purchase Plan
  • Wellbeing: Wellness platform with incentives, Headspace app subscription, Employee Assistance and Time‑off Programs
  • Short‑and‑Long Term Disability, Life and Accidental Death Insurance, Critical Illness, and Hospital Indemnity
  • Family Benefits, including bonding and family care leaves, adoption and surrogacy benefits
  • Health Savings, Health Care, Dependent Care and Commuter Spending Accounts
  • In addition to annual Paid Time Off, we offer up to two days of paid leave each to participate in Employee Resource Groups and to volunteer with your charity of choice

U.S. National Base Pay Range: $104,900 - $174,700. Geographic differentials may apply in some locations to better reflect local market rates.

This job is eligible for an annual incentive bonus.

We are an equal opportunity employer: qualified applicants are considered for and treated during employment without regard to race, color, creed, religion, sex, national origin, citizenship status, disability status, protected veteran status, age, marital status, sexual orientation, gender identity, genetic information, or any other characteristic protected by law.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal Site Reliability Engineer, Infrastructure Observability
Principal Site Reliability Engineer, Infrastructure Observability

United States Digital Space LLC • Greater London

On-site
GBP 120,000 - 170,000
Hybrid work up to 3 days per week
Site Reliability Engineer
Site Reliability Engineer

United States Digital Space LLC • Greater London

On-site
GBP 90,000 - 130,000
Daily catered lunches
Modern office environment
Tech talks and knowledge sharing
Site Reliability Engineer
Site Reliability Engineer

Insight International (UK) Ltd • Bournemouth

On-site
GBP 55,000 - 75,000
Site Reliability Engineer
Site Reliability Engineer

Biometric Talent Ltd • Manchester

On-site
GBP 40,000 - 65,000
Performance-Based Bonus
Pension Scheme
Hybrid Working
+2
Site Reliability Engineer
Site Reliability Engineer

SR2 | Socially Responsible Recruitment | Certified B Corporation • Slough

Hybrid
GBP 65,000 - 90,000
Site Reliability Engineer
Site Reliability Engineer

ScaleneWorks People Solutions LLP • Bournemouth

On-site
GBP 60,000 - 80,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Relx Plc • Greater London

Hybrid
GBP 90,000 - 120,000
Comprehensive Pension Plan
Generous vacation entitlement
Family leave (Maternity/Paternity/Adop
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Tenth Revolution Group • Knutsford

On-site
GBP 70,000 - 90,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Imanage • Belfast City District

On-site
GBP 85,000 - 120,000
Private medical insurance
Pension contributions matching
Annual performance bonus
+2
Sr. Site Reliability Engineer/ SWE
Sr. Site Reliability Engineer/ SWE

Visa • Basingstoke

Hybrid
GBP 70,000 - 120,000