Site Reliability Manager - Environment Strategy

EPAM Systems

Greater London

Hybrid

GBP 110,000 - 150,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

EPAM Systems is seeking a Site Reliability Manager in London to lead a team focused on environment strategy, automation, patch governance and operational reliability for AWS-based platforms.

You will define roadmaps, drive best practices, and ensure resilience across production and non-production systems while managing budgets and schedules.

Qualifications

  • 8+ years of experience in site reliability, DevOps or platform engineering.
  • Strong knowledge of AWS platforms and cloud infrastructure principles.
  • Experience implementing automation using Terraform or CloudFormation.
  • Experience with CI/CD pipelines and monitoring with CloudWatch, Prometheus, Grafana.
  • Background in risk management, compliance frameworks and incident response processes.
  • Excellent leadership, communication and stakeholder management skills.

Responsibilities

  • Define and own the vision and roadmap for site reliability and environment strategy.
  • Lead, mentor and develop a team of DevOps and environment engineers.
  • Set and enforce standards for environment provisioning, lifecycle management and patch governance.
  • Drive adoption of Infrastructure as Code and automation-first practices across environments.
  • Oversee monitoring, alerting and operational readiness to meet availability and performance objectives.
  • Partner with security, engineering and operations teams to reduce risk and ensure compliance.
  • Build continuous improvement processes and establish key performance metrics for reliability.

Skills

AWS
DevOps leadership
CI/CD
Monitoring
Infrastructure as Code
Security and compliance
Leadership and stakeholders

Tools

Terraform
CloudFormation
CloudWatch
Prometheus
Grafana

Job description

We're looking for a Site Reliability Manager to join our team in London, United Kingdom in a hybrid working mode. In this role, you will lead a team focused on environment strategy, automation, patch governance and operational reliability for AWS-based platforms. Your responsibilities include setting roadmaps, driving best practices and delivering consistency across complex environments. You will manage people, budgets and schedules while ensuring strong technical standards, compliance and resilience across all production and non-production systems.

Responsibilities
  • Define and own the vision and roadmap for site reliability and environment strategy
  • Lead, mentor and develop a team of DevOps and environment engineers
  • Set and enforce standards for environment provisioning, lifecycle management and patch governance
  • Drive adoption of Infrastructure as Code and automation-first practices across environments
  • Oversee monitoring, alerting and operational readiness to meet availability and performance objectives
  • Partner with security, engineering and operations teams to reduce risk and ensure compliance
  • Build continuous improvement processes and establish key performance metrics for reliability
Requirements
  • 8+ years of experience in site reliability, DevOps or platform engineering roles
  • Strong knowledge of AWS platforms and cloud infrastructure principles
  • Proven ability to manage teams and operational priorities in complex environments
  • Experience implementing automation using Terraform or CloudFormation
  • Experience with CI/CD pipelines and operational monitoring with tools such as CloudWatch, Prometheus and Grafana
  • Understanding of OS patching for Windows Server and RHEL as part of governance practices
  • Background in risk management, compliance frameworks and incident response processes
  • Excellent leadership, communication and stakeholder management skills
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineering Manager
Site Reliability Engineering Manager

Gravitas Recruitment Group (Global) Ltd • Greater London

Hybrid
GBP 75,000 - 100,000
Site Reliability Engineer
Site Reliability Engineer

SR2 | Socially Responsible Recruitment | Certified B Corporation • Slough

Hybrid
GBP 65,000 - 90,000
Site Reliability Engineer
Site Reliability Engineer

Incite-Insight.co.uk • West of England

On-site
GBP 70,000 - 95,000
Director of Site Reliability Engineering
Director of Site Reliability Engineering

EPAM Systems • Greater London

Hybrid
GBP 180,000 - 240,000
ESPP
Life Assurance
Income protection
+14
Site Reliability Engineer
Site Reliability Engineer

Falconsmartit • Hove

Hybrid
GBP 90,000 - 130,000
Site Reliability Engineer (SRE) – Cloud Platforms
Site Reliability Engineer (SRE) – Cloud Platforms

Talenzon group • Greater London

Hybrid
GBP 70,000 - 110,000
Site Reliability Engineer
Site Reliability Engineer

ReVybe IT Recruitment Limited • Greater London

Hybrid
GBP 51,000 - 85,000
Bonus
Benefits
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Spectrum IT Recruitment • Southampton

Hybrid
GBP 70,000 - 110,000
Life Insurance 4x salary
Private Medical Insurance
Employee Assistance Programme
+2
IT Infrastructure, Platform & Engineering Manager
IT Infrastructure, Platform & Engineering Manager

SoCode Recruitment • Huntingdon

Hybrid
GBP 90,000 - 140,000
Site Reliability Engineer - Negotiable
Site Reliability Engineer - Negotiable

Alchemy • Reading

Hybrid
GBP 60,000 - 80,000
Competitive salary
Healthcare benefits