Director of Site Reliability Engineering

EPAM Systems

Greater London

Hybrid

GBP 180,000 - 240,000

Full time

32 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

ESPP
Life Assurance
Income protection
Critical illness cover
Private medical
Dental care
Employee Assistance Program
Pension plan
Cyclescheme
Techscheme
Season ticket loans
Free Wednesday lunch in-office
On-site massages
Regular social events
In-house training and certifications
Discretionary annual bonus program
Long-Term Incentive (LTI) Program

Job summary

EPAM Systems in London, United Kingdom, is seeking a Director of Site Reliability Engineering to lead a global SRE organization in a hybrid work setting. The role focuses on reliability, operational excellence, and governance across mission‑critical platforms, with an emphasis on AI‑enabled automation and improving engineering standards.

The successful candidate will drive resilience, define KPIs, and collaborate across product, platform, operations, and security teams to embed reliability

Qualifications

  • Strong background in Site Reliability Engineering, DevOps, or platform operations in complex, distributed environments.
  • Expertise in observability platforms, troubleshooting distributed systems, and telemetry‑driven insights.
  • Hands‑on experience with automation, Infrastructure as Code (Terraform or CloudFormation), and CI/CD practices.
  • Deep understanding of incident management processes, ITSM standards, and ITIL principles.
  • Knowledge of resilience design patterns, high availability, and fault‑tolerant architectures.
  • Familiarity with AI/ML‑driven approaches for operational efficiency and system reliability.
  • Ability to lead transformation, influence across teams, and foster continuous improvement in culture.

Responsibilities

  • Lead and scale a global SRE organization, focusing on engineering excellence and team empowerment
  • Collaborate with product, platform, operations, and security teams to embed reliability within SDLC practices
  • Define and monitor KPIs for system reliability, performance, and operational efficiency
  • Advance automation, Infrastructure as Code approaches, and promote self‑healing systems using AI/ML techniques
  • Develop robust incident management frameworks and lead major incident response activities for critical systems
  • Implement blameless postmortems and deliver systemic improvements across production environments
  • Establish observability strategies with standardized tooling for metrics, logs, and tracing to support distributed systems
  • Adopt and enforce SRE practices, including SLIs, SLOs, SLAs, and error budgets across services
  • Drive resilience strategies with highly available architectures and disaster recovery readiness
  • Champion an automation‑first culture, leveraging CI/CD pipelines and operational tooling to reduce manual processes

Job description

We're looking for a Director of Site Reliability Engineering to join our team in London, United Kingdom in a hybrid working mode. This role is responsible for driving reliability engineering and operational excellence across global technology platforms while leading the adoption of AI-enabled solutions for automation and efficiency. The position combines strategic leadership with hands‑on governance to ensure highly available, resilient systems that align with business and regulatory requirements. As a technology thought leader, you will influence engineering standards, enhance operational frameworks, and foster a culture of continuous improvement across mission‑critical environments.

Responsibilities
  • Lead and scale a global SRE organization, focusing on engineering excellence and team empowerment
  • Collaborate with product, platform, operations, and security teams to embed reliability within SDLC practices
  • Define and monitor KPIs for system reliability, performance, and operational efficiency
  • Advance automation, Infrastructure as Code approaches, and promote self‑healing systems using AI/ML techniques
  • Develop robust incident management frameworks and lead major incident response activities for critical systems
  • Implement blameless postmortems and deliver systemic improvements across production environments
  • Establish observability strategies with standardized tooling for metrics, logs, and tracing to support distributed systems
  • Adopt and enforce SRE practices, including SLIs, SLOs, SLAs, and error budgets across services
  • Drive resilience strategies with highly available architectures and disaster recovery readiness
  • Champion an automation‑first culture, leveraging CI/CD pipelines and operational tooling to reduce manual processes
Requirements
  • Strong background in Site Reliability Engineering, DevOps, or platform operations in complex, distributed environments
  • Expertise in observability platforms, troubleshooting distributed systems, and telemetry‑driven insights
  • Hands‑on experience with automation, Infrastructure as Code (Terraform or CloudFormation), and CI/CD practices
  • Deep understanding of incident management processes, ITSM standards, and ITIL principles
  • Knowledge of resilience design patterns, high availability, and fault‑tolerant architectures
  • Familiarity with AI/ML‑driven approaches for operational efficiency and system reliability
  • Ability to lead transformation, influence across teams, and foster continuous improvement in culture
Nice to have
  • Experience in financial services or other highly regulated, mission‑critical environments
  • Certifications in cloud technologies, such as AWS
  • Exposure to AIOps platforms or advanced observability tooling
We offer
  • EPAM Employee Stock Purchase Plan (ESPP)
  • Protection benefits including life assurance, income protection and critical illness cover
  • Private medical insurance and dental care
  • Employee Assistance Program
  • Competitive group pension plan
  • Cyclescheme, Techscheme and season ticket loans
  • Various perks such as free Wednesday lunch in-office, on‑site massages and regular social events
  • Learning and development opportunities including in‑house training and coaching, professional certifications, and courses
  • If otherwise eligible, participation in the discretionary annual bonus program
  • If otherwise eligible and hired into a qualifying level, participation in the discretionary Long-Term Incentive (LTI) Program
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Manager - Environment Strategy
Site Reliability Manager - Environment Strategy

EPAM Systems • Greater London

Hybrid
GBP 110,000 - 150,000
EPAM ESPP
Life assurance
Income protection
+8
Senior Site Reliability Engineer
Senior Site Reliability Engineer

LSEG • Nottingham

On-site
GBP 70,000 - 90,000
Healthcare
Retirement planning
Paid volunteering days
+1
SRE Architect
SRE Architect

Hitachi • Greater London

On-site
GBP 42,000 - 70,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

LSEG • Nottingham

On-site
GBP 80,000 - 100,000
Healthcare
Retirement Planning
Paid Volunteering Days
+1
Site Reliability Engineering Manager
Site Reliability Engineering Manager

Gravitas Recruitment Group (Global) Ltd • Greater London

Hybrid
GBP 75,000 - 100,000
Senior SRE
Senior SRE

Pulse Recruit • Greater London

Hybrid
GBP 65,000 - 85,000
Site Reliability Engineer – NS London
Site Reliability Engineer – NS London

BAE Systems • Greater London

Hybrid
GBP 45,000 - 70,000
Hybrid working flexibility
On-call allowances
Overtime benefits
Site Reliability Engineer - NS London
Site Reliability Engineer - NS London

BAE Systems Digital Intelligence • Greater London

Hybrid
GBP 50,000 - 70,000
Hybrid working environment
On-call allowances
Overtime benefits for night shifts
SRE Director — AI-Driven Reliability & Scale
SRE Director — AI-Driven Reliability & Scale

EPAM Systems • Greater London

Hybrid
GBP 180,000 - 240,000
ESPP
Life Assurance
Income protection
+14
Lead Site Reliability Engineer
Lead Site Reliability Engineer

McLaren Automotive Ltd • Woking

On-site
GBP 70,000 - 90,000
25 days holiday plus bank holidays
Enhanced company pension scheme
Discretionary annual bonus
+6