SRE Director — AI-Driven Reliability & Scale

EPAM Systems

Greater London

Hybrid

GBP 180,000 - 240,000

Full time

20 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

ESPP
Life Assurance
Income protection
Critical illness cover
Private medical
Dental care
Employee Assistance Program
Pension plan
Cyclescheme
Techscheme
Season ticket loans
Free Wednesday lunch in-office
On-site massages
Regular social events
In-house training and certifications
Discretionary annual bonus program
Long-Term Incentive (LTI) Program

Job summary

EPAM Systems in London, United Kingdom, is seeking a Director of Site Reliability Engineering to lead a global SRE organization in a hybrid work setting. The role focuses on reliability, operational excellence, and governance across mission‑critical platforms, with an emphasis on AI‑enabled automation and improving engineering standards.

The successful candidate will drive resilience, define KPIs, and collaborate across product, platform, operations, and security teams to embed reliability

Qualifications

  • Strong background in Site Reliability Engineering, DevOps, or platform operations in complex, distributed environments.
  • Expertise in observability platforms, troubleshooting distributed systems, and telemetry‑driven insights.
  • Hands‑on experience with automation, Infrastructure as Code (Terraform or CloudFormation), and CI/CD practices.
  • Deep understanding of incident management processes, ITSM standards, and ITIL principles.
  • Knowledge of resilience design patterns, high availability, and fault‑tolerant architectures.
  • Familiarity with AI/ML‑driven approaches for operational efficiency and system reliability.
  • Ability to lead transformation, influence across teams, and foster continuous improvement in culture.

Responsibilities

  • Lead and scale a global SRE organization, focusing on engineering excellence and team empowerment
  • Collaborate with product, platform, operations, and security teams to embed reliability within SDLC practices
  • Define and monitor KPIs for system reliability, performance, and operational efficiency
  • Advance automation, Infrastructure as Code approaches, and promote self‑healing systems using AI/ML techniques
  • Develop robust incident management frameworks and lead major incident response activities for critical systems
  • Implement blameless postmortems and deliver systemic improvements across production environments
  • Establish observability strategies with standardized tooling for metrics, logs, and tracing to support distributed systems
  • Adopt and enforce SRE practices, including SLIs, SLOs, SLAs, and error budgets across services
  • Drive resilience strategies with highly available architectures and disaster recovery readiness
  • Champion an automation‑first culture, leveraging CI/CD pipelines and operational tooling to reduce manual processes

Job description

EPAM Systems in London, United Kingdom, is seeking a Director of Site Reliability Engineering to lead a global SRE organization in a hybrid work setting. The role focuses on reliability, operational excellence, and governance across mission‑critical platforms, with an emphasis on AI‑enabled automation and improving engineering standards.

The successful candidate will drive resilience, define KPIs, and collaborate across product, platform, operations, and security teams to embed reliability

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Director of Site Reliability Engineering
Director of Site Reliability Engineering

EPAM Systems • Greater London

Hybrid
GBP 180,000 - 240,000
ESPP
Life Assurance
Income protection
+14
Senior SRE Manager - Hybrid London (AWS & Automation)
Senior SRE Manager - Hybrid London (AWS & Automation)

EPAM Systems • Greater London

Hybrid
GBP 110,000 - 150,000
EPAM ESPP
Life assurance
Income protection
+8
SRE Manager: Reliability & Incident Leadership (Hybrid London)
SRE Manager: Reliability & Incident Leadership (Hybrid London)

Gravitas Recruitment Group (Global) Ltd • Greater London

Hybrid
GBP 75,000 - 100,000
SRE & Reliability Lead — AI-Ops & Observability
SRE & Reliability Lead — AI-Ops & Observability

LexisNexis Risk Solutions • Carshalton

On-site
GBP 90,000 - 130,000
SRE for AI-Driven Financial Infrastructure
SRE for AI-Driven Financial Infrastructure

United States Digital Space LLC • Greater London

On-site
GBP 90,000 - 130,000
Daily catered lunches
Modern office environment
Tech talks and knowledge sharing
SRE & Operations Manager — AI-Driven Reliability
SRE & Operations Manager — AI-Driven Reliability

LexisNexis Risk Solutions • Sutton

Hybrid
GBP 90,000 - 130,000
SRE: Drive Reliability & Automation at Scale
SRE: Drive Reliability & Automation at Scale

Xpertise Recruitment • West Drayton

On-site
GBP 60,000 - 80,000
site reliability engineer
site reliability engineer

Enfint • Greater London

On-site
GBP 120,000 - 180,000
SRE Architect
SRE Architect

Hitachi • Greater London

On-site
GBP 42,000 - 70,000
SRE
SRE

Technopride Ltd • Hove

Hybrid
GBP 60,000 - 80,000