site reliability engineer

Enfint

Greater London

On-site

GBP 120,000 - 180,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

EPAM приглашает опытного руководителя SRE в Лондон для ведения глобальной команды инженеров, ориентированной на устойчивость систем, автоматизацию и масштабируемость.

Задачи включают интеграцию надёжности в SDLC, сотрудничество с продуктовой, операционной и безопасной командами, а также развитие процессов инцидент-менеджмента и мониторинга.

Qualifications

  • Опыт работы в SRE/DevOps в распределённых системах.
  • Развертывание инфраструктуры как код (IaC) и CI/CD.
  • Управление инцидентами, знание ITIL.

Responsibilities

  • Руководить и масштабировать глобальную SRE-команду, ориентированную на инженерное качество.
  • Сотрудничать с продуктом, платформой, операциями и безопасностью, внедряя надёжность в SDLC.
  • Определять и мониторить KPI надёжности, производительности и эффективности.
  • Развивать автоматизацию, IaC и самоисцеление через AI/ML.
  • Разрабатывать устойчивые фреймворки инцидент-менеджмента и возглавлять реакции на крупные инциденты.
  • Внедрять безвиновные постмортемы и системные улучшения в продакшне.
  • Установить стратегии наблюдаемости с унифицированными инструментами для метрик, логов и трассировки.
  • Внедрять SRE-практики: SLI, SLO, SLA и бюджеты ошибок.
  • Продвигать стратегии устойчивости, высокую доступность и DR readiness.
  • Способствовать культуре автоматизации: CI/CD и эксплуатационные инструменты.

Skills

SRE leadership
DevOps
Distributed systems
Observability
Automation
IaC
CI/CD
Incident management
ITIL
AI for operations

Tools

Terraform
CloudFormation
CI/CD pipelines
Observability platforms

Job description

Описание

EPAM is a global provider of digital engineering, cloud, and AI-enabled transformation services, focusing on complex software product development and digital platform engineering.

Задачи
  • Lead and scale a global SRE organization, focusing on engineering excellence and team empowerment
  • Collaborate with product, platform, operations, and security teams to embed reliability within SDLC practices
  • Define and monitor KPIs for system reliability, performance, and operational efficiency
  • Advance automation, Infrastructure as Code approaches, and promote self-healing systems using AI/ML techniques
  • Develop robust incident management frameworks and lead major incident response activities for critical systems
  • Implement blameless postmortems and deliver systemic improvements across production environments
  • Establish observability strategies with standardized tooling for metrics, logs, and tracing to support distributed systems
  • Adopt and enforce SRE practices, including SLIs, SLOs, SLAs, and error budgets across services
  • Drive resilience strategies with highly available architectures and disaster recovery readiness
  • Champion an automation-first culture, leveraging CI/CD pipelines and operational tooling to reduce manual processes
Требования
  • Strong background in Site Reliability Engineering, DevOps, or platform operations in complex, distributed environments
  • Expertise in observability platforms, troubleshooting distributed systems, and telemetry-driven insights
  • Hands‑on experience with automation, Infrastructure as Code (Terraform or CloudFormation), and CI/CD practices
  • Deep understanding of incident management processes, ITSM standards, and ITIL principles
  • Knowledge of resilience design patterns, high availability, and fault‑tolerant architectures
  • Familiarity with AI/ML‑driven approaches for operational efficiency and system reliability
  • Ability to lead transformation, influence across teams, and foster continuous improvement in culture
  • Nice to have: Experience in financial services or other highly regulated, mission‑critical environments, Certifications in cloud technologies such as AWS, Exposure to AIOps platforms or advanced observability tooling
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Falconsmartit • Hove

Hybrid
GBP 90,000 - 130,000
SRE Architect (68019)
SRE Architect (68019)

Hitachi Digital Services • Greater London

On-site
GBP 90,000 - 150,000
project manager for enterprise technology initiatives
project manager for enterprise technology initiatives

Enfint • Greater London

On-site
GBP 90,000 - 130,000
Principal Site Reliability Engineer, Infrastructure Observability
Principal Site Reliability Engineer, Infrastructure Observability

United States Digital Space LLC • Greater London

Hybrid
GBP 120,000 - 170,000
Hybrid work up to 3 days per week
SRE Architect
SRE Architect

Hitachi • Greater London

On-site
GBP 42,000 - 70,000
SRE Architect (68019) (DEAI DS) Cloud & Data Engineering United Kingdom
SRE Architect (68019) (DEAI DS) Cloud & Data Engineering United Kingdom

Hitachids • Greater London

On-site
GBP 90,000 - 140,000
Principal Site Reliability Engineer, Infrastructure Observability
Principal Site Reliability Engineer, Infrastructure Observability

T. Rowe Price • Greater London

Hybrid
GBP 120,000 - 180,000
Hybrid work
On-call rotation
Senior Site Reliability Engineer
Senior Site Reliability Engineer

LSEG • Nottingham

On-site
GBP 70,000 - 90,000
Healthcare
Retirement planning
Paid volunteering days
+1
Site Reliability Engineer
Site Reliability Engineer

Insight International (UK) Ltd • Bournemouth

On-site
GBP 55,000 - 75,000
Lead SRE - AWS Platform
Lead SRE - AWS Platform

JPMorgan Chase & Co. • Glasgow

On-site
GBP 90,000 - 130,000