Director, Platform Operations & Reliability

Epsilon

Greater London

Hybrid

GBP 90,000 - 120,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Competitive compensation
Comprehensive benefits package
Opportunities for career advancement
Hybrid working opportunities

Job summary

A subsidiary of a global technology company is seeking a System and Platform Operations Director to manage the reliability and support of production systems. This leadership role involves orchestrating team operations and ensuring customer support excellence. Candidates should have over 5 years of experience in Site Reliability Engineering, a strong understanding of containerization technologies, and effective communication skills. The company promotes employee development and offers a hybrid working environment.

Qualifications

  • At least 5 years of hands-on experience in Site Reliability focused positions.
  • Strong knowledge of containerization technologies (Docker, Kubernetes).
  • Experience with infrastructure as code (Terraform).
  • Solid understanding of networking, security, and system architecture.
  • Proficient in scripting languages (Java, Golang, Python, Bash, or similar).
  • Experience with monitoring and observability tools (DataDog, Prometheus, Grafana).
  • Knowledge of database management systems (PostgreSQL, Bigtable).
  • Understanding of API and microservices architecture.
  • Strong people leadership skills with at least a year in leading high-performance technical teams.

Responsibilities

  • Establish and manage operational practices for support model design.
  • Own incident management processes and on-call response.
  • Work with Engineering, Product, Delivery, and Security teams on system reliability.
  • Execute Service Management processes to deliver high support levels.
  • Identify capabilities needed to meet current and emerging business needs.

Skills

Site Reliability Engineering
Containerization technologies (Docker, Kubernetes)
Infrastructure as code (Terraform)
Scripting languages (Java, Golang, Python, Bash)
Monitoring tools (DataDog, Prometheus, Grafana)
Database management systems (PostgreSQL, Bigtable)
API and microservices architecture
IT Service Management principles
Communication and written skills
Service Delivery strategies

Education

Bachelor's degree or equivalent

Job description

A subsidiary of a global technology company is seeking a System and Platform Operations Director to manage the reliability and support of production systems. This leadership role involves orchestrating team operations and ensuring customer support excellence. Candidates should have over 5 years of experience in Site Reliability Engineering, a strong understanding of containerization technologies, and effective communication skills. The company promotes employee development and offers a hybrid working environment.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Hybrid Platform Operations & Reliability Director
Hybrid Platform Operations & Reliability Director

Publicis Groupe Holdings B.V • Greater London

Hybrid
GBP 80,000 - 110,000
Competitive compensation
Great benefits package
Endless opportunities for career advancement
+1
Global Director, Cloud Reliability & SRE
Global Director, Cloud Reliability & SRE

Omnicell • Manchester

On-site
GBP 150,000 - 190,000
DevOps & SRE Lead: Platform Reliability & CI/CD
DevOps & SRE Lead: Platform Reliability & CI/CD

OCU • Preston

On-site
GBP 70,000 - 90,000
SRE Manager: Reliability & Incident Leadership (Hybrid London)
SRE Manager: Reliability & Incident Leadership (Hybrid London)

Gravitas Recruitment Group (Global) Ltd • Greater London

Hybrid
GBP 75,000 - 100,000
Global SVP, Platform Operations & AI-Driven Reliability
Global SVP, Platform Operations & AI-Driven Reliability

Trading Technologies International • Greater London

Hybrid
GBP 180,000 - 280,000
Hybrid in-office 3 days/week
25 days PTO per year
Volunteer day
+3
Senior SRE: Observability & Platform Reliability
Senior SRE: Observability & Platform Reliability

Selby Jennings • Greater London

On-site
GBP 70,000 - 90,000
Senior DevOps Engineer - Platform & SaaS Reliability
Senior DevOps Engineer - Platform & SaaS Reliability

Few&Far • United Kingdom

On-site
GBP 98,000 - 137,000
Production Reliability Lead
Production Reliability Lead

Complexio • Bournemouth

On-site
GBP 90,000 - 120,000
Platform Operations Lead – Java & AWS Reliability
Platform Operations Lead – Java & AWS Reliability

develop • Greater London

Hybrid
GBP 72,000 - 88,000
Hybrid working (3 days per week in the
Annual performance bonus
Career progression opportunities
SRE Director — AI-Driven Reliability & Scale
SRE Director — AI-Driven Reliability & Scale

EPAM Systems • Greater London

Hybrid
GBP 180,000 - 240,000
ESPP
Life Assurance
Income protection
+14