Site Reliability Engineer – Cloud, Observability & Resilience

IMG

Greater London

Hybrid

GBP 70,000 - 110,000

Full time

3 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

IMG is seeking a Site Reliability Engineer to design, build, operate, and continuously improve resilient platforms underpinning our digital, cloud, and broadcast-adjacent services. The role combines strong infrastructure and software engineering with an operational mindset to embed reliability across live, business-critical environments.

The successful candidate will enhance service reliability, observability, incident response, automation, and disaster recovery readiness while collaborating

Qualifications

  • Proven experience in a Site Reliability Engineer, DevOps Engineer, Platform Engineer, or similar role.
  • Strong knowledge of Linux and operating system fundamentals.
  • Hands-on experience with cloud platforms such as AWS, Azure, or Google Cloud.
  • Experience with containerisation and orchestration technologies such as Docker and Kubernetes.
  • Strong experience with CI/CD tooling and modern software delivery practices.
  • Hands-on experience with Infrastructure as Code tools such as Terraform or CloudFormation.
  • Experience with monitoring, logging, and alerting tooling, and with designing actionable observability solutions.
  • Solid understanding of networking, security, system architecture, and distributed systems principles.
  • Strong scripting or programming capability in Python, Bash, or similar languages.
  • Experience working in high-availability, live production, or other business-critical operational environments.
  • Strong troubleshooting skills, calm decision-making under pressure, and a continuous improvement mindset.
  • Excellent communication and collaboration skills, including the ability to work effectively with technical and non-technical stakeholders.
  • Experience supporting media, broadcast, streaming, or live event platforms.
  • Familiarity with incident management, postmortem practice, and error-budget based operational models.
  • Experience with resilience engineering, multi-site failover, and disaster recovery testing.
  • Exposure to event-driven or low-latency systems, media transport, or hybrid on-prem/cloud architectures.
  • Understanding of compliance, operational risk management, and support processes in client-facing environments.

Responsibilities

  • Design, build, and maintain reliable, scalable infrastructure and platform services across on-premises and cloud environments.
  • Improve service availability, latency, performance, and operational efficiency through reliability practices.
  • Build and enhance observability across services and infrastructure, including monitoring, logging, alerting, dashboards, and service health indicators.
  • Define and maintain SLIs, SLOs, alerting standards, and operational runbooks for critical services.
  • Automate infrastructure provisioning, configuration, deployment, and recovery processes using IaC and scripting.
  • Partner with software, platform, broadcast engineering, and operational teams to improve release quality, resilience, and supportability.
  • Act as an escalation point for production incidents, guiding diagnosis, mitigation, communication, and post-incident follow-up.
  • Drive root cause analysis and corrective actions following incidents to prevent recurrence.
  • Support design, testing, and documentation of high availability, backup, failover, and disaster recovery arrangements.
  • Help enforce security, access control, patching, and operational best practices across infrastructure and services.
  • Optimise system capacity, cost, and performance across environments.
  • Produce and maintain clear technical documentation, operational procedures, and handover materials.
  • Support live event workflows where reliability and rapid response are essential.
  • Contribute to technical planning for new services, migrations, and platform enhancements.

Skills

Linux fundamentals
Cloud platforms
Docker/Kubernetes
CI/CD tooling
Infrastructure as Code
Observability/monitoring
Scripting (Python/Bash)
HA/Disaster recovery
Incident management
Communication & collaboration

Tools

Terraform
CloudFormation
Prometheus/Grafana
ELK/Logging
Docker
Kubernetes
Jenkins
Git

Job description

IMG is seeking a Site Reliability Engineer to design, build, operate, and continuously improve resilient platforms underpinning our digital, cloud, and broadcast-adjacent services. The role combines strong infrastructure and software engineering with an operational mindset to embed reliability across live, business-critical environments.

The successful candidate will enhance service reliability, observability, incident response, automation, and disaster recovery readiness while collaborating

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer - Live Ops & Cloud Resilience
Site Reliability Engineer - Live Ops & Cloud Resilience

World Wrestling Entertainment, Inc. • Greater London

Hybrid
GBP 70,000 - 110,000
Site Reliability Engineer — Cloud & Live Ops (Hybrid)
Site Reliability Engineer — Cloud & Live Ops (Hybrid)

IMG • Uxbridge

Hybrid
GBP 70,000 - 100,000
Senior SRE, Observability & Cloud Reliability
Senior SRE, Observability & Cloud Reliability

United States Digital Space LLC • Greater London

Hybrid
GBP 120,000 - 170,000
Hybrid work up to 3 days per week
Site Reliability Engineer - Data Platform Reliability & Automation
Site Reliability Engineer - Data Platform Reliability & Automation

Intuition IT Solutions Ltd • Ouzlewell Green

On-site
GBP 70,000 - 110,000
Senior Site Reliability Engineer - Cloud Observability & Automation
Senior Site Reliability Engineer - Cloud Observability & Automation

Omilia • Greater London

On-site
GBP 90,000 - 120,000
Fixed compensation
Long-term vacation
Professional growth
+3
Senior Site Reliability Engineer — Automation & Reliability Leader
Senior Site Reliability Engineer — Automation & Reliability Leader

RELX Group • Greater London

On-site
GBP 90,000 - 140,000
Comprehensive Pension Plan
Generous vacation entitlement
Family leave (Maternity/Paternity/Adop
+1
Senior SRE: Cloud Reliability & Observability
Senior SRE: Cloud Reliability & Observability

Renesas Electronics Corp. • Cambridge

On-site
GBP 90,000 - 120,000
Site Reliability Engineer
Site Reliability Engineer

The Business Connection Group • Greater London

Remote
GBP 130,000 - 150,000
Site Reliability Engineer - Hybrid, Observability
Site Reliability Engineer - Hybrid, Observability

bet365 Group • United Kingdom

Hybrid
GBP 75,000 - 110,000
Eye care and Flu Vaccinations
Life Assurance
Site Reliability Engineer (DV Security Clearance)
Site Reliability Engineer (DV Security Clearance)

Onyx-Conseil • Manchester

On-site
GBP 90,000 - 120,000