Senior Site Reliability Engineer - Cloud Observability & Automation

Omilia

Greater London

On-site

GBP 90,000 - 120,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Fixed compensation
Long-term vacation
Professional growth
Global impact products
Great colleagues
Apple gear

Job summary

Omilia is seeking a Senior Site Reliability Engineer with Cloud platform experience to join our production reliability team. You will operate and maintain production clusters, develop observability solutions, and drive automation across monitoring, alerting, and runbooks.

You will collaborate with engineering and cloud teams to embed reliability and performance into the software delivery lifecycle and participate in on-call rotations, championing a culture of reliability and continuous

Qualifications

  • Bachelor's Degree or MS in Engineering or equivalent.
  • Experience in operating at least one container orchestration cluster (Kubernetes, Docker Swarm).
  • Experience developing or maintaining software for production services at scale.
  • Experience with ELK.
  • Experience with AWS.
  • Experience with Grafana/Prometheus stack.
  • Strong scripting skills (Bash, Python or Go).
  • Excellent communication skills.
  • Thinking out of the box and anticipating challenges. It is imperative we are not simply reactive; we must expect challenges and question technologies, procedures and thinking already in place. You will be expected to constantly review and challenge at all levels.
  • Versatility. We work with agile/lean methods. We'd much rather iterate and learn than assume we know all the answers.
  • Being a team player. You dont(always) work in isolation and are excited by the thought of using your team whilst involving product, experience design, engineering, and more in the process.

Responsibilities

  • Ensure platform reliability and availability across production and pre-production environments through proactive monitoring, alerting, and automation.
  • First response for incidents, contribute to problem management and root cause analysis.
  • Supporting the development team's effort towards reliability, creating a solid reliability culture within the development lifecycle.
  • Develop troubleshooting documentation for production support resources.
  • Collaborate with Engineering teams to develop optimised and productive runbooks, operational documentation and automation of operational tasks.
  • Collaborate with development and cloud engineering teams to embed reliability and performance into the software delivery lifecycle.
  • Design, implement, and evolve observability solutions (metrics, logs, traces, dashboards) using tools such as Prometheus, Grafana, and ELK.
  • Participate in on-call rotations and continuously improve alert quality and response processes.
  • Champion a culture of reliability, performance, and continuous improvement across teams.

Skills

Kubernetes
Docker Swarm
ELK Stack
AWS
Grafana/Prometheus
Bash
Python
Go
Communication
Agile/Lean
Teamwork

Education

Bachelor's or MS in Engineering

Tools

Prometheus
Grafana

Job description

Omilia is seeking a Senior Site Reliability Engineer with Cloud platform experience to join our production reliability team. You will operate and maintain production clusters, develop observability solutions, and drive automation across monitoring, alerting, and runbooks.

You will collaborate with engineering and cloud teams to embed reliability and performance into the software delivery lifecycle and participate in on-call rotations, championing a culture of reliability and continuous

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: Cloud Reliability & Observability Leader
Senior SRE: Cloud Reliability & Observability Leader

Omilia • United Kingdom

On-site
GBP 90,000 - 130,000
Fixed compensation
Long-term employment
Professional growth
+1
Senior SRE: Cloud Reliability, Observability & Automation
Senior SRE: Cloud Reliability, Observability & Automation

Omilia Natural Language Solutions Ltd • United Kingdom

On-site
GBP 70,000 - 90,000
Fixed compensation
Long-term employment with vacation
Professional growth opportunities
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Omilia • Greater London

On-site
GBP 90,000 - 120,000
Fixed compensation
Long-term vacation
Professional growth
+3
Site Reliability Engineer - Live Ops & Cloud Resilience
Site Reliability Engineer - Live Ops & Cloud Resilience

World Wrestling Entertainment, Inc. • Greater London

Hybrid
GBP 70,000 - 110,000
Platform Reliability Engineer - Automation & Observability
Platform Reliability Engineer - Automation & Observability

Onyx-Conseil • Manchester

On-site
GBP 90,000 - 120,000
Senior SRE, Observability & Cloud Reliability
Senior SRE, Observability & Cloud Reliability

United States Digital Space LLC • Greater London

Hybrid
GBP 120,000 - 170,000
Hybrid work up to 3 days per week
Senior SRE — Build a Global Cloud Reliability Practice
Senior SRE — Build a Global Cloud Reliability Practice

Omnicell • Manchester

Hybrid
GBP 90,000 - 130,000
Senior Site Reliability Engineer - Cloud & Observability
Senior Site Reliability Engineer - Cloud & Observability

London Stock Exchange Group • Greater London

On-site
GBP 80,000 - 100,000
Healthcare
Retirement planning
Paid volunteering days
Senior Production Engineer, Reliability & Automation
Senior Production Engineer, Reliability & Automation

Clear Street • Greater London

On-site
GBP 90,000 - 130,000
Global Director, Cloud Reliability & SRE
Global Director, Cloud Reliability & SRE

Omnicell • Manchester

On-site
GBP 150,000 - 190,000