Senior SRE: Cloud Reliability & Observability Lead

Omilia Natural Language Solutions Ua Ltd

United Kingdom

On-site

GBP 85,000 - 120,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Fixed compensation
Long-term employment with the working
Development in professional growth (ca
Apple gear

Job summary

Omilia Natural Language Solutions Ua Ltd seeks a Senior Site Reliability Engineer with cloud experience to operate and maintain production clusters and observability tooling. You will collaborate to automate tasks, design runbooks, and improve monitoring across environments.

You will contribute to reliability, performance, and a culture of continuous improvement, working with Kubernetes, AWS, and the Grafana/Prometheus/ELK stack to deliver scalable services.

Qualifications

  • Bachelor's Degree or MS in Engineering or equivalent.
  • Experience operating at least one container orchestration cluster (Kubernetes, Docker Swarm).
  • Experience developing or maintaining software for production services at scale.
  • Experience with ELK.
  • Experience with AWS.
  • Experience with Grafana/Prometheus stack.
  • Strong scripting skills (Bash, Python or Go).
  • Excellent communication skills.
  • Thinking out of the box and anticipating challenges; proactive mindset.
  • Versatility and ability to work with agile/lean methods.
  • Being a team player across product, experience design and engineering.

Responsibilities

  • Ensure platform reliability and availability through proactive monitoring, alerting, and automation.
  • First response for incidents and contribute to problem management and root cause analysis.
  • Support development teams toward reliability and a strong reliability culture.
  • Develop troubleshooting documentation for production support resources.
  • Collaborate to develop runbooks, operational docs and automation.
  • Embed reliability and performance into the software delivery lifecycle.
  • Design and evolve observability solutions using Prometheus, Grafana and ELK.
  • Participate in on-call rotations and improve alert quality and response processes.
  • Champion a culture of reliability, performance, and continuous improvement.

Skills

Kubernetes
Docker Swarm
ELK
AWS
Grafana
Prometheus
Scripting
Bash
Python
Go
Communication
Problem solving
Team collaboration

Education

Bachelor's degree in Engineering
MS in Engineering

Tools

Prometheus
Grafana
ELK
AWS

Job description

Omilia Natural Language Solutions Ua Ltd seeks a Senior Site Reliability Engineer with cloud experience to operate and maintain production clusters and observability tooling. You will collaborate to automate tasks, design runbooks, and improve monitoring across environments.

You will contribute to reliability, performance, and a culture of continuous improvement, working with Kubernetes, AWS, and the Grafana/Prometheus/ELK stack to deliver scalable services.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE: Cloud Reliability & Observability Leader
Senior SRE: Cloud Reliability & Observability Leader

Omilia • United Kingdom

On-site
GBP 90,000 - 130,000
Fixed compensation
Long-term employment
Professional growth
+1
Senior SRE: Cloud Reliability, Observability & Automation
Senior SRE: Cloud Reliability, Observability & Automation

Omilia Natural Language Solutions Ltd • United Kingdom

On-site
GBP 70,000 - 90,000
Fixed compensation
Long-term employment with vacation
Professional growth opportunities
+1
Senior Site Reliability Engineer - Cloud Observability & Automation
Senior Site Reliability Engineer - Cloud Observability & Automation

Omilia • Greater London

On-site
GBP 90,000 - 120,000
Fixed compensation
Long-term vacation
Professional growth
+3
Senior SRE, Observability & Cloud Reliability
Senior SRE, Observability & Cloud Reliability

United States Digital Space LLC • Greater London

Hybrid
GBP 120,000 - 170,000
Hybrid work up to 3 days per week
Senior SRE — Build a Global Cloud Reliability Practice
Senior SRE — Build a Global Cloud Reliability Practice

Omnicell • Manchester

Hybrid
GBP 90,000 - 130,000
Senior SRE: Cloud Reliability & Observability
Senior SRE: Cloud Reliability & Observability

Renesas Electronics Corp. • Cambridge

On-site
GBP 90,000 - 120,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Omilia • Greater London

On-site
GBP 90,000 - 120,000
Fixed compensation
Long-term vacation
Professional growth
+3
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Omilia Natural Language Solutions Ua Ltd • United Kingdom

On-site
GBP 85,000 - 120,000
Fixed compensation
Long-term employment with the working
Development in professional growth (ca
+1
Senior SRE: Azure Cloud Platform & Observability Leader
Senior SRE: Azure Cloud Platform & Observability Leader

Spectrum IT Recruitment • Southampton

Hybrid
GBP 80,000 - 110,000
Senior SRE: Cloud Platform Reliability & Observability
Senior SRE: Cloud Platform Reliability & Observability

Renesas Electronics Corporation • Cambridge

Hybrid
GBP 90,000 - 120,000