Senior Site Reliability Engineer

Omilia Natural Language Solutions Ua Ltd

United Kingdom

On-site

GBP 85,000 - 120,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Fixed compensation
Long-term employment with the working
Development in professional growth (ca
Apple gear

Job summary

Omilia Natural Language Solutions Ua Ltd seeks a Senior Site Reliability Engineer with cloud experience to operate and maintain production clusters and observability tooling. You will collaborate to automate tasks, design runbooks, and improve monitoring across environments.

You will contribute to reliability, performance, and a culture of continuous improvement, working with Kubernetes, AWS, and the Grafana/Prometheus/ELK stack to deliver scalable services.

Qualifications

  • Bachelor's Degree or MS in Engineering or equivalent.
  • Experience operating at least one container orchestration cluster (Kubernetes, Docker Swarm).
  • Experience developing or maintaining software for production services at scale.
  • Experience with ELK.
  • Experience with AWS.
  • Experience with Grafana/Prometheus stack.
  • Strong scripting skills (Bash, Python or Go).
  • Excellent communication skills.
  • Thinking out of the box and anticipating challenges; proactive mindset.
  • Versatility and ability to work with agile/lean methods.
  • Being a team player across product, experience design and engineering.

Responsibilities

  • Ensure platform reliability and availability through proactive monitoring, alerting, and automation.
  • First response for incidents and contribute to problem management and root cause analysis.
  • Support development teams toward reliability and a strong reliability culture.
  • Develop troubleshooting documentation for production support resources.
  • Collaborate to develop runbooks, operational docs and automation.
  • Embed reliability and performance into the software delivery lifecycle.
  • Design and evolve observability solutions using Prometheus, Grafana and ELK.
  • Participate in on-call rotations and improve alert quality and response processes.
  • Champion a culture of reliability, performance, and continuous improvement.

Skills

Kubernetes
Docker Swarm
ELK
AWS
Grafana
Prometheus
Scripting
Bash
Python
Go
Communication
Problem solving
Team collaboration

Education

Bachelor's degree in Engineering
MS in Engineering

Tools

Prometheus
Grafana
ELK
AWS

Job description

We are looking for a Senior Site Reliability Engineer with Cloud platform experience. This individual will be part of a team responsible for operating and maintaining production clusters and developing our observability solutions; they will collaborate with team members to develop automation strategies, monitoring & alerting, and ensuring overall platform reliability. Your goal will be to become an integral part of the team, making every challenge of the platform – your own challenge, and solving them accordingly.

Responsibilities
  • Ensure platform reliability and availability across production and pre-production environments through proactive monitoring, alerting, and automation.
  • First response for incidents, contribute to problem management and root cause analysis.
  • Supporting the development team's effort towards reliability, creating a solid reliability culture within the development lifecycle.
  • Develop troubleshooting documentation for production support resources.
  • Collaborate with Engineering teams to develop optimised and productive runbooks, operational documentation and automation of operational tasks.
  • Collaborate with development and cloud engineering teams to embed reliability and performance into the software delivery lifecycle.
  • Design, implement, and evolve observability solutions (metrics, logs, traces, dashboards) using tools such as Prometheus, Grafana, and ELK.
  • Participate in on-call rotations and continuously improve alert quality and response processes.
  • Champion a culture of reliability, performance, and continuous improvement across teams.
Requirements
  • Bachelor's Degree or MS in Engineering or equivalent.
  • Experience in operating at least one container orchestration cluster (Kubernetes, Docker Swarm).
  • Experience developing or maintaining software for production services at scale.
  • Experience with ELK.
  • Experience with AWS.
  • Experience with Grafana/Prometheus stack.
  • Strong scripting skills (Bash, Python or Go).
  • Excellent communication skills.
  • Thinking out of the box and anticipating challenges. It is imperative we are not simply reactive; we must expect challenges and question technologies, procedures and thinking already in place. You will be expected to constantly review and challenge at all levels.
  • Versatility. We work with agile/lean methods. We'd much rather iterate and learn than assume we know all the answers.
  • Being a team player. You don't (always) work in isolation and are excited by the thought of using your team whilst involving product, experience design, engineering, and more in the process.
Will be considered as a plus:
  • Telephony knowledge (SIP, VoIP);
  • Experience in Linux Administration (RedHat, CentOS, AL);
  • Working knowledge in Configuration Management tools (Terraform, Ansible);
  • Experience with TCP/IP and general networking concepts;
  • RDBMS knowledge (MySQL, Postgres);
  • NoSQL knowledge (Redis).
Benefits
  • Fixed compensation;
  • Long-term employment with the working days vacation;
  • Development in professional growth (courses, training, etc);
  • Being part of successful cutting-edge technology products that are making a global impact in the service industry;
  • Proficient and fun-to-work-with colleagues;
  • Apple gear.

Omilia is proud to be an equal opportunity employer and is dedicated to fostering a diverse and inclusive workplace. We believe that embracing diversity in all its forms enriches our workplace and drives our collective success. We are committed to creating an environment where everyone feels welcomed, valued, and empowered to contribute their unique perspectives without regard to factors such as race, color, religion, gender, gender identity or expression, sexual orientation, national origin, heredity, disability, age, or veteran status, all eligible candidates will be given consideration for employment.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Omilia • Greater London

On-site
GBP 90,000 - 120,000
Fixed compensation
Long-term vacation
Professional growth
+3
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Symphony • Belfast City District

On-site
GBP 60,000 - 70,000
Regional specific competitive benefits
Build your own Benefits (BYOB) perk
Local events, team building, and devop
Senior AWS Site Reliability Engineer
Senior AWS Site Reliability Engineer

Spectrum IT Recruitment • City Of London

Hybrid
GBP 65,000 - 120,000
Life Insurance 4x Annual Salary
Private Medical Insurance
Bonus Scheme
+3
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Spectrum IT Recruitment • Southampton

Hybrid
GBP 55,000 - 90,000
Life Insurance
Private Medical Insurance
Employee Assistance Programme
+3
Senior Site Reliability Engineer - Cloud Observability & Automation
Senior Site Reliability Engineer - Cloud Observability & Automation

Omilia • Greater London

On-site
GBP 90,000 - 120,000
Fixed compensation
Long-term vacation
Professional growth
+3
Senior SRE: Cloud Reliability & Observability Lead
Senior SRE: Cloud Reliability & Observability Lead

Omilia Natural Language Solutions Ua Ltd • United Kingdom

On-site
GBP 85,000 - 120,000
Fixed compensation
Long-term employment with the working
Development in professional growth (ca
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Understanding Recruitment • United Kingdom

On-site
GBP 90,000 - 120,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Imanage • Belfast City District

On-site
GBP 85,000 - 120,000
Private medical insurance
Pension contributions matching
Annual performance bonus
+2
Principal Site Reliability Engineer, Infrastructure Observability
Principal Site Reliability Engineer, Infrastructure Observability

United States Digital Space LLC • Greater London

On-site
GBP 120,000 - 170,000
Hybrid work up to 3 days per week
Site Reliability Engineer
Site Reliability Engineer

Insight International (UK) Ltd • Bournemouth

On-site
GBP 55,000 - 75,000