Site Reliability Engineer III

Talent Octopusventures

Dundrum

On-site

GBP 44,000 - 73,000

Full time

6 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Country-specific benefits

Job summary

LexisNexis Risk Solutions is seeking a Site Reliability Engineer to design, operate and improve secure, highly available Azure platforms for large-scale production environments. Lead reliability initiatives and partner with engineering, security, architecture and product teams to enhance resilience.

You will implement Terraform IaC, automate processes with Python/PowerShell/Go, and drive CI/CD with GitHub Actions and Azure DevOps, supporting on-call rotations and incident response.

Qualifications

  • Experience in Site Reliability Engineering, Cloud Engineering, DevOps or Platform Engineering.
  • Expert knowledge of Microsoft Azure and strong experience with Kubernetes, preferably Azure Kubernetes Service.
  • Advanced Terraform skills and experience designing reusable Infrastructure as Code components.
  • Strong Git and GitOps practices, with experience using CI/CD platforms such as GitHub Actions and Azure DevOps.
  • Strong Linux administration, scripting and programming skills using languages such as Python, PowerShell, Go or Bash.
  • Experience with monitoring and observability platforms, distributed tracing, application performance monitoring and proactive alerting.
  • Skills in incident management, root cause analysis, capacity planning, performance optimisation, disaster recovery testing and production operations.
  • Technical leadership, strategic thinking, problem solving, effective communication, collaboration, customer focus and an automation-first approach. Azure, Kubernetes or Terraform certifications are welcomed.

Responsibilities

  • Design, implement and maintain highly available Azure infrastructure, including Azure Kubernetes Service, Virtual Machines, Functions, App Services, Storage, Networking, Key Vault and Azure Monitor.
  • Define and maintain Service Level Indicators, Service Level Objectives and Error Budgets, driving continuous improvements in platform reliability and availability.
  • Lead root cause analysis and post-incident reviews, developing resilience patterns such as auto-scaling, self-healing, disaster recovery, multi-region failover and high-availability architectures.
  • Design and maintain Infrastructure as Code using Terraform, build reusable platform components and automate manual operational processes using Python, PowerShell, Go or Bash.
  • Operate and optimise AKS clusters, establish container security standards, implement GitOps practices using tools such as ArgoCD and manage upgrades, capacity and workloads.
  • Develop monitoring, logging, alerting and distributed-tracing strategies, maintaining dashboards through Grafana and Azure Monitor to identify service degradation proactively.
  • Build and support reliable, repeatable and auditable CI/CD pipelines using GitHub Actions and Azure DevOps, including blue/green and canary deployment strategies.
  • Provide technical leadership, mentor SRE I and SRE II engineers, promote SRE practices across teams and participate in on-call rotations and major incident management.

Skills

Azure
Kubernetes
Terraform
GitHub Actions
Azure DevOps
Python
PowerShell
Go
Bash

Tools

Azure Kubernetes Service

Job description

About the Business

LexisNexis Risk Solutions is the essential partner in the assessment of risk. Within our Insurance vertical, we provide customers with solutions and decision tools that combine public and industry specific content with advanced technology and analytics to assist them in evaluating and predicting risk and enhancing operational efficiency. Our insurance risk solutions help drive better data-driven decisions across the insurance policy lifecycle, all while reducing risk. You can learn more about LexisNexis Risk at https://risk.lexisnexis.com/insurance.

About the Role

As a Site Reliability Engineer, you will design, operate and continuously improve secure, highly available Azure platforms that support large-scale production environments. You will provide technical leadership for reliability initiatives, develop reusable automation and partner with engineering, security, architecture and product teams to improve resilience and operational performance.

Responsibilities
  • Design, implement and maintain highly available Azure infrastructure, including Azure Kubernetes Service, Virtual Machines, Functions, App Services, Storage, Networking, Key Vault and Azure Monitor.
  • Define and maintain Service Level Indicators, Service Level Objectives and Error Budgets, driving continuous improvements in platform reliability and availability.
  • Lead root cause analysis and post-incident reviews, developing resilience patterns such as auto-scaling, self-healing, disaster recovery, multi-region failover and high-availability architectures.
  • Design and maintain Infrastructure as Code using Terraform, build reusable platform components and automate manual operational processes using Python, PowerShell, Go or Bash.
  • Operate and optimise AKS clusters, establish container security standards, implement GitOps practices using tools such as ArgoCD and manage upgrades, capacity and workloads.
  • Develop monitoring, logging, alerting and distributed-tracing strategies, maintaining dashboards through Grafana and Azure Monitor to identify service degradation proactively.
  • Build and support reliable, repeatable and auditable CI/CD pipelines using GitHub Actions and Azure DevOps, including blue/green and canary deployment strategies.
  • Provide technical leadership, mentor SRE I and SRE II engineers, promote SRE practices across teams and participate in on-call rotations and major incident management.
Requirements
  • Experience in Site Reliability Engineering, Cloud Engineering, DevOps or Platform Engineering, including supporting large-scale production environments.
  • Expert knowledge of Microsoft Azure and strong experience with Kubernetes, preferably Azure Kubernetes Service.
  • Advanced Terraform skills and experience designing reusable Infrastructure as Code components.
  • Strong Git and GitOps practices, with experience using CI/CD platforms such as GitHub Actions and Azure DevOps.
  • Strong Linux administration, scripting and programming skills using languages such as Python, PowerShell, Go or Bash.
  • Experience with monitoring and observability platforms, distributed tracing, application performance monitoring and proactive alerting.
  • Skills in incident management, root cause analysis, capacity planning, performance optimisation, disaster recovery testing and production operations.
  • Technical leadership, strategic thinking, problem solving, effective communication, collaboration, customer focus and an automation-first approach. Azure, Kubernetes or Terraform certifications are welcomed.
Benefits

We know your well-being and happiness are key to a long and successful career. We are delighted to offer country specific benefits.

Primary Location Base Pay Range: Ireland - Dublin (Rockfield Central) €51,400 - €85,700.

We are an equal opportunity employer: qualified applicants are considered for and treated during employment without regard to race, color, creed, religion, sex, national origin, citizenship status, disability status, protected veteran status, age, marital status, sexual orientation, gender identity, genetic information, or any other characteristic protected by law.

USA Job Seekers: EEO Know Your Rights.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer III
Site Reliability Engineer III

LexisNexis Risk Solutions • Dublin

On-site
EUR 51,000 - 86,000
Site Reliability Engineer III
Site Reliability Engineer III

LexisNexis Risk Solutions (Europe) Limited Company • Dublin

On-site
EUR 51,000 - 86,000
Country-specific benefits
Site Reliability Engineer III
Site Reliability Engineer III

LexisNexis Special Services Inc. • Dublin

On-site
EUR 51,000 - 86,000
Site Reliability Engineer III
Site Reliability Engineer III

RELX INC • Dublin

On-site
EUR 51,000 - 86,000
Site Reliability Engineer
Site Reliability Engineer

Talent Octopusventures • Ireland

Remote
EUR 49,000 - 82,000
Site Reliability Engineer III
Site Reliability Engineer III

Relx Plc • Dublin

On-site
EUR 51,000 - 86,000
Senior Azure SRE & Reliability Engineer
Senior Azure SRE & Reliability Engineer

LexisNexis Special Services Inc. • Dublin

On-site
EUR 51,000 - 86,000
Azure Site Reliability Engineer III - Reliability Leader
Azure Site Reliability Engineer III - Reliability Leader

LexisNexis Risk Solutions • Dublin

On-site
EUR 51,000 - 86,000
Senior Site Reliability Engineer (Azure/Kubernetes)
Senior Site Reliability Engineer (Azure/Kubernetes)

Talent Octopusventures • Dundrum

On-site
GBP 44,000 - 73,000
Country-specific benefits
Senior Site Reliability Engineer - Azure/Kubernetes
Senior Site Reliability Engineer - Azure/Kubernetes

Relx Plc • Dublin

On-site
EUR 51,000 - 86,000