Site Reliability Engineer

DOCOsoft

Dublin

On-site

EUR 90,000 - 130,000

Full time

48 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

25 days Annual Leave
Private pension
Bonus scheme
Private health
Life assurance

Job summary

DOCOsoft is seeking a Senior Azure Site Reliability Engineer to ensure reliability, availability and performance of our SaaS platform on Microsoft Azure. You will collaborate with Engineering, DevOps and Infrastructure to design scalable, resilient systems and drive automation, observability and security across the stack.

Responsibilities include building fault-tolerant architectures, enhancing CI/CD, leading incident responses, and embedding security and compliance practices as the platform

Qualifications

  • Proven experience as a Site Reliability Engineer or similar reliability-focused role in SaaS/cloud environments.
  • Hands-on experience operating production workloads on Microsoft Azure (compute, networking, storage, monitoring).
  • Infrastructure as Code expertise using Bicep, ARM, Terraform, or similar tools.
  • Strong automation and scripting capability (PowerShell essential).
  • Experience with containerised environments (Docker) and Kubernetes orchestration.
  • Experience with observability tooling (Azure Monitor, Grafana, Prometheus, Datadog, OpenTelemetry).
  • Knowledge of reliability engineering practices, SLA/SLO concepts, and incident response.
  • Understanding of security and compliance standards (ISO27001, SOC 2, GDPR).
  • Strong problem-solving and cross-functional collaboration abilities.
  • Azure certifications (e.g., Azure Administrator Associate, Azure Solutions Architect Expert) desirable.

Responsibilities

  • Design, implement, and operate highly available, scalable Azure systems.
  • Define and track reliability metrics and performance reporting.
  • Develop and maintain IaC using Bicep, ARM, Terraform or similar tools.
  • Build automation for provisioning, deployment, scaling, and workflows.
  • Enhance CI/CD pipelines with DevOps collaboration for safer deployments.
  • Implement monitoring, logging, tracing, and alerting for real-time visibility.
  • Define alerting strategies to reduce noise and improve response.
  • Lead incident response including troubleshooting, RCA, and post-incident reviews.
  • Strengthen incident and problem management to improve SLA adherence.
  • Embed security and compliance best practices across infrastructure.
  • Drive continuous service improvement for performance and reliability.
  • Collaborate with Development and QA to improve resilience.

Skills

Azure
SRE / Reliability
IaC (Terraform/Bicep/ARM)
PowerShell
Docker
Kubernetes
Observability tools
Security & compliance
Azure certifications

Tools

Terraform
Bicep
ARM
Grafana
Prometheus
Datadog
OpenTelemetry

Job description

As a Senior Azure Site Reliability Engineer, you will play a critical role in ensuring the reliability, availability, and performance of our Vew SaaS platform hosted on Microsoft Azure.

You will work closely with Engineering, DevOps, and Infrastructure teams to design and operate scalable, resilient systems. This role is focused on automation, observability, incident response, and continuous improvement, ensuring our platform remains stable, secure, and operationally mature as it scales.

Responsibilities
  • Design, implement, and operate highly available, scalable, and fault-tolerant systems on Microsoft Azure.
  • Define, track, and improve reliability metrics, including service health indicators and operational performance reporting.
  • Develop and maintain Infrastructure as Code using Bicep, ARM, Terraform, or similar tooling to ensure consistent and reproducible environments.
  • Build automation for provisioning, deployment, scaling, and operational workflows, reducing manual intervention and operational toil.
  • Enhance CI/CD pipelines in collaboration with DevOps to improve deployment safety, reliability, and efficiency.
  • Implement and maintain monitoring, logging, tracing, and alerting solutions to ensure real-time visibility and rapid issue detection.
  • Define meaningful alerting strategies that reduce noise and improve response effectiveness.
  • Lead incident response activities, including structured troubleshooting, stakeholder communication, root cause analysis, and post-incident reviews.
  • Strengthen incident and problem management processes to improve SLA adherence and customer impact mitigation.
  • Implement systemic improvements to prevent repeat incidents rather than applying short-term workarounds.
  • Embed security and compliance best practices across infrastructure, including access control, encryption, and policy enforcement.
  • Drive continuous service improvement initiatives to enhance performance, reliability, efficiency, and operational maturity.
  • Collaborate closely with Development and QA teams to improve application resilience and supportability.
Key Requirements
  • Proven experience as a Site Reliability Engineer or in a similar reliability-focused role within a SaaS or cloud-native environment.
  • Strong hands-on experience operating production workloads on Microsoft Azure across compute, networking, storage, and monitoring services.
  • Infrastructure as Code expertise using Bicep, ARM, Terraform, or similar tools.
  • Strong automation and scripting capability (PowerShell essential; additional scripting languages advantageous).
  • Experience working with containerised environments (Docker) and orchestration concepts such as Kubernetes.
  • Practical experience with observability tooling such as Azure Monitor, Grafana, Prometheus, Datadog, or OpenTelemetry.
  • Strong understanding of structured incident response, root cause analysis, SLA/SLO concepts, and reliability engineering practices.
  • Knowledge of security best practices and compliance standards such as ISO27001, SOC 2, and GDPR.
  • Strong problem-solving capability with the ability to troubleshoot complex, distributed systems.
  • Effective communication skills and ability to collaborate across engineering, operations, and business stakeholders.
  • Azure certifications (e.g., Azure Administrator Associate, Azure Solutions Architect Expert) are desirable.
Who We Are

DOCOsoft is a leading software and services provider to Lloyds of London and the broader London insurance market. Since our foundation, we have grown to become one of the leading insurance software specialists in the London Insurance Market. We are a growing team of over 105 colleagues based in Dublin, London, Tokyo, Portugal, Spain, India and Poland.

Here’s What We Have To Offer

DOCOsoft aspires to be a market leader in the technology sector and we are always looking for new ways to improve how we deliver value. We hire people who bring hard work, enthusiasm and their own ideas.

We Offer
  • 25 days Annual Leave
  • Private pension
  • Bonus scheme
  • Private health
  • Life assurance
Equal Opportunity Employer

DOCOsoft is committed to building an inclusive and diverse team that represents a variety of backgrounds, experiences and perspectives. We welcome applications from all suitably qualified candidates and do not discriminate on any legally protected grounds. If you require reasonable accommodation during any stage of the recruitment process, please let us know.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

DevOps Engineer (Azure)
DevOps Engineer (Azure)

DOCOsoft • Dublin

On-site
EUR 90,000 - 130,000
25 days Annual Leave
Private pension
Bonus scheme
+2
Senior Software Developer (CMS)
Senior Software Developer (CMS)

Docosoft Ltd. • Leinster

On-site
EUR 90,000 - 120,000
25 days Annual Leave
Private pension
Bonus scheme
+2
Site Reliability Engineer
Site Reliability Engineer

Jobtailor • Dublin

On-site
EUR 90,000 - 130,000
Senior Software Developer (CMS)
Senior Software Developer (CMS)

DOCOsoft • Dublin

Hybrid
EUR 70,000 - 90,000
25 days Annual Leave
Private pension
Bonus scheme
+2
Senior Azure SRE: Scale, Reliability & Automation
Senior Azure SRE: Scale, Reliability & Automation

DOCOsoft • Dublin

On-site
EUR 90,000 - 130,000
25 days Annual Leave
Private pension
Bonus scheme
+2
Technical Lead (C#/ .NET)
Technical Lead (C#/ .NET)

Docosoft Ltd. • Leinster

On-site
EUR 90,000 - 120,000
25 days Annual Leave
Private pension
Bonus scheme
+2
Data Migration Engineer
Data Migration Engineer

Docosoft Ltd. • Leinster

On-site
EUR 70,000 - 110,000
25 days Annual Leave
Private pension
Bonus scheme
+2
Data Migration Engineer
Data Migration Engineer

DOCOsoft • Dublin

On-site
EUR 70,000 - 100,000
25 days Annual Leave
Private pension
Bonus scheme
+2
Staff Site Reliability Engineer
Staff Site Reliability Engineer

United States Digital Space LLC • Dublin

On-site
EUR 92,000 - 127,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Acuity Brands, Inc. • Cork

Hybrid
EUR 90,000 - 130,000