Site Reliability Engineer

Jobtailor

Dublin

On-site

EUR 90,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor is seeking a Site Reliability Engineer to design, implement, and operate highly available systems on Microsoft Azure with a focus on reliability metrics and automation. You will build IaC using Bicep, ARM, and Terraform, enhance CI/CD pipelines, and implement robust monitoring and alerting.

Collaboration with Development and QA teams will be key to improving resilience across services. The role demands strong Azure experience, containerization skills with Docker/Kubernetes, and a

Qualifications

  • Proven experience as a Site Reliability Engineer or in a similar reliability-focused role within a SaaS or cloud-native environment.
  • Strong hands-on experience operating production workloads on Microsoft Azure across compute, networking, storage, and monitoring services.
  • Infrastructure as Code expertise using Bicep, ARM, Terraform, or similar tools.
  • Strong automation and scripting capability (PowerShell essential; additional scripting languages advantageous).
  • Experience working with containerised environments (Docker) and orchestration concepts such as Kubernetes.
  • Practical experience with observability tooling such as Azure Monitor, Grafana, Prometheus, Datadog, or OpenTelemetry.
  • Strong understanding of structured incident response, SLA/SLO concepts, and reliability engineering practices.
  • Knowledge of security best practices and compliance standards such as ISO27001, SOC 2, and GDPR.
  • Strong problem-solving capability with the ability to troubleshoot complex, distributed systems.
  • Effective communication skills and ability to collaborate across engineering, operations, and business stakeholders.
  • Azure certifications (e.g., Azure Administrator Associate, Azure Solutions Architect Expert) are desirable.

Responsibilities

  • Design, implement, and operate highly available, scalable, and fault‑tolerant systems on Microsoft Azure.
  • Define, track, and improve reliability metrics, including service health indicators and operational performance reporting.
  • Develop and maintain Infrastructure as Code using Bicep, ARM, Terraform, or similar tooling to ensure consistent and reproducible environments.
  • Build automation for provisioning, deployment, scaling, and operational workflows, reducing manual intervention and operational toil.
  • Enhance CI/CD pipelines in collaboration with DevOps to improve deployment safety, reliability, and efficiency.
  • Implement and maintain monitoring, logging, tracing, and alerting solutions to ensure real‑time visibility and rapid issue detection.
  • Define meaningful alerting strategies that reduce noise and improve response effectiveness.
  • Lead incident response activities, including structured troubleshooting, stakeholder communication, root cause analysis, and post‑incident reviews.
  • Strengthen incident and problem management processes to improve SLA adherence and customer impact mitigation.
  • Implement systemic improvements to prevent repeat incidents rather than applying short‑term workarounds.
  • Embed security and compliance best practices across infrastructure, including access control, encryption, and policy enforcement.
  • Drive continuous service improvement initiatives to enhance performance, reliability, efficiency, and operational maturity.
  • Collaborate closely with Development and QA teams to improve application resilience and supportability.

Skills

Site Reliability Engineering
Cloud infrastructure
PowerShell
Automation & Scripting
Communication & Collaboration

Tools

Docker
Kubernetes
Terraform
Bicep
ARM templates
Azure Monitor
Grafana
Prometheus
Datadog
OpenTelemetry

Job description

Responsibilities
  • Design, implement, and operate highly available, scalable, and fault‑tolerant systems on Microsoft Azure.
  • Define, track, and improve reliability metrics, including service health indicators and operational performance reporting.
  • Develop and maintain Infrastructure as Code using Bicep, ARM, Terraform, or similar tooling to ensure consistent and reproducible environments.
  • Build automation for provisioning, deployment, scaling, and operational workflows, reducing manual intervention and operational toil.
  • Enhance CI/CD pipelines in collaboration with DevOps to improve deployment safety, reliability, and efficiency.
  • Implement and maintain monitoring, logging, tracing, and alerting solutions to ensure real‑time visibility and rapid issue detection.
  • Define meaningful alerting strategies that reduce noise and improve response effectiveness.
  • Lead incident response activities, including structured troubleshooting, stakeholder communication, root cause analysis, and post‑incident reviews.
  • Strengthen incident and problem management processes to improve SLA adherence and customer impact mitigation.
  • Implement systemic improvements to prevent repeat incidents rather than applying short‑term workarounds.
  • Embed security and compliance best practices across infrastructure, including access control, encryption, and policy enforcement.
  • Drive continuous service improvement initiatives to enhance performance, reliability, efficiency, and operational maturity.
  • Collaborate closely with Development and QA teams to improve application resilience and supportability.
Requirements
  • Proven experience as a Site Reliability Engineer or in a similar reliability‑focused role within a SaaS or cloud‑native environment.
  • Strong hands‑on experience operating production workloads on Microsoft Azure across compute, networking, storage, and monitoring services.
  • Infrastructure as Code expertise using Bicep, ARM, Terraform, or similar tools.
  • Strong automation and scripting capability (PowerShell essential; additional scripting languages advantageous).
  • Experience working with containerised environments (Docker) and orchestration concepts such as Kubernetes.
  • Practical experience with observability tooling such as Azure Monitor, Grafana, Prometheus, Datadog, or OpenTelemetry.
  • Strong understanding of structured incident response, root cause analysis, SLA/SLO concepts, and reliability engineering practices.
  • Knowledge of security best practices and compliance standards such as ISO27001, SOC 2, and GDPR.
  • Strong problem‑solving capability with the ability to troubleshoot complex, distributed systems.
  • Effective communication skills and ability to collaborate across engineering, operations, and business stakeholders.
  • Azure certifications (e.g., Azure Administrator Associate, Azure Solutions Architect Expert) are desirable.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Site Reliability Engineer
Staff Site Reliability Engineer

United States Digital Space LLC • Dublin

Hybrid
EUR 90,000 - 130,000
Site Reliability Engineer III - Eng
Site Reliability Engineer III - Eng

UKG • Leinster

On-site
EUR 60,000 - 80,000
Platform Engineer, Azure SaaS
Platform Engineer, Azure SaaS

Jobtailor • Dublin

On-site
EUR 90,000 - 120,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobtailor • Dublin

On-site
EUR 100,000 - 140,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobtailor • Ireland

On-site
EUR 70,000 - 120,000
Staff Site Reliability Engineer (Site Experience)
Staff Site Reliability Engineer (Site Experience)

Reddit • Ireland

On-site
EUR 80,000 - 110,000
Comprehensive health benefits
Flexible vacation
Paid parental leave
+3
DevOps Engineer, Azure
DevOps Engineer, Azure

Jobtailor • Dublin

On-site
EUR 90,000 - 120,000
Azure SRE: Scalable Cloud Reliability & Automation
Azure SRE: Scalable Cloud Reliability & Automation

Jobtailor • Dublin

On-site
EUR 90,000 - 130,000
Site Reliability Engineer
Site Reliability Engineer

Back4good • Ireland

On-site
EUR 60,000 - 80,000
Strong and attractive package
Relocation assistance
Site Reliability Engineering Technical Lead
Site Reliability Engineering Technical Lead

AMCS Group • Dublin

On-site
EUR 110,000 - 150,000