Senior Site Reliability Engineer (R-19383)

Dun & Bradstreet

Dublin

On-site

EUR 110,000 - 150,000

Full time

41 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Dun & Bradstreet is seeking a Senior Site Reliability Engineer to ensure reliability, availability, performance, and operability of production systems across our platforms in Dublin. You will apply software engineering practices to operations with a focus on automation, observability, and incident response.

The role requires hands-on experience with GCP, Kubernetes (GKE), SRE incident management, and building scalable, resilient cloud architectures.

Qualifications

  • Bachelor's degree in Computer Science, Information Technology or related field.
  • Experience with cloud-native concepts and technologies, with strong preference for GCP and Kubernetes (GKE).
  • Proven experience in Site Reliability Engineering and incident management.
  • Experience with monitoring and observability tools (metrics, logs, traces, synthetics).
  • Exposure to resilience testing or cost optimisation initiatives.
  • Analytical and problem-solving skills for diagnosing production issues quickly.
  • Software development or automation using Python, shell scripts, or similar.

Responsibilities

  • Own and improve reliability, availability, and performance of production services in Google Cloud (GCP).
  • Participate in incident management: detection, triage, mitigation, escalation, recovery.
  • Use and improve incident workflows and tooling (e.g., ServiceNow).
  • Design, implement, and operate observability solutions including monitoring, logging, tracing, synthetics, and dashboards (Splunk Observability, OpenTelemetry).
  • Reduce operational toil through automation and engineering-led solutions.
  • Support on-call rotations across multiple time zones for 24/7 coverage.
  • Define, monitor, and report SLIs, SLOs, and error budgets for critical services.
  • Drive best-in-class availability through SRE principles and proactive reliability engineering.

Skills

GCP
Kubernetes (GKE)
SRE mgmt
Observability tools
Python
Shell scripting
Multi-region HA
Cloud infrastructure
OpenTelemetry
ServiceNow

Education

Bachelor's degree in Computer Science, Information Technology or related field

Tools

Splunk Observability
OpenTelemetry
ServiceNow

Job description

The Senior Site Reliability Engineer (SRE) is responsible for ensuring the reliability, availability, performance, and operability of production systems across our platforms, by applying software engineering practices to operations, with a focus on automation, observability, and incident response.

The Senior Site Reliability Engineer (SRE) is responsible for ensuring the reliability, availability, performance, and operability of production systems across our platforms, by applying software engineering practices to operations, with a focus on automation, observability, and incident response.

Responsibilities:
  • Own and improve the reliability, availability, and performance of production services in Google Cloud (GCP).
  • Participate in incident management, including detection, triage, mitigation, escalation, and recovery.
  • Use and improve incident workflows and tooling (e.g., ServiceNow) to ensure clear ownership and timely communication.
  • Design, implement, and operate observability solutions including monitoring, logging, tracing, synthetics, and dashboards (e.g., Splunk Observability, OpenTelemetry).
  • Reduce operational toil through automation and engineering-led solutions, proactively introducing and driving SRE best practices.
  • Support on-call rotations across multiple time zones, contributing to a sustainable 24/7 support model.
  • Define, monitor, and report SLIs, SLOs, and error budgets for critical services.
  • Drive and be accountable for best-in-class service availability through SRE principles, automation, and proactive reliability engineering.
Essential skills and/or Certifications:
  • Bachelor’s degree in Computer Science, Information Technology or related field
  • Strong experience with cloud-native concepts and technologies, with a strong preference for Google Cloud Platform (GCP) and Kubernetes (GKE).
  • Proven experience with Site Reliability Engineering and production incident management, ideally using platforms such as ServiceNow.
  • Experience with monitoring and observability tools, including metrics, logs, traces, and synthetics (e.g., Splunk Observability, OpenTelemetry).
  • Exposure to reliability testing, resilience engineering, or cost optimisation initiatives.
  • Excellent analytical and problem-solving skills, with the ability to diagnose complex production issues quickly.
  • Software development or automation experience using Python, shell scripts, or similar languages.
  • Hands-on experience operating production cloud infrastructure at scale.
  • Experience managing multi-region, high-availability production systems with a focus on scalability, resilience, and minimising service disruption during failures.
  • Proficiency in Microsoft Office Suites Skills
  • Show an ownership mindset in everything you do; be a problem solver, be curious and be inspired to take action, be proactive, seek ways to collaborate and connect with people and teams in support of driving success.
  • Continuous growth mindset, keep learning through social experiences and relationships with stakeholders, experts, colleagues and mentors as well as widen and broaden your competencies through structural courses and programs.

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please visit https://bit.ly/3LMn4CQ.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer (R-19383)
Senior Site Reliability Engineer (R-19383)

Dnb • Dublin

On-site
EUR 70,000 - 90,000
Software Engineering Manager, Site Reliability Engineering, Turnup Org
Software Engineering Manager, Site Reliability Engineering, Turnup Org

Google • Dublin

On-site
EUR 150,000 - 153,000
Senior SRE: Cloud Reliability, Automation & Observability
Senior SRE: Cloud Reliability, Automation & Observability

Dnb • Dublin

On-site
EUR 70,000 - 90,000
Staff Software Engineer, Site Reliability Engineering, AlphaNet SRE
Staff Software Engineer, Site Reliability Engineering, AlphaNet SRE

Google • Dublin

On-site
EUR 120,000 - 180,000
Equity
Bonus target
Benefits
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Harvey Nash • Dublin

On-site
EUR 90,000 - 130,000
Software Engineering Manager, Site Reliability Engineering, Turnup Org
Software Engineering Manager, Site Reliability Engineering, Turnup Org

Google Inc. • Leinster

On-site
EUR 150,000 - 153,000
Equity
Benefits
Bonus target
Software Engineering Manager, Site Reliability Engineering, Turnup Org
Software Engineering Manager, Site Reliability Engineering, Turnup Org

Google • Ireland

On-site
EUR 120,000 - 160,000
Equity
Bonus
Benefits
Engineering Manager, Site Reliability Engineering, Alphanet SRE
Engineering Manager, Site Reliability Engineering, Alphanet SRE

Google • Ireland

On-site
EUR 140,000 - 200,000
Equity
Benefits
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobtailor • Ireland

On-site
EUR 70,000 - 120,000
Senior SRE: Cloud Reliability, Automation & Observability
Senior SRE: Cloud Reliability, Automation & Observability

Dun & Bradstreet • Dublin

On-site
EUR 110,000 - 150,000