Lead, Site Reliability Engineering

Mastercard

Dundrum

In loco

EUR 90.000 - 135.000

Tempo pieno

8 giorni fa
Generatore di candidature

Trasforma questa posizione in un colloquio — un curriculum e una lettera di presentazione creati in base a ciò questo datore di lavoro sta cercando.

Supera i filtri ATS

Descrizione del lavoro

Mastercard seeks a senior Site Reliability Engineer – Generalist to lead reliability across platforms, translating incidents into preventative actions and improving observability. You will partner with application, platform, and infrastructure teams to continually reduce outages and operational toil.

You will mentor engineers, drive automation, and champion data-driven decision making to raise the reliability baseline while delivering scalable, robust services for Mastercard's global platforms.

Competenze

  • Strong ability to reason about systems end to end and connect application behavior to infrastructure performance.
  • Expertise in observability, monitoring and troubleshooting tools.
  • Proficiency in scripting and automation to operationalize reliability improvements.
  • Broad infrastructure knowledge including networking, Linux, databases, containers and storage.
  • Strong data analysis and storytelling skills to communicate insights.
  • Working knowledge of machine learning concepts for predictive operational problem solving.

Mansioni

  • Anticipate reliability risks by analyzing application behavior and historical incidents to identify systemic weaknesses.
  • Translate reliability knowledge into architectural guidance and infrastructure improvements.
  • Continuously assess health and debt to increase service robustness over time.
  • Lead incident response efforts and perform root cause analysis to prevent recurrence.
  • Design and evolve observability strategies, dashboards and alerts for early indicators.
  • Drive automation to reduce manual toil and implement self-healing capabilities.
  • Mentor engineers and influence design decisions with a reliability perspective.

Conoscenze

Observability
Automation scripting
Systems thinking
Networking & Linux
Data storytelling

Descrizione del lavoro

Our Purpose

Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential.

Title and Summary

Lead, Site Reliability Engineering

Site Reliability Engineer (SRE) – Generalist

Role Summary

The Site Reliability Engineer (SRE) – Generalist is a senior level engineer and cross stack reliability expert who proactively ensures system stability, performance, and operational resilience by deeply understanding application behavior and how it manifests across infrastructure.

This role emphasizes anticipation over reaction. While the SRE Generalist participates in incident response, their primary value is in converting operational signals, incidents, and patterns into preventative actions—improving observability, reducing risk, and eliminating classes of failure before they impact customers. They partner closely with application, platform, and infrastructure teams to continuously reduce mean time to detect (MTTD), mean time to resolve (MTTR), and overall incident frequency through data driven insight, automation, and engineering rigor.

Key Responsibilities
Proactive Reliability Engineering
  • Anticipate reliability risks by analyzing application behavior, system signals, and historical incidents to identify failure patterns and systemic weaknesses before they result in outages.
  • Translate deep application knowledge into reliability requirements, architectural guidance, and infrastructure improvements that prevent incidents rather than simply respond to them.
  • Continuously assess system health, resiliency gaps, and operational debt, driving improvements that increase service robustness over time.
Incident Response as an Input to Prevention
  • Participate in and lead troubleshooting efforts during high severity and cross domain incidents, applying structured, data driven investigation techniques.
  • Use incidents as learning opportunities—performing root cause analysis that focuses on why systems allowed failure, not just what broke.
  • Ensure incident outcomes result in concrete, measurable improvements such as better instrumentation, safer defaults, automation, or architectural changes.
Observability, Monitoring & Signal Quality
  • Proactively design and evolve observability strategies by onboarding new data sources and improving signal quality across logs, metrics, traces, and events.
  • Build dashboards, alerts, and monitors that surface early indicators of degradation, not just failure states.
  • Apply analytical techniques to detect emerging trends, weak signals, and anomalous behavior before customers are impacted.
  • Communicate insights through clear data storytelling that enables engineering teams and leaders to act decisively and early.
Automation & Continuous Improvement
  • Lead automation efforts that reduce manual intervention, shorten feedback loops, and eliminate repetitive operational work.
  • Convert operational learnings into reusable tools, standards, documentation, and patterns that raise the reliability baseline across teams.
  • Actively reduce operational toil and risk by improving system defaults, guardrails, and self healing capabilities.
Collaboration, Influence & Mentorship
  • Partner across application, infrastructure, and platform teams to drive shared ownership of reliability outcomes and proactive operational thinking.
  • Influence design and delivery decisions by representing the reliability perspective early in the development lifecycle.
  • Mentor engineers by modeling proactive troubleshooting, systems thinking, and data driven decision making.
Knowledge, Skills & Abilities
  • Strong ability to reason about systems end to end, connecting application behavior to infrastructure performance and failure modes.
  • Expertise in observability, monitoring, and troubleshooting tools, with a focus on signal quality and actionable insight.
  • Proficiency in scripting and automation to operationalize reliability improvements and accelerate learning.
  • Broad infrastructure knowledge (networking, Linux, databases, containers, storage), with depth in at least one domain.
  • Strong data analysis and storytelling skills, enabling proactive identification of risks and clear communication of technical insights.
  • Working knowledge of machine learning concepts and their application to predictive and proactive operational problem solving.
  • Curiosity, ownership, and a mindset oriented toward preventing tomorrow’s incidents, not just fixing today’s.
What Defines Success in This Role
  • Sees incidents as signals, not endpoints.
  • Uses observability and data to shift reliability work left and upstream.
  • Reduces incident frequency and impact over time—not just MTTR.
  • Acts as a connective force across teams, turning complexity into clarity and prevention.
Corporate Security Responsibility

All activities involving access to Mastercard assets, information, and networks comes with an inherent risk to the organization and, therefore, it is expected that every person working for, or on behalf of, Mastercard is responsible for information security and must:

  • Abide by Mastercard’s security policies and practices;
  • Ensure the confidentiality and integrity of the information being accessed;
  • Report any suspected information security violation or breach, and
  • Complete all periodic mandatory security trainings in accordance with Mastercard’s guidelines.
Ottieni la revisione del curriculum gratis e riservata.

o trascina qui il file.

Similar jobs

Offerte di lavoro simili che vale la pena confrontare

Lead, Site Reliability Engineering
Lead, Site Reliability Engineering

Mastercard • Bray

In loco
EUR 100.000 - 130.000
Lead, Site Reliability Engineering
Lead, Site Reliability Engineering

Mastercard • Celbridge

In loco
EUR 100.000 - 140.000
Lead, Site Reliability Engineering
Lead, Site Reliability Engineering

Mastercard • Howth

In loco
EUR 100.000 - 150.000
Lead, Site Reliability Engineering
Lead, Site Reliability Engineering

Mastercard • Blackrock

In loco
EUR 120.000 - 180.000
Lead, Site Reliability Engineering
Lead, Site Reliability Engineering

Mastercard • Lusk

In loco
EUR 90.000 - 120.000
Lead, Site Reliability Engineering
Lead, Site Reliability Engineering

Mastercard • Swords

In loco
EUR 90.000 - 140.000
Lead, Site Reliability Engineering
Lead, Site Reliability Engineering

Mastercard • Dublin

In loco
EUR 110.000 - 150.000
Site Reliability Engineer II
Site Reliability Engineer II

Mastercard • Malahide

In loco
EUR 90.000 - 130.000
Site Reliability Engineer II
Site Reliability Engineer II

Mastercard • Donabate

In loco
EUR 85.000 - 120.000
Site Reliability Engineer II
Site Reliability Engineer II

Mastercard • Rathcoole

In loco
EUR 75.000 - 110.000