Senior SRE Generalist: Proactive Reliability & Automation

Mastercard

Lusk

On-site

EUR 90,000 - 120,000

Full time

8 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Mastercard is seeking a Lead, Site Reliability Engineer – Generalist to proactively ensure system stability, performance, and resilience across infrastructure. The role emphasizes prevention, observability, and data-driven improvements, partnering with application, platform, and infrastructure teams to reduce outages and incidents.

The Senior SRE Generalist will drive reliability as a core capability, translating knowledge into architectural guidance and safer defaults, while mentoring peers and

Qualifications

  • Strong ability to reason about systems end to end, connecting application behavior to infrastructure performance and failure modes.
  • Expertise in observability, monitoring, and troubleshooting tools, with a focus on signal quality and actionable insight.
  • Proficiency in scripting and automation to operationalize reliability improvements and accelerate learning.
  • Broad infrastructure knowledge (networking, Linux, databases, containers, storage), with depth in at least one domain.
  • Strong data analysis and storytelling skills, enabling proactive identification of risks and clear communication of technical insights.
  • Working knowledge of machine learning concepts and their application to predictive and proactive operational problem solving.
  • Curiosity, ownership, and a mindset oriented toward preventing tomorrow’s incidents, not just fixing today’s.

Responsibilities

  • Anticipate reliability risks by analyzing application behavior, system signals, and historical incidents to identify failure patterns and systemic weaknesses before they result in outages.
  • Translate deep application knowledge into reliability requirements, architectural guidance, and infrastructure improvements that prevent incidents rather than simply respond to them.
  • Continuously assess system health, resiliency gaps, and operational debt, driving improvements that increase service robustness over time.
  • Participate in and lead troubleshooting efforts during high severity and cross domain incidents, applying structured, data driven investigation techniques.
  • Use incidents as learning opportunities—performing root cause analysis that focuses on why systems allowed failure, not just what broke.
  • Ensure incident outcomes result in concrete, measurable improvements such as better instrumentation, safer defaults, automation, or architectural changes.
  • Proactively design and evolve observability strategies by onboarding new data sources and improving signal quality across logs, metrics, traces, and events.
  • Build dashboards, alerts, and monitors that surface early indicators of degradation, not just failure states.
  • Apply analytical techniques to detect emerging trends, weak signals, and anomalous behavior before customers are impacted.
  • Communicate insights through clear data storytelling that enables engineering teams and leaders to act decisively and early.
  • Lead automation efforts that reduce manual intervention, shorten feedback loops, and eliminate repetitive operational work.
  • Convert operational learnings into reusable tools, standards, documentation, and patterns that raise the reliability baseline across teams.
  • Actively reduce operational toil and risk by improving system defaults, guardrails, and self healing capabilities.
  • Partner across application, infrastructure, and platform teams to drive shared ownership of reliability outcomes and proactive operational thinking.
  • Influence design and delivery decisions by representing the reliability perspective early in the development lifecycle.
  • Mentor engineers by modeling proactive troubleshooting, systems thinking, and data driven decision making.
  • Strong ability to reason about systems end to end, connecting application behavior to infrastructure performance and failure modes.
  • Expertise in observability, monitoring, and troubleshooting tools, with a focus on signal quality and actionable insight.
  • Proficiency in scripting and automation to operationalize reliability improvements and accelerate learning.
  • Broad infrastructure knowledge (networking, Linux, databases, containers, storage), with depth in at least one domain.
  • Strong data analysis and storytelling skills, enabling proactive identification of risks and clear communication of technical insights.
  • Working knowledge of machine learning concepts and their application to predictive and proactive operational problem solving.
  • Curiosity, ownership, and a mindset oriented toward preventing tomorrow’s incidents, not just fixing today’s.

Skills

Systems thinking
Observability
Automation scripting
Infrastructure knowledge
Data storytelling
ML concepts
Ownership mindset

Job description

Mastercard is seeking a Lead, Site Reliability Engineer – Generalist to proactively ensure system stability, performance, and resilience across infrastructure. The role emphasizes prevention, observability, and data-driven improvements, partnering with application, platform, and infrastructure teams to reduce outages and incidents.

The Senior SRE Generalist will drive reliability as a core capability, translating knowledge into architectural guidance and safer defaults, while mentoring peers and

Get your free, confidential resume review.

or drag and drop your file here.