Senior Site Reliability Engineer — Proactive Observability

Mastercard

Bray

On-site

EUR 100,000 - 130,000

Full time

8 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Mastercard is seeking a senior Site Reliability Engineer (SRE) – Generalist to own end-to-end reliability across platforms. You will anticipate risks, drive observability, and partner with cross-functional teams to prevent outages.

In this role you will lead automation, improve monitoring, mentor engineers, and translate incidents into concrete improvements that reduce MTTD and MTTR while maintaining service quality.

Qualifications

  • Strong ability to reason about systems end to end, connecting application behavior to infrastructure performance.
  • Expertise in observability, monitoring, and troubleshooting tools, with a focus on signal quality and actionable insight.
  • Proficiency in scripting and automation to operationalize reliability improvements.
  • Broad infrastructure knowledge (networking, Linux, databases, containers, storage).
  • Strong data analysis and storytelling skills to communicate technical insights clearly.

Responsibilities

  • Proactively design and evolve observability strategies by onboarding new data sources and improving signal quality across logs, metrics, traces, and events.
  • Build dashboards, alerts, and monitors that surface early indicators of degradation.
  • Lead automation efforts that reduce manual intervention and shorten feedback loops.
  • Mentor engineers by modeling proactive troubleshooting and data-driven decision making.
  • Collaborate across application, infrastructure, and platform teams to drive shared ownership of reliability outcomes.

Skills

Observability & Monitoring
Troubleshooting
Automation & Scripting
Data Analysis
Systems Thinking
Cross-team Collaboration

Tools

Linux
Kubernetes
Databases
Containers

Job description

Mastercard is seeking a senior Site Reliability Engineer (SRE) – Generalist to own end-to-end reliability across platforms. You will anticipate risks, drive observability, and partner with cross-functional teams to prevent outages.

In this role you will lead automation, improve monitoring, mentor engineers, and translate incidents into concrete improvements that reduce MTTD and MTTR while maintaining service quality.

Get your free, confidential resume review.

or drag and drop your file here.