Systems Reliability Engineer

Mphasis

Toronto

On-site

CAD 110,000 - 150,000

Full time

8 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Mphasis is seeking a Systems Reliability Engineer to join the Technical Operations team. You will own the reliability and resiliency of a large-scale enterprise platform, designing monitoring and alerting frameworks and leading incident response.

You will define SLOs/SLIs, drive RCA, build automation to reduce toil, and collaborate across engineering, product and risk teams to embed reliability best practices.

Qualifications

  • Experience in reliability engineering for large-scale platforms.
  • Proven ability to design and implement monitoring and alerting strategies.
  • Strong RCA and incident response experience.
  • Solid understanding of observability concepts and metrics.
  • Excellent written and verbal communication, with cross-team collaboration.

Responsibilities

  • Own the reliability, resiliency and availability of the platform.
  • Design and maintain monitoring and alerting with Splunk, Dynatrace, Grafana and Datadog.
  • Define SLOs, SLIs and error budgets to measure platform health.
  • Lead incident response and coordinate remediation with technical and product teams.
  • Drive RCA processes and document root causes to prevent recurrence.
  • Manage incident tickets and reporting (e.g., ServiceNow).
  • Build automation to reduce toil and improve MTTD/MTTR.
  • Collaborate with engineers, product and risk teams to embed reliability practices.

Skills

Reliability engineering
Incident response
Root cause analysis
Observability
Communication
Collaboration

Tools

Splunk
Dynatrace
Grafana
Datadog
ServiceNow

Job description

We are seeking a Systems Reliability Engineer to join the Technical Operations team. In this role, you will own the reliability and resiliency of a large-scale enterprise platform, ensuring that our services remain highly available, performant and secure. You will design and implement monitoring and alerting frameworks, lead incident response and drive the root cause analysis (RCA) process to continuously improve platform stability. This is an opportunity to be a Partner in Possibility — helping our clients deliver financial services experiences that are essential to everyday life.

What you will do:
  • Own the reliability, resiliency and availability of the Embedded Finance platform, proactively identifying and mitigating risks to service continuity.
  • Design, implement and maintain comprehensive monitoring and alerting frameworks leveraging Splunk, Dynatrace, Grafana and Datadog to provide end-to-end observability across the platform.
  • Define and track service level objectives (SLOs), service level indicators (SLIs) and error budgets to measure and improve platform health.
  • Lead and participate in incident response, serving as a technical driver during remediation calls and coordinating with impacted and impacting technical and product teams.
  • Own and advance the root cause analysis (RCA) process — investigating incidents, documenting the sequence of events and remediating actions, and clearly identifying underlying root causes to prevent recurrence.
  • Ensure timely creation and management of incident tickets (e.g., ServiceNow) and accurate incident tracking, aging and reporting.
  • Build automation and tooling to reduce toil, improve mean time to detection (MTTD) and mean time to resolution (MTTR), and increase operational efficiency.
  • Collaborate with engineering, product and risk stakeholders to embed reliability best practices into the platform lifecycle.
What you will need to have:
  • Hands-on experience with monitoring, observability and alerting tools, specifically Splunk, Dynatrace, Grafana and Datadog.
  • Proven experience operating and supporting a large-scale enterprise platform environment.
  • Demonstrated experience with incident response and leading or contributing to root cause analysis (RCA) processes.
  • Strong understanding of reliability engineering principles, including availability, resiliency, monitoring and alerting best practices.
  • Experience with ticketing and incident management workflows (e.g., ServiceNow).
  • Excellent communication skills, with the ability to drive remediation efforts and collaborate across technical, product and risk teams.
What would be great to have:
  • Experience in financial services, payments or embedded finance environments.
  • Proficiency with scripting or programming languages for automation (e.g., Python, Go, Bash).
  • Familiarity with cloud platforms, containerization and CI/CD pipelines.
  • Experience defining and managing SLOs, SLIs and error budgets
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Sr Incident & Reliability Manager
Sr Incident & Reliability Manager

Paymentus • Richmond Hill

On-site
CAD 140,000 - 210,000
Staff Developer - Incident Command
Staff Developer - Incident Command

IBM • Vancouver

On-site
CAD 150,000 - 190,000
Staff Developer - Incident Command
Staff Developer - Incident Command

IBM • Bedford

On-site
CAD 140,000 - 200,000
Staff Developer - Incident Command
Staff Developer - Incident Command

IBM • Ottawa

On-site
CAD 120,000 - 160,000
Staff Developer - Incident Command
Staff Developer - Incident Command

IBM • Edmonton

On-site
CAD 120,000 - 180,000
Staff Developer - Incident Command
Staff Developer - Incident Command

IBM • Toronto

On-site
CAD 140,000 - 190,000
Staff Developer - Incident Command
Staff Developer - Incident Command

IBM • Markham

On-site
CAD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Compunnel, Inc. • Montreal (administrative region)

On-site
CAD 90,000 - 130,000
Site Reliability Engineer
Site Reliability Engineer

ALLTECH CONSULTING SVC INC • Quebec

On-site
CAD 90,000 - 130,000
Production & Reliability Management Expert
Production & Reliability Management Expert

ALLTECH CONSULTING SVC INC • Quebec

On-site
CAD 70,000 - 90,000
Competitive compensation and benefits package
Health insurance coverage
Professional training and certifications
+3