Sr Incident & Reliability Manager

Paymentus

Richmond Hill

On-site

CAD 140,000 - 210,000

Full time

9 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Paymentus in Canada seeks a highly technical Senior Major Incident & Reliability Manager to lead live incident command, automate operational workflows, and drive blameless post-mortems for a cloud-native, multi-tenant fintech platform serving 2,000+ enterprise clients.

You will direct Sev1/2 response, enforce fixes to remove SPOFs, and partner with product and engineering to monitor SLOs and strengthen overall reliability.

Qualifications

  • 5+ years in Site Reliability Engineering, DevOps, or technical incident command in cloud-native SaaS.
  • Deep knowledge of microservices, Kubernetes, Kafka/RabbitMQ, and databases.
  • Proficiency with modern observability, logging, and APM platforms.
  • Ability to drive high-severity incident response with clarity and authority.

Responsibilities

  • Active Incident Command: direct Sev1/2 incident response and drive rapid MTTR.
  • Closed-Loop Problem Management: own blameless post-mortems and backlog fixes.
  • Reliability Governance: track SLAs/SLOs and enforce reliability thresholds.
  • Operational Automation: design event-driven automations to reduce toil.

Skills

Incident command
SRE/DevOps leadership
Kubernetes
Distributed systems
Observability

Tools

Elastic Stack
Grafana
CloudWatch
Datadog
Opsgenie

Job description

We are seeking a highly technical Senior Major Incident & Reliability Manager to lead active incident command, automated operational workflows, and proactive problem management for our cloud-native, multi-tenant fintech platform servicing 2,000+ enterprise clients. This is not an administrative reporting role. You will actively drive live triage in Zoom war rooms, challenge engineering teams on root causes, and enforce fixes that eliminate single points of failure (SPOFs).

Essential Functions/ Responsibilities
  • Active Incident Command: Direct high-severity (Sev 1/2) incident response. Probe engineers on microservices dependencies, database locks, queuing delays, and cloud infra bottlenecks to drive rapid MTTR.
  • Closed-Loop Problem Management: Own blameless post-mortems. Translate technical root causes into non-negotiable engineering sprint backlog items to prevent reoccurrence.
  • Reliability Governance: Track SLAs and service level objectives (SLOs). Partner with product and engineering teams to enforce reliability thresholds.
  • Operational Automation: Collaborate with workflow specialists to design event-driven automations (Opsgenie, Slack ChatOps, Statuspage triggers) that eliminate manual operational toil.
Supervisory Responsibility

This role will have direct reports.

Education and Experience
  • 5+ years experience in Site Reliability Engineering (SRE), DevOps, Systems Engineering, or Technical Incident Command within a cloud-native SaaS environment (AWS/GCP/Azure).
  • Deep fluency in distributed systems architectures: microservices, container orchestration (Kubernetes), message brokers (Kafka/RabbitMQ), and databases.
  • Proficiency in modern observability, logging, and APM platforms (Elastic Stack/APM, Grafana, CloudWatch, Datadog).
  • Demonstrated ability to maintain absolute authority, clarity, and analytical focus during major multi-tenant outages.

This job operates in a professional office environment. This role routinely uses standard office equipment such as laptop computers, photocopiers and smartphones.

Physical Demands

This role requires extended periods of sitting or standing at a computer workstation.

Position Type/Expected Hours of Work

This is a full-time position with standard working hours from Monday through Friday during normal business hours.

Because this role oversees critical platform reliability, it requires on-call availability as necessary to respond to overnight and off-hours incidents.

Travel
Other Duties

Please note this job description is not designed to cover or contain a comprehensive listing of activities, duties or responsibilities that are required of the employee for this job. Duties, responsibilities and activities may change at any time with or without notice.

EEO Statement

Paymentus is an equal opportunity employer. We enthusiastically accept our responsibility to make employment decisions without regard to race, religious creed, color, age, sex, sexual orientation, national origin, ancestry, citizenship status, religion, marital status, disability, military service or veteran status, genetic information, medical condition including medical characteristics, or any other classification protected by applicable federal, state, and local laws and ordinances. Our management is dedicated to ensuring the fulfillment of this policy with respect to hiring, placement, promotion, transfer, demotion, layoff, termination, recruitment advertising, pay, and other forms of compensation, training, and general treatment during employment.

Reasonable Accommodation

Paymentus recognizes and supports its obligation to endeavor to accommodate job applicants and employees with known physical or mental disabilities who are able to perform the essential functions of the position, with or without reasonable accommodation. Paymentus will endeavor to provide reasonable accommodations to otherwise qualified job applicants and employees with known physical or mental disabilities, unless doing so would impose an undue hardship on the Company or pose a direct threat of substantial harm to the employee or others.

An applicant or employee who believes he or she needs a reasonable accommodation of a disability should discuss the need for possible accommodation with the Human Resources Department, or his or her direct supervisor.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Manager, Technical Project Management
Sr. Manager, Technical Project Management

Paymentus • Richmond Hill

On-site
CAD 140,000 - 190,000
Sr Mgr, Technical Program Management
Sr Mgr, Technical Program Management

Paymentus Holdings Inc. • Richmond Hill

On-site
CAD 153,000 - 210,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

iManage • Toronto

Hybrid
CAD 90,000 - 120,000
Market-competitive salary
Annual performance-based bonus
Comprehensive Health, Vision, Dental, and Life insurance
+4
Cloud Security Engineer
Cloud Security Engineer

Paymentus • Richmond Hill

On-site
CAD 110,000 - 150,000
Production Support Engineer
Production Support Engineer

Paymentus Holdings Inc. • Richmond Hill

Hybrid
CAD 125,000 - 153,000
Senior Application Security Engineer
Senior Application Security Engineer

Paymentus • Richmond Hill

On-site
CAD 130,000 - 170,000
Senior Manager, Site Reliability Engineering (SRE)
Senior Manager, Site Reliability Engineering (SRE)

Dawninfotek • Toronto

On-site
CAD 150,000 - 210,000
Manager, Site Reliability Engineering (SRE)
Manager, Site Reliability Engineering (SRE)

Quantum Technology Recruiting Inc. (QTR) • Toronto

On-site
CAD 155,000 - 165,000
Site Reliability Engineer
Site Reliability Engineer

Gemini Solutions Pvt Ltd • Toronto

On-site
CAD 120,000 - 170,000
Director, Technical Operations - Infrastructure
Director, Technical Operations - Infrastructure

Accommodations Plus International • Markham

On-site
CAD 160,000 - 180,000