Senior Site Reliability Engineer

Socket.dev

Toronto

On-site

CAD 120,000 - 180,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Total Rewards Program
Bonuses and stock where applicable
World-class training
Career development opportunities

Job summary

RBC Wealth Management Technology seeks a Senior Site Reliability Engineer to strengthen reliability across critical wealth management platforms. You will design and implement self-healing automation, scalable observability, and ML-based anomaly detection, collaborating with development and platform teams.

You will drive incident response, define SLIs/SLOs, and evolve runbooks into automation-first workflows while supporting deployments with robust reliability standards.

Qualifications

  • 5+ years in SRE/Production/DevOps or similar with strong ops depth.
  • Bachelor’s degree in CS, Engineering, or related field.
  • Experience with infrastructure automation and config management (Ansible).
  • Scripting and automation skills in Bash, Python, PowerShell, or equivalent.
  • Hands-on with observability tools like Elasticsearch, Dynatrace, Kubernetes, OpenShift, Kafka, PagerDuty, Moogsoft.

Responsibilities

  • Build and enhance the SRE product base with intelligent monitoring and automated remediation.
  • Implement modern observability across apps: metrics, logs, traces, dashboards, alerting.
  • Design ML-based anomaly detection to move to predictive operations.
  • Architect self-healing solutions with governance.
  • Standardize telemetry and instrumentation across platforms.
  • Lead incident management and participate in on-call rotations.
  • Automate workflows using Ansible, GitHub Actions, and scripting.

Skills

SRE/DevOps experience
Cloud-native architectures
Automation scripting
Elasticsearch
Dynatrace
Kubernetes
OpenShift
Kafka
PagerDuty
Moogsoft
Ansible
GitHub Actions
Python

Education

Bachelor’s degree in Computer Science or Engineering

Tools

Elasticsearch
Dynatrace
Kubernetes
OpenShift
Kafka
PagerDuty
Moogsoft
Ansible
GitHub Actions
Python

Job description

Job Description

RBC Wealth Management Technology is seeking a Senior Site Reliability Engineer to join its Wealth Management SRE Team. This team is responsible for ensuring the performance, availability, resilience, and operational excellence of critical applications and platforms that support the wealth management business. Working at the intersection of software engineering, cloud-native operations, observability, and automation, the team plays a central role in delivering reliable digital services for both internal users and clients.

As a Senior Site Reliability Engineer, you will bring an engineering-first mindset, strong operational judgment, and a passion for automation to improve system reliability at scale. You will work closely with development, infrastructure, platform, and support teams to build modern observability practices, improve incident response, strengthen reliability engineering standards, and drive the evolution toward intelligent, self-healing operations.

This role is ideal for a hands‑on engineer who is equally comfortable improving production resilience, building automation, defining service-level objectives, and shaping the future of AI-enhanced operations. You will help design and implement scalable SRE solutions across the technology estate using tools and platforms such as Elasticsearch, Ansible, GitHub Actions, Dynatrace, PagerDuty, Moogsoft, Kubernetes, OpenShift, Kafka, and emerging AIOps capabilities.

What will you do?
  • Build and enhance the SRE product base, including intelligent monitoring, alerting, reliability testing, anomaly detection, and automated remediation.
  • Implement modern observability practices across supported applications, including metrics, logs, traces, dashboards, and actionable alerting.
  • Design and pilot machine learning-based anomaly detection capabilities to improve signal quality and move from reactive to predictive operations.
  • Architect and implement self-healing solutions that automatically remediate recurring operational issues with appropriate controls and governance.
  • Design human-in-the-loop workflows that balance automation speed with accountability, risk management, and operational oversight.
  • Standardize telemetry and instrumentation across platforms to improve visibility, coverage, and correlation of operational signals.
  • Contribute to the centralization and evolution of observability and monitoring backends to enable deeper analytics and faster incident triage.
  • Partner with cross-functional teams to improve monitoring, logging, alerting, incident response, and production readiness practices.
  • Automate operational workflows and platform tasks using Ansible, GitHub Actions, and scripting languages such as Bash, Python, and PowerShell.
  • Develop and maintain custom tooling that improves operational efficiency, reliability, and scale.
  • Work closely with development teams to understand application changes, production risks, and release readiness, ensuring services meet reliability standards before and after deployment.
  • Define, track, and improve SLIs, SLOs, error budgets, and other critical service health indicators.
  • Evolve runbooks into automation-first remediation patterns and intelligent operational workflows.
  • Support production deployments by advocating for reliability, resilience, and performance improvements.
  • Lead or contribute to incident management, problem management, and root cause analysis, ensuring corrective actions are implemented and sustained.
  • Troubleshoot production issues across application, middleware, infrastructure, and platform layers.
  • Participate in an on-call rotation and provide senior operational support for business-critical systems.
  • Continuously identify opportunities to simplify, automate, and modernize operations using engineering and AI-driven approaches.
What do you need to succeed?
Must-have
  • 5+ years of experience in Site Reliability Engineering, Production Engineering, DevOps, Platform Engineering, or Systems Engineering roles with strong operational depth.
  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • Strong experience with infrastructure automation and configuration management, particularly Ansible.
  • Strong scripting and automation skills in Bash, Python, PowerShell, or similar languages.
  • Hands‑on experience with modern reliability and observability tooling such as Elasticsearch, Dynatrace, GitHub, Kubernetes, OpenShift, Kafka, PagerDuty, Moogsoft, or related platforms.
  • Strong understanding of production operations, incident management, root cause analysis, and reliability engineering practices.
  • Experience defining and operating SLIs, SLOs, alerting strategies, and service health metrics.
  • Knowledge of cloud-native and distributed systems concepts, including resiliency, scalability, fault isolation, and performance tuning.
  • Understanding of AIOps, AI/ML concepts, or intelligent automation as applied to observability and operations.
  • Ability to work across teams, influence engineering practices, and communicate clearly with technical and non-technical stakeholders.
Nice-to-have
  • Experience in financial services, wealth management, banking, insurance, or other highly regulated environments.
  • Experience with OpenTelemetry and telemetry standardization across distributed systems.
  • Hands‑on experience with Prometheus, Grafana, Splunk, Catchpoint, Azure Automation, or similar SRE and observability platforms.
  • Experience with CI/CD and developer platform tools such as Jenkins, Artifactory, and Vault.
  • Familiarity with containerization and cloud platform patterns, including Docker and Kubernetes-based deployments.
  • Experience building or operating anomaly detection, predictive alerting, or self-healing automation solutions.
  • Familiarity with AI governance, model validation, and operational controls in regulated environments.
  • Experience with reliability testing, chaos engineering, or resilience validation practices.
What’s in it for you?

We thrive on the challenge to be our best, progressive thinking to keep growing, and working together to deliver trusted advice to help our clients thrive and communities prosper. We care about each other, reaching our potential, making a difference to our communities, and achieving success that is mutual.

  • A comprehensive Total Rewards Program including bonuses and flexible benefits, competitive compensation, commissions, and stock where applicable
  • Leaders who support your development through coaching and managing opportunities
  • Ability to make a difference and lasting impact
  • Work in a dynamic, collaborative, progressive, and high-performing team
  • A world‑class training program in financial services
  • Opportunities to do challenging work
Job Skills

Agile Methodology, Group Problem Solving, IT Systems Integration, Organizational Leadership, Product Services, Software Development Life Cycle (SDLC), System Applications, System Integration Testing (SIT), Systems Software

Additional Job Details
Address

RBC CENTRE, 155 WELLINGTON ST W:TORONTO

City

Toronto

Country

Canada

Work hours/week

37.5

Employment Type

Full time

Platform

TECHNOLOGY AND OPERATIONS

Job Type

Regular

Pay Type

Salaried

Posted Date

2026-07-15

Application Deadline

2026-09-07

Note

Applications will be accepted until 11:59 PM on the day prior to the application deadline date above

Our Employment Opportunities

At RBC, we are guided by living shared values of Client First, Integrity, Collaboration, Respect and Excellence and winning together as One RBC. We believe an inclusive workplace that has diverse perspectives is core to our continued growth as one of the largest and most successful banks in the world. Maintaining a workplace where our employees feel supported to perform at their best, effectively collaborate, drive innovation, and grow professionally helps to bring our Purpose to life and create value for our clients and communities. RBC strives to deliver this through policies and programs intended to foster a workplace based on respect, belonging and opportunity for all.

Join our Talent Community

Stay in-the-know about great career opportunities at RBC. Sign up and get customized info on our latest jobs, career tips and Recruitment events that matter to you.
Expand your limits and create a new future together at RBC. Find out how we use our passion and drive to enhance the well‑being of our clients and communities at jobs.rbc.com.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE/AIOps Engineer
Senior SRE/AIOps Engineer

Socket.dev • Toronto

On-site
CAD 90,000 - 140,000
Site Reliability Engineer - SRE
Site Reliability Engineer - SRE

RBC • Toronto

On-site
CAD 110,000 - 140,000
Total rewards
Coaching & development
Impact
+3
Site Reliability Engineer - SRE
Site Reliability Engineer - SRE

Socket.dev • Toronto

On-site
CAD 90,000 - 130,000
Staff, Site Reliability Engineer(Global Security)
Staff, Site Reliability Engineer(Global Security)

RBC • Toronto

On-site
CAD 120,000 - 150,000
Staff, Site Reliability Engineer(Global Security)
Staff, Site Reliability Engineer(Global Security)

RBC • Vancouver

On-site
CAD 120,000 - 180,000
Bonuses & flexible benefits
Coaching & development
Impactful work
+3
Staff, Site Reliability Engineer(Global Security)
Staff, Site Reliability Engineer(Global Security)

RBC • Halifax

On-site
CAD 140,000 - 190,000
Application Support Engineer
Application Support Engineer

RBC • Toronto

On-site
CAD 67,000 - 110,000
Total rewards
Flexible benefits
Coaching and development
Senior Director, Head, SRE and Production Operations
Senior Director, Head, SRE and Production Operations

RBC • Toronto

On-site
CAD 180,000 - 240,000
Bonuses
Stock options
Flexible benefits
Staff Technical Lead – AI & Platform
Staff Technical Lead – AI & Platform

RBC • Toronto

On-site
CAD 140,000 - 190,000
Total rewards program
Hybrid/Flexible work options
Industry training & development
Principal Engineer, Cyber Technology Operations SRE (Global Security)
Principal Engineer, Cyber Technology Operations SRE (Global Security)

RBC • Toronto

On-site
CAD 140,000 - 210,000
Total rewards program
Bonuses and stock where applicable
Flexible benefits