Site Reliability Engineer (SRE), Cloud Operations

RBC

Toronto

On-site

CAD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Bonuses
Stock options where applicable
Flexible work options

Job summary

RBC is seeking a Platform Reliability Engineer to join the Platform Engineering & AI Operations team in Toronto. You will shape how RBC operates its private and public cloud platforms, building self-healing automation and reliable operational practices across Kubernetes/OpenShift, Kafka environments, and automation workflows.

You will automate tasks, participate in design reviews, and drive CI/CD and IaC practices while focusing on security, reliability, and regulatory requirements with on-call

Qualifications

  • 5+ years of hands-on experience in Site Reliability Engineering, DevOps, or infrastructure operations.
  • Strong working knowledge of Kubernetes/OpenShift administration and troubleshooting in enterprise environments.
  • Hands-on experience with Ansible and Terraform for Infrastructure as Code and automation.
  • Proficiency in Python scripting (core to infrastructure automation and platform development).
  • Hands-on experience with monitoring and observability stacks (Prometheus, Grafana, ELK, or equivalent).
  • Experience with incident management processes, on-call rotations, and post-incident review practices.

Responsibilities

  • Support scalable, secure architectures across private and public cloud platforms (Kubernetes/OpenShift, ECE, Confluent Kafka).
  • Automate infrastructure workflows and reduce toil with automation pipelines across data centers, branches, and NOC operations.
  • Extend self-healing automation to reduce manual intervention.
  • Lead design reviews for platform features and changes with security and regulatory alignment.
  • Collaborate with platform teams to implement data standards and pipelines (ServiceNow, Prometheus).
  • Drive CI/CD and Infrastructure as Code practices using Ansible and Terraform.

Skills

Kubernetes administration
Python scripting
Monitoring & observability
Incident management
Security & compliance fundamentals
AI/ML in operations

Tools

Ansible
Terraform
Prometheus
Grafana
ELK stack
ServiceNow

Job description

What is the opportunity?

Join the Platform Engineering & AI Operations team within OTK0, where you'll sit at the intersection of Site Reliability Engineering and intelligent infrastructure operations. This role offers the chance to shape how the bank operates, monitors, and self-heals its private and public cloud platforms — from OpenShift clusters and Kafka environments to self-healing automation systems. You'll work on real problems at enterprise scale: reducing toil for NOC, Data Center, and Branch teams, building automation that eliminates manual work, and establishing reliable operational practices. If you want to move beyond traditional ops into the future of intelligent, autonomous infrastructure operations — this is the role.

What will you do?
  • Support highly scalable, secure, and highly available architectures across private and public cloud platforms (Kubernetes/OpenShift, ECE, Confluent Kafka).
  • Write code and scripts to automate infrastructure workflows and eliminate toil, including automation pipelines that reduce manual intervention across Data Center, Branch, and NOC operations.
  • Extend self-healing automation capabilities built on Ansible, automating routine operational tasks (e.g., CPU remediation) to reduce manual intervention.
  • Participate in and lead design reviews for new platform features, infrastructure changes, and operational integration points, ensuring alignment with security, reliability, and regulatory requirements.
  • Collaborate with platform teams to provide technical feedback, contribute code changes to shared repositories, and establish data standards and pipelines (e.g., ServiceNow, Prometheus) that support operational excellence.
  • Drive automation, CI/CD, and Infrastructure as Code practices across the team, leveraging Ansible and Terraform for deployment validation and self-healing remediation workflows.
  • Minimize risk of reliability failures related to durability, availability, performance, and correctness, leveraging proactive alerting and anomaly detection.
  • Participate in on-call rotation for platform support, incident management, and troubleshooting, triaging incidents via Grafana, Prometheus, Dynatrace, and PagerDuty.
What do you need to succeed?
Must-have
  • 5+ years of hands-on experience in Site Reliability Engineering, DevOps, or infrastructure operations.
  • Strong working knowledge of Kubernetes/OpenShift administration and troubleshooting in enterprise environments.
  • Hands-on experience with Ansible and Terraform for Infrastructure as Code and automation.
  • Proficiency in Python scripting (core to infrastructure automation and platform development).
  • Hands-on experience with monitoring and observability stacks (Prometheus, Grafana, ELK, or equivalent).
  • Experience with incident management processes, on-call rotations, and post-incident review practices.
  • Familiarity with capacity planning, threshold-based alerting, and performance trend analysis.
  • Understanding of security and compliance fundamentals, including vulnerability assessment and remediation tracking.
  • Experience with AI/ML concepts applied to operations (anomaly detection, intelligent alerting, predictive capacity planning).
Nice-to-have
  • Experience with AI/ML concepts applied to operations (anomaly detection, intelligent alerting, predictive capacity planning).
  • Hands-on experience with public cloud platforms (AWS, Azure, GCP) in hybrid or multi-cloud environments.
  • Experience with GPU/compute infrastructure for ML inference workloads.
What’s in it for you?
  • A comprehensive Total Rewards Program including bonuses and flexible benefits, competitive compensation, commissions, and stock where applicable
  • Leaders who support your development through coaching and managing opportunities
  • Ability to make a difference and lasting impact
  • Work in a dynamic, collaborative, progressive, and high-performing team
  • Flexible work/life balance options
  • Opportunities to do challenging work
  • Opportunities to take on progressively greater accountabilities
  • Access to a variety of job opportunities across business

Additional Job Details

Address: RBC CENTRE, 155 WELLINGTON ST W

City: Toronto

Country: Canada

Work hours/week: 37.5

Employment Type: Full time

Platform: TECHNOLOGY AND OPERATIONS

Job Type: Regular

Pay Type: Salaried

Posted Date: 2026-07-31

Application Deadline: 2026-08-28—Note: Applications will be accepted until 11:59 PM on the day prior to the application deadline date above

Our Employment Opportunities

At RBC, we are guided by living shared values of Client First, Integrity, Collaboration, Respect and Excellence and winning together as One RBC. We believe an inclusive workplace that has diverse perspectives is core to our continued growth as one of the largest and most successful banks in the world. Maintaining a workplace where our employees feel supported to perform at their best, effectively collaborate, drive innovation, and grow professionally helps to bring our Purpose to life and create value for our clients and communities. RBC strives to deliver this through policies and programs intended to foster a workplace based on respect, belonging and opportunity for all.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer (SRE), Cloud Operations
Site Reliability Engineer (SRE), Cloud Operations

United States Digital Space LLC • Toronto

On-site
CAD 110,000 - 140,000
Bonuses
Flexible benefits
Stock options
+1
Senior Director, Head, SRE and Production Operations
Senior Director, Head, SRE and Production Operations

Socket.dev • Toronto

On-site
CAD 180,000 - 280,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

RBC • Toronto

On-site
CAD 110,000 - 160,000
Senior Engineer, Application Maintenance & Transformation
Senior Engineer, Application Maintenance & Transformation

RBC • Toronto

On-site
CAD 110,000 - 140,000
Lead Platform Engineer
Lead Platform Engineer

Socket.dev • Toronto

On-site
CAD 140,000 - 190,000
Bonuses
Flexible benefits
Stock options
+2
Senior Platform Engineer
Senior Platform Engineer

RBC • Toronto

On-site
CAD 110,000 - 150,000
Bonuses & flexible benefits
Stock options
Career development
Senior Analyst, Operational Innovation
Senior Analyst, Operational Innovation

Socket.dev • Toronto

On-site
CAD 95,000 - 135,000
Total rewards program
Leadership development opportunities
Dynamic, collaborative team
+2
AI Engineer - AI Platform Engineering
AI Engineer - AI Platform Engineering

Socket.dev • Vancouver

On-site
CAD 120,000 - 180,000
Total rewards program
Bonuses
Stock options
+2
Senior Platform Engineer
Senior Platform Engineer

Socket.dev • Toronto

On-site
CAD 85,000 - 125,000
Bonus program
Flexible benefits
Stock options
+1
Director, Data Engineer
Director, Data Engineer

ODAIA • Toronto

On-site
CAD 150,000 - 230,000
Bonuses
Stock options
Flexible benefits