Staff, Site Reliability Engineer(Global Security)

rbc

Toronto

On-site

CAD 140,000 - 190,000

Full time

3 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

RBC is seeking an experienced Staff Site Reliability Engineer focused on IAM reliability and scalable cloud-native architectures. You will drive SLOs, SLIs, and robust incident response while building self-service automation for IAM services.

You will own CI/CD pipelines, implement resilience strategies, and collaborate across security, infrastructure, and application teams to ensure high availability and performance across hybrid-cloud environments.

Qualifications

  • 5+ years in Site Reliability Engineering or similar with leadership across reliability metrics.
  • Strong software fundamentals and production-grade service design capabilities.
  • Experience designing and operating highly available, fault-tolerant systems in hybrid/multi-cloud.

Responsibilities

  • Set architecture direction, reliability standards, and lead hands-on design and operations.
  • Own IAM service reliability end-to-end, defining SLOs/SLIs and error budgets.
  • Design resilient IAM infrastructure across multi-region and hybrid-cloud environments.
  • Develop production-grade code and tooling for reliability, not just scripts.
  • Build self-service platforms and automation for provisioning IAM services.
  • Champion IaC and GitOps to enable safe, repeatable deployments.
  • Own CI/CD pipelines and release engineering for IAM services.
  • Lead incident response and postmortems for high-severity outages.
  • Improve observability with metrics, logs, and traces; validate DR plans.

Skills

SRE leadership
SLO/SLI design
Programming: Python/Go/Java
CI/CD pipelines
Kubernetes
Docker
Observability
Incident management
Chaos engineering
Security integration

Tools

Terraform
Ansible
Kubernetes
GitHub Actions
Puppet
Helm
Docker

Job description

What is the opportunity?

We are seeking an experienced and hands-on Staff Site Reliability Engineer (SRE) who is passionate about building scalable systems, automating infrastructure, and improving the reliability of production environments. You will work closely with system support, engineering and infrastructure teams to ensure the availability, performance, and efficiency of our IAM systems and services. This is a technical, execution-focused role with deep engineering work.

What will you do?
  • Serve as the senior-most technical voice for IAM reliability - setting architecture direction and reliability standards, and leading by example through hands-on design, coding, and operations
  • Own service reliability for IAM systems end-to-end by defining and maintaining SLOs, SLIs, and error budgets, and using them to drive prioritization and continuous improvement
  • Design and implement resilient, highly available IAM infrastructure and services - spanning authentication, authorization, identity lifecycle, and privileged access - across multi-region and hybrid-cloud architectures
  • Write and review production-grade code (services, APIs, automation frameworks, and internal tools), applying software engineering rigor - testing, code review, version control, modular design - to reliability work rather than treating it as throwaway scripting
  • Build self-service platforms, reusable modules, and golden-path automation that let application teams provision, deploy, and safely operate IAM services with less hand-holding from the production support team
  • Champion Infrastructure as Code and GitOps practices (Terraform, Ansible, Puppet, Kubernetes/Helm) to eliminate manual toil, enforce consistency, and enable safe, repeatable deployments at scale
  • Own and continuously improve CI/CD pipelines and release engineering for IAM services, embedding reliability, security, and rollback safety directly into the delivery pipeline
  • Lead incident response and on-call operations for high-severity IAM outages and performance degradations, driving root cause analysis, blameless postmortems, and long-term structural remediation
  • Build and evolve observability, monitoring, and alerting pipelines (metrics, logs, traces) using modern tooling to proactively detect, investigate, and resolve availability and security issues before they impact users
  • Develop and test failover strategies and recovery procedures, including chaos engineering exercises, backup validation, and disaster recovery simulations to validate IAM system readiness
  • Orchestrate workload automation, scheduling, and release pipelines across enterprise systems (e.g., Stonebranch, CI/CD platforms) to streamline delivery and reduce manual operational overhead
  • Partner with security, infrastructure, application, and compliance teams to embed IAM into enterprise-wide business continuity and resilience strategy, ensuring alignment with risk and regulatory mandates
  • Mentor and coach engineers on reliability, automation-first thinking, and sound software engineering practice; help raise the bar for how the broader team builds and operates IAM services
What do you need to succeed?

Must Have:

  • 5+ years of experience in Site Reliability Engineering, DevOps or Platform Engineering with demonstrated staff/senior-level technical leadership across SLOs/SLIs, error budgets, and driving continuous reliability improvement at scale
  • Solid software engineering fundamentals - proficient in at least one modern language (Python, Go, Java, or similar), with the ability to design and build production-quality services and tooling, not just automation scripts
  • Proven experience designing, implementing, and operating highly available, fault-tolerant, and scalable systems in production, including hybrid and multi-cloud environments
  • Strong DevOps foundation - experience owning CI/CD pipelines (e.g., Jenkins, GitLab CI, GitHub Actions) and release engineering practices that build reliability into the delivery process itself
  • Platform-engineering mindset - a track record of turning recurring operational work into self-service tooling, reusable modules, or golden-path automation that other engineering teams can adopt independently
  • Experience with containerization and orchestration (Docker, Kubernetes) in production environments
  • Deep knowledge of building and operating monitoring, alerting, and observability platforms (e.g., Prometheus, Grafana, Dynatrace, ELK, Splunk, SIEM) to enable proactive incident detection and response
  • Proven incident management skills - able to lead high-severity incident response, root cause analysis, and postmortem processes, and to drive long-term fixes
  • Proficient in disaster recovery, failover strategies, and resilience testing (e.g., chaos engineering, tabletop exercises) to validate system readiness
  • Solid understanding of cloud platforms (AWS, Azure) and hybrid environments, including experience supporting production workloads at scale
  • Excellent collaboration
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

LanceSoft, Inc. • Montreal (administrative region)

On-site
CAD 110,000 - 140,000
Senior Infrastructure SRE
Senior Infrastructure SRE

PointClickCare • Mississauga

On-site
CAD 110,000 - 150,000
Senior Site Reliability Engineer, SRE
Senior Site Reliability Engineer, SRE

Jobtailor • Toronto

On-site
CAD 120,000 - 180,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

twentysix • Vancouver

On-site
CAD 90,000 - 130,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

iManage • Toronto

On-site
CAD 90,000 - 120,000
Market-competitive salary
Annual performance-based bonus
Comprehensive Health, Vision, Dental, and Life insurance
+4
Senior IAM Reliability Engineer - Global SRE
Senior IAM Reliability Engineer - Global SRE

rbc • Toronto

On-site
CAD 140,000 - 190,000
Staff Developer - Incident Command
Staff Developer - Incident Command

IBM • Toronto

On-site
CAD 140,000 - 190,000
Staff Developer - Incident Command
Staff Developer - Incident Command

IBM • Markham

On-site
CAD 120,000 - 180,000
Staff Developer - Incident Command
Staff Developer - Incident Command

IBM • Calgary

On-site
CAD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

ALLTECH CONSULTING SVC INC • Quebec

On-site
CAD 90,000 - 130,000