Get more replies from employers
Send a job-specific resume in minutes.
RBC is seeking a founding member for its Site Reliability Engineering team to build the bank’s first-ever Agentic AI platform for software reliability and resiliency. You will design end-to-end AI-driven solutions, develop intelligent automation frameworks, and implement production-grade infrastructure at enterprise scale.
The role emphasizes cutting-edge AI at scale, incident prevention, and reducing toil while collaborating with cross-functional teams.
Job Description
Join RBC’s Site Reliability Engineering team as a founding member building the bank’s first-ever Agentic AI platform for Software reliability and resiliency. You’ll pioneer intelligent automation systems that autonomously prevent incidents, accelerate response times, and transform how we maintain resilience across enterprise systems. This is a rare opportunity to shape the future of AI-driven reliability at scale. Your innovations will protect millions of daily customer transactions and sign-ins. With a clear technical leadership trajectory, you’ll architect cutting-edge solutions at the intersection of AI and infrastructure, setting the standard for autonomous operations in financial services.
Design and implement end-to-end Agentic AI solutions that autonomously detect anomalies, identify root causes, and resolve incidents with minimal human intervention
Develop intelligent automation frameworks using LangChain and LangGraph to create context-aware agents that learn from incident patterns and continuously improve response strategies
Build ML-powered monitoring and alerting systems that distinguish signal from noise, dramatically reducing false positives and improving MTTD (Mean Time to Detect) and MTTI (Mean Time to Identify)
Architect scalable, production-grade solutions on OpenShift and Kubernetes that process real-time system metrics and telemetry data at enterprise scale
Implement infrastructure-as-code using Ansible and containerization (Docker) to ensure reproducibility, consistency, and rapid deployment across environments
Partner with incident management and operations teams to translate operational pain points into AI-driven automation opportunities that measurably reduce toil
Establish and track KPIs focused on reducing MTTR (Mean Time to Resolve), MTTD, and MTTI while improving system reliability
Lead technical design discussions and contribute to architectural decisions that shape RBC’s AI-powered reliability strategy
We thrive on the challenge to be our best, progressive thinking to keep growing, and working together to deliver trusted advice to help our clients thrive and communities prosper. We care about each other, reaching our potential, making a difference to our communities, and achieving success that is mutual.
#LI-POST
#TECHPJ
Docker Kubernetes Architecture, LangChain (FrameWork), LangGraph, Machine Learning (ML), Python (Programming Language), Red Hat Ansible, Red Hat OpenShift
RBC WATERPARK PLACE, 88 QUEENS QUAY W:TORONTO
Toronto
Canada
37.5
Full time
TECHNOLOGY AND OPERATIONS
Regular
Salaried
2026-04-27
2026-08-21
Note** : Applications will be accepted until 11:59 PM on the day prior to the application deadline date above
At RBC, we are guided by living shared values of Client First, Integrity, Collaboration, Respect and Excellence and winning together as One RBC. We believe an inclusive workplace that has diverse perspectives is core to our continued growth as one of the largest and most successful banks in the world. Maintaining a workplace where our employees feel supported to perform at their best, effectively collaborate, drive innovation, and grow professionally helps to bring our Purpose to life and create value for our clients and communities. RBC strives to deliver this through policies and programs intended to foster a workplace based on respect, belonging and opportunity for all.