Senior Principal Site Reliability Engineer

Questrade, Inc.

Canada

On-site

CAD 210,000 - 266,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health & wellbeing resources
Hybrid work environment
Competitive compensation

Job summary

Questrade Financial Group is seeking a Senior Principal Site Reliability Engineer to own the reliability of brokerage back-end systems across hybrid on‑prem and cloud environments.

You will drive SLOs/SLIs, observability, and incident response while contributing hands-on code across languages to raise the operational bar and reduce toil. This role emphasizes automation, AI-driven insights, and scalable architectures in a regulated financial services setting.

Qualifications

  • Proven ability to design reliable, scalable systems for hybrid on-prem and cloud environments.
  • Experience leading blameless post-incident reviews and RCA processes.
  • Strong collaboration with multiple teams to improve operational excellence.

Responsibilities

  • Own end-to-end reliability for brokerage back-end apps across on‑prem and cloud.
  • Define and track SLOs/SLIs, and manage error budgets with application teams.
  • Lead incident response and drive remediation to prevent recurrence.
  • Develop and mature observability with metrics, logs, and traces.

Skills

Observability
Incident response
SLOs/SLIs
Capacity planning
Automation
Hybrid cloud
CI/CD

Education

Bachelor's or Master's in CS/Engineering/IS

Tools

Terraform

Job description

Senior Principal Site Reliability Engineer

5700 Yonge St, North York, ON M2M 4K2, Canada

Job Description

Posted Wednesday, August 5, 2026 at 4:00 AM

Questrade Financial Group (QFG), through its companies - Questrade, Questbank, Questrade Wealth Management, Community Trust Company, Zolo, and Flexiti, provides securities and foreign currency investment, professionally managed investment portfolios, mortgages, real estate services, financial services and more. We use cutting-edge technology to help Canadians become much more financially successful and secure.

At QFG, we combine human-centric collaboration with AI-driven innovation to redefine financial services. The ideal candidate will be a catalyst for change, using AI to transform and deliver unparalleled customer experiences and shaping a future where AI empowers our teams to do their best work.

Join our diverse, inclusive, and hybrid workplace to unleash your creativity and nurture your curiosity without limits. If you share this sense of infinite possibility, come shape your future at QFG.

What’s in it for you as an employee of QFG?

Health & wellbeing resources and programs

Paid vacation, personal, and sick days for work-life balance

Competitive compensation and benefits packages

Work-life balance in a hybrid environment with at least 3 days in office

Career growth and development opportunities

Opportunities to contribute to community causes

Work with diverse team members in an inclusive and collaborative environment

This job posting is for an existing vacancy

We’re looking for our next Senior Principal Site Reliability Engineer. Could It Be You?

The Senior Principal Site Reliability Engineer is directly responsible for the stability, resiliency, and scalability of business-critical brokerage back-end applications running across a hybrid on-premises and cloud architecture.

This individual drives reliability engineering practices — SLOs/SLIs, observability, incident response, and capacity planning — while also making hands‑on, code contributions directly into multiple applications across different technology stacks using an inner‑source model. The role blends deep technical execution with cross‑team influence: this person is expected to identify systemic reliability risks, fix them where they appear, and raise the operational bar for every team they touch.

This position is a strong fit for a hands‑on senior principal engineer who is energized by fixing production reliability at the source, is comfortable navigating multiple codebases and cloud/on‑prem environments, and wants to have an outsized impact on the resiliency of a regulated, high‑availability brokerage platform.

Application Stability, Reliability & Growth

Own the end‑to‑end reliability posture of critical brokerage back‑end applications, driving measurable improvements in availability, latency, and error budgets across on‑premises and cloud environments.

Define and track Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets in partnership with application teams; use them to prioritize reliability work over feature work when warranted.

Lead root cause analysis and blameless post‑incident reviews for high‑severity production incidents; drive remediation items to closure and identify systemic patterns across applications.

Establish and mature observability practices (metrics, logging, tracing, alerting) so that failures are detected proactively and diagnosed quickly across a heterogeneous, multi‑stack estate.

Build capacity planning, load testing, and chaos/failure‑injection practices to validate resilience before incidents occur.

Champion a culture of operational excellence, toil reduction, and AI and automation‑first thinking across engineering teams.

Make hands‑on, code contributions directly into multiple applications spanning different languages, frameworks, and stacks, using an inner‑source model to fix reliability defects, add instrumentation, and improve resiliency patterns.

Partner with individual application teams to raise pull requests, follow their contribution standards, and pair with owning engineers so fixes land safely and are properly reviewed and owned long‑term.

Identify recurring reliability anti‑patterns across codebases (e.g., missing timeouts/retries, unbounded queues, improper connection pooling) and drive standardized, reusable fixes or shared libraries.

Contribute to and help govern internal reliability tooling, shared SDKs, and common patterns (circuit breakers, backoff/retry, health checks) that can be inner‑sourced across teams.

Cloud Scalability & Hybrid Architecture

Design and advise on cloud scalability strategies (auto‑scaling, load balancing, multi‑region/multi‑AZ failover, caching, queuing) for workloads that span on‑premises data centers and public cloud.

Guide capacity and cost‑aware scaling decisions, balancing performance, resiliency, and cloud spend across hybrid deployments.

Evaluate and recommend cloud‑native and hybrid resiliency patterns (e.g., disaster recovery, active‑active/active‑passive architectures, data replication strategies) appropriate for regulated brokerage workloads.

Organizational Awareness & Risk

Bring strong organizational awareness of the operational, financial, regulatory, and reputational risk that production incidents pose to a brokerage business, and factor that into prioritization.

Participate in risk assessments related to system reliability, availability, and disaster recovery, partnering with Risk, Compliance, and Information Security as needed.

Contribute to change management and release governance practices that reduce the likelihood and blast radius of production incidents.

Promptly identify, escalation, and help remediate reliability or security‑related incidents in accordance with company policy.

Leadership & Influence

Act as a technical reference and mentor for reliability engineering practices, coaching application teams on operational excellence without formal direct reports.

Influence architecture and design decisions across multiple teams by bringing a reliability and scalability lens to reviews and planning.

Document and evangelize reliability standards, runbooks, and best practices; lead or contribute to internal tech talks and communities of practice.

Partner with engineering leadership to define the reliability roadmap and report on progress against stability goals

So are YOU our next Senior Principal Site Reliability Engineer. You are if you…

Bachelor's or Master's degree in Computer Science, Information Systems, Engineering, or a related field, or equivalent combination of education and experience.

8+ years of software engineering and/or site reliability engineering experience, including production ownership of business‑critical applications; financial services or brokerage experience strongly preferred.

Demonstrated ability to read, debug, and make minor‑to‑moderate code changes across multiple languages/stacks (e.g., Java, .NET, Node.js/TypeScript, Python) in an inner‑source or cross‑team contribution model.

Deep experience with cloud scalability strategies on one or more major providers (AWS, Azure, GCP), including auto‑scaling, load balancing, multi‑region resiliency, and cost‑aware capacity planning.

Experience operating and supporting hybrid architectures spanning on‑premises data centers and cloud environments.

Strong background in observability tooling (e.g., Prometheus/Grafana, Datadog, Splunk, ELK, AppDynamics, Dynatrace) and building actionable alerting and dashboards.

Practical experience defining and operating against SLOs/SLIs/error budgets and running blameless post‑incident reviews.

Experience with CI/CD pipelines and infrastructure‑as‑code (e.g., Terraform, Ansible, CloudFormation) in support of reliable, repeatable deployments.

Solid understanding of microservices architecture, distributed systems failure modes, and resiliency patterns (circuit breakers, retries/backoff, bulkheads, timeouts).

Familiarity with relational and NoSQL data stores and their operational/scaling characteristics.

Experience with incident management and on‑call practices (e.g., PagerDuty, Opsgenie) including leading major incident response

Knowledge of security, audit, and regulatory considerations relevant to brokerage / financial services production systems.

Excellent communication skills, with the ability to influence engineers and stakeholders across many teams without direct authority.

Strong documentation, analytical, and problem‑solving skills

Compensation Information:

Base salary range: $150,000 - $190,000

The final compensation package will be commensurate with the successful candidate's experience, skills, and geographic location (Canada). It includes a comprehensive benefits plan and a competitive incentive (bonus) program for Full‑Time Permanent roles.

#LI-Hybrid

At Questrade Financial Group of Companies, with multiple office locations around the world, we are committed to fostering a diverse, inclusive and accessible work environment. This is an environment where individuals are treated with dignity and respect. Here, the unique skills and experience you bring will be valued. You will be supported and motivated, so that you can harness your unlimited potential. Our team reflects the diversity of the communities we serve and operate in. Having a collaborative and diverse team helps us push boundaries to bring the future of fintech into existence—not only for the benefit of our customers, but for those who build their career with us.

Questrade Financial Group of companies Applicant Tracking System utilizes artificial intelligence (AI) for application screening. The AI system operates on predetermined criteria, with final decisions subject to human review.

Candidates selected for an interview will be contacted directly. If you require accommodation during the recruitment/selection process, please let us know and we will work with you to meet your needs.

5700 Yonge St, North York, ON M2M 4K2, Canada

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Cloud Platform Engineer
Senior Cloud Platform Engineer

Community Trust Company • Canada

Hybrid
CAD 120,000 - 140,000
Hybrid work environment
Comprehensive benefits package
Paid vacation and sick days
+1
Principal Software Engineer
Principal Software Engineer

Community Trust Company • Canada

Hybrid
CAD 135,000 - 165,000
Health & wellbeing resources
Paid vacation and sick days
Hybrid work environment with 3 days in
+2
Principal AI Engineer - AI Engineering & Enablement
Principal AI Engineer - AI Engineering & Enablement

Community Trust Company • Canada

Hybrid
CAD 115,000 - 170,000
Health & wellbeing resources
Hybrid work model
Competitive compensation and benefits
Director, Customer Engagement Platforms
Director, Customer Engagement Platforms

Community Trust Company • Canada

Hybrid
CAD 140,000 - 190,000
Health benefits
Hybrid work
Paid vacation
+3
Senior Full Stack API Engineer
Senior Full Stack API Engineer

Community Trust Company • Canada

Hybrid
CAD 110,000 - 140,000
Hybrid work environment
Comprehensive benefits
Bonus program
Senior Software Engineer
Senior Software Engineer

Community Trust Company • Toronto

On-site
CAD 120,000 - 140,000
Benefits plan
Bonus program
Hybrid work environment
Senior Manager, Learning, Quality & Enablement
Senior Manager, Learning, Quality & Enablement

Community Trust Company • Canada

Hybrid
CAD 105,000 - 130,000
Hybrid work environment
Career growth
Benefits package
Principal Front End Engineer
Principal Front End Engineer

Community Trust Company • Canada

Hybrid
CAD 130,000 - 170,000
Health & wellbeing resources
Paid vacation and sick days
Hybrid work environment
Senior Principal AI Engineer
Senior Principal AI Engineer

Questrade Financial Group • Toronto

Hybrid
CAD 165,000 - 180,000
Team Lead, Software Engineering (Full-Stack)
Team Lead, Software Engineering (Full-Stack)

Community Trust Company • Canada

Hybrid
CAD 150,000 - 180,000
Hybrid workplace in Canada
Comprehensive benefits
Bonus program