Staff Data Scientist, AI Evaluations Platform

RBC

Toronto

On-site

CAD 130,000 - 190,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Total rewards program
Career development opportunities
Training and development

Job summary

RBC’s AI Group is seeking a Staff Data Scientist, AI Evaluations, to lead the design and execution of enterprise-wide evaluation frameworks for model and agent quality, safety, and performance. You will mentor other data scientists and partner with research, engineering, product, and governance teams to embed rigorous evaluation into development and release workflows.

You will own end-to-end evaluation datasets, scorecards, and measurement strategies.

Qualifications

  • 8+ years of experience in data science, applied ML, AI evaluation, ML quality, or related field with leadership/mentorship experience.
  • Experience designing evaluation frameworks for ML or generative/agentic AI systems.
  • Strong foundation in data science, statistics, ML, experimentation, data curation, Python, SQL, and modern AI/ML practices.

Responsibilities

  • Design and drive advanced model and agent evaluation methodologies, datasets, rubrics, and human evaluation protocols.
  • Define evaluation science standards translating risk, governance, and business needs into measurable criteria.
  • Own end-to-end lifecycle for evaluation datasets and scorecards including sourcing, validation, and lineage.
  • Design scalable evaluation approaches for generative AI and agentic systems across multiple levels.
  • Partner with cross-functional teams to embed evaluations into build, release, and monitoring workflows.
  • Establish human evaluation and review protocols with adjudication and audit-ready evidence.
  • Measure scorer accuracy, calibration, and robustness across automated and human evaluation methods.
  • Provide technical leadership and mentorship to junior data scientists.
  • Communicate with RBC partners across Canada/worldwide.

Skills

Data science leadership
ML evaluation
Python
SQL
Communication
Stakeholder management
CI/CD
Observability tooling

Tools

MLflow
Langfuse
LangSmith
OpenTelemetry
Grafana
CI/CD pipelines

Job description

What is the opportunity?

RBC’s AI Group is building trusted AI capabilities for the enterprise, and evaluation is one of the core controls that makes that possible. As Staff Data Scientist, AI Evaluations, you will be a senior technical leader in the data science function responsible for how RBC measures model and agent quality, safety, risk, and performance. You will design, build, and continuously improve evaluation datasets, sourcing methods, LLM judge design, deterministic scorers, human evaluation protocols, quality benchmarks, and measurement frameworks that help AI systems move from experimentation to production with evidence and control, while providing technical guidance and mentorship to other data scientists on the team.

What will you do?
  • Design and drive advanced model and agent evaluation methodologies, including evaluation datasets, rubrics, LLM-as-judge methods, deterministic scorers, human evaluation, and measurement frameworks, providing technical guidance to junior team members.
  • Define evaluation science standards that translate model risk, responsible AI, product quality, safety, and business expectations into measurable criteria, repeatable methods, and clear evidence.
  • Own the end-to-end lifecycle for evaluation datasets and scorecards, including sourcing, curation, validation, quality checks, versioning, lineage, reuse, and ongoing improvement.
  • Design scalable evaluation approaches for generative AI and agentic systems, including task-level, workflow-level, trajectory-level, and runtime evaluation methods.
  • Partner with AI research, platform engineering, product, risk, governance, and business teams to embed evaluations into build, release, certification, monitoring, and recertification workflows.
  • Establish human evaluation and review protocols that produce reliable labels, reviewer guidance, adjudication processes, quality controls, and audit-ready evidence.
  • Measure and improve scorer accuracy, calibration, robustness, failure-mode coverage, and explainability across automated and human evaluation approaches.
  • Provide clear technical leadership, executive-ready communication, and mentorship to junior data scientists to help RBC scale trusted AI with speed, rigor, and control.
  • In this role, you will communicate and interact frequently with RBC partners and/or employees located across Canada and/or worldwide.
What do you need to succeed?
Must Have
  • 8+ years of experience in data science, applied machine learning, AI evaluation, ML quality, or a related technical field, including experience providing technical leadership and mentorship within a team.
  • Strong experience designing evaluation frameworks for ML, generative AI, or agentic AI systems, including metrics, datasets, benchmarks, rubrics, and quality measurement.
  • Practical experience with LLM evaluation methods such as LLM-as-judge, deterministic scoring, human evaluation, hallucination assessment, factuality assessment, safety evaluation, or model quality benchmarking.
  • Strong technical foundation in data science, statistics, machine learning, experimentation, data curation, Python, SQL, and modern AI/ML development practices.
  • Proven ability to translate governance, model risk, responsible AI, and business requirements into measurable controls, repeatable evaluation processes, and decision-ready evidence.
  • Strong communication and stakeholder management skills, with the ability to influence senior leaders across research, engineering, product, governance, risk, and business teams.
Nice to Have
  • Experience evaluating agentic AI systems, tool-calling workflows, multi-step reasoning, runtime traces, trajectory scoring, or workflow-level performance.
  • Experience in financial services, regulated AI, model risk management, responsible AI, enterprise governance, or audit-ready evidence processes.
  • Familiarity with tools and platforms such as MLflow, Langfuse, LangSmith, OpenTelemetry, Grafana, CI/CD pipelines, or comparable evaluation and observability tooling.
  • Publications, patents, open-source contributions, or industry work related to AI evaluation, ML quality, AI safety, applied research, or responsible AI.
What’s in it for you?

We thrive on the challenge to be our best, progressive thinking to keep growing, and working together to deliver trusted advice to help our clients thrive and communities prosper. We care about each other, reaching our potential, making a difference to our communities, and achieving success that is mutual.

  • A comprehensive Total Rewards Program including bonuses and flexible benefits, competitive compensation, commissions, and stock where applicable
  • Leaders who support your development through coaching and managing opportunities
  • Ability to make a difference and lasting impact
  • Work in a dynamic, collaborative, progressive, and high-performing team
  • A world-class training program in financial services
  • Opportunities to do challenging work.
Job Skills

Big Data Analytics, Critical Thinking, Decision Making, Industry Knowledge, Machine Learning (ML), Results-Oriented, Software Engineering, Software Product Design

Additional Job Details

Address: RBC WATERPARK PLACE, 88 QUEENS QUAY W:TORONTO

City: Toronto

Country: Canada

Work hours/week: 37.5

Employment Type: Full time

Platform:

Job Type: Regular

Pay Type: Salaried

Posted Date: 2026-07-13

Application Deadline: 2026-08-20

Note: Applications will be accepted until 11:59 PM on the day prior to the application deadline date above

Our Employment Opportunities

At RBC, we are guided by living shared values of Client First, Integrity, Collaboration, Respect and Excellence and winning together as One RBC. We believe an inclusive workplace that has diverse perspectives is core to our continued growth as one of the largest and most successful banks in the world. Maintaining a workplace where our employees feel supported to perform at their best, effectively collaborate, drive innovation, and grow professionally helps to bring our Purpose to life and create value for our clients and communities. RBC strives to deliver this through policies and programs intended to foster a workplace based on respect, belonging and opportunity for all.

Join our Talent Community

Stay in-the-know about great career opportunities at RBC. Sign up and get customized info on our latest jobs, career tips and Recruitment events that matter to you.

Expand your limits and create a new future together at RBC. Find out how we use our passion and drive to enhance the well‑being of our clients and communities at jobs.rbc.com

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Data Scientist, AI Evaluations Platform
Staff Data Scientist, AI Evaluations Platform

Socket.dev • Toronto, Quebec

On-site
CAD 140,000 - 190,000
Total rewards program
Coaching and development
Stock where applicable
+1
Staff AI Engineer
Staff AI Engineer

RBC • Toronto

On-site
CAD 140,000 - 190,000
Staff Data/AI Engineer
Staff Data/AI Engineer

RBC • Toronto

On-site
CAD 140,000 - 190,000
Bonuses
Flexible benefits
Stock options
+1
Senior Director, AI Research
Senior Director, AI Research

ODAIA • Toronto

On-site
CAD 220,000 - 320,000
Senior Software Engineer
Senior Software Engineer

RBC • Toronto

On-site
CAD 120,000 - 190,000
Senior Director, Responsible AI
Senior Director, Responsible AI

RBC • Toronto

On-site
CAD 180,000 - 260,000
Director, Program Delivery
Director, Program Delivery

ODAIA • Toronto

On-site
CAD 150,000 - 210,000
Lead Data and AI Quality Engineer
Lead Data and AI Quality Engineer

RBC • Toronto

On-site
CAD 140,000 - 180,000
Data Scientist, AI Model Risk
Data Scientist, AI Model Risk

RBC • Toronto

On-site
CAD 110,000 - 170,000
Staff Machine Learning Software Engineer
Staff Machine Learning Software Engineer

RBC • Vancouver

On-site
CAD 180,000 - 230,000