Principal AI/ML Engineer

Vanguard

Toronto

On-site

CAD 170,000 - 250,000

Full time

37 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Vanguard is seeking a Principal AI Engineer to lead the transformation of cutting-edge AI research into production-ready capabilities across enterprise environments. You will own the AI architecture, oversee engineering, reliability, and scalable deployment while mentoring engineers and shaping an engineering culture focused on operational excellence and responsible AI.

Collaborate with researchers, product leaders and security teams to deliver end-to-end AI solutions that meet performance,

Qualifications

  • 10+ years of software/AI engineering experience in related fields.
  • Experience deploying large-scale AI/ML systems in production.
  • Proven track record leading complex technical initiatives from concept to deployment.
  • Strong knowledge of architecture, reliability engineering, ML Ops, DevOps, and cloud tech.
  • Ability to mentor engineers and lead teams through complex technical challenges.

Responsibilities

  • Define and lead architecture for enterprise-scale AI and ML platforms.
  • Design scalable, resilient AI systems for mission-critical workloads.
  • Drive AI deployment best practices including MLOps and AI observability.
  • Mentor and lead junior engineers; promote operational ownership.
  • Collaborate with researchers, product leaders and engineering teams to move from prototype to production.

Skills

AI architecture
MLOps
DevOps
Cloud technologies
Leadership
Observability
Software engineering
Mentoring

Job description

As a Principal AI Engineer, you will serve as a senior technical leader responsible for transforming state-of-the-art AI research into scalable, production-ready capabilities that create measurable value for our clients. You will lead the architecture, engineering, operationalization, and ongoing reliability of advanced AI systems, ensuring they can scale across enterprise environments while meeting rigorous standards for performance, security, resilience, and responsible AI.

This role sits at the critical intersection of AI research, engineering, product development, and operations. You will partner closely with world-class AI researchers, product leaders, and engineering teams to accelerate the journey from prototype to production. Your work will span some of the most advanced areas of AI, including Large Language Models (LLMs), Trustworthy AI, agentic systems, and emerging AI technologies.

You will mentor engineers, shape architecture, guide production support strategy, and serve as a thought leader for scaling AI across the organization. In addition to building and scaling AI solutions, you will help establish an engineering culture that emphasizes operational excellence, ownership, reliability, and continuous improvement.

Key Responsibilities:
AI Architecture & Technical Leadership
  • Define and lead the technical architecture for enterprise-scale AI and ML platforms.
  • Design scalable, resilient, and reusable AI systems capable of supporting mission-critical workloads.
  • Establish architectural standards, engineering patterns, and best practices for AI deployment and operations.
  • Drive technical decisions around model serving, inference optimization, agent architectures, orchestration frameworks, observability, and AI infrastructure.
Productize AI Research
  • Partner closely with AI researchers to transform cutting‑edge prototypes into production‑grade solutions.
  • Lead efforts to operationalize advanced AI capabilities across areas such as:
    • Large Language Models (LLMs)
    • Trustworthy and Responsible AI
    • Agentic AI Systems
  • Establish repeatable pathways that accelerate innovation‑to‑production cycles.
  • Ensure production solutions maintain scientific rigor while meeting enterprise engineering standards.
  • Bridge the gap between research breakthroughs and sustainable business value.
Engineering Excellence & Scalability
  • Solve the organization's most complex AI engineering and scalability challenges.
  • Design systems that operate reliably at enterprise scale while balancing performance, latency, governance, security, and cost.
  • Drive adoption of MLOps, LLMOps, and AI platform engineering best practices.
  • Improve the robustness, maintainability, observability, and operational readiness of our AI products.
  • Identify and eliminate architectural bottlenecks that impact scale, reliability, or client experience.
  • Raise standards through coaching, architecture reviews, design guidance, and technical leadership.
Production Reliability & Operational Leadership
  • Own the operational excellence, reliability, performance and availability of our products.
  • Lead technical response and resolution efforts for complex production incidents, performance degradation, model failures, and system outages.
  • Serve as the senior technical escalation point for the team's most challenging production challenges.
  • Establish best practices for AI system monitoring, observability, alerting, incident management, capacity planning, and service-level objectives (SLOs).
  • Mentor and lead junior engineers in troubleshooting, root cause analysis, operational decision‑making, and incident response.
  • Drive post‑incident reviews focused on learning, continuous improvement, and long‑term corrective actions.
  • Develop operational processes that ensure AI solutions remain secure, scalable, performant, and reliable for business‑critical use cases.
  • Partner with product, infrastructure, security, and support teams to proactively identify operational risks and continuously improve service reliability.
Mentorship & Thought Leadership
  • Mentor AI and ML engineers within the team.
  • Foster a culture of technical excellence and operational ownership where engineers are accountable not only for building systems, but also for running and supporting them successfully in production.
  • Represent our team as a thought leader in scalable AI deployment, operational excellence, and responsible AI practices.
Required Qualifications
  • 10+ years of experience in software engineering, machine learning engineering, AI engineering, or related technical disciplines.
  • Deep expertise designing, deploying, and supporting large-scale AI and ML systems in production environments.
  • Demonstrated success leading complex technical initiatives from concept through deployment and ongoing operations.
  • Strong knowledge of software architecture, reliability engineering, observability, ML Ops, DevOps, and cloud technologies.
  • Proven ability to mentor engineers and lead teams through highly complex technical and operational ch
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal AI/ML Engineer
Principal AI/ML Engineer

The Vanguard Group • Toronto

Hybrid
CAD 180,000 - 260,000
Principal AI Architect
Principal AI Architect

Harnham • Toronto

Hybrid
CAD 200,000 - 220,000
Flexible hybrid working environment
Competitive CAD 200k–220k salary
Staff Engineer (AI & Engineering)
Staff Engineer (AI & Engineering)

Kinvie • Toronto

On-site
CAD 140,000 - 190,000
Staff Engineer (AI & Engineering)
Staff Engineer (AI & Engineering)

EQ Bank | Canada's Challenger Bank • Toronto

On-site
CAD 150,000 - 210,000
Principal Software Engineer, AI (Web & Data)
Principal Software Engineer, AI (Web & Data)

Okta • Toronto

On-site
CAD 125,000 - 150,000
AI/ML Engineer
AI/ML Engineer

BrainWave Professionals • Canada

On-site
CAD 120,000 - 160,000
Staff Engineer (AI & Engineering)
Staff Engineer (AI & Engineering)

EQ Bank • Toronto

On-site
CAD 140,000 - 190,000
Principal Software Development Engineer - AI
Principal Software Development Engineer - AI

Quarry Consulting • Canada

On-site
CAD 120,000 - 180,000
AI Architect
AI Architect

BrainRidge Consulting • Toronto

Hybrid
CAD 140,000 - 190,000
AI Architect
AI Architect

HCLTech • Vancouver

On-site
CAD 180,000 - 260,000