Lead AI Platform Engineer

EQ Bank

Toronto

On-site

CAD 140,000 - 180,000

Full time

4 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

EQ Bank is seeking a Lead AI Platform Engineer to provide technical leadership across the design, implementation, and operation of enterprise AI platforms. You will drive reliability, observability, and security, guiding cross‑functional teams to deliver scalable, secure AI capabilities.

You will mentor engineers, establish engineering patterns, and coordinate complex delivery across architecture, security, and cloud teams to enable safe AI adoption at scale.

Qualifications

  • University degree or equivalent practical experience in CS/Engineering/IT.
  • 7+ years of experience in platform engineering, SRE/DevOps, cloud operations, or production platform support.
  • Experience leading technical delivery, production readiness, incident response, and operational reporting for enterprise platforms.

Responsibilities

  • Lead engineering, configuration, and operation of enterprise AI platforms for availability, performance, resilience, and scale.
  • Define and implement platform engineering patterns, standards, reusable components, and guardrails for secure AI delivery.
  • Provide technical leadership for incident triage, escalation, and post-incident reviews to improve reliability.
  • Report service reliability indicators, incident trends, and operational improvements.
  • Lead enablement of approved AI use cases into production environments with readiness and transition planning.
  • Collaborate with architecture, security, cloud, and infra teams to translate requirements into secure, scalable platform implementations.
  • Guide platform lifecycle management including release coordination and maintenance planning.
  • Ensure platform changes meet engineering, security, and control readiness prior to release.
  • Mentor engineers on observability, automation, troubleshooting, and reliability practices.

Skills

Cloud platforms
Observability
Automation
CI/CD
Infrastructure as Code
Azure
AKS
Security patterns

Education

University degree in Computer Science, Engineering, IT, or related field

Tools

Azure DevOps
GitHub Actions
Terraform
Bicep
Logic Apps
Key Vault
App Service

Job description

Purpose Of The Job

The Lead AI Platform Engineer is accountable for technical leadership, engineering excellence, reliability, operability, and the controlled enablement of the organization’s enterprise AI platforms.

This role provides hands‑on technical leadership across AI platform design, implementation, automation, observability, and production readiness. The incumbent ensures AI platform services and solutions are secure, resilient, observable, supportable, and compliant with enterprise standards for reliability, security, platform management, monitoring, incident coordination, and governance control enforcement. The incumbent acts as a senior technical lead for platform engineering activities, guiding implementation decisions, establishing engineering patterns, mentoring team members, and partnering with cross‑functional stakeholders to enable the safe and scalable adoption of AI across the enterprise.

Main Activities
AI Platform Engineering Leadership, Reliability and Operations
  • Lead the engineering, configuration, and operation of enterprise AI platforms to ensure availability, performance, resilience, and scalability across environments, integrations, and supporting infrastructure.
  • Define and implement platform engineering patterns, standards, reusable components, and operational guardrails that support secure and reliable AI solution delivery.
  • Provide technical leadership for platform triage, incident resolution, escalation coordination, and post‑incident reviews to strengthen service stability and resilience.
  • Track and report on service reliability indicators, incident trends, engineering risks, and operational performance improvements.
AI Platform Enablement, Solution Readiness, Production Readiness and Technical Delivery
  • Lead technical enablement of approved AI use cases into non‑production and production environments by ensuring environment readiness, dependency validation, release readiness, operational supportability, and service transition planning.
  • Partner with architecture, security, cloud, infrastructure, delivery, and application teams to translate solution requirements into secure, supportable, and scalable platform implementations.
  • Guide platform lifecycle management through release coordination, change readiness validation, maintenance planning, capacity planning, and technical risk mitigation.
  • Ensure AI platform changes meet defined engineering, operational, security, and control readiness criteria prior to release.
Observability, Automation and AI Ops Engineering
  • Design, implement, and continuously improve observability capabilities, including telemetry, logging, metrics, traces, dashboards, and alerting required for enterprise AI operations.
  • Lead automation initiatives using approved tools and practices to reduce manual effort, improve reliability, and standardize repeatable operational activities.
  • Analyze operational data to identify anomalies, recurring issues, root‑cause patterns, performance bottlenecks, and opportunities for proactive service improvement.
  • Implement AI Ops use cases such as alert correlation, anomaly detection, forecasting, root‑cause support, knowledge retrieval, and automation of repetitive operational tasks.
  • Mentor engineers on observability, automation, troubleshooting, and service reliability practices.
Governance, Risk and Control Engineering Execution
  • Embed governance, security, privacy, auditability, traceability, and human oversight requirements into AI platform engineering and operational practices.
  • Ensure platform implementations align with enterprise security policies, risk controls, compliance requirements, architecture standards, and operational readiness expectations.
  • Partner with security, risk, compliance, architecture, and data teams to assess implementation risks, close control gaps, and enable responsible deployment of AI capabilities.
  • Maintain documentation and evidence required for audit, governance reviews, production readiness checkpoints, and control validation.
  • Identify technical and control risks, recommend mitigation options, and escalates appropriately to relevant governance and risk stakeholders.
AI Asset Visibility, Platform Integrity, Operational Integrity and Engineering Standards
  • Maintain engineering and operational visibility of AI platform assets required for monitoring, support, ownership, lifecycle management, and cost alignment.
  • Validate asset ownership, relationships, configuration integrity, and lifecycle status in collaboration with application, platform, architecture, and infrastructure owners.
  • Establish and promote engineering standards, reusable patterns, documentation, and technical practices that improve platform supportability and operational integrity.
Knowledge/Skill Requirements
  • University degree in Computer Science, Engineering, Information Technology, or a related field, or equivalent practical experience.
  • 7+ years of experience in platform engineering, site reliability engineering, DevOps, cloud operations, enterprise IT operations, or production platform support.
  • Demonstrated experience leading technical delivery, engineering standards, production readiness, incident response, problem management, service restoration, and operational reporting for enterprise platforms.
Technical Expertise
  • Advanced experience with cloud platforms, observability, automation, configuration management, and integration patterns, including Azure Automation runbooks, Azure AI, Copilot integrations, AKS, virtual networks, App Service, and supporting Azure services.
  • Strong expertise with observability tools such as Azure Monitor, Application Insights, Log Analytics, Grafana, dashboards, alerting, and operational telemetry design.
  • Strong experience with CI/CD, automation, and infrastructure-as-code tools such as Azure DevOps, GitHub Actions, Logic Apps, Bicep, Terraform, Azure Policy, Key Vault, and related open‑source technologies.
  • Knowledge of integration and event‑driven technologies such as API Management, open‑source API tools, Service Bus, Event Grid, and Apache Kafka.
  • Working knowledge of platform‑supporting data and search services such as Elastic, Azure AI Search, Cosmos DB, and related data platform capabilities.
  • Knowledge of enterprise network, edge security, identity, access management, and related internal platforms such as DNA, Fortinet, and Akamai is an asset.
Additional Capabilities
  • Strong working knowledge of AI/ML operational concepts, including model lifecycle support, platform telemetry, governance controls, human‑in‑the‑loop practices, responsible AI considerations, and production monitoring.
  • Strong understanding of ITIL/ITSM processes, including change, release, incident, problem, configuration, service reporting, and operational risk practices.
  • Proven ability to provide technical leadership, mentor engineers, influence standards, guide implementation decisions, and coordinate complex cross‑functional delivery.
  • Analytical and structured thinker with advanced troubleshooting, root‑cause analysis, prioritization, risk assessment, and continuous improvement skills.
  • Strong service orientation, professional maturity, and the ability to collaborate effectively across operations, engineering, security, risk, data, architecture, and business teams.
  • Experience creating technical documentation, engineering patterns, operational procedures, support playbooks, dashboards, and user guidance materials.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead AI Platform Engineer
Lead AI Platform Engineer

Eqbank • Toronto

On-site
CAD 140,000 - 190,000
Lead AI Platform Engineer
Lead AI Platform Engineer

Kinvie • Toronto

Hybrid
CAD 140,000 - 190,000
Lead AI Platform Engineer
Lead AI Platform Engineer

EQ Bank | Canada's Challenger Bank • Toronto

On-site
CAD 120,000 - 190,000
Senior AI Platform Operations Engineer
Senior AI Platform Operations Engineer

EQ Bank • Toronto

On-site
CAD 120,000 - 160,000
Senior AI Platform Engineer
Senior AI Platform Engineer

Luxoft • Toronto

On-site
CAD 120,000 - 180,000
Staff Engineer (AI & Engineering)
Staff Engineer (AI & Engineering)

EQ Bank | Canada's Challenger Bank • Toronto

On-site
CAD 150,000 - 210,000
Staff Engineer (AI & Engineering)
Staff Engineer (AI & Engineering)

Kinvie • Toronto

On-site
CAD 140,000 - 190,000
Forward Deployed AI Engineer
Forward Deployed AI Engineer

Kinvie • Toronto

On-site
CAD 90,000 - 120,000
Lead AI Platform Owner
Lead AI Platform Owner

PowerToFly • Toronto

Hybrid
CAD 150,000 - 210,000
Hybrid Work Model
Career Development and Growth
Mental Health Days
+2
AI Platform Engineer with DevOps
AI Platform Engineer with DevOps

Apptoza Inc. • Halifax

On-site
CAD 110,000 - 140,000