Senior AI Platform Operations Engineer

EQ Bank

Toronto

On-site

CAD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

EQ Bank is seeking a Senior AI Platform Operations Engineer to ensure reliability, security, and operability of its AI platform across production environments. You will drive observability, incident response, and governance while enabling scalable AI adoption and continuous improvement in line with enterprise standards for reliability, security, and compliance.

The role emphasizes automation, release readiness, change validation, and platform lifecycle management to support trusted AI services

Qualifications

  • University degree or equivalent practical experience in a relevant field.
  • Experience in platform operations, site reliability, DevOps, cloud operations, or enterprise IT operations.

Responsibilities

  • Administer and operate the AI platform to ensure availability, performance, and resilience across environments.
  • Lead operational triage, escalation coordination, and post-incident reviews to strengthen services stability.
  • Enable approved AI use cases into production with readiness checklists and structured service transitions.
  • Implement observability, telemetry, logging, metrics, and traces for enterprise AI operations.
  • Ensure governance controls for AI solutions including data privacy, auditability, and human oversight.

Skills

Platform operations
Site reliability engineering
DevOps
Cloud operations
Operational reporting
Incident response
Problem management

Education

University degree in Computer Science, Engineering, Information Technology, or related field

Tools

Azure Monitor
Grafana
Azure DevOps
GitHub Actions
Terraform
Azure AI
Kafka
Cosmos DB
API Management

Job description

Purpose of Job

The Senior AI Platform Operations Engineer is accountable for the reliability, operability, and controlled enablement of the organization’s AI platform. This role ensures that AI Platform services and solutions are production-ready, secure, observable, and compliant by executing disciplined operational practices, across platform management, monitoring, incident coordination, and governance control enforcement. The incumbent plays a key role in enabling the safe and scalable adoption of AI by ensuring that AI solutions are deployed, monitored, supported, and continuously improved in line with enterprise standards for reliability, security, and compliance.

Main Activities
AI Platform Reliability and Operations
  • Administer and operate the AI platform to ensure availability, performance, and resilience across environments, integrations, and supporting infrastructure.
  • Monitor platform health using dashboards, logs, metrics, and alerts, and coordinate incident and service restoration activities.
  • Lead operational triage, escalation coordination, and post-incident reviews to strengthen services stability and resilience.
  • Track and report on service reliability indicators, incident trends, and operational performance.
AI Platform Enablement & Production Readiness
  • Enable approved AI use cases into production by ensuring:
    • Environment readiness
    • Dependency validation
    • Completion of operational readiness checklists
    • Structured service transition activities
  • Support platform lifecycle management through:
    • Release coordination
    • Change readiness validation
    • Maintenance and capacity planning
  • Ensure AI platform changes meet defined operational and control readiness criteria prior to release.
Observability, Automation & AI Ops
  • Implement and maintain observability capabilities, including telemetry, logging, metrics, and traces required for enterprise AI operations.
  • Analyze operational data to identify anomalies, recurring issues, root‑cause patterns.
  • Implement AI Ops use cases such as:
    • Alert correlation
    • Anomaly detection
    • Root‑cause support
    • Forecasting and predictive insights
    • Automation of repetitive operational tasks
  • Continuously improve operational efficiency through targeted automation and process optimization.
Governance, Risk, & Control Execution
  • Execute governance controls for AI solutions, including:
    • Usage and access controls
    • Data privacy considerations
    • Auditability and traceability
    • Human oversight requirements
  • Ensure operational practices align with enterprise security policies, risk controls, and compliance requirements.
  • Maintain documentation and evidence required for audit, governance reviews, production readiness checkpoints, and control validation.
  • Identify control gaps and elevate risks appropriately to relevant governance and risk stakeholders.
AI Asset Visibility & Operational Integrity
  • Maintain operational visibility of AI platform assets required for monitoring, support, and cost alignment.
  • Validate asset ownership, relationships, and lifecycle status in collaboration with application and platform owners.
  • Support ongoing audits to ensure AI assets and associated cost attribution remain accurate and current.
Knowledge/Skill Requirements
  • University degree in Computer Science, Engineering, Information Technology, or a related field, or equivalent practical experience.
  • 5–7 years of experience in platform operations, site reliability engineering, DevOps, cloud operations, or enterprise IT operations.
  • Strong experience supporting production platforms and services, including monitoring, incident response, problem management, service restoration, and operational reporting.
Technical Expertise
  • Experience with cloud platforms, observability, automation, configuration management, and integration patterns, including Azure Automation runbooks (PowerShell/Python), Azure AI, Copilot integrations, AKS, virtual networks (hub‑and‑spoke), and App Service.
  • Expertise with observability tools such as Azure Monitor, Application Insights, and Grafana.
  • Experience with CI/CD and automation tools such as Azure DevOps, GitHub Actions, and Logic Apps.
  • Knowledge of configuration management and infrastructure‑as‑code tools such as Bicep, Terraform, Azure Policy, Key Vault, and relevant open‑source technologies.
  • Knowledge of integration and event‑driven technologies such as API Management, open‑source API tools, Service Bus, Event Grid, and Apache Kafka.
  • Working knowledge of platform‑supporting data and search services such as Elastic, Azure AI Search, and Cosmos DB.
  • Knowledge of enterprise network, edge security, and related internal platforms such as DNA, Fortinet, and Akamai is an asset.
Additional Capabilities
  • Working knowledge of AI/ML operational concepts, including model lifecycle support, telemetry, governance controls, human‑in‑the‑loop practices, and production monitoring.
  • Strong understanding of ITIL/ITSM processes, including change, release, incident, problem, configuration, and service reporting practices.
  • Analytical and structured thinker with strong troubleshooting, root‑cause analysis, prioritization, and continuous improvement skills.
  • Strong service orientation, professional maturity, and the ability to collaborate effectively across operations, engineering, security, risk, data, and business teams.
  • Experience creating technical documentation, operational procedures, support playbooks, dashboards, and user guidance materials.
  • Knowledge of security, privacy, audit, and compliance considerations relevant to enterprise AI and platform operations.
Job Complexities / Thinking Challenges

This role requires balancing platform reliability, operational efficiency, and governance discipline in a rapidly evolving AI environment.

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Platform Engineer – Senior/Principal
AI Platform Engineer – Senior/Principal

Jobtailor • Toronto

On-site
CAD 150,000 - 230,000
Senior AI Platform Engineer
Senior AI Platform Engineer

Luxoft • Toronto

On-site
CAD 120,000 - 180,000
Senior DevOps Engineer
Senior DevOps Engineer

Tru India • Toronto

Hybrid
CAD 100,000 - 120,000
Staff Engineer (AI & Engineering)
Staff Engineer (AI & Engineering)

EQ Bank • Toronto

On-site
CAD 140,000 - 190,000
Forward Deployed AI Engineer
Forward Deployed AI Engineer

EQ Bank • Toronto

On-site
CAD 140,000 - 190,000
Staff Platform Engineer
Staff Platform Engineer

Robots and Pencils • Calgary

Hybrid
CAD 96,000 - 138,000
Forward Deployed AI Engineer
Forward Deployed AI Engineer

Kinvie • Toronto

On-site
CAD 90,000 - 120,000
AI Platform Ops Engineer – Azure, CI/CD & Automation
AI Platform Ops Engineer – Azure, CI/CD & Automation

Compunnel, Inc. • Toronto

On-site
CAD 80,000 - 110,000
Staff Engineer (AI & Engineering)
Staff Engineer (AI & Engineering)

EQ Bank | Canada's Challenger Bank • Toronto

On-site
CAD 150,000 - 210,000
Staff Engineer (AI & Engineering)
Staff Engineer (AI & Engineering)

Kinvie • Toronto

On-site
CAD 140,000 - 190,000