Lead Platform Reliability Engineer, Global AI Platform & Solutions

United States Digital Space LLC

Toronto

On-site

CAD 113,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

United States Digital Space LLC seeks a Lead Platform Reliability Engineer to ensure stability and performance of the shared AI platform. You’ll blend software engineering with SRE practices to keep systems reliable and developer-friendly.

Responsibilities include defining SLOs/SLIs, observability, incident response, automation, and IaC management with Terraform/Ansible. You’ll collaborate across global teams on security governance and cost optimization.

Qualifications

  • Bachelor’s in Computer Science/Engineering or equivalent experience.
  • 5–8 years in DevOps/Platform Engineering or Production Operations.
  • Experience operating large-scale distributed systems with on-call.
  • Cloud-native development: Azure, Kubernetes, containers, CI/CD, observability stacks.
  • Knowledge of Python and/or Java/Scala/TypeScript for backend/services automation.
  • Understanding AI solutions, LLMs, retrieval architectures, vector stores, and orchestration.
  • API design, asynchronous workflows, concurrency, SREs and reliability.
  • Security, governance, and compliance for AI/data systems.

Responsibilities

  • Define SLOs/SLIs, track budgets, plan capacity, tune autoscaling.
  • Build and maintain logging, metrics, tracing, alerting; dashboards.
  • On-call incident response; triage, root-cause, postmortems.
  • Develop self-service capabilities and CI/CD pipelines; tooling automation.
  • Manage clusters, networks, storage, policies via Terraform/Ansible; prevent drift.
  • Enforce RBAC, secrets management, supply chain security; audits.
  • Optimize resource usage and cost; rightsizing and reservations.

Skills

SRE & DevOps
Cloud-native
Python/Java/TypeScript
AI/LLM awareness
API design
Security & governance

Education

Bachelor’s in CS/Engineering or equivalent

Tools

Azure
Kubernetes
CI/CD
Terraform
Ansible

Job description

The Lead Platform Reliability Engineer (PRE) ensures the stability, performance, and scalability of the shared platform that supports internal AI solution development. It combines software engineering, SRE practices, and operations to keep the platform reliable and developer-friendly.

Position Responsibilities:
  • Reliability and performance: Define SLOs/SLIs, track operations budgets, reduce MTTR, capacity plan, and tune autoscaling.
  • Observability: Build and maintain logging, metrics, tracing, and alerting; instrument platform components; create runbooks and dashboards.
  • Incident response: On-call for platform incidents; triage, mitigate, root-cause, and drive postmortems and corrective actions.
  • Automation and tooling: Develop self-service capabilities, AIOps/MLOps/GitOps/CICD pipelines, and operational automations (provisioning, upgrades, backups).
  • Infrastructure as code: Manage clusters, networks, storage, and policies via Terraform/Ansible; prevent configuration drift.
  • Security and compliance: Enforce identity/RBAC, secrets management, supply chain security, and regulatory controls; collaborate with risk and audit.
  • Scalability and cost: Optimize resource usage, plan capacity, control spend (rightsizing, autoscaling, reservations/spot).
  • Change management: Safe rollouts, progressive delivery, and policy-as-code guardrails.
  • Platform productization: Treat the platform as a product, define operations SLAs in alignment to product roadmap, service catalog, and developer experience.
  • Collaborate with global engineering, security, and AI governance teams to ensure compliance with cross-geo regulations and Asia’s data residency requirements.
  • Operate scalable backend services supporting high-traffic agent interactions, retrieval operations, and real-time execution flows.
  • Maintain AI services runbooks, playbooks, and enablement for GOCC
Required Qualifications:
  • Bachelor’s in Computer Science/Engineering or equivalent experience (not strictly required if skills demonstrated).
  • 5-8 years experience in DevOps/Platform Engineering or Production Operations.
  • Proven track record operating large-scale distributed systems and running on-call.
  • Operational experience with cloud-native development: Azure, Kubernetes, containers, CI/CD, and observability stacks.
  • Knowledge with Python and/or Java/Scala/TypeScript for building backend services and automation.
  • Understanding of AI solution, LLM systems, retrieval architectures, embeddings, vector stores, prompt/tool orchestration, and agent workflow fundamentals.
  • Knowledge of API design, asynchronous workflows, concurrency, reliability engineering (SLOs, error budgets), and performance tuning.
  • Familiarity with security, governance, and compliance for AI/data systems (authN/authZ, data protection, audit logging, model governance).
  • Ability to collaborate across global teams and translate business requirements into platform capabilities and operational SLAs.
Preferred Qualifications:
  • ITIL & ITSM certification
  • Azure Administrator/DevOps certificate (nice to have)
  • Kubernetes: CKA/CKS certificate (nice to have)
  • HashiCorp Terraform Associate certificate (nice to have)
When you join our team:
  • We’ll empower you to learn and grow the career you want.
  • We’ll recognize and support you in a flexible environment where well-being and inclusion are more than just words.

As part of our global team, we’ll support you in shaping the future you want to see.

#LI-Hybrid

The role being advertised is an existing vacancy.

About the company and John Hancock

the company Financial Corporation is a leading international financial services provider, helping people make their decisions easier and lives better. To learn more about us, visit https://www.the company.com/en/about/our-story.html .

the company is an Equal Opportunity Employer

At the company/John Hancock, we embrace our diversity. We strive to attract, develop and retain a workforce that is as diverse as the customers we serve and to foster an inclusive work environment that embraces the strength of cultures and individuals. We are committed to fair recruitment, retention, advancement and compensation, and we administer all of our practices and programs without discrimination on the basis of race, ancestry, place of origin, colour, ethnic origin, citizenship, religion or religious beliefs, creed, sex (including pregnancy and pregnancy-related conditions), sexual orientation, genetic characteristics, veteran status, gender identity, gender expression, age, marital status, family status, disability, or any other ground protected by applicable law.

It is our priority to remove barriers to provide equal access to employment. A Human Resources representative will work with applicants who request a reasonable accommodation during the application process. All information shared during the accommodation request process will be stored and used in a manner that is consistent with applicable laws and the company/John Hancock policies. To request a reasonable accommodation in the application process, contact hr@unitedstatesdigital.space .

Referenced Salary Location

Toronto, Ontario

Working Arrangement

Hybrid

Salary range is expected to be between

$113,260.00 CAD - $210,340.00 CADEmployees also have the opportunity to participate in incentive programs and earn incentive compensation tied to business and individual performance. The actual salary will vary depending on local market conditions, geography and relevant job-related factors such as knowledge, skills, qualifications, experience, and education/training. If you are applying for this role outside of the primary location, please contact hr@unitedstatesdigital.space for the salary range for your location.

the company offers eligible employees a wide array of customizable benefits, including health, dental, mental health, vision, short- and long-term disability, life and AD\&D insurance coverage, adoption/surrogacy and wellness benefits, and employee/family assistance plans. We also offer eligible employees various retirement savings plans (including pension and a global share ownership plan with employer matching contributions) and financial education and counseling resources. Our generous paid time off program in Canada includes holidays, vacation, personal, and sick days, and we offer the full range of statutory leaves of absence. If you are applying for this role in the U.S., please contact hr@unitedstatesdigital.space for more information about U.S.-specific paid time off provisions.

We use data and analytics technologies, such as artificial intelligence (AI), and automated processing tools, to analyze and process the information you provide to us or third parties in the application process. For more information, please refer to our personal information collection statement .

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Back-End Software Engineer
Back-End Software Engineer

United States Digital Space LLC • Toronto

Hybrid
CAD 86,000 - 136,000
Health benefits
Pension & share plan
Paid time off
Lead Platform Reliability Engineer, Global AI Platform & Solutions
Lead Platform Reliability Engineer, Global AI Platform & Solutions

Manulife Financial • Toronto

Hybrid
CAD 113,000 - 210,000
Hybrid work model
Benefits package
Career development opportunities
Senior Full-Stack Software Engineer - Global AI Platform
Senior Full-Stack Software Engineer - Global AI Platform

Manulife Financial • Toronto

Hybrid
CAD 113,000 - 163,000
Back-end Software Engineer (Kotlin)
Back-end Software Engineer (Kotlin)

United States Digital Space LLC • Toronto

Hybrid
CAD 86,000 - 137,000
Health insurance
Dental insurance
Mental health benefits
+6
Cloud/DevOps Engineer, Manulife Bank Technology
Cloud/DevOps Engineer, Manulife Bank Technology

Manulife Financial • Waterloo

Hybrid
CAD 113,000 - 163,000
Health
Dental
Vision
+4
Director, IT Systems and Operations
Director, IT Systems and Operations

Manulife • Toronto

Hybrid
CAD 125,000 - 175,000
Incentive programs
Health benefits
Solution Architect - Hybrid
Solution Architect - Hybrid

Manulife • Toronto

Hybrid
CAD 86,000 - 136,000
AVP, AI
AVP, AI

Manulife • Toronto

Hybrid
CAD 130,000 - 241,000
Health and dental benefits
Pension plan
Generous paid time off
Software Engineer(s), Manulife Bank Technology Team
Software Engineer(s), Manulife Bank Technology Team

Socket.dev • Southwestern Ontario

Hybrid
CAD 86,000 - 136,000
AVP, AI Delivery
AVP, AI Delivery

Manulife • Toronto

Hybrid
CAD 130,000 - 241,000