Ingénieur principal, Opérations de plateforme et observabilité /Senior Engineer, Platform Operations & Observabilityobservabilité

McKesson’s Corporate

Montreal (administrative region)

Hybrid

CAD 130,000 - 217,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Hybrid work model

Job summary

McKesson Canada is seeking a Senior Lead Engineer for Platform Operations & Observability to drive reliability, observability, and operational excellence across healthcare platforms. You will lead monitoring strategies, incident response, and RCA initiatives while guiding engineering teams to scalable, secure systems.

You will collaborate with software, platform, security, and operations teams to automate processes, improve service reliability, and establish best practices for production

Qualifications

  • 7+ years of professional experience in Software Engineering, Site Reliability Engineering, Platform Engineering, DevOps, or related technical roles.
  • Bachelor's degree in Computer Science, Engineering, IT, or equivalent.
  • Experience supporting large-scale production environments and enterprise applications.
  • Hands-on experience with monitoring, observability, logging, alerting, and APM tools.
  • Proven experience leading incident management and production support activities.
  • Experience with CI/CD, automation, DevOps practices, and software delivery pipelines.
  • Experience with microservices, APIs, distributed systems, and cloud architectures.

Responsibilities

  • Lead monitoring and observability strategies across enterprise applications, platforms, and services.
  • Design, implement, and optimize dashboards, alerts, telemetry, logging, and performance monitoring solutions.
  • Drive incident management processes, major incident response, escalation coordination, service restoration activities and postmortem reviews.
  • Conduct RCA investigations and lead corrective and preventive action planning.
  • Partner with engineering teams to improve platform reliability, resiliency, scalability, and operational readiness.
  • Lead change management reviews and promote safe deployment practices.
  • Provide technical leadership, coaching, and mentoring to engineers.
  • Influence architecture, automation, CI/CD, and operational excellence initiatives.

Skills

Software engineering
Site Reliability Engineering
DevOps
Cloud architectures
Lead mentorship

Education

Bachelor's degree in CS/Engineering

Tools

Dynatrace
Prometheus
Kubernetes
Azure

Job description

McKesson is an impact-driven, Fortune 10 company that touches virtually every aspect of healthcare. We are known for delivering insights, products, and services that make quality care more accessible and affordable. Here, we focus on the health, happiness, and well-being of you and those we serve – we care. What you do at McKesson matters. We foster a culture where you can grow, make an impact, and are empowered to bring new ideas. Together, we thrive as we shape the future of health for patients, our communities, and our people.

McKesson is an impact-driven, Fortune 10 company that touches virtually every aspect of healthcare. We are known for delivering insights, products, and services that make quality care more accessible and affordable. Here, we focus on the health, happiness, and well-being of you and those we serve – we care. What you do at McKesson matters. We foster a culture where you can grow, make an impact, and are empowered to bring new ideas. Together, we thrive as we shape the future of health for patients, our communities, and our people.

À propos du poste

Équipe / Projet : Solution numérique B2C Canada. La principale application du portefeuille est une plateforme B2C destinée aux patients en pharmacie. L’équipe est composée d’environ 20 personnes. McKesson est à la recherche d’un Ingénieur principal responsable, Opérations de plateforme et observabilité pour diriger la fiabilité, l’observabilité et l’excellence opérationnelle des plateformes technologiques de santé de l’entreprise. Dans ce rôle, en collaboration avec le responsable de la surveillance applicative et de l’observabilité, vous dirigerez les stratégies de surveillance, les pratiques de gestion des incidents, la gouvernance des changements et les initiatives d’analyse des causes fondamentales, tout en contribuant à la création de systèmes évolutifs, sécurisés et résilients. Vous collaborerez avec les équipes d’ingénierie logicielle, de plateforme, de sécurité et d’exploitation afin d’améliorer la fiabilité des services, d’automatiser les processus opérationnels et d’établir les meilleures pratiques en matière de préparation à la production. Ce poste offre également un leadership technique et du mentorat aux équipes d’ingénierie tout en influençant les normes de fiabilité et la stratégie à long terme des plateformes.

About the Role Team/Project: Canada B2C Digital Solution.

Main application in the portfolio is a B2C Platform for pharmacy patients. Team of around 20 people. McKesson is seeking a Senior Lead Engineer, Platform Operations & Observability to lead the reliability, observability, and operational excellence of enterprise healthcare technology platforms. In this role, in accordance with Application Monitoring & Observability Lead, you will drive monitoring strategies, incident management practices, change governance, and root cause analysis initiatives while helping build scalable, secure, and resilient systems. You will collaborate with software engineering, platform, security, and operations teams to improve service reliability, automate operational processes, and establish best practices for production readiness. This position also provides technical leadership and mentorship to engineering teams while influencing reliability standards and long-term platform strategy.

What You'll Do
  • Lead monitoring and observability strategies across enterprise applications, platforms, and services.
  • Design, implement, and optimize dashboards, alerts, telemetry, logging, and performance monitoring solutions.
  • Drive incident management processes, major incident response, escalation coordination, service restoration activities and postmortem incident.
  • Conduct root cause analysis (RCA) investigations and lead corrective and preventive action planning.
  • Partner with engineering teams to improve platform reliability, resiliency, scalability, and operational readiness.
  • Lead change management reviews and promote safe deployment and release practices.
  • Ability to execute regression and validation test plans following each production deployment.
  • Provide technical leadership, coaching, and mentoring to engineers while establishing engineering best practices.
  • Influence architecture, automation, CI/CD, and operational excellence initiatives supporting enterprise platforms.
Basic Requirements
  • 7+ years of professional experience in Software Engineering, Site Reliability Engineering, Platform Engineering, DevOps, or related technical roles.
  • Bachelor's degree in Computer Science, Engineering, Information Technology, or equivalent experience.
  • Experience supporting large-scale production environments and enterprise applications.
  • Hands-on experience with monitoring, observability, logging, alerting, and application performance monitoring tools.
  • Proven experience leading incident management and production support activities.
  • Experience performing root cause analysis and implementing preventive solutions.
  • Experience with CI/CD, automation, DevOps practices, and software delivery pipelines.
  • Experience with microservices, APIs, distributed systems, and cloud-based architectures.
  • Lead production readiness reviews and operational acceptance activities prior to major releases.
Preferred Skills / Experience
  • Experience with tools such as Dynatrace, Prometheus, Dotcom Monitor or similar observability platforms.
  • Experience with Kubernetes, containers, and cloud platforms such as Azure.
  • Knowledge of ITIL-aligned incidents, problems, and change management practices.
  • Experience defining SLAs, MTTR and service reliability metrics.
  • Demonstrated technical leadership and mentorship of engineering teams.
  • Experience operating in regulated or highly compliant environments.
  • Experience driving platform modernization and operational excellence initiatives.
  • Strong analytical, troubleshooting, continuous improvement and stakeholder communication skills.
  • Fair understanding and mastery of AI tools (Copilot, Rovo) and Ai agents.
Travel / Work Environment / Physical Requirements
  • Hybrid, two mandatory days at the Dobrin Office (Usually on Monday and Wednesday).
  • Ability to work at a computer for extended periods and participate in virtual collaboration activities.
  • Participation in on-call support and prod deployment rotations out of business hours may be required based on organizational needs.

We are proud to offer a competitive compensation package at McKesson as part of our Total Rewards. This is determined by several factors, including performance, experience and skills, equity, regular job market evaluations, and geographical markets. The pay range shown below is aligned with McKesson's pay philosophy, and pay will always be compliant with any applicable regulations. In addition to base pay, other compensation, such as an annual bonus or long-term incentive opportunities may be offered.

Our Base Pay Range for this position $94,400 - $157,300

In addition to base pay, other compensation, such as an annual bonus or long-term incentive opportunities may be offered.

McKesson is an Equal Opportunity Employer

McKesson provides equal employment opportunities to applicants and employees, without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, protected veteran status, disability, age, genetic information, or any other legally protected category.

McKesson is an Equal Employment Opportunity Employer and offers opportunities to all job seekers including job seekers with disabilities.

For additional information on McKesson's full Equal Employment Opportunity policies, visit our Equal Employment Opportunity page.

If you need a reasonable accommodation to assist with your job search or application for employment, please contact us by sending an email to (United States) Disability_Accommodation@McKesson.com or (Canada) Accessibility@mckesson.ca.

McKesson is an impact-driven, Fortune 10 company that touches virtually every aspect of healthcare. We are known for delivering insights, products, and services that make quality care more accessible and affordable. Here, we focus on the health, happiness, and well-being of you and those we serve – we care. What you do at McKesson matters. We foster a culture where you can grow, make an impact, and are empowered to bring new ideas. Together, we thrive as we shape the future of health for patients, our communities, and our people. If you want to be part of tomorrow's health today, we want to hear from you.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Lead Engineer, Platform Operations & Observability
Senior Lead Engineer, Platform Operations & Observability

Mckesson Corporation • Canada

Remote
CAD 131,000 - 218,000
Rémunération compétitive
Travail hybride
Senior Lead Engineer, Platform Operations & Observability/Ingénieur principal, Opérations de pl[...]
Senior Lead Engineer, Platform Operations & Observability/Ingénieur principal, Opérations de pl[...]

McKesson • Montreal (administrative region)

On-site
CAD 131,000 - 218,000
Spécialiste principal des opérations de plateforme et de l’observabilité. /Senior Specialist, Platform Operations & Observability
Spécialiste principal des opérations de plateforme et de l’observabilité. /Senior Specialist, Platform Operations & Observability

Mckesson Corporation • Montreal (administrative region)

On-site
CAD 110,000 - 150,000
Lead Business Systems Analyst
Lead Business Systems Analyst

McKesson • Montreal (administrative region)

On-site
CAD 162,000 - 270,000
Flex & Connect
In-office 2 days/week
Sr. Business Systems Analyst
Sr. Business Systems Analyst

McKesson • Montreal (administrative region)

Hybrid
CAD 116,000 - 194,000
Flex & Connect
Two days in office
Competitive compensation
Analyste principal(e) des systèmes d’affaires / Sr. Business Systems Analyst
Analyste principal(e) des systèmes d’affaires / Sr. Business Systems Analyst

McKesson • Canada

Hybrid
CAD 130,000 - 217,000
Analyste des systèmes d’affaires, Solutions de transport/Business Systems Analyst, Transportation Solutions
Analyste des systèmes d’affaires, Solutions de transport/Business Systems Analyst, Transportation Solutions

McKesson • Montreal (administrative region)

Hybrid
CAD 110,000 - 183,000
Gestionnaire principal(e), IA, Innovation et Opérations / Sr. Manager, AI, Innovation & Operation
Gestionnaire principal(e), IA, Innovation et Opérations / Sr. Manager, AI, Innovation & Operation

McKesson • Montreal (administrative region)

Hybrid
CAD 140,000 - 190,000
Directeur, Planification de la chaîne d’approvisionnement et transport/Director, Supply Chain Planning & Transportation
Directeur, Planification de la chaîne d’approvisionnement et transport/Director, Supply Chain Planning & Transportation

McKesson • Montreal (administrative region)

On-site
CAD 121,000 - 201,000
Flex & Connect
Total Rewards
Spécialiste principal(e), Stratégie d’affaires/ Senior Specialist, Business Strategy
Spécialiste principal(e), Stratégie d’affaires/ Senior Specialist, Business Strategy

McKesson • Montreal (administrative region)

On-site
CAD 110,000 - 170,000