Associate Director, Observability and Service Reliability

Kyndryl Inc.

Toronto

Hybrid

CAD 130,000 - 190,000

Full time

11 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Be Well programs
Hybrid-friendly culture
Extensive learning opportunities

Job summary

Kyndryl Inc. seeks a seasoned Enterprise Observability Leader to own the strategy for monitoring and service reliability across our enterprise tech stack.

You will drive architecture, standards, and governance, ensuring observability is measurable, scalable, and aligned with business goals. You will partner with application, infrastructure, cloud, DevOps, security, and operations teams to define indicators, improve root-cause analysis, and enable proactive remediation at scale.

Qualifications

  • Bachelor’s degree in information technology, computer science, engineering, or related field.
  • 8+ years in observability, application performance, service reliability, or enterprise tech operations.
  • 3+ years of technical leadership or people leadership experience.
  • Experience designing and operating monitoring in complex enterprise environments.
  • Familiarity with Dynatrace and Nexthink or similar Digital Employee Experience tech.
  • Knowledge of monitoring architectures and SRE concepts.

Responsibilities

  • Lead enterprise observability strategy and standards across apps, infra, cloud, and networks.
  • Define SLIs/SLOs, availability targets, performance thresholds, and health metrics.
  • Guide teams on observability methods and tools, ensuring measurable value.
  • Collaborate with application, cloud, DevOps, and security teams to improve telemetry.
  • Drive AIOps and automation to reduce noise and accelerate incident response.

Skills

Observability
Leadership
Architecture
Analytical thinking
Communication
Stakeholder management
Cloud technologies
Incident management

Education

Bachelor’s degree in IT/CS/Engineering

Tools

Dynatrace
Nexthink
ServiceNow
OpenTelemetry
AIOps

Job description

Who We Are

At Kyndryl, we run and reimagine the mission‑critical technology systems that drive advantage for the world's leading businesses. We are at the heart of progress; with proven expertise and a continuous flow of AI‑powered insight, enabling smarter decisions, faster innovation, and a lasting competitive edge. For our people—Kyndryls—that means doing purposeful work that powers human progress. Join us and experience a flexible, supportive environment where your well‑being is prioritized and your potential can thrive.

The Role
Enterprise Observability Strategy
  • Own and mature the enterprise observability and service reliability strategy.
  • Define standards for monitoring applications, infrastructure, cloud platforms, networks, endpoints, APIs, databases, middleware, and other critical technology services.
  • Establish expectations for metrics, logs, traces, events, synthetic monitoring, real user monitoring, digital experience, service health, and business transaction visibility.
  • Create a consistent enterprise approach while allowing teams to use monitoring technologies suited to their platforms and services.
  • Identify monitoring gaps, redundant capabilities, excessive alerting, and opportunities to improve visibility.
  • Move the organization from traditional monitoring toward proactive, predictive, and automated operations.
Service Reliability Engineering
  • Establish and mature the organization’s service reliability framework.
  • Partner with technical service owners to define monitoring requirements for critical applications and services.
  • Ensure monitoring reflects the complete service, including application performance, infrastructure, dependencies, integrations, user experience, business transactions, capacity, and failure conditions.
  • Define minimum observability requirements based on service criticality and business impact.
  • Help teams establish meaningful Service Level Indicators, Service Level Objectives, availability targets, performance thresholds, and health measures.
  • Use reliability data, incidents, problem records, capacity trends, and telemetry to identify systemic weaknesses and prioritize improvements.
Technical Thought Leadership
  • Serve as the enterprise technical authority for observability, monitoring, and service reliability.
  • Provide architectural guidance to application, infrastructure, cloud, engineering, DevOps, SRE, cybersecurity, and operations teams.
  • Influence solution design so services are observable, measurable, supportable, and resilient by design.
  • Develop enterprise monitoring patterns, reference architectures, standards, and reusable capabilities.
  • Guide technical teams in selecting appropriate monitoring methods and technologies for specific platforms and use cases.
  • Evaluate emerging observability, AIOps, automation, analytics, and service reliability capabilities for measurable operational value.
Observability Platform Leadership
  • Provide strategic oversight for the enterprise observability and monitoring tool ecosystem.
  • Lead the strategy, architecture, governance, adoption, and optimization of major platforms, including Dynatrace, Nexthink, and related enterprise monitoring technologies.
  • Ensure monitoring tools operate as an integrated ecosystem rather than isolated platforms.
  • Establish standards for instrumentation, tagging, alerting, dashboards, integrations, service mapping, ownership, and data quality.
  • Partner with technical teams to maximize platform value while reducing tooling duplication and complexity.
  • Manage strategic technology and vendor relationships to ensure observability investments deliver measurable operational value.
Dynatrace Platform Strategy
  • Provide strategic leadership for enterprise use of Dynatrace across applications, infrastructure, cloud, and digital services.
  • Drive adoption of application performance monitoring, distributed tracing, real user monitoring, synthetic monitoring, infrastructure monitoring, logs, topology, service health, and intelligent problem detection.
  • Partner with technical owners to improve application instrumentation and ensure Dynatrace provides meaningful service visibility.
  • Use Dynatrace capabilities to improve root cause identification, dependency visibility, anomaly detection, and incident response.
Nexthink and Digital Employee Experience
  • Lead the enterprise strategy for Nexthink and Digital Employee Experience.
  • Provide visibility into endpoint health, application performance, employee technology experience, device reliability, and technology friction.
  • Partner with Digital Workplace, Service Desk, application teams, and infrastructure teams to identify and remediate issues proactively.
  • Use experience data to reduce support demand, improve employee productivity, and identify systemic technology issues.
  • Expand automation and targeted remediation to resolve employee experience issues at scale.
Event Management, AIOps, and Automation
  • Improve operational signal quality by reducing noise and increasing alert relevance.
  • Drive event correlation, enrichment, anomaly detection, automated diagnostics, and proactive remediation.
  • Integrate observability platforms with ITSM, incident management, automation, collaboration, configuration, and operational data platforms.
  • Develop capabilities that help support teams detect degradation earlier and understand business impact faster.
  • Use AI, machine learning, analytics, and automation where they improve outcomes and reduce manual effort.
Governance, Metrics, and Continuous Improvement

Establish governance for enterprise monitoring standards, tooling, integrations, data quality, licensing, and adoption.

Develop executive and operational reporting that provides insight into service health, reliability, performance, and user experience.

Use service measures to identify reliability risks, guide investment decisions, and prioritize service improvements.

Key measures may include:

  • Service availability and reliability
  • Mean Time to Detect, Mean Time to Identify, and Mean Time to Restore
  • Monitoring coverage for critical services
  • Service Level Objective attainment
  • Alert quality and noise reduction
  • Proactive issue detection and automated remediation
  • Digital Employee Experience Recurring service degradation
  • Observability maturity
  • Tool adoption and value realization
Leadership Responsibilities
  • Lead and develop observability, monitoring, reliability, and platform engineering professionals.
  • Create an engineering culture focused on proactive operations, technical excellence, automation, and measurable service outcomes.
  • Build strong partnerships with application owners, platform owners, infrastructure, cloud, network, cybersecurity, Digital Workplace, DevOps, SRE, Service Management, and business teams.
  • Influence technical teams without relying solely on direct reporting authority.
  • Create clear accountability for monitoring coverage and service reliability across the enterprise.
  • Partner with Incident, Problem, Change, Major Incident Management, and Operational Resilience teams to improve detection, recovery, and prevention.
Who You Are

Required Qualifications

  • Bachelor’s degree in Information Technology, Computer Science, Engineering, or a related discipline, or equivalent professional experience.
  • 8 or more years of experience in observability, application performance, service reliability, infrastructure, cloud, engineering, or enterprise technology operations.
  • 3 or more years of technical leadership or people leadership experience.
  • Experience designing and operating monitoring and observability capabilities in complex enterprise environments.
  • Experience with Dynatrace and familiarity with Nexthink or comparable Digital Employee Experience technologies.
  • Knowledge of observability technologies, monitoring architectures, application performance management, infrastructure monitoring, event management, logging, tracing, synthetic monitoring, and user experience monitoring.
  • Understanding of modern enterprise applications, cloud technologies, containers, APIs, networks, databases, infrastructure, and distributed architectures.
  • Experience defining service health, reliability measures, Service Level Indicators, Service Level Objectives, and operational performance standards.
  • Ability to influence senior technical leaders and translate complex technical concepts into business risk and operational outcomes.
  • Strong leadership, architecture, analytical, problem-solving, communication, and stakeholder management skills.

Preferred Qualifications

  • Advanced Dynatrace experience in a large enterprise environment.
  • Experience implementing or scaling Nexthink.
  • Experience with ServiceNow and enterprise event management platforms.
  • Experience with Site Reliability Engineering, DevOps, AIOps, automation, OpenTelemetry, cloud native monitoring, and modern observability architectures.
  • Experience establishing enterprise monitoring standards or observability reference architectures.
  • Experience within complex, global, or highly regulated enterprise environments.
Being You

The "Kyn" in Kyndryl means kinship, which represents the strong bonds we have with each other, our customers and our communities. We focus on ensuring all Kyndryls feel included and we welcome people of all cultures, backgrounds, and experiences. Even if you don’t meet every requirement, we encourage you to apply. We believe in growth, and we’re excited to see what you can bring. At Kyndryl, employee feedback has told us that our number one driver of employee engagement is belonging. That sense of belonging - being a valued, respected, trusted member of the team - is fundamental to our culture and fueling great experiences for our customers. This dedication to welcoming everyone into our company means that Kyndryl gives you the ability to thrive and contribute to our culture of empathy and shared success. That’s The Kyndryl Way.

What You Can Expect
  • Your career with us isn’t just a job—it’s an adventure with purpose.
  • We offer a dynamic, hybrid-friendly culture that supports your well-being and empowers you to grow.
  • Our Be Well programs are thoughtfully designed to support your financial, mental, physical, and social health - because we know that when you feel your best, you do your best.
  • From your very first day, you’ll dive into impactful work that powers the systems our customers rely on every day.
  • You won’t just contribute - you’ll make a difference, tackling meaningful projects that sharpen your skills and fuel your growth.
  • We’re here to champion your journey.
  • With powerful tools to chart your career path, personalized development goals aligned with your ambitions, and continuous feedback to keep you inspired and on track, you’ll have everything you need to thrive and evolve.
  • You’ll develop in-demand skills to grow your career and achieve your ambitions with access to cutting-edge learning opportunities - from certifications with Microsoft, Google, and Amazon to coaching and hands-on experiences.
  • And through it all, you’ll be part of a culture that values empathy, restless learning, and a devotion to shared success.
  • We want you to thrive here - and we’re committed to helping you do just that.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Consult Associate Partner Retail/CPG
Consult Associate Partner Retail/CPG

Kyndryl • Toronto

Hybrid
CAD 150,000 - 210,000
Be Well program
Hybrid-friendly culture
Senior Cloud Consultant
Senior Cloud Consultant

Kyndryl Inc. • Markham

On-site
CAD 161,000 - 231,000
Strategic Sales Role (Complex Seller)
Strategic Sales Role (Complex Seller)

Kyndryl • Toronto

Hybrid
CAD 90,000 - 150,000
Be Well program
Early Career Consult Program – Cybersecurity Defense Associate
Early Career Consult Program – Cybersecurity Defense Associate

Kyndryl • Pickering

Hybrid
CAD 70,000 - 110,000
None
Senior Cloud Consultant
Senior Cloud Consultant

Kyndryl • Markham

Hybrid
CAD 115,000 - 165,000
Associate Director, Data Center Management
Associate Director, Data Center Management

Kyndryl Inc. • Toronto

Hybrid
CAD 138,000 - 189,000
Be Well program
Hybrid-friendly culture
Console Operator
Console Operator

1120 ISM Information Systems Management Canada Corporation • Regina

On-site
CAD 63,000 - 76,000
Mainframe Security Architect
Mainframe Security Architect

Kyndryl Inc. • Toronto

Hybrid
CAD 194,000 - 265,000
Customer Unit Leader
Customer Unit Leader

Kyndryl • Ottawa

Hybrid
CAD 180,000 - 240,000
Be Well program
Hybrid-friendly culture
Vice President, Delivery Leadership
Vice President, Delivery Leadership

Kyndryl • Toronto

On-site
CAD 120,000 - 150,000
Employee learning programs
Access to industry certifications
Volunteering and giving platform
+1