Senior Observability Engineer

Astra-North Infoteck Inc. ~ Conquering today’s challenges, achieving tomorrow’s vision!

Montreal (administrative region)

Hybrid

CAD 120,000 - 160,000

Full time

35 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Astra-North Infoteck Inc. in Montreal, QC, is seeking a Senior Observability Engineer to design and implement observability pipelines, dashboards, and alerting across distributed systems. You will drive real-time performance insights using Dynatrace, ELK, and Splunk, and lead incident response in a hybrid workplace.

The role requires extensive hands-on expertise with AKS, Terraform, and Azure managed services, plus strong collaboration with cross-functional teams to meet SLIs/SLOs.

Qualifications

  • 8+ years hands-on experience in observability, SRE, or DevOps roles.
  • Deep expertise in Dynatrace, ELK, Splunk, and PagerDuty; SLI/SLO knowledge.
  • Advanced proficiency with AKS, Terraform, and Azure managed services.
  • Strong hands-on experience instrumenting apps for observability in Node.js and .NET.
  • Proven troubleshooting of distributed systems in production.
  • Excellent incident management with PagerDuty and ServiceNow.
  • Knowledge of incident, problem, and change management, chaos engineering.
  • Exceptional communication and leadership.

Responsibilities

  • Design and implement observability-as-code solutions using Terraform to deploy monitoring pipelines, dashboards, and alerting strategies across distributed systems.
  • Drive observability improvements leveraging Dynatrace, ELK, Splunk, and PagerDuty for real-time performance insights.
  • Instrument applications for end-to-end observability across Node.js and .NET microservices.
  • Troubleshoot incidents in production across service layers, databases, caches, and APIs under load.
  • Investigate AKS infrastructure, ensuring reliability and scalability of containerized workloads.
  • Translate business requirements into observable, resilient systems meeting SLIs/SLOs.
  • Automate operational tasks via infrastructure-as-code and CI/CD.
  • Lead incident response and postmortems, building resilience through chaos engineering.
  • Collaborate with development and platform teams to improve availability and scalability.

Skills

Observability
SRE/DevOps
Incident management
Distributed systems troubleshooting
Leadership

Tools

AKS
Terraform
Dynatrace
ELK
Splunk
PagerDuty
ServiceNow
Azure SQL MI
Redis
Azure Functions
Event Grid

Job description

Job Description
Senior Observability Engineer

Key Skills: AKS + Terraform, Dynatrace/Splunk/ELK, PagerDuty/ServiceNow

Montreal, QC - Hybrid (2-4 Days WFO

Role Descriptions:
What will you do?
  • Design and implement observability-as-code solutions using Terraform to deploy monitoring pipelines, dashboards, and alerting strategies across distributed systems.
  • Drive observability improvements leveraging industry-leading tools (Dynatrace, ELK, Splunk, PagerDuty) to achieve real-time performance insights and comprehensive system visibility.
  • Instrument applications for end-to-end observability implementing distributed tracing, metrics collection, and log aggregation across Node.js and .NET microservices and event-driven architectures.
  • Troubleshoot complex incidents in production environments, diagnosing root causes across multiple service layers, databases, caches, and APIs under load using SLI/SLO frameworks.
  • Investigate and resolve Azure Kubernetes Service (AKS) infrastructure, ensuring reliability and scalability of containerized workloads with deep proficiency in Terraform and Azure managed services (SQL MI, Redis, Functions, Event Grid).
  • Translate business requirements into observable, resilient systems that meet defined SLIs/SLOs and drive reliability improvements.
  • Automate operational tasks to reduce toil and improve system resilience through infrastructure-as-code and CI/CD best practices.
  • Lead incident response and remediation for mission-critical systems, conducting blameless postmortems and building resilience through chaos engineering and tabletop exercises.
  • Collaborate cross-functionally with development, platform, and business teams to improve service availability, scalability, and operational excellence.
What do you need to succeed?
Must-have:
  • 8+ years hands-on experience in observability, SRE, or DevOps roles with proven expertise across infrastructure and application-level reliability.
  • Deep expertise in observability tooling: Dynatrace, ELK, Splunk, and PagerDuty; demonstrated understanding of observability principles (instrumentation, correlation IDs, SLI/SLO frameworks).
  • Advanced proficiency with Azure Kubernetes Service (AKS), Terraform, and Azure managed services (SQL MI, Redis, Functions, Event Grid); proven ability to design and implement infrastructure-as-code solutions.
  • Strong hands-on experience instrumenting applications for comprehensive observability: distributed tracing, metrics collection, and log aggregation across Node.js and .NET applications in microservices and event-driven architectures.
  • Proven troubleshooting expertise in distributed systems diagnosing root causes across multiple service layers, databases, caches, and APIs in production environments.
  • Excellent incident management skills: hands‑on experience with PagerDuty and ServiceNow; ability to resolve high‑severity incidents rapidly and conduct effective root cause analysis.
  • Knowledge of incident, problem, and change management processes, including SRE principles, blameless postmortems, and chaos engineering practices.
  • Exceptional communication and leadership.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr Support Engineer
Sr Support Engineer

Apptoza Inc. • Montreal (administrative region)

On-site
CAD 120,000 - 160,000
Senior SRE — Observability & Reliability
Senior SRE — Observability & Reliability

Apptoza Inc. • Montreal (administrative region)

On-site
CAD 120,000 - 160,000
Senior Site Reliability Engineer, SRE
Senior Site Reliability Engineer, SRE

Jobtailor • Toronto

On-site
CAD 120,000 - 180,000
Platform Engineer
Platform Engineer

LanceSoft, Inc. • Montreal (administrative region)

On-site
CAD 80,000 - 120,000
Senior Infrastructure SRE
Senior Infrastructure SRE

PointClickCare • Mississauga

Hybrid
CAD 110,000 - 150,000
Senior Support Engineer
Senior Support Engineer

Tata Consultancy Services • Toronto

On-site
CAD 100,000 - 120,000
Site Reliability Engineer (SRE) – Observability
Site Reliability Engineer (SRE) – Observability

Astra-North Infoteck Inc. ~ Conquering today’s challenges, achieving tomorrow’s vision! • Toronto

Hybrid
CAD 75,000 - 95,000
Manager, Site Reliability Engineering (SRE)
Manager, Site Reliability Engineering (SRE)

Quantum Technology Recruiting Inc. (QTR) • Toronto

On-site
CAD 155,000 - 165,000
Site Reliability Expert
Site Reliability Expert

NEPSE Trading • Canada

On-site
CAD 100,000 - 150,000
Comprehensive insurance plan (Gold/SIl
Virtual healthcare via Sun Life
Employee and Family Assistance Program
+2
GCP Observability Engineer
GCP Observability Engineer

ALLTECH CONSULTING SVC INC • Quebec

On-site
CAD 85,000 - 115,000