Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Apptoza Inc. in Montreal seeks an experienced Observability/SRE/DevOps engineer to design and implement observability-as-code with Terraform, deploying monitoring pipelines across distributed systems.
You will instrument Node.js and .NET microservices and lead incident response to ensure reliability and rapid root-cause analysis. Expertise with Dynatrace, ELK, Splunk, Pager Duty, AKS, Terraform, and Azure is required, along with strong collaboration across teams to meet SLIs/SLOs and drive
What will you do? Design and implement observability-as-code solutions using Terraform to deploy monitoring pipelines| dashboards| and alerting strategies across distributed systems. Drive observability improvements leveraging industry-leading tools (Dynatrace| ELK| Splunk| Pager Duty) to achieve real-time performance insights and comprehensive system visibility. Instrument applications for end-to-end observability implementing distributed tracing| metrics collection| and log aggregation across Node.js and .NET microservices and event-driven architectures. Troubleshoot complex incidents in production environments| diagnosing root causes across multiple service layers| databases| caches| and APIs under load using SLI/SLO frameworks. Investigate and resolve Azure Kubernetes Service (AKS) infrastructure| ensuring reliability and scalability of containerized workloads with deep proficiency in Terraform and Azure managed services (SQL MI| Redis | Functions| Event Grid).Translate business requirements into observable| resilient systems that meet defined SLIs/SLOs and drive reliability improvements. Automate operational tasks to reduce toil and improve system resilience through infrastructure-as-code and CI/CD best practices. Lead incident response and remediation for mission-critical systems| conducting blameless postmortems and building resilience through chaos engineering and tabletop exercises. Collaborate cross-functionally with development| platform| and business teams to improve service availability| scalability| and operational excellence. What do you need to succeed?Must-have:8+ years hands-on experience in observability| SRE| or DevOps roles with proven expertise across infrastructure and application-level reliability .Deep expertise in observability tooling: Dynatrace| ELK| Splunk| and Pager Duty; demonstrated understanding of observability principles (instrumentation| correlation IDs| SLI/SLO frameworks).Advanced proficiency with Azure Kubernetes Service (AKS)| Terraform| and Azure managed services (SQL MI| Redis| Functions| Event Grid); proven ability to design and implement infrastructure-as-code solutions. Strong hands-on experience instrumenting applications for comprehensive observability: distributed tracing| metrics collection| and log aggregation across Node.js and .NET applications in microservices and event-driven architectures. Proven troubleshooting expertise in distributed systems diagnosing root causes across multiple service layers| databases| caches| and APIs in production environments. Excellent incident management skills: hands-on experience with Pager Duty and ServiceNow; ability to resolve high-severity incidents rapidly and conduct effective root cause analysis. Knowledge of incident| problem| and change management processes| including SRE principles| blameless postmortems| and chaos engineering practices. Exceptional communication and leadership