Cloud SRE Engineer: Observability, Automation & Reliability

AXON-Networks

Toronto

Remote

CAD 166,000 - 229,000

Part time

12 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

AXON Networks is seeking a Site Reliability Engineer to improve cloud service reliability, observed metrics and incident response. You will define SLIs/SLOs, automate NOC tasks and lead during incidents, partnering with Support, Operations, DevOps and firmware teams.

You will own dashboards, capacity planning and runbooks, driving continuous improvement across cloud, Kubernetes and device-management paths in a contract setting. Strong automation and observability skills are required.

Qualifications

  • 5+ years of experience in site reliability engineering, production engineering, DevOps, cloud infrastructure, systems engineering or a closely related role.
  • Strong software or automation skills in Python, Go, Java, Bash or a comparable language, with experience producing maintainable operational code.
  • Hands-on experience operating distributed production systems in a public cloud environment and troubleshooting across application, infrastructure, network and device-integration layers.
  • Experience with Google Cloud Platform, Oracle Cloud Infrastructure and production Kubernetes environments.
  • Experience with infrastructure as code and delivery tooling such as Terraform, Helm, Git-based CI/CD and policy-as-code.
  • Strong Linux, containers and Kubernetes fundamentals, including deployment behavior, resource management, networking and failure diagnosis.
  • Strong troubleshooting & debugging skills in Kubernetes platforms.
  • Experience with modern observability practices and tools across metrics, logs, traces, alerting, dashboards and synthetic monitoring.
  • Familiarity with Prometheus, Grafana, OpenTelemetry or equivalent observability ecosystems.
  • Familiarity with Apache Pulsar or similar distributed messaging and streaming platforms handling requests from millions of devices.
  • Experience participating in an on-call rotation and responding effectively to high-severity, customer-impacting production incidents.
  • Working knowledge of SLOs, error budgets, capacity planning, resilience engineering, change safety and blameless incident learning.
  • Strong networking knowledge, including TCP/IP, DNS, DHCP, TLS, routing, NAT, load balancing and systematic packet-or-session-level troubleshooting.
  • Clear communication, disciplined documentation and the ability to collaborate across NOC, cloud, DevOps, firmware and service-provider teams.
  • Bachelor’s degree in computer science, engineering or equivalent practical experience.

Responsibilities

  • Own reliability outcomes for assigned cloud services.
  • Improve observability, capacity, resilience and recovery.
  • Define and operationalize service-level indicators, service-level objectives and actionable alerting.
  • Automate repetitive NOC work and create safe, testable mechanisms for diagnosis, recovery, device operations and routine production changes.
  • Lead technically during incidents, drive evidence-based learning and ensure high-value corrective actions are completed.
  • Establish reliability baselines, SLIs, SLOs and error budgets for cloud services and critical device-management workflows.
  • Trace failures across the end-to-end service path: cloud APIs and microservices, Kubernetes and infrastructure, databases and messaging, internet and access-network dependencies, device-management protocols and the devices.
  • Identify fleet-wide and customer-specific failure patterns involving device reachability, session stability, configuration drift, command latency, telemetry gaps, firmware behavior and cloud capacity.
  • Maintain NOC dashboards for service health, device reachability, provisioning success, command and telemetry performance, firmware adoption and customer impact.
  • Participate in the NOC production on-call rotation and serve as a technical incident lead or senior troubleshooter when appropriate.
  • Diagnose complex failures across applications, cloud infrastructure, Kubernetes, APIs, networking, DNS/TLS, databases, messaging platforms, device-management sessions and CPE behavior.
  • Coordinate evidence gathering and technical escalation with service-provider customers, Engineering, firmware, DevOps and vendors while maintaining clear mitigation, recovery and handoff.
  • Lead or contribute to post-incident reviews; convert recurring device, platform and process failures into prioritized and measurable corrective actions.
  • Develop production-grade software, scripts and workflows for diagnosis, remediation, deployment safety, fleet analysis, scaling, maintenance and recovery.
  • Improve CI/CD and GitOps practices for operational software and infrastructure, including automated testing, release validation, progressive delivery and rollback readiness.
  • Manage or contribute to infrastructure as code, configuration as code and reusable self-service patterns for cloud and NOC operations.
  • Measure NOC toil and partner with Automation & Tools Engineers to prioritize durable platform capabilities instead of fragmented one-off scripts.
  • Develop capacity models for service-provider growth, managed-device populations, telemetry volume, messaging throughput, API demand and rollout events.
  • Create and maintain runbooks, troubleshooting decision trees, service maps, device and cloud dependency records, known-error guidance and operational knowledge.
  • Coach NOC and Support personnel on diagnosis, safe mitigation, evidence capture and escalation across cloud, network and CPE layers.
  • Build self-service diagnostic views and tools that help the NOC determine scope, affected customers, device cohorts, likely fault domain and next action.
  • Share reliability insights with Engineering and Product and contribute to reliability reviews, operational-readiness reviews and continuous-improvement priorities.

Skills

Python
Go
Java
Bash
Cloud experience
Kubernetes
Linux
Observability
Incident management
Networking

Education

Bachelor's degree in computer science, engineering or equivalent practical experience

Tools

Terraform
Helm
Git CI/CD
OpenTelemetry
Prometheus
Grafana
Apache Pulsar/Kafka

Job description

AXON Networks is seeking a Site Reliability Engineer to improve cloud service reliability, observed metrics and incident response. You will define SLIs/SLOs, automate NOC tasks and lead during incidents, partnering with Support, Operations, DevOps and firmware teams.

You will own dashboards, capacity planning and runbooks, driving continuous improvement across cloud, Kubernetes and device-management paths in a contract setting. Strong automation and observability skills are required.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliable Engineer (Canada - Remote)
Site Reliable Engineer (Canada - Remote)

AXON-Networks • Toronto

Remote
CAD 166,000 - 229,000
Senior SRE: Cloud, Automation & Observability
Senior SRE: Cloud, Automation & Observability

RXinsider LTD. • Montreal (administrative region)

Hybrid
CAD 100,000 - 150,000
Site Reliability Engineer (SRE) – Observability
Site Reliability Engineer (SRE) – Observability

Astra-North Infoteck Inc. ~ Conquering today’s challenges, achieving tomorrow’s vision! • Toronto

Hybrid
CAD 75,000 - 95,000
Network Operations Manager (Canada - Remote)
Network Operations Manager (Canada - Remote)

AXON-Networks • Toronto

Hybrid
CAD 120,000 - 180,000
Lead Site Reliability & Observability Engineer
Lead Site Reliability & Observability Engineer

CRC Group • New Brunswick

On-site
CAD 120,000 - 160,000
Medical insurance
401(k) plan with company match
Paid time off
Founding SRE: Cloud Reliability & Platform Lead
Founding SRE: Cloud Reliability & Platform Lead

Katalyze AI, Inc. • Toronto

On-site
CAD 120,000 - 180,000
Lead Azure SRE for Cloud-Native Platforms
Lead Azure SRE for Cloud-Native Platforms

S I Systems • Toronto

On-site
CAD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Mantu • Montreal (administrative region)

On-site
CAD 90,000 - 130,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

twentysix • Vancouver

On-site
CAD 90,000 - 130,000
Platform Engineer
Platform Engineer

LanceSoft, Inc. • Montreal (administrative region)

On-site
CAD 80,000 - 120,000