Lead SRE for Imunify Reliability Platform (Remote)

Jobgether

Canada

Remote

CAD 170,000 - 230,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Fully remote work
Vacation days: 24 paid per year
Holidays: 10 paid national holidays
Unlimited sick leave
Private medical insurance
Co-working space reimbursement
Gym reimbursement
Professional development opportunities
Patent reward program
Remote-first, asynchronous environment

Job summary

Jobgether is partnering to fill a Lead Site Reliability Engineer role for Imunify Reliability Platform based in Canada. The role focuses on defining healthy metrics, building telemetry, and establishing reliable practices across cloud services and customer-hosted agents.

You will drive SLIs/SLOs, error budgets, and governance for an observability platform; work with multiple engineering squads; and shape an async, remote-first reliability culture across time zones.

Qualifications

  • Substantial production engineering or SRE experience with SLO framework ownership.
  • Strong Python skills; read/modify Go or Rust for instrumentation.
  • Experience with time-series telemetry at scale and high-cardinality data stores.
  • Experience debugging distributed systems beyond Kubernetes-centric scope.
  • Experience with production-scale configuration management and CI/CD tooling.
  • Understanding telemetry for non-scrapable, customer-managed infra, including privacy.
  • Excellent written/asynchronous communication to align teams around health metrics.
  • Strong engineering judgment on observability, reliability, and ownership.
  • Security product familiarity is a strong advantage (WAF/EDR/AV).
  • Familiarity with SOC 2, ISO 27001, NIST frameworks; OpenTelemetry/eBPF advantageous.

Responsibilities

  • Define and establish meaningful SLIs for ~70 components with squad leads.
  • Develop a reliability taxonomy covering availability, latency, and telemetry health.
  • Ensure indicators are independently measurable and tamper-resistant.
  • Design and build telemetry pipelines for customer-hosted agents and cloud services.
  • Extend instrumentation across Python, Go, and Rust components.
  • Consolidate dashboards and reporting into a lean observability platform.
  • Implement SLO-driven alerting with multi-window burn-rate and clear classifications.
  • Assign ownership and runbooks to every production alert.
  • Strengthen incident response with blameless postmortems and follow-through.
  • Coach teams to own operational responsibilities and on-call duties.
  • Deliver measurable reliability outcomes and reduced degradation detection time.

Skills

Python
Go
Rust
Prometheus/OpenMetrics
Grafana
Alertmanager
SRE leadership
Telemetry
OpenTelemetry
eBPF
Sentry
Kubernetes
CI/CD tooling
Ansible
GitLab CI/Jenkins
Time-series data
Security product awareness

Tools

ClickHouse
Prometheus
Grafana
Terraform
CI/CD pipelines

Job description

Jobgether is partnering to fill a Lead Site Reliability Engineer role for Imunify Reliability Platform based in Canada. The role focuses on defining healthy metrics, building telemetry, and establishing reliable practices across cloud services and customer-hosted agents.

You will drive SLIs/SLOs, error budgets, and governance for an observability platform; work with multiple engineering squads; and shape an async, remote-first reliability culture across time zones.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer - Imunify Reliability Platform
Lead Site Reliability Engineer - Imunify Reliability Platform

Jobgether • Canada

Remote
CAD 170,000 - 230,000
Fully remote work
Vacation days: 24 paid per year
Holidays: 10 paid national holidays
+7
Senior Cloud SRE & Reliability Leader
Senior Cloud SRE & Reliability Leader

Jobgether • Canada

Hybrid
CAD 151,000 - 200,000
Health benefits
Equity stock options
Remote-friendly environment
+2
Founding SRE: Cloud Reliability & Platform Lead
Founding SRE: Cloud Reliability & Platform Lead

Katalyze AI, Inc. • Toronto

On-site
CAD 120,000 - 180,000
Manager, Site Reliability Engineering (SRE)
Manager, Site Reliability Engineering (SRE)

Quantum Technology Recruiting Inc. (QTR) • Toronto

On-site
CAD 155,000 - 165,000
Senior Site Reliability Engineer - Global Infra & CI/CD Impact
Senior Site Reliability Engineer - Global Infra & CI/CD Impact

CloudFactory Limited • Canada

Hybrid
CAD 120,000 - 160,000
Hybrid Working Model
Comprehensive medical cover
Group life insurance
+3
Remote Senior Site Reliability Engineer – AWS
Remote Senior Site Reliability Engineer – AWS

Linux User Group • Saskatoon

On-site
CAD 110,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Thinkific • Canada

Remote
CAD 111,000 - 167,000
Fair and transparent pay
Inclusive work culture
Remote work flexibility
Site Reliability Engineer
Site Reliability Engineer

TELUS Digital • Canada

Remote
CAD 90,000 - 120,000
Senior SRE Leader: Scale Reliability & Observability
Senior SRE Leader: Scale Reliability & Observability

Rootly • Toronto

On-site
CAD 120,000 - 180,000
Competitive compensation
Comprehensive medical coverage
3 weeks of vacation
+2
Senior SRE & DevOps Engineer - Hybrid (Mississauga)
Senior SRE & DevOps Engineer - Hybrid (Mississauga)

RE Partners • Mississauga

Hybrid
CAD 110,000 - 170,000
Hybrid work model