Senior Site Reliability Engineer (Observability & Analytics, Platform Infra)

Elastic

Ottawa

Hybrid

CAD 120,000 - 170,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Health coverage for you and family
Flexible location and schedule
Generous vacation days
Parental leave 16+ weeks
Volunteer hours 40 per year
Charitable matching up to USD 1500

Job summary

Elastic is seeking a Senior SRE/Platform Engineer to own end-to-end delivery of complex projects on the observability platform. You will help harden Elastic Cloud infrastructure as code and contribute to reliable, secure production systems across multiple regions.

You will mentor engineers, review production changes, and continuously improve runbooks, docs, and operational processes to reduce on-call load while supporting 24/7 reliability.

Qualifications

  • 5+ years in SRE, platform engineering, or infra.
  • Strong Python skills; Go is a plus.
  • Experience carrying 24/7 on-call and writing RCAs.
  • Proficient with Terraform and large multi-workspace configs.
  • Deep Linux knowledge and production container ops.

Responsibilities

  • Own end-to-end delivery of moderate-to-high complexity projects with minimal direction.
  • Run and harden Elastic Cloud infrastructure as code (Terraform, Python, Go).
  • Review code and designs, acting as trusted reviewer for production changes.
  • Mentor less experienced engineers and surface risks and improvements.
  • Improve runbooks, docs, and operational processes to reduce on-call load.

Skills

SRE/Platform
Python
Go
Terraform
Linux
Kubernetes
GitOps
On-call
Security-minded
Code review

Tools

ArgoCD
Helm
Vault
Kubernetes

Job description

  • Platform Observability & Analytics runs the infrastructure that tells Elastic the truth about its own platform
  • The observability clusters show Cloud engineers how production is behaving right now, and the analytics pipelines show the business how the platform and the products get used over time
  • This role sits on the observability side. We run 200+ hosted deployments across every supported cloud region, ingesting logs, metrics and traces for all of Elastic Cloud, plus the SLA and SLO monitoring for ESS and Serverless
  • When Cloud engineering needs to know what production is doing, they’re looking at something we run
  • Owning end-to-end delivery of moderate-to-high complexity projects on the team’s roadmap, with minimal day-to-day direction
  • Operating and hardening shared Elastic Cloud infrastructure (ECH, ECE, and ECK) as Infrastructure as Code - writing and reviewing the Terraform, Python, and Go that other engineers depend on
  • Carrying a 24/7 on-call rotation: responding to incidents, driving them to resolution, and writing clear RCAs/postmortems that lead to lasting fixes rather than repeat pages
  • Reviewing others’ code and designs, and being a trusted second set of eyes on production changes to critical infrastructure
  • Mentoring less experienced engineers, and proactively raising risks, ideas, and improvements in team discussions
  • Improving runbooks, documentation, and operational processes so the on-call load gets lighter over time
Benefits
  • Toast to your health: Fully paid health coverage for you and your family, in many locations.
  • Craft your calendar: Flexible location and schedule for most roles.
  • Create space for you: Distributed by design workforce, plus generous number of vacation days each year.
  • Embrace parenthood: Minimum of 16 weeks of parental leave, plus generous family formation benefits.
  • Give back your time: 40 hours each year to use toward volunteering with organizations and causes you’re passionate about.
  • Amplify your impact: Double your charitable giving — we match donations up to $1500 USD (or local currency equivalent).
  • A track record of consistently delivering end-to-end projects of moderate-to-high complexity with minimal oversight, and being a valuable code/design reviewer for your team5+ years of SRE, platform engineering, or infrastructure engineering experience
  • Strong software engineering fundamentals in Python; comfort with Go is a plus
  • Comfort working across time zones, in both real-time and asynchronous contexts
  • Comfort thinking about the security implications of the infrastructure you build, not just its reliability - you don't need to be a security specialist, but you default to a security-conscious mindset
  • Experience carrying a 24/7 on-call rotation, resolving incidents under pressure, and writing RCAs that hold up under review
  • Clear written and verbal communication - you document what you build and can explain it to both engineers and non-engineers
  • A pattern of mentoring less experienced engineers and speaking up with ideas and concerns in team discussions
  • Proficiency with Terraform; comfortable owning large, multi-workspace configurations in a team setting
  • Deep Linux systems knowledge and experience operating containerized workloads in production
  • Experience with the Elastic Stack (Elasticsearch, Logstash, Beats, Kibana) in production
  • Experience with GitOps-style deployment tooling (ArgoCD, Helm) or policy-as-code frameworks (e.g., Kyverno) on Kubernetes
  • Experience with secrets management (Vault) or access-control/bastion tooling (Teleport)
  • Experience with configuration management tools (e.g., Puppet, Ansible) at fleet scale
  • Exposure to FedRAMP, GovCloud, or other regulated/compliance-driven infrastructure
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

LanceSoft, Inc. • Montreal (administrative region)

On-site
CAD 110,000 - 140,000
Senior Site Reliability Engineer (Observability & Analytics) – Platform Infra
Senior Site Reliability Engineer (Observability & Analytics) – Platform Infra

Engg • Canada

On-site
CAD 138,000 - 186,000
Health coverage
Flexible locations and schedules
Generous vacation days
+1
Senior Infrastructure SRE
Senior Infrastructure SRE

PointClickCare • Mississauga

On-site
CAD 110,000 - 150,000
Senior DevOps Engineer
Senior DevOps Engineer

MarkiTech.AI • Toronto

On-site
CAD 140,000 - 180,000
Senior Site Reliability Engineer (Observability & Analytics) – Platform Infra
Senior Site Reliability Engineer (Observability & Analytics) – Platform Infra

Elasticsearch B.V. • Canada

Hybrid
CAD 138,000 - 186,000
Competitive pay
Health coverage
Flexible schedules
+4
Senior Platform SRE: Observability, Infra & On-Call
Senior Platform SRE: Observability, Infra & On-Call

Elastic • Ottawa

Hybrid
CAD 120,000 - 170,000
Health coverage for you and family
Flexible location and schedule
Generous vacation days
+3
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

twentysix • Vancouver

On-site
CAD 90,000 - 130,000
Senior DevOps Engineer
Senior DevOps Engineer

MarkiTech • Toronto

On-site
CAD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

PagerDuty • Toronto

Hybrid
CAD 85,000 - 120,000
Volunteer time off
Health insurance
Wellness days
+4
Senior Site Reliability Engineer
Senior Site Reliability Engineer

iManage • Toronto

On-site
CAD 90,000 - 120,000
Market-competitive salary
Annual performance-based bonus
Comprehensive Health, Vision, Dental, and Life insurance
+4