Site Reliability Engineer (SRE)

twentysix

Vancouver

On-site

CAD 90,000 - 130,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

twentysix is seeking an SRE/Platform Operations professional in Vancouver to monitor health, respond to incidents, and improve observability. You will maintain dashboards, runbooks, and postmortems while supporting the Service Desk and collaborating with the EMEA team.

The role requires 2–3 years in SRE or a similar function, strong English, and experience with Grafana/Prometheus, Docker, and Kubernetes. On-call weekends and Agile work culture are part of the job.

Qualifications

  • Strong proficiency in English
  • 2–3 years of experience in SRE, platform operations, or a technical Service Desk role
  • Experience with monitoring and observability tools such as Grafana, Prometheus, or equivalent
  • Solid understanding of incident management processes (triaging, escalation, postmortems)
  • Experience supporting ecommerce platforms
  • Scripting skills in shell and/or Python for automation and operational tasks
  • Familiarity with containerization concepts (Docker, Kubernetes) at an operational level
  • Experience working in Agile environments and using ticketing tools (e.g. Jira)
  • Comfort working independently during early-morning shifts with minimal supervision
  • Strong documentation habits and attention to detail
  • Experience with Agile processes, testing, and code review
  • Strong experience with scripting - shell, Python, etc.
  • Excellent customer service attitude, communication skills (written and verbal), and interpersonal skills
  • Excellent analytical and problem-solving skills
  • Ability to communicate effectively with technical and non-technical stakeholders
  • Experience working in fast-paced, Agile environments, balancing priorities across multiple projects
  • Nice to have: Experience with Google Cloud Platform (GCP) or other major cloud providers (AWS, Azure)
  • Familiarity with CI/CD pipelines (GitHub Actions, GitLab CI)
  • Basic experience with Infrastructure as Code tools such as Terraform
  • Basic knowledge of AIOps concepts and their application in operational workflows
  • SRE or cloud certifications (Google Cloud, AWS, Kubernetes)

Responsibilities

  • Monitor platform health across multiple client environments using Grafana and Prometheus, or other monitoring tools
  • Respond to and triage incidents, following established runbooks and escalation paths
  • Participate in post-incident reviews and contribute to postmortem documentation
  • Support the Service Desk team with technical triaging, incident classification, and resolution
  • Maintain and improve observability dashboards, alerts, and SLI, and SLO tracking
  • Write and maintain runbooks, operational documentation, and knowledge base articles
  • Identify recurring issues and propose automation or process improvements to reduce toil
  • Participate in on-call rotation covering weekends (alternating schedule — one weekend on, one weekend off)
  • Collaborate with the EMEA team during shift overlap to ensure smooth handoffs and continuity
  • Support root cause analysis and contribute to continuous improvement initiatives

Skills

Grafana
Prometheus
SRE
Incident management
Python
Shell scripting
Docker
Kubernetes
Agile
Jira
On-call
Documentation

Tools

Grafana
Prometheus
Jira
GitHub Actions
GitLab CI
Terraform
GCP

Job description

WHAT YOU’LL DO

Monitor platform health across multiple client environments using tools like Grafana and Prometheus, or other monitoring tools


Respond to and triage incidents, following established runbooks and escalation paths


Participate in post-incident reviews and contribute to postmortem documentation


Support the Service Desk team with technical triaging, incident classification, and resolution


Maintain and improve observability dashboards, alerts, and SLI, and SLO tracking


Write and maintain runbooks, operational documentation, and knowledge base articles


Identify recurring issues and propose automation or process improvements to reduce toil


Participate in on-call rotation covering weekends (alternating schedule — one weekend on, one weekend off)


Collaborate with the EMEA team during shift overlap to ensure smooth handoffs and continuity


Support root cause analysis and contribute to continuous improvement initiatives


WHAT WE’RE LOOKING FOR:

Strong proficiency in English (written and verbal communication) is required


2–3 years of experience in SRE, platform operations, or a technical Service Desk role


Experience with monitoring and observability tools such as Grafana, Prometheus, or equivalent


Solid understanding of incident management processes (triaging, escalation, postmortems)


Experience supporting e-commerce platforms


Scripting skills in shell and/or Python for automation and operational tasks


Familiarity with containerization concepts (Docker, Kubernetes) at an operational level


Experience working in Agile environments and using ticketing tools (e.g. Jira)


Comfort working independently during early-morning shifts with minimal supervision


Strong documentation habits and attention to detail


Experience with Agile processes, testing, and code review


Strong experience with scripting - shell, Python, etc.


Excellent customer service attitude, communication skills (written and verbal), and interpersonal skills


Excellent analytical and problem-solving skills


Ability to communicate effectively with technical and non-technical stakeholders. You should feel comfortable explaining technical concepts in simple terms


Experience working in fast-paced, Agile environments, balancing priorities across multiple projects


**NICE TO HAVE:**


Experience with Google Cloud Platform (GCP) or other major cloud providers (AWS, Azure)


Familiarity with CI/CD pipelines (GitHub Actions, GitLab CI)


Basic experience with Infrastructure as Code tools such as Terraform


Basic knowledge of AIOps concepts and their application in operational workflows


SRE or cloud certifications (Google Cloud, AWS, Kubernetes)

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Gemini Solutions Pvt Ltd • Toronto

On-site
CAD 120,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

Mantu • Montreal (administrative region)

On-site
CAD 90,000 - 130,000
Level 3 Support and SRE
Level 3 Support and SRE

ALLTECH CONSULTING SVC INC • Quebec

On-site
CAD 75,000 - 95,000
Senior Site Reliability Engineer (SRE) – Kubernetes
Senior Site Reliability Engineer (SRE) – Kubernetes

Software Mind Americas • Montreal (administrative region)

On-site
CAD 110,000 - 170,000
Competitive salary
Laptop provided
Professional development
+2
Site Reliability Engineer
Site Reliability Engineer

ALLTECH CONSULTING SVC INC • Quebec

On-site
CAD 90,000 - 130,000
Manager, Site Reliability Engineering (SRE)
Manager, Site Reliability Engineering (SRE)

Quantum Technology Recruiting Inc. (QTR) • Toronto

On-site
CAD 155,000 - 165,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

iManage • Toronto

Hybrid
CAD 90,000 - 120,000
Market-competitive salary
Annual performance-based bonus
Comprehensive Health, Vision, Dental, and Life insurance
+4
Site Reliability Engineer (Linux / Cloud Infrastructure)
Site Reliability Engineer (Linux / Cloud Infrastructure)

Atlantis IT Group • Montreal

On-site
CAD 80,000 - 100,000
Site Reliability Engineer
Site Reliability Engineer

Compunnel, Inc. • Montreal (administrative region)

Hybrid
CAD 90,000 - 130,000
Canada- Jr Software Engineer (SRE)
Canada- Jr Software Engineer (SRE)

Embedded Shishya • Mississauga

Hybrid
CAD 60,000 - 85,000