Senior Site Reliability Engineer, SRE

Jobtailor

Toronto

On-site

CAD 120,000 - 180,000

Full time

3 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Jobtailor in Toronto is seeking a Senior Site Reliability Engineer to design, build, and maintain scalable distributed systems, guiding architectural decisions with reliability at the forefront. You will partner with development teams, automate operations and CI/CD pipelines, and develop internal observability tools while joining a 24/7 on-call rotation as incident commander during disruptions.

Candidates should have 5+ years in SRE/DevOps, strong Rust/Go/Python/Typescript skills, deep

Qualifications

  • Bachelor's degree in CS or related field or equivalent practical experience.
  • 5+ years in SRE/DevOps/Software Engineering focusing on infra and operations.
  • Proficient in Rust, Go, Python, or Typescript and able to write production-grade code.
  • Deep knowledge of AWS services and Kubernetes in production.

Responsibilities

  • Design, build, and maintain scalable distributed systems.
  • Partner with development teams as a reliability consultant and influence architectural decisions.
  • Write code to automate operational tasks and CI/CD pipelines.
  • Build internal tools, libraries, and frameworks for self-service observability.
  • Participate in a 24/7 on-call rotation and act as incident commander during critical disruptions.
  • Conduct blameless root cause analyses and implement corrective actions.
  • Monitor, measure, and optimize system performance, latency, and capacity.
  • Forecast capacity needs using usage patterns and historical data.
  • Build and integrate AIOps solutions, including automated responses and self-healing systems.
  • Develop and document runbooks for the observability knowledge base.
  • Analyze telemetry data, build predictive capacity models, and identify bottlenecks and failure modes.

Skills

Rust
Go
Python
Typescript
Automation
Observability
Incident Management

Education

Bachelor's degree in Computer Science

Tools

Kubernetes
AWS
EKS
CloudWatch
Route 53
ALBs/NLBs

Job description

  • Design, build, and maintain scalable, highly available, and fault-tolerant distributed systems
  • Partner with development teams as a reliability consultant and influence architectural decisions
  • Write code to automate operational tasks and CI/CD pipelines
  • Build internal tools, libraries, and frameworks for self-service observability
  • Participate in a 24/7 on-call rotation and act as incident commander during critical disruptions
  • Conduct blameless root cause analyses and implement corrective actions
  • Monitor, measure, and optimize system performance, latency, and capacity
  • Forecast capacity needs using usage patterns and historical data
  • Build and integrate AIOps solutions, including automated responses and self-healing systems
  • Use AI-assisted coding tools such as Claude Code and Cursor
  • Develop and document runbooks and procedural guides for the observability knowledge base
  • Analyze telemetry data, build predictive capacity models, and identify bottlenecks and failure modes
Requirements
  • Bachelor's degree in Computer Science, a related technical field, or equivalent practical experience
  • 5+ years of professional experience in a Site Reliability Engineering, DevOps, or Software Engineering role focused on infrastructure and operations
  • Strong programming proficiency in one or more high-level languages such as Rust, Go, Python, or Typescript
  • Comfortable writing, testing, and deploying production-grade code
  • Deep knowledge of AWS services, especially networking, IAM, EKS, ALBs/NLBs, Route 53, and CloudWatch
  • Proven experience with Kubernetes in production, including service exposure, networking, and availability engineering
  • Solid understanding of Linux/Unix operating systems, TCP/IP, DNS, HTTP, and modern distributed systems architecture
Core Competencies

Demonstrates expertise in designing and maintaining scalable distributed systems, with a strong focus on automation, observability, and incident management. Proficient in programming and cloud services, particularly in AWS and Kubernetes, to optimize system performance and reliability.

Highest-signal resume keywords
  • Site Reliability Engineering
  • AWS Services
  • Kubernetes
  • Programming Proficiency
  • Automation
Hard Skills
  • Rust
  • Go
  • Python
  • Typescript
  • Linux/Unix
  • TCP/IP
  • DNS
  • HTTP
  • Distributed Systems Architecture
  • CI/CD
Soft Skills
  • Incident Management
  • Root Cause Analysis
  • Collaboration
Certifications & Qualifications
  • Bachelor's Degree in Computer Science
Industry Keywords
  • Infrastructure
  • Operations
  • Observability
  • Capacity Forecasting
  • Self-Healing Systems
Tools & Technologies
  • AIOps Solutions
  • Claude Code
  • Cursor
  • CloudWatch
  • EKS
  • ALBs/NLBs
  • Route 53
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Infrastructure SRE
Senior Infrastructure SRE

Jobtailor • Mississauga

On-site
CAD 110,000 - 160,000
Senior Software Engineer, Infrastructure
Senior Software Engineer, Infrastructure

Jobtailor • Toronto

On-site
CAD 110,000 - 160,000
Senior Infrastructure SRE
Senior Infrastructure SRE

PointClickCare • Mississauga

Hybrid
CAD 110,000 - 150,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

twentysix • Vancouver

On-site
CAD 90,000 - 130,000
Staff Developer - Incident Command
Staff Developer - Incident Command

IBM • Toronto

On-site
CAD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Mantu • Montreal (administrative region)

On-site
CAD 90,000 - 130,000
Staff Developer - Incident Command
Staff Developer - Incident Command

IBM • Markham

On-site
CAD 120,000 - 180,000
Staff Developer - Incident Command
Staff Developer - Incident Command

IBM • Ottawa

Hybrid
CAD 120,000 - 160,000
Staff Developer - Incident Command
Staff Developer - Incident Command

IBM • Bedford

On-site
CAD 140,000 - 200,000
Staff Developer - Incident Command
Staff Developer - Incident Command

IBM • Edmonton

On-site
CAD 120,000 - 180,000