Senior SRE – AI-Driven Cloud & Reliability Leader

Cover Genius

Vancouver

Hybrid

CAD 115,000 - 145,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Cover Genius is building an AI-first infrastructure platform and seeks a Site Reliability Engineer to lead reliability initiatives across multiple teams. You will shape system design, tooling, and process to scale production systems while reducing operational risk.

You will apply cloud expertise (AWS/GCP), implement observability standards, and mentor engineers on best practices. The role blends hands-on coding with strategic reliability ownership in a hybrid work environment.

Qualifications

  • 3+ years in SRE, Platform Engineering, DevOps or related roles.
  • Experience applying SRE principles to lead cross-team projects.
  • Experience with observability tools (Datadog, Elasticsearch, Prometheus, Grafana).
  • Experience with cloud native tech: Docker, Kubernetes; Terraform.
  • Scripting in Bash; Python or Go.
  • AI-driven development environments and using AI in workflows.
  • Experience on Linux, networking, distributed systems.
  • Deep AWS and/or GCP expertise.
  • Bachelor's in CS/Engineering; postgraduate degrees desirable.

Responsibilities

  • Analyze, design, and implement reliability and infrastructure improvements across multi-team projects.
  • Apply AWS and GCP expertise to architect reliable, highly-available cloud infra.
  • Develop observability standards and dashboards used by other teams.
  • Champion SLOs and incident response practices to protect service quality.
  • Lead blameless post-mortems and continuous improvement.
  • Build automation and self-service tooling to reduce toil.
  • Mentor engineers on systems thinking and production ownership.
  • Ensure security best practices in infra and pipelines.

Skills

SRE concepts
Platform engineering
Observability tools
Docker
Kubernetes
Terraform
Bash
Python/Go
AI tools
Linux
AWS/GCP
Networking
Distributed systems

Education

BS in Computer Science/Engineering
Postgraduate degree desirable

Tools

Terraform

Job description

Cover Genius is building an AI-first infrastructure platform and seeks a Site Reliability Engineer to lead reliability initiatives across multiple teams. You will shape system design, tooling, and process to scale production systems while reducing operational risk.

You will apply cloud expertise (AWS/GCP), implement observability standards, and mentor engineers on best practices. The role blends hands-on coding with strategic reliability ownership in a hybrid work environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer - AI-Driven Infra (Hybrid)
Senior Site Reliability Engineer - AI-Driven Infra (Hybrid)

Socket.dev • Vancouver

Hybrid
CAD 115,000 - 145,000
Hybrid work environment
Hybrid work schedule with WFH midweek
Founding SRE: Cloud Reliability & Platform Lead
Founding SRE: Cloud Reliability & Platform Lead

Katalyze AI, Inc. • Toronto

On-site
CAD 120,000 - 180,000
Senior SRE: Cloud, Automation & Observability
Senior SRE: Cloud, Automation & Observability

RXinsider LTD. • Montreal (administrative region)

Hybrid
CAD 100,000 - 150,000
Senior Incident Command & Reliability Engineer
Senior Incident Command & Reliability Engineer

IBM • Vancouver

On-site
CAD 150,000 - 190,000
Senior SRE & Cloud Platform Manager
Senior SRE & Cloud Platform Manager

TekRek • Vancouver

On-site
CAD 150,000 - 210,000
Remote Senior Backend Engineer - AI-Driven Reliability
Remote Senior Backend Engineer - AI-Driven Reliability

Affirm • Kelowna

On-site
CAD 153,000 - 213,000
Health coverage
Flexible Spending Wallets
Time off
+1
Senior SRE: CI/CD Automation for Cloud & On-Prem Infra
Senior SRE: CI/CD Automation for Cloud & On-Prem Infra

HRB • Montreal (administrative region)

On-site
CAD 110,000 - 170,000
Azure SRE Lead — Cloud Reliability & Automation
Azure SRE Lead — Cloud Reliability & Automation

SimCorp • Toronto

Hybrid
CAD 113,000 - 142,000
Health and dental care
Group RRSP/TFSA
Hybrid work policy
Senior Incident Command Engineer – Cloud Reliability
Senior Incident Command Engineer – Cloud Reliability

IBM • Ottawa

Hybrid
CAD 120,000 - 160,000
Senior Software Engineer (Site Reliability Engineering - SRE Automation)
Senior Software Engineer (Site Reliability Engineering - SRE Automation)

Linux User Group • Saskatoon

Hybrid
CAD 167,000 - 222,000