Senior Site Reliability Engineer – AI-First Cloud

Socket.dev

Vancouver

Hybrid

CAD 115,000 - 145,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Hybrid work environment

Job summary

Cover Genius is seeking a Site Reliability Engineer to lead reliability and infrastructure initiatives across multiple teams. You’ll shape system design, tooling, and processes to operate at scale, with emphasis on cloud infrastructure, IaC, observability, and security.

You’ll work with AWS/GCP, Docker, Kubernetes, and AI-assisted development tools to reduce operational risk and improve product reliability while collaborating with security and software teams to deliver robust, scalable platforms.

Qualifications

  • 3+ years of experience in SRE/Platform Eng/DevOps.
  • Strong SRE and platform engineering principles.
  • Experience with modern observability tools (Datadog, Elasticsearch, Prometheus, Grafana).
  • Experience with Docker and Kubernetes in production.
  • Terraform for IaC and cloud automation.
  • Scripting in Bash and at least one language (Python/Go).
  • AI-driven development environments integration into pipelines.
  • Linux proficiency and strong networking knowledge.
  • AWS and/or GCP experience; Bachelor's degree in CS/Engineering.

Responsibilities

  • Analyze, test, and evolve systems for reliability and performance.
  • Architect and build cloud infrastructure across products.
  • Implement observability standards and tooling.
  • Lead incident response and blameless postmortems.
  • Reduce toil with automation and self-service tooling.
  • Develop runbooks and design standards for teams.

Skills

SRE/Platform Eng
Observability tools
AWS/GCP
Docker
Kubernetes
Terraform
Scripting (Bash)
Python/Go
AI tools for development
Linux
Networking & distributed systems

Education

Bachelor's degree in CS/Engineering

Tools

Datadog
Elasticsearch
Prometheus
Grafana

Job description

Cover Genius is seeking a Site Reliability Engineer to lead reliability and infrastructure initiatives across multiple teams. You’ll shape system design, tooling, and processes to operate at scale, with emphasis on cloud infrastructure, IaC, observability, and security.

You’ll work with AWS/GCP, Docker, Kubernetes, and AI-assisted development tools to reduce operational risk and improve product reliability while collaborating with security and software teams to deliver robust, scalable platforms.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE – AI-Driven Cloud & Reliability Leader
Senior SRE – AI-Driven Cloud & Reliability Leader

Cover Genius • Vancouver

Hybrid
CAD 115,000 - 145,000
Site Reliability Engineer — Kubernetes & Terraform
Site Reliability Engineer — Kubernetes & Terraform

Future Secure AI • Toronto

On-site
CAD 90,000 - 130,000
Senior Incident Command Engineer – Cloud Reliability
Senior Incident Command Engineer – Cloud Reliability

IBM • Ottawa

Hybrid
CAD 120,000 - 160,000
Remote Senior Backend Engineer - AI-Driven Reliability
Remote Senior Backend Engineer - AI-Driven Reliability

Affirm • Kelowna

On-site
CAD 153,000 - 213,000
Health coverage
Flexible Spending Wallets
Time off
+1
Senior Incident Command & Reliability Engineer
Senior Incident Command & Reliability Engineer

IBM • Vancouver

On-site
CAD 150,000 - 190,000
Platform Engineer - Databricks & ML Ops, CI/CD Focus
Platform Engineer - Databricks & ML Ops, CI/CD Focus

Jarvis Consulting Group • Toronto

On-site
CAD 125,000 - 160,000
Flexible paid time off
Professional development support
Comprehensive benefits package
+1
Site Reliability Engineer
Site Reliability Engineer

Future Secure AI • Toronto

On-site
CAD 90,000 - 130,000
Founding SRE: Cloud Reliability & Platform Lead
Founding SRE: Cloud Reliability & Platform Lead

Katalyze AI, Inc. • Toronto

On-site
CAD 120,000 - 180,000
Senior Site Reliability Engineer – Cloud, CI/CD & Automation
Senior Site Reliability Engineer – Cloud, CI/CD & Automation

Electronic Arts (EA) • Edmonton

On-site
CAD 122,000 - 171,000
Vacation 3 weeks (Canada)
Sick time 10 days
EI/QPIP top-up
+3
Senior Site Reliability Engineer — Cloud Automation & CI/CD
Senior Site Reliability Engineer — Cloud Automation & CI/CD

Electronic Arts (EA) • Victoria

On-site
CAD 122,000 - 171,000