Senior Site Reliability Engineer - AI-Driven Infra (Hybrid)

Socket.dev

Vancouver

Hybrid

CAD 115,000 - 145,000

Full time

10 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Hybrid work environment
Hybrid work schedule with WFH midweek

Job summary

Cover Genius is seeking a Site Reliability Engineer to lead reliability and infrastructure initiatives across multiple teams. You’ll shape system design, tooling, and processes to scale production systems and reduce operational risk.

You should have strong cloud expertise (AWS/GCP), IaC (Terraform), CI/CD, observability, security, and disaster recovery, with hands-on experience deploying and monitoring web apps and databases at scale.

Qualifications

  • 3+ years in SRE, Platform Engineering, or DevOps.
  • Strong SRE and platform engineering principles with cross-team delivery.
  • Experience with observability tools (Datadog, Prometheus, Grafana).
  • Hands-on with cloud native tech (Docker, Kubernetes).
  • IaC modules (Terraform) and scripting (Bash/Python).
  • Fluent in AI-driven development environments is a plus.
  • Experience with Linux, networking, and large-scale systems.
  • Bachelor's degree in CS/Engineering; advanced degrees desirable.

Responsibilities

  • Lead reliability and infra projects across teams from design to delivery.
  • Apply AWS/GCP expertise to build reliable cloud infrastructure.
  • Implement observability strategy and tooling standards.
  • Champion SLOs and error budgets for services in your area.
  • Act as incident commander for major production incidents.
  • Automate toil with self-service tooling and automation.
  • Develop runbooks and standards for engineering teams.
  • Apply AI-assisted development to infra problems and workflows.
  • Contribute to capacity planning and cost optimization.
  • Mentor engineers in systems thinking and production ownership.
  • Enforce security best practices like least-privilege and policy-as-code.

Skills

SRE
Platform Engineering
DevOps
Observability
AWS/GCP
Kubernetes
Terraform
CI/CD
Bash/Python
AI tooling

Education

Bachelor's degree in CS/Engineering
Postgraduate degree desirable

Tools

Docker
Kubernetes
Terraform
Datadog
Grafana
Elasticsearch
Prometheus

Job description

Cover Genius is seeking a Site Reliability Engineer to lead reliability and infrastructure initiatives across multiple teams. You’ll shape system design, tooling, and processes to scale production systems and reduce operational risk.

You should have strong cloud expertise (AWS/GCP), IaC (Terraform), CI/CD, observability, security, and disaster recovery, with hands-on experience deploying and monitoring web apps and databases at scale.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE – AI-Driven Cloud & Reliability Leader
Senior SRE – AI-Driven Cloud & Reliability Leader

Cover Genius • Vancouver

Hybrid
CAD 115,000 - 145,000
Site Reliability Engineer — Kubernetes & Terraform
Site Reliability Engineer — Kubernetes & Terraform

Future Secure AI • Toronto

On-site
CAD 90,000 - 130,000
Senior Incident Command & Reliability Engineer
Senior Incident Command & Reliability Engineer

IBM • Vancouver

On-site
CAD 150,000 - 190,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

TekRek • Vancouver

On-site
CAD 150,000 - 210,000
Senior Incident Command Engineer – Cloud Reliability
Senior Incident Command Engineer – Cloud Reliability

IBM • Ottawa

Hybrid
CAD 120,000 - 160,000
Remote Senior Backend Engineer - AI-Driven Reliability
Remote Senior Backend Engineer - AI-Driven Reliability

Affirm • Kelowna

On-site
CAD 153,000 - 213,000
Health coverage
Flexible Spending Wallets
Time off
+1
Senior Site Reliability Engineer, SRE
Senior Site Reliability Engineer, SRE

Jobtailor • Toronto

On-site
CAD 120,000 - 180,000
Founding SRE: Cloud Reliability & Platform Lead
Founding SRE: Cloud Reliability & Platform Lead

Katalyze AI, Inc. • Toronto

On-site
CAD 120,000 - 180,000
Senior Site Reliability Engineer – Cloud, CI/CD & Automation
Senior Site Reliability Engineer – Cloud, CI/CD & Automation

Electronic Arts (EA) • Edmonton

On-site
CAD 122,000 - 171,000
Vacation 3 weeks (Canada)
Sick time 10 days
EI/QPIP top-up
+3
Senior Site Reliability Engineer — Cloud Automation & CI/CD
Senior Site Reliability Engineer — Cloud Automation & CI/CD

Electronic Arts (EA) • Victoria

On-site
CAD 122,000 - 171,000