Staff Site Reliability Engineer (Copy)

Socket.dev

Toronto

On-site

CAD 170,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Socket.dev is seeking a Staff Software Engineer (SRE) to own reliability, scalability, and performance of our cloud infrastructure across GCP. You will architect high-availability systems and automate operations to support millions of voice AI interactions.

You’ll lead incident response, drive toil reduction, and collaborate with cross-functional teams to align roadmaps, security standards, and compliance initiatives such as PCI and SOC.

Qualifications

  • 12+ years of software engineering experience.
  • Expert-level experience with Google Cloud Platform (GCP) services.
  • Proficient in Infrastructure as Code (Terraform or Pulumi).
  • Deep experience with Kubernetes and service meshes.
  • Strong observability background with monitoring tools.
  • Experience designing high-throughput distributed systems.
  • Excellent communication and mentoring abilities.

Responsibilities

  • Design, build, and maintain highly available, scalable infrastructure on GCP.
  • Architect and automate CI/CD pipelines for rapid, reliable deployments.
  • Implement robust monitoring, alerting, and observability strategies.
  • Partner with engineering teams to optimize performance, cost, and reliability.
  • Drive incident response, post-mortem analysis, and remediation efforts.
  • Identify and eliminate toil, promoting self-service and maturity.
  • Collaborate to align infrastructure roadmaps and security standards.
  • Lead department-wide compliance initiatives (PCI, SOC).

Skills

GCP expertise
Terraform/Pulumi
Kubernetes
Monitoring: Datadog
Distributed systems
Mentoring
Communication
Cost optimization

Tools

Terraform
Pulumi
Kubernetes
Datadog
Prometheus
Grafana
Cloud Monitoring

Job description

The Opportunity

We’re looking for a Staff Software Engineer (SRE) to join our Retail and Restaurants AI team. You will be responsible for the reliability, scalability, and performance of our infrastructure, with a deep focus on Google Cloud Platform (GCP). You will architect and maintain high-availability systems, automate operational tasks, and ensure our services can handle the demands of millions of voice AI interactions.

What You’ll Do
  • Design, build, and maintain highly available and scalable infrastructure on Google Cloud Platform.
  • Architect and automate CI/CD pipelines to ensure rapid, reliable deployments.
  • Implement robust monitoring, alerting, and observability strategies to proactively identify and resolve system issues.
  • Partner with engineering teams to optimize performance, cost, and reliability of backend services.
  • Drive incident response, post-mortem analysis, and long-term remediation efforts.
  • Identify and eliminate sources of toil, promoting operational maturity and self-service capabilities.
  • Collaborate with cross-functional teams to ensure alignment on infrastructure roadmaps and security standards.
  • Lead department wide compliance (PCI, SOC) initiatives.
What You’ll Bring
  • 12+ years of software engineering experience, with significant experience in Site Reliability Engineering or DevOps roles.
  • Expert-level experience with Google Cloud Platform (GCP) services (e.g., GKE, Compute Engine, Cloud Run, Pub/Sub).
  • Proficient in Infrastructure as Code (IaC) tools like Terraform or Pulumi.
  • Deep experience with Kubernetes, container orchestration, and service mesh architectures.
  • Strong background in monitoring and observability tools (e.g., Datadog, Prometheus, Grafana, Cloud Monitoring).
  • Experience designing and managing high-throughput, distributed systems.
  • Strong problem-solving skills and a growth mindset—comfortable with ambiguity and making high-stakes technical trade-offs.
  • Excellent communication skills and a demonstrated ability to mentor engineers.
Preferred Qualifications
  • Experience working in a high-velocity, customer-focused environment.
  • Familiarity with functional programming paradigms (e.g., Clojure/ClojureScript).
  • Prior experience in the restaurant technology, hospitality, or AI-driven SaaS space.
  • Experience implementing security and compliance best practices in the cloud.
Workplace & Compensation

This role is available throughout Canada.

Compensation includes salary, equity, comprehensive healthcare, paid time off, and other benefits.

Our recruiting team will provide a specific salary range based on location and years of experience.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Site Reliability Developer, Google Unified Security and Threat Operations
Staff Site Reliability Developer, Google Unified Security and Threat Operations

Google Inc. • Southwestern Ontario

On-site
CAD 216,000 - 221,000
Staff Site Reliability Developer, Google Unified Security and Threat Operations
Staff Site Reliability Developer, Google Unified Security and Threat Operations

Google • Southwestern Ontario

On-site
CAD 216,000 - 221,000
Staff Site Reliability Developer, Protected Data SRE
Staff Site Reliability Developer, Protected Data SRE

Google • Southwestern Ontario

On-site
CAD 216,000 - 221,000
Bonus
Equity
Benefits
Staff Site Reliability Developer, Protected Data SRE
Staff Site Reliability Developer, Protected Data SRE

Google Inc. • Southwestern Ontario

Hybrid
CAD 216,000 - 221,000
Equity
Bonus target (20%)
Senior Software Developer, Site Reliability
Senior Software Developer, Site Reliability

Google • Southwestern Ontario

On-site
CAD 182,000 - 186,000
Site Reliability Manager, Data center Networking, SRE
Site Reliability Manager, Data center Networking, SRE

Google Inc. • Southwestern Ontario

On-site
CAD 216,000 - 221,000
Staff SRE: Cloud Reliability & Scale on GCP
Staff SRE: Cloud Reliability & Scale on GCP

SoundHound AI • Toronto

On-site
CAD 140,000 - 180,000
Equity
Healthcare
Paid time off
Site Reliability Manager, Data center Networking, SRE
Site Reliability Manager, Data center Networking, SRE

Google • Southwestern Ontario

On-site
CAD 216,000 - 221,000
Equity
Bonus target
Staff Devops Engineer
Staff Devops Engineer

HRB • Ottawa

On-site
CAD 150,000 - 190,000
Senior Software Developer, Site Reliability
Senior Software Developer, Site Reliability

Google Inc. • Southwestern Ontario

On-site
CAD 182,000 - 186,000
Equity
Benefits