Staff Site Reliability Engineer Toronto, Canada (remote) •

SoundHound Inc.

Toronto

Hybrid

CAD 140,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Healthcare
Paid time off
Remote work

Job summary

SoundHound Inc. is seeking a Staff Software Engineer (SRE) to own the reliability, scalability, and performance of our infrastructure on Google Cloud Platform.

You will architect high-availability systems, automate operations, and ensure our services can handle millions of voice AI interactions. You will lead incident response, improve observability, and partner with engineering teams to optimize cost and security.

Qualifications

  • 12+ years of software engineering experience with a focus on SRE/DevOps.
  • Expert-level experience with Google Cloud Platform services (GKE, Compute Engine, Cloud Run, Pub/Sub).
  • Proficient in Infrastructure as Code tools like Terraform or Pulumi.
  • Deep experience with Kubernetes and container orchestration.

Responsibilities

  • Design, build, and maintain highly available and scalable infrastructure on Google Cloud Platform.
  • Architect and automate CI/CD pipelines for rapid, reliable deployments.
  • Implement robust monitoring and observability strategies to proactively resolve issues.
  • Collaborate with engineering teams to optimize performance, cost, and reliability of backend services.
  • Lead incident response, post-mortem analysis, and remediation efforts.

Skills

SRE/DevOps
GCP
Kubernetes
Terraform
Pulumi
Observability
CI/CD
Security
Mentoring
Communication

Tools

Datadog
Prometheus
Grafana
Cloud Monitoring

Job description

The Opportunity

We’re looking for a Staff Software Engineer (SRE) to join our Retail and Restaurants AI team. You will be responsible for the reliability, scalability, and performance of our infrastructure, with a deep focus on Google Cloud Platform (GCP). You will architect and maintain high-availability systems, automate operational tasks, and ensure our services can handle the demands of millions of voice AI interactions.

What You’ll Do
  • Design, build, and maintain highly available and scalable infrastructure on Google Cloud Platform.
  • Architect and automate CI/CD pipelines to ensure rapid, reliable deployments.
  • Implement robust monitoring, alerting, and observability strategies to proactively identify and resolve system issues.
  • Partner with engineering teams to optimize performance, cost, and reliability of backend services.
  • Drive incident response, post-mortem analysis, and long-term remediation efforts.
  • Identify and eliminate sources of toil, promoting operational maturity and self-service capabilities.
  • Collaborate with cross-functional teams to ensure alignment on infrastructure roadmaps and security standards.
  • Lead department wide compliance (PCI, SOC) initiatives.
What You’ll Bring
  • 12+ years of software engineering experience, with significant experience in Site Reliability Engineering or DevOps roles.
  • Expert-level experience with Google Cloud Platform (GCP) services (e.g., GKE, Compute Engine, Cloud Run, Pub/Sub).
  • Proficient in Infrastructure as Code (IaC) tools like Terraform or Pulumi.
  • Deep experience with Kubernetes, container orchestration, and service mesh architectures.
  • Strong background in monitoring and observability tools (e.g., Datadog, Prometheus, Grafana, Cloud Monitoring).
  • Experience designing and managing high-throughput, distributed systems.
  • Strong problem-solving skills and a growth mindset—comfortable with ambiguity and making high-stakes technical trade-offs.
  • Excellent communication skills and a demonstrated ability to mentor engineers.
Preferred Qualifications
  • Experience working in a high-velocity, customer-focused environment.
  • Familiarity with functional programming paradigms (e.g., Clojure/ClojureScript).
  • Prior experience in the restaurant technology, hospitality, or AI-driven SaaS space.
  • Experience implementing security and compliance best practices in the cloud.
Workplace & Compensation

This role is available throughout Canada.

Compensation includes salary, equity, comprehensive healthcare, paid time off, and other benefits. Our recruiting team will provide a specific salary range based on location and years of experience.

#LI-MQ1 #LI-REMOTE

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Site Reliability Developer, Protected Data SRE
Staff Site Reliability Developer, Protected Data SRE

Google • Southwestern Ontario

On-site
CAD 216,000 - 221,000
Bonus
Equity
Benefits
Site Reliability Manager, Data center Networking, SRE
Site Reliability Manager, Data center Networking, SRE

Google Inc. • Southwestern Ontario

On-site
CAD 216,000 - 221,000
Manager, Site Reliability Engineering (SRE)
Manager, Site Reliability Engineering (SRE)

Quantum Technology Recruiting Inc. (QTR) • Toronto

On-site
CAD 155,000 - 165,000
Site Reliability Engineer
Site Reliability Engineer

TELUS Digital • Canada

Remote
CAD 90,000 - 120,000
Staff SRE: Cloud Reliability & Scale on GCP
Staff SRE: Cloud Reliability & Scale on GCP

SoundHound AI • Toronto

On-site
CAD 140,000 - 180,000
Equity
Healthcare
Paid time off
Site Reliability Consultant
Site Reliability Consultant

Pythian • Vancouver

On-site
CAD 90,000 - 100,000
Paid vacation
Training allowance
Wellness budget
+1
Site Reliability Manager, Data center Networking, SRE
Site Reliability Manager, Data center Networking, SRE

Google • Southwestern Ontario

On-site
CAD 216,000 - 221,000
Equity
Bonus target
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Thinkific • Canada

Remote
CAD 111,000 - 167,000
Fair and transparent pay
Inclusive work culture
Remote work flexibility
Software Developer III, Site Reliability
Software Developer III, Site Reliability

Google Inc. • Southwestern Ontario

On-site
CAD 150,000 - 153,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Sage Recruiting Inc. • Canada

On-site
CAD 180,000 - 200,000
Unlimited vacation
Comprehensive health and dental benefits