Staff SRE: Cloud Reliability & Scale on GCP

SoundHound AI

Toronto

On-site

CAD 140,000 - 180,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Healthcare
Paid time off

Job summary

SoundHound AI is hiring a Staff Software Engineer (SRE) to secure reliability, scalability, and performance of our cloud infrastructure across Canada. You’ll design and maintain robust systems on GCP, automate CI/CD, and drive incident response with a focus on security and cost efficiency.

You will mentor engineers, collaborate with cross-functional teams, and lead PCI/SOC compliance initiatives while pushing for operational maturity and self-service tooling.

Qualifications

  • Minimum 12+ years in software engineering with SRE/DevOps background.
  • Proficient with Google Cloud Platform services (GKE, Compute Engine, Cloud Run, Pub/Sub).
  • Strong IaC experience with Terraform or Pulumi.
  • Deep knowledge of Kubernetes, containers, and service mesh concepts.
  • Strong monitoring/observability using Datadog, Prometheus, Grafana, or Cloud Monitoring.
  • Experience designing high-throughput distributed systems.
  • Excellent communication and mentoring abilities.

Responsibilities

  • Design, build, and maintain scalable infra on GCP.
  • Architect and automate CI/CD pipelines for reliable deployments.
  • Implement robust monitoring, alerts, and observability strategies.
  • Work with engineering to optimize performance, cost, and reliability.
  • Lead incident response, post-mortems, and remediation efforts.
  • Eliminate toil and promote self-service capabilities.
  • Collaborate on infra roadmaps and security standards.
  • Lead department-wide PCI/SOC compliance initiatives.

Skills

GCP
Kubernetes
IaC (Terraform/Pulumi)
Observability/Monitoring
SRE/DevOps
Distributed systems
Incident response
Mentoring engineers
Communication

Tools

Terraform
Pulumi

Job description

SoundHound AI is hiring a Staff Software Engineer (SRE) to secure reliability, scalability, and performance of our cloud infrastructure across Canada. You’ll design and maintain robust systems on GCP, automate CI/CD, and drive incident response with a focus on security and cost efficiency.

You will mentor engineers, collaborate with cross-functional teams, and lead PCI/SOC compliance initiatives while pushing for operational maturity and self-service tooling.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Cloud SRE & Reliability Leader
Senior Cloud SRE & Reliability Leader

Jobgether • Canada

Hybrid
CAD 151,000 - 200,000
Health benefits
Equity stock options
Remote-friendly environment
+2
Staff Site Reliability Developer, Protected Data SRE
Staff Site Reliability Developer, Protected Data SRE

Google • Southwestern Ontario

On-site
CAD 216,000 - 221,000
Bonus
Equity
Benefits
Staff Site Reliability Developer, Google Unified Security and Threat Operations
Staff Site Reliability Developer, Google Unified Security and Threat Operations

Google • Southwestern Ontario

On-site
CAD 216,000 - 221,000
Staff Site Reliability Developer, Google Unified Security and Threat Operations
Staff Site Reliability Developer, Google Unified Security and Threat Operations

Google Inc. • Southwestern Ontario

On-site
CAD 216,000 - 221,000
Senior Site Reliability Engineer - Global Infra & CI/CD Impact
Senior Site Reliability Engineer - Global Infra & CI/CD Impact

CloudFactory Limited • Canada

Hybrid
CAD 120,000 - 160,000
Hybrid Working Model
Comprehensive medical cover
Group life insurance
+3
Senior SRE: Kubernetes Reliability for Cloud UI Services
Senior SRE: Kubernetes Reliability for Cloud UI Services

Worky • Montreal (administrative region)

On-site
CAD 120,000 - 170,000
Laptop
Flexible work arrangements
Professional development and training
Staff Site Reliability Developer, Protected Data SRE
Staff Site Reliability Developer, Protected Data SRE

Google Inc. • Southwestern Ontario

Hybrid
CAD 216,000 - 221,000
Equity
Bonus target (20%)
Site Reliability Manager, Data center Networking, SRE
Site Reliability Manager, Data center Networking, SRE

Google • Southwestern Ontario

On-site
CAD 216,000 - 221,000
Equity
Bonus target
Senior SRE: AI/ML HPC Infra & GPU Cluster
Senior SRE: AI/ML HPC Infra & GPU Cluster

Boson AI • Toronto

On-site
CAD 100,000 - 130,000
Site Reliability Manager, Data center Networking, SRE
Site Reliability Manager, Data center Networking, SRE

Google Inc. • Southwestern Ontario

On-site
CAD 216,000 - 221,000