Real-Time Cloud SRE & Reliability Engineer

Zoom

San Jose (CA)

Hybrid

USD 99,000 - 229,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Zoom is hiring a DevOps/Site Reliability Engineer to ensure reliability, scalability, and operational excellence for a real-time communications platform. The role spans incident response, SLO/SLI ownership, and cross-region collaboration with teams across time zones.

The ideal candidate has extensive experience in cloud infrastructure, Kubernetes, and observability stacks, with a focus on real-time media workloads and reliable deployment practices.

Qualifications

  • 5+ years in DevOps, SRE, or infrastructure roles, with staff/principal level experience.
  • Owning reliability for large-scale, latency-sensitive systems in production.
  • Experience with real-time or media-heavy platforms (video, streaming, etc.).
  • Ability to lead cross-functional initiatives without direct authority.
  • Understanding of real-time protocols: WebRTC, RTP/RTCP, TURN/STUN, SDP.
  • Cloud infra (AWS, GCP, or Azure) and Kubernetes expertise.
  • IaC tooling experience: Terraform, Pulumi, or equivalent.
  • Observability stacks: Prometheus, Grafana, Datadog, Jaeger, OpenTelemetry.
  • Networking fundamentals: BGP, DNS, load balancing, CDN.
  • CI/CD: GitHub Actions, Jenkins, Spinnaker.
  • Canary releases, feature flags, blue/green strategies.
  • Python, Bash, or Go for automation and incident response.
  • Occasional weekend work; able to work across time zones.

Responsibilities

  • Own the SLO/SLI framework for real-time services and track latency and availability.
  • Lead incident response for outages across the real-time platform.
  • Promote blameless postmortems and drive reliability improvements.
  • Implement chaos engineering and game day exercises to identify failure modes.
  • Build observability tools—dashboards, alerts, and distributed tracing.
  • Serve as architectural authority on deployment patterns and infra design.
  • Review system designs for scalability, fault tolerance, and complexity.
  • Drive capacity planning, traffic modeling, and cost optimization.
  • Evaluate infrastructure tools, platforms, and vendors (media servers, CDN, cloud).
  • Ensure CI/CD standards and progressive rollout strategies.
  • Collaborate with network, security, product, and data teams.
  • Advise product leaders on reliability trade-offs and choices.
  • Advocate DevOps best practices: IaC, GitOps, automated testing, deployment automation.
  • Guide senior engineers on SRE principles and operational discipline.
  • Bridge US/China/India teams; English and Mandarin as appropriate.
  • Maintain runbooks and architectural decision records for global teams.

Skills

DevOps
SRE
Cloud infrastructure
Kubernetes
Terraform
CI/CD
Observability
Python
Bash
Go
Real-time media
Networking
Incident response

Tools

Prometheus
Grafana
Datadog
Jaeger
OpenTelemetry
GitHub Actions
Jenkins
Spinnaker
ArgoCD
Helm

Job description

Zoom is hiring a DevOps/Site Reliability Engineer to ensure reliability, scalability, and operational excellence for a real-time communications platform. The role spans incident response, SLO/SLI ownership, and cross-region collaboration with teams across time zones.

The ideal candidate has extensive experience in cloud infrastructure, Kubernetes, and observability stacks, with a focus on real-time media workloads and reliable deployment practices.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Staff DevOps/SRE for Real-Time Media Platform
Senior Staff DevOps/SRE for Real-Time Media Platform

Maven • San Jose (CA)

Hybrid
USD 124,000 - 271,000
Real-Time Platform Reliability Engineer
Real-Time Platform Reliability Engineer

Socket.dev • San Jose (CA)

Hybrid
USD 99,000 - 229,000
Senior Real-Time Cloud Reliability Engineer
Senior Real-Time Cloud Reliability Engineer

cloudzero • Boston (MA)

On-site
USD 150,000 - 210,000
Senior Real-Time Cloud Reliability Engineer
Senior Real-Time Cloud Reliability Engineer

CloudZero • San Francisco (CA)

On-site
USD 180,000 - 260,000
Cloud Operations Engineer
Cloud Operations Engineer

Zoom • San Jose (CA)

Hybrid
USD 99,000 - 229,000
Staff DevOps Engineer
Staff DevOps Engineer

Maven • San Jose (CA)

Hybrid
USD 124,000 - 271,000
Cloud Operations Engineer
Cloud Operations Engineer

Socket.dev • San Jose (CA)

Hybrid
USD 99,000 - 229,000
Lead SRE: Hybrid Cloud, Kubernetes & Security
Lead SRE: Hybrid Cloud, Kubernetes & Security

Zoom • San Jose (CA), Northern (KY)

Hybrid
USD 124,000 - 271,000
Remote Senior DevOps Engineer - SRE & Cloud Automation
Remote Senior DevOps Engineer - SRE & Cloud Automation

ZoomInfo Technologies LLC • United States

On-site
USD 133,000 - 209,000
Remote Senior SRE: Own Cloud Reliability & Automation
Remote Senior SRE: Own Cloud Reliability & Automation

SLAMcore • United States

On-site
USD 110,000 - 150,000