Senior Real-Time DevOps SRE — Global Reliability

Zoom

San Jose (CA)

Hybrid

USD 99,000 - 229,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Hybrid work model
Competitive compensation

Job summary

Zoom is seeking a Senior DevOps/SRE Engineer to own reliability for a real-time communications platform, including audio/video, recording, and live streaming. You will lead SLO/SLI, incident response, and chaos testing while shaping deployment patterns and observability for global, latency-sensitive systems.

Ideal candidates have 5+ years in DevOps/SRE, strong cloud/IaC skills, and experience with real-time media stacks.

Qualifications

  • 5+ years in DevOps, SRE, or infrastructure roles.
  • Staff or principal level experience.
  • Experience owning reliability for large-scale systems.
  • Experience with real-time or media platforms.
  • Lead cross-functional technical initiatives without direct authority.
  • Understanding of real-time protocols: WebRTC, RTP/RTCP, TURN/STUN, SDP.
  • Cloud infrastructure experience: AWS, GCP, or Azure.
  • IaC tooling: Terraform, Pulumi, or equivalent.
  • Observability stacks: Prometheus, Grafana, Datadog, Jaeger, OpenTelemetry.
  • Networking fundamentals: BGP, anycast, DNS, load balancing, CDN.
  • CI/CD: GitHub Actions, Jenkins, Spinnaker.
  • Canary/release, feature flags, blue/green deployment.
  • Automation: Ansible, Python, Bash, or Go.
  • On-call rotation is required.
  • Ability to work across time zones.
  • Mandarin speaker preferred.
  • US citizenship preferred.

Responsibilities

  • Own the SLO/SLI framework for real-time services.
  • Lead incident response for outages across time zones.
  • Promote blameless postmortems with actionable improvements.
  • Implement chaos engineering and game day exercises.
  • Build observability dashboards and distributed tracing.
  • Architect deployment patterns and infra design for real-time services.
  • Review system designs for scalability and fault tolerance.
  • Drive capacity planning, traffic modeling, and cost optimization.
  • Evaluate infra tools, platforms, and vendors.
  • Ensure CI/CD standards and progressive rollout strategies.
  • Primary SRE partner for multiple teams across engineering.
  • Collaborate with network, security, product, and data teams.
  • Translate constraints into actionable recommendations for product leaders.
  • Establish DevOps best practices: IaC, GitOps, automated testing.
  • Guide senior engineers on SRE principles and patterns.
  • Serve as liaison between US and China/India teams.
  • Conduct architecture reviews and postmortems.
  • Maintain runbooks and documentation for global teams.

Skills

Reliability engineering
Distributed systems
Cross-functional leadership
DevOps practices
Incident response
SRE principles
Time zone coordination

Tools

Terraform
Pulumi
Prometheus
Grafana
Datadog
Jaeger
OpenTelemetry
GitHub Actions
Jenkins
Spinnaker

Job description

Zoom is seeking a Senior DevOps/SRE Engineer to own reliability for a real-time communications platform, including audio/video, recording, and live streaming. You will lead SLO/SLI, incident response, and chaos testing while shaping deployment patterns and observability for global, latency-sensitive systems.

Ideal candidates have 5+ years in DevOps/SRE, strong cloud/IaC skills, and experience with real-time media stacks.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior DevOps\\SRE Engineer
Senior DevOps\\SRE Engineer

Zoom • San Jose (CA)

Hybrid
USD 99,000 - 229,000
Hybrid work model
Competitive compensation
Staff Infra Engineer, SRE for Real-Time Voice Platform
Staff Infra Engineer, SRE for Real-Time Voice Platform

Vapi • San Francisco (CA)

On-site
USD 280,000 - 314,000
Base salary + equity
Comprehensive health coverage
Quarterly off-sites
+2
Senior Tech Ops & Reliability Leader 24/7 SRE (VoIP/Cloud)
Senior Tech Ops & Reliability Leader 24/7 SRE (VoIP/Cloud)

ClearCaptions, LLC • United States

On-site
USD 170,000 - 180,000
Comprehensive benefits package
Remote work flexibility
10% performance-based compensation
Senior DevOps & SRE Lead - 24/7 Telecom Platform (Remote)
Senior DevOps & SRE Lead - 24/7 Telecom Platform (Remote)

ClearCaptions, LLC. • United States

Remote
USD 170,000 - 180,000
Comprehensive benefits program
Flexible work hours
Senior SRE Lead — Real-Time Healthcare Platform
Senior SRE Lead — Real-Time Healthcare Platform

Kontakt.io • New York (NY)

Hybrid
USD 200,000 - 250,000
Hybrid work 3 days/week in NYC office.
Equity in a high-growth company
Health, dental, vision insurance
+1
Senior SRE: Global Cloud Reliability & Remote Flex
Senior SRE: Global Cloud Reliability & Remote Flex

Unifonic, Inc. • United States

Remote
USD 140,000 - 210,000
Competitive salary and bonus
Unifonic share scheme
30 holiday days after first year
+3
CloudDevs: Senior Site Reliability Engineer (SRE)
CloudDevs: Senior Site Reliability Engineer (SRE)

Breakout Tools • San Francisco (CA)

On-site
USD 120,000 - 160,000
Remote Senior Network Reliability Engineer (SRE)
Remote Senior Network Reliability Engineer (SRE)

Gainbridge • Zionsville (IN), Northern (KY)

On-site
USD 135,000 - 190,000
Health Insurance
Dental Insurance
Vision Insurance
+4
SRE Leader: Real-Time Healthcare Reliability
SRE Leader: Real-Time Healthcare Reliability

Kontakt Micro-Location Sp. Z.o.o. • New York (NY)

Hybrid
USD 180,000 - 260,000
Equity in a high-growth company
Health, dental, and vision coverage
401k
+3
Senior SRE & Software Engineer: Domain Reliability Lead
Senior SRE & Software Engineer: Domain Reliability Lead

Hispanic Alliance for Career Enhancement • Richardson (TX)

On-site
USD 93,000 - 204,000