Real-Time Platform Reliability Engineer

Socket.dev

San Jose (CA)

Hybrid

USD 99,000 - 229,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Zoom is seeking a DevOps/SRE engineer to strengthen reliability for its real-time communications platform, including audio/video streaming, recording, and live events. You will own SLO/SLI, drive incident response, and champion chaos engineering and observability across global teams.

You will collaborate with network, security, product, and data teams, shaping deployment patterns and IaC practices to ensure scalable, fault-tolerant operations.

Qualifications

  • 5+ years in DevOps, SRE, or infrastructure engineering roles, with at least 3 years at a staff or principal level scope.
  • Experience owning reliability for large-scale, distributed, latency-sensitive systems in production.
  • Experience supporting real-time or media-heavy platforms (video conferencing, live streaming, gaming, trading systems, or similar).
  • Ability to lead cross-functional technical initiatives without direct authority, driving alignment across teams.
  • Conceptual and architectural understanding of real-time communication protocols: WebRTC, RTP/RTCP, TURN/STUN, SDP, SFU/MCU.
  • Solid expertise in cloud infrastructure (AWS, GCP, or Azure) and container orchestration (Kubernetes, Helm, ArgoCD).
  • Proficiency with infrastructure-as-code tooling: Terraform, Pulumi, or equivalent.
  • Observability stacks: Prometheus, Grafana, Datadog, Jaeger, OpenTelemetry, or equivalent.
  • Understanding of networking fundamentals: BGP, anycast, DNS, load balancing, CDN architecture.
  • CI/CD tooling: GitHub Actions, Jenkins, Spinnaker.

Responsibilities

  • Own the SLO/SLI framework for real-time services and optimize latency, availability, jitter, and packet loss.
  • Lead incident response for critical outages across the real-time platform across multiple time zones.
  • Promote blameless postmortems and ensure action items yield reliability improvements.
  • Implement chaos engineering and game-day exercises to identify failure modes before user impact.
  • Build and evolve observability dashboards, alerts, and distributed tracing for real-time media infra.
  • Serve as architectural authority on deployment patterns, infra design, and operational readiness for real-time services.
  • Review system design proposals, feedback on scalability, fault tolerance, and operational complexity.
  • Drive capacity planning, traffic modeling, and cost optimization across globally distributed infra.
  • Evaluate infra tools/platforms/vendors including media servers, CDN, cloud-native services, and edge networking.
  • Ensure CI/CD standards, deployment safety, and progressive rollout strategies across teams.
  • Act as primary SRE partner for multiple teams building real-time features; align with product and ops.
  • Collaborate with network, security, product, and data teams on reliability requirements.
  • Advocate infrastructure-as-code, GitOps, automated testing, and deployment automation across teams.
  • Guide senior engineers on SRE principles and reliability patterns.
  • Liaise between US-based and China/India-based engineering teams; communicate across time zones.

Skills

DevOps
SRE
Cloud (AWS/GCP/Azure)
Kubernetes
CI/CD
IaC
Observability
Incident response
Automation (Python/Bash/Go)
Chaos engineering

Tools

Terraform
Pulumi
Prometheus
Grafana
Datadog
Jaeger
OpenTelemetry
GitHub Actions
Jenkins
Spinnaker
ArgoCD

Job description

Zoom is seeking a DevOps/SRE engineer to strengthen reliability for its real-time communications platform, including audio/video streaming, recording, and live events. You will own SLO/SLI, drive incident response, and champion chaos engineering and observability across global teams.

You will collaborate with network, security, product, and data teams, shaping deployment patterns and IaC practices to ensure scalable, fault-tolerant operations.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Real-Time Cloud SRE & Reliability Engineer
Real-Time Cloud SRE & Reliability Engineer

Zoom • San Jose (CA)

Hybrid
USD 99,000 - 229,000
Senior Staff DevOps/SRE for Real-Time Media Platform
Senior Staff DevOps/SRE for Real-Time Media Platform

Maven • San Jose (CA)

Hybrid
USD 124,000 - 271,000
Cloud Operations Engineer
Cloud Operations Engineer

Zoom • San Jose (CA)

Hybrid
USD 99,000 - 229,000
Staff DevOps Engineer
Staff DevOps Engineer

Maven • San Jose (CA)

Hybrid
USD 124,000 - 271,000
Cloud Operations Engineer
Cloud Operations Engineer

Socket.dev • San Jose (CA)

Hybrid
USD 99,000 - 229,000
Senior Real-Time Cloud Reliability Engineer
Senior Real-Time Cloud Reliability Engineer

cloudzero • Boston (MA)

On-site
USD 150,000 - 210,000
Senior Real-Time Cloud Reliability Engineer
Senior Real-Time Cloud Reliability Engineer

CloudZero • San Francisco (CA)

On-site
USD 180,000 - 260,000
Senior Escalation Engineer - AI-Powered Reliability
Senior Escalation Engineer - AI-Powered Reliability

Zoom • Boise (ID)

Hybrid
USD 98,000 - 226,000
Benefits package
Production Platform Reliability Engineer
Production Platform Reliability Engineer

United States Digital Space LLC • United States

Hybrid
USD 140,000 - 180,000
Equity participation
Health insurance
401(k) retirement plan
+1
Remote Platform Reliability Engineer
Remote Platform Reliability Engineer

Redwood Logistics LLC • Northern (KY)

Hybrid
USD 160,000 - 175,000
Health insurance
401k with match
Paid time off
+1