Senior Production Engineer - SRE

Zoom

San Jose (CA)

Hybrid

USD 99,000 - 229,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Zoom is seeking a hands-on Platform Engineer to own the reliability, scalability, and operational excellence of government-facing products. You will own the observability platform, work across Kubernetes, Terraform, and cloud infra, and respond to incidents with urgency, delivering permanent improvements.

You will design Terraform modules, runbooks, and SLOs; operate production Kubernetes workloads at scale; collaborate to reduce toil and raise resilience across teams.

Qualifications

  • US citizenship or lawful permanent resident status.
  • 4–5+ years of hands-on Platform Engineering, SRE, DevOps, or Production Engineering in live production.
  • Production-grade Kubernetes experience at scale.
  • Authors Terraform modules and providers from scratch.
  • Owns incidents end-to-end with on-call, triage, and root-cause analysis.
  • Deploys or improves observability/monitoring with Python or Bash.
  • Clear written communication; runbooks and incident reports.

Responsibilities

  • Architect end-to-end deployment lifecycle of the observability platform across cloud providers.
  • Design and build Terraform modules and providers; automate provisioning and config management.
  • Operate and improve production Kubernetes workloads; lead incident response and permanent fixes.
  • Develop reusable runbooks and standards for observability onboarding and alerting.
  • Collaborate to define and track SLOs/SLIs; reduce toil via automation.

Skills

Kubernetes operations
On-call incident experience
Terraform modules & providers
Python / Bash scripting
Technical writing / runbooks

Tools

Datadog
Oracle OCI
AWS GovCloud
GitLab CI/CD

Job description

What You Can Expect

In this role, you will be responsible for the reliability, scalability, and operational excellence of Zoom's government and military-facing products. A core focus is deploying and evolving our observability platform, working across Kubernetes, Terraform, and cloud infrastructure to build the systems and frameworks that set the standard for how the team operates. This is a hands-on, on-call engineering role: you will own the systems you build, respond to production incidents with urgency, and deliver permanent improvements.

About the Team

We build and own the cloud infrastructure behind Zoom's government products, operating with full end-to-end accountability. Our team moves fast, solves hard problems, and makes a direct impact where reliability is non-negotiable.

Responsibilities
  • Architecting and owning the end-to-end deployment lifecycle of Zoom's observability platform across AWS GovCloud and Oracle OCI, streamlining processes, eliminating manual steps, and setting the standard for how teams deploy at scale

  • Designing and building Terraform modules and providers from scratch, alongside GitLab CI/CD pipelines, to automate infrastructure provisioning and configuration management across development and production environments

  • Operating and improving production Kubernetes workloads at scale, leading incident response, driving root cause analysis, and delivering permanent fixes that reduce outage risk and improve system resilience

  • Developing reusable frameworks, runbooks, and operational standards that engineering teams across the organization can adopt — raising the bar on observability onboarding, alerting coverage, and operational readiness

  • Collaborating with engineering teams to define and track SLOs and SLIs, contribute to architecture and design reviews, and systematically eliminate toil through automation and thoughtful infrastructure design

What We're Looking For
  • Holds U.S. Citizenship or Lawful Permanent Resident (Green Card) status

  • Brings 4–5+ years of hands-on experience in Platform Engineering, Site Reliability Engineering, DevOps, or Production Engineering in a live production environment

  • Demonstrates production-grade Kubernetes experience, including workload operations, troubleshooting, and cluster-level architectural design at scale

  • Authors Terraform modules and providers from scratch, not limited to applying or executing existing configurations

  • Owns incidents end-to-end - proven on-call experience including structured triage, resolution, and post-incident review with documented root cause analysis

  • Deploys, operates, or improves observability or monitoring platforms at production scale, with scripting proficiency in Python and/or Bash for automation and operational tooling

  • Communicates clearly in writing - produces runbooks, incident reports, methods of procedure, and architectural documentation to a high standard

Preferred:
  • Has hands-on Datadog experience building observability frameworks in a production environment, ideally within a government, FedRAMP, or high-compliance cloud context

  • Experience operating within FedRAMP, AWS GovCloud, or other regulated or compliance-driven cloud environments

  • Hands-on experience with popular open-source observability tooling in a production context

  • Background supporting 24/7 or mission-critical production operations

  • Experience driving cross-team adoption of operational standards or tooling frameworks

  • Oracle OCI hands-on operational experience

  • Familiarity with container image security and CVE remediation workflows

Salary Range or On Target Earnings:

Minimum:

$98,900.00

Maximum:

$228,700.00

In addition to the base salary and/or OTE listed Zoom has a Total Direct Compensation philosophy that takes into consideration; base salary, bonus and equity value.

Note: Starting pay will be based on a number of factors and commensurate with qualifications & experience.

We also have a location based compensation structure; there may be a different range for candidates in this and other locations

Anticipated Position Close Date: 09/17/26

Ways of Working

Our structured hybrid approach is centered around our offices and remote work environments. The work style of each role, Hybrid, Remote, or In-Person is indicated in the job description/posting.

Benefits

As part of our award-winning workplace culture and commitment to delivering happiness, our benefits program offers a variety of perks, benefits, and options to help employees maintain their physical, mental, emotional, and financial health; support work-life balance; and contribute to their community in meaningful ways. Click Learn for more information.

About Us

Zoomies help people stay connected so they can get more done together. We set out to build the best collaboration platform for the enterprise, and today help people communicate better with products like Zoom Contact Center, Zoom Phone, Zoom Events, Zoom Apps, Zoom Rooms, and Zoom Webinars. We’re problem-solvers, working at a fast pace to design solutions with our customers and users in mind. Find room to grow with opportunities to stretch your skills and advance your career in a collaborative, growth-focused environment.

Our Commitment

At Zoom, we believe great work happens when people feel supported and empowered. We’re committed to fair hiring practices that ensure every candidate is evaluated based on skills, experience, and potential. If you require an accommodation during the hiring process, let us know—we’re here to support you at every step.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer, Platforms
Lead Site Reliability Engineer, Platforms

Zoom • San Jose (CA)

Hybrid
USD 124,000 - 271,000
Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

Socket.dev • San Jose (CA)

Hybrid
USD 147,000 - 339,000
Benefits program
Flexible work options
Security DevOps Engineer
Security DevOps Engineer

Zoom • San Jose (CA)

Hybrid
USD 87,600 - 186,000
Cloud Operations Engineer
Cloud Operations Engineer

Socket.dev • San Jose (CA)

Hybrid
USD 99,000 - 229,000
DevOps Engineer
DevOps Engineer

Zoom Video Communications • Northern (KY)

Hybrid
USD 99,000 - 229,000
Manager of Platform DevOps
Manager of Platform DevOps

Zoom • Seattle (WA)

Hybrid
USD 124,000 - 271,200
Hybrid work model
Competitive compensation
Cloud Operations Engineer
Cloud Operations Engineer

Zoom • San Jose (CA)

Hybrid
USD 99,000 - 229,000
Principal DevOps Engineer
Principal DevOps Engineer

Socket.dev • San Jose (CA)

Hybrid
USD 147,000 - 339,000
Software Architect
Software Architect

Zoom • San Jose (CA)

Hybrid
USD 221,000 - 229,000
Senior Software Engineer
Senior Software Engineer

Zoom • Jackson (MS)

Hybrid
USD 99,000 - 229,000