Lead Site Reliability Engineer, Platforms

Zoom

United States

Hybrid

USD 124,000 - 271,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Zoom is seeking a Lead Staff Site Reliability Engineer to serve as a technical lead for the DevOps Platforms group. You will guide cross‑team initiatives, shape roadmaps, and drive reliability across Kubernetes, cloud, and on‑prem environments including Zoom for Government (ZfG).

The role emphasizes SRE best practices, automation, monitoring, and incident management with broad impact, mentoring, and strong collaboration with security teams and leadership.

Qualifications

  • 8+ years of SRE or DevOps experience building and operating production infrastructure at scale.
  • Code proficiency in at least one language beyond scripting (e.g., Python, Go, Java).
  • Deploy and manage CI/CD pipelines using Git, Jenkins, Argo CD, or JFrog.
  • Operate cloud infrastructure on AWS, OCI, or similar platforms using Terraform and Kubernetes.
  • Implement observability with logging/monitoring tools like ELK, Prometheus, or Grafana.
  • Communicate complex technical concepts clearly to diverse audiences including security teams and leadership.
  • Participate in on‑call rotations and lead incident response to maintain reliability.
  • Hold a degree in Computer Science or related field, or equivalent practical experience.
  • Hold US citizenship, or Greencard status

Responsibilities

  • Design and scale DevOps platform services including Kubernetes infrastructure, cloud systems, and compliance-ready environments.
  • Define technical roadmaps and architectural direction for infrastructure automation and security.
  • Partner with service teams to understand platform needs and deliver solutions improving reliability and efficiency.
  • Establish and advocate SRE best practices including infrastructure as code, monitoring, and incident management.
  • Mentor team members through design, implementation, and production deployment of complex systems.

Skills

SRE/DevOps experience
Programming in Python/Go/Java
CI/CD pipelines

Education

Bachelor's degree in Computer Science or related field

Tools

Git
Jenkins
Argo CD
JFrog
Terraform
Kubernetes
ELK
Prometheus
Grafana

Job description

What You Can Expect

As a Lead Staff Site Reliability Engineer, you will be one of the technical leads for our DevOps Platforms organization. This group is responsible for DevOps Platforms including cloud infrastructure, physical data center orchestration, critical security services, and our Zoom for Government (ZfG) environment. You will be an uber tech lead working across a broad area, defining projects and guiding work across various teams. Your scope of work is wide and you will have the opportunity to improve our datacenter kubernetes infrastructure, our cloud infrastructure, our security posture, and our operation of ZfG environments. Broadly speaking, you are an exemplary SRE and you will guide all of our teams toward SRE best practices (automation, monitoring, infrastructure as code, etc).

About the Team

The DevOps Platforms organization owns the full infrastructure stack: cloud infrastructure on AWS and OCI, physical data center orchestration, critical security services (identity, authentication, authorization), and Zoom's FedRAMP-rated federal environment, Zoom for Government (ZfG). The team is currently working on one of the most technically interesting work in the org, hardening our security posture, and building the automation and reliability systems that underpin Zoom's global services. If you want broad visibility, real cross-team influence, and the chance to define how infrastructure gets built, this is the seat.

Responsibilities
  • Design and scale DevOps platform services including Kubernetes infrastructure, cloud systems, and compliance-ready environments
  • Define technical roadmaps and architectural direction for infrastructure automation and security
  • Partner with service teams to understand platform needs and deliver solutions that improve reliability and efficiency
  • Establish and advocate for SRE best practices including infrastructure as code, monitoring, and incident management
  • Mentor team members through design, implementation, and production deployment of complex systems
What We’re Looking For
  • Bring 8+ years of SRE or DevOps experience building and operating production infrastructure at scale
  • Code proficiently in at least one programming language beyond scripting (e.g., Python, Go, Java)
  • Deploy and manage CI/CD pipelines using tools like Git, Jenkins, Argo CD, or JFrog
  • Operate cloud infrastructure on AWS, OCI, or similar platforms using Terraform and Kubernetes
  • Implement observability solutions with logging and monitoring tools such as ELK, Prometheus, or Grafana
  • Communicate complex technical concepts clearly to diverse audiences including security teams, senior leadership, and external auditors
  • Participate in on‑call rotations and lead incident response to maintain system reliability
  • Hold a degree in Computer Science or related field, or equivalent practical experience
  • Hold US citizenship, or Greencard status
Preferred
  • Have experience with security from an SRE perspective
  • Have experience with Identity security (e.g. IAM, workload identity, zero trust) and tools (e.g. Teleport, Okta)
  • Have experience operating Government environments and understanding their compliance requirements
  • Have experience with system design and distributed computing at scale
  • Ability to speak Chinese/Mandarin is a plus, but not required
Salary Range or On Target Earnings

Minimum: $124,000.00

Maximum: $271,200.00

In addition to the base salary and/or OTE listed Zoom has a Total Direct Compensation philosophy that takes into consideration; base salary, bonus and equity value.

Note: Starting pay will be based on a number of factors and commensurate with qualifications & experience.

We also have a location based compensation structure; there may be a different range for candidates in this and other locations

Anticipated Position Close Date

08/31/26

Ways of Working

Our structured hybrid approach is centered around our offices and remote work environments. The work style of each role, Hybrid, Remote, or In-Person is indicated in the job description/posting.

Benefits

As part of our award-winning workplace culture and commitment to delivering happiness, our benefits program offers a variety of perks, benefits, and options to help employees maintain their physical, mental, emotional, and financial health; support work-life balance; and contribute to their community in meaningful ways.

About Us

Zoomies help people stay connected so they can get more done together. We set out to build the best collaboration platform for the enterprise, and today help people communicate better with products like Zoom Contact Center, Zoom Phone, Zoom Events, Zoom Apps, Zoom Rooms, and Zoom Webinars. We're problem-solvers, working at a fast pace to desi

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer, Platforms
Lead Site Reliability Engineer, Platforms

Socket.dev • San Jose (CA)

Hybrid
USD 124,000 - 271,000
Lead Site Reliability Engineer, Platforms
Lead Site Reliability Engineer, Platforms

Zoom • San Jose (CA), Northern (KY)

On-site
USD 124,000 - 271,000
Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

Socket.dev • San Jose (CA)

Hybrid
USD 147,000 - 339,000
Senior Platform Engineer (SRE) - Government Cloud Operations
Senior Platform Engineer (SRE) - Government Cloud Operations

Zoom • United States

Hybrid
USD 99,000 - 229,000
Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

Zoom • San Jose (CA), Northern (KY)

On-site
USD 147,000 - 339,000
Manager of Platform DevOps
Manager of Platform DevOps

Zoom • San Jose (CA)

Hybrid
USD 124,000 - 272,000
Senior Platform Engineer (SRE) - Government Cloud Operations
Senior Platform Engineer (SRE) - Government Cloud Operations

Pantera Capital • San Jose (CA)

Hybrid
USD 99,000 - 229,000
Manager of Platform DevOps
Manager of Platform DevOps

Zoom • Seattle (WA)

Hybrid
USD 124,000 - 272,000
Hybrid work model
Competitive compensation
Security DevOps Engineer
Security DevOps Engineer

Zoom • United States

Hybrid
USD 88,000 - 186,000
Staff DevOps Engineer
Staff DevOps Engineer

Socket.dev • San Jose (CA)

Hybrid
USD 124,000 - 271,000