Senior Lead Site Reliability Engineer

Socket.dev

San Jose (CA)

Hybrid

USD 147,000 - 339,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Zoom in San Jose seeks an experienced Senior Lead Site Reliability Engineer to drive reliability across hybrid systems worldwide. You will install, configure, monitor systems, patch thousands of devices, and build automation to cut repetitive tasks while addressing performance bottlenecks.

Lead cross-team initiatives, mentor engineers, partner with Security, Networking, and Platform teams, design self-healing platforms, optimize Linux at scale, and oversee on-call incidents and after-hours

Qualifications

  • 10+ years in SRE, production engineering, or large-scale systems administration.
  • Strong Linux system administration experience (systemd, cgroups, networking, filesystems, performance analysis).
  • Proficient in at least one programming language such as Python.
  • Experience with Ansible, Terraform, Packer, Jenkins/GitLab, Kubernetes and Docker.
  • Proven incident response experience in mission-critical environments.
  • Security-first mindset including TPM, secure boot, identity and secrets management.
  • Networking expertise: BGP, load balancing, DNS, TLS, traffic engineering.
  • Experience with chaos engineering and resilience testing; Ceph knowledge is a plus.

Responsibilities

  • Provide technical direction for cross-team initiatives and major incidents.
  • Mentor SREs and developers; define best practices and design patterns.
  • Partner with Security, Networking, and Platform teams on roadmaps.
  • Influence vendor and hardware strategy for on-prem and cloud workloads.
  • Design self-healing platforms using automation and fault-tolerant patterns.
  • Optimize Linux systems at scale: performance tuning, kernel, networking, storage, security hardening.
  • Lead by example in adopting best practices across the company.
  • Communicate effectively and drive cross-team projects as a technical lead.
  • Participate in on-call shifts and post-hours deployments as needed.

Skills

SRE / production engineering
Linux system administration
Python
Incident response
Security mindset
Networking expertise
Chaos engineering
Distributed storage Ceph
On-call experience

Tools

Ansible
Terraform
Packer
Jenkins
GitLab
Kubernetes
Docker

Job description

What you can expect

As a Senior Lead Site Reliability Engineer, you can anticipate opportunities to work on our hybrid systems across the globe. You will be responsible for installing, configuring, and monitoring new systems within a network of global data centers. Additionally, you will patch and maintain thousands of physical and cloud systems worldwide. To streamline operations, you will develop automation to reduce repetitive tasks and analyze and address performance bottlenecks. Furthermore, you will update and troubleshoot user access permissions, resolve network connectivity issues, and maintain system firewalls.

About the Team

Zoom's SRE team is committed to delivering customer happiness, improving business efficiency, and promoting agility through innovation, data-driven insights, and automation. Our impact is reflected in smooth user experiences, optimized processes, and support for Zoom's expansion in the realm of communication and collaboration.

Responsibilities

Providing technical direction for cross-team initiatives and major incidents. Mentor SRE's and developers; define best practices and design patterns. Partner with Security, Networking, and Platform teams on architecture roadmaps. Influence vendor and hardware strategy for on-prem and cloud workloads. Design self-healing platforms using automation, chaos engineering, and fault-tolerant patterns. Optimize Linux systems at scale: performance tuning, kernel parameters, networking, storage, and security hardening. Define best practices and advocate for them across the company. Excellent communication skills and experience driving cross team projects as a technical lead. Able to participate in on-call shifts and incident management and work after hours/weekends for application releases/deployments.

What we’re looking for
  • 10+ years in SRE, production engineering, or large-scale systems administration
  • Have experience of Linux system administration (systemd, cgroups, networking, filesystems, performance analysis)
  • Demonstrate coding ability with at least one programming language e.g. Python
  • Have experience with configuration management (Ansible), IaC (Terraform, Packer), CI/CD pipelines (Jenkins, GitLab), container orchestration (k8s, Docker) and observability platforms.
  • Have experience with incident response for mission-critical environments.
  • Possess a security -first mindset (TPM, secure boot, identity, secrets management).
  • Demonstrate networking expertise: BGP, load balancing, DNS, TLS, traffic engineering.
  • Have experience with chaos engineering and resilience testing. Have experience with distributed storage systems such as Ceph
  • Occasional weekend work may be required
  • Ability to work across the globe or multiple time zones
Salary Range or On Target Earnings:

Minimum:

$146,700.00

Maximum:

$339,300.00

In addition to the base salary and/or OTE listed Zoom has a Total Direct Compensation philosophy that takes into consideration; base salary, bonus and equity value.

Note: Starting pay will be based on a number of factors and commensurate with qualifications & experience.

We also have a location based compensation structure; there may be a different range for candidates in this and other locations

Anticipated Position Close Date:

08/26/26

Ways of Working

Our structured hybrid approach is centered around our offices and remote work environments. The work style of each role, Hybrid, Remote, or In-Person is indicated in the job description/posting.

Benefits

As part of our award-winning workplace culture and commitment to delivering happiness, our benefits program offers a variety of perks, benefits, and options to help employees maintain their physical, mental, emotional, and financial health; support work-life balance; and contribute to their community in meaningful ways. Click Learnfor more information.

About Us

Zoomies help people stay connected so they can get more done together. We set out to build the best collaboration platform for the enterprise, and today help people communicate better with products like Zoom Contact Center, Zoom Phone, Zoom Events, Zoom Apps, Zoom Rooms, and Zoom Webinars. We’re problem-solvers, working at a fast pace to design solutions with our customers and users in mind. Find room to grow with opportunities to stretch your skills and advance your career in a collaborative, growth-focused environment.

Our Commitment

At Zoom, we believe great work happens when people feel supported and empowered. We’re committed to fair hiring practices that ensure every candidate is evaluated based on skills, experience, and potential. If you require an accommodation during the hiring process, let us know—we’re here to support you at every step.

If you need assistance navigating the interview process due to a medical disability, please submit an Accommodations Request Form and someone from our team will reach out soon. This form is solely for applicants who require an accommodation due to a qualifying medical disability. Non-accommodation-related requests, such as application follow-ups or technical issues, will not be addressed.

Our interviews are supported by BrightHire, a tool that helps us create a consistent and thoughtful interview experience and may include recordings. Please refer to our candidate privacy statement for more information of how we use your data.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

Zoom • San Jose (CA), Northern (KY)

On-site
USD 147,000 - 339,000
Lead Site Reliability Engineer, Platforms
Lead Site Reliability Engineer, Platforms

Zoom • San Jose (CA), Northern (KY)

On-site
USD 124,000 - 271,000
Lead Site Reliability Engineer, Platforms
Lead Site Reliability Engineer, Platforms

Socket.dev • San Jose (CA)

Hybrid
USD 124,000 - 271,000
Senior Platform Engineer (SRE) - Government Cloud Operations
Senior Platform Engineer (SRE) - Government Cloud Operations

Pantera Capital • San Jose (CA)

Hybrid
USD 99,000 - 229,000
Manager of Platform DevOps
Manager of Platform DevOps

Zoom • San Jose (CA)

Hybrid
USD 124,000 - 272,000
Manager of Platform DevOps
Manager of Platform DevOps

Zoom • Seattle (WA)

Hybrid
USD 124,000 - 272,000
Hybrid work model
Competitive compensation
Lead Site Reliability Engineer, Platforms
Lead Site Reliability Engineer, Platforms

Zoom • United States

Hybrid
USD 124,000 - 271,000
Staff DevOps Engineer
Staff DevOps Engineer

Socket.dev • San Jose (CA)

Hybrid
USD 124,000 - 271,000
Principal DevOps Engineer
Principal DevOps Engineer

Zoom • San Jose (CA)

Hybrid
USD 147,000 - 339,000
Head of Infrastructure and Employee Services
Head of Infrastructure and Employee Services

Zoom • San Jose (CA)

Hybrid
USD 194,000 - 427,000