Staff Site Reliability Engineer - High-Scale Production

Zscaler

Virginia (MN)

Hybrid

USD 180,000 - 240,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Zscaler is seeking a Staff Site Reliability Engineer (Production Engineer) for a hybrid role based in the San Jose area. You will own the systems-level reliability of high-throughput bare-metal and cloud infrastructure, debugging issues at OS and network level, and driving platform resilience.

Responsibilities include incident response, automation of lifecycle management, telemetry orchestration with Prometheus/OpenTelemetry, and CI/CD validation to ensure seamless production deployments.

Qualifications

  • 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms.
  • Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation.
  • Deep knowledge of Linux OS internals and kernel troubleshooting and networking concepts; comfortable debugging at OS/network levels.
  • US Citizenship is required due to sensitive customer engagements.

Responsibilities

  • Maintain high availability across large-scale bare-metal or cloud-native infrastructure and Kubernetes clusters.
  • Lead incident response with cross-stack troubleshooting using low-level OS and network tools.
  • Automate infrastructure lifecycle management, service provisioning, configurations, and release deployments.
  • Operate telemetry pipelines (metrics, logs, traces) using Prometheus and OpenTelemetry; define/review SLOs and error budgets.
  • Perform capacity planning, OS/kernel upgrades, and CI/CD validation prior to production rollouts.

Skills

SRE
Python
Go
Bash
Linux
Kubernetes
Ansible
Telemetry

Tools

Prometheus
OpenTelemetry
Ansible

Job description

Zscaler is seeking a Staff Site Reliability Engineer (Production Engineer) for a hybrid role based in the San Jose area. You will own the systems-level reliability of high-throughput bare-metal and cloud infrastructure, debugging issues at OS and network level, and driving platform resilience.

Responsibilities include incident response, automation of lifecycle management, telemetry orchestration with Prometheus/OpenTelemetry, and CI/CD validation to ensure seamless production deployments.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Site Reliability Engineer - Production (Hybrid)
Staff Site Reliability Engineer - Production (Hybrid)

Zscaler • United States

Hybrid
USD 119,000 - 170,000
Time off plans for vacation and sick
Parental leave options
Retirement options
+1
Staff Production Engineer (SRE) (Federal)
Staff Production Engineer (SRE) (Federal)

Zscaler • Virginia (MN)

Hybrid
USD 180,000 - 240,000
Lead Production Engineer: Automation & Reliability (Remote)
Lead Production Engineer: Automation & Reliability (Remote)

Zscaler • San Jose (CA)

Hybrid
USD 130,000 - 170,000
Time off plans for vacation
Parental leave options
Retirement options
+1
Automation & Reliability Engineer — Hybrid/Remote
Automation & Reliability Engineer — Hybrid/Remote

Zscaler • San Jose (CA)

Hybrid
USD 102,400 - 128,000
Various health plans
Time off plans for vacation and sick time
Parental leave options
+3
Staff Site Reliability Engineer (Production Engineer)- Federal
Staff Site Reliability Engineer (Production Engineer)- Federal

Zscaler • United States

Hybrid
USD 119,000 - 170,000
Time off plans for vacation and sick
Parental leave options
Retirement options
+1
Senior Site Reliability Engineer: Scalable Hybrid Infra
Senior Site Reliability Engineer: Scalable Hybrid Infra

Redwood Materials • Nevada (IA)

On-site
USD 140,000 - 180,000
Remote Site Reliability Engineer - Production Support
Remote Site Reliability Engineer - Production Support

Bayside Solutions • Cupertino (CA)

Remote
USD 83,000 - 96,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Axiom Pursuits • San Francisco (CA)

On-site
USD 150,000 - 180,000
Senior Site Reliability Engineer — Remote Production Reliability
Senior Site Reliability Engineer — Remote Production Reliability

Fingerprint • Chicago (IL)

Remote
USD 152,000 - 205,000
Lead Site Reliability Engineer - Architect & Own Production
Lead Site Reliability Engineer - Architect & Own Production

Optimal Market Technologies • New York (NY)

On-site
USD 175,000 - 200,000