Staff Site Reliability Engineer - Production (Hybrid)

Zscaler

United States

Hybrid

USD 119,000 - 170,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Time off plans for vacation and sick
Parental leave options
Retirement options
In-office perks

Job summary

Zscaler is seeking a Staff Site Reliability Engineer (Production Engineer) to own reliability for a high-scale global platform. This hybrid role in the US focuses on software-driven SRE, building automation with Python/Go, and improving incident response and telemetry across bare-metal and cloud infrastructure.

You will work on Linux/BSD fleets, Kubernetes, and custom routing stacks, writing production-grade code, and driving SRE practices with strong collaboration with Engineering and

Qualifications

  • US citizenship is required due to the nature of assigned customers.
  • 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering on high-scale, low-latency production platforms.
  • Proficient in writing and debugging executable code (Python/Go/Bash) with hands-on Ansible automation.

Responsibilities

  • Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and routing stacks in collaboration with Engineering and Networking teams.
  • Lead full-cycle incident response with cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb).
  • Automate infrastructure lifecycle management, service provisioning, configuration workflows, and releases using Ansible, Python, and Go; measure toil and convert to durable automation.
  • Own end-to-end telemetry with Prometheus and OpenTelemetry; define and enforce SLOs and reduce alert noise.
  • Perform architectural reviews, OS/kernel upgrades, capacity/performance tuning, and CI/CD validation prior to production rollouts; embed operability standards into service design.

Skills

Python
Go
Bash
Ansible
Linux internals
Networking
Observability

Tools

Kubernetes
Prometheus
OpenTelemetry
tcpdump
strace
lsof
gdb

Job description

Zscaler is seeking a Staff Site Reliability Engineer (Production Engineer) to own reliability for a high-scale global platform. This hybrid role in the US focuses on software-driven SRE, building automation with Python/Go, and improving incident response and telemetry across bare-metal and cloud infrastructure.

You will work on Linux/BSD fleets, Kubernetes, and custom routing stacks, writing production-grade code, and driving SRE practices with strong collaboration with Engineering and

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer — Automation & Cloud
Senior Site Reliability Engineer — Automation & Cloud

Socket.dev • Bellevue (WA)

Hybrid
USD 104,000 - 148,000
Health plans
Time off
Parental leave
+3
Staff Site Reliability Engineer — AI-Driven Zero Trust
Staff Site Reliability Engineer — AI-Driven Zero Trust

Zscaler • San Jose (CA)

Hybrid
USD 119,000 - 170,000
Health plans
Vacation & sick time
Parental leave
+3
Senior Production Engineer - Automation & Global Reliability
Senior Production Engineer - Automation & Global Reliability

Zscaler • Bellevue (WA)

Hybrid
USD 140,000 - 175,000
Automation & Reliability Engineer — Hybrid/Remote
Automation & Reliability Engineer — Hybrid/Remote

Zscaler • San Jose (CA)

Hybrid
USD 102,000 - 128,000
Various health plans
Time off plans for vacation and sick time
Parental leave options
+3
Senior Staff Production Engineer — Automation-First SRE Leader
Senior Staff Production Engineer — Automation-First SRE Leader

Socket.dev • Bellevue (WA)

Hybrid
USD 144,000 - 205,000
Remote-Eligible Senior Production Engineer, Automation-Driven SRE
Remote-Eligible Senior Production Engineer, Automation-Driven SRE

Zscaler • San Jose (CA)

Hybrid
USD 104,000 - 148,000
Lead Production Engineer: Automation & Reliability (Remote)
Lead Production Engineer: Automation & Reliability (Remote)

Zscaler • San Jose (CA)

Hybrid
USD 130,000 - 170,000
Time off plans for vacation
Parental leave options
Retirement options
+1
Automation-First Production Engineer – Cloud Infrastructure
Automation-First Production Engineer – Cloud Infrastructure

Zscaler • Bellevue (WA)

Hybrid
USD 102,000 - 128,000
Health plans
Time off plans
Parental leave options
+3
Staff Site Reliability Engineer (Production Engineer)- Federal
Staff Site Reliability Engineer (Production Engineer)- Federal

Zscaler • United States

Hybrid
USD 119,000 - 170,000
Time off plans for vacation and sick
Parental leave options
Retirement options
+1
Staff Site Reliability Engineer (Production Engineer)- Federal
Staff Site Reliability Engineer (Production Engineer)- Federal

Zscaler • San Jose (CA)

Hybrid
USD 119,000 - 170,000
Health plans
Vacation & sick time
Parental leave
+3