Senior Production SRE: Scale Reliability Engineer (Hybrid)

Zscaler

New York (NY)

Hybrid

USD 119,000 - 170,000

Full time

7 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Health plans
Time off plans for vacation and sick
Parental leave options
Retirement options
Education reimbursement
In-office perks, and more!

Job summary

Zscaler seeks a Staff Site Reliability Engineer (Production Engineer) to join our Zero Trust Exchange team in a hybrid role. You will own reliability of large-scale bare-metal and cloud infrastructure, write production-grade code, and drive proactive incident management across global fleets.

You should excel at OS/network debugging, automation with Ansible/Python/Go, and shaping SRE practices with telemetry, SLOs, and CI/CD readiness to keep services stable at scale.

Qualifications

  • US Citizenship is required.
  • Foundational understanding of AI/ML technologies and experience leveraging AI-driven solutions to optimize outcomes.
  • 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms.
  • Proven ability to write and debug executable code live (Python, Go, or Bash) and hands-on experience writing Ansible playbooks/tasks for infrastructure automation.
  • Deep knowledge of Linux OS internals and kernel troubleshooting.
  • Comprehensive understanding of networking protocols and packet-level analysis, including DNS, TLS handshakes, TCP/IP, and tcpdump.

Responsibilities

  • Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams.
  • Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb).
  • Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; reduce recurring manual work through automation.
  • Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry; define and enforce SLOs/error budgets to reduce alert noise.
  • Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, CI/CD validation before production rollouts; embed operability standards into service design.

Skills

Python
Go
Bash
Ansible
Linux
Networking
tcpdump

Tools

Kubernetes
Prometheus
OpenTelemetry
Ansible
Python
Go

Job description

Zscaler seeks a Staff Site Reliability Engineer (Production Engineer) to join our Zero Trust Exchange team in a hybrid role. You will own reliability of large-scale bare-metal and cloud infrastructure, write production-grade code, and drive proactive incident management across global fleets.

You should excel at OS/network debugging, automation with Ansible/Python/Go, and shaping SRE practices with telemetry, SLOs, and CI/CD readiness to keep services stable at scale.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Site Reliability Engineer - Production (Hybrid)
Staff Site Reliability Engineer - Production (Hybrid)

Zscaler • United States

Hybrid
USD 119,000 - 170,000
Time off plans for vacation and sick
Parental leave options
Retirement options
+1
Staff Site Reliability Engineer — Production & Automation
Staff Site Reliability Engineer — Production & Automation

Zscaler • Virginia (MN)

Hybrid
USD 119,000 - 170,000
Health plans
Parental leave options
Retirement options
+2
Staff Site Reliability Engineer — AI-Driven Zero Trust
Staff Site Reliability Engineer — AI-Driven Zero Trust

Zscaler • San Jose (CA)

Hybrid
USD 119,000 - 170,000
Health plans
Vacation & sick time
Parental leave
+3
Senior Production Engineer - Automation & Global Reliability
Senior Production Engineer - Automation & Global Reliability

Zscaler • Bellevue (WA)

Hybrid
USD 140,000 - 175,000
Remote-Eligible Senior Production Engineer, Automation-Driven SRE
Remote-Eligible Senior Production Engineer, Automation-Driven SRE

Zscaler • San Jose (CA)

Hybrid
USD 104,000 - 148,000
Senior SRE: Scale Reliability Leader (Hybrid)
Senior SRE: Scale Reliability Leader (Hybrid)

Plenful • San Francisco (CA)

Hybrid
USD 180,000 - 250,000
Healthcare Coverage
401(k) with Company Match
Equity
+5
Automation & Reliability Engineer — Hybrid/Remote
Automation & Reliability Engineer — Hybrid/Remote

Zscaler • San Jose (CA)

Hybrid
USD 102,000 - 128,000
Various health plans
Time off plans for vacation and sick time
Parental leave options
+3
Senior Production SRE: Cloud & On-Prem Reliability
Senior Production SRE: Cloud & On-Prem Reliability

Weights & Biases • New York (NY)

On-site
USD 140,000 - 180,000
Medical Insurance
Dental Insurance
Vision Insurance
+15
Lead Production Engineer: Automation & Reliability (Remote)
Lead Production Engineer: Automation & Reliability (Remote)

Zscaler • San Jose (CA)

Hybrid
USD 130,000 - 170,000
Time off plans for vacation
Parental leave options
Retirement options
+1
Staff Site Reliability Engineer (Production Engineer)- Federal
Staff Site Reliability Engineer (Production Engineer)- Federal

Zscaler • San Jose (CA)

Hybrid
USD 119,000 - 170,000
Health plans
Vacation & sick time
Parental leave
+3