Automation-First Production Engineer – Cloud Infrastructure

Zscaler

Bellevue (WA)

Hybrid

USD 102,000 - 128,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Health plans
Time off plans
Parental leave options
Retirement options
Education reimbursement
In-office perks

Job summary

Zscaler is seeking a Production Engineer for a hybrid role in San Jose, CA or Bellevue, WA. You will drive an automation-first culture, build scalable multi-cloud infrastructure, and mature observability to reduce MTTR across a global platform processing massive daily transactions.

You will lead incident response, collaborate with engineering, and define SLIs/SLOs, error budgets, and post-incident analyses to improve reliability and performance. Hybrid schedule with strong security focus.

Qualifications

  • Demonstrated curiosity and active exploration of AI tools, with a proven history of integrating new technologies to enhance daily workflows and augment problem-solving
  • 1-3 years of experience managing reliability, scalability, and availability for large-scale production services
  • Deep expertise in programming (e.g., Python, Go, or C/C++)
  • Strong background in networking protocols, Linux/RHEL systems, and distributed architecture
  • Experience in high-stakes incident management and participation in a 24/7 on-call rotation
  • Proficiency in leveraging ITIL frameworks and incident data to drive service maturity through systematic problem management and technical operability reviews

Responsibilities

  • Implement highly available, scalable infrastructure across AWS, GCP, and bare-metal environments
  • Drive an "automation-first" culture by writing code (Python/Go) to eliminate manual toil and build self-healing systems
  • Implement and maintain sophisticated observability (Prometheus, Grafana, OpenTelemetry), define SLIs/SLOs, and establish error budgets
  • Act as a lead Incident Commander (TDO on-call), develop response playbooks, and conduct deep-dive post-incident analyses
  • Partner with Engineering and partner teams to conduct operability reviews

Skills

Python
Go
C/C++
Networking
Linux/RHEL
Incident management
ITIL

Tools

Prometheus
Grafana
OpenTelemetry
Ansible
Terraform
Helm
Temporal
HAProxy

Job description

Zscaler is seeking a Production Engineer for a hybrid role in San Jose, CA or Bellevue, WA. You will drive an automation-first culture, build scalable multi-cloud infrastructure, and mature observability to reduce MTTR across a global platform processing massive daily transactions.

You will lead incident response, collaborate with engineering, and define SLIs/SLOs, error budgets, and post-incident analyses to improve reliability and performance. Hybrid schedule with strong security focus.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Production Engineer - Automation & Global Reliability
Senior Production Engineer - Automation & Global Reliability

Zscaler • Bellevue (WA)

Hybrid
USD 140,000 - 175,000
Senior Staff Production Engineer — Automation-First SRE Leader
Senior Staff Production Engineer — Automation-First SRE Leader

Socket.dev • Bellevue (WA)

Hybrid
USD 144,000 - 205,000
Senior Site Reliability Engineer — Automation & Cloud
Senior Site Reliability Engineer — Automation & Cloud

Socket.dev • Bellevue (WA)

Hybrid
USD 104,000 - 148,000
Health plans
Time off
Parental leave
+3
Remote-Eligible Senior Production Engineer, Automation-Driven SRE
Remote-Eligible Senior Production Engineer, Automation-Driven SRE

Zscaler • San Jose (CA)

Hybrid
USD 104,000 - 148,000
Automation & Reliability Engineer — Hybrid/Remote
Automation & Reliability Engineer — Hybrid/Remote

Zscaler • San Jose (CA)

Hybrid
USD 102,000 - 128,000
Various health plans
Time off plans for vacation and sick time
Parental leave options
+3
Lead Production Engineer: Automation & Reliability (Remote)
Lead Production Engineer: Automation & Reliability (Remote)

Zscaler • San Jose (CA)

Hybrid
USD 130,000 - 170,000
Time off plans for vacation
Parental leave options
Retirement options
+1
Staff Site Reliability Engineer - Production (Hybrid)
Staff Site Reliability Engineer - Production (Hybrid)

Zscaler • United States

Hybrid
USD 119,000 - 170,000
Time off plans for vacation and sick
Parental leave options
Retirement options
+1
Remote AI-Driven DevOps Engineer - Multi-Cloud
Remote AI-Driven DevOps Engineer - Multi-Cloud

Zscaler, Inc. • San Jose (CA)

Hybrid
USD 140,000 - 175,000
Time off plans
Parental leave
Retirement options
+1
Staff Production Engineer
Staff Production Engineer

Zscaler • Bellevue (WA)

Hybrid
USD 140,000 - 175,000
Staff Site Reliability Engineer — AI-Driven Zero Trust
Staff Site Reliability Engineer — AI-Driven Zero Trust

Zscaler • San Jose (CA)

Hybrid
USD 119,000 - 170,000
Health plans
Vacation & sick time
Parental leave
+3