Senior SRE Manager: Reliability, Automation & Leadership

Palo Alto Networks, Inc.

Santa Clara (CA)

On-site

USD 180,000 - 260,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Palo Alto Networks, Inc. is seeking an experienced Site Reliability Engineer Manager to lead a team responsible for the reliability of Cortex services and infrastructure. You will provide technical direction, mentorship, and ensure operational health across highly available systems.

You will collaborate with engineering teams to improve architecture, implement monitoring and automation, and drive incident response excellence while shaping a culture of ownership and continuous improvement.

Qualifications

  • 10+ years of experience in SRE/DevOps or related, with leadership responsibilities.
  • Strong knowledge of cloud platforms, particularly GCP.
  • Experience with large-scale distributed systems and reliability engineering practices.
  • Proven ability to lead engineers and manage multiple initiatives across teams.
  • Excellent cross-functional collaboration and communication skills.

Responsibilities

  • Lead, mentor, and develop a team of Site Reliability/Production Engineers, providing technical direction and career growth.
  • Own the reliability, availability, and operational health of Cortex services and infrastructure.
  • Drive improvements in monitoring, alerting, incident management, and observability.
  • Lead major production incidents, perform root cause analysis, and drive corrective actions.
  • Partner with Engineering teams to improve service architecture, production readiness, scalability, and operability.
  • Drive automation and self-healing solutions to reduce operational toil and improve efficiency.
  • Establish priorities, processes, and measurable reliability goals for the team.
  • Collaborate with Production Engineering teams across regions for follow-the-sun operations.
  • Evaluate new technologies and drive adoption to improve reliability.

Skills

SRE leadership
Cloud platforms (GCP)
Kubernetes
Incident management
Automation & IaC
Communication

Tools

Prometheus
Grafana
OpenTelemetry
Terraform
Ansible
GitOps tooling
Python scripting

Job description

Palo Alto Networks, Inc. is seeking an experienced Site Reliability Engineer Manager to lead a team responsible for the reliability of Cortex services and infrastructure. You will provide technical direction, mentorship, and ensure operational health across highly available systems.

You will collaborate with engineering teams to improve architecture, implement monitoring and automation, and drive incident response excellence while shaping a culture of ownership and continuous improvement.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE – Observability & Cloud Reliability
Senior SRE – Observability & Cloud Reliability

Palo Alto Networks • Santa Clara (CA)

On-site
USD 133,000 - 215,000
Senior Staff DevOps Engineer - Cloud Reliability Automation
Senior Staff DevOps Engineer - Cloud Reliability Automation

Jobs Paloaltonetworks • California (MO)

On-site
USD 133,000 - 215,000
Senior Staff DevOps Engineer - Cloud Reliability
Senior Staff DevOps Engineer - Cloud Reliability

Palo Alto Networks, Inc. • Santa Clara (CA)

On-site
USD 133,000 - 215,000
Principal Cloud SRE & Automation Engineer
Principal Cloud SRE & Automation Engineer

Palo Alto Networks, Inc. • Santa Clara (CA)

On-site
USD 151,000 - 246,000
Senior Observability SRE – GCP, Kubernetes, Automation
Senior Observability SRE – GCP, Kubernetes, Automation

Palo Alto Networks, Inc. • Santa Clara (CA)

On-site
USD 133,000 - 215,000
Lead Cloud SRE & AI-Driven Infra Architect
Lead Cloud SRE & AI-Driven Infra Architect

Palo Alto Networks, Inc. • Santa Clara (CA)

On-site
USD 130,000 - 170,000
Employee benefits
Diverse workplace
Cloud SRE Architect — Automation & Reliability Lead
Cloud SRE Architect — Automation & Reliability Lead

Jobs Paloaltonetworks • California (MO)

On-site
USD 152,000 - 246,000
Principal SRE — Cloud Reliability & Self-Healing AI
Principal SRE — Cloud Reliability & Self-Healing AI

Palo Alto Networks • California (MO)

On-site
USD 152,000 - 246,000
Manager, Site Reliability Engineering
Manager, Site Reliability Engineering

Palo Alto Networks, Inc. • Santa Clara (CA)

On-site
USD 180,000 - 260,000
Senior Principal AI-Driven Platform Architect
Senior Principal AI-Driven Platform Architect

Palo Alto Networks • Santa Clara (CA)

On-site
USD 230,000 - 340,000