Senior SRE – Observability & Cloud Reliability

Palo Alto Networks

Santa Clara (CA)

On-site

USD 133,000 - 215,000

Full time

12 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Palo Alto Networks in Santa Clara, CA, is seeking a Senior Staff SRE on the Cortex Observability team to design, operate and scale our observability stack across GCP, with an emphasis on high reliability, performance, and scalable automation.

You will lead monitoring, incident response, auto‑remediation, and cross‑functional collaboration with software engineers, SREs, and product teams to deliver measurable improvements in uptime and customer experience for our security platform.

Qualifications

  • 5+ years of experience as a DevOps/SRE engineer with a passion for technology and a strong motivation for high reliability at the service level.
  • High proficiency with Prometheus, Grafana, Open Telemetry and other monitoring tools.
  • Clear understanding of incident and alerts management using tools like Pagerduty and Prometheus Alert Manager.
  • High proficiency in Google Cloud Platform or Amazon Web Services.
  • High proficiency with Kubernetes and Docker for container orchestration.
  • High proficiency in Python programming and Linux Shell commands. Experience with Ansible and Terraform for infrastructure as code.
  • Effective communication and interpersonal skills, with the ability to work and coordinate between multiple teams in different timezones.
  • Ability to operate independently, make decisions, take action, and take responsibility.

Responsibilities

  • Cloud expertise: Utilize your expertise in monitoring cloud platforms, particularly GCP, to optimize our infrastructure, leveraging cloud‑native technologies.
  • Monitoring expertise: Improve monitoring processes, alerts, and metrics. Work with development teams to ensure that all of our services have the right monitoring and metrics in place so that we detect problems before our customers do.
  • Incident Management: Leverage incident management processes to ensure efficient resolution of system issues and minimal impact on services.
  • Automation: Automate complex monitoring and alerting tasks by building tools for cloud operations, such as automated remediation of known issues and auto‑scaling.
  • Continuously Improve: Stay up‑to‑date with cutting‑edge technologies, evaluate their potential impact on our operations, and implement them when appropriate.
  • On‑Call: Provide follow‑the‑sun operational coverage in the production of our Observability infrastructure.
  • Collaborate: Work with our Engineering team to influence the operability of the product and ensure the reliability and availability of our services.

Skills

DevOps/SRE experience
Observability
Incident management
Cloud platforms
Python scripting
Linux shell
Communication
Troubleshooting
Independent work

Tools

Kubernetes
Docker
Ansible
Terraform

Job description

Palo Alto Networks in Santa Clara, CA, is seeking a Senior Staff SRE on the Cortex Observability team to design, operate and scale our observability stack across GCP, with an emphasis on high reliability, performance, and scalable automation.

You will lead monitoring, incident response, auto‑remediation, and cross‑functional collaboration with software engineers, SREs, and product teams to deliver measurable improvements in uptime and customer experience for our security platform.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Observability SRE – GCP, Kubernetes, Automation
Senior Observability SRE – GCP, Kubernetes, Automation

Palo Alto Networks, Inc. • Santa Clara (CA)

On-site
USD 133,000 - 215,000
Senior Staff DevOps Engineer - Cloud Reliability
Senior Staff DevOps Engineer - Cloud Reliability

Palo Alto Networks, Inc. • Santa Clara (CA)

On-site
USD 133,000 - 215,000
Senior SRE Manager: Reliability, Automation & Leadership
Senior SRE Manager: Reliability, Automation & Leadership

Palo Alto Networks, Inc. • Santa Clara (CA)

On-site
USD 180,000 - 260,000
Senior SRE: Cloud Reliability & Automation
Senior SRE: Cloud Reliability & Automation

Palo Alto Networks, Inc. • Santa Clara (CA)

On-site
USD 152,000 - 246,000
Senior Staff DevOps Engineer - Cloud Reliability Automation
Senior Staff DevOps Engineer - Cloud Reliability Automation

Jobs Paloaltonetworks • California (MO)

On-site
USD 133,000 - 215,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Palo Alto Networks, Inc. • Santa Clara (CA)

On-site
USD 133,000 - 215,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobs Paloaltonetworks • California (MO)

Hybrid
USD 133,000 - 215,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Palo Alto Networks • Santa Clara (CA)

On-site
USD 133,000 - 215,000
Lead Cloud SRE & AI-Driven Infra Architect
Lead Cloud SRE & AI-Driven Infra Architect

Palo Alto Networks, Inc. • Santa Clara (CA)

On-site
USD 130,000 - 170,000
Employee benefits
Diverse workplace
Principal SRE — Cloud Reliability & Self-Healing AI
Principal SRE — Cloud Reliability & Self-Healing AI

Palo Alto Networks • California (MO)

On-site
USD 152,000 - 246,000