Senior SRE, Observability & Cloud Reliability

Palo Alto Networks, Inc.

Santa Clara, Northern (CA, KY)

Hybrid

USD 151,000 - 244,000

Full time

5 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Palo Alto Networks is seeking a Senior Staff SRE for the Cortex Observability team to design, operate and scale a large GCP-based observability platform. You’ll implement monitoring, tracing, and logging at scale and partner with engineering to improve reliability and performance across services.

Your responsibilities include incident management, automation of remediation and alerts, follow-the-sun coverage, and staying current with cloud-native tooling to drive reliability.

Qualifications

  • DevOps/SRE Expertise: 5+ years
  • Observability Tools: Prometheus, Grafana, OpenTelemetry
  • Incident and Alerts Management: Pagerduty, Prometheus Alert Manager
  • Cloud Proficiency: GCP or AWS
  • Kubernetes and Docker: container orchestration
  • Scripting and Automation: Python and Linux shell; Ansible and Terraform

Responsibilities

  • Cloud expertise to optimize infrastructure on GCP.
  • Improve monitoring, alerts, and metrics across services.
  • Incident management processes for efficient resolution.
  • Automate complex monitoring tasks and remediation.
  • Stay up-to-date with cloud-native tech and apply when useful.
  • Provide follow-the-sun operational coverage for Observability infra.
  • Collaborate with Engineering to improve product operability and reliability.

Skills

DevOps/SRE
Troubleshooting
Communication
Independence
Team coordination

Tools

Prometheus
Grafana
OpenTelemetry
Pagerduty
Terraform
Ansible
Docker
Kubernetes
Python
Linux shell

Job description

Palo Alto Networks is seeking a Senior Staff SRE for the Cortex Observability team to design, operate and scale a large GCP-based observability platform. You’ll implement monitoring, tracing, and logging at scale and partner with engineering to improve reliability and performance across services.

Your responsibilities include incident management, automation of remediation and alerts, follow-the-sun coverage, and staying current with cloud-native tooling to drive reliability.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE, Observability & Cloud Reliability
Senior SRE, Observability & Cloud Reliability

CamWebDir • California (MO)

On-site
USD 133,000 - 215,000
Senior Observability SRE – GCP, Kubernetes, Automation
Senior Observability SRE – GCP, Kubernetes, Automation

Palo Alto Networks, Inc. • Santa Clara (CA)

On-site
USD 133,000 - 215,000
Senior SRE Manager: Reliability, Automation & Leadership
Senior SRE Manager: Reliability, Automation & Leadership

Palo Alto Networks, Inc. • Santa Clara (CA)

On-site
USD 180,000 - 260,000
Cloud & SRE Principal — Cortex Platform Architect
Cloud & SRE Principal — Cortex Platform Architect

Palo Alto Networks • California (MO)

Hybrid
USD 147,000 - 238,000
Senior Staff DevOps Engineer: Cloud Reliability
Senior Staff DevOps Engineer: Cloud Reliability

Palo Alto Networks • California (MO)

On-site
USD 133,000 - 215,000
Senior Staff DevOps Engineer: Scale Secure Cloud Platforms
Senior Staff DevOps Engineer: Scale Secure Cloud Platforms

Palo Alto Networks • Santa Clara (CA)

On-site
USD 133,000 - 215,000
Senior DevOps Architect – Global GCP, IaC & SRE
Senior DevOps Architect – Global GCP, IaC & SRE

Palo Alto Networks • Santa Clara (CA)

On-site
USD 147,000 - 238,000
Senior SRE: Cloud Reliability & Automation
Senior SRE: Cloud Reliability & Automation

Palo Alto Networks, Inc. • Santa Clara (CA)

On-site
USD 152,000 - 246,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Palo Alto Networks, Inc. • Santa Clara (CA)

On-site
USD 133,000 - 215,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

CamWebDir • California (MO)

On-site
USD 133,000 - 215,000