Senior Observability SRE – GCP, Kubernetes, Automation

Palo Alto Networks, Inc.

Santa Clara (CA)

On-site

USD 133,000 - 215,000

Full time

13 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Palo Alto Networks, in Santa Clara, CA, seeks a Senior Staff SRE for the Cortex Observability team to operate and optimize a large-scale GCP observability stack and deliver reliable services. You will drive incident response, automation, and scalable monitoring across our Cortex ecosystem.

You will collaborate with engineering across time zones, continuously improve observability practices, and maintain high availability for critical security platforms.

Qualifications

  • 5+ years of DevOps/SRE experience with a passion for technology and a strong motivation for high reliability at the service level.
  • High proficiency with Prometheus, Grafana, Open Telemetry and other monitoring tools.
  • Clear understanding of incident and alerts management using tools like Pagerduty and Prometheus Alert Manager.
  • High proficiency in either Google Cloud Platform or Amazon Web Services.
  • High proficiency with Kubernetes and Docker for container orchestration.
  • High proficiency in Python programming and Linux Shell commands. Experience with Ansible and Terraform for infrastructure as code.
  • Effective communication and interpersonal skills, with the ability to work and coordinate between multiple teams in different timezones.
  • Ability to effectively troubleshoot and address emerging and complex problems.
  • Ability to operate independently, make decisions, take action, and take responsibility.

Responsibilities

  • Cloud Expertise: Utilize your expertise in monitoring cloud platforms, particularly GCP, to optimize our infrastructure, leveraging cloud-native technologies.
  • Monitoring Expertise: Improve monitoring processes, alerts, and metrics. Work with development teams to ensure that all of our services have the right monitoring and metrics in place so that we detect problems before our customers do.
  • Incident Management: Leverage incident management processes to ensure efficient resolution of system issues and minimal impact on services.
  • Automation: Automate complex monitoring and alerting tasks by building tools for cloud operations, such as automated remediation of known issues and auto-scaling.
  • Continuously Improve: Stay up-to-date with cutting-edge technologies, evaluate their potential impact on our operations, and implement them when appropriate.
  • On-Call: Provide follow-the-sun operational coverage in the production of our Observability infrastructure..
  • Collaborate: Work with our Engineering team to influence the operability of the product and ensure the reliability and availability of our services.

Skills

DevOps/SRE Expertise
Observability Tools
Incident and Alerts Management
Cloud Proficiency
Kubernetes and Docker
Scripting and Automation
Communication Skills
Troubleshooting
Independence

Tools

Prometheus
Grafana
OpenTelemetry

Job description

Palo Alto Networks, in Santa Clara, CA, seeks a Senior Staff SRE for the Cortex Observability team to operate and optimize a large-scale GCP observability stack and deliver reliable services. You will drive incident response, automation, and scalable monitoring across our Cortex ecosystem.

You will collaborate with engineering across time zones, continuously improve observability practices, and maintain high availability for critical security platforms.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE – Observability & Cloud Reliability
Senior SRE – Observability & Cloud Reliability

Palo Alto Networks • Santa Clara (CA)

On-site
USD 133,000 - 215,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Palo Alto Networks, Inc. • Santa Clara (CA)

On-site
USD 133,000 - 215,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobs Paloaltonetworks • California (MO)

Hybrid
USD 133,000 - 215,000
Senior Staff DevOps Engineer - Cloud Reliability
Senior Staff DevOps Engineer - Cloud Reliability

Palo Alto Networks, Inc. • Santa Clara (CA)

On-site
USD 133,000 - 215,000
Senior SRE Manager: Reliability, Automation & Leadership
Senior SRE Manager: Reliability, Automation & Leadership

Palo Alto Networks, Inc. • Santa Clara (CA)

On-site
USD 180,000 - 260,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Palo Alto Networks • Santa Clara (CA)

On-site
USD 133,000 - 215,000
Cloud & SRE Principal — Cortex Platform Architect
Cloud & SRE Principal — Cortex Platform Architect

Palo Alto Networks • California (MO)

Hybrid
USD 147,000 - 238,000
Senior Staff DevOps Engineer - Cloud Reliability Automation
Senior Staff DevOps Engineer - Cloud Reliability Automation

Jobs Paloaltonetworks • California (MO)

On-site
USD 133,000 - 215,000
Senior DevOps Architect, Global Cloud Platform
Senior DevOps Architect, Global Cloud Platform

Jobs Paloaltonetworks • California (MO)

Hybrid
USD 147,000 - 238,000
Senior Observability & SRE Engineer — GCP/Kubernetes
Senior Observability & SRE Engineer — GCP/Kubernetes

Ontrac Solutions • United States

On-site
USD 120,000 - 180,000