Principal SRE — Cloud Reliability & Self-Healing AI

Palo Alto Networks

California (MO)

On-site

USD 152,000 - 246,000

Full time

4 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Palo Alto Networks is seeking a Principal Site Reliability Engineer to join the ADEM team, supporting services that provide end-to-end visibility and self-healing capabilities for global customers.

You will contribute to CI/CD, IaC/Monitoring as Code, and RCA leadership, ensuring scalable, secure cloud infrastructure using Kubernetes, Terraform, GCP/AWS, and a strong focus on SLOs and reliability.

Qualifications

  • U.S. citizenship required due to federal government requirements.
  • 7+ years of infra/DevOps/system engineering experience.
  • Proficiency with AI productivity tools (e.g., Claude code, Cursor, Windsurf, GitHub Copilot).
  • Expertise in cloud-native apps on GCP (preferred) or AWS.
  • Expertise in Terraform, Helm, Ansible.
  • Strong programming in Python, Go, or Java; Kafka/Pulsar a plus.
  • Deep Kubernetes (GKE/EKS), container networking, Linux internals.
  • Experience with GitOps and GitLab CI and ArgoCD.
  • Familiarity with FedRAMP, SOC2 and policy-as-code automation.
  • BS or MS in Computer Science or equivalent experience.

Responsibilities

  • Drive SRE/DevOps success with CI/CD and AIOps initiatives.
  • Architect Golden Paths for service delivery with SLOs, error budgets, and canaries.
  • Design, build, and operate secure cloud infrastructure for high-scale monitoring.
  • Ensure production-ready, scalable, and resilient apps; collaborate with researchers and data scientists.
  • Develop IaC and Monitoring as Code tooling.
  • Lead root cause analysis of critical issues and drive preventative improvements.

Skills

AI productivity tools usage

Education

BS or MS in Computer Science

Tools

Kubernetes (GKE/EKS)
GitLab CI
ArgoCD
Terraform
Helm
Ansible
Kafka

Job description

Palo Alto Networks is seeking a Principal Site Reliability Engineer to join the ADEM team, supporting services that provide end-to-end visibility and self-healing capabilities for global customers.

You will contribute to CI/CD, IaC/Monitoring as Code, and RCA leadership, ensuring scalable, secure cloud infrastructure using Kubernetes, Terraform, GCP/AWS, and a strong focus on SLOs and reliability.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: Cloud Reliability & Automation
Senior SRE: Cloud Reliability & Automation

Palo Alto Networks, Inc. • Santa Clara (CA)

On-site
USD 152,000 - 246,000
Cloud SRE Architect — Automation & Reliability Lead
Cloud SRE Architect — Automation & Reliability Lead

Jobs Paloaltonetworks • California (MO)

On-site
USD 152,000 - 246,000
Principal Cloud SRE & Automation Engineer
Principal Cloud SRE & Automation Engineer

Palo Alto Networks, Inc. • Santa Clara (CA)

On-site
USD 151,000 - 246,000
Lead Cloud SRE & AI-Driven Infra Architect
Lead Cloud SRE & AI-Driven Infra Architect

Palo Alto Networks, Inc. • Santa Clara (CA)

On-site
USD 130,000 - 170,000
Employee benefits
Diverse workplace
Senior Principal Platform Architect - AI-Driven SRE
Senior Principal Platform Architect - AI-Driven SRE

Palo Alto Networks, Inc. • Santa Clara (CA)

On-site
USD 170,000 - 277,000
Senior SRE Manager: Reliability, Automation & Leadership
Senior SRE Manager: Reliability, Automation & Leadership

Palo Alto Networks, Inc. • Santa Clara (CA)

On-site
USD 180,000 - 260,000
Senior Principal AI-Driven Platform Architect
Senior Principal AI-Driven Platform Architect

Palo Alto Networks • Santa Clara (CA)

On-site
USD 230,000 - 340,000
Principal SRE: Hybrid Cloud Reliability & Observability
Principal SRE: Hybrid Cloud Reliability & Observability

Hewlett Packard Enterprise • San Juan (PR)

Hybrid
USD 120,000 - 160,000
AI-Driven Principal SRE — Remote & Cloud Reliability
AI-Driven Principal SRE — Remote & Cloud Reliability

UnitedHealth Group • Eden Prairie (MN)

Hybrid
Confidential
Remote work options
Principal SRE - Enterprise Reliability Leader (Hybrid)
Principal SRE - Enterprise Reliability Leader (Hybrid)

Early Warning Services LLC • San Francisco (CA)

Hybrid
USD 207,000 - 276,000
Healthcare coverage
401(k) match
Paid time off
+1