Sr Staff Site Reliability Engineer

Palo Alto Networks

United States

Remote

USD 140,000 - 200,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Palo Alto Networks is seeking an experienced Site Reliability Engineer to join our distributed and highly available engineering team. You will operate multi-cloud production environments, build robust monitoring, and drive incident response across geographies.

You will collaborate with cross-functional teams, own automation in Python, and contribute to a culture of operational excellence in a fast-paced cybersecurity-focused organization.

Qualifications

  • 5+ years in SRE/production environments at scale.
  • Hands-on with Kubernetes and Terraform in cloud ecosystems.
  • Experience building/ configuring monitoring with Prometheus & Grafana.
  • Proficiency in Python scripting and automation.

Responsibilities

  • Own and operate large-scale, multi-cloud production environments (GCP/AWS/Azure).
  • Monitor, investigate, and resolve incidents from automated alerting systems.
  • Design, deploy, and improve observability and monitoring stacks.
  • Collaborate with CX/CS/Engineering to ensure reliability and performance.
  • Develop tooling and automation in Python; participate in on-call rotations.

Skills

Collaboration
Communication
Ownership
Problem-solving
Self-management
Ambiguity handling

Education

Bachelor's degree in CS/Engineering

Tools

Kubernetes
Terraform
Prometheus
Grafana
GitLab CI
GitHub Actions
Jenkins
Flux

Job description

Our Mission

At Palo Alto Networks®, we’re united by a shared mission—to protect our digital way of life. We thrive at the intersection of innovation and impact, solving real-world problems with cutting-edge technology and bold thinking. Here, everyone has a voice, and every idea counts. If you’re ready to do the most meaningful work of your career alongside people who are just as passionate as you are, you’re in the right place.

Who We Are

In order to be the cybersecurity partner of choice, we must trailblaze the path and shape the future of our industry. This is something our employees work at each day and is defined by our values: Disruption, Collaboration, Execution, Integrity, and Inclusion. We weave AI into the fabric of everything we do and use it to augment the impact every individual can have. If you are passionate about solving real-world problems and ideating beside the best and the brightest, we invite you to join us!

This role is remote, but distance is no barrier to impact. Our hybrid teams collaborate across geographies to solve big problems, stay close to our customers, and grow together. You will be part of a culture that values trust, accountability, and shared success where your work truly matters.

Job Summary
The Team

Engineering - Our engineering team is at the core of our products and connected directly to the mission of preventing cyberattacks. We are constantly innovating - challenging the way we, and the industry, think about cybersecurity. Our engineers don’t shy away from building products to solve problems no one has pursued before. We define the industry instead of waiting for directions. We need individuals who feel comfortable in ambiguity, excited by the prospect of a challenge, and empowered by the unknown risks facing our everyday lives that are only enabled by a secure digital environment.

Your Career:

Join a team of senior engineers operating in a large-scale, multi-cloud production environment supporting tens of thousands of enterprise customers worldwide. This is not a typical SRE role - you’ll work at the core of a complex, high-impact system alongside experienced DevOps professionals in a fast-paced, cybersecurity-focused organization.

Your Impact:
  • Own and operate large-scale, global production environments across multiple cloud providers (GCP, AWS, Azure)
  • Actively monitor, investigate, and resolve incidents triggered by automated alerting systems (PagerDuty / Incident Response)
  • Drive end-to-end troubleshooting across complex, distributed systems with high context switching
  • Design, deploy, and improve monitoring and observability systems (e.g., Prometheus, Grafana) - not just react to alerts
  • Collaborate closely with internal teams (CX, CS, Engineering) to ensure system reliability and performance
  • Work hands-on with modern DevOps and infrastructure tools including Kubernetes, Terraform, CI/CD pipelines, and GitOps workflows
  • Develop and maintain automation and tooling (primarily in Python)
  • Gain deep understanding of system architecture and interconnected services
  • Contribute to a culture of operational excellence in a high-scale, high-availability environment
  • Champion asynchronous communication, documentation, and tooling standards to ensure seamless collaboration across time zones and fully distributed teams.
  • On call responsibilities:
    Daytime hours (12:00-20:00 CET/CEST, based on candidate location and team coverage needs)

Occasional weekends and holidays (rotation-based)

Your experience:
  • 5+ years of experience in SRE roles in production environments at scale
  • Strong hands-on experience with Kubernetes and Terraform
  • Strong hands-on experience with at least one major cloud platform (GCP or AWS required)
  • Experience building and configuring monitoring systems (e.g., Prometheus, Grafana)
  • Familiarity with CI/CD and GitOps tools (GitLab CI, GitHub Actions, Jenkins, Flux)
  • Proficiency in Python for scripting and automation
  • Proven success in a fully remote or distributed team environment, demonstrating strong self-management and time organization.
  • Strong troubleshooting and problem-solving skills with a passion for incident handling
  • Ability to work in fast-paced environments with high context switching
  • Highly responsive, proactive, and ownership-driven
  • Strong collaboration and communication skills
  • Curious mindset and eagerness to learn
Our Commitment

We’re trailblazers that dream big, take risks, and challenge cybersecurity’s status quo. It’s simple: we can’t accomplish our mission without diverse teams innovating, together.

We are committed to providing reasonable accommodations for all qualified individuals with a disability. If you require assistance or accommodation due to a disability or special need, please contact us at accommodations@paloaltonetworks.com.

Palo Alto Networks is an equal opportunity employer. We celebrate diversity in our workplace, and all qualified applicants will receive consideration for employment without regard to age, ancestry, color, family or medical care leave, gender identity or expression, genetic information, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran status, race, religion, sex (including pregnancy), sexual orientation, or other legally protected characteristics.

All your information will be kept confidential according to EEO guidelines.

Is role eligible for Immigration Sponsorship?: Yes

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Manager, Site Reliability Engineering
Manager, Site Reliability Engineering

Palo Alto Networks, Inc. • Santa Clara (CA)

On-site
USD 180,000 - 260,000
Sr Site Reliability Engineer (Prisma Access)
Sr Site Reliability Engineer (Prisma Access)

Palo Alto Networks • Santa Clara (CA)

On-site
USD 120,000 - 200,000
Principal Cloud Infrastructure Engineer (Advanced Threat Protection)
Principal Cloud Infrastructure Engineer (Advanced Threat Protection)

Palo Alto Networks, Inc. • Santa Clara (CA)

On-site
USD 130,000 - 170,000
Employee benefits
Diverse workplace
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Palo Alto Networks, Inc. • Santa Clara (CA)

On-site
USD 133,000 - 215,000
Principal Site Reliability Engineer Santa Clara, California, United States
Principal Site Reliability Engineer Santa Clara, California, United States

Palo Alto Networks, Inc. • Santa Clara (CA), Northern (KY)

Hybrid
USD 151,000 - 244,000
Principal SRE Engineer (US Citizen)
Principal SRE Engineer (US Citizen)

Palo Alto Networks, Inc. • Santa Clara (CA)

On-site
USD 147,000 - 237,500
Principal Site Reliability Engineer
Principal Site Reliability Engineer

CamWebDir • California (MO)

On-site
USD 133,000 - 215,000
Principal Site Reliability Engineer (ADEM)
Principal Site Reliability Engineer (ADEM)

Palo Alto Networks, Inc. • Santa Clara (CA)

On-site
USD 152,000 - 246,000
Principal Site Reliability Engineer (ADEM)
Principal Site Reliability Engineer (ADEM)

Palo Alto Networks • California (MO)

On-site
USD 152,000 - 246,000
Sr Staff Engineer Software
Sr Staff Engineer Software

Palo Alto Networks, Inc. • Santa Clara (CA)

On-site
USD 126,000 - 205,000