Sr Staff Site Reliability Engineer (Wildfire) - NetSec - Bangalore

Palo Alto Networks

Bengaluru

On-site

INR 4,000,000 - 7,000,000

Full time

13 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Palo Alto Networks is seeking a Senior Staff Site Reliability Engineer to drive reliability for WildFire analytics and related services. You will own cross‑functional resilience across cloud, private, and hybrid deployments, bridging threat analysis pipelines with high throughput sandboxing infrastructure.

You will lead on-call rotations, automate via IaC (Terraform/Ansible, ArgoCD), and optimize performance and observability to ensure sub‑second telemetry and 24/7 availability for global

Qualifications

  • 4–9 years in SRE, DevOps, platform engineering or infra roles.
  • Production workloads on cloud platforms (AWS, GCP, Azure or OCI) and hybrid/on‑prem.
  • Strong Linux internals, networking, and kernel tuning.
  • Scripting in Python or Bash for tooling and automation.
  • Kubernetes, Docker, and GitOps tooling experience.
  • IaC with Terraform and Ansible; scalable infra management.
  • Excellent communication and cross‑functional collaboration.

Responsibilities

  • Architect and scale resilient hybrid cloud infra across multi‑tenant environments.
  • Own automation and IaC pipelines to reduce toil and improve availability.
  • Define and monitor SLIs/SLOs, manage incident response and RCA.
  • Lead on-call rotations and implement post‑mortems with systemic fixes.
  • Tune performance, capacity planning, and resource allocation at scale.
  • Mentor engineers and drive SRE/DevSecOps best practices.

Skills

Experience 4-9y
SRE/DevOps
Cloud & Hybrid Systems
Linux internals
Python & Bash
Kubernetes
Terraform & Ansible
GitOps (ArgoCD)
Distributed systems
Communication & Leadership

Education

Bachelor's degree in CS/Engineering

Tools

Kubernetes
Terraform
Ansible
GitOps (ArgoCD)

Job description

Our Mission

At Palo Alto Networks®, we’re united by a shared mission—to protect our digital way of life. We thrive at the intersection of innovation and impact, solving real-world problems with cutting- edge technology and bold thinking. Here, everyone has a voice, and every idea counts. If you’re ready to do the most meaningful work of your career alongside people who are just as passionate as you are, you’re in the right place.

Who We Are

In order to be the cybersecurity partner of choice, we must trailblaze the path and shape the future of our industry. This is something our employees work at each day and is defined by our values: Disruption, Collaboration, Execution, Integrity, and Inclusion. We weave AI into the fabric of everything we do and use it to augment the impact every individual can have. If you are passionate about solving real-world problems and ideating beside the best and the brightest, we invite you to join us! We believe collaboration thrives in person. That’s why most of our teams work from the office full time, with flexibility when it’s needed. This model supports real-time problem-solving, stronger relationships, and the kind of precision that drives great outcomes.

The Team

Palo Alto Networks’ Cloud-Delivered Security Services (CDSS) is the intelligence engine of our Next-Generation Security platform. We provide a suite of AI- driven, subscription-based services—including Advanced Wildfire, Advanced Threat Prevention, DNS Security, URL Filtering, and IoT Security—integrated natively into our firewalls. Our infrastructure processes trillions of events daily, delivering real- time protection to over 85,000 global enterprises. Working in CDSS means building the backbone of global cybersecurity at a scale few companies in the world ever reach.

Job Summary

We are seeking an ambitious, technically sharp Senior Staff Site Reliability Engineer to drive the reliability, operational scalability, and infrastructure engineering for Palo Alto Networks’ WildFire malware analysis platform. In this high-impact role, you will take technical ownership of system resilience across WildFire services and appliance platforms. You will bridge the gap between threat analysis pipelines, low-level system execution, and high-throughput malware sandboxing infrastructure – ensuring sub-second detection telemetry and high availability for enterprise deployments worldwide.

Key Responsibilities
  • Hybrid & Cloud Infrastructure Resilience: Architect, scale, and maintain operational reliability across WildFire’s multi-tenant public clouds (AWS/GCP/Azure/OCI), private cloud appliances, and hybrid inspection pipelines.
  • Infrastructure Resilience: Architect, scale, and maintain the overarching operational reliability for WildFire’s cloud analysis engines, virtualized sandboxes, and distributed appliance infrastructure.
  • Autonomous Operations & IaC: Spearhead the transition to fully automated operational workflows using Terraform, Ansible, and GitOps (ArgoCD) to eliminate operational toil across multi-tenant and edge environments.
  • SLO & Error Budget Governance: Define and enforce SLIs, SLOs, and SLAs across malware inspection pipelines. Partner with security engineering leads on Error Budget strategies to balance rapid threat signature deployment with platform stability.
  • Observability & Threat Telemetry: Architect end- to- end observability stacks (Prometheus, Grafana, OpenTelemetry, Datadog/ELK) optimized for low- latency tracing, high- concurrency sample processing monitoring, and rapid MTTR.
  • Cloud & Platform Release Engineering: Build and scale enterprise CI/CD automation (GitHub Actions / GitLab CI) empowering engineering teams to safely deploy cloud microservices, threat analysis engines, and platform firmware updates.
  • On-Call & Incident Response: Lead production on- call rotations, establishing escalation paths, automated alerting, and incident response procedures to ensure 24/7 reliability for mission- critical WildFire services.
  • Incident Leadership & RCA: Lead critical incident response for high- severity platform outages. Conduct blameless post- mortems and implement systemic prevention measures across OS, container, and network layers.
  • Performance Tuning & Sandboxing Efficiency: Perform capacity modeling, Linux kernel tuning, and resource optimization for virtualized sandboxing environments and high- volume malware sample processing pipelines.
  • Capacity Planning & Performance Tuning: Perform capacity modeling, cost optimization, Linux kernel tuning, and resource allocation for cloud compute clusters and virtualized sandboxing environments.
  • Technical Mentorship: Mentor engineers across teams, championing SRE and DevSecOps best practices while conducting rigorous operational reviews.
Qualifications
Required Qualifications
  • Experience: 4 - 9 years of professional experience in SRE, DevOps, Platform Engineering, or Infrastructure Engineering roles.
  • Cloud & Hybrid Systems: Strong hands- on experience architecting and managing production workloads in major Cloud Platforms (AWS, GCP, Azure, or OCI) alongside hybrid or on- premise infrastructure.
  • Linux Systems Internals: Advanced proficiency in Linux/Unix administration, kernel mechanics, process/memory isolation, and low- level networking (TCP/IP, DNS, Load Balancing).
  • Development & Scripting: Strong scripting skills in Python and Bash for building operational tooling, custom Kubernetes operators, and platform automation.
  • Containers & Orchestration: Hands- on experience with container orchestrators (Kubernetes, K3s, Docker) and GitOps frameworks (ArgoCD, Helm).
  • Infrastructure as Code: Proven expertise with Terraform, Ansible, and declarative infrastructure management at scale.
  • Distributed Systems Debugging: Exceptional ability to dissect and troubleshoot performance bottlenecks in complex, high- throughput distributed systems.
  • Communication & Leadership: Strong written and verbal communication skills with a track record of driving cross- functional alignment across software, security, and infrastructure teams.
Preferred Qualifications
  • AI/ML SRE Tooling: Familiarity with leveraging AI/LLM tools for intelligent log analysis, automated incident routing, or predictive capacity planning.
  • Complex Systems Debugging: Proven track record of dissecting, diagnosing, and resolving critical issues within complex, multi- threaded distributed systems processing high- volume, mission- critical workloads.
  • Cloud Platform Expertise: Hands- on experience architecting or operating production environments on GCP / AWS / OCI / Azure.
  • High- Throughput Engineering: Direct experience building, tuning, or operating systems designed for high- throughput, low- latency sample processing and data pipelines.
  • Stakeholder Communication: Exceptional communication skills with the demonstrated ability to translate complex architectural trade- offs to both deep technical teams and non- technical stakeholders.
Our Commitment

We’re trailblazers that dream big, take risks, and challenge cybersecurity’s status quo. It’s simple: we can’t accomplish our mission without diverse teams innovating, together.

We are committed to providing reasonable accommodations for all qualified individuals with a disability. If you require assistance or accommodation due to a disability or special need, please contact us at accommodations@paloaltonetworks.com.

Palo Alto Networks is an equal opportunity employer. We celebrate diversity in our workplace, and all qualified applicants will receive consideration for employment without regard to age, ancestry, color, family or medical care leave, gender identity or expression, genetic information, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran status, race, religion, sex (including pregnancy), sexual orientation, or other legally protected characteristics.

All your information will be kept confidential according to EEO guidelines.

Is role eligible for Immigration Sponsorship? No. Please note that we will not sponsor applicants for work visas for this position.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Site Reliability Engineer (Wildfire) - NetSec - Bangalore
Staff Site Reliability Engineer (Wildfire) - NetSec - Bangalore

Palo Alto Networks • Bengaluru

On-site
INR 1,500,000 - 2,800,000
Staff Site Reliability Engineer (Wildfire) - NetSec - Bangalore
Staff Site Reliability Engineer (Wildfire) - NetSec - Bangalore

Palo Alto Networks, Inc. • Bengaluru

On-site
INR 3,000,000 - 4,200,000
Sr Staff Software Engineer (Wildfire) - NetSec - Bangalore
Sr Staff Software Engineer (Wildfire) - NetSec - Bangalore

Palo Alto Networks, Inc. • Bengaluru

Hybrid
INR 3,500,000 - 7,000,000
Staff DevOps Engineer - NetSec - Bangalore
Staff DevOps Engineer - NetSec - Bangalore

Palo Alto Networks • Bengaluru

Hybrid
INR 4,000,000 - 6,000,000
Staff DevOps Engineer (Cloud NGFW)- NetSec - Bangalore
Staff DevOps Engineer (Cloud NGFW)- NetSec - Bangalore

Palo Alto Networks, Inc. • Bengaluru

On-site
INR 2,600,000 - 4,200,000
Principal Engineer Software (Wildfire) - NetSec - Bangalore
Principal Engineer Software (Wildfire) - NetSec - Bangalore

Palo Alto Networks, Inc. • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Sr. Staff DevOps Engineer (SD WAN) - NetSec - Bangalore
Sr. Staff DevOps Engineer (SD WAN) - NetSec - Bangalore

Palo Alto Networks, Inc. • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Sr Staff Engineer Software (SD WAN, Controller) - NetSec
Sr Staff Engineer Software (SD WAN, Controller) - NetSec

Palo Alto Networks • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Senior Staff Software Engineer
Senior Staff Software Engineer

Palo Alto Networks • Hyderabad

On-site
INR 4,000,000 - 7,000,000
Staff MDR Analyst
Staff MDR Analyst

Palo Alto Networks • Bengaluru

On-site
INR 3,500,000 - 5,500,000