Staff Site Reliability Engineer

Palo Alto Networks

Bengaluru

On-site

INR 1,800,000 - 2,600,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Palo Alto Networks is seeking a Staff AI-SRE to build and operate scalable, secure cloud platforms in a multi‑cloud production environment. You will drive AI‑enabled automation and observability, joining senior engineers to modernize incident response and proactive capacity planning.

You will manage Kubernetes (GKE) and Kong API Gateway, adopt AIOps, and partner with DevOps to deliver reliable, scalable infrastructure for tens of thousands of customers.

Qualifications

  • 6–10 years of experience in SRE/DevOps at scale.
  • Linux administration experience (RHEL/Ubuntu) at scale.
  • Hands-on with AIOps platforms and tools.
  • Proficiency with GCP; AWS/Azure a plus.
  • Strong scripting in Python.

Responsibilities

  • Operate and scale multi-cloud production environments with AI-driven automation.
  • Lead incident bridges for P1/P2 and post-mortems on complex stacks.
  • Design self-healing infrastructure with AI for proactive remediation.
  • Implement and maintain IaC with Terraform and Helm.

Skills

Kubernetes
GKE
Kong API Gateway
Terraform
Python
CI/CD
GitOps
GCP
AI/ML integration
Incident management

Education

Bachelor's degree in CS/CE or equivalent
CKA/CKAD/Google Cloud certs

Tools

Kubernetes (GKE)
Kong
Terraform
GitHub Actions
GitLab CI
ArgoCD
Flux
Istio (Nice to have)

Job description

Our Mission At Palo Alto Networks®, we’re united by a shared mission—to protect our digital way of life.

We thrive at the intersection of innovation and impact, solving real-world problems with cutting-edge technology and bold thinking.

Here, everyone has a voice, and every idea counts.

If you’re ready to do the most meaningful work of your career alongside people who are just as passionate as you are, you’re in the right place.

Who We Are In order to be the cybersecurity partner of choice, we must trailblaze the path and shape the future of our industry. This is something our employees work at each day and is defined by our values: Disruption, Collaboration, Execution, Integrity, and Inclusion. We weave AI into the fabric of everything we do and use it to augment the impact every individual can have.

If you are passionate about solving real-world problems and ideating beside the best and the brightest, we invite you to join us! We believe collaboration thrives in person. That’s why most of our teams work from the office full time, with flexibility when it’s needed. This model supports real-time problem-solving, stronger relationships, and the kind of precision that drives great outcomes.

Job Summary

Palo Alto Networks is looking for a cloud infrastructure and observability professional to help build and operate reliable, scalable, and secure technology platforms. You will combine expertise in cloud operations, automation, and observability to improve infrastructure performance and reliability, while exploring opportunities to apply AI and machine learning to IT operations.

Career

Join a team of senior engineers operating in a large-scale, multi-cloud production environment supporting tens of thousands of enterprise customers worldwide. This is not a typical SRE role — you'll work at the intersection of Site Reliability Engineering and AI-driven automation, pioneering the next generation of intelligent infrastructure operations.

As a Staff AI-SRE, you will be managing cloud infrastructure comprising Kubernetes (GKE) and API Gateway (Kong), while leading the adoption of AIOps, Large Language Models (LLMs), and AI‑assisted workflows to transform how we build, monitor, and operate systems. You'll work alongside experienced DevOps professionals in a fast‑paced, cybersecurity-focused organization committed to AI‑first operations.

Own and operate large‑scale, global production environments with an AI focus — leveraging AIOps and machine learning to drive autonomous operations

Pioneer AI‑driven automation: Implement LLM‑powered runbooks, AI‑assisted incident diagnostics, and intelligent alerting systems that predict issues before they impact customers

Design self‑healing infrastructure: Build systems that leverage AI for automated incident remediation, anomaly detection, and proactive service monitoring

Lead architecture, deployment, and operations of Kong API Gateway infrastructure — including rate limiting, authentication, traffic management, and plugin customization

Design, deploy, and manage production‑grade GKE clusters — cluster upgrades, node pool management, workload optimization, and multi‑tenancy

Implement predictive scalability: Use AI modeling to forecast resource needs, prevent bottlenecks, and optimize capacity planning across GKE and cloud infrastructure

Actively monitor, investigate, and resolve P1/P2 incidents using AI‑assisted diagnostics and automated playbooks

Drive end‑to‑end troubleshooting across complex, distributed systems — augmented by AI tools for faster root cause analysis

Implement and maintain Infrastructure as Code using Terraform and Helm — integrating with AI APIs for intent‑based infrastructure management

Champion “vibe coding” workflows: Leverage AI coding assistants (GitHub Copilot, Claude) to accelerate development — focusing on high‑level architecture and intent while AI handles boilerplate

Develop and maintain automation and tooling (Python, Bash, Go) with AI‑augmented development practices

Create AI‑powered operational runbooks using Generative AI for faster documentation, knowledge sharing, and incident response

Contribute to a culture of AI‑first operational excellence in a high‑scale, high‑availability environment

On‑call responsibilities: Daytime hours with occasional weekends and holidays (rotation‑based)

Qualifications
Your Experience
  • 6–10 years of experience in SRE/DevOps roles in production environments at scale
  • Linux Expert: Deep hands‑on experience managing and supporting Linux (RHEL/Ubuntu) at scale, including kernel tuning and system internals
  • AI/ML Integration Experience: Hands‑on experience or strong interest in AIOps platforms and tools
  • Strong hands‑on experience with Kubernete
  • Strong hands‑on experience with API Gateway (Kong)
  • Strong hands‑on experience with GCP (required); AWS or Azure experience is a plus
  • Mastery of Terraform for infrastructure provisioning and management
  • Familiarity with CI/CD and GitOps tools (GitLab CI, GitHub Actions, ArgoCD, Flux)
  • Proficiency in Python for scripting, automation, and AI/ML integrations
  • Proven track record leading P1/P2 incident bridges and post‑mortems for complex, global application stacks
  • Strong troubleshooting and problem‑solving skills with a passion for incident handling
  • Fluent in security, vulnerability management, and CIS compliance
  • Highly responsive, proactive, and ownership‑driven

The AI‑First Mindset: You possess the “vibe coding” spirit — leveraging AI to solve complex problems faster without losing the critical eye of a senior engineer. You're excited about transforming traditional SRE practices with intelligent automation.

Nice to Have Experience

Experience with Service Mesh (Istio, Kong Mesh)

Certifications
  • CKA, CKAD, Google Cloud Professional Experience building or integrating with AIOps platforms
Why This Role is Different

This isn't just about keeping systems running — it's about reinventing how SRE is done. You'll be at the forefront of: AI‑Assisted Incident Response: Using LLMs to accelerate diagnostics and automate remediation Predictive Operations: Moving from reactive firefighting to proactive, AI‑driven prevention Intelligent Automation: Building systems that learn, adapt, and self‑heal Next‑Gen Tooling: Shaping how AI transforms infrastructure engineering

Our Commitment

We’re trailblazers that dream big, take risks, and challenge cybersecurity’s status quo. It’s simple: we can’t accomplish our mission without diverse teams innovating, together. We are committed to providing reasonable accommodations for all qualified individuals with a disability. If you require assistance or accommodation due to a disability or special need, please contact us at accommodations@paloaltonetworks.com. Palo Alto Networks is an equal opportunity employer. We celebrate diversity in our workplace, and all qualified applicants will receive consideration for employment without regard to age, ancestry, color, family or medical care leave, gender identity or expression, genetic information, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran status, race, religion, sex (including pregnancy), sexual orientation or other legally protected characteristics. All your information will be kept confidential according to EEO guidelines.

Is role eligible for Immigration Sponsorship? No. Please note that we will not sponsor applicants for work visas for this position.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Site Reliability Engineer
Staff Site Reliability Engineer

Palo Alto Networks, Inc. • Bengaluru

On-site
INR 2,400,000 - 5,400,000
Staff Devops Engineer
Staff Devops Engineer

Palo Alto Networks • Hyderabad

On-site
INR 1,800,000 - 3,000,000
Staff DevOps Engineer
Staff DevOps Engineer

Palo Alto Networks, Inc. • Hyderabad

On-site
INR 1,800,000 - 3,200,000
Staff AI Engineer
Staff AI Engineer

Palo Alto Networks • Bengaluru

On-site
INR 3,000,000 - 6,000,000
Senior Staff Software Engineer
Senior Staff Software Engineer

Palo Alto Networks • Hyderabad

On-site
INR 1,800,000 - 2,400,000
Staff AI Engineer
Staff AI Engineer

Palo Alto Networks, Inc. • Bengaluru

On-site
INR 4,000,000 - 7,000,000
AI Engineer, Agentic Systems (Quality Engineering)
AI Engineer, Agentic Systems (Quality Engineering)

CamWebDir • Bengaluru

On-site
INR 2,500,000 - 5,000,000
Staff Professional Services Consultant, Prisma AIRS
Staff Professional Services Consultant, Prisma AIRS

Palo Alto Networks, Inc. • Karnataka

On-site
INR 3,000,000 - 5,000,000
Staff DevOps Engineer (Cortex)
Staff DevOps Engineer (Cortex)

Palo Alto Networks • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Staff DevOps Engineer (Cortex)
Staff DevOps Engineer (Cortex)

Jobs Paloaltonetworks • Bengaluru

Hybrid
INR 2,000,000 - 4,000,000