Principal Engineer Software

Palo Alto Networks

Santa Clara (CA)

On-site

USD 147,000 - 237,500

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salary
Comprehensive health, dental, and vision insurance
401(k) retirement plan with company match

Job summary

Palo Alto Networks is seeking a Senior Staff Data Center & OpenShift Operations Engineer to ensure high availability of our infrastructure. The role involves monitoring systems, implementing disaster recovery strategies, and managing OpenShift clusters.

Qualified candidates will have 5+ years of experience in Red Hat OpenShift, deep knowledge of GPU hardware, and strong skills in automation and scripting.

The position offers competitive compensation and comprehensive benefits including health, dental, and vision insurance.

Qualifications

  • 5+ years of experience with Red Hat OpenShift in a production environment.
  • Advanced proficiency in automation tools like Ansible.
  • Experience with GPU systems and specialized AI/ML hardware.

Responsibilities

  • Monitor and maintain high-availability infrastructure.
  • Implement automated failover strategies and backup procedures.
  • Resolve hardware and software issues in production environments.

Skills

Red Hat OpenShift (OCP) expertise
Deep experience with high-density GPU systems
Ansible or Pulumi proficiency
Strong Python skills
CoreOS and RHEL administration
Understanding of BGP and VLAN tagging

Education

Bachelor’s degree in Computer Science or IT

Tools

Prometheus
Grafana
vSphere or KVM
DCIM tools (Netbox)

Job description

Our Mission

At Palo Alto Networks®, we’re united by a shared mission—to protect our digital way of life. We thrive at the intersection of innovation and impact, solving real-world problems with cutting‑edge technology and bold thinking. Here, everyone has a voice, and every idea counts. If you’re ready to do the most meaningful work of your career alongside people who are just as passionate as you are, you’re in the right place.

Who We Are

In order to be the cybersecurity partner of choice, we must trailblaze the path and shape the future of our industry. This is something our employees work at each day and is defined by our values: Disruption, Collaboration, Execution, Integrity, and Inclusion. We weave AI into the fabric of everything we do and use it to augment the impact every individual can have. If you are passionate about solving real-world problems and ideating beside the best and the brightest, we invite you to join us!

We believe collaboration thrives in person. That’s why most of our teams work from the office full time, with flexibility when it’s needed. This model supports real‑time problem‑solving, stronger relationships, and the kind of precision that drives great outcomes.

Job Summary
Sr Staff Data Center & OpenShift Operations Engineer
Position Overview

The Senior Data Center Operations Engineer is responsible for the bedrock of our high‑availability infrastructure. This role bridges the gap between physical hardware and the Red Hat OpenShift Container Platform (OCP). Your mission is to ensure 99.99% availability by architecting resilient physical layouts and automating the deployment, scaling, and self‑healing capabilities of our production clusters.

Key Responsibilities
  • High‑Availability (HA) Infrastructure: Monitor and maintain data center systems with a focus on “Zero Single Point of Failure” (ZSPoF) architecture for OpenShift control planes and worker nodes.
  • Cluster Reliability Engineering: Implement and manage OpenShift 4.x clusters across multiple power and cooling zones to ensure 99.99% uptime.
  • Disaster Recovery & Business Continuity: Design, test, and execute automated failover strategies and backup/restore procedures using tools like OADP (Velero) and Red Hat ACM.
  • Automated Maintenance: Perform routine maintenance and upgrades using GitOps (ArgoCD) and the Machine Config Operator to ensure zero‑downtime node evacuations and patching.
  • Complex Troubleshooting: Resolve deep‑stack hardware and software issues, from faulty GPU firmware to OpenShift SDN (OVN‑Kubernetes) network latencies.
  • Vendor & Lifecycle Management: Coordinate with vendors for specialized hardware (e.g., NVIDIA, Dell, Cisco) while maintaining strict security and firmware compliance.
  • Efficiency & Capacity Architecture: Optimize rack density for high‑performance GPU clusters while managing thermal loads and power distribution (PDU) to prevent circuit‑trip outages.
  • Observability Implementation: Maintain accurate documentation and integrate hardware health metrics (IPMI/SNMP) into Prometheus/Grafana for proactive alerting.
  • Physical Deployment: Rack and stack high‑density GPU servers, ensuring redundant power‑pathing and high‑speed (100G/200G) InfiniBand or Ethernet cabling.
  • Hardware Lifecycle: Perform precision physical installation and replacement of critical components (CPUs, GPUs, NVMe storage) in a live production environment without impacting cluster quorum.
Qualifications
  • Bachelor’s degree in Computer Science, IT, or equivalent experience.
  • 5+ years of experience specifically operating Red Hat OpenShift (OCP) in a production environment.
  • Deep experience racking/stacking and cabling high‑density GPU systems (e.g., NVIDIA DGX or similar) and specialized AI/ML hardware.
  • Advanced proficiency in Ansible or Pulumi for automating bare‑metal provisioning and cluster configuration.
  • Strong Python and Bash skills for developing custom health‑check scripts and API integrations.
  • Expert‑level CoreOS and RHEL administration, including kernel tuning and systemd management.
  • Solid understanding of BGP, VLAN tagging, LACP, and Load Balancing (F5/NGINX) essential for cluster ingress.
  • Experience with vSphere or KVM, and persistent storage solutions like OpenShift Data Foundation (ODF) or Ceph.
  • Familiarity with DCIM tools (Netbox) and monitoring stacks (ELK, Loki, etc.).
Physical Requirements
  • Lifting: Ability to lift and move equipment up to 50 pounds (e.g., high‑density 2U/4U servers).
  • Environment: Comfortable working in high‑decibel, climate-controlled data center aisles.
  • Dexterity: Capable of standing, walking, and performing precision cabling in tight rack spaces for extended periods.
  • Travel: May require occasional travel to remote data center sites or edge locations.
Benefits
  • Competitive salary commensurate with high‑availability expertise.
  • Comprehensive health, dental, and vision insurance.
  • 401(k) retirement plan with company match.
Compensation Disclosure

The compensation offered for this position will depend on qualifications, experience, and work location. The offered compensation may also include restricted stock units and a bonus. $147,000.00 – $237,500.00/yr

Our Commitment

We’re trailblazers that dream big, take risks, and challenge cybersecurity’s status quo. It’s simple: we can’t accomplish our mission without diverse teams innovating, together.

We are committed to providing reasonable accommodations for all qualified individuals with a disability. If you require assistance or accommodation due to a disability or special need, please contact us at accommodations@paloaltonetworks.com.

Palo Alto Networks is an equal opportunity employer. We celebrate diversity in our workplace, and all qualified applicants will receive consideration for employment without regard to age, ancestry, color, family or medical care leave, gender identity or expression, genetic information, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran status, race, religion, sex (including pregnancy), sexual orientation, or other legally protected characteristics.

All your information will be kept confidential according to EEO guidelines.

Is role eligible for Immigration Sponsorship? No.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Engineer Software
Principal Engineer Software

Palo Alto Networks, Inc. • Santa Clara (CA)

On-site
USD 147,000 - 238,000
Comprehensive health insurance
401(k) retirement plan
Competitive salary
Principal Cloud Infrastructure Engineer (Advanced Threat Protection)
Principal Cloud Infrastructure Engineer (Advanced Threat Protection)

Palo Alto Networks • United States

On-site
USD 150,000 - 230,000
DevOps Architect, Remote, Eastern or Central TimeZone
DevOps Architect, Remote, Eastern or Central TimeZone

Palo Alto Networks • Massachusetts

On-site
USD 173,000 - 280,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Palo Alto Networks • United States

Hybrid
USD 152,000 - 245,000
DevOps Architect Massachusetts, United States
DevOps Architect Massachusetts, United States

Palo Alto Networks, Inc. • Massachusetts

On-site
USD 173,000 - 279,500
Manager, Site Reliability Engineering
Manager, Site Reliability Engineering

Palo Alto Networks, Inc. • Santa Clara (CA)

On-site
USD 180,000 - 260,000
Manager, Site Reliability Engineering
Manager, Site Reliability Engineering

Socket.dev • California (MO)

On-site
USD 150,000 - 230,000
Sr. Principal Software Engineer (L7 Security)
Sr. Principal Software Engineer (L7 Security)

Palo Alto Networks, Inc. • San Francisco (CA)

On-site
USD 170,000 - 277,000
Equity and stock options
Comprehensive health benefits
Work-life balance initiatives
Manager, Site Reliability Engineering
Manager, Site Reliability Engineering

Palo Alto Networks • United States

On-site
USD 180,000 - 240,000
Principal DevOps Engineer (US Citizen)
Principal DevOps Engineer (US Citizen)

Palo Alto Networks, Inc. • Santa Clara (CA)

Hybrid
USD 147,000 - 238,000