Principal Site Reliability Engineer

Zerto

Puerto Rico

Hybrid

Confidential

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Hybrid work model
San Juan office

Job summary

Hewlett Packard Enterprise seeks a Principal Site Reliability Engineer to design, build, and optimize cloud infrastructure, ensuring scalability, security, and reliability across platforms. You will lead enhancements to CI/CD, monitoring, and security programs while collaborating with development and security teams.

The role requires strong Linux, cloud, and IaC skills, plus experience with Kafka/Cassandra. This is a hybrid role with San Juan office attendance twice weekly.

Qualifications

  • 10+ years hands-on experience in Infra Ops, DevOps, or Site Reliability Engineering (SRE).
  • Proficiency with Linux systems, especially Debian-based distributions.
  • Strong experience with cloud platforms such as AWS and GCP.
  • Expertise in Infrastructure as Code tools like Terraform, Packer, and Ansible.
  • Solid programming skills in Python and/or Golang.
  • Deep understanding of containerization (Docker, Container) and orchestration tools (AWS EKS, GCP GKE).
  • Experience with GitOps workflows.
  • Proven track record in implementing and maintaining CI/CD pipelines.
  • Strong background in security and familiarity with security programs.
  • Experience with monitoring and logging tools (Prometheus, Grafana, ELK).
  • Knowledge of both relational (SQL) and non-relational databases.
  • Excellent problem-solving and debugging skills with a strong sense of ownership.
  • Experience managing distributed systems like Apache Kafka and Cassandra.
  • Effective communicator and collaborative team player.
  • It is mandatory to attend to San Juan office twice a week.

Responsibilities

  • Design, build, and optimize cloud infrastructure.
  • Improve CI/CD pipelines with FluxCD and Jenkins.
  • Address container image vulnerabilities and remediation.
  • Build AMIs aligned with CIS and STIG.
  • Strengthen monitoring with Prometheus, Grafana, ELK.
  • Troubleshoot production issues for reliability.
  • Collaborate with dev, security, and ops teams.
  • Enhance IAC and enforce best practices.

Skills

Infra Ops
DevOps
SRE
Linux Debian
AWS
GCP
Terraform
Packer
Ansible
Python
Golang
Docker
Kubernetes
GitOps
CI/CD
Security
Prometheus
Grafana
ELK
SQL/NoSQL

Tools

AWS EKS
GCP GKE
Terraform
Packer
Ansible
FluxCD
Jenkins

Job description

Principal Site Reliability Engineer

This role has been designed as ‘Hybrid’ with an expectation that you will work on average 2 days per week from an HPE office.

Who We Are:

Hewlett Packard Enterprise is the global edge-to-cloud company advancing the way people live and work. We help companies connect, protect, analyze, and act on their data and applications wherever they live, from edge to cloud, so they can turn insights into outcomes at the speed required to thrive in today’s complex world. Our culture thrives on finding new and better ways to accelerate what’s next. We know varied backgrounds are valued and succeed here. We have the flexibility to manage our work and personal needs. We make bold moves, together, and are a force for good. If you are looking to stretch and grow your career our culture will embrace you. Open up opportunities with HPE.

Job Description:
Job Family Definition:

Designs, develops, troubleshoots and debugs software programs for software enhancements and new products. Develops software including operating systems, compilers, routers, networks, utilities, databases and Internet-related tools. Determines hardware compatibility and/or influences hardware design.

Management Level Definition:

Contributions impact technical components of HPE products, solutions, or services regularly and sustainable. Applies advanced subject matter knowledge to solve complex business issues and is regarded as a subject matter expert. Provides expertise and partnership to functional and technical project teams and may participate in cross-functional initiatives. Exercises significant independent judgment to determine best method for achieving objectives. May provide team leadership and mentoring to others.

In a typical day as a Principal Site Reliability Engineer, you would...

As a Principal Site Reliability Engineer, you will play a key role in designing, building, and optimizing cloud infrastructure and deployment systems. Your work will directly impact scalability, security, and operational efficiency across our platforms. Key responsibilities include:

  • Enhance Infrastructure as Code (IAC) and enforce best practices.
  • Optimize cloud infrastructure for scalability, security, and cost-effectiveness.
  • Develop internal tools to support and streamline cloud platform operations.
  • Improve CI/CD pipelines and deployment workflows using FluxCD and Jenkins.
  • Address container image vulnerabilities and standardize remediation processes.
  • Build Amazon Machine Images (AMIs) aligned with CIS and STIG benchmarks.
  • Strengthen monitoring, alerting, and observability using Prometheus, Grafana, and logging tools.
  • Troubleshoot complex production issues to ensure system reliability and customer satisfaction.
  • Fine-tune distributed systems such as Apache Kafka and Cassandra.
  • Collaborate with development, security, and operations teams to align infrastructure with application needs.
What you need to bring:
  • Minimum of 10 years of hands-on experience in Infra Ops, Dev Ops, or Site Reliability Engineering (SRE).
  • Proficiency with Linux systems, especially Debian-based distributions.
  • Strong experience with cloud platforms such as AWS and GCP.
  • Expertise in Infrastructure as Code tools like Terraform, Packer, and Ansible.
  • Solid programming skills in Python and/or Golang.
  • Deep understanding of containerization (Docker, Container) and orchestration tools (AWS EKS, GCP GKE).
  • Experience with GitOps workflows.
  • Proven track record in implementing and maintaining CI/CD pipelines.
  • Strong background in security and familiarity with security programs.
  • Experience with monitoring and logging tools (Prometheus, Grafana, ELK).
  • Knowledge of both relational (SQL) and non-relational databases.
  • Excellent problem-solving and debugging skills with a strong sense of ownership.
  • Experience managing distributed systems like Apache Kafka and Cassandra.
  • Effective communicator and collaborative team player.
  • It is mandatory to attend to San Juan office twice a week.
Preferred Qualifications
  • Experience contributing to open-source projects.
  • Background in security engineering or related disciplines.
What We Can Offer You:
Health & Wellbeing

We strive to provide our team members and their loved ones with a comprehensive suite of benefits that supports their physical, financial and emotional wellbeing.

Personal & Professional Development

We also invest in your career because the better you are, the better we all are. We have specific programs catered to helping you reach any career goals you have — whether you want to become a knowledge expert in your field or apply your skills to another division.

Unconditional Inclusion

We are unconditionally inclusive in the way we work and celebrate individual uniqueness. We know varied backgrounds are valued and succeed here. We have the flexibility to manage our work and personal needs. We make bold moves, together, and are a force for good.

Let's Stay Connected:

Follow @HPECareers on Instagram to see the latest on people, culture and tech at HPE.

#puertorico#networking

Job:

Engineering

Job Level:

TCP_05

Hewlett Packard Enterprise is EEO Protected Veteran/ Individual with Disabilities.

HPE will comply with all applicable laws related to employer use of arrest and conviction records, including laws requiring employers to consider for employment qualified applicants with criminal histories.

HPE is an Equal Employment Opportunity/ Veterans/Disabled/LGBT employer. We do not discriminate on the basis of race, gender, or any other protected category, and all decisions we make are made on the basis of qualifications, merit, and business need. Our goal is to be one global team that is representative of our customers, in an inclusive environment where we can continue to innovate and grow together. Please click here: Equal Employment Opportunity.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal Site Reliability Engineer
Principal Site Reliability Engineer

Hewlett Packard Enterprise • San Juan (PR)

On-site
USD 120,000 - 160,000
Comprehensive benefits suite
Personal & professional development opportunities
Unconditional inclusion in the workplace
Site Reliability Engineer Staff
Site Reliability Engineer Staff

Zerto • Puerto Rico

Hybrid
Confidential
Site Reliability Engineer Sr. Staff
Site Reliability Engineer Sr. Staff

Zerto • Puerto Rico

Hybrid
Confidential
DevOps Engineer – Generative AI & Enterprise Web Platforms
DevOps Engineer – Generative AI & Enterprise Web Platforms

Hewlett Packard Enterprise • San Juan (PR)

Hybrid
USD 120,000 - 150,000
DevOps Engineer - Generative AI & Enterprise Web Platforms
DevOps Engineer - Generative AI & Enterprise Web Platforms

Hewlett Packard Enterprise Company in • San Juan (PR)

Hybrid
USD 100,000 - 150,000
Systems/Software Engineer
Systems/Software Engineer

Hewlett Packard Enterprise • Sunnyvale (CA)

On-site
USD 106,000 - 214,000
Principal Sustain Engineer - PCE (Houston, TX)
Principal Sustain Engineer - PCE (Houston, TX)

Hewlett Packard Enterprise Company • Spring (TX)

Hybrid
USD 152,000 - 349,000
Hybrid work model
AI Ops Engineer
AI Ops Engineer

Hewlett Packard Enterprise • San Juan (PR)

On-site
USD 90,000 - 150,000
Hybrid work model (2 days in-office)
Relocation support
Senior Software Developer Cloud & Distributed Systems
Senior Software Developer Cloud & Distributed Systems

Hewlett Packard Enterprise Company in • San Juan (PR)

On-site
USD 120,000 - 170,000
DevOps & Tooling Manager
DevOps & Tooling Manager

Hewlett Packard Enterprise Company • San Juan (PR)

Hybrid
USD 120,000 - 180,000
Relocation support