Principal SRE: Hybrid Cloud Reliability & Observability

Hewlett Packard Enterprise

San Juan (PR)

Hybrid

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Comprehensive benefits suite
Personal & professional development opportunities
Unconditional inclusion in the workplace

Job summary

A leading technology company is seeking a Principal Site Reliability Engineer to optimize and manage cloud infrastructure. The role requires over 10 years of experience in Infra Ops or SRE, with proficiency in cloud platforms like AWS and GCP. Key responsibilities include enhancing Infrastructure as Code, improving CI/CD pipelines, and troubleshooting production issues. This position operates within a hybrid model, requiring attendance at the San Juan office twice a week. Join a forward-thinking team focused on operational excellence.

Qualifications

  • Minimum of 10 years of hands-on experience in Infra Ops, Dev Ops, or Site Reliability Engineering (SRE).
  • Proficiency with Linux systems, especially Debian-based distributions.
  • Strong experience with cloud platforms such as AWS and GCP.
  • Expertise in Infrastructure as Code tools like Terraform, Packer, and Ansible.
  • Solid programming skills in Python and/or Golang.
  • Deep understanding of containerization (Docker) and orchestration tools (AWS EKS, GCP GKE).
  • Experience with GitOps workflows.
  • Proven track record in implementing and maintaining CI/CD pipelines.

Responsibilities

  • Enhance Infrastructure as Code (IAC) and enforce best practices.
  • Optimize cloud infrastructure for scalability, security, and cost-effectiveness.
  • Develop internal tools to support and streamline cloud platform operations.
  • Improve CI/CD pipelines and deployment workflows using FluxCD and Jenkins.
  • Address container image vulnerabilities and standardize remediation processes.
  • Build Amazon Machine Images (AMIs) aligned with CIS and STIG benchmarks.
  • Strengthen monitoring, alerting, and observability using Prometheus and Grafana.
  • Troubleshoot complex production issues to ensure system reliability.

Skills

Infrastructure as Code (IAC)
Cloud platforms such as AWS and GCP
Linux systems proficiency
Containerization and orchestration tools
Python and/or Golang programming
GitOps workflows
CI/CD pipeline implementation
Monitoring tools (Prometheus, Grafana)

Education

10+ years experience in Infra Ops or SRE

Tools

Terraform
Ansible
Docker
Prometheus
Grafana
Apache Kafka
Cassandra

Job description

A leading technology company is seeking a Principal Site Reliability Engineer to optimize and manage cloud infrastructure. The role requires over 10 years of experience in Infra Ops or SRE, with proficiency in cloud platforms like AWS and GCP. Key responsibilities include enhancing Infrastructure as Code, improving CI/CD pipelines, and troubleshooting production issues. This position operates within a hybrid model, requiring attendance at the San Juan office twice a week. Join a forward-thinking team focused on operational excellence.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal SRE: Infrastructure Observability & Reliability
Principal SRE: Infrastructure Observability & Reliability

T. Rowe Price • Washington

Hybrid
USD 175,000 - 299,000
Hybrid work schedule
Competitive compensation
Annual bonus eligibility
+4
Principal SRE: Hybrid Cloud-Native Reliability Leader
Principal SRE: Hybrid Cloud-Native Reliability Leader

ViziRecruiter,LLC. • Salisbury (NC)

Hybrid
USD 146,000 - 221,000
Principal SRE — Cloud Platform & Reliability Lead (Hybrid)
Principal SRE — Cloud Platform & Reliability Lead (Hybrid)

ShipperHQ • Austin (TX)

Hybrid
USD 140,000 - 200,000
22 days PTO
401k Match
Medical, Dental, Vision Insurance
+3
Senior Site Reliability Engineer - Build SRE Ops (Hybrid)
Senior Site Reliability Engineer - Build SRE Ops (Hybrid)

Mission Staffing • New York (NY)

Hybrid
USD 140,000 - 200,000
Principal SRE & Reliability Architect (Hybrid)
Principal SRE & Reliability Architect (Hybrid)

Milwaukee Succeeds • Milwaukee (WI)

Hybrid
USD 120,000 - 150,000
Opportunities to innovate in fintech
Tools for personal and professional growth
Inclusive and diverse work environment
+1
Hybrid SRE: Cloud Infrastructure & Automation
Hybrid SRE: Cloud Infrastructure & Automation

TikTok • Seattle (WA)

Hybrid
USD 112,000 - 178,000
Principal SRE: Cloud Identity & Trust
Principal SRE: Cloud Identity & Trust

5 Star Recruitment • San Jose (CA)

On-site
USD 120,000 - 150,000
Lead Principal SRE: Cloud Reliability & DevEx
Lead Principal SRE: Cloud Reliability & DevEx

iSpot.tv • Bellevue (WA)

Hybrid
USD 163,000 - 213,000
Equity options
Flexible work hours
Comprehensive benefits package
Principal Site Reliability Engineer – Hybrid Multi‑Cloud
Principal Site Reliability Engineer – Hybrid Multi‑Cloud

PowerPlan, Inc. • Smyrna (GA)

Hybrid
USD 140,000 - 200,000
Hybrid work model
Onsite office in Smyrna, GA
Senior SRE: Architect Scalable, Reliable Cloud Infra
Senior SRE: Architect Scalable, Reliable Cloud Infra

The ReWork Group • New York (NY)

On-site
USD 120,000 - 160,000