Principal SRE: Hybrid Cloud Reliability & Observability
Hewlett Packard Enterprise
San Juan (PR)
Hybrid
USD 120,000 - 160,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Benefits offered by this job
Comprehensive benefits suite
Personal & professional development opportunities
Unconditional inclusion in the workplace
Job summary
A leading technology company is seeking a Principal Site Reliability Engineer to optimize and manage cloud infrastructure. The role requires over 10 years of experience in Infra Ops or SRE, with proficiency in cloud platforms like AWS and GCP. Key responsibilities include enhancing Infrastructure as Code, improving CI/CD pipelines, and troubleshooting production issues. This position operates within a hybrid model, requiring attendance at the San Juan office twice a week. Join a forward-thinking team focused on operational excellence.
Qualifications
Minimum of 10 years of hands-on experience in Infra Ops, Dev Ops, or Site Reliability Engineering (SRE).
Proficiency with Linux systems, especially Debian-based distributions.
Strong experience with cloud platforms such as AWS and GCP.
Expertise in Infrastructure as Code tools like Terraform, Packer, and Ansible.
Solid programming skills in Python and/or Golang.
Deep understanding of containerization (Docker) and orchestration tools (AWS EKS, GCP GKE).
Experience with GitOps workflows.
Proven track record in implementing and maintaining CI/CD pipelines.
Responsibilities
Enhance Infrastructure as Code (IAC) and enforce best practices.
Optimize cloud infrastructure for scalability, security, and cost-effectiveness.
Develop internal tools to support and streamline cloud platform operations.
Improve CI/CD pipelines and deployment workflows using FluxCD and Jenkins.
Address container image vulnerabilities and standardize remediation processes.
Build Amazon Machine Images (AMIs) aligned with CIS and STIG benchmarks.
Strengthen monitoring, alerting, and observability using Prometheus and Grafana.
Troubleshoot complex production issues to ensure system reliability.
Skills
Infrastructure as Code (IAC)
Cloud platforms such as AWS and GCP
Linux systems proficiency
Containerization and orchestration tools
Python and/or Golang programming
GitOps workflows
CI/CD pipeline implementation
Monitoring tools (Prometheus, Grafana)
Education
10+ years experience in Infra Ops or SRE
Tools
Terraform
Ansible
Docker
Prometheus
Grafana
Apache Kafka
Cassandra
Job description
A leading technology company is seeking a Principal Site Reliability Engineer to optimize and manage cloud infrastructure. The role requires over 10 years of experience in Infra Ops or SRE, with proficiency in cloud platforms like AWS and GCP. Key responsibilities include enhancing Infrastructure as Code, improving CI/CD pipelines, and troubleshooting production issues. This position operates within a hybrid model, requiring attendance at the San Juan office twice a week. Join a forward-thinking team focused on operational excellence.