Senior Site Reliability Engineer

Socket.dev

Denver (CO)

Hybrid

USD 140,000 - 165,000

Full time

6 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Colorado PERA is seeking an experienced Senior Site Reliability Engineer to own production reliability, observability, and platform automation across a hybrid cloud and on‑premises environment. You will drive incident management, scale CI/CD pipelines, and partner with cross‑functional teams to achieve uptime for critical services.

The ideal candidate excels at automating reliability improvements, is fluent in AWS and Azure, and can lead architecture discussions for scalable, cost‑efficient

Qualifications

  • Bachelor’s degree in Computer Science, Information Technology, or related field preferred, or an equivalent combination of education and experience.
  • Relevant certifications such as AWS DevOps Engineer, Solutions Architect, CKAs (CKA), Terraform Associate are preferred.
  • Production experience with AWS and Azure across compute, networking, and managed services.

Responsibilities

  • Own PERA’s observability platform and drive enterprise rollout including licensing, architecture, and integration strategy.
  • Design and maintain monitoring, logging, and alerting standards across AWS, Azure and on‑premises infrastructure, Kubernetes workloads, APIs, and services.
  • Build automation with AI-driven integrations to reduce manual work and drive improvement of SLO/KPI metrics.
  • Serve as senior technical responder for production incidents and lead blameless postmortems.
  • Define and maintain SLOs and error budgets for critical production services.
  • Collaborate to design for scalability and fault tolerance across hybrid Kubernetes clusters.

Skills

AWS
Azure
Kubernetes
IaC
Python
PowerShell
Linux
CI/CD
Observability

Education

Bachelor's degree in Computer Science / Information Technology or related field

Tools

Terraform
ServiceNow

Job description

Job Details: Level: Experienced, Job Location: Penn Center - Denver, CO, Position Type: Full Time, Salary Range: $140000.00 - $165000.00Salary, Job Shift: Day

Summary of Job Responsibilities The Senior Site Reliability Engineer (SRE) is Colorado PERA’s technical owner for production reliability, observability, and platform automation across a hybrid AWS, Azure and on-premises environment. This position leads the design, implementation, and ongoing operation of PERA’s enterprise observability platform, drives incident management practices that reduce time to resolution, scales CI/CD delivery pipelines, and partners with Infrastructure, Application Development, Investment Technical Services and Product teams to achieve uptime for critical services. The role is accountable for turning monitoring data into measurable improvements in reliability, delivery speed, and cloud cost efficiency and serves as the senior technical voice for reliability engineering practices across the organization.

Ideal Candidate Statement

The ideal candidate is a senior, hands-on reliability engineer who thinks in systems, automates and treats every incident as a chance to remove future value leakage. They are fluent across AWS and Azure cloud, Kubernetes, and IaC from build to configuration management utilizing AI tooling to speed MTTR and improve uptime by consolidating metrics, logs, and distributed trace collection into observability tools with well-defined dashboards and actionable integrations into ITSM Service Now management in an environment primarily on premises and rapidly moving to cloud platform and infrastructure services.

Essential Duties and Responsibilities
Observability Platform & Automation Tooling
  • Lead the evaluation, selection, and enterprise rollout of PERA’s observability platform, including licensing, architecture, and integration strategy.
  • Design and maintain monitoring, logging, and alerting standards across AWS, Azure and on-premises infrastructure, Kubernetes workloads, APIs, services and microservices.[DS1.1]
  • Tune alerting to reduce noise and false positives while improving detection of real degradation.[DS1.1]
  • Build automation with AI-driven integrations to eliminate manual, repetitive operational work, including self-healing and auto-remediation driving improvement of SLO/KPI metrics where appropriate.[DS2.1]
  • Develop dashboards and reporting that translate technical telemetry into business-relevant reliability and cost metrics for IT leadership.
Incident Management & Reliability Engineering
  • Serve as a senior technical responder and escalation point for Priority 1/Priority 2 production incidents, driving time-to-resolution down through structured incident command practices.
  • Own PERA’s blameless postmortem process: facilitate root-cause reviews, document findings, and track corrective actions to closure.
  • Define and maintain Service Level Objectives (SLOs) and error budgets for critical production services.
  • Build and maintain runbooks, escalation paths, and on-call procedures that reduce reliance on tribal knowledge.
  • Analyze incident trends to identify systemic reliability risks and prioritize remediation work.
Architecture & Scaling
  • Partner with Infrastructure, Application Development and Investment Technical Services to design for scalability, resiliency, and fault tolerance across hybrid Kubernetes cluster [MP3.1][DS3.2]orchestration.
  • Conduct capacity planning and load/performance analysis to proactively identify scaling risks before they affect uptime.
  • Contribute reliability and observability requirements into architecture reviews for new services and major changes.
CI/CD & Release Engineering
  • Assess and scale PERA’s CI/CD pipelines to support faster, safer, and more frequent deployments.
  • Implement methodologies to safeguard modern hybrid DevSecOps, and Kubernetes environments through security controls in CI/CD pipelines.
  • Integrate observability and automated testing gates into the CI/CD pipeline so reliability issues are caught before production.
Cost Optimization & Vendor/Tooling Governance
  • Own the business case, budget, and ongoing vendor relationship for the selected observability platform.
  • Identify and implement cloud cost optimization opportunities (rightsizing, autoscaling, reserved capacity) surfaced through observability data.
  • Evaluate emerging SRE/observability tooling and recommend investments aligned with PERA’s hybrid technology roadmap.
Job Qualifications
  • 6+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or a closely related discipline.
  • Bachelor’s degree in Computer Science, Information Technology, or related field preferred, or an equivalent combination of education and experience.
  • Relevant certifications preferred: AWS Certified DevOps Engineer or Solutions Architect, Certified Kubernetes Administrator (CKA), HashiCorp Terraform Associate.[MP4.1][DS4.2]
  • Production experience with AWS and Azure across compute, networking, and managed services.
  • Hands-on experience operating and troubleshooting Kubernetes in a hybrid (on-premises and cloud) environment.
  • Strong Python, Powershell and Linux [MP5.1][DS5.2][BP5.3][BP5.4]shell scripting/automation skills; ability to build tooling, not just run it.
  • Experience designing and operating an enterprise observability/monitoring platforms.
  • Demonstrated experience reducing MTTR/MTTD through tooling, automation, and process not headcount.
  • Experience with CI/CD tooling and modern release practices.
  • Working knowledge of microservices architecture and the reliability challenges specific to distributed systems.
  • Experience defining and reporting on SLOs, error budgets, and uptime commitments (99.9%+ environments).
  • Experience participating in an on-call rotation and leading through live production incidents.
Preferred Leadership & Strategic Experience
  • Experience leading a vendor evaluation and selection process for enterprise tooling, including business case development.
  • Experience establishing reliability standards or practices adopted across multiple engineering teams without direct reporting authority.
  • Experience mentoring engineers on reliability, automation, or observability practices.
  • Experience presenting reliability and cost metrics to IT leadership or executive audiences.

Working Conditions The physical demands described here are representative of those that must be met by an employee to successfully perform the essential functions of this job. Reasonable accommodation may be made to enable individuals with disabilities to perform the essential functions. Standard office environment with frequent computer operation and use of collaboration/communication tools. Participation in an on-call rotation, including evenings, weekends, and holidays as needed to support 99.99% uptime commitments. Ability to remain calm, clear, and decisive under pressure during live production incidents. Ability to sit for prolonged periods of time and operate standard PC equipment. Ability to handle stress associated with production incidents, tight deadlines, and competing priorities.

Hybrid Work Option Opportunity to work from home up to three days per week. Eligibility dependent upon factors detailed in PERA’s Work from Home Policy and on-call coverage needs.

Job Description Disclaimer This job description is not designed to cover or contain a comprehensive listing of activities, duties, or responsibilities that are required of an employee. Duties, responsibilities, and activities may change or be assigned with or without notice. Unfortunately, at this time, PERA cannot consider candidates that require sponsorship (now or in the future), or are located outside of the US. All Colorado PERA employees are subject to PERA’s Ethics Policy and some employees are subject to the Personal Trading Policy. These policies include restrictions on outside business activities and employment and have certain requirements on personal trading. You may request copies of these policies from PERA’s talent acquisition team and any questions can be answered by PERA’s Investment Administration team.

Why Work at PERA Colorado PERA offers more than a traditional pension career. We are a mission-driven organization with a growing focus on technology and modernization. Employees have the opportunity to do meaningful work with a real impact on over 700,000 members. From enhancing digital tools and member experiences to supporting major enterprise initiatives, employees are part of meaningful, future-focused work that blends public service with innovation. We take pride in our inclusive culture, career development opportunities, and our consistent recognition as a Top Workplace based on employee feedback. PERA is a place to connect, contribute, and be part of something bigger. A Culture That Cares At PERA, leveraging employee strengths and supporting strong engagement are the foundations of our culture. Employees are encouraged to grow through development opportunities and internal career movement, while also being part of an organization where people feel respected, informed, and included. Collaboration is encouraged across teams, allowing employees to learn from one another and contribute in ways that maximize their unique skills. We strive to foster an environment where employees can do their best work and feel supported by the people around them An Employer that Invests in You PERA invests in our employees in ways that matter; from comprehensive benefits and generous paid time-off to thoughtful everyday amenities that enhance the office experience. Employees are encouraged to continue learning through training, mentoring, and development at every stage of their careers. We champion a workplace where people feel valued, inspired, and equipped to grow. Join Us If you’re energized by meaningful work, motivated by an organization that evolves with a changing world and looking for an employer that invests in you, then consider joining PERA’s team. Learn more about careers at PERA at copera.org/careers.

Position Title: Senior Site Reliability Engineer Division: Information Technology Reports to: Infrastructure Manager Job Status: Full-time, Exempt Salary: $140,000.00 to $165,000.00 Annual, Commensurate with experience Posting Dates: 09/18/2026 to 10/04/2026

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Colorado-Public-Employees • Denver (CO)

Hybrid
USD 140,000 - 165,000
Hybrid work option
On-call rotation
Work from home eligibility
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Colorado PERA • Denver (CO)

Hybrid
USD 140,000 - 190,000
Hybrid work option
Investment Systems Engineer
Investment Systems Engineer

Colorado-Public-Employees • Denver (CO)

Hybrid
USD 108,000 - 134,000
Hybrid work option
Work-from-home days after orientation
End-User Computing Systems Senior Engineer
End-User Computing Systems Senior Engineer

Colorado-Public-Employees • Denver (CO)

Hybrid
USD 120,000 - 160,000
Hybrid work option
Retirement Payroll Specialist
Retirement Payroll Specialist

Colorado Public Employees' Retirement Association • Denver (CO)

Hybrid
USD 43,000 - 48,000
Hybrid work option
Member Services Manager
Member Services Manager

Colorado Public Employees' Retirement Association in • Westminster (CO)

Hybrid
USD 96,000 - 110,000
Hybrid work option
Work from home 2 days per week
External Job Posting Title Site Reliability Engineer
External Job Posting Title Site Reliability Engineer

Peraton • Washington

On-site
USD 112,000 - 179,000
External Job Posting Title Site Reliability Engineer (SRE) – Evening Shift
External Job Posting Title Site Reliability Engineer (SRE) – Evening Shift

Peraton • Northern (KY)

Hybrid
USD 104,000 - 166,000
Senior Investment Stewardship Analyst
Senior Investment Stewardship Analyst

Colorado Public Employees' Retirement Association in • Denver (CO)

Hybrid
USD 130,000 - 155,000
Senior SRE: Observability, Automation & Hybrid Cloud
Senior SRE: Observability, Automation & Hybrid Cloud

Colorado-Public-Employees • Denver (CO)

Hybrid
USD 140,000 - 165,000
Hybrid work option
On-call rotation
Work from home eligibility