Principal Site Reliability Engineer

Hirebridge LLC

Denver (CO)

On-site

USD 160,000 - 180,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Vertafore seeks a Principal Site Reliability Engineer to define strategic vision and own enterprise reliability, scalability, and performance for production services. You will set architectural standards and drive proactive, engineering-first operations across AWS, hybrid data centers, and customer-hosted environments.

You will lead cross‑departmental initiatives, shape observability and incident response culture, and govern SLOs with executive alignment, delivering robust platform stability and

Qualifications

  • Hands-on cloud operations and reliability engineering experience.
  • End-to-end service ownership experience across large orgs.
  • Strong software engineering background in modern languages.
  • Expertise in SRE principles: SLIs, SLOs, error budgets.
  • Proven ability to operate at Principal/Architect scope.

Responsibilities

  • Define enterprise-wide standards for end-to-end service ownership.
  • Lead cross-functional architectural initiatives across cloud platforms.
  • Drive observability strategy with actionable metrics and four golden signals.
  • Govern SLOs/SLIs and error budgets across multiple product lines.
  • Lead incident response with blameless postmortems and timely root cause analyses.

Skills

SRE & cloud ops
Architecture leadership
Observability
Incident management
Executive collaboration

Education

Bachelor’s or Master’s degree in Computer Science or related field

Tools

AWS
Kubernetes
CI/CD
Infrastructure as Code

Job description

$160,000 - $180,000 / year + Bonus

The insurance industry runs on Vertafore. We equip agencies, MGAs, and carriers with the core digital systems, specialized AI, and data-driven foundation to eliminate distribution drag across the insurance lifecycle, spanning sales, servicing, and back-office operations.

Underpinned by unmatched speed and performance power, we are the trusted backbone that’s taking the insurance industry from friction to flow with Distribution Velocity – speed, performance, and trust - to drive growth at scale.

With over 95% of the top agencies and insurers and 50% of industry compliance transactions running through Vertafore, we lead at the intersection of innovation and trust, giving insurance professionals the confidence to transform and win in the AI era.

Our reach is global, with headquarters in Denver, Colorado, and offices across the U.S., Canada, and India.

We are seeking a Principal Site Reliability Engineer to define the strategic vision and own the enterprise-wide reliability, scalability, and performance of our critical production services. As a foundational pillar of our engineering organization, this role drivesarchitecturalstandards for the full-service lifecycle—from initial design and deployment readiness to proactive production operations. At Vertafore, we view reliability as a core engineering responsibility. You will operate autonomously across AWS, hybrid data centers, and customer-hosted environments, setting the technical direction for how we treat operations as a software engineering challenge. This role is pivotal in transitioning cross-departmental teams toward a highly proactive, engineering-first culture.

Roles and Responsibilities:
Strategic Leadership & Reliability Architecture
  • Enterprise-Wide Ownership:Define the standards for end-to-end service ownership, holding the organization accountable for availability, performance, and overall operational health.

  • Architectural Influence:Lead cross-departmental initiatives to influence system design at the architectural level, driving fault tolerance, strict compliance, and operational sustainability across public and private clouds.

  • Advanced Observability Vision:Dictate the enterprise strategy for observability frameworks, ensuring the Four Golden Signals (Latency, Traffic, Errors, and Saturation) provide actionable, predictive insights across all platforms.

Strategic Leadership & Reliability Architecture
  • Enterprise-Wide Ownership:Define the standards for end-to-end service ownership, holding the organization accountable for availability, performance, and overall operational health.

  • Architectural Influence:Lead cross-departmental initiatives to influence system design at the architectural level, driving fault tolerance, strict compliance, and operational sustainability across public and private clouds.

  • Advanced Observability Vision:Dictate the enterprise strategy for observability frameworks, ensuring the Four Golden Signals (Latency, Traffic, Errors, and Saturation) provide actionable, predictive insights across all platforms.

Data-Driven Reliability Governance
  • SLO & Error Budget Authority:Establish the governance models for defining and managing SLIs and SLOs across multiple product lines.

  • Delivery Alignment:Champion Error Budgets as the ultimate technical arbiter at theexecutive level, balancing feature velocity with the absolute requirement for platform stability.

Incident Management & Cultural Transformation
  • Enterprise Incident Command:Lead incident response for the most critical, high‑severity events.

  • Blameless Culture Champion:Foster a "Win Together" environment by championing a Blameless Postmortem culture globally, ensuring root cause analyses focus strictly on systemic and process improvements rather than individual error.

Qualifications & Requirements
  • Experience:12 to 15+ years of hands‑on Cloud Operations, SRE, or reliability‑focused engineering experience, with a proven track record of end‑to‑end enterprise service ownership.

  • Proven Scope:Demonstrated ability to operate at a Principal/Architect scope, driving large‑scale reliability outcomes and operational excellence across global organizations.

  • Software Engineering:Expert‑level software engineering skills in C#, .NET, Java, Python, or React.

  • Principles:Deep expertise in scaling core SRE principles (SLIs, SLOs, error budgets) across complex, distributed systems.

  • Technical Stack:Mastery of AWS, Kubernetes, CI/CD pipelines, Infrastructure‑as‑Code, and extensive knowledge of Linux and Windows environments and relational databases.

  • Education:Bachelor’s orMaster’s degree in Computer Scienceor a related technical field.

  • Commitment:Participation in an executive on‑call rotation with flexible hours as required

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal Site Reliability Engineer
Principal Site Reliability Engineer

Vertafore Career Center • Denver (CO)

On-site
USD 160,000 - 180,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Hirebridge • Northern (KY)

On-site
USD 110,000 - 145,000
Bonus
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Vertafore Career Center • Denver (CO)

On-site
USD 110,000 - 145,000
Director, Site Reliability Engineering
Director, Site Reliability Engineering

Vertafore Career Center • Denver (CO)

On-site
USD 175,000 - 220,000
Senior SRE Architect: Enterprise Reliability & Scale
Senior SRE Architect: Enterprise Reliability & Scale

Vertafore • Denver (CO)

On-site
USD 160,000 - 180,000
Medical, Vision & Dental
401(k) + Employer Match
Parental Leave
+2
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Engg • Tempe (AZ)

On-site
USD 140,000 - 190,000
Enterprise SRE Architect – Reliability & Observability
Enterprise SRE Architect – Reliability & Observability

Vertafore Career Center • Denver (CO)

On-site
USD 160,000 - 180,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Vertafore • Denver (CO)

On-site
USD 160,000 - 180,000
Medical, Vision & Dental
401(k) + Employer Match
Parental Leave
+2
Senior Site Reliability Engineer — Flexible Hours, Scale
Senior Site Reliability Engineer — Flexible Hours, Scale

Vertafore • Denver (CO)

On-site
USD 110,000 - 145,000
Senior SRE: Scale, Reliability & Observability Leader
Senior SRE: Scale, Reliability & Observability Leader

Vertafore • Colorado

On-site
USD 110,000 - 145,000
Medical, Vision & Dental plans
401(k) with employer match
Parental Leave & Adoption Assistance
+3