Principal SRE: Cloud Platform & Reliability Leader

ShipperHQ

Austin (TX)

Hybrid

USD 180,000 - 240,000

Full time

4 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

22 days PTO + public holidays
401k Match
Medical, Dental, and Vision Insurance
Maternity and Paternity Leave
Hybrid Austin, TX office

Job summary

ShipperHQ is looking for a Principal Site Reliability Engineer to lead evolution of our cloud platform, reliability strategy, and infrastructure architecture. This is a highly technical, hands-on leadership role focused on scalable, resilient systems and engineering best practices across teams.

As the Principal SRE, you will own the cloud infra roadmap in AWS, drive observability and platform reliability, and partner with Security, QA, and Product to build secure, highly available services that

Qualifications

  • 10+ years of experience in Site Reliability Engineering or related fields.
  • Proven experience designing and operating large-scale AWS infrastructure.
  • Strong software engineering background with automation skills.
  • Expert-level experience with Infrastructure as Code, preferably Terraform.
  • Deep experience with modern CI/CD pipelines using GitLab or similar.
  • Strong knowledge of Kubernetes and cloud-native architectures.
  • Extensive experience with observability platforms, distributed tracing, logging, monitoring, and incident response.
  • Experience defining and implementing SLOs, SLIs, and reliability engineering best practices.
  • Strong understanding of networking, security, Linux systems administration, and cloud architecture.
  • Experience supporting high-traffic SaaS applications and mission-critical production environments.
  • Excellent problem-solving skills with the ability to influence technical direction.
  • Demonstrated ability to influence technical direction without direct authority while mentoring engineers across multiple teams.
  • Experience working in Agile development environments and partnering closely with cross-functional engineering teams.

Responsibilities

  • Own the technical vision and roadmap for ShipperHQ's cloud infrastructure, reliability, and platform engineering initiatives.
  • Design, build, and maintain highly available, scalable, and secure cloud infrastructure in AWS.
  • Architect and evolve Infrastructure as Code (Terraform) standards across all environments.
  • Design and optimize CI/CD pipelines that enable fast, reliable, and repeatable software delivery.
  • Define and implement reliability standards, SLOs, SLIs, error budgets, and incident management best practices.
  • Lead the design and implementation of observability, monitoring, logging, and alerting across the platform.
  • Build self-service platform capabilities and automation that empower engineering teams and reduce operational overhead.
  • Drive infrastructure modernization initiatives, including containerization, orchestration, and platform scalability.
  • Partner with Security to implement cloud security best practices, compliance controls, and governance.
  • Collaborate with Engineering teams to improve application reliability, performance, and operational excellence.
  • Lead technical decision-making for infrastructure architecture and serve as a trusted advisor across engineering.
  • Mentor engineers and promote best practices in cloud architecture, automation, reliability, and operational excellence.
  • Evaluate and introduce new technologies that improve scalability, reliability, developer productivity, and operational efficiency.
  • Participate in incident response, root cause analysis, and continuous improvement efforts for production systems.

Job description

ShipperHQ is looking for a Principal Site Reliability Engineer to lead evolution of our cloud platform, reliability strategy, and infrastructure architecture. This is a highly technical, hands-on leadership role focused on scalable, resilient systems and engineering best practices across teams.

As the Principal SRE, you will own the cloud infra roadmap in AWS, drive observability and platform reliability, and partner with Security, QA, and Product to build secure, highly available services that

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal SRE: Cloud Reliability & Observability Leader
Principal SRE: Cloud Reliability & Observability Leader

Gen Digital Inc. • Tempe (AZ)

On-site
USD 180,000 - 230,000
Senior SRE Leader: Cloud, Reliability & Scale
Senior SRE Leader: Cloud, Reliability & Scale

AVG • Tempe (AZ), Northern (KY)

Hybrid
USD 140,000 - 190,000
Senior SRE & Cloud Reliability Architect
Senior SRE & Cloud Reliability Architect

United States Digital Space LLC • United States

Remote
USD 180,000 - 250,000
Remote Principal SRE Team Lead | Global Cloud Reliability
Remote Principal SRE Team Lead | Global Cloud Reliability

Jobgether • United States

Hybrid
USD 170,000 - 210,000
Annual bonus
Medical, dental, and vision insurance
Life and disability insurance
+3
Principal SRE: Enterprise Reliability & Automation
Principal SRE: Enterprise Reliability & Automation

Early Warning • San Francisco (CA)

Hybrid
USD 194,000 - 284,000
Healthcare Coverage
401(k) Retirement Plan
Paid Time Off
+2
Principal SRE: Hybrid Cloud Reliability & Observability
Principal SRE: Hybrid Cloud Reliability & Observability

Hewlett Packard Enterprise • San Juan (PR)

Hybrid
USD 120,000 - 160,000
Senior Cloud SRE & Reliability Architect
Senior Cloud SRE & Reliability Architect

Apply • Tempe (AZ), Northern (KY)

Hybrid
USD 180,000 - 240,000
Principal Site Reliability Engineer - Austin, Texas
Principal Site Reliability Engineer - Austin, Texas

ShipperHQ • Austin (TX)

Hybrid
USD 180,000 - 240,000
22 days PTO + public holidays
401k Match
Medical, Dental, and Vision Insurance
+2
Senior Principal SRE – Cloud Reliability Leader
Senior Principal SRE – Cloud Reliability Leader

Habitat For Humanity Of Durham • Durham (NC)

On-site
USD 140,000 - 170,000
On-site health centers
Fully paid parental leave
Principal SRE — Observability & Cloud Reliability
Principal SRE — Observability & Cloud Reliability

T. Rowe Price • Washington

Hybrid
USD 159,000 - 339,000
Competitive compensation
Annual bonus eligibility
Hybrid work schedule
+2