Principal Site Reliability Engineer

Prosum

Scottsdale (AZ)

On-site

USD 150,000 - 190,000

Full time

5 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Prosum is seeking a Principal Site Reliability Engineer to provide technical leadership across large-scale production environments. You will drive reliability, observability, and automation across software, cloud, and DevOps teams.

The role requires deep SRE and software engineering experience, strong leadership, and a focus on scalable, resilient architectures in a modern cloud environment. Scottsdale-based on-site work is expected with enterprise impact.

Qualifications

  • 15+ years of relevant experience in SRE, software engineering, or related fields.
  • Experience with production reliability, distributed systems, and cloud technologies.
  • Strong knowledge of Linux/Unix, networking, and observability.
  • Experience with AWS and cloud-native architectures; Linux, scripting, debugging.

Responsibilities

  • Lead reliability engineering efforts across large-scale production systems.
  • Define and improve SLIs, SLOs, error budgets, and service health metrics.
  • Design observability using metrics, logs, tracing, dashboards; drive incident response and RCA.
  • Continuously improve CI/CD, IaC, deployment practices, automation, testing, and capacity planning.
  • Mentor engineers and influence architecture and engineering practices across teams.

Skills

SRE
Software engineering
DevOps
Cloud computing
Observability
Automation
Technical leadership

Education

Bachelor's degree

Tools

AWS
Terraform
Kubernetes
Docker

Job description

We are seeking an experienced Principal Site Reliability Engineer (SRE) to provide technical leadership across highly available, large-scale production environments. This role combines software engineering, systems engineering, cloud infrastructure, automation, DevOps, and production reliability to improve the resilience, scalability, performance, observability, and operational health of critical services.

The Principal SRE will partner closely with Software Engineering, Platform Engineering, Cloud Infrastructure, DevOps, and other technology teams to ensure reliability and operational readiness are incorporated throughout the software development lifecycle.

This is a senior individual contributor position with enterprise-level influence. The successful candidate will identify systemic reliability risks, establish technical direction, influence architecture and engineering practices, and help improve reliability capabilities across multiple engineering teams.

Key Responsibilities

  • Apply Site Reliability Engineering (SRE), software engineering, automation, and DevOps principles to improve how production services are built, tested, deployed, monitored, operated, and recovered.
  • Establish and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), error budgets, availability metrics, and service-health measurements.
  • Design and enhance observability capabilities using metrics, logging, distributed tracing, monitoring, alerting, dashboards, and service-health instrumentation.
  • Drive continuous improvement across CI/CD pipelines, Infrastructure as Code (IaC), cloud infrastructure, deployment practices, automation, testing, incident management, capacity planning, resilience, disaster recovery, and operational readiness.
  • Analyze production environments to identify systemic reliability risks, performance bottlenecks, recurring incidents, and opportunities for automation.
  • Translate production and operational experience into improvements in application code, architecture, infrastructure, tooling, automation, and engineering standards.
  • Partner with Software Engineering teams to incorporate reliability, resiliency, scalability, performance, observability, recoverability, and operational readiness throughout the development lifecycle.
  • Lead or participate in production incident response, troubleshooting, root cause analysis, service restoration, and blameless post-incident reviews.
  • Provide technical leadership during critical production incidents and help improve incident response, escalation procedures, service restoration, and sustainable on-call practices.
  • Reduce operational toil and manual intervention through software development, scripting, automation, reusable tooling, platforms, and engineering patterns.
  • Apply data-driven analysis, experimentation, and engineering principles to validate assumptions and guide technical decisions.
  • Establish and influence enterprise-level SRE, DevOps, cloud, reliability, and operational engineering standards and best practices.
  • Mentor engineers and technical leaders while promoting knowledge sharing and sustainable engineering capabilities across the organization.
  • Operate independently across complex, business-critical reliability and infrastructure challenges.

Required Qualifications

  • 15+ years of relevant professional experience in one or more of the following areas:
  • Site Reliability Engineering (SRE)
  • Software Engineering
  • Strong experience with software development and/or scripting using one or more modern programming languages.
  • Advanced understanding of software engineering principles, distributed systems, production environments, troubleshooting, automation, and observability.
  • Experience designing, supporting, or improving highly available, scalable production systems and distributed applications.
  • Experience with public cloud platforms and cloud-native architectures, preferably AWS.
  • Strong knowledge of Linux/Unix systems, networking, infrastructure, application architecture, and production operations.
  • Demonstrated experience diagnosing complex production issues and implementing sustainable technical solutions.
  • Strong analytical, troubleshooting, problem-solving, communication, and cross-functional collaboration skills.
  • Ability to provide technical direction and influence engineering practices across multiple teams and organizational boundaries.

Preferred Qualifications

  • Extensive hands-on experience with Amazon Web Services (AWS) or another major cloud platform, including Microsoft Azure, Google Cloud Platform (GCP), or Oracle Cloud Infrastructure (OCI).
  • Experience with CI/CD pipelines and software delivery automation.
  • Experience with Infrastructure as Code (IaC) technologies and practices.
  • Experience with containers and container orchestration technologies.
  • Strong experience with monitoring, logging, distributed tracing, dashboards, alerting, and observability platforms.
  • Experience defining and managing SLIs, SLOs, error budgets, availability targets, and reliability metrics.
  • Experience with incident management, root cause analysis, performance engineering, capacity planning, resilience testing, disaster recovery, and operational readiness.
  • Experience creating reusable automation, tooling, platforms, frameworks, engineering patterns, or standards that improve engineering productivity and system reliability.
  • Experience influencing architecture and technical strategy for large-scale or business-critical production systems.
  • Demonstrated ability to mentor senior engineers and improve technical capabilities across engineering organizations.
  • Bachelor's degree in Computer Science, Software Engineering, Computer Engineering, Information Systems, or a related technical discipline, or equivalent practical experience.

Key Technical Skills / ATS Keywords

Site Reliability Engineering (SRE), AWS, Cloud Computing, DevOps, Software Engineering, Systems Engineering, Platform Engineering, Distributed Systems, Production Reliability, High Availability, Scalability, Resilience, Observability, Infrastructure as Code (IaC), CI/CD, Automation, Linux, Unix, Networking, Containers, Container Orchestration, Monitoring, Logging, Distributed Tracing, Alerting, SLIs, SLOs, Error Budgets, Incident Management, Root Cause Analysis, Production Support, Performance Engineering, Capacity Planning, Disaster Recovery, Resilience Testing, Operational Readiness, Cloud Architecture, Production Operations, Software Development, Scripting, Troubleshooting, Technical Leadership.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer Lead
Site Reliability Engineer Lead

Good co India • United States

Remote
USD 120,000 - 160,000
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

On-site
USD 130,000 - 180,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Engg • Tempe (AZ)

On-site
USD 140,000 - 190,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

AVG • Tempe (AZ), Northern (KY)

Hybrid
USD 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

SDI International • Chicago (IL)

On-site
USD 130,000 - 180,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Apply • Tempe (AZ), Northern (KY)

Hybrid
USD 180,000 - 240,000
Site Reliability Engineer
Site Reliability Engineer

Harvey Nash • United States

Remote
USD 120,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

Brooksource • San Antonio (TX)

On-site
USD 80,000 - 120,000
Site Reliability Engineer -Jersey City, NJ & Dallas, TX
Site Reliability Engineer -Jersey City, NJ & Dallas, TX

Stradit LLC • Dallas (TX), Northern (KY)

On-site
USD 140,000 - 190,000