Principal Site Reliability Engineer

Jobtailor

Arizona

On-site

USD 180,000 - 240,000

Full time

3 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Jobtailor in the United States is seeking a senior Site Reliability Engineer with 15+ years of experience to lead reliability across services, drive incidents response, and set standards in observability, automation, and cloud architectures.

You will collaborate with Software Engineering teams to improve SLIs/SLOs, implement CI/CD and IaC, and reduce toil through reusable patterns while ensuring production readiness.

Qualifications

  • 15+ years of professional experience in software engineering, SRE, or related fields.
  • Bachelor's degree in CS or a related technical field.

Responsibilities

  • Apply software engineering, automation, and DevOps to improve how services are built, tested, deployed, observed, operated, and recovered.
  • Use data, evidence, experimentation, and engineering analysis to identify reliability risks and guide decisions.
  • Define, implement, or improve SLIs, SLOs, error budgets, and service-health measures.
  • Improve observability through metrics, logging, tracing, monitoring, alerting, dashboards, and instrumentation.
  • Drive CI/CD, IaC, automation, testing, incident response, capacity management, resilience, and operational readiness.
  • Identify recurring or systemic production issues and translate experience into code, architecture, tooling, and practices.
  • Partner with software teams to improve reliability, resiliency, scalability, performance, observability, recoverability, and on-call operations.
  • Participate in or lead incident response and blameless post-incident learning.
  • Provide enterprise-level technical leadership for critical production incidents and impact engineering practices.
  • Reduce operational toil through reusable patterns and automation.
  • Establish enterprise direction and multiply engineering capability.

Skills

15+ years exp
AWS
CI/CD
SLIs/SLOs
Incident management
Observability
Automation
Distributed systems
Cloud architectures
Scripting

Education

Bachelor's degree in CS/Engineering

Tools

CI/CD Tools
Monitoring Tools
Logging Tools
Dashboards
Containers
Orchestration

Job description

  • Apply software engineering, automation, and DevOps principles to improve how services are built, tested, deployed, observed, operated, and recovered
  • Use data, evidence, experimentation, and rigorous engineering analysis to identify reliability risks, test assumptions, and guide technical decisions
  • Define, implement, or improve SLIs, SLOs, error budgets, and service-health measures
  • Improve observability through metrics, logging, tracing, monitoring, alerting, dashboards, and service-health instrumentation
  • Drive continuous improvement across CI/CD, observability, deployment practices, Infrastructure as Code, automation, testing, incident response, capacity management, resilience and operational readiness
  • Identify recurring or systemic production issues and translate operational experience into improvements in code, architecture, automation, tooling, and engineering practices
  • Partner with Software Engineering teams to incorporate reliability, resiliency, scalability, performance, observability, recoverability, and operational readiness throughout the development lifecycle
  • Participate in or lead incident response, troubleshooting, service restoration, and blameless post-incident learning
  • Provide enterprise-level technical leadership for critical production incidents and influence engineering practices that improve incident response, escalation, service restoration, and sustainable on-call operations
  • Reduce operational toil and unnecessary manual intervention through software, automation, reusable patterns, and better engineering practices
  • Establish enterprise technical direction, develop senior technical leaders, and multiply the capability of the broader engineering organization
Requirements
  • Typically 15+ years of relevant professional experience in Software Engineering, Site Reliability Engineering, Systems Engineering, Cloud/Platform Engineering, DevOps, Infrastructure Engineering, Architecture, or a comparable technical discipline
  • Experience with software development or scripting using one or more modern programming languages
  • Experience with software engineering principles, distributed systems, production troubleshooting, automation, and observability appropriate to the level
  • Experience with public cloud technologies and architectures, preferably AWS, along with infrastructure, networking, Linux/Unix, and modern application architectures
  • Demonstrated analytical, problem-solving, communication, and collaboration skills appropriate to the scope of the role
  • Preferred hands-on experience with AWS or comparable experience with Microsoft Azure, Google Cloud Platform (GCP), or Oracle Cloud Infrastructure (OCI)
  • Experience developing, deploying, operating, or improving highly available production software or distributed systems
  • Experience with CI/CD, Infrastructure as Code, containers or orchestration, observability, monitoring, alerting, and software-delivery automation
  • Experience with SLIs, SLOs, error budgets, incident management, performance analysis, capacity management, resilience testing, disaster recovery, or operational readiness
  • Experience creating reusable automation, tooling, platforms, patterns, or practices that improve engineering effectiveness
  • Bachelor's degree in Computer Science, Software Engineering, Computer Engineering, Information Systems, or a related technical field, or equivalent practical experience
  • Must independently possess eligibility to work in the United States at the date of hire
  • Position is ineligible for employment Visa sponsorship
Core Competencies

Demonstrates extensive experience in Software Engineering, Site Reliability Engineering, and DevOps principles, with a strong focus on automation, observability, and incident management. Proficient in leveraging cloud technologies, particularly AWS, to enhance system reliability and operational readiness.

Highest-signal resume keywords
  • 15+ Years Experience in Software Engineering
  • Expertise in AWS or Comparable Cloud Technologies
  • Proficient in CI/CD and Infrastructure as Code
  • Experience with SLIs, SLOs, and Incident Management
  • Strong Analytical and Problem-Solving Skills
Hard Skills
  • Software Development
  • Scripting
  • Distributed Systems
  • Automation
  • Observability
  • Incident Management
  • Performance Analysis
  • Capacity Management
  • Disaster Recovery
  • Resilience Testing
Soft Skills
  • Analytical Skills
  • Problem-Solving
  • Communication
  • Collaboration
Industry Keywords
  • Site Reliability Engineering
  • DevOps
  • Infrastructure Engineering
  • Cloud/Platform Engineering
  • Software Engineering Principles
Tools & Technologies
  • AWS
  • Microsoft Azure
  • Google Cloud Platform
  • Oracle Cloud Infrastructure
  • CI/CD Tools
  • Monitoring Tools
  • Logging Tools
  • Dashboards
  • Containers
  • Orchestration
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer – Digital Assets
Senior Site Reliability Engineer – Digital Assets

Jobtailor • Arizona

On-site
USD 120,000 - 170,000
AVP SRE, Cloud Solutions
AVP SRE, Cloud Solutions

Jobtailor • Arlington (TX)

On-site
USD 180,000 - 240,000
Site Reliability Engineer – Lead
Site Reliability Engineer – Lead

Jobtailor • Arizona

On-site
USD 140,000 - 230,000
Staff Site Reliability Engineer, SRE
Staff Site Reliability Engineer, SRE

Jobtailor • California (MO)

On-site
USD 120,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Director, Site Reliability Engineering
Director, Site Reliability Engineering

Jobtailor • California (MO)

On-site
USD 180,000 - 260,000
Senior Availability & Reliability Engineering Manager
Senior Availability & Reliability Engineering Manager

Jobtailor • Charlotte (NC)

On-site
USD 140,000 - 190,000
Senior Staff Engineer – DevOps
Senior Staff Engineer – DevOps

Jobtailor • California (MO)

On-site
USD 140,000 - 180,000
Director of Engineering – Quality & Reliability
Director of Engineering – Quality & Reliability

Jobtailor • New York (NY)

On-site
USD 180,000 - 260,000
Site Reliability Engineer
Site Reliability Engineer

Jobtailor • New York (NY)

On-site
USD 140,000 - 190,000