Site Reliability Engineer – Lead

Jobtailor

Arizona

On-site

USD 140,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor is seeking a Senior Site Reliability Engineer/Platform Engineering leader to shape the technology strategy, build a high-performing team, and drive reliability across distributed environments. You will partner with engineering leadership to define standards, governance, and scalable operations.

You will own incident response, capacity planning, and telemetry-driven improvements, mentor engineers, and advance modern practices such as IaC, Kubernetes, and comprehensive observability

Qualifications

  • 10+ years of experience in systems engineering, DevOps, or SRE roles in large-scale environments.
  • Deep understanding of Linux/Unix & Windows systems, networking, and distributed computing.
  • Proven experience with observability stacks (e.g., Dynatrace, Grafana, Splunk, OpenTelemetry).
  • Expertise in infrastructure-as-code and automation tools (e.g., Terraform, Ansible, Python).
  • Strong knowledge of cloud platforms and container orchestration (Kubernetes).
  • Demonstrated success in leading incident response and driving systemic improvements.
  • Experience with capacity planning, performance tuning, and cost optimization.
  • Excellent communication and stakeholder management skills, including executive engagement.

Responsibilities

  • Build and lead a team to deliver technology products and services
  • Develop a technology strategy and ensure technology solutions comply with standards
  • Promote design, engineering, and organizational practices
  • Advocate and advance modern, Agile solution delivery practices
  • Define and implement SRE frameworks, including SLIs/SLOs/SLAs, error budgets, and incident response protocols
  • Establish governance models for reliability engineering across distributed teams
  • Champion a culture of observability and proactive monitoring
  • Lead root cause analysis (RCA) and post-incident reviews
  • Implement proactive problem detection using telemetry
  • Develop and maintain capacity models and monitor performance trends
  • Drive automation of operational tasks including deployments and scaling
  • Oversee major incident response and communication processes
  • Serve as a senior technical advisor and thought leader in SRE and platform engineering
  • Mentor SRE teams and partner with engineering leaders across the enterprise

Job description

  • Build and lead a team to deliver technology products and services
  • Develop a technology strategy and ensure technology solutions comply with standards
  • Promote design, engineering, and organizational practices
  • Advocate and advance modern, Agile solution delivery practices
  • Define and implement SRE frameworks, including SLIs/SLOs/SLAs, error budgets, and incident response protocols
  • Establish governance models for reliability engineering across distributed teams
  • Champion a culture of observability and proactive monitoring
  • Lead root cause analysis (RCA) and post-incident reviews
  • Implement proactive problem detection using telemetry
  • Develop and maintain capacity models and monitor performance trends
  • Drive automation of operational tasks including deployments and scaling
  • Oversee major incident response and communication processes
  • Serve as a senior technical advisor and thought leader in SRE and platform engineering
  • Mentor SRE teams and partner with engineering leaders across the enterprise
Requirements
  • 10+ years of experience in systems engineering, DevOps, or SRE roles in large-scale environments
  • Deep understanding of Linux/Unix & Windows systems, networking, and distributed computing
  • Proven experience with observability stacks (e.g., Dynatrace, Grafana, Splunk, OpenTelemetry)
  • Expertise in infrastructure-as-code and automation tools (e.g., Terraform, Ansible, Python)
  • Strong knowledge of cloud platforms and container orchestration (Kubernetes)
  • Demonstrated success in leading incident response and driving systemic improvements
  • Experience with capacity planning, performance tuning, and cost optimization
  • Excellent communication and stakeholder management skills, including executive engagement.
Core Competencies

Demonstrates extensive expertise in Site Reliability Engineering (SRE) and platform engineering, with a strong focus on developing technology strategies, implementing observability practices, and leading incident response efforts. Proven ability to mentor teams and drive automation in large-scale environments while ensuring compliance with industry standards.

Highest-signal resume keywords
  • Site Reliability Engineering (SRE)
  • Infrastructure-as-Code
  • Observability Stacks
  • Cloud Platforms
  • Incident Response Leadership
Hard Skills
  • Systems Engineering
  • DevOps
  • Linux/Unix Systems
  • Windows Systems
  • Networking
  • Distributed Computing
  • Capacity Planning
  • Performance Tuning
  • Cost Optimization
  • Automation Tools
Soft Skills
  • Excellent Communication
  • Stakeholder Management
  • Executive Engagement
Industry Keywords
  • Agile Solution Delivery
  • Governance Models
  • Proactive Monitoring
  • Root Cause Analysis
  • Telemetry
Tools & Technologies
  • Terraform
  • Ansible
  • Python
  • Kubernetes
  • Dynatrace
  • Grafana
  • Splunk
  • OpenTelemetry
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Director, Site Reliability Engineering
Director, Site Reliability Engineering

Jobtailor • California (MO)

On-site
USD 180,000 - 260,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Luxoft • Wilmington (DE)

On-site
USD 140,000 - 190,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Veriipro • Washington

On-site
USD 120,000 - 180,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Luxoft • Buffalo (NY)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobtailor • Town of Florida (NY)

Hybrid
USD 150,000 - 190,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 260,000
Site Reliability Engineer
Site Reliability Engineer

SCIGON • Naperville (IL)

Hybrid
USD 110,000 - 170,000
Software Engineering Manager – Site Reliability Center
Software Engineering Manager – Site Reliability Center

Jobtailor • Alabama

On-site
USD 120,000 - 160,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Virtual Tech Gurus • Puerto Rico

On-site
USD 140,000 - 210,000