Systems Engineer – SRE Enablement

AutoZone

Memphis (TN)

Hybrid

USD 120,000 - 160,000

Full time

5 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Hybrid work model

Job summary

AutoZone's Site Reliability Engineering team seeks a Systems Engineer focused on SRE Enablement to promote reliability across the engineering organization. The role emphasizes standards, shared tools, and guidance for teams, with a hybrid GCP and on‑prem footprint.

The successful candidate will collaborate with application, infrastructure, and architecture teams to embed SRE practices early in development and ensure production readiness. Strong communication and consulting skills are essential.

Qualifications

  • Bachelor's degree in computer science, MIS, information technology, or related field, or equivalent practical experience.
  • 4 to 7 years of experience in Systems Engineering, DevOps, or SRE-related roles.
  • Strong understanding of SRE principles, including SLOs, SLIs, error budgets, and TOIL reduction.
  • Hands-on experience building, administering, and optimizing observability and APM pipelines with Dynatrace.
  • Experience deploying and supporting workloads in Google Cloud Platform (GCP) and on-premises environments.
  • Experience with container orchestration platforms (Kubernetes) and IaC tools (Terraform, Ansible).
  • Excellent communication and consulting skills to influence architecture decisions.

Responsibilities

  • Define enterprise-wide reliability standards, SLO frameworks, and error budget policies.
  • Establish production readiness criteria and conduct readiness reviews across teams.
  • Own and maintain the internal SRE handbook and reliability playbooks.
  • Build, maintain, and standardize shared observability platforms, leveraging Dynatrace.
  • Provide templates for alerting, dashboards, and runbooks for cloud and on-prem workloads.
  • Participate in incident management, including post-mortems to improve reliability.
  • Run SRE training programs and reliability workshops for engineering teams.
  • Coach teams on SLO-based thinking and error budget management.
  • Embed proactive SRE practices and a continuous improvement mindset.
  • Track and report reliability metrics across the enterprise.
  • Identify systemic reliability gaps and report health to leadership.
  • Act as an internal consultant during architecture and system design reviews.
  • Advise on reliability design patterns for hybrid GCP/on-prem environments.
  • Engage early in new product development to influence system reliability.

Skills

SRE principles
Dynatrace
GCP
Kubernetes
Python
Golang
Java
Observability
Incident management
SLOs/SLIs/Error budgets
TOIL reduction
IaC (Terraform/Ansible)

Education

Bachelor's degree in CS / IT / MIS or related field

Tools

Kubernetes
Terraform
Ansible
Dynatrace

Job description

AutoZone's Site Reliability Engineering (SRE) team is seeking a Systems Engineer with a focus on SRE Enablement. This position is responsible for promoting reliability and operational excellence throughout the engineering organization. The successful candidate will play a key role in establishing standards, developing shared tools, providing guidance to development teams, and cultivating a culture of reliability across our hybrid infrastructure, which includes a primary emphasis on the Google Cloud Platform (GCP), as well as on-premises servers and applications.

The SRE Enablement Engineer collaborates closely with application, infrastructure, and architecture teams to integrate SRE best practices early in the software development lifecycle and ensure platforms adhere to rigorous production readiness standards.

  • Define enterprise-wide reliability standards, Service-Level Objective (SLO) frameworks, and error budget policies.
  • Establish production readiness criteria that teams must meet prior to launching and conduct production readiness reviews across teams.
  • Own, document, and maintain the internal SRE handbook and reliability playbooks.
  • Build, maintain, and standardize shared observability platforms, specifically leveraging Dynatrace to be consumed by all engineering teams.
  • Provide templates for alerting, dashboards, and runbooks across both cloud and on-premises application workloads.
  • Participate in the incident management process, including post-mortem analysis, to continuously strengthen systemic reliability.
  • Run SRE training programs and reliability workshops for engineering teams.
  • Coach and mentor teams on SLO-based thinking and error budget management.
  • Embed proactive SRE practices and a continuous improvement mindset into the broader engineering culture.
  • Track and report reliability metrics across the enterprise, rather than just a single service.
  • Identify systemic reliability gaps and trends across cross-functional teams.
  • Report organizational reliability health to leadership and hold teams accountable to agreed-upon operational standards.
  • Act as an internal consultant during architecture and system design reviews.
  • Advise development teams on reliability design patterns (e.g., circuit breakers, retries, graceful degradation) suitable for a hybrid GCP and on-premises environment.
  • Engage early in new product development to influence system reliability from the outset.
  • Bachelor's degree in computer science, MIS, Information Technology, or a related field, or equivalent practical experience. 4 to 7 years of experience in Systems Engineering, DevOps, or SRE-related roles.
  • Deep understanding of Site Reliability Engineering principles, particularly regarding SLOs, SLIs, error budgets, and TOIL reduction.
  • Hands‑on experience building, administering, and optimizing observability and APM pipelines, with a strong focus on Dynatrace.
  • Strong experience deploying and supporting workloads in Google Cloud Platform (GCP), as well as maintaining legacy on‑premises servers and applications.
  • Experience with container orchestration platforms (e.g., Kubernetes).
  • Strong programming/scripting skills (e.g., Python, Golang, Java) and experience with IaC tools (e.g., Terraform, Ansible).
  • Exceptional communication and consulting skills, with the ability to influence architecture decisions and translate technical concepts to non‑technical leadership.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE Enablement Engineer: Reliability & Observability Lead
SRE Enablement Engineer: Reliability & Observability Lead

AutoZone • Memphis (TN)

Hybrid
USD 120,000 - 160,000
Hybrid work model
Site Reliability Engineer
Site Reliability Engineer

Compunnel, Inc. • New Jersey

On-site
USD 120,000 - 150,000
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

Hybrid
USD 130,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

JPS Tech Solutions • San Jose (CA)

On-site
USD 130,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Observability Engineer / Site Reliability Engineer
Observability Engineer / Site Reliability Engineer

Ontrac Solutions • New York (NY)

On-site
USD 120,000 - 190,000
Senior Site Reliability Engineer NEX
Senior Site Reliability Engineer NEX

NexTier Completion Solutions Inc. • Houston (TX)

On-site
USD 110,000 - 150,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Bank of America • Charlotte (NC)

On-site
USD 152,000 - 192,000
Industry-leading benefits
Paid time off
Discretionary incentive eligibility
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Bank of America • Plano (TX)

On-site
USD 152,600 - 191,500
Industry-leading benefits
Paid time off
Access to resources and support
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 240,000