Senior Site Reliability Engineer

Inspire

Atlanta (GA)

On-site

USD 140,000 - 200,000

Full time

5 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Inspire Brands is seeking two Senior Site Reliability Engineers to build, scale, and operate reliable, observable platforms that power high-traffic digital experiences. You will blend software engineering with reliability practices to reduce toil, improve resilience, and prevent incidents at scale.

The role focuses on defining SLIs/SLOs, incident response, and automation, with an 80% on-site presence in Atlanta and collaboration with cross-functional teams to raise reliability across the

Qualifications

  • 5+ years in Site Reliability Engineering, Software Engineering, or Platform Engineering.

Responsibilities

  • Define and manage SLIs, SLOs, and Error Budgets for critical services.

Skills

SRE fundamentals
Kubernetes
Programming: Python/Go/Java/Node.js
Incident response leadership
Distributed systems

Education

Bachelor's degree in CS or related

Tools

Observability tooling
Cloud platforms (Azure/AWS/GCP)

Job description

Inspire Brands is hiring two Senior Site Reliability Engineers to help build and scale reliable, resilient, and observable systems supporting high-traffic, customer-facing digital platforms. These role blends software engineering, systems thinking, and operational excellence to reduce toil, prevent incidents, and improve system reliability at scale.

Responsibilities
Reliability Engineering
  • Define and manage SLIs, SLOs, and Error Budgets for critical services
  • Drive production readiness reviews and reliability requirements into architecture and design
  • Perform capacity planning, failure mode analysis, and dependency risk assessments
  • Identify systemic reliability risks and drive remediation before they cause customer impact
Observability
  • Design monitoring, alerting, logging, and tracing solutions using modern observability tooling
  • Improve signal-to-noise ratio and reduce alert fatigue
  • Build dashboards and telemetry that reflect true service health, not just infrastructure metrics
Incident Management
  • Lead technical response for high-severity incidents
  • Drive blameless postmortems and root cause analysis focused on systemic fixes
  • Continuously improve detection, response, and recovery processes
  • Participate in an on-call rotation
Automation & Toil Reduction
  • Identify and eliminate manual, repetitive operational work through automation
  • Build self-healing systems, tooling, and scripts to reduce human intervention
  • Improve CI/CD pipelines and deployment safety (canary, rollback, blue-green)
  • Support Infrastructure as Code (Terraform, Bicep, or similar)
Performance & Scalability
  • Conduct load testing, performance benchmarking, and bottleneck analysis
  • Partner with engineering to design systems for horizontal scalability and fault tolerance
Collaboration & Culture
  • Partner with engineering teams to implement resiliency patterns (circuit breakers, retries, graceful degradation, rate limiting)
  • Mentor engineers on SRE best practices
  • Promote a culture of engineering-driven reliability over reactive operations
Education And Experience Qualifications
Required Qualifications
  • 5+ years experience in Site Reliability Engineering, Software Engineering, or Platform Engineering
  • 2+ years experience with Kubernetes and containerized workloads
  • 4-year degree in Computer Science or related field
Preferred Qualifications
  • Experience with chaos engineering or resiliency testing
  • Experience with high-volume, high-availability transactional systems
  • Experience with AI-assisted observability or operational automation
  • Experience making meaningful contributions to internal SRE tooling, frameworks, or platforms
Required Knowledge, Skills, Or Abilities
  • Strong programming/scripting skills (Python, Go, Java, or Node.js)
  • Demonstrated experience defining and operating against SLOs/Error Budgets
  • Strong skills in leading incident response and root cause analysis for production systems
  • Solid understanding of distributed systems and microservices architecture
  • Deep knowledge and expertise in at least one major cloud platform (Azure, AWS, or GCP)
  • Expertise with observability platforms and monitoring strategy

This position is based in our Atlanta Support Center, with an expected on-site presence of 80%.

Inspire is a multi-brand restaurant company whose portfolio includes more than 33,300 Arby’s, Baskin-Robbins, Buffalo Wild Wings, Dunkin’, Jimmy John’s, and SONIC restaurants worldwide. We’re made up of some of the world’s most iconic restaurant brands, but we’re much more than just a restaurant company. We’re a team of hundreds of thousands who individually and collectively are changing the way people eat, drink, and gather around the table. We know that food is much more than a staple—it’s an experience. At Inspire, that’s our purpose: to ignite and nourish flavorful experiences.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Inspire Brands, Inc. • Atlanta (GA)

On-site
USD 120,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

IRB USA Inspire Resources • Atlanta (GA)

On-site
USD 130,000 - 180,000
Senior SRE: Reliability & Observability Lead
Senior SRE: Reliability & Observability Lead

Inspire • Atlanta (GA)

On-site
USD 140,000 - 200,000
Senior Site Reliability Engineer: Build Resilient Systems
Senior Site Reliability Engineer: Build Resilient Systems

IRB USA Inspire Resources • Atlanta (GA)

On-site
USD 130,000 - 180,000
Lead Technology Operations Analyst-Sonic
Lead Technology Operations Analyst-Sonic

Inspire • Atlanta (GA)

On-site
USD 95,000 - 130,000
Senior Site Reliability Engineer – Scale & Observability
Senior Site Reliability Engineer – Scale & Observability

Inspire Brands, Inc. • Atlanta (GA)

On-site
USD 120,000 - 180,000
Staff Tech Operations Engineer-Payment & POPs
Staff Tech Operations Engineer-Payment & POPs

Inspire • Atlanta (GA)

On-site
USD 120,000 - 160,000
Senior Manager - Technology Operations, SONIC
Senior Manager - Technology Operations, SONIC

IRB USA Inspire Resources • Atlanta (GA)

On-site
USD 150,000 - 190,000
Incident and Problem Management Analyst
Incident and Problem Management Analyst

Inspire-Brands • Atlanta (GA)

On-site
USD 90,000 - 130,000
Vice President - Restaurant Technology & Operations - Arby’s & SONIC
Vice President - Restaurant Technology & Operations - Arby’s & SONIC

Inspire • Atlanta (GA)

On-site
USD 230,000 - 360,000