Senior Site Reliability Engineer

Inspire Brands, Inc.

Atlanta (GA)

On-site

USD 120,000 - 180,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Inspire Brands is hiring two Senior Site Reliability Engineers to build and scale reliable, observable systems for high-traffic digital platforms. The role blends software engineering with reliability practices to reduce toil and incidents, and to improve system resilience at scale.

Required 5+ years in SRE/engineering, 2+ years with Kubernetes, and a 4-year CS degree. On-site presence in Atlanta, with strong collaboration across engineering teams to implement resiliency patterns and robust

Qualifications

  • 5+ years experience in Site Reliability Engineering, Software Engineering, or Platform Engineering.
  • 2+ years experience with Kubernetes and containerized workloads.
  • 4-year degree in Computer Science or related field.
  • Experience with chaos engineering or resiliency testing (preferred).
  • Experience with AI-assisted observability or operational automation (preferred).

Responsibilities

  • Define and manage SLIs, SLOs, and Error Budgets for critical services.
  • Drive production readiness reviews and reliability requirements into architecture and design.
  • Perform capacity planning, failure mode analysis, and dependency risk assessments.
  • Identify systemic reliability risks and drive remediation before customer impact.
  • Lead technical response for high-severity incidents and drive blameless postmortems.

Skills

SRE principles
Kubernetes
Cloud platforms
Incident response
SLOs & budgets
Scripting: Python/Go/JS
Distributed systems

Education

4-year degree in Computer Science or related field

Job description

Inspire Brands is hiring two Senior Site Reliability Engineers to help build and scale reliable, resilient, and observable systems supporting high-traffic, customer-facing digital platforms. These role blends software engineering, systems thinking, and operational excellence to reduce toil, prevent incidents, and improve system reliability at scale.

The ideal candidate has hands-on experience applying and implementing SRE principles — not just supporting production systems, but engineering reliability into them.

RESPONSIBILITIES
Reliability Engineering
  • Define and manage SLIs, SLOs, and Error Budgets for critical services
  • Drive production readiness reviews and reliability requirements into architecture and design
  • Perform capacity planning, failure mode analysis, and dependency risk assessments
  • Identify systemic reliability risks and drive remediation before they cause customer impact
Observability
  • Design monitoring, alerting, logging, and tracing solutions using modern observability tooling
  • Improve signal-to-noise ratio and reduce alert fatigue
  • Build dashboards and telemetry that reflect true service health, not just infrastructure metrics
Incident Management
  • Lead technical response for high-severity incidents
  • Drive blameless postmortems and root cause analysis focused on systemic fixes
  • Continuously improve detection, response, and recovery processes
  • Participate in an on-call rotation
Automation & Toil Reduction
  • Identify and eliminate manual, repetitive operational work through automation
  • Build self-healing systems, tooling, and scripts to reduce human intervention
  • Improve CI/CD pipelines and deployment safety (canary, rollback, blue-green)
  • Support Infrastructure as Code (Terraform, Bicep, or similar)
Performance & Scalability
  • Conduct load testing, performance benchmarking, and bottleneck analysis
  • Partner with engineering to design systems for horizontal scalability and fault tolerance
Collaboration & Culture
  • Partner with engineering teams to implement resiliency patterns (circuit breakers, retries, graceful degradation, rate limiting)
  • Mentor engineers on SRE best practices
  • Promote a culture of engineering-driven reliability over reactive operations
EDUCATION AND EXPERIENCE QUALIFICATIONS
Required Qualifications
  • 5+ years experience in Site Reliability Engineering, Software Engineering, or Platform Engineering
  • 2+ years experience with Kubernetes and containerized workloads
  • 4-year degree in Computer Scienceor related field
Preferred Qualifications
  • Experience with chaos engineering or resiliency testing
  • Experience with high-volume, high-availability transactional systems
  • Experience with AI-assisted observability or operational automation
  • Experience making meaningful contributions to internal SRE tooling, frameworks, or platforms
REQUIRED KNOWLEDGE, SKILLS, OR ABILITIES
  • Strong programming/scripting skills (Python, Go, Java, or Node.js)
  • Demonstrated experience defining and operating against SLOs/Error Budgets
  • Strong skills in leading incident response and root cause analysis for production systems
  • Solid understanding of distributed systems and microservices architecture
  • Deep knowledge and expertise in at least one major cloud platform (Azure, AWS, or GCP)
  • Expertise with observability platforms and monitoring strategy

This position is based in our Atlanta Support Center, with an expected on-site presence of 80%.

Inspire is a multi-brand restaurant company whose portfolio includes more than 33,300 Arby’s, Baskin-Robbins, Buffalo Wild Wings, Dunkin’, Jimmy John’s, and SONIC restaurants worldwide. We’re made up of some of the world’s most iconic restaurant brands, but we’re much more than just a restaurant company. We’re a team of hundreds of thousands who individually and collectively are changing the way people eat, drink, and gather around the table. We know that food is much more than a staple—it’s an experience. At Inspire, that’s our purpose: to ignite and nourish flavorful experiences.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Inspire • Atlanta (GA)

On-site
USD 140,000 - 200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

IRB USA Inspire Resources • Atlanta (GA)

On-site
USD 130,000 - 180,000
Senior SRE: Reliability & Observability Lead
Senior SRE: Reliability & Observability Lead

Inspire • Atlanta (GA)

On-site
USD 140,000 - 200,000
Senior Site Reliability Engineer: Build Resilient Systems
Senior Site Reliability Engineer: Build Resilient Systems

IRB USA Inspire Resources • Atlanta (GA)

On-site
USD 130,000 - 180,000
Senior Site Reliability Engineer – Scale & Observability
Senior Site Reliability Engineer – Scale & Observability

Inspire Brands, Inc. • Atlanta (GA)

On-site
USD 120,000 - 180,000
Lead Technology Operations Analyst-Sonic
Lead Technology Operations Analyst-Sonic

Inspire • Atlanta (GA)

On-site
USD 95,000 - 130,000
Incident and Problem Management Analyst
Incident and Problem Management Analyst

Inspire-Brands • Atlanta (GA)

On-site
USD 90,000 - 130,000
Vice President - Restaurant Technology & Operations - Arby’s & SONIC
Vice President - Restaurant Technology & Operations - Arby’s & SONIC

Inspire • Atlanta (GA)

On-site
USD 230,000 - 360,000
Lead, Site Reliability Engineer
Lead, Site Reliability Engineer

CardWorks • Pittsburgh

Hybrid
USD 146,000 - 163,000
Competitive Pay
Medical, Dental, and Vision Benefits
401(k) Plan with Company Match
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mission Staffing • New York (NY)

Hybrid
USD 140,000 - 200,000