Senior Site Reliability Engineer

Inspire-Brands

Atlanta (GA)

On-site

USD 140,000 - 190,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Inspire Brands is seeking a Senior Site Reliability Engineer to build, scale, and improve reliability for our high-traffic, customer-facing digital platforms. You will blend software engineering, systems thinking, and operational excellence to reduce toil and incidents while engineering reliability into systems.

You will define SLIs/SLOs, drive incident responses, automate toil, and collaborate with engineering to design scalable, observable architectures across cloud platforms.

Qualifications

  • 5+ years experience in Site Reliability Engineering, Software Engineering, or Platform Engineering.
  • 2+ years experience with Kubernetes and containerized workloads.
  • 4-year degree in Computer Science or related field.
  • Strong programming/scripting skills (Python, Go, Java, or Node.js).
  • Experience with chaos engineering or resiliency testing is a plus.

Responsibilities

  • Define and manage SLIs, SLOs, and Error Budgets for critical services.
  • Drive production readiness reviews and reliability requirements into architecture and design.
  • Perform capacity planning, failure mode analysis, and dependency risk assessments.
  • Identify systemic reliability risks and drive remediation before customer impact.
  • Design monitoring, alerting, logging, and tracing solutions using modern tooling.
  • Lead technical response for high-severity incidents and conduct blameless postmortems.
  • Continuously improve detection, response, and recovery processes.
  • Automate toil reduction and build self-healing capabilities.
  • Support Infrastructure as Code (Terraform, Bicep or similar) and CI/CD improvements.
  • Partner with engineering for horizontal scalability and fault tolerance.

Skills

Python/Go/Java/Node.js
SRE principles
Incident response
Distributed systems
Cloud platforms

Education

4-year degree in Computer Science or related field

Tools

Kubernetes
Terraform
Bicep
Azure/AWS/GCP

Job description

Inspire Brands is hiring two Senior Site Reliability Engineers to help build and scale reliable, resilient, and observable systems supporting high-traffic, customer-facing digital platforms. These role blends software engineering, systems thinking, and operational excellence to reduce toil, prevent incidents, and improve system reliability at scale.The ideal candidate has hands-on experience applying and implementing SRE principles — not just supporting production systems, but engineering reliability into them.RESPONSIBILITIESReliability EngineeringDefine and manage SLIs, SLOs, and Error Budgets for critical servicesDrive production readiness reviews and reliability requirements into architecture and designPerform capacity planning, failure mode analysis, and dependency risk assessmentsIdentify systemic reliability risks and drive remediation before they cause customer impactObservabilityDesign monitoring, alerting, logging, and tracing solutions using modern observability toolingImprove signal-to-noise ratio and reduce alert fatigueBuild dashboards and telemetry that reflect true service health, not just infrastructure metricsIncident ManagementLead technical response for high-severity incidentsDrive blameless postmortems and root cause analysis focused on systemic fixesContinuously improve detection, response, and recovery processesParticipate in an on-call rotationAutomation & Toil ReductionIdentify and eliminate manual, repetitive operational work through automationBuild self-healing systems, tooling, and scripts to reduce human interventionImprove CI/CD pipelines and deployment safety (canary, rollback, blue-green)Support Infrastructure as Code (Terraform, Bicep, or similar)Performance & ScalabilityConduct load testing, performance benchmarking, and bottleneck analysisPartner with engineering to design systems for horizontal scalability and fault toleranceCollaboration & CulturePartner with engineering teams to implement resiliency patterns (circuit breakers, retries, graceful degradation, rate limiting)Mentor engineers on SRE best practicesPromote a culture of engineering-driven reliability over reactive operationsEDUCATION AND EXPERIENCE QUALIFICATIONSRequired Qualifications5+ years experience in Site Reliability Engineering, Software Engineering, or Platform Engineering2+ years experience with Kubernetes and containerized workloads4-year degree in Computer Science or related fieldPreferred QualificationsExperience with chaos engineering or resiliency testingExperience with high-volume, high-availability transactional systemsExperience with AI-assisted observability or operational automationExperience making meaningful contributions to internal SRE tooling, frameworks, or platformsREQUIRED KNOWLEDGE, SKILLS, OR ABILITIESStrong programming/scripting skills (Python, Go, Java, or Node.js)Demonstrated experience defining and operating against SLOs/Error BudgetsStrong skills in leading incident response and root cause analysis for production systemsSolid understanding of distributed systems and microservices architectureDeep knowledge and expertise in at least one major cloud platform (Azure, AWS, or GCP)Expertise with observability platforms and monitoring strategyThis position is based in our Atlanta Support Center, with an expected on-site presence of 80%.Inspire is a multi-brand restaurant company whose portfolio includes more than 33,300 Arby’s, Baskin-Robbins, Buffalo Wild Wings, Dunkin’, Jimmy John’s, and SONIC restaurants worldwide. We’re made up of some of the world’s most iconic restaurant brands, but we’re much more than just a restaurant company. We’re a team of hundreds of thousands who individually and collectively are changing the way people eat, drink, and gather around the table. We know that food is much more than a staple—it’s an experience. At Inspire, that’s our purpose: to ignite and nourish flavorful experiences.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Inspire • Atlanta (GA)

On-site
USD 140,000 - 200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Inspire Brands, Inc. • Atlanta (GA)

On-site
USD 120,000 - 180,000
Senior Site Reliability Engineer: Build Resilient Systems
Senior Site Reliability Engineer: Build Resilient Systems

IRB USA Inspire Resources • Atlanta (GA)

On-site
USD 130,000 - 180,000
Senior Site Reliability Engineer – Scale & Observability
Senior Site Reliability Engineer – Scale & Observability

Inspire Brands, Inc. • Atlanta (GA)

On-site
USD 120,000 - 180,000
Lead Technology Operations Analyst - SONIC
Lead Technology Operations Analyst - SONIC

Inspire Brands, Inc. • Atlanta (GA), Northern (KY)

Hybrid
USD 100,000 - 140,000
Senior SRE: Reliability & Observability Lead
Senior SRE: Reliability & Observability Lead

Inspire • Atlanta (GA)

On-site
USD 140,000 - 200,000
Lead Technology Operations Analyst - SONIC
Lead Technology Operations Analyst - SONIC

Inspire • Atlanta (GA)

On-site
USD 90,000 - 130,000
Senior Site Reliability Engineer: Build Resilient, Scalable Systems
Senior Site Reliability Engineer: Build Resilient, Scalable Systems

Inspire-Brands • Atlanta (GA)

On-site
USD 140,000 - 190,000
Technology Operations Senior Analyst
Technology Operations Senior Analyst

Inspire Brands, Inc. • Atlanta (GA)

On-site
USD 120,000 - 180,000
Technology Operations Senior Analyst
Technology Operations Senior Analyst

Inspire • Atlanta (GA)

On-site
USD 110,000 - 160,000