Inspire Brands is hiring two Senior Site Reliability Engineers to help build and scale reliable, resilient, and observable systems supporting high-traffic, customer-facing digital platforms. These role blends software engineering, systems thinking, and operational excellence to reduce toil, prevent incidents, and improve system reliability at scale.The ideal candidate has hands-on experience applying and implementing SRE principles — not just supporting production systems, but engineering reliability into them.RESPONSIBILITIESReliability EngineeringDefine and manage SLIs, SLOs, and Error Budgets for critical servicesDrive production readiness reviews and reliability requirements into architecture and designPerform capacity planning, failure mode analysis, and dependency risk assessmentsIdentify systemic reliability risks and drive remediation before they cause customer impactObservabilityDesign monitoring, alerting, logging, and tracing solutions using modern observability toolingImprove signal-to-noise ratio and reduce alert fatigueBuild dashboards and telemetry that reflect true service health, not just infrastructure metricsIncident ManagementLead technical response for high-severity incidentsDrive blameless postmortems and root cause analysis focused on systemic fixesContinuously improve detection, response, and recovery processesParticipate in an on-call rotationAutomation & Toil ReductionIdentify and eliminate manual, repetitive operational work through automationBuild self-healing systems, tooling, and scripts to reduce human interventionImprove CI/CD pipelines and deployment safety (canary, rollback, blue-green)Support Infrastructure as Code (Terraform, Bicep, or similar)Performance & ScalabilityConduct load testing, performance benchmarking, and bottleneck analysisPartner with engineering to design systems for horizontal scalability and fault toleranceCollaboration & CulturePartner with engineering teams to implement resiliency patterns (circuit breakers, retries, graceful degradation, rate limiting)Mentor engineers on SRE best practicesPromote a culture of engineering-driven reliability over reactive operationsEDUCATION AND EXPERIENCE QUALIFICATIONSRequired Qualifications5+ years experience in Site Reliability Engineering, Software Engineering, or Platform Engineering2+ years experience with Kubernetes and containerized workloads4-year degree in Computer Science or related fieldPreferred QualificationsExperience with chaos engineering or resiliency testingExperience with high-volume, high-availability transactional systemsExperience with AI-assisted observability or operational automationExperience making meaningful contributions to internal SRE tooling, frameworks, or platformsREQUIRED KNOWLEDGE, SKILLS, OR ABILITIESStrong programming/scripting skills (Python, Go, Java, or Node.js)Demonstrated experience defining and operating against SLOs/Error BudgetsStrong skills in leading incident response and root cause analysis for production systemsSolid understanding of distributed systems and microservices architectureDeep knowledge and expertise in at least one major cloud platform (Azure, AWS, or GCP)Expertise with observability platforms and monitoring strategyThis position is based in our Atlanta Support Center, with an expected on-site presence of 80%.Inspire is a multi-brand restaurant company whose portfolio includes more than 33,300 Arby’s, Baskin-Robbins, Buffalo Wild Wings, Dunkin’, Jimmy John’s, and SONIC restaurants worldwide. We’re made up of some of the world’s most iconic restaurant brands, but we’re much more than just a restaurant company. We’re a team of hundreds of thousands who individually and collectively are changing the way people eat, drink, and gather around the table. We know that food is much more than a staple—it’s an experience. At Inspire, that’s our purpose: to ignite and nourish flavorful experiences.