Overview
An experienced and pragmatic Senior Site Reliability Engineer is sought to own the reliability, design, implementation, and continuous improvement of the infrastructure that powers restaurant technology. This includes cloud platforms, CI/CD pipelines, Kubernetes-based edge systems deployed in restaurants, networks, MDM platforms, and automation tooling. This is a senior-level, hands-on engineering role requiring a foundation in cloud infrastructure and reliability engineering, along with experience supporting edge technologies, networking, and operational automation. The ideal candidate can define reliability standards, measure system performance, and implement scalable solutions that continuously improve operational excellence. You will work with Reliability Engineering, DevOps, and Security teams while balancing strategic architecture responsibilities with hands-on engineering execution.
Job #
3054890
Job Description
Senior Site Reliability Engineer
Location
Louisville, Kentucky (Hybrid)
Key Responsibilities
- Define and maintain Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets to balance reliability with velocity.
- Design, build, and maintain monitoring, alerting, and observability platforms that provide meaningful signals.
- Lead blameless post-incident reviews, drive root cause analysis, and ensure permanent corrective actions are implemented.
- Design, implement, and support secure, scalable, and highly available cloud solutions primarily within AWS, with working knowledge of Azure.
- Architect and manage containerized workloads using Kubernetes, including both cloud and edge-based deployments.
- Develop and maintain Infrastructure as Code (IaC) using Terraform.
- Build and optimize CI/CD pipelines using GitLab CI/CD and modern DevOps practices.
- Develop automation solutions, internal tools, and middleware integrations to eliminate repetitive operational work.
- Design and support Kubernetes-based edge systems deployed in restaurant locations.
- Manage and optimize Mobile Device Management (MDM) platforms.
- Establish and enforce cloud governance, security policies, and architectural standards.
- Provide technical leadership and mentorship to engineering teams.
Required Qualifications
- 6+ years in IT infrastructure, with 3+ years focused on site reliability engineering, cloud architecture, or platform engineering.
- Hands-on experience with AWS services (e.g., VPC, EC2, ECS, EKS, Fargate, Lambda, IAM, RDS, S3).
- Working experience with Microsoft Azure.
- Strong expertise in Kubernetes and container orchestration, including edge or distributed deployments.
- Experience with GitLab CI/CD and CI/CD pipeline design.
- Solid experience with Terraform for infrastructure provisioning.
- Experience with enterprise networking, including switch configuration, VLANs, and network troubleshooting.
- Familiarity with Mobile Device Management (MDM) platforms.
- Experience with automation and scripting (Python, Bash, Go, or equivalent).
- Proven ability to build internal tooling and API integrations.
- Experience defining and operating against SLOs, SLIs, and error budgets.
- Comfortable working in Linux command-line environments.
Preferred Qualifications
- Experience designing serverless architectures (AWS Lambda, Fargate, API Gateway, EventBridge).
- Experience with DevSecOps tooling (SAST, DAST, container scanning, IaC scanning).
- Familiarity with security frameworks (CIS, NIST, ISO 27001, SOC 2).
- AWS and/or Azure certifications.
- Experience with monitoring and observability tools (CloudWatch, Prometheus, Grafana, or SIEM solutions).
- Background in restaurant, retail, or distributed edge technology environments.
- Experience using AI-assisted development tools for scripting, troubleshooting, and documentation.
Key Competencies
- Generalist mindset, comfortable moving between cloud, edge, networking, and device management.
- Reliability-first thinking, treating toil reduction and error budgets as core engineering disciplines.
- Bias toward permanent fixes and automation over repeated manual intervention.
- Ability to balance strategic architecture with hands-on execution.
- Strong communication and cross-functional collaboration skills.
- A proactive and detail-oriented approach to systems design and operations.
Equal Opportunity Employer Statement
Everforth Apex Systems is an equal opportunity employer. We do not discriminate or allow discrimination on the basis of race, color, religion, creed, sex (including pregnancy, childbirth, breastfeeding, or related medical conditions), age, sexual orientation, gender identity, national origin, ancestry, citizenship, genetic information, registered domestic partner status, marital status, disability, status as a crime victim, protected veteran status, political affiliation, union membership, or any other characteristic protected by law. Everforth Apex will consider qualified applicants with criminal histories in a manner consistent with the requirements of applicable law.
Accommodation Statement
If you require an accommodation under the Americans with Disabilities Act to participate in an interview with a virtual recruiter or to use our website for a search or application, please contact our Benefits Department at [email protected] or 804-523-8228. Please note that this contact information is strictly to be used for medical ADA accommodations and that no other inquiries will be answered.