Site Reliability Engineering (SRE)/Dev Ops

Spectraforce

Louisville (KY)

Hybrid

USD 80,000 - 90,000

Full time

14 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Spectraforce is seeking a Senior Site Reliability Engineer to own reliability across cloud, edge, and network stacks for restaurant technology. You will shape SLOs/SLIs, build observability, and drive automation across Kubernetes, CI/CD, and IaC platforms.

The role requires strong cloud fundamentals, incident response, and hands-on execution in a hybrid environment in Louisville, KY. 6-month assignment with conversion potential, 40h/wk.

Qualifications

  • 6+ years in IT infrastructure, with 3+ years focused on site reliability engineering, cloud architecture, or platform engineering.
  • Hands-on experience with AWS services: VPC, EC2, ECS, EKS, Fargate, Lambda, IAM, RDS, S3.
  • Working experience with Microsoft Azure.
  • Strong expertise in Kubernetes and container orchestration, including edge or distributed deployments.

Responsibilities

  • Define and own SLOs/SLIs for restaurant technology systems and use error budgets to balance reliability with velocity.
  • Build and maintain monitoring, alerting, and observability platforms that provide meaningful signal.
  • Lead blameless post-incident reviews and drive root cause analysis with permanent corrective actions.
  • Proactively identify reliability risks across the stack before incidents.
  • Architect and manage cloud architectures on AWS with knowledge of Azure; manage edge Kubernetes deployments.
  • Develop IaC using Terraform and support containerized workloads (EKS/AKS).
  • Own architecture across cloud, edge, networking, and device management for restaurant tech.

Skills

Cloud architecture
Site reliability engineering
Kubernetes
AWS
Azure
CI/CD
Terraform
Monitoring/observability
Linux
Networking
Automation scripting
Incident response

Tools

GitLab CI/CD
Terraform
Kubernetes
AWS CloudWatch/Grafana

Job description

Role: Site Reliability Engineering (SRE)/Dev Ops

Location: Louisville, KY 40213

Duration: 6 months assignment with possibility of conversion

Schedule: Hybrid - Tue & Thur onsite, 40hrs/week

Job Description:

We are seeking an experienced and pragmatic Senior Site Reliability Engineer to own the reliability, design, implementation, and continuous improvement of the infrastructure that powers restaurant technology — from the cloud platforms and CI/CD (Continuous Integration/Continuous Deployment) pipelines we build on, to the Kubernetes-based edge systems deployed in restaurants, to the networks, MDM platforms, and automation tooling that keeps everything running.

This is a generalist role at a senior level. The right candidate brings strong cloud and reliability engineering fundamentals but is equally comfortable working across edge infrastructure, enterprise networking, mobile deployments, and operational automation. You will define what "reliable" means for our systems, measure it, and engineer solutions to continuously improve it. You'll work closely with the Global Reliability Engineering (GRE), DevOps, and security teams, and you'll be expected to shift fluidly between strategic architecture and hands‑on execution.

Key Responsibilities:
Reliability & Observability
  • Define and own Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for restaurant technology systems; use error budgets to balance reliability with velocity.
  • Build and maintain monitoring, alerting, and observability platforms that provide meaningful signal — not noise.
  • Lead blameless post-incident reviews; drive root cause analysis and ensure permanent corrective actions are implemented.
  • Proactively identify reliability risks across the stack before they become incidents.
  • Establish and track reliability metrics; report on system health to engineering leadership.
  • Design and implement scalable, secure, and highly available cloud architectures primarily on AWS, with working knowledge of Azure.
  • Architect and manage containerized workloads using Kubernetes (EKS, AKS), including edge Kubernetes deployments in restaurant environments.
  • Design and implement serverless and container‑based solutions, including AWS Fargate and other managed services.
  • Develop and maintain Infrastructure as Code (IaC) using Terraform.
  • Own architecture across the full restaurant technology stack — cloud, edge, networking, and device management — not just the cloud layer.
DevOps, Automation & Tooling
  • Build and optimize CI/CD pipelines using GitLab CI/CD and modern DevOps practices.
  • Build internal tools, automations, and middleware integrations that eliminate repetitive operational work.
  • Use AI‑assisted development to accelerate scripting, troubleshooting, and documentation.
  • Champion a culture of engineering solutions over repeated manual fixes — if something is done twice, it should be automated.
  • Design and support Kubernetes‑based edge systems deployed in restaurant locations.
  • Support mobile application deployments and troubleshoot deployment issues across restaurant endpoints.
  • Manage and optimize Mobile Device Management (MDM) platforms covering the restaurant device fleet.
  • Configure and troubleshoot enterprise networking — primarily switches and restaurant‑facing network infrastructure.
  • Lead and participate in incident response for restaurant technology systems, including on‑call coverage and post‑incident review.
  • Reduce mean time to detection (MTTD) and mean time to resolution (MTTR) through better tooling, runbooks, and automation.
Security & Governance
  • Establish and enforce cloud governance, security policies, and architectural standards.
  • Implement cloud security best practices: IAM strategy, network segmentation, encryption, and secrets management.
  • Conduct security architecture reviews; identify vulnerabilities, misconfigurations, and compliance gaps.
  • Integrate security into CI/CD pipelines (DevSecOps — Development, Security, and Operations), including automated scanning, policy validation, and vulnerability management.
  • Collaborate with engineering, DevOps, and security teams to ensure secure‑by‑design solutions across cloud and restaurant tech.
  • Provide technical leadership and mentorship to engineering teams.
  • Create and maintain documentation, runbooks, and architectural decision records.
  • Continuously evaluate emerging technologies and recommend improvements.
Required Qualifications
  • 6+ years in IT infrastructure, with 3+ years focused on site reliability engineering, cloud architecture, or platform engineering.
  • Hands‑on experience with AWS (VPC, EC2, ECS, EKS, Fargate, Lambda, IAM, RDS, S3).
  • Working experience with Microsoft Azure.
  • Strong expertise in Kubernetes and container orchestration, including edge or distributed deployments.
  • Experience with GitLab CI/CD and CI/CD pipeline design.
  • Solid experience with Terraform for infrastructure provisioning.
  • Experience with enterprise networking — switch configuration, VLANs, network troubleshooting.
  • Familiarity with Mobile Device Management (MDM) platforms.
  • Experience with automation and scripting (Python, Bash, Go, or equivalent).
  • Proven ability to build internal tooling and API integrations, not just configure managed services.
  • Experience defining and operating against SLOs, SLIs, and error budgets.
  • Comfortable working in Linux command‑line environments; Windows familiarity a plus where restaurant endpoints require it.
Preferred Qualifications
  • Experience designing serverless architectures (AWS Lambda, Fargate, API Gateway, EventBridge).
  • Experience with DevSecOps tooling (SAST — Static Application Security Testing, DAST — Dynamic Application Security Testing, container scanning, IaC scanning).
  • Familiarity with security frameworks (CIS, NIST, ISO 27001, SOC 2).
  • AWS and/or Azure certifications.
  • Experience with monitoring and observability tools (CloudWatch, Prometheus, Grafana, or SIEM solutions).
  • Background in restaurant, retail, or distributed edge technology environments.
  • Experience using AI‑assisted development tools for scripting, troubleshooting, and documentation.
  • Generalist mindset — comfortable moving between cloud, edge, networking, and device management in the same week.
  • Reliability‑first thinking — treats toil reduction, error budgets, and post‑incident learning as core engineering disciplines, not afterthoughts.
  • Bias toward permanent fixes and automation over repeated manual intervention.
  • Ability to balance strategic architecture with hands‑on execution.
  • Strong communication and cross‑functional collaboration skills.
  • Curious, proactive, and detail‑oriented approach to systems design and operations.

At SPECTRAFORCE, we are committed to maintaining a workplace that ensures fair compensation and wage transparency in adherence with all applicable state and local laws. This position's pay range is $58.00/hr - $65.00/hr.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineering (SRE)/Dev Ops
Site Reliability Engineering (SRE)/Dev Ops

Pyramid Consulting, Inc • Louisville (KY)

On-site
USD 83,000 - 90,000
Site Reliability Engineer
Site Reliability Engineer

Staffworxs • Louisville (KY)

Hybrid
USD 120,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Experis Technology Group • Louisville (KY)

Hybrid
USD 65,000 - 87,000
Medical and Prescription Drug Plans
Dental Plan
Vision Plan
+2
SRE/Devops Engineer
SRE/Devops Engineer

KellyMitchell Group • Louisville (KY)

On-site
USD 80,000 - 114,000
Medical insurance
Dental insurance
Vision insurance
+1
Senior SRE: Cloud, Edge & CI/CD Reliability
Senior SRE: Cloud, Edge & CI/CD Reliability

Spectraforce • Louisville (KY)

Hybrid
USD 80,000 - 90,000
Site Reliability Engineer II
Site Reliability Engineer II

Restaurant365 • Denver (CO)

On-site
USD 98,583 - 138,016
100% employee medical benefits
401k + matching
Equity Option Grant
+2
Site Reliability Engineer II
Site Reliability Engineer II

Restaurant365 • San Francisco (CA)

On-site
USD 98,583 - 138,016
Comprehensive medical benefits, 100% paid for employee
401k + matching
Equity Option Grant
+2
Sr Engineer, Site Reliability Engineer
Sr Engineer, Site Reliability Engineer

Devopsroles • St. Louis (MO)

Hybrid
USD 112,000 - 134,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

GovCIO • Arlington (VA)

On-site
USD 210,000 - 230,000
Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

VITG • Ellicott City (MD)

On-site
USD 90,000 - 120,000
401(k) with employer contribution
Medical/Dental/Vision insurance
Paid vacation (PTO)