Senior Site Reliability Engineer

ViziRecruiter,LLC.

Chicago (IL)

Hybrid

USD 125,040 - 187,560

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading food retail company is looking for a Site Reliability Engineer III in Chicago. This role involves ensuring system reliability, implementing automation, and leading operational processes in a distributed cloud-native environment. Candidates should have a Bachelor's in Computer Science, extensive experience in SRE or DevOps, and proficiency in relevant programming and tools. The position offers competitive salaries and a hybrid work schedule.

Qualifications

  • 5+ years of experience in Site Reliability Engineering or DevOps.
  • Proficiency in programming and scripting languages including Python, Java, Bash, or Go.
  • Proven experience managing production-grade systems on AKS/Kubernetes.

Responsibilities

  • Design and implement infrastructure solutions for cloud-native environments.
  • Develop automation for deployment and incident remediation.
  • Monitor production environments and improve performance and reliability.

Skills

Site Reliability Engineering
DevOps
Java programming
Linux environments
Observability stacks
Containerization
Automation
Problem-solving

Education

Bachelor's Degree in Computer Science or related field

Tools

Terraform
GitHub Actions
Docker
Kubernetes
Datadog
Spring Boot
Kafka

Job description

Introduction

Ahold Delhaize USA, a division of global food retailer Ahold Delhaize, is part of the U.S. family of brands, which also includes five leading omnichannel grocery brands – Food Lion, Giant Food, The GIANT Company, Hannaford and Stop & Shop. Ahold Delhaize USA associates support the brands with a wide range of services, including Finance, Legal, Sustainability, Commercial, Digital and E-commerce, Technology and more.

Overview

The Site Reliability Engineer (SRE) III is responsible for ensuring the scalability, reliability, and performance of production systems through automation, observability, incident response, and infrastructure engineering. This role involves designing and implementing robust operational processes and tooling to support highly available, fault-tolerant systems in a cloud-native environment. The SRE III collaborates closely with engineering squads, product teams, and stakeholders to embed reliability best practices across the software delivery lifecycle. The role includes ownership of system uptime, service level objectives (SLOs), and operational excellence, along with mentoring junior engineers and leading cross-functional initiatives that improve system resilience.

Applicants must be currently authorized to work in the United States on a full-time basis.

Our flexible/hybrid work schedule includes 3 in-person days at our Chicago office and 2 remote days.

Responsibilities
  • Design and implement infrastructure solutions that ensure system availability, scalability, and reliability across cloud-native environments like AKS and Kubernetes.
  • Develop automation for provisioning, deployment, configuration, monitoring, and incident remediation using tools such as Terraform, ArgoCD, and GitHub Actions.
  • Collaborate with engineering teams to define and track service level objectives (SLOs) and service level indicators (SLIs).
  • Build and manage microservices-based platforms leveraging Spring Boot, Java, Tomcat, and Redis.
  • Monitor production environments using Datadog and proactively address performance and reliability issues.
  • Perform root cause analysis and lead post-incident reviews to drive continual improvement.
  • Manage CI/CD pipelines and deployment automation using GitHub, Docker, and container orchestration technologies.
  • Create and maintain infrastructure as code (IaC) using Terraform, with deployment pipelines integrated into GitOps workflows.
  • Lead and support operational readiness reviews, game days, chaos engineering practices, and failure mode analysis.
  • Build scalable observability and alerting frameworks with Datadog.
  • Implement resilient, asynchronous architectures using Kafka for event-driven services.
  • Reduce operational toil through self-healing automation and proactive system tuning.
  • Troubleshoot Linux-based environments such as Ubuntu and optimize them for reliability.
  • Provide on-call support and ensure 24/7/365 system reliability for mission-critical applications.
  • Collaborate with the security team to enforce secure operational practices and cloud compliance.
  • Mentor junior engineers and contribute to documentation, technical design, and knowledge-sharing across the organization.
Requirements
  • Bachelor's Degree in Computer Science, Information Systems, or a related technical field; equivalent training, certifications, or experience will be considered.
  • 5+ years of experience in a Site Reliability Engineering, or DevOps, or Java programming role.
  • Experience managing production-grade systems and services on AKS/Kubernetes in distributed environments.
  • Proficiency in programming and scripting languages including Python, Java, Bash, or Go.
  • Proven experience with Spring Boot, Tomcat, Redis, and microservices architecture.
  • Hands‑on experience in managing Linux environments, particularly Ubuntu.
  • Proficiency with observability stacks and performance monitoring using Datadog, Prometheus, and ELK.
  • Deep understanding of containerization and orchestration using Docker, Kubernetes, and ArgoCD.
  • Experience managing event‑driven systems using Kafka.
  • Expertise in IaC and automation using Terraform and GitHub Actions.
  • Familiarity with networking concepts, DNS, load balancing, and cloud infrastructure (AWS, Azure, or GCP).
  • Strong analytical, debugging, and problem‑solving skills.
  • Excellent verbal and written communication skills and the ability to collaborate effectively across teams.

Salary Range: $125,040 - $187,560

Actual compensation offered to a candidate may vary based on their unique qualifications and experience, internal equity, and market conditions. Final compensation decisions will be made in accordance with company policies and applicable laws.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

ViziRecruiter,LLC. • Quincy (MA)

Hybrid
USD 125,000 - 188,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

ViziRecruiter,LLC. • Salisbury (NC)

Hybrid
USD 125,000 - 188,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

ViziRecruiter,LLC. • Quincy (MA)

Hybrid
USD 146,000 - 221,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

ViziRecruiter,LLC. • Salisbury (NC)

Hybrid
USD 146,000 - 221,000
Site Reliability Engineer
Site Reliability Engineer

Request Technology, LLC • Chicago (IL)

Hybrid
USD 150,000 - 155,000
Lead, Site Reliability Engineer
Lead, Site Reliability Engineer

CardWorks • Pittsburgh

Hybrid
USD 146,000 - 163,000
Competitive Pay
Medical, Dental, and Vision Benefits
401(k) Plan with Company Match
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mission Staffing • New York (NY)

Hybrid
Site Reliability Engineer
Site Reliability Engineer

SRE • Puerto Rico

Hybrid
USD 120,000 - 180,000
Site Reliability Engineering Lead
Site Reliability Engineering Lead

RELX INC • Chicago (IL)

On-site
USD 124,200 - 230,800
Health Benefits
401(k) with match
Wellness platform
+1
Senior Site Reliability Engineer II
Senior Site Reliability Engineer II

RELX • New York (NY)

Hybrid
USD 104,000 - 175,000
Competitive salary
Hybrid or remote options
Career growth in SRE and DevOps
+1