Principal Site Reliability Engineer

ViziRecruiter,LLC.

Salisbury (NC)

Hybrid

USD 146,960 - 220,440

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading recruitment firm is seeking a Senior Site Reliability Engineer (SRE) IV to drive operational excellence across complex systems. This hybrid role requires expertise in Java, Kubernetes, and DevOps practices. Candidates should possess over 8 years of experience in SRE or related fields and a degree in Computer Science. The compensation ranges from $146,960 to $220,440 based on qualifications and experience.

Qualifications

  • 8+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering.
  • Expertise in building and maintaining Java-based microservices.
  • Proven ability to implement observability platforms using Datadog.

Responsibilities

  • Lead the design and implementation of SRE frameworks.
  • Automate infrastructure provisioning and application deployment workflows.
  • Mentor junior SREs and engineers to foster a culture of continuous improvement.

Skills

Site Reliability Engineering
Java programming
Kubernetes orchestration
Microservices architecture
Automation with Terraform
Observability with Datadog
Containerization with Docker
Python scripting

Education

Bachelor's or Master's degree in Computer Science

Tools

Spring Boot
Tomcat
Redis
GitHub Actions
ArgoCD
Kafka
Terraform
Ubuntu/Linux

Job description

Introduction

Ahold Delhaize USA, a division of global food retailer Ahold Delhaize, is part of the U.S. family of brands, which also includes five leading omnichannel grocery brands – Food Lion, Giant Food, The GIANT Company, Hannaford and Stop & Shop. Ahold Delhaize USA associates support the brands with a wide range of services, including Finance, Legal, Sustainability, Commercial, Digital and E-commerce, Technology and more.

Overview

The Site Reliability Engineer (SRE) IV is a senior technical leader responsible for designing, guiding, and scaling site reliability engineering practices across complex, distributed systems. This role plays a crucial part in driving operational excellence, ensuring system resiliency, and fostering a high-performing engineering culture. The SRE IV works closely with senior leadership, engineering, and product teams to set strategic goals around availability, performance, and incident response while leading large‑scale reliability initiatives.

This position emphasizes deep technical expertise in platforms such as Spring Boot, Java, Tomcat, Redis, and Kafka, along with infrastructure tooling like AKS, Kubernetes, ArgoCD, Terraform, GitHub Actions, and observability platforms like Datadog. The ideal candidate will also bring strong experience working with Ubuntu/Linux environments, containerization with Docker, and automation of operational workflows across a modern DevOps toolchain.

Our flexible/hybrid work schedule includes 3 in-person days at one of our Chicago, IL office and 2 remote days.

Applicants must be currently authorized to work in the United States on a full-time basis.

Responsibilities
  • Architect, evolve, and lead implementation of enterprise-level SRE frameworks, tools, and cloud-native reliability strategies.
  • Build, scale, and manage microservices platforms using Spring Boot, Java, Tomcat, and Redis with Kubernetes and AKS.
  • Lead technical design reviews, chaos testing, and infrastructure planning with an emphasis on scalability, high availability, and fault tolerance.
  • Define, implement, and refine SLOs/SLIs and operational health indicators for business-critical services.
  • Automate infrastructure provisioning and application deployment workflows using Terraform, GitHub Actions, and ArgoCD.
  • Drive observability and telemetry adoption using Datadog, including dashboards, alerts, custom metrics, and distributed tracing.
  • Act as incident commander during critical production issues; conduct blameless postmortems and guide root cause remediation.
  • Lead cross-team efforts in reducing mean time to detect (MTTD) and resolve (MTTR), and promoting self-healing systems.
  • Partner with security and compliance teams to ensure that systems are secure, auditable, and operationally compliant.
  • Enhance service resiliency through strategies including Kafka-based event‑driven architecture, retries, rate limiting, and circuit breakers.
  • Mentor junior SREs and engineers, lead technical communities of practice, and promote a culture of continuous improvement.
  • Maintain and improve Ubuntu-based production systems and containerized workloads with Docker.
  • Evaluate and integrate emerging DevOps technologies to support scalability and reliability objectives.
Requirements
  • Bachelor's or Master's degree in Computer Science, Engineering, or a related technical field; equivalent practical experience may be considered.
  • 8+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering roles in large‑scale production environments.
  • Expertise in building and maintaining Java-based microservices using Spring Boot, Tomcat, and Redis in containerized deployments.
  • Strong hands‑on experience with Kubernetes, AKS, and ArgoCD for orchestration and GitOps deployment workflows.
  • Proficiency in Python, Java, Bash, or Go for automation, scripting, and infrastructure tooling.
  • Proven ability to implement observability platforms and practices using Datadog (metrics, logs, traces, dashboards, alerts).
  • Advanced experience working with CI/CD pipelines using GitHub and GitHub Actions.
  • Deep understanding of networking, Linux (especially Ubuntu), distributed systems, and container security.
  • Experience operating message‑driven architectures using Kafka, with an emphasis on throughput, retries, and resilience.
  • Solid knowledge of Terraform and infrastructure as code best practices.
  • Excellent communication, collaboration, and stakeholder alignment skills across engineering and business teams.

Salary Range: $146,960 ‑ $220,440

Actual compensation offered to a candidate may vary based on their unique qualifications and experience, internal equity, and market conditions. Final compensation decisions will be made in accordance with company policies and applicable laws.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Site Reliability Engineer
Principal Site Reliability Engineer

ViziRecruiter,LLC. • Quincy (MA)

Hybrid
USD 146,000 - 221,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

ViziRecruiter,LLC. • Chicago (IL)

Hybrid
USD 125,000 - 188,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

ViziRecruiter,LLC. • Salisbury (NC)

Hybrid
USD 125,000 - 188,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

ViziRecruiter,LLC. • Quincy (MA)

Hybrid
USD 125,000 - 188,000
Lead, Site Reliability Engineer
Lead, Site Reliability Engineer

CardWorks • Pittsburgh

Hybrid
USD 146,000 - 163,000
Competitive Pay
Medical, Dental, and Vision Benefits
401(k) Plan with Company Match
+1
SRE
SRE

Benton Partners • Chicago (IL), New York (NY)

On-site
USD 175,000 - 225,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mission Staffing • New York (NY)

Hybrid
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

GovCIO • Arlington (VA)

Hybrid
USD 210,000 - 230,000
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000