Site Reliability Engineer

United States Cold Storage, Inc.

Philadelphia (Philadelphia County)

Hybrid

USD 130,000 - 150,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

US Cold owns and operates one of the most complex temperature-controlled logistics networks in North America, coordinating the storage and movement of food at national scale across automated facilities.

The Site Reliability Engineer will be a founding member of US Cold’s SRE practice, driving engineered reliability, observability, and automation to reduce incidents and toil while shaping SRE practices across cloud and on‑prem environments.

Qualifications

  • 3+ years of experience in SRE, DevOps, Systems Engineering, or related roles.
  • Strong Linux and Windows systems administration and troubleshooting skills.
  • Hands-on experience with automation and scripting.
  • Experience designing and operating monitoring, alerting, and observability solutions.
  • Practical experience working in Azure environments.
  • Strong analytical skills and a bias toward eliminating root causes, not symptoms.
  • Ability to collaborate across application, infrastructure, and operations teams.
  • Experience supporting warehouse management systems or industrial automation platforms.
  • Exposure to Kubernetes, microservices, or container orchestration.
  • Hands on experience with infrastructure-as-code tools such as Terraform or Ansible.
  • Understanding of distributed systems and high-availability design.
  • Experience with SRE practices such as SLO-based operations, runbook automation, or chaos testing.

Responsibilities

  • Ensure reliability of the Phenix WMS and its integration with facility automation systems (robotics, conveyors, and control interfaces).
  • Define and implement SLIs and SLOs that measure meaningful system health, not just availability.
  • Establish observability across the full stack—cloud services, APIs, and on-premise operations.
  • Automate to eliminate toil, including patching, data corrections, restarts, and recovery tasks.
  • Develop self-healing behaviors for common failure modes.
  • Participate in on-call rotations and lead blameless post-incident reviews.
  • Design and execute disaster recovery tests across SaaS, cloud, and on-premise environments.

Skills

SRE experience
Linux administration
Automation scripting
Observability
Azure
Kubernetes
IaC (Terraform/Ansible)
Distributed systems
Runbook automation

Tools

Terraform
Ansible
Kubernetes
Azure

Job description

Site Reliability Engineer (SRE)

Engineer Reliability into the Systems That Move the Nation’s Food Supply

Who We Are

US Cold owns and operates one of the most complex temperature-controlled logistics networks in North America. Every day, our systems coordinate the storage and movement of food at national scale across a network of state-of-the-art distribution centers, including multiple highly automated warehouse facilities. We continue to advance our core warehouse and logistics platforms. Our current focus is on modular, event-driven, API-first andcloudarchitectures. We continue to enhance reliability and accelerate engineering productivity by strengthening our SRE and AI practices. This is a large investment in innovation to continue to drive operational excellence at our facilities. If you want to build durable systems that operate in the physical world at scale, this is that opportunity.

The Role

The Site Reliability Engineer is a founding member of US Cold’s SRE practice. This role exists to move the organization from reactive operations to engineered reliability. You will study how our most critical systems fail – particularly our Phenix WMS and facility automation interfaces – and design controls, automation, and observability that reduce incidents over time. Success in this role means fewer false alerts, faster recovery, less manual intervention, and systems that heal themselves when possible. You will work closely with application, infrastructure, and operations teams and participate directly in on‑call and incident response.

What You Will Own
  • Reliability of the Phenix WMS and its integration with facility automation systems (robotics, conveyors, and control interfaces)
  • Definition and implementation of SLIs and SLOs that measure meaningful system health, not just availability
  • Observability across the full stack, correlating cloud services, APIs, and on‑premise facility operations
  • Automation to eliminate operational toil, including patching, data corrections, restarts, and recovery tasks
  • Development of self‑healing behaviors for common failure modes
  • Participation in on‑call rotations and leadership of blameless post‑incident reviews
  • Design and execution of disaster recovery tests across SaaS, cloud, and on‑premise environments
Technical Environment
  • Hybrid environments spanning cloud and on‑premise infrastructure
  • Azure cloud services
  • Warehouse Management Systems (Phenix WMS) and facility automation interfaces
  • Java Development
  • Observability tooling across logs, metrics, and alerting
  • Automation using Python, PowerShell, Bash, or Ansible
  • CI/CD tools and modern deployment practices
  • Exposure to containerized and distributed systems environments
What We’re Looking For
  • 3+ years of experience in SRE, DevOps, Systems Engineering, or related roles
  • Strong Linux and Windows systems administration and troubleshooting skills
  • Hands‑on experience with automation and scripting
  • Experience designing and operating monitoring, alerting, and observability solutions
  • Practical experience working in Azure environments
  • Strong analytical skills and a bias toward eliminating root causes, not symptoms
  • Ability to collaborate across application, infrastructure, and operations teams
  • Experience supporting warehouse management systems or industrial automation platforms
  • Exposure to Kubernetes, microservices, or container orchestration
  • Hands on experience with infrastructure‑as‑code tools such as Terraform or Ansible
  • Understanding of distributed systems and high‑availability design
  • Experience with SRE practices such as SLO‑based operations, runbook automation, or chaos testing
Why This Role Is Different

This is not an inherited SRE function. There is no mature framework to maintain. You will:

  • Help define what reliability means at US Cold
  • Work on systems that operate in the physical world
  • Engineer solutions that reduce toil and operational load
  • See the direct impact of your work on warehouse uptime and performance
  • Build practices that scale as the platform modernizes
Compensation & Structure
  • Location:Hybrid - Camden NJ
  • Reports to: IT - Site Reliability Engineering Manager
  • Salary Range: $130,000- $150,000
Operational Context
  • Systems operate continuously across warehouse facilities
  • Reliability failures have physical and operational consequences
  • On‑call participation is part of the role
  • Work occurs across cloud, SaaS, and on‑premise environments

Equal Opportunity Employer

This employer is required to notify all applicants of their rights pursuant to federal employment laws. For further information, please review the Know Your Rights notice from the Department of Labor.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

United States Cold Storage, Inc. • Camden (NJ)

Hybrid
USD 130,000 - 150,000
Senior DevOps Engineer
Senior DevOps Engineer

United States Cold Storage, Inc. • Camden (NJ)

Hybrid
USD 125,000 - 145,000
Medical
Dental
Vision
+2
Senior DevOps Engineer
Senior DevOps Engineer

United States Cold Storage Inc • Camden (NJ)

Hybrid
USD 125,000 - 145,000
Medical benefits
Dental & Vision
Hybrid work model
+1
Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

VITG • Ellicott City (MD)

Hybrid
USD 90,000 - 120,000
401(k) with employer contribution
Medical/Dental/Vision insurance
Paid vacation (PTO)
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Core Scientific, Inc • Austin (TX)

Hybrid
USD 140,000 - 190,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

ViziRecruiter,LLC. • Salisbury (NC)

Hybrid
USD 146,000 - 221,000
Site Reliability Engineer -- SINDC5717546
Site Reliability Engineer -- SINDC5717546

Compunnel Inc. • Denton (TX)

Hybrid
USD 120,000 - 150,000
SRE: Build Self-Healing, Scalable Systems (Hybrid)
SRE: Build Self-Healing, Scalable Systems (Hybrid)

United States Cold Storage, Inc. • Philadelphia

Hybrid
USD 130,000 - 150,000
Lead, Site Reliability Engineer
Lead, Site Reliability Engineer

CardWorks • Pittsburgh

Hybrid
USD 146,000 - 163,000
Competitive Pay
Medical, Dental, and Vision Benefits
401(k) Plan with Company Match
+1
Sr Software Engineer - Reliability Engineering
Sr Software Engineer - Reliability Engineering

Cox Enterprises • Village of North Hills (NY)

On-site
USD 150,000 - 185,000