Senior Site Reliability Engineer

stord

United States

On-site

USD 140,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Stord is building the operating system for modern supply chains. The SRE team owns the infrastructure that keeps our platform running, primarily on Google Cloud Platform, across GKE, Cloud Run, AlloyDB, and the networking that connects our services. This role emphasizes reliability, automation, and scalable design in a fast-growing startup environment.

You will partner with development, improve deployment practices, and lead post-incident reviews to raise the bar for reliability and performance.

Qualifications

  • Experience building scalable, reliable infrastructure on GCP.
  • Strong knowledge of Kubernetes (GKE) and container orchestration.
  • Proficient with Terraform for IaC.
  • Experience with Datadog for monitoring and alerting.
  • Hands-on with CI/CD pipelines, preferably GitHub Actions.
  • Ability to diagnose incidents and implement durable fixes.

Responsibilities

  • Own architecture and implementation of scalable, reliable infrastructure on GCP, including GKE, Cloud Run, AlloyDB.
  • Manage Terraform modules and IaC patterns.
  • Configure and optimize Kubernetes workloads and networking.
  • Build monitoring and alerting to catch issues before customers.
  • Design disaster recovery and business-continuity strategies.
  • Develop CI/CD pipelines and automate operational workflows.
  • Collaborate with data and development teams and lead post-incident reviews.

Skills

Automation
Incident response
On-call readiness
Collaboration

Tools

GKE
Cloud Run
AlloyDB
Terraform
Datadog

Job description

Stord is The Consumer Experience Company, powering seamless checkout through delivery for today's leading brands. Stord is rapidly growing and is on track to double our revenue in the next 18 months. To meet and exceed this target, Stord is strategically scaling teams across the entire company, and seeking energetic experts to help us achieve our mission.

By combining comprehensive commerce‑enablement technology with high‑volume fulfillment services, Stord provides brands a platform to compete with retail giants. Stord manages over $10 billion of commerce annually through its fulfillment, warehousing, transportation, and operator‑built software suite—including OMS, Pre‑ and Post‑Purchase, and WMS platforms. Stord is leveling the playing field for all brands to deliver the best consumer experience at scale.

With Stord, brands can increase cart conversion, improve unit economics, and drive sustained customer loyalty. Stord's end‑to‑end commerce solutions combine best‑in‑class omnichannel fulfillment and shipping with leading technology to ensure fast shipping, reliable delivery promises, easy access to more channels, and improved margins on every order.

Hundreds of leading DTC and B2B companies like AG1, True Classic, Native, Seed Health, quip, goodr, Sundays for Dogs, and more trust Stord to deliver industry‑leading consumer experiences on every order. Stord is headquartered in Atlanta with facilities across the United States, Canada, and Europe. Stord is backed by top‑tier investors including Kleiner Perkins, Franklin Templeton, Founders Fund, Strike Capital, Baillie Gifford, and Salesforce Ventures.

Stord is building the operating system for modern supply chain: a unified platform handling Order Management, Warehouse Management, Transportation, and Consumer Experience for brands doing over $10B in commerce annually.

The SRE Team

The SRE team is small, fast‑moving, and owns the infrastructure that keeps that platform running, primarily on Google Cloud Platform, across GKE, Cloud Run, AlloyDB, and the networking that connects our services. This is a high‑autonomy environment with a wide surface area and few layers between you and the systems you’re responsible for. The work is real infrastructure engineering: reliability, scale, cost, and the automation that lets a lean team punch well above its weight.

Why This Role

This role is for an engineer who wants to help a startup grow up. We're maturing fast, and SRE is central to that path: turning ad hoc fixes into repeatable processes, replacing toil with automation, and building the reliability practices a scaling platform depends on. You’ll work as part of a team that values clear communication and shared ownership. You’ll also act as the technical bridge between development and operations. If you take pride in building durable systems and raising the bar for how a team operates, this role was built for you.

What You’ll Build
Infrastructure & Platform
  • Own architecture and implementation of scalable, reliable infrastructure on GCP, including GKE, Cloud Run, AlloyDB, and networking.
  • Own Infrastructure as Code in Terraform: modules, org policies, and the patterns the team builds on.
  • Manage containerized workloads on Kubernetes, including performance tuning, capacity planning, and resource optimization.
  • Drive down cost and toil through better defaults, right‑sizing, and automation rather than manual intervention.
Reliability & Observability
  • Build monitoring, alerting, and observability in Datadog (APM, logs, RUM) that catches problems before customers do.
  • Define the reliability signals that matter for the services you own, and hold the line on them.
  • Develop and maintain disaster recovery and business‑continuity strategies, and prove they work.
Automation & Delivery
  • Design and maintain CI/CD pipelines in GitHub Actions, including runner strategy and deployment safety.
  • Automate operational workflows and infrastructure provisioning so the platform scales smoothly as the team grows.
  • Build custom tooling and scripts that remove recurring operational pain.
Collaboration & Incident Response
  • Partner with data and development teams to improve deployment practices and application reliability.
  • Provide escalation support for production incidents, help lead post‑incident reviews, and turn findings into durable fixes.
  • Participate in technical design reviews and offer architectural input across teams.
  • Help improve SRE and infrastructure best practices across the team, and participate in on‑call for critical systems.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Stord • United States

On-site
USD 180,000 - 240,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The Consensus • United States

On-site
USD 140,000 - 190,000
Senior SRE: Scale, Automation & Cloud Reliability
Senior SRE: Scale, Automation & Cloud Reliability

PVH (Tommy Hilfiger/Calvin Klein) • United States

On-site
USD 140,000 - 190,000
Senior SRE: Scalable, Reliable Cloud Platform
Senior SRE: Scalable, Reliable Cloud Platform

Stord • United States

On-site
USD 180,000 - 240,000
Forward Deployed Engineer, AI Enablement
Forward Deployed Engineer, AI Enablement

Stord • United States

On-site
USD 180,000 - 280,000
Engineering Manager, WMS
Engineering Manager, WMS

慨正橡扯 • Atlanta (GA)

On-site
USD 180,000 - 240,000
Director Engineering, Logistics
Director Engineering, Logistics

B Capital • United States

On-site
USD 180,000 - 260,000
IT Operations Specialist
IT Operations Specialist

Stord • Irvine (CA)

On-site
USD 85,000 - 120,000
Director Engineering, Logistics
Director Engineering, Logistics

Stord • Atlanta (GA)

On-site
USD 210,000 - 320,000
Engineering Manager, WMS
Engineering Manager, WMS

B Capital • United States

On-site
USD 180,000 - 240,000