Senior Site Reliability Engineer

Stord

United States

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Stord is seeking an experienced Site Reliability Engineer to own scalable infra on GCP, manage Terraform modules, and optimize Kubernetes workloads. You will implement observability with Datadog, design robust CI/CD pipelines, and drive cost-efficient, reliable platform improvements.

You will collaborate with cross‑functional teams to improve deployment practices, participate in incident post‑mortems, and raise the bar for reliability as the company scales.

Qualifications

  • 5+ years in SRE, platform, or infrastructure engineering: you've owned complex systems and driven technical work forward with minimal supervision.
  • GCP depth: strong hands‑on experience with GCP core services (GKE, Cloud Run, AlloyDB, networking, IAM).
  • Containers & orchestration: you're fluent in Docker and Kubernetes and can debug, tune, and scale real workloads.
  • Infrastructure as Code: deep Terraform experience. You write reusable modules, reason about state, and treat infra changes with the same rigor as application code.
  • A programming language you're genuinely productive in: TypeScript, Python, Go, or similar, used to build tooling and automation, not just glue scripts.
  • Observability: you build monitoring and alerting that's actionable (Datadog, or equivalents like Prometheus/Grafana), and you know the difference between a noisy dashboard and a useful one.
  • Distributed systems fundamentals: failure modes, consistency, and how systems break at scale.
  • Git and collaborative development workflows: you work in shared codebases and review others' changes well.
  • Incident management: you've run incidents and post‑mortems and can stay calm and methodical when production is on fire.
  • Ownership & Accountability: You own features end‑to‑end and take pride in what you ship.
  • Strong Communication: You can explain technical decisions and trade‑offs to engineers, PMs, and stakeholders.
  • Collaborative Approach: You work well with others, give constructive code review feedback, and actively seek input from teammates.
  • Production Mindset: You prioritize reliability and user impact.
  • Learning Agility: You're comfortable with rapidly evolving AI/ML technologies and tools.
  • Directed AI‑Assisted Development: You know how to use AI coding tools as a productivity multiplier while maintaining quality.
  • Database operations depth: PostgreSQL internals or migrations and scaling.
  • Event‑driven systems: Kafka/Redpanda or Pub/Sub, schema registries, and streaming at scale.
  • Cost engineering: you've meaningfully reduced cloud or observability spend without sacrificing reliability.
  • GCP certifications or equivalent depth
  • Cloudflare: experience with Workers and other Cloudflare services
  • Multi‑cloud or hybrid architecture exposure

Responsibilities

  • Own architecture and implementation of scalable, reliable infrastructure on GCP, including GKE, Cloud Run, AlloyDB, and networking.
  • Own Infrastructure as Code in Terraform: modules, org policies, and the patterns the team builds on.
  • Manage containerized workloads on Kubernetes, including performance tuning, capacity planning, and resource optimization.
  • Drive down cost and toil through better defaults, right‑sizing, and automation rather than manual intervention.
  • Build monitoring, alerting, and observability in Datadog (APM, logs, RUM) that catches problems before customers do.
  • Define the reliability signals that matter for the services you own, and hold the line on them.
  • Develop and maintain disaster recovery and business‑continuity strategies, and prove they work.
  • Design and maintain CI/CD pipelines in GitHub Actions, including runner strategy and deployment safety.
  • Automate operational workflows and infrastructure provisioning so the platform scales smoothly as the team grows.
  • Build custom tooling and scripts that remove recurring operational pain.
  • Partner with data and development teams to improve deployment practices and application reliability.
  • Provide escalation support for production incidents, help lead post‑incident reviews, and turn findings into durable fixes.
  • Participate in technical design reviews and offer architectural input across teams.
  • Help improve SRE and infrastructure best practices across the team, and participate in on‑call for critical systems.

Skills

5+ years SRE
GCP core services
Docker & Kubernetes
Terraform / IaC
TS/Python/Go tooling
Observability (Datadog)
Distributed systems
Git workflows
Incident management

Tools

Docker Tools
Kubernetes
Terraform
GitHub Actions
Datadog

Job description

Stord is The Consumer Experience Company, powering seamless checkout through delivery for today's leading brands. Stord is rapidly growing and is on track to double our revenue in the next 18 months. To meet and exceed this target, Stord is strategically scaling teams across the entire company, and seeking energetic experts to help us achieve our mission.

By combining comprehensive commerce-enablement technology with high-volume fulfillment services, Stord provides brands a platform to compete with retail giants. Stord manages over $10 billion of commerce annually through its fulfillment, warehousing, transportation, and operator-built software suite including OMS, Pre- and Post-Purchase, and WMS platforms. Stord is leveling the playing field for all brands to deliver the best consumer experience at scale.

With Stord, brands can increase cart conversion, improve unit economics, and drive sustained customer loyalty. Stord’s end-to-end commerce solutions combine best-in-class omnichannel fulfillment and shipping with leading technology to ensure fast shipping, reliable delivery promises, easy access to more channels, and improved margins on every order.

Hundreds of leading DTC and B2B companies like AG1, True Classic, Native, Seed Health, quip, goodr, Sundays for Dogs, and more trust Stord to deliver industry‑leading consumer experiences on every order. Stord is headquartered in Atlanta with facilities across the United States, Canada, and Europe. Stord is backed by top‑tier investors including Kleiner Perkins, Franklin Templeton, Founders Fund, Strike Capital, Baillie Gifford, and Salesforce Ventures.

Stord is building the operating system for modern supply chain: a unified platform handling Order Management, Warehouse Management, Transportation, and Consumer Experience for brands doing over $10B in commerce annually.

The SRE team is small, fast‑moving, and owns the infrastructure that keeps that platform running, primarily on Google Cloud Platform, across GKE, Cloud Run, AlloyDB, and the networking that connects our services. This is a high‑autonomy environment with a wide surface area and few layers between you and the systems you're responsible for. The work is real infrastructure engineering: reliability, scale, cost, and the automation that lets a lean team punch well above its weight.

This role is for an engineer who wants to help a startup grow up. We're maturing fast, and SRE is central to that path: turning ad‑hoc fixes into repeatable process, replacing toil with automation, and building the reliability practices a scaling platform depends on. You'll work as part of a team that values clear communication and shared ownership. You'll also act as the technical bridge between development and operations. If you take pride in building durable systems and raising the bar for how a team operates, this role was built for you.

Why This Role
  • You’ll own high‑impact infrastructure work directly, in a small team where your contributions are visible and your judgment is trusted. The path from decision to production is short.
  • The surface area is broad and real: GKE, Cloud Run, GCP core services, AlloyDB, and the CI/CD and observability tooling that ties it together.
  • You’ll shape reliability and automation practices during a period of real growth, working alongside engineers who care about doing it well.
What You’ll Build
Infrastructure & Platform
  • Own architecture and implementation of scalable, reliable infrastructure on GCP, including GKE, Cloud Run, AlloyDB, and networking.
  • Own Infrastructure as Code in Terraform: modules, org policies, and the patterns the team builds on.
  • Manage containerized workloads on Kubernetes, including performance tuning, capacity planning, and resource optimization.
  • Drive down cost and toil through better defaults, right‑sizing, and automation rather than manual intervention.
Reliability & Observability
  • Build monitoring, alerting, and observability in Datadog (APM, logs, RUM) that catches problems before customers do.
  • Define the reliability signals that matter for the services you own, and hold the line on them.
  • Develop and maintain disaster recovery and business‑continuity strategies, and prove they work.
Automation & Delivery
  • Design and maintain CI/CD pipelines in GitHub Actions, including runner strategy and deployment safety.
  • Automate operational workflows and infrastructure provisioning so the platform scales smoothly as the team grows.
  • Build custom tooling and scripts that remove recurring operational pain.
Collaboration & Incident Response
  • Partner with data and development teams to improve deployment practices and application reliability.
  • Provide escalation support for production incidents, help lead post‑incident reviews, and turn findings into durable fixes.
  • Participate in technical design reviews and offer architectural input across teams.
  • Help improve SRE and infrastructure best practices across the team, and participate in on‑call for critical systems.
What We’re Looking For
Required
  • 5+ years in SRE, platform, or infrastructure engineering: you've owned complex systems and driven technical work forward with minimal supervision.
  • GCP depth: strong hands‑on experience with GCP core services (GKE, Cloud Run, AlloyDB, networking, IAM). You know how these fit together in production, not just in a certification.
  • Containers & orchestration: you're fluent in Docker and Kubernetes and can debug, tune, and scale real workloads.
  • Infrastructure as Code: deep Terraform experience. You write reusable modules, reason about state, and treat infra changes with the same rigor as application code.
  • A programming language you're genuinely productive in: TypeScript, Python, Go, or similar, used to build tooling and automation, not just glue scripts.
  • Observability: you build monitoring and alerting that's actionable (Datadog, or equivalents like Prometheus/Grafana), and you know the difference between a noisy dashboard and a useful one.
  • Distributed systems fundamentals: failure modes, consistency, and how systems break at scale.
  • Git and collaborative development workflows: you work in shared codebases and review others' changes well.
  • Incident management: you've run incidents and post‑mortems and can stay calm and methodical when production is on fire.
Required Soft Skills
  • Ownership & Accountability: You own features end‑to‑end and take pride in what you ship. You follow through from design to production and don't drop things.
  • Strong Communication: You can explain technical decisions and trade‑offs to engineers, PMs, and stakeholders. You ask good questions and listen well.
  • Collaborative Approach: You work well with others, give constructive code review feedback, and actively seek input from teammates.
  • Production Mindset: You prioritize reliability and user impact. You think about failure modes, monitoring, and operational concerns as part of your design process.
  • Learning Agility: You're comfortable with rapidly evolving AI/ML technologies and tools. You stay current without chasing hype.
  • Directed AI‑Assisted Development: You know how to use AI coding tools as a productivity multiplier while maintaining quality and your own technical judgment.
Strongly Preferred
  • Database operations depth: PostgreSQL internals (logical replication, vacuuming, lock contention) or experience with migrations and database scaling. Familiarity with Redis, ClickHouse, or analytical stores is a plus.
  • Event‑driven systems: Kafka/Redpanda or Pub/Sub, schema registries, and the operational realities of streaming at scale.
  • Cost engineering: you've meaningfully reduced cloud or observability spend without sacrificing reliability.
Nice to Have
  • GCP certifications (Cloud Architect, Cloud DevOps Engineer) or demonstrably equivalent depth.
  • Cloudflare: experience with Workers and other Cloudflare services.
  • Multi‑cloud or hybrid architecture exposure.
What Success Looks Like
  • 30 days: You've ramped on our GCP environment and Terraform setup, shipped your first infrastructure or automation change to production, and are contributing in incident and design discussions.
  • 90 days: You independently own a meaningful slice of the infrastructure, have improved a reliability, cost, or automation pain point that was slowing the team down, and teammates lean on you in your area.
  • 6 months: You're a trusted owner of your part of the infrastructure, consistently delivering improvements that make the platform more reliable and efficient.
About Stord

Stord is a cloud‑based supply chain platform that enables brands to compete and grow through end‑to‑end logistics solutions. We process over $10B in commerce annually and operate across Order Management (OMS), Warehouse Management (WMS), Transportation Management (TMS), Consumer Experience, and Demand Planning. We are backed by leading investors and are rapidly scaling our engineering organization to match our ambitions.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

stord • United States

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The Consensus • United States

On-site
USD 140,000 - 190,000
Senior Engineer, Platform Services
Senior Engineer, Platform Services

Stord • United States

On-site
USD 120,000 - 180,000
Engineering & Product Senior Engineer, Platform Services Remote, United States
Engineering & Product Senior Engineer, Platform Services Remote, United States

Stord • Northern (KY)

Hybrid
USD 160,000 - 230,000
Forward Deployed Engineer, AI Enablement
Forward Deployed Engineer, AI Enablement

Stord • United States

On-site
USD 180,000 - 280,000
Director Engineering, Logistics
Director Engineering, Logistics

Stord • Atlanta (GA)

On-site
USD 210,000 - 320,000
Engineering Manager, SDET
Engineering Manager, SDET

Stord • United States

On-site
USD 180,000 - 240,000
Senior Software Engineer
Senior Software Engineer

Stord • United States

On-site
USD 120,000 - 190,000
Implementation Engineer
Implementation Engineer

Stord, Inc. • Hebron Estates (KY)

Hybrid
USD 90,000 - 130,000
Implementation Engineer
Implementation Engineer

Stord-Warehous • Hebron Estates (KY)

On-site
USD 90,000 - 130,000