Senior SRE: Scalable, Reliable Cloud Platform

Stord

United States

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Stord is seeking an experienced Site Reliability Engineer to own scalable infra on GCP, manage Terraform modules, and optimize Kubernetes workloads. You will implement observability with Datadog, design robust CI/CD pipelines, and drive cost-efficient, reliable platform improvements.

You will collaborate with cross‑functional teams to improve deployment practices, participate in incident post‑mortems, and raise the bar for reliability as the company scales.

Qualifications

  • 5+ years in SRE, platform, or infrastructure engineering: you've owned complex systems and driven technical work forward with minimal supervision.
  • GCP depth: strong hands‑on experience with GCP core services (GKE, Cloud Run, AlloyDB, networking, IAM).
  • Containers & orchestration: you're fluent in Docker and Kubernetes and can debug, tune, and scale real workloads.
  • Infrastructure as Code: deep Terraform experience. You write reusable modules, reason about state, and treat infra changes with the same rigor as application code.
  • A programming language you're genuinely productive in: TypeScript, Python, Go, or similar, used to build tooling and automation, not just glue scripts.
  • Observability: you build monitoring and alerting that's actionable (Datadog, or equivalents like Prometheus/Grafana), and you know the difference between a noisy dashboard and a useful one.
  • Distributed systems fundamentals: failure modes, consistency, and how systems break at scale.
  • Git and collaborative development workflows: you work in shared codebases and review others' changes well.
  • Incident management: you've run incidents and post‑mortems and can stay calm and methodical when production is on fire.
  • Ownership & Accountability: You own features end‑to‑end and take pride in what you ship.
  • Strong Communication: You can explain technical decisions and trade‑offs to engineers, PMs, and stakeholders.
  • Collaborative Approach: You work well with others, give constructive code review feedback, and actively seek input from teammates.
  • Production Mindset: You prioritize reliability and user impact.
  • Learning Agility: You're comfortable with rapidly evolving AI/ML technologies and tools.
  • Directed AI‑Assisted Development: You know how to use AI coding tools as a productivity multiplier while maintaining quality.
  • Database operations depth: PostgreSQL internals or migrations and scaling.
  • Event‑driven systems: Kafka/Redpanda or Pub/Sub, schema registries, and streaming at scale.
  • Cost engineering: you've meaningfully reduced cloud or observability spend without sacrificing reliability.
  • GCP certifications or equivalent depth
  • Cloudflare: experience with Workers and other Cloudflare services
  • Multi‑cloud or hybrid architecture exposure

Responsibilities

  • Own architecture and implementation of scalable, reliable infrastructure on GCP, including GKE, Cloud Run, AlloyDB, and networking.
  • Own Infrastructure as Code in Terraform: modules, org policies, and the patterns the team builds on.
  • Manage containerized workloads on Kubernetes, including performance tuning, capacity planning, and resource optimization.
  • Drive down cost and toil through better defaults, right‑sizing, and automation rather than manual intervention.
  • Build monitoring, alerting, and observability in Datadog (APM, logs, RUM) that catches problems before customers do.
  • Define the reliability signals that matter for the services you own, and hold the line on them.
  • Develop and maintain disaster recovery and business‑continuity strategies, and prove they work.
  • Design and maintain CI/CD pipelines in GitHub Actions, including runner strategy and deployment safety.
  • Automate operational workflows and infrastructure provisioning so the platform scales smoothly as the team grows.
  • Build custom tooling and scripts that remove recurring operational pain.
  • Partner with data and development teams to improve deployment practices and application reliability.
  • Provide escalation support for production incidents, help lead post‑incident reviews, and turn findings into durable fixes.
  • Participate in technical design reviews and offer architectural input across teams.
  • Help improve SRE and infrastructure best practices across the team, and participate in on‑call for critical systems.

Skills

5+ years SRE
GCP core services
Docker & Kubernetes
Terraform / IaC
TS/Python/Go tooling
Observability (Datadog)
Distributed systems
Git workflows
Incident management

Tools

Docker Tools
Kubernetes
Terraform
GitHub Actions
Datadog

Job description

Stord is seeking an experienced Site Reliability Engineer to own scalable infra on GCP, manage Terraform modules, and optimize Kubernetes workloads. You will implement observability with Datadog, design robust CI/CD pipelines, and drive cost-efficient, reliable platform improvements.

You will collaborate with cross‑functional teams to improve deployment practices, participate in incident post‑mortems, and raise the bar for reliability as the company scales.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: Cloud Platform Reliability & Automation
Senior SRE: Cloud Platform Reliability & Automation

Stord, Inc. • United States

Remote
USD 140,000 - 180,000
Senior Site Reliability Engineer — Remote, Cloud Infra Lead
Senior Site Reliability Engineer — Remote, Cloud Infra Lead

The Consensus • United States

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

stord • United States

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Stord • United States

On-site
USD 180,000 - 240,000
Senior SRE: Cloud Reliability, Automation & Incidents
Senior SRE: Cloud Reliability, Automation & Incidents

Stelvio Inc. • Town of Texas (WI)

On-site
USD 125,000 - 145,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The Consensus • United States

On-site
USD 140,000 - 190,000
Senior SRE — Scale Reliability for Healthcare Data Platform
Senior SRE — Scale Reliability for Healthcare Data Platform

CertifyOS • United States

On-site
USD 120,000 - 170,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Kovoro • Denver (CO), Northern (KY)

Hybrid
USD 150,000 - 190,000
Senior Observability & SRE Engineer (GCP/Kubernetes)
Senior Observability & SRE Engineer (GCP/Kubernetes)

Ontrac Solutions • New York (NY)

On-site
USD 120,000 - 190,000
Senior SRE Lead: Cloud Reliability & Automation
Senior SRE Lead: Cloud Reliability & Automation

Oracle • Vienna (VA)

On-site
USD 96,000 - 265,000
Medical, dental, vision insurance
401(k) with company match
Paid time off and holidays
+1