Platform SRE: Own & Automate Scalable Infra

Visa Hunt

San Francisco (CA)

On-site

USD 150,000 - 210,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

PostHog is seeking a Senior SRE to own and evolve a scalable, stateful platform hosted on AWS and Kubernetes. You will lead provisioning, scaling, recovery, and automation to reduce operational load while delivering reliable services.

You will tackle multi-region, multi-account environments with GitOps (ArgoCD), Terraform/Terragrunt, and modern tooling. Join a remote-first team that ships quickly and thinks deeply about reliability.

Qualifications

  • Deep hands-on experience with Kubernetes in production (EKS preferred).
  • Strong experience operating production infrastructure on AWS with multiple accounts.
  • Experience automating infrastructure using Terraform or Terragrunt at scale.
  • Solid understanding of Linux systems (disk, memory, networking, failure modes).
  • Experience supporting stateful systems (databases, queues, storage systems).
  • Ability to debug and reason about performance and reliability issues in production.
  • You're comfortable owning systems end-to-end, including on-call responsibilities.

Responsibilities

  • Turn a fast-growing, stateful system into a predictable, well-automated platform (provisioning, scaling, rebalancing, recovery).
  • Reduce operational stress by designing safe automation for traffic-heavy workloads and building tooling for scale.
  • Work on problems at large scale (petabytes of data, thousands of cores) across a multi-region AWS platform running Kubernetes.
  • Operate EKS clusters across environments with Karpenter autoscaling and ArgoCD-driven GitOps deployments.
  • Manage and evolve a multi-AWS account organization including networking and cross-account connectivity.
  • Maintain the Terraform/Terragrunt IaC platform and safe patterns for shared infrastructure.
  • Improve operational tooling around deployments, schema changes, backups, restores, and incident response.
  • Identify repeat pain points and eliminate them through code and self-healing automation.
  • Help optimize cloud spend as you go.
  • Participate in on-call and incident response with a focus on reducing incidents over time.

Skills

Kubernetes in prod
AWS expertise
On-call experience
Incident response
Linux administration
Troubleshooting at scale

Tools

Terraform
Terragrunt
Karpenter
ArgoCD
GitHub Actions
IAM & Networking (AWS)

Job description

PostHog is seeking a Senior SRE to own and evolve a scalable, stateful platform hosted on AWS and Kubernetes. You will lead provisioning, scaling, recovery, and automation to reduce operational load while delivering reliable services.

You will tackle multi-region, multi-account environments with GitOps (ArgoCD), Terraform/Terragrunt, and modern tooling. Join a remote-first team that ships quickly and thinks deeply about reliability.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote SRE – Scale, Automation & Ownership (US CT/ET)
Remote SRE – Scale, Automation & Ownership (US CT/ET)

PostHog • United States

On-site
USD 140,000 - 210,000
Remote SRE: Blockchain Infrastructure & Scale
Remote SRE: Blockchain Infrastructure & Scale

Embedded Shishya • United States

On-site
USD 150,000 - 210,000
Remote‑first global workforce
Professional reimbursement program
Medical, dental & vision coverage
+3
Senior Cloud SRE: Kubernetes, GitOps & Automation
Senior Cloud SRE: Kubernetes, GitOps & Automation

Pinterest • San Francisco (CA)

On-site
USD 140,000 - 288,000
Senior Infrastructure Architect & SRE
Senior Infrastructure Architect & SRE

Hanover Park • Northern (KY), New York (NY)

Hybrid
USD 180,000 - 240,000
Senior SRE — Postgres Infra for Cloud Data Platform
Senior SRE — Postgres Infra for Cloud Data Platform

ClickHouse, Inc. • Northern (KY)

Hybrid
USD 150,000 - 210,000
Flexible work environment
Healthcare
Equity in the company
+3
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Kovoro • Denver (CO), Northern (KY)

Hybrid
USD 150,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000
Site Reliability Engineer — Scale Data Infra on AWS & ClickHouse
Site Reliability Engineer — Scale Data Infra on AWS & ClickHouse

PostHog • United States

On-site
USD 140,000 - 210,000
Senior SRE: Automate, Own, and Stabilize AI Platform
Senior SRE: Automate, Own, and Stabilize AI Platform

Nscale • San Francisco (CA)

On-site
USD 130,000 - 200,000
Equity
Career growth
Flexible schedule
Platform SRE - Scale Kubernetes, Cloud Infra & Reliability
Platform SRE - Scale Kubernetes, Cloud Infra & Reliability

Specter • San Francisco (CA)

On-site
USD 180,000 - 230,000