Senior SRE: Scale, Reliability & Platform Observability

Sanity

United States

On-site

USD 140,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive stock options program
Comprehensive health plans
Flexible, trust-based environment
Location-based salary

Job summary

Sanity is building the future of AI-powered Content Operations. We seek a seasoned SRE to design, build, and operate scalable infrastructure across GCP and Kubernetes, ensuring high reliability and real-time content processing for global customers.

You will partner with development teams, improve observability, and drive incident response, automation, and deployment best practices to keep our platform fast, secure, and easy to use.

Qualifications

  • 5+ years of experience as part of an SRE on-call rotation.
  • Experience with managing scalable, highly available cloud-based applications.
  • Experience with Kubernetes for orchestrating and scaling containerized apps.
  • Experience building CI/CD pipelines.
  • Experience with an observability stack (Prometheus, etc.).

Responsibilities

  • Design, build, and operate core platform foundations on GCP, Kubernetes, networking, CI/CD, and observability.
  • Diagnose and troubleshoot complex distributed systems under high load.
  • Ensure observability and analyze stack behavior.
  • Contribute to edge, caching, gateway modernization on Fastly.
  • Raise reliability with dashboards, alerts, paging, on-call readiness, and incident response.
  • Make deployments boring by producing golden paths, readiness checks, and safe rollouts.
  • Mentor engineers and elevate technical standards through code reviews and pairing.
  • Participate in on-call rotation and ensure smooth on-call handoffs.

Skills

SRE/DevOps
Kubernetes
CI/CD pipelines
Observability
Cloud platforms
On-call
Incident response
Networking

Tools

Kubernetes
Prometheus
ElasticSearch
PostgreSQL
NATS
Kong
Fastly
Google Cloud Platform

Job description

Sanity is building the future of AI-powered Content Operations. We seek a seasoned SRE to design, build, and operate scalable infrastructure across GCP and Kubernetes, ensuring high reliability and real-time content processing for global customers.

You will partner with development teams, improve observability, and drive incident response, automation, and deployment best practices to keep our platform fast, secure, and easy to use.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: Scale AI Content Ops with Kubernetes
Senior SRE: Scale AI Content Ops with Kubernetes

Sanity CMS • United States

Remote
USD 100,000 - 140,000
Comprehensive health plans
Competitive stock options
Positive work environment
Senior SRE — Scale Reliability for Healthcare Data Platform
Senior SRE — Scale Reliability for Healthcare Data Platform

CertifyOS • United States

On-site
USD 120,000 - 170,000
Senior SRE: Scale & Reliability for AI-Driven SaaS Platform
Senior SRE: Scale & Reliability for AI-Driven SaaS Platform

Instrumental Inc. • Palo Alto (CA)

On-site
USD 175,000 - 229,000
Health benefits
Commuter plans
Parental leave
Senior SRE: Scale, Reliability & Observability
Senior SRE: Scale, Reliability & Observability

Google • Raleigh (NC)

On-site
USD 174,000 - 252,000
Senior SRE - Hybrid, Platform Reliability Lead
Senior SRE - Hybrid, Platform Reliability Lead

TransUnion LLC • Reston (VA)

Hybrid
USD 112,000 - 188,000
Day-one medical, dental, vision
Company-paid basic life/AD&D
12 weeks paid parental leave
+2
Senior SRE: Scale Reliability & Observability
Senior SRE: Scale Reliability & Observability

Megaport • Abbeyville (CO)

On-site
USD 130,000 - 190,000
Contractor (PJ)
Paid Time Off
Competitive Compensation
+4
Senior SRE: Scalable Infra, Observability & Automation
Senior SRE: Scalable Infra, Observability & Automation

Early Warning • Scottsdale (AZ)

Hybrid
USD 106,000 - 156,000
Healthcare Coverage
401(k) Company Match
Paid Time Off
+2
Senior SRE: Scale Infra, Automate, Elevate Reliability
Senior SRE: Scale Infra, Automate, Elevate Reliability

Fathom.ai • United States

On-site
USD 100,000 - 130,000
Competitive compensation
Supportive environment for personal growth
Dynamic and collaborative team
Senior SRE: Scale Reliability, Observability & CI/CD
Senior SRE: Scale Reliability, Observability & CI/CD

Breakout Tools • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior SRE - Cloud & Observability
Senior SRE - Cloud & Observability

Ridgeline • Reno (NV)

Hybrid
USD 153,000 - 210,000
Unlimited vacation
Education reimbursement
Wellness reimbursement
+1