Staff Site Reliability Engineer - Platform & Observability

Pantheon

United States

Remote

USD 150,000 - 220,000

Full time

8 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Pantheon is seeking a Staff Software Engineer with a strong SRE orientation to elevate platform reliability. You will shape SRE practices, define SLO/SLI frameworks, and standardize observability across teams.

The role centers on Go, cloud primitives, and a Grafana-centered stack to support fast, safe shipping for 15+ engineering teams. The ideal candidate leads incident response, champions reliability culture, and collaborates across the Internal Platform Group to raise the bar for product

Qualifications

  • Staff level software engineer with a strong SRE orientation.
  • Experience building reliable platforms and incident response culture.
  • Familiarity with Go and cloud-native observability patterns.

Responsibilities

  • Establish SRE as a discipline across PIE and the Internal Platform Group.
  • Define and drive adoption of SLO/SLI frameworks and reliability standards.
  • Improve overall platform reliability and observability practices.

Skills

Go
SRE
Observability
Prometheus
OpenTelemetry
Terraform
GCP
Distributed Systems

Tools

Grafana
Cloud Run
GKE

Job description

Pantheon is seeking a Staff Software Engineer with a strong SRE orientation to elevate platform reliability. You will shape SRE practices, define SLO/SLI frameworks, and standardize observability across teams.

The role centers on Go, cloud primitives, and a Grafana-centered stack to support fast, safe shipping for 15+ engineering teams. The ideal candidate leads incident response, champions reliability culture, and collaborates across the Internal Platform Group to raise the bar for product

Get your free, confidential resume review.

or drag and drop your file here.