Staff SRE: Scale, Observability & Kubernetes

Replit

Foster City (CA)

On-site

USD 180,000 - 260,000

Full time

3 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Competitive Salary & Equity
401(k) with 4% match
Health, Dental, Vision and Life Ins.
Short/Long Disability
Parental, Medical, Caregiver Leave

Job summary

Replit is seeking aStaff Site Reliability Engineer to ensure reliability, scalability, and performance of its global infrastructure. You will lead incident response, design robust observability, and drive automation across Kubernetes-based deployments.

You will mentor engineers while collaborating with product teams to maintain high availability and drive reliability as a core company value.

Qualifications

  • 8–10 years of experience in SRE or similar roles.
  • Strong programming in Python or Go with well-tested code.
  • Deep understanding of distributed systems and SOA.
  • Experience with Kubernetes and cloud-native tech.
  • Proven observability and incident-management track record.
  • Experience with Terraform, Pulumi, and IaC
  • Excellent written and verbal communication; mentoring skills.

Responsibilities

  • Architect and implement observability, dashboards, and metrics.
  • Define SLOs/SLIs and monitor reliability standards.
  • Lead incident management with blameless post-mortems.
  • Automate operations and build CI/CD pipelines.
  • Optimize Kubernetes deployments and latency globally.
  • Debug, harden, and design long-term fixes for distributed systems.
  • Provide staff-level design review and guidance.
  • Educate and mentor engineers on reliability practices.
  • Write high-quality Python/Go code for internal tools.

Skills

Python
Go
Distributed systems
Kubernetes
Incident management
Terraform
Pulumi
Observability
Communication
Mentoring

Tools

Prometheus
Grafana
Datadog
OpenTelemetry
GCP

Job description

Replit is seeking aStaff Site Reliability Engineer to ensure reliability, scalability, and performance of its global infrastructure. You will lead incident response, design robust observability, and drive automation across Kubernetes-based deployments.

You will mentor engineers while collaborating with product teams to maintain high availability and drive reliability as a core company value.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff SRE: Scale, Observability & Automation Leader
Staff SRE: Scale, Observability & Automation Leader

Replit • Northern (KY)

Hybrid
USD 180,000 - 260,000
Salary & equity
401(k) matching
Health, dental, vision, life
+9
Senior SRE: Observability, Automation & Scalable Systems
Senior SRE: Observability, Automation & Scalable Systems

Replit • Northern (KY)

Hybrid
USD 140,000 - 190,000
Competitive Salary & Equity
401(k) 4% match (US)
Health, Dental, Vision & Life
+7
Staff SRE: Scale, Resilience & Observability
Staff SRE: Scale, Resilience & Observability

Early Warning Services LLC • Scottsdale (AZ)

On-site
USD 120,000 - 160,000
Healthcare Coverage
401(k) Retirement Plan
Flexible Time Off
+2
Staff SRE: Scale, Resilience & Observability
Staff SRE: Scale, Resilience & Observability

Early Warning Services LLC • San Francisco (CA)

On-site
USD 130,000 - 160,000
Healthcare Coverage
401(k) Retirement Plan
Paid Time Off
+2
Senior Site Reliability Engineer – Scale & Observability
Senior Site Reliability Engineer – Scale & Observability

Inspire Brands, Inc. • Atlanta (GA)

On-site
USD 120,000 - 180,000
Senior SRE: Scale Reliability, Observability & Resilience
Senior SRE: Scale Reliability, Observability & Resilience

Early Warning Services LLC • Scottsdale (AZ)

Hybrid
USD 106,000 - 130,000
Healthcare Coverage
401(k) Retirement Plan
Paid Time Off
+2
Staff SRE: Scale, Resilience & Observability
Staff SRE: Scale, Resilience & Observability

Early Warning Services LLC • Chicago (IL)

On-site
USD 120,000 - 150,000
Healthcare Coverage
401(k) Retirement Plan
Paid Time Off
+1
Staff SRE — Cloud Reliability & Kubernetes Leader
Staff SRE — Cloud Reliability & Kubernetes Leader

Socket.dev • San Mateo (CA)

On-site
USD 240,000 - 300,000
Stock options
Health insurance
401K savings plan
Remote SRE — Scale, Resilience & Observability
Remote SRE — Scale, Resilience & Observability

Bright Vision Technologies • United States

On-site
USD 100,000 - 150,000
Competitive base salary
Health benefits
Long-term stability
Senior SRE: Scale Reliability, Observability & CI/CD
Senior SRE: Scale Reliability, Observability & CI/CD

Breakout Tools • San Francisco (CA)

On-site
USD 120,000 - 160,000