Staff SRE: Scale, Observability & Automation Leader

Replit

Northern (KY)

Hybrid

USD 180,000 - 260,000

Full time

2 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Salary & equity
401(k) matching
Health, dental, vision, life
Disability insurance
Parental, medical, caregiver leave
Flexible time off
Commuter benefits
Wellness stipend
Autonomous work environment
Office setup reimbursement
Quarterly team gatherings
Office amenities

Job summary

Replit, a leading software creation platform, is seeking a Staff Site Reliability Engineer to ensure reliability, scale, and performance of our global infrastructure. You will bridge development and operations, architect observability, and drive automation across Kubernetes, cloud-native stacks, and CI/CD pipelines.

You will lead incident management, mentor engineers, and design robust, self-healing systems while maintaining high availability for millions of developers worldwide.

Qualifications

  • 8-10 years of SRE/DevOps related experience.
  • Strong programming in Python or Go.
  • Deep understanding of distributed systems and SOA concepts.
  • Experience with Kubernetes and cloud-native tech.
  • Proven monitoring/observability design and implementation.
  • IaC experience (Terraform, Pulumi).
  • Excellent communication and mentorship skills.

Responsibilities

  • Architect and implement comprehensive observability and dashboards.
  • Define and track SLOs/SLIs with cross-functional teams.
  • Lead incident management and post-mortems.
  • Automate infrastructure and CI/CD pipelines.
  • Optimize performance of large-scale Kubernetes deployments.
  • Debug and harden distributed systems across stack.
  • Provide staff-level design reviews and mentorship.

Skills

Python/Go
Distributed systems
Kubernetes
Observability
Incident management
IaC Terraform/Pulumi
Communication
Mentoring
Debugging

Tools

Kubernetes
Terraform
Pulumi
Prometheus/Grafana/OpenTelemetry
Datadog

Job description

Replit, a leading software creation platform, is seeking a Staff Site Reliability Engineer to ensure reliability, scale, and performance of our global infrastructure. You will bridge development and operations, architect observability, and drive automation across Kubernetes, cloud-native stacks, and CI/CD pipelines.

You will lead incident management, mentor engineers, and design robust, self-healing systems while maintaining high availability for millions of developers worldwide.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: Observability, Automation & Scalable Systems
Senior SRE: Observability, Automation & Scalable Systems

Replit • Northern (KY)

Hybrid
USD 140,000 - 190,000
Competitive Salary & Equity
401(k) 4% match (US)
Health, Dental, Vision & Life
+7
Staff SRE — Scale, Reliability & Platform Leadership
Staff SRE — Scale, Reliability & Platform Leadership

Attentive • Wilmington (DE)

On-site
USD 180,000 - 240,000
Health & wellness
Equity
Staff SRE: Scale, Resilience & Observability
Staff SRE: Scale, Resilience & Observability

Early Warning Services LLC • Scottsdale (AZ)

On-site
USD 120,000 - 160,000
Healthcare Coverage
401(k) Retirement Plan
Flexible Time Off
+2
Staff SRE: Scale, Resilience & Observability
Staff SRE: Scale, Resilience & Observability

Early Warning Services LLC • San Francisco (CA)

On-site
USD 130,000 - 160,000
Healthcare Coverage
401(k) Retirement Plan
Paid Time Off
+2
Senior SRE: Scale Reliability, Observability & Resilience
Senior SRE: Scale Reliability, Observability & Resilience

Early Warning Services LLC • Scottsdale (AZ)

Hybrid
USD 106,000 - 130,000
Healthcare Coverage
401(k) Retirement Plan
Paid Time Off
+2
Senior SRE: Scale Reliability & Observability
Senior SRE: Scale Reliability & Observability

Megaport • Abbeyville (CO)

On-site
USD 130,000 - 190,000
Contractor (PJ)
Paid Time Off
Competitive Compensation
+4
Senior Site Reliability Engineer: Scale, Automate, Observe
Senior Site Reliability Engineer: Scale, Automate, Observe

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Replit • Northern (KY)

Hybrid
USD 140,000 - 190,000
Competitive Salary & Equity
401(k) 4% match (US)
Health, Dental, Vision & Life
+7
Senior SRE: Scale Reliability, Observability & CI/CD
Senior SRE: Scale Reliability, Observability & CI/CD

Breakout Tools • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior Site Reliability Engineer – Scale & Observability
Senior Site Reliability Engineer – Scale & Observability

Inspire Brands, Inc. • Atlanta (GA)

On-site
USD 120,000 - 180,000