Senior SRE: Scale systems, automate with IaC

Replit

United States

Remote

USD 120,000 - 210,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Competitive Salary
Equity
401(k) Match
Health Insurance
Dental Insurance
Vision Insurance
Life Insurance
Short-term Disability
Long-term Disability
Parental Leave
Caregiver Leave
Flexible Time Off
Holidays
Commuter Benefits
Wellness Stipend
Autonomous Work Environment
In-Office Setup Reimbursement
Quarterly Team Gatherings
In-Office Amenities

Job summary

Replit is seeking a Site Reliability Engineer to join our SRE team and help ensure the reliability, scalability, and performance of our global infrastructure. You will design observability, automate infrastructure, and drive reliability improvements across services used by millions of developers.

You will build monitoring, define SLOs/SLIs, implement IaC with Terraform/Ansible/Pulumi, and lead incident response to reduce MTTR. Strong programming in Python/Go and Kubernetes experience required.

Qualifications

  • 4-8 years of experience in Site Reliability Engineering or similar roles (DevOps, Systems Engineering, Infrastructure Engineering)
  • Strong programming skills in languages commonly used for automation (Python, Go, or similar)
  • Deep understanding of distributed systems
  • Experience with container orchestration platforms (Kubernetes) and cloud-native technologies
  • Proven track record of implementing and maintaining monitoring/observability solutions
  • Strong incident management skills with experience leading incident response
  • Experience with infrastructure as code and configuration management tools

Responsibilities

  • Design and Implement Observability Solutions: Develop monitoring and alerting, dashboards, metrics, and logging strategies.
  • Drive Automation and Infrastructure as Code: Architect automation with Terraform/Ansible/Pulumi; maintain CI/CD pipelines; build self-healing systems.
  • Establish SLOs and SLIs: Define and track metrics with product/engineering teams.
  • Incident Management and Response: Lead incident response, post-mortems, and runbooks to reduce MTTR.
  • Performance Optimization: Identify bottlenecks and optimize resource usage across regions.

Skills

Python
Go
Distributed systems
Observability
Incident response
Automation
Kubernetes
IaC

Tools

Kubernetes
Terraform
Ansible
Pulumi

Job description

Replit is seeking a Site Reliability Engineer to join our SRE team and help ensure the reliability, scalability, and performance of our global infrastructure. You will design observability, automate infrastructure, and drive reliability improvements across services used by millions of developers.

You will build monitoring, define SLOs/SLIs, implement IaC with Terraform/Ansible/Pulumi, and lead incident response to reduce MTTR. Strong programming in Python/Go and Kubernetes experience required.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE: Observability, Automation & Scalable Systems
Senior SRE: Observability, Automation & Scalable Systems

Replit • Northern (KY)

Hybrid
USD 140,000 - 190,000
Competitive Salary & Equity
401(k) 4% match (US)
Health, Dental, Vision & Life
+7
Staff SRE: Scale, Observability & Automation Leader
Staff SRE: Scale, Observability & Automation Leader

Replit • Northern (KY)

Hybrid
USD 180,000 - 260,000
Salary & equity
401(k) matching
Health, dental, vision, life
+9
Senior SRE: Scale Resilient AI Platforms & Automation
Senior SRE: Scale Resilient AI Platforms & Automation

Relx Plc • Philadelphia

Hybrid
USD 95,000 - 159,000
Senior SRE: Scale Reliability, Observability & CI/CD
Senior SRE: Scale Reliability, Observability & CI/CD

Breakout Tools • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior Site Reliability Engineer: Scale, Automate, Observe
Senior Site Reliability Engineer: Scale, Automate, Observe

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Senior SRE: Scale, Reliability & Observability Lead
Senior SRE: Scale, Reliability & Observability Lead

Alien Blue • Chicago (IL)

On-site
USD 190,800 - 267,100
Comprehensive Healthcare Benefits
401k Matching
Flexible Vacation
+2
Senior SRE: Scale, Reliability & Observability Leader
Senior SRE: Scale, Reliability & Observability Leader

Brez Technology Inc. • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Private Medical, Dental and Vision Benefits
Retirement Savings plan with matching contributions
Workspace benefits for your home office
+4
Senior SRE: Scale Infra, Automate, Elevate Reliability
Senior SRE: Scale Infra, Automate, Elevate Reliability

Fathom.ai • United States

On-site
USD 100,000 - 130,000
Competitive compensation
Supportive environment for personal growth
Dynamic and collaborative team
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Kovoro • Denver (CO), Northern (KY)

On-site
USD 150,000 - 190,000
Senior SRE II — Scale Systems with AI-Driven Reliability
Senior SRE II — Scale Systems with AI-Driven Reliability

Juniper Square • United States

On-site
USD 165,000 - 195,000
Health, dental, and vision care
Life insurance
Mental wellness coverage
+3