Senior SRE: Observability, Automation & Scalable Systems

Replit

Northern (KY)

Hybrid

USD 140,000 - 190,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Competitive Salary & Equity
401(k) 4% match (US)
Health, Dental, Vision & Life
Paid Parental, Medical, CaregiverLeave
Flexible Time Off (FTO) + Holidays
Commuter Benefits
Wellness stipend
In Office Set-Up Reimbursement
Quarterly Team Gatherings
In Office Amenities

Job summary

Replit is seeking a Site Reliability Engineer to help ensure reliability, scalability, and performance of its global infrastructure. You will bridge development and operations, building automation, monitoring, and self-healing systems at scale.

The role emphasizes implementing robust observability, incident response leadership, and capacity planning to reduce MTTR while sustaining rapid growth across regions.

Qualifications

  • 4–8 years of experience in Site Reliability Engineering or similar roles (DevOps, Systems Engineering)
  • Strong programming skills in Python, Go, or similar for automation
  • Deep understanding of distributed systems and cloud-native tech
  • Experience with container orchestration (Kubernetes) and observability tools
  • Proven track record in monitoring/observability, incident management, and IaC
  • Experience with infrastructure as code and configuration management tools

Responsibilities

  • Design and implement observability solutions with real-time dashboards and alerts
  • Drive automation and infrastructure as code across deployments and CI/CD pipelines
  • Establish SLOs/SLIs with product and engineering teams and monitor them
  • Lead incident response, post-mortems, and runbook development
  • Identify and resolve performance bottlenecks, optimize resource usage across regions

Skills

Python
Go
Distributed Systems
Observability
Incident Management
Automation

Tools

Kubernetes
Terraform
Ansible
Pulumi

Job description

Replit is seeking a Site Reliability Engineer to help ensure reliability, scalability, and performance of its global infrastructure. You will bridge development and operations, building automation, monitoring, and self-healing systems at scale.

The role emphasizes implementing robust observability, incident response leadership, and capacity planning to reduce MTTR while sustaining rapid growth across regions.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff SRE: Scale, Observability & Automation Leader
Staff SRE: Scale, Observability & Automation Leader

Replit • Northern (KY)

Hybrid
USD 180,000 - 260,000
Salary & equity
401(k) matching
Health, dental, vision, life
+9
Senior SRE: Scalable Infra, Observability & Automation
Senior SRE: Scalable Infra, Observability & Automation

Early Warning • Scottsdale (AZ)

Hybrid
USD 106,000 - 156,000
Healthcare Coverage
401(k) Company Match
Paid Time Off
+2
Senior SRE: Scale Reliability, Observability & Resilience
Senior SRE: Scale Reliability, Observability & Resilience

Early Warning Services LLC • Scottsdale (AZ)

Hybrid
USD 106,000 - 130,000
Healthcare Coverage
401(k) Retirement Plan
Paid Time Off
+2
Senior SRE: Automate Reliability & Observability
Senior SRE: Automate Reliability & Observability

United States Digital Space LLC • Charlotte (TX)

On-site
USD 153,000 - 192,000
Discretionary incentive eligible
Benefits package
Senior SRE - Hybrid, Observability & Reliability
Senior SRE - Hybrid, Observability & Reliability

Early Warning Services LLC • Chicago (IL)

Hybrid
USD 106,000 - 130,000
Healthcare Coverage
401(k) Plan with match
PTO and Holidays
+1
Senior SRE: Reliability & Observability Lead
Senior SRE: Reliability & Observability Lead

Inspire • Atlanta (GA)

On-site
USD 140,000 - 200,000
Senior SRE: Scale Reliability & Observability
Senior SRE: Scale Reliability & Observability

Megaport • Abbeyville (CO)

On-site
USD 130,000 - 190,000
Contractor (PJ)
Paid Time Off
Competitive Compensation
+4
Senior SRE Engineer: Reliability, Observability
Senior SRE Engineer: Reliability, Observability

Bloomerang • United States

On-site
USD 115,000 - 150,000
Health insurance
PTO and holidays
401(k) match
+2
Senior SRE: Observability & Automation Lead
Senior SRE: Observability & Automation Lead

ISO New England Inc. • Holyoke (MA)

Hybrid
USD 134,000 - 170,000
Enhanced 401(k) and financial planning
Tuition reimbursement and professional
Onsite gym and wellness programs
+4
Senior SRE: Build Resilient, Scalable Systems
Senior SRE: Build Resilient, Scalable Systems

Methodic • San Francisco (CA)

On-site
USD 140,000 - 210,000