Engineering Manager, Site Reliability Engineering

Replit

Foster City (CA)

On-site

USD 190,000 - 230,000

Full time

6 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Competitive salary
Equity
401(k) match
Health insurance
Parental leave
Flexible time off
Wellness stipend

Job summary

Replit is seeking an Engineering Manager to lead SRE across observability, incident management, load testing, performance engineering, cloud cost and capacity, and rollout infrastructure.

You will grow a distributed team that builds and operates production platforms, collaborating across application and infrastructure boundaries to drive reliable, scalable software at scale.

Qualifications

  • You have led engineers, set priorities, and delivered through a team.
  • You have built and operated distributed systems with deployment awareness.
  • You can lead migrations and incidents while measuring reliability and performance.
  • You build platform capabilities adopted by other teams and balance tradeoffs.

Responsibilities

  • Own observability strategy: metrics, logs, traces, and alerting with meaningful SLOs.
  • Lead incident tooling, cross-team response, and engineering improvements.
  • Build and maintain load/failure testing capabilities for critical paths.
  • Drive end-to-end performance improvements with service owners.
  • Review designs and production changes; debug failures; use AI tooling for improvements.
  • Coach engineers, grow leaders, manage performance, and hire to needs.
  • Measure rollout safety, MTTR, latency, and capacity improvements.

Skills

Engineering management
Distributed systems
Kubernetes
Observability
Performance engineering
Incident management
Cloud cost
Rollout infra
OpenTelemetry (observability)

Tools

OpenTelemetry
GitOps
AI tooling

Job description

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation.

About the Role

Replit enables people to build software with AI. The systems underneath that experience must support safe production changes, measurable reliability, and predictable performance as usage grows.

This Engineering Manager will lead SRE across observability, incident management, load testing, performance engineering, cloud cost and capacity, and rollout infrastructure. You'll lead and grow an existing team that builds and operates production platforms and works hands-on across application and infrastructure boundaries.

This is a software-building leadership role, not simply an incident-management function. You'll help teams ship safely, understand production behavior, and remove performance bottlenecks through concrete engineering improvements. You should be comfortable going deep on a rollout failure or performance investigation while developing technical leaders and sustainable ownership across a distributed team.

What You'll Do
  • Observability. Build and operate metrics, logs, traces, and alerting capabilities. Help teams establish meaningful SLOs and use production telemetry to diagnose problems and verify improvements.

  • Incident Management. Own incident tooling and practices, coordinate cross-team response, and turn incident reviews into engineering improvements that reduce recovery time and repeat failures.

  • Load Testing. Build and maintain load/failure testing capabilities. Validate critical paths under expected demand, quantify headroom, and test recovery and production readiness with service owners.

  • Performance Engineering. Lead deep engagements with internal teams on SLOs and end-to-end performance. Use profiling, telemetry, and load tests to identify bottlenecks and deliver improvements with service owners—not just recommendations.

  • Stay technically engaged. Review designs and production changes, debug difficult failure modes, and use AI coding tools—including Replit—to prototype and automate. Apply rigorous review and verification to AI-generated changes.

  • Build and grow a high-ownership engineering team. Coach engineers, develop technical leaders, manage performance, and hire against agreed needs. Make distributed collaboration, mentoring, and backup coverage deliberate rather than relying on a few permanent escalation points.

  • Measure outcomes and close the loop. Track rollout safety, recovery time, repeat incidents, critical-path latency/throughput, test coverage, and improvements arising from cost/capacity analysis. Agree success measures and continuing ownership with partner teams.

What You'll Bring
  • Demonstrated engineering management. You have led and developed engineers, made prioritization and performance decisions, hired thoughtfully, and delivered through a team—not only acted as its strongest individual contributor.

  • Software-oriented production systems depth. You have built and operated distributed systems or reliability platforms and can reason across deployment behavior, Kubernetes, telemetry, service dependencies, and recovery mechanisms.

  • Safe-change and performance judgment. You have led consequential migrations or incidents and used measurement to diagnose reliability or performance problems. You can distinguish symptoms from causes and validate fixes under realistic conditions.

  • Platform-product and cross-team judgment. You can build capabilities other teams adopt, lead hands-on engagements without absorbing every service's operations, and make clear tradeoffs among reliability, performance, engineering effort, and cost.

Nice to Have
  • Experience with GitOps or progressive-delivery platforms such as Harness, ArgoCD, or Kargo.

  • Experience with observability, profiling, load-testing, and failure-testing systems, including OpenTelemetry or comparable tooling.

  • Experience with cloud cost attribution, capacity planning, and provider coordination, particularly on GCP.

  • Experience growing distributed teams and using AI tools to increase engineering output while preserving production safeguards.

Full-Time Employee Benefits Include:

Competitive Salary & Equity

401(k) Program with a 4% match (US Only)

Health, Dental, Vision and Life Insurance

Short Term and Long Term Disability

Paid Parental, Medical, Caregiver Leave

Flexible Time Off (FTO) + Holidays

Commuter Benefits (In-Office & US Only)

Monthly Wellness Stipend

Autonomous Work Environment

In Office Set-Up Reimbursement (In-Office Only)

Quarterly Team Gatherings

In Office Amenities (In-Office Only)

Want to learn more about what we are up to?

  • Self-driving Company
  • Replit Agent at Scale
  • AI Adoption
  • Build Open-Source Apps

Interviewing + Culture at Replit

  • Operating Principles
  • Reasons not to work at Replit

To achieve our mission of making programming more accessible around the world, we need our team to be representative of the world. We welcome your unique perspective and experiences in shaping this product. We encourage people from all kinds of backgrounds to apply, including and especially candidates from underrepresented and non-traditional backgrounds.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Replit • United States

Remote
USD 120,000 - 210,000
Competitive Salary
Equity
401(k) Match
+16
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Replit • Northern (KY)

Hybrid
USD 140,000 - 190,000
Competitive Salary & Equity
401(k) 4% match (US)
Health, Dental, Vision & Life
+7
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Replit • Foster City (CA)

On-site
USD 180,000 - 260,000
Competitive Salary & Equity
401(k) with 4% match
Health, Dental, Vision and Life Ins.
+2
Engineering Manager, Cloud Infrastructure
Engineering Manager, Cloud Infrastructure

Replit • Foster City (CA)

On-site
USD 180,000 - 280,000
Competitive salary
Equity
401(k) with match
+8
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Replit • Northern (KY)

Hybrid
USD 180,000 - 260,000
Salary & equity
401(k) matching
Health, dental, vision, life
+9
Staff Infrastructure Engineer
Staff Infrastructure Engineer

Replit, Inc. • Foster City (CA)

On-site
USD 140,000 - 180,000
Competitive Salary & Equity
401(k) Program
Health, Dental, Vision and Life Insurance
+2
Field Engineer
Field Engineer

Replit • Foster City (CA)

Hybrid
USD 140,000 - 180,000
Competitive Salary & Equity
4% 401(k) match
Health, Dental, Vision
+2
Software Engineer, Replit Cloud
Software Engineer, Replit Cloud

Georgian • Foster City (CA)

Hybrid
USD 150,000 - 210,000
Competitive salary
Equity
Health, Dental, Vision
+2
Senior Software Engineer, Compute Platform
Senior Software Engineer, Compute Platform

Replit • Foster City (CA)

Hybrid
USD 140,000 - 210,000
Salary & Equity
401k match (US)
Health/Dental/Vision + Life Insurance
+9
Senior Software Engineer, Enterprise Platform
Senior Software Engineer, Enterprise Platform

Replit, Inc. • Foster City (CA)

On-site
USD 120,000 - 160,000
Competitive Salary & Equity
401(k) Program
Health, Dental, Vision and Life Insurance
+9