Senior SRE Engineering Manager – Reliability at Scale

Upstart

United States

On-site

USD 180,000 - 260,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Annual equity grants
401(k) retirement match
ESPP (US only)
Health coverage (medical, dental, and
Wellness resources
Paid time off
Parental leave

Job summary

Upstart is seeking a Senior Engineering Manager for Site Reliability Engineering to lead reliability improvements across incident management, observability, and operational readiness. You will guide a hands-on team, translate strategy into actionable plans, and balance immediate needs with durable, scalable improvements in a fast‑growing, digital‑first environment.

You will partner with engineering leaders to set reliability expectations, identify systemic risks, and build scalable capabilities

Qualifications

  • 5+ years of reliability engineering management experience and 7+ years of experience in software engineering, site reliability engineering, infrastructure, or platform engineering
  • Significant hands-on experience in Site Reliability Engineering, Production Engineering, or an equivalent role responsible for operating and improving production systems
  • Direct experience managing an SRE, Production Engineering, or equivalent reliability function, including ownership of its strategy, roadmap, operating model, and outcomes
  • Strong technical depth in distributed systems, cloud infrastructure, observability, and production operations
  • Experience leading high severity incident response and improving incident management practices at scale
  • Demonstrated ability to translate strategy into focused, capacity aware plans and deliver measurable outcomes
  • Track record of hiring, developing, and retaining high performing engineers and engineering leaders
  • Strong cross-functional leadership and communication, with the ability to turn complex operational data into clear decisions and drive alignment across teams

Responsibilities

  • Manage incident management, observability, operational readiness, and reliability engineering initiatives
  • Define charter, priorities, roadmap, and measurable outcomes for the SRE function
  • Translate strategy into capacity-aware plans with clear ownership and milestones
  • Maintain delivery health visibility and intervene when needed
  • Build a resilient operating model through cross-training and strong ownership
  • Set a high bar for technical quality and executive communication
  • Develop engineers and leaders who can independently own complex reliability initiatives
  • Evolve incident management programs to improve detection, response, and learning
  • Improve postmortem quality and drive durable engineering improvements
  • Define service level objectives and align reliability efforts with product goals
  • Establish scalability standards for new services and major launches
  • Drive automation and resilience to reduce toil and risk
  • Embed reliability into standard engineering workflows

Skills

Distributed systems
SRE management
Incident response
Observability
Engineering leadership

Tools

Datadog
Grafana
Prometheus
OpenTelemetry
Kubernetes
AWS

Job description

Upstart is seeking a Senior Engineering Manager for Site Reliability Engineering to lead reliability improvements across incident management, observability, and operational readiness. You will guide a hands-on team, translate strategy into actionable plans, and balance immediate needs with durable, scalable improvements in a fast‑growing, digital‑first environment.

You will partner with engineering leaders to set reliability expectations, identify systemic risks, and build scalable capabilities

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: Scale Reliability, Observability & Resilience
Senior SRE: Scale Reliability, Observability & Resilience

Early Warning Services LLC • Scottsdale (AZ)

Hybrid
USD 106,000 - 130,000
Healthcare Coverage
401(k) Retirement Plan
Paid Time Off
+2
Senior SRE: Reliability & Observability Lead
Senior SRE: Reliability & Observability Lead

Inspire • Atlanta (GA)

On-site
USD 140,000 - 200,000
Remote SRE Manager: Lead Reliability & Automation
Remote SRE Manager: Lead Reliability & Automation

NationsBenefits, LLC • Plantation (FL)

Remote
USD 140,000 - 180,000
Fully remote
Unlimited PTO
Competitive compensation
+1
Senior Site Reliability Lead: Scale & Reliability Champion
Senior Site Reliability Lead: Scale & Reliability Champion

Twitter • San Francisco (CA)

On-site
USD 130,000 - 160,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 240,000
Senior SRE & Performance Engineering Lead
Senior SRE & Performance Engineering Lead

Omnissa, LLC in • Mountain View (CA)

Hybrid
USD 223,000 - 310,000
Employee ownership
Health insurance
401k with matching contributions
+3
Senior SRE: Automate Reliability & Observability
Senior SRE: Automate Reliability & Observability

United States Digital Space LLC • Charlotte (TX)

On-site
USD 153,000 - 192,000
Discretionary incentive eligible
Benefits package
SRE Manager: Lead Reliability & Observability at Scale
SRE Manager: Lead Reliability & Observability at Scale

Iac/interactivecorp • Sacramento (CA)

On-site
USD 150,000 - 210,000
Collaborative work environment
Commitment to carbon‑reduction mission
Flex schedule
+2
Senior SRE: Scalable Infra, Observability & Automation
Senior SRE: Scalable Infra, Observability & Automation

Early Warning • Scottsdale (AZ)

Hybrid
USD 106,000 - 156,000
Healthcare Coverage
401(k) Company Match
Paid Time Off
+2
Senior SRE - Hybrid, Observability & Reliability
Senior SRE - Hybrid, Observability & Reliability

Early Warning Services LLC • Chicago (IL)

Hybrid
USD 106,000 - 130,000
Healthcare Coverage
401(k) Plan with match
PTO and Holidays
+1