Senior Manager SRE

Expedite Talent Solutions

United States

Hybrid

USD 130,000 - 160,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Expedite Talent Solutions is looking for an experienced professional to drive the SRE Activation & Operating Model. This role focuses on ensuring that SRE practices are integrated throughout the development lifecycle and that SLOs, error budgets, and incident management are standardized and effective.

The ideal candidate has significant experience in engineering and leading SRE teams, as well as proven expertise with observability tools and cloud environments. A strong focus on strategic thinking and communication is also essential for success in this position.

Qualifications

  • 10+ years in engineering, operations, or SRE roles.
  • 5+ years leading SRE, platform, or reliability-focused teams.
  • Proven experience implementing SRE practices at scale.

Responsibilities

  • Drive adoption of the SRE operating model across application teams.
  • Define and enforce SLIs, SLOs, and Error Budgets.
  • Ensure SRE practices are embedded into the development lifecycle.

Skills

SRE practices
Cloud environments (AWS, Azure, GCP)
Incident management
Observability tools (Splunk, AppDynamics, Prometheus)
Strategic thinking
Communication and influence skills

Job description

SRE Activation & Operating Model
Key Responsibilities
  • Drive adoption of the SRE operating model across application teams.
  • Establish clarity in roles between:
    • SRE
    • Production Support Engineering (PSE)
    • Application teams
  • Ensure SRE practices are embedded into the development lifecycle, not treated as post-production activities.
  • Define and enforce:
    • SLIs, SLOs, and Error Budgets
    • Production readiness criteria
    • Reliability best practices
  • Lead SLO adoption and compliance reviews across the organization.
  • Establish governance frameworks to ensure consistent application of standards.
  • Partner with:
    • Application product teams
    • Production Support Engineering (MG team)
    • Platform / Infrastructure / Observability teams
  • Drive alignment and reduce friction between engineering and operations.
  • Ensure clear handoffs, escalation models, and operational ownership.
  • Lead adoption of centralized observability standards across:
    • Metrics
    • Logging
    • Tracing
  • Align tooling (AppDynamics, Splunk, Prometheus, etc.).
  • Ensure monitoring and alerting are SLO-driven and actionable, not noise-based.
  • Partner with PSE to strengthen:
    • Incident management processes
    • RCA (Root Cause Analysis) standards
  • Drive identification of patterns and systemic issues.
  • Ensure learnings translate into engineering improvements and automation.
  • Identify opportunities to:
    • Reduce manual operational work
    • Improve system resilience
    • Enable self-healing capabilities
  • Promote a culture of engineering over reaction.
  • Define and track reliability metrics across FS&I.
  • Build reporting that provides visibility into:
    • System health
    • Incident trends
    • SLO performance
  • Translate technical data into actionable business insights.
Required Qualifications
  • 10+ years in engineering, operations, or SRE roles.
  • 5+ years leading SRE, platform, or reliability-focused teams.
  • Proven experience implementing SRE practices at scale (SLIs, SLOs, error budgets).
  • Strong background in cloud environments (AWS, Azure, GCP).
  • Hands‑on experience with observability tools (Splunk, AppDynamics, Prometheus, etc.).
  • Experience in incident management and production operations at scale.
  • Ability to operate effectively in high‑pressure and complex enterprise environments.
Preferred Qualifications
  • Experience driving organizational transformation (not just technical implementation).
  • Strong understanding of CI/CD, DevOps, and automation practices.
  • Experience working in regulated or large enterprise environments.
  • Familiarity with AIOps or advanced automation strategies.
Key Success Indicators
  • Increased adoption of SLOs and reliability standards.
  • Reduction in high‑severity incidents over time.
  • Improved MTTR and operational efficiency.
  • Increased adoption of standardized observability practices.
  • Reduction in reactive, ticket‑driven work across teams.
  • Clear alignment between SRE, PSE, and application teams.
Core Competencies
  • Strategic thinking with strong execution focus.
  • Ability to drive alignment across multiple teams and stakeholders.
  • Strong communication and influence skills.
  • Bias toward structure, clarity, and accountability.
  • Ability to operate with urgency and discipline in complex environments.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Jobtailor • New Jersey

On-site
USD 120,000 - 180,000
SRE Support Engineer - Observability
SRE Support Engineer - Observability

Gigster • Austin (TX)

Remote
USD 80,000 - 100,000
High autonomy in a remote-first environment
Real technical problem solving
Opportunity for scaling support
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

Hybrid
USD 130,000 - 180,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

Harvey Nash • Charlotte (NC)

On-site
USD 100,000 - 130,000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Practice by Numbers • United States

On-site
USD 120,000 - 160,000
High ownership and autonomy
Strong engineering culture
Impactful work on healthcare infrastructure
Staff Site Reliability Engineer, SRE
Staff Site Reliability Engineer, SRE

Jobtailor • California (MO)

On-site
USD 120,000 - 210,000
Senior Engineer - Site Reliability Engineering
Senior Engineer - Site Reliability Engineering

LSEG (London Stock Exchange Group) • Allen (TX)

On-site
USD 140,000 - 190,000
Lead, Site Reliability Engineer
Lead, Site Reliability Engineer

CardWorks • Pittsburgh

Hybrid
USD 146,000 - 163,000
Competitive Pay
Medical, Dental, and Vision Benefits
401(k) Plan with Company Match
+1
Lead SRE
Lead SRE

JPMorgan Chase & Co. • Plano (TX)

On-site
USD 150,000 - 190,000
Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

JPMorgan Chase & Co. • Jersey City (NJ)

On-site
USD 150,000 - 210,000