Site Reliability Engineering Manager

O.C. Tanner

Salt Lake City (UT)

On-site

USD 180,000 - 240,000

Full time

22 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

O.C. Tanner in Salt Lake City is seeking a Manager of Site Reliability Engineering to lead the strategy, execution, and evolution of reliability for our employee recognition platform.

You will mentor a team of SREs and partner with Engineering, Product, and Support to deliver highly available, scalable services serving millions of users. This role emphasizes automation, observability, and a reliability-first culture through incident management, post-incident reviews, and 24x7 coverage in a

Qualifications

  • 5+ years in Site Reliability Engineering, DevOps, Platform Engineering, or related disciplines.
  • 2+ years in technical leadership or people management.
  • Experience leading production operations, reliability engineering, incident management, and operational excellence.
  • Experience operating large-scale, customer-facing SaaS platforms with high availability.
  • Hands-on with observability platforms and cloud environments.

Responsibilities

  • Lead, mentor, and develop a team of Site Reliability Engineers, fostering reliability and continual improvement.
  • Define and execute the reliability strategy to improve availability, scalability, and resilience.
  • Set priorities, goals, and success metrics aligned with business objectives and platform health.
  • Drive shared ownership of production services across Engineering, Product, and Support.
  • Build observability capabilities and establish standards for metrics, logs, traces, and SLOs.
  • Oversee incident response, root cause analysis, and blameless post-incident reviews.
  • Champion automation and shift-left quality practices to reduce toil.
  • Support follow-the-sun operations with seamless 24x7 coverage and effective handoffs.
  • Own on-call programs, incident management practices, and operational health metrics.
  • Manage team capacity, hiring, performance, budgeting, and workforce planning.
  • Report reliability trends and risks to engineering and executive leadership.

Job description

As the Manager of Site Reliability Engineering, you will lead the strategy, execution, and evolution of reliability for our world-class employee recognition platform. You will build, mentor, and empower a team of Site Reliability Engineers while partnering closely with Engineering, Product, and Support organizations to deliver highly available, scalable, and resilient services that serve millions of users. We are seeking a leader who is passionate about operational excellence, continuous improvement, and fostering a reliability-first culture through automation, observability, and shared ownership. In this role, you will champion the development of self-healing platforms, drive incident and operational maturity, and enable engineering teams to innovate faster while delivering exceptional customer experiences.

Key Responsibilities
:
  • Lead, mentor, and develop a team of Site Reliability Engineers, fostering a culture of reliability, accountability, operational excellence, and continuous improvement.
  • Define and execute the organization's reliability strategy, improving availability, scalability, performance, and resilience through automation and engineering best practices.
  • Establish team priorities, goals, and success metrics aligned with business objectives, customer needs, and platform health.
  • Partner with Engineering, Product, and Support leaders to drive shared ownership of production services and embed reliability, observability, and operational excellence throughout the software development lifecycle.
  • Build and evolve observability capabilities using OpenTelemetry, Datadog, Coralogix, or similar tools, establishing enterprise standards for metrics, logs, traces, alerting, and Service Level Objectives (SLOs).
  • Oversee production triage, incident response, and escalation processes, ensuring timely service restoration, effective root cause analysis, and blameless post-incident reviews.
  • Champion a reliability-first engineering culture focused on automation, proactive risk reduction, operational readiness, shift-left quality practices, and continuous improvement.
  • Collaborate with global engineering teams in a follow-the-sun support model, ensuring seamless 24x7 coverage, effective operational handoffs, and consistent service ownership.
  • Own on-call programs, incident management practices, and operational health metrics, driving improvements in alert quality, operational efficiency, and toil reduction.
  • Manage team capacity, hiring, performance management, career development, budgeting, and workforce planning to ensure effective support of business-critical services.
  • Provide regular reporting to engineering and executive leadership on reliability trends, incidents, risks, performance metrics, and strategic initiatives.
Required Qualifications
:
  • 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or related disciplines, including 2+ years in a technical leadership or people management role.
  • Proven experience leading teams responsible for production operations, reliability engineering, incident management, and operational excellence.
  • Experience designing and implementing SRE practices, reliability programs, or operational maturity initiatives within growing engineering organizations.
  • Experience operating large-scale, customer-facing SaaS platforms with high availability, performance, and scalability requirements.
  • Strong understanding of modern software engineering practices and partnering with development teams to build reliable, resilient systems.
  • Hands‑on experience with observability platforms such as OpenTelemetry, Datadog, Coralogix, or similar technologies.
  • Strong knowledge of AWS and Kubernetes in production environments.
  • Deep understanding of monitoring, logging, distributed tracing, SLIs, SLOs, error budgets, and reliability engineering principles.
  • Demonstrated ability to lead cross‑functional initiatives and influence stakeholders across Engineering, Product, and Support organizations.
  • Experience developing engineering roadmaps, defining team objectives, aligning reliability investments with business priorities, and driving continuous operational improvement through incident learning and post‑incident reviews.
Preferred Qualifications
:
  • Experience leading distributed or globally dispersed engineering teams.
  • Experience with multiple cloud providers or cloud‑agnostic platform architectures.
  • Familiarity with security, compliance, governance, and operational risk management frameworks.
  • Proficiency with modern Infrastructure‑as‑Code and technologies such as Terraform, Golang, Python, Playwright, and Performance Monitoring tools.
  • Experience with relational and distributed data technologies such as PostgreSQL, OpenSearch, Redis/ElastiCache, or Aurora.
  • Experience with messaging and streaming platforms such as Kafka, ActiveMQ, SNS/SQS, or similar event‑driven technologies.
  • Strong understanding of cost optimization, platform sustainability, and engineering efficiency metrics.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Sr. Director, Site Reliability and Platform Engineering
Sr. Director, Site Reliability and Platform Engineering

Optomi • Tacoma (WA)

On-site
USD 150,000 - 200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mission Staffing • New York (NY)

Hybrid
USD 140,000 - 200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

SDI International • Chicago (IL)

Hybrid
USD 130,000 - 180,000
Lead, Site Reliability Engineer
Lead, Site Reliability Engineer

CardWorks • Pittsburgh

Hybrid
USD 146,000 - 163,000
Competitive Pay
Medical, Dental, and Vision Benefits
401(k) Plan with Company Match
+1
Senior Vice President, Site Reliability Engineer
Senior Vice President, Site Reliability Engineer

BNY Mellon • Town of Florida (NY)

On-site
USD 170,000 - 230,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

United States Digital Space LLC • Charlotte (TX)

On-site
USD 153,000 - 192,000
Discretionary incentive eligible
Benefits package
Site Reliability Engineer
Site Reliability Engineer

Request Technology, LLC • Chicago (IL)

Hybrid
USD 150,000 - 155,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Jobgether • United States

Remote
USD 150,000 - 200,000
Competitive salary
Comprehensive healthcare coverage
401(k) plan with company matching
+3
Site Reliability Engineer
Site Reliability Engineer

Stelvio Inc. • Town of Texas (WI)

On-site
USD 125,000 - 145,000