Senior SRE: Cloud-Native Reliability & Observability

O.C. Tanner

Salt Lake City (UT)

On-site

USD 130,000 - 180,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

O.C. Tanner in Salt Lake City seeks a Senior Site Reliability Engineer to drive reliability for our cloud-native, SaaS platform.

You will build self-healing systems, improve observability, and collaborate with product and engineering teams to move faster without compromising quality. You will own incident response, design scalable infrastructure, lead on-call rotations, and promote best practices in automation, monitoring, and deployment.

Qualifications

  • 5+ years of experience in SRE/DevOps or related roles with strong incident response background.
  • Experience operating large-scale SaaS platforms with high availability requirements.
  • Proficiency in Go, Python, or Java, with automation experience.
  • Hands-on experience with IaC and GitOps tools (Terraform/OpenTofu/CDKTF/Pulumi).
  • Experience with observability platforms (OpenTelemetry, Datadog, Coralogix).
  • Strong knowledge of AWS services and Kubernetes in production.
  • Understanding of monitoring, logging, and distributed tracing for complex systems.
  • Ability to collaborate with software engineering and testing teams to improve reliability.

Responsibilities

  • Improve availability, scalability, and performance of cloud-native apps via automation and best practices.
  • Build and evolve observability platforms and establish SLOs and monitoring standards.
  • Lead production triage, incident response, RCA, and blameless post-incident reviews.
  • Partner with Eng/Support/Product to embed reliability and observability in lifecycle.
  • Promote automation, monitoring, shift-left quality, and shared ownership of services.
  • Collaborate across global teams to provide 24x7 support and smooth handoffs.
  • Participate in on-call rotation to maintain service health and reduce toil.

Skills

SRE/DevOps experience
Production triage
High availability
On-call experience
Collaboration
Monitoring & tracing concepts
Reliability culture

Tools

Terraform
OpenTofu
CDKTF
Pulumi
ArgoCD
Helm
Kubernetes
OpenTelemetry
Datadog
Coralogix
AWS
PostgreSQL
OpenSearch
Redis/ElastiCache
Aurora
Kafka
ActiveMQ
SNS/SQS

Job description

O.C. Tanner in Salt Lake City seeks a Senior Site Reliability Engineer to drive reliability for our cloud-native, SaaS platform.

You will build self-healing systems, improve observability, and collaborate with product and engineering teams to move faster without compromising quality. You will own incident response, design scalable infrastructure, lead on-call rotations, and promote best practices in automation, monitoring, and deployment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE Manager: Lead Reliability & Observability at Scale
SRE Manager: Lead Reliability & Observability at Scale

O.C. Tanner • Salt Lake City (UT)

On-site
USD 150,000 - 190,000
SRE Manager: Drive Reliability, Observability & Scale
SRE Manager: Drive Reliability, Observability & Scale

1 O.C. Tanner Company • Salt Lake City (UT)

On-site
USD 120,000 - 180,000
Head of SRE & Reliability Engineering
Head of SRE & Reliability Engineering

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 260,000
Senior SRE Lead: Cloud Reliability & Automation
Senior SRE Lead: Cloud Reliability & Automation

Oracle • Vienna (VA)

On-site
USD 96,000 - 265,000
Medical, dental, vision insurance
401(k) with company match
Paid time off and holidays
+1
Senior SRE - Cloud & Observability
Senior SRE - Cloud & Observability

Ridgeline • Reno (NV)

Hybrid
USD 153,000 - 210,000
Unlimited vacation
Education reimbursement
Wellness reimbursement
+1
Senior Cloud SRE: Automation, Observability & Reliability
Senior Cloud SRE: Automation, Observability & Reliability

Salve.Lab • Atlanta (GA)

On-site
USD 140,000 - 190,000
Remote-friendly
Professional development
Apple equipment
+1
Senior SRE: Scale Reliability & Observability
Senior SRE: Scale Reliability & Observability

Megaport • Abbeyville (CO)

On-site
USD 130,000 - 190,000
Contractor (PJ)
Paid Time Off
Competitive Compensation
+4
Manager, Site Reliability Engineering
Manager, Site Reliability Engineering

1 O.C. Tanner Company • Salt Lake City (UT)

On-site
USD 120,000 - 180,000
Senior SRE: Cloud Automation & AI-Driven Reliability
Senior SRE: Cloud Automation & AI-Driven Reliability

Ll Oefentherapie • Frankfort (KY)

On-site
USD 120,000 - 160,000
Senior Cloud SRE: Secure, Scalable Infra & Automation
Senior Cloud SRE: Secure, Scalable Infra & Automation

Okta • Bellevue (WA)

On-site
USD 147,000 - 202,000
Amazing Benefits
Making Social Impact
Fostering Diversity, Equity, Inclusion