Senior Site Reliability Engineer (SRE)

PrincePerelson and Associates

Salt Lake City (UT)

Hybrid

USD 120,000 - 190,000

Full time

39 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Hybrid work schedule (4 days in-office
Medical, dental, vision
Retirement plan
Generous paid time off
Parental leave
Wellness benefits

Job summary

PrincePerelson & Associates seeks a senior SRE/Platform Engineer to drive reliability across cloud-native apps in AWS and Kubernetes in a hybrid Salt Lake City environment.

You will define SLOs/SLIs, build observability pipelines, and lead incident response while collaborating with software and platform teams to deliver scalable, resilient infrastructure.

Qualifications

  • 7+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or similar infra-focused roles.
  • Strong experience supporting production workloads in AWS and Kubernetes.
  • Demonstrated success implementing SLOs, SLIs, and reliability engineering best practices.
  • Experience designing observability solutions that provide actionable insights while minimizing operational noise.
  • Strong Infrastructure as Code experience, preferably with Terraform.
  • Experience automating operational workflows and building reusable engineering standards.
  • Familiarity with distributed systems and modern data platforms (PostgreSQL, MongoDB, DynamoDB or similar).
  • Experience with observability platforms such as Prometheus, Grafana, New Relic, CloudWatch, ELK, or similar.

Responsibilities

  • Define and evolve SLOs and SLIs to align platform performance with customer and business expectations.
  • Build and standardize monitoring, alerting, and observability using IaC (Terraform preferred).
  • Develop scalable observability solutions leveraging metrics, logs, traces, and OpenTelemetry.
  • Evaluate and implement modern observability platforms and integrations.
  • Establish reliability standards for Kubernetes-based applications including scaling and deployment safety.
  • Design automation to reduce manual effort and improve incident response.
  • Lead high-severity incident response, drive postmortems, and measurable reliability improvements.
  • Participate in an on-call rotation and reduce alert noise.
  • Partner with software engineers and platform teams to build resilient cloud infrastructure.

Skills

SRE
Platform Eng
DevOps
AWS
Kubernetes
Terraform
Observability
Distributed systems
Incident response
Communication

Tools

Prometheus
Grafana
New Relic
CloudWatch
ELK
OpenTelemetry

Job description

What You'll Do
  • Define and evolve Service Level Objectives (SLOs) and Service Level Indicators (SLIs) that align platform performance with customer and business expectations.
  • Build and standardize monitoring, alerting, and observability practices across engineering teams using Infrastructure as Code (Terraform preferred).
  • Develop scalable observability solutions leveraging metrics, logs, traces, and distributed tracing technologies such as OpenTelemetry.
  • Evaluate and implement modern observability platforms and integrations, leading proof-of-concepts and defining adoption strategies.
  • Establish reliability standards for Kubernetes-based applications, including scaling strategies, deployment safety, resource optimization, dashboards, and alerting.
  • Design automation that reduces manual effort, streamline operations, and improves incident response and recovery.
  • Lead high-severity incident response efforts, facilitate postmortems, and drive long-term reliability improvements through measurable action plans.
  • Participate in an on-call rotation while continually improving monitoring quality and reducing unnecessary alert noise.
  • Partner with software engineers and platform teams to build resilient, scalable cloud infrastructure and operational best practices.
What We're Looking For
  • 7+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or a similar infrastructure-focused engineering role.
  • Strong experience supporting production workloads in AWS and Kubernetes environments.
  • Demonstrated success implementing SLOs, SLIs, and reliability engineering best practices.
  • Experience designing observability solutions that provide actionable insights while minimizing operational noise.
  • Strong Infrastructure as Code experience, preferably with Terraform.
  • Experience automating operational workflows and building reusable engineering standards.
  • Familiarity with distributed systems and modern data platforms including technologies such as PostgreSQL, MongoDB, DynamoDB, or similar.
  • Experience with multiple observability platforms such as Prometheus, Grafana, New Relic, Splunk, CloudWatch, ELK, or comparable technologies.
  • Strong troubleshooting and debugging skills across distributed production environments.
  • Experience leading incident response and driving continuous operational improvements.
  • Excellent communication skills with the ability to influence engineering teams through collaboration, technical leadership, and practical solutions.
Why This Opportunity?
  • Join a high-performing engineering organization where reliability is a strategic priority.
  • Work with modern cloud-native technologies including AWS, Kubernetes, Terraform, and OpenTelemetry.
  • Influence engineering standards and platform architecture across multiple teams.
  • Solve complex technical challenges at scale while helping shape the future of the platform.
  • Collaborative hybrid environment with four days in the office and one remote day each week.
  • Comprehensive benefits package including medical, dental, vision, retirement plan, generous paid time off, parental leave, and additional wellness benefits.

PrincePerelson & Associates is an Equal Opportunity Employer and complies with all provisions of the EEO and ADA laws. We do not discriminate in our employment practices on the basis of race, color, religion, national origin, sex (including sexual orientation and sexual identity), age, genetic information, parental status, military status, disability, or any non-merit-based factors or other federal, state, or locally protected class. All applicants applying for U.S. job openings must be authorized to work in the United States.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mission Staffing • New York (NY)

Hybrid
USD 140,000 - 200,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 240,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

New York Technology Partners • Chicago (IL)

On-site
USD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

Veritas Search Group • Tustin (CA)

On-site
USD 140,000 - 190,000
Lead, Site Reliability Engineer
Lead, Site Reliability Engineer

CardWorks • Pittsburgh

Hybrid
USD 146,000 - 163,000
Competitive Pay
Medical, Dental, and Vision Benefits
401(k) Plan with Company Match
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Storm2 • Scottsdale (AZ)

Hybrid
USD 140,000 - 150,000
Competitive healthcare, dental, and vision coverage
401(k) with company match
Generous PTO and paid holidays
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

SDI International • Chicago (IL)

Hybrid
USD 130,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The ReWork Group • New York (NY)

On-site
USD 120,000 - 160,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

OutSolve • Mission (KS)

Remote
USD 90,000 - 130,000
100% remote work environment
Competitive compensation
Professional development opportunities
+1