Sr Engineer, Site Reliability Engineer

Devopsroles

St. Louis (MO)

Hybrid

USD 112,000 - 134,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Devopsroles in Saint Louis, MO is seeking a Sr Engineer, Site Reliability Engineer to lead reliability practices for critical platforms. You will partner with Engineering, Security, and Product teams to improve reliability, scalability, and performance through automation and observability.

The role focuses on defining SLOs/SLIs, driving resiliency, and maintaining IaC and CI/CD pipelines. You will mentor engineers and influence architecture decisions across teams.

Qualifications

  • 10+ years in Site Reliability/DevOps or related fields.
  • Deep experience with cloud-native distributed systems (public cloud).
  • Strong SRE concepts: SLOs/SLIs, error budgets, availability modeling.
  • Hands-on IaC (Terraform/CloudFormation).
  • Kubernetes and container platforms expertise.
  • Observability proficiency: metrics, logs, traces, APM.
  • Software scripting in Python/Go/Java/Bash.
  • CI/CD pipelines and modern delivery practices.
  • Networking, OS, security basics, databases, app architectures.
  • Excellent communication and collaboration across teams.

Responsibilities

  • Design and evolve SRE practices for mission-critical apps and distributed systems.
  • Define and govern SLOs/SLIs and error budgets to improve reliability.
  • Lead automation, self-healing, and toil reduction initiatives.
  • Embed reliability requirements into architecture and software lifecycles.
  • Develop and maintain IaC, platform automation, CI/CD pipelines, and tooling.
  • Lead reliability reviews, architecture assessments, and resiliency tests.
  • Develop observability solutions with metrics, logs, traces, and synthetic monitoring.
  • Analyze production data to identify risks and drive improvements.
  • Collaborate with development to optimize performance and cloud usage.
  • Provide leadership during incidents with RCA and improvements.
  • Promote high availability, disaster recovery, and business continuity.

Skills

SRE
Cloud
Kubernetes
IaC
Observability
Automation
Python/Go

Job description

Sr Engineer, Site Reliability Engineer
Job Purpose

In this role, you will serve as a senior technical leader responsible for designing, building, and evolving reliability engineering practices across critical platforms and services. You will partner closely with Engineering, Architecture, Security, and Product teams to improve system reliability, scalability, resiliency, and performance through automation, observability, and engineering excellence.

You will establish reliability standards, define SLOs and SLIs, drive adoption of reliability best practices, and leverage data-driven insights to continuously improve customer experience and platform health. As a technical expert, you will influence architecture decisions, lead reliability initiatives, and help engineering teams build highly available, fault-tolerant systems at scale.

Location: St. Louis, MO (hybrid)

Responsibilities
  • Design, implement, and evolve Site Reliability Engineering practices for mission-critical applications and distributed systems.
  • Define, implement, and govern Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets to measure and improve service reliability.
  • Drive reliability improvements through automation, self-healing capabilities, resiliency engineering, and reduction of operational toil.
  • Partner with engineering teams to embed reliability, scalability, security, and observability requirements into system architecture and software development lifecycles.
  • Develop and maintain Infrastructure as Code (IaC), platform automation, CI/CD pipelines, and reliability tooling to improve deployment consistency and operational efficiency.
  • Lead reliability reviews, architecture assessments, and resiliency testing efforts, including failure mode analysis, fault injection, and chaos engineering practices.
  • Design and implement comprehensive observability solutions utilizing metrics, logs, traces, synthetic monitoring, and real user monitoring.
  • Analyze production performance, reliability trends, and failure patterns to proactively identify systemic risks and drive long-term improvements.
  • Collaborate with development teams to optimize application performance, scalability, resource utilization, and cloud infrastructure efficiency.
  • Provide technical leadership during complex production incidents by supporting root cause analysis and identifying opportunities for reliability improvements.
  • Drive adoption of engineering best practices for high availability, disaster recovery, fault tolerance, and business continuity.
  • Establish reliability standards, reference architectures, engineering patterns, and platform capabilities to enable scalable and resilient service delivery.
  • Mentor engineers and act as a subject matter expert in SRE, cloud architecture, observability, automation, and distributed systems engineering.
Requirements
  • 10+ years of experience in Site Reliability Engineering, Software Engineering, Platform Engineering, Infrastructure Engineering, or DevOps, supporting large-scale, highly available systems.
  • Deep experience designing, operating, and optimizing cloud-native and distributed systems in public cloud environments (Preferably GCP).
  • Strong expertise in reliability engineering concepts, including SLOs, SLIs, error budgets, availability modeling, scalability, resiliency, and performance engineering.
  • Hands-on experience building infrastructure and platform automation using Infrastructure as Code tools such as Terraform or CloudFormation.
  • Strong experience with Kubernetes, container platforms, orchestration technologies, and modern cloud-native architectures.
  • Experience implementing observability strategies utilizing monitoring, logging, tracing, and application performance management platforms.
  • Strong software engineering and scripting skills using languages such as Python, Go, Java, Bash, or similar.
  • Experience building and supporting CI/CD pipelines and modern software delivery practices.
  • Strong understanding of networking, operating systems, security principles, databases, and application architectures.
  • Demonstrated ability to perform deep technical troubleshooting and root cause analysis in complex distributed environments.
  • Excellent communication and collaboration skills, with proven ability to influence technical decisions across multiple engineering teams.
Preferred Qualifications
  • Experience designing and operating large-scale distributed systems supporting millions of transactions or users.
  • Strong background in observability and APM platforms such as Dynatrace, Datadog, New Relic, Splunk, Grafana, OpenTelemetry, AppDynamics, or equivalent technologies.
  • Experience implementing advanced reliability practices such as chaos engineering, fault injection testing, resiliency validation, and performance benchmarking.
  • Deep expertise in cloud platform architecture, including compute, networking, storage, service mesh, and container ecosystem technologies.
  • Experience building internal developer platforms, self-service infrastructure capabilities, and platform engineering solutions.
  • Knowledge of disaster recovery, business continuity, and multi-region architecture design, including RTO/RPO strategies.
  • Experience operating in regulated environments with security and compliance requirements such as PCI, SOX, SOC2, or ISO 27001.
  • Demonstrated technical leadership and ability to influence engineering strategy, architecture decisions, and organizational adoption of SRE principles.
  • Experience leveraging OpenTelemetry and modern observability frameworks to correlate business, application, and infrastructure telemetry.
  • Passion for automation, engineering excellence, continuous improvement, and building resilient systems at enterprise scale.

Equal Opportunity Employer: Disabled/VeteransCompetitive Pay $111,808 - $134,169 annually

The actual pay offered will be determined by multiple factors, including but not limited to the candidate’s relevant experience, job-related knowledge, skills, and geographical location. Individual compensation decisions are dependent upon the facts and circumstances of each position and candidate.

Saint Louis Support Center

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Bank of America • Charlotte (TX)

On-site
USD 153,000 - 192,000
Discretionary incentive eligible
Benefits package
Senior Site Reliability Engineer
Senior Site Reliability Engineer

HITEC • Jersey City (NJ)

On-site
USD 153,000 - 192,000
Benefits eligible
Discretionary incentive plan
Senior Site Reliability Engineer
Senior Site Reliability Engineer

NBMBAA • Jersey City (NJ)

On-site
USD 153,000 - 192,000
Benefits eligible
Annual discretionary plan
Lead, Site Reliability Engineer
Lead, Site Reliability Engineer

CardWorks • Pittsburgh

On-site
USD 146,032 - 162,257
Competitive Pay
Medical, Dental, and Vision Benefits
401(k) Plan with Company Match
+1
Senior Site Reliability Engineer II
Senior Site Reliability Engineer II

LexisNexis Risk Solutions • San Jose (CA), Northern (KY)

Hybrid
USD 105,000 - 175,000
401(k) with match
Wellbeing programs
Life Insurance
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mission Staffing • New York (NY)

On-site
USD 140,000 - 200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Storm2 • Scottsdale (AZ)

On-site
USD 140,000 - 150,000
Competitive healthcare, dental, and vision coverage
401(k) with company match
Generous PTO and paid holidays
+1
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Brooksource • San Antonio (TX)

On-site
USD 80,000 - 120,000
Technical Operations Lead
Technical Operations Lead

First Citizens Bank • Phoenix (AZ)

On-site
USD 140,000 - 190,000
Benefits program