Staff Site Reliability Engineer, SRE

Jobtailor

California (MO)

On-site

USD 120,000 - 210,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor is seeking an experienced SRE leader to co-found the reliability practice and drive measurable outcomes across engineering systems. You will design tooling for SLOs, error budgets, and post-incident learning, partnering with Incident Management to standardize responses.

You will transform manual toil into reusable software and establish robust production reliability architecture across hybrid cloud environments, influencing multiple teams and services.

Qualifications

  • Significant experience as an SRE or Production Engineer for large-scale distributed systems.
  • Hands-on in defining SLIs/SLOs and operationalizing error budgets.

Responsibilities

  • Co-found the SRE practice by translating reliability charter into engineering systems and measurable outcomes.
  • Design and build shared reliability tooling and automation (SLOs, readiness evidence, toil reduction).
  • Establish risk-based production reliability architecture and minimum standards.
  • Partner with Incident Management to operationalize severity, command, and response standards.
  • Turn repetitive operational work into reusable software to protect engineering capacity.
  • Help teams classify services, map dependencies, and establish meaningful SLIs/SLOs.

Skills

SRE / Production Engineering
Go / Python / Java / C++
Linux / Kubernetes / Cloud
SLIs / SLOs definition
Error budgets / on-call health
Automation / tools development

Tools

Kubernetes
CI/CD Tools
Observability
Hybrid Cloud

Job description

  • Co-found the SRE practice by translating the reliability charter into engineering systems, adoption paths, and measurable outcomes
  • Design and build shared reliability tooling and automation for service cataloging, SLOs, error budgets, readiness evidence, and toil reduction
  • Establish risk-based production reliability architecture, minimum standards, golden paths, and a time-bound exception process
  • Partner with Incident Management to operationalize severity, command, and response standards
  • Establish blameless post-incident practices and ensure corrective actions are tracked and completed
  • Help service teams classify critical services, map dependencies, identify failure modes, and establish meaningful SLIs/SLOs
  • Turn repeated manual operational work into reusable software and protect engineering capacity
Requirements
  • Significant experience as an SRE, Production Engineer, or Infrastructure Software Engineer operating large-scale distributed systems
  • Demonstrated experience taking an SRE or operational excellence program from 0 to 1, or materially improving a program in an organization with uneven reliability maturity
  • Strong software engineering ability in at least one general-purpose language (Go, Python, Java, C++, etc.)
  • Deep systems knowledge across several domains: Linux, networking, Kubernetes, cloud infrastructure, CI/CD, or observability
  • Practical, hands-on experience defining SLIs/SLOs, using error budgets, designing actionable alerts, and improving on-call health
  • Ability to build consensus and drive adoption of standards across teams without taking away ownership of underlying services
  • Sound judgment under ambiguity
  • Experience in autonomous vehicles, robotics, safety-relevant systems, automotive software, or another high-consequence production environment preferred
  • Experience with hybrid cloud/on-premises environments and foundational platform services preferred
  • Experience with data or ML platforms preferred
  • Experience building service catalogs, production-readiness automation, reliability scorecards, or policy-as-code preferred
  • Experience facilitating game days, failure injection, regional failover, or disaster-recovery exercises preferred
Core Competencies

Demonstrates expertise in Site Reliability Engineering (SRE) by establishing reliability practices, defining SLIs/SLOs, and implementing automation for operational excellence. Proficient in software engineering and systems knowledge across cloud infrastructure, Kubernetes, and CI/CD processes.

Highest-signal resume keywords
  • Site Reliability Engineering (SRE)
  • Software Engineering (Go, Python, Java, C++)
  • SLIs/SLOs Definition
  • Cloud Infrastructure
  • Kubernetes
Hard Skills
  • Site Reliability Engineering
  • Software Engineering
  • SLIs/SLOs Definition
  • Error Budgets
  • Production Reliability Architecture
  • Automation
  • Service Cataloging
  • CI/CD
  • Observability
  • Incident Management
Soft Skills
  • Consensus Building
  • Sound Judgment
  • Collaboration
Industry Keywords
  • Distributed Systems
  • Operational Excellence
  • High-Consequence Production Environment
  • Data Platforms
  • ML Platforms
Tools & Technologies
  • Kubernetes
  • Cloud Infrastructure
  • CI/CD Tools
  • Observability Tools
  • Hybrid Cloud
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Jobtailor • New Jersey

On-site
USD 120,000 - 180,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

JPS Tech Solutions • Colorado

On-site
USD 160,000 - 230,000
Site Reliability Engineer
Site Reliability Engineer

SCIGON • Naperville (IL)

Hybrid
USD 110,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

JobCubby • Barrington (RI), Northern (KY)

On-site
USD 110,000 - 170,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Virtual Tech Gurus • Puerto Rico

On-site
USD 140,000 - 210,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Kovoro • Denver (CO), Northern (KY)

Hybrid
USD 150,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 260,000
Senior SRE Engineer
Senior SRE Engineer

Compunnel, Inc. • Alpharetta (GA)

On-site
USD 140,000 - 190,000