Senior Staff SRE: Reliability & Observability Leader

Early Warning Services, LLC

United States

Hybrid

USD 150,000 - 200,000

Full time

7 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Healthcare coverage
401(k)
Paid time off
Parental leave
Maven family planning

Job summary

The Senior Staff Site Reliability Engineer at Early Warning Services, LLC will apply software and systems engineering practices to enhance the reliability, resilience, and performance of production services across multiple domains. You will lead collaboration with Software Engineering and other technology teams to embed observability, recoverability, and operational readiness throughout the lifecycle.

This role provides broad technical leadership across a major technology domain, guiding

Qualifications

  • Typicall y 12+ years of relevant professional experience in Software Engineering, Site Reliability Engineering, Systems Engineering, Cloud/Platform Engineering, DevOps, Infrastructure Engineering, Architecture where applicable, or a comparable technical discipline.
  • Experience with software development or scripting using one or more modern programming languages.
  • Experience with software engineering principles, distributed systems, production troubleshooting, automation, and observability appropriate to the level.
  • Experience with public cloud technologies and architectures, preferably AWS, along with infrastructure, networking, Linux/Unix, and modern application architectures appropriate to the level.
  • Demonstrated analytical, problem-solving, communication, and collaboration skills appropriate to the scope of the role.

Responsibilities

  • Use software engineering, automation, and DevOps principles and practices to continually improve how services are built, tested, deployed, observed, operated, and recovered.
  • Use data, evidence, experimentation, and rigorous engineering analysis appropriate to the level to identify reliability risks, test assumptions, and guide technical decisions.
  • Define, implement, or improve SLIs, SLOs, error budgets, and other service-health measures appropriate to the scope of responsibility.
  • Improve observability through metrics, logging, tracing, monitoring, alerting, dashboards, and service-health instrumentation.
  • Drive continuous improvement across CI/CD, observability, deployment practices, Infrastructure as Code, automation, testing, incident response, capacity management, resilience, and operational readiness.
  • Identify recurring or systemic production issues and translate operational experience into improvements in code, architecture, automation, tooling, and engineering practices.
  • Partner with Software Engineering teams to incorporate reliability, resiliency, scalability, performance, observability, recoverability, and operational readiness throughout the development lifecycle.
  • Participate in or lead incident response, troubleshooting, service restoration, and blameless post-incident learning appropriate to the level.
  • Provides senior technical leadership and escalation support for significant production incidents and drives improvements to incident response and sustainable on-call practices across a pillar or broad technical domain.
  • Reduce operational toil and unnecessary manual intervention through software, automation, reusable patterns, and better engineering practices.

Skills

Software engineering
Site reliability engineering
Cloud/Platform engineering
DevOps
Automation
Observability
Leadership

Education

Bachelor’s degree in Computer Science or related field

Tools

AWS
Kubernetes
Terraform
CI/CD
Monitoring tools

Job description

The Senior Staff Site Reliability Engineer at Early Warning Services, LLC will apply software and systems engineering practices to enhance the reliability, resilience, and performance of production services across multiple domains. You will lead collaboration with Software Engineering and other technology teams to embed observability, recoverability, and operational readiness throughout the lifecycle.

This role provides broad technical leadership across a major technology domain, guiding

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Staff SRE: Reliability Platform Leader
Senior Staff SRE: Reliability Platform Leader

Socket.dev • Scottsdale (AZ)

On-site
USD 150,000 - 200,000
Healthcare coverage
401(k) match
Paid time off
+2
Senior Staff SRE — Reliability, Observability & Automation
Senior Staff SRE — Reliability, Observability & Automation

Early Warning Services LLC • Scottsdale (AZ)

Hybrid
USD 150,000 - 200,000
Healthcare coverage
401(k) with company match
Paid time off & holidays
+2
Senior SRE: Scale Reliability, Observability & Resilience
Senior SRE: Scale Reliability, Observability & Resilience

Early Warning Services LLC • Scottsdale (AZ)

Hybrid
USD 106,000 - 130,000
Healthcare Coverage
401(k) Retirement Plan
Paid Time Off
+2
Senior SRE - Hybrid, Observability & Reliability
Senior SRE - Hybrid, Observability & Reliability

Early Warning Services LLC • Chicago (IL)

Hybrid
USD 106,000 - 130,000
Healthcare Coverage
401(k) Plan with match
PTO and Holidays
+1
Enterprise Site Reliability Engineer Lead
Enterprise Site Reliability Engineer Lead

Early Warning® • Scottsdale (AZ)

Hybrid
USD 194,000 - 284,000
Health coverage
401(k) match
PTO & Holidays
+2
Principal SRE - Enterprise Reliability Leader (Hybrid)
Principal SRE - Enterprise Reliability Leader (Hybrid)

Early Warning Services LLC • San Francisco (CA)

Hybrid
USD 207,000 - 276,000
Healthcare coverage
401(k) match
Paid time off
+1
Principal SRE: Enterprise Reliability & Automation
Principal SRE: Enterprise Reliability & Automation

Early Warning Services LLC • Chicago (IL)

Hybrid
USD 194,000 - 284,000
Healthcare Coverage
401(k) Retirement Plan
Paid Time Off
+1
Senior Site Reliability Engineer - Reliability Leader
Senior Site Reliability Engineer - Reliability Leader

Early Warning Services LLC • San Francisco (CA)

Hybrid
USD 128,000 - 156,000
Healthcare Coverage
401(k) Match
Paid Time Off
+2
Lead Enterprise SRE & Reliability Engineering
Lead Enterprise SRE & Reliability Engineering

Early Warning • Chicago (IL)

Hybrid
USD 194,000 - 237,000
Healthcare coverage
401(k) matching
Paid time off
+2
Principal SRE — Enterprise Reliability & Automation
Principal SRE — Enterprise Reliability & Automation

Early Warning • San Francisco (CA)

Hybrid
USD 207,000 - 276,000
Healthcare Coverage
401(k) Retirement Match
Paid Time Off
+2