Director, Site Reliability Engineering

Jobtailor

California (MO)

On-site

USD 180,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor seeks a senior Site Reliability Engineering leader to shape the long-term SRE strategy, build scalable platforms, and drive reliability across the stack. You will lead a high-performing team, advance observability, and automate operations with AI-assisted workflows.

You will partner with executives to balance business impact with resilience, defining SLIs/SLOs, incident management, and readiness practices for major launches and events.

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, Software Engineering, or a related technical field.
  • 10+ years of progressive engineering experience.
  • 5+ years in engineering leadership managing SRE, Platform, or Systems Engineering teams.
  • Proven experience building or transforming a reliability or operational engineering organization.
  • Strong understanding of distributed systems, cloud architecture, application architecture, networking, infrastructure, and software delivery.
  • Experience establishing observability, incident-management, service-level objective, and production-readiness practices.
  • Demonstrated ability to improve reliability through engineering and automation.
  • Proven experience leading teams responsible for highly available, customer-facing, or business-critical systems.
  • Strong understanding of metrics, logs, distributed tracing, synthetic monitoring, and real-user monitoring.
  • Demonstrated success driving alignment across cross-functional engineering teams and executive stakeholders.

Responsibilities

  • Define and execute the long-term SRE strategy and roadmap.
  • Establish the SRE operating model, team scope, engagement models, ownership boundaries, and success measures.
  • Build and develop a high-performing team of site reliability and operations engineers.
  • Modernize SRE through automation, AI-assisted operations, self-service capabilities, and engineering-first practices.
  • Translate business priorities and customer impact into reliability investments and engineering outcomes.
  • Advise senior technology leaders on operational risk, resilience, capacity, and reliability tradeoffs.
  • Establish service-level indicators, service-level objectives, error budgets, and reliability standards.
  • Partner with engineering teams to design reliability, scalability, recoverability, and graceful degradation into systems.
  • Define production operational and observability readiness, including readiness reviews and certification practices.
  • Drive improvements in availability, performance, resiliency, and recovery.
  • Define enterprise observability strategy across metrics, logs, traces, events, synthetics, real-user monitoring, and business telemetry.
  • Establish instrumentation, telemetry, dashboards, alerting, and service-health standards.
  • Improve incident detection, response, mitigation, communication, learning, and on-call practices.
  • Lead automated detection, diagnosis, remediation, and incident creation.
  • Establish blameless post-incident reviews and eliminate recurring operational toil.
  • Develop intelligent-operations roadmaps covering anomaly detection, event correlation, automated triage, root-cause analysis, and remediation.
  • Evaluate AI-assisted workflows across observability, incident response, capacity planning, and operational support.
  • Build safe, measurable, auditable automation with appropriate human oversight.
  • Partner with engineering, DevOps, and business stakeholders to build shared accountability for production outcomes.
  • Support major launches and critical business events through readiness planning, risk assessment, testing, and operational coordination.

Skills

Distributed systems
Cloud architecture
Observability
Automation
Leadership

Education

Bachelor's degree in Computer Science/Computer Engineering/Software Engineering
Master's degree or MBA (preferred)

Tools

New Relic
Splunk
Datadog
Sentry
Grafana
Prometheus
OpenTelemetry

Job description

  • Define and execute the long-term Site Reliability Engineering strategy and roadmap
  • Establish the SRE operating model, team scope, engagement models, ownership boundaries, and success measures
  • Build and develop a high-performing team of site reliability and operations engineers
  • Modernize SRE through automation, AI-assisted operations, self-service capabilities, and engineering-first practices
  • Translate business priorities and customer impact into reliability investments and engineering outcomes
  • Advise senior technology leaders on operational risk, resilience, capacity, and reliability tradeoffs
  • Establish service-level indicators, service-level objectives, error budgets, and reliability standards
  • Partner with engineering teams to design reliability, scalability, recoverability, and graceful degradation into systems
  • Define production operational and observability readiness, including readiness reviews and certification practices
  • Drive improvements in availability, performance, resiliency, and recovery
  • Define enterprise observability strategy across metrics, logs, traces, events, synthetics, real-user monitoring, and business telemetry
  • Establish instrumentation, telemetry, dashboards, alerting, and service-health standards
  • Improve incident detection, response, mitigation, communication, learning, and on-call practices
  • Lead automated detection, diagnosis, remediation, and incident creation
  • Establish blameless post-incident reviews and eliminate recurring operational toil
  • Develop intelligent-operations roadmaps covering anomaly detection, event correlation, automated triage, root-cause analysis, and remediation
  • Evaluate AI-assisted workflows across observability, incident response, capacity planning, and operational support
  • Build safe, measurable, auditable automation with appropriate human oversight
  • Partner with engineering, DevOps, and business stakeholders to build shared accountability for production outcomes
  • Support major launches and critical business events through readiness planning, risk assessment, testing, and operational coordination
Requirements
  • Bachelor's degree in Computer Science, Computer Engineering, Software Engineering, or a related technical field
  • 10+ years of progressive engineering experience
  • 5+ years in engineering leadership managing SRE, Platform, or Systems Engineering teams
  • Proven experience building or transforming a reliability or operational engineering organization
  • Strong understanding of distributed systems, cloud architecture, application architecture, networking, infrastructure, and software delivery
  • Experience establishing observability, incident-management, service-level objective, and production-readiness practices
  • Demonstrated ability to improve reliability through engineering and automation
  • Proven experience leading teams responsible for highly available, customer-facing, or business-critical systems
  • Strong understanding of metrics, logs, distributed tracing, synthetic monitoring, and real-user monitoring
  • Demonstrated success driving alignment across cross-functional engineering teams and executive stakeholders
  • Ability to balance immediate operational needs with long-term engineering transformation
  • Strong written, verbal, and executive communication skills
  • Preferred: Master's degree or MBA
  • Preferred: Experience operating large-scale systems in AWS or another major cloud environment
  • Preferred: Experience with New Relic, Splunk, Datadog, Sentry, Honeycomb, Grafana, Prometheus, or OpenTelemetry
  • Preferred: Experience implementing OpenTelemetry or common instrumentation standards
  • Preferred: Experience building internal developer platforms, paved roads, or self-service reliability capabilities
  • Preferred: Experience applying AI, machine learning, or agent-based automation to operational workflows
  • Preferred: Experience with chaos engineering, resilience testing, disaster recovery, capacity planning, and performance engineering
  • Preferred: Software engineering experience and ability to engage in architecture and design discussions
  • Preferred: Experience supporting high-profile launches, events, or systems with significant customer and business impact
Core Competencies

Demonstrates expertise in Site Reliability Engineering, focusing on automation, observability, and operational excellence. Proven ability to lead high-performing teams and drive engineering transformations that enhance system reliability and performance.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer – Lead
Site Reliability Engineer – Lead

Jobtailor • Arizona

On-site
USD 140,000 - 230,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Luxoft • Wilmington (DE)

On-site
USD 140,000 - 190,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 260,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Luxoft • Buffalo (NY)

On-site
USD 140,000 - 190,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Virtual Tech Gurus • Puerto Rico

On-site
USD 140,000 - 210,000
Software Engineering Manager – Site Reliability Center
Software Engineering Manager – Site Reliability Center

Jobtailor • Alabama

On-site
USD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

Jobtailor • New Hampshire

On-site
USD 110,000 - 160,000
Senior Manager SRE
Senior Manager SRE

Expedite Talent Solutions • United States

Hybrid
USD 130,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Veloc Inc • Coppell (TX)

On-site
USD 140,000 - 190,000