Lead Site Reliability Engineer: Reliability & Incidents

Referrals Only

Cincinnati (OH)

On-site

USD 120,000 - 180,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Thoughtworks is seeking a Service Reliability Engineer: Lead to drive reliability and performance across complex systems. You will lead incident response, champion DevOps and GitOps practices, and guide cross-functional teams toward scalable, resilient solutions.

You will mentor junior SREs, collaborate with development leads, and engage with client stakeholders to ensure service level objectives are met. This role requires strong English communication and crisis management skills.

Qualifications

  • Proficiency in one or more high-level languages (Python, Golang, Shell, Ruby or Java).
  • Experience with DevOps and GitOps practices across CI/CD pipelines.
  • Deep knowledge of IaC and configuration management tools for provisioning infrastructure.
  • Expertise in observability, logs, tracing and monitoring tooling.
  • Strong container orchestration and microservices architecture experience.
  • Experience tuning performance and scaling for heavy load scenarios.
  • Understanding of SLI/SLO/SLA, chaos engineering and related concepts.
  • Experience with TLS, certificate management and basic networking.

Responsibilities

  • Understand SRE goals from both technical and business perspectives.
  • Identify and implement reliability improvements and fault-tolerance architectures.
  • Enhance incident management, triage, communication, and post-mortem actions.
  • Interface with client stakeholders and provide remediation plans for incidents.
  • Collaborate with client engineering teams and leadership to drive reliability improvements.
  • Guide SRE teams in system performance enhancements aligned with SLAs/SLOs.
  • Collaborate with development leads and architects to adopt reliability best practices.
  • Mentor other SREs and contribute to their growth.

Skills

Python
Golang
Shell scripting
Ruby
Java
DevOps practices
GitOps practices
Observability automation
Capacity planning
Problem-solving
Analytical skills
Communication
Crisis management
Requirements analysis

Tools

Terraform
Ansible
ARM
CloudFormation
Grafana
Prometheus
Graylog
Jaeger
Zipkin
ELK stack
Kubernetes
AWS EKS
Docker Swarm
Nomad

Job description

Thoughtworks is seeking a Service Reliability Engineer: Lead to drive reliability and performance across complex systems. You will lead incident response, champion DevOps and GitOps practices, and guide cross-functional teams toward scalable, resilient solutions.

You will mentor junior SREs, collaborate with development leads, and engage with client stakeholders to ensure service level objectives are met. This role requires strong English communication and crisis management skills.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote SRE Lead: Reliability & Incident Response
Remote SRE Lead: Reliability & Incident Response

Thoughtworksreferral • Cincinnati (OH)

On-site
USD 155,000 - 249,000
Lead Site Reliability Engineer - Scale & Reliability
Lead Site Reliability Engineer - Scale & Reliability

Doist • United States

Remote
USD 170,000 - 230,000
Remote Lead Site Reliability Engineer: Performance & Scaling
Remote Lead Site Reliability Engineer: Performance & Scaling

Techholding • United States

Remote
USD 140,000 - 190,000
Lead Site Reliability Engineer - Architect & Own Production
Lead Site Reliability Engineer - Architect & Own Production

Optimal Market Technologies • New York (NY)

On-site
USD 175,000 - 200,000
Remote Lead SRE - Performance & Scalability
Remote Lead SRE - Performance & Scalability

Tech Holding • Northern (KY)

Hybrid
USD 150,000 - 210,000
Lead Site Reliability Engineer: AWS Cloud & Automation
Lead Site Reliability Engineer: AWS Cloud & Automation

Selby Jennings • Wilmington (NC)

On-site
USD 140,000 - 200,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Senior Site Reliability Lead: Scale & Reliability Champion
Senior Site Reliability Lead: Scale & Reliability Champion

Twitter • San Francisco (CA)

On-site
USD 130,000 - 160,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 260,000
Site Reliability Engineer Lead
Site Reliability Engineer Lead

TechDigital Group • Fairfax (VA)

On-site
USD 100,000 - 130,000