Senior SRE: Automation, Reliability & Self-Healing Expert

TechDigital Group

Pittsburgh (Allegheny County)

On-site

USD 120,000 - 160,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

TechDigital Group is seeking an experienced Reliability Engineer to automate high-volume workflows, implement self-healing capabilities, and define SLOs for critical services. You will drive observability, runbooks, and self-service tooling to reduce manual intervention.

The role focuses on automating top support types, improving batch reliability, and tightening incident response with data-driven decision making in a growth-focused environment.

Qualifications

  • 8–10 years of experience in reliability engineering or DevOps roles.
  • Proficient in Java Spring Boot, Kafka, and automation of CI/CD pipelines.
  • Experience building self-healing and observable systems.
  • Ability to define SLOs and manage incident response.

Responsibilities

  • Automate the top 5 high-volume support and request types.
  • Build self-service and agent-driven solutions to reduce manual work.
  • Harden operational workflows for consistency, auditability, and resilience.
  • Implement auto-retry and backoff for recurring failure patterns.
  • Define and manage SLOs for critical services and batch processes.
  • Apply error budget concepts to guide reliability and release decisions.
  • Improve batch reliability through standardized recovery patterns and monitoring.
  • Build reliability dashboards tracking incidents, repeat issues, failure rates, and automation coverage.
  • Improve operational reporting and visibility across incidents, problems, and changes.
  • Develop and expand runbooks for key production scenarios.
  • Convert runbooks into automated remediation workflows.
  • Enable self-service for repeat operational requests.
  • Drive conversion of repeat incidents into permanent fixes and known problems.
  • Implement self-healing capabilities to minimize manual intervention.
  • Optimize alerting systems (Moogsoft) to reduce noise and improve signal quality.
  • Leverage automation and AI to resolve recurring issues with minimal human involvement.

Skills

Java Spring Boot
Apache Kafka
DevOps
CI/CD automation

Job description

TechDigital Group is seeking an experienced Reliability Engineer to automate high-volume workflows, implement self-healing capabilities, and define SLOs for critical services. You will drive observability, runbooks, and self-service tooling to reduce manual intervention.

The role focuses on automating top support types, improving batch reliability, and tightening incident response with data-driven decision making in a growth-focused environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: Self-Healing Automation & Runbooks Expert
Senior SRE: Self-Healing Automation & Runbooks Expert

100 Eli Lilly and Company • Indianapolis (IN)

On-site
USD 66,000 - 158,000
401(k) plan
Pension
Vacation benefits
+1
Senior SRE: Automate Reliability & Observability
Senior SRE: Automate Reliability & Observability

United States Digital Space LLC • Charlotte (TX)

On-site
USD 153,000 - 192,000
Discretionary incentive eligible
Benefits package
Associate SRE: Automation & Reliability
Associate SRE: Automation & Reliability

Socket.dev • United States

Remote
USD 90,000 - 120,000
Automation SRE Lead - Self-Healing & Observability
Automation SRE Lead - Self-Healing & Observability

Mphasis • New Jersey

On-site
USD 140,000 - 180,000
Senior SRE: Automation, AI-Driven Reliability (Enterprise)
Senior SRE: Automation, AI-Driven Reliability (Enterprise)

ManpowerGroup Global, Inc. • Austin (TX)

On-site
USD 66,000 - 90,000
Senior SRE & Software Engineer: Domain Reliability Lead
Senior SRE & Software Engineer: Domain Reliability Lead

Hispanic Alliance for Career Enhancement • Richardson (TX)

On-site
USD 93,000 - 204,000
Senior SRE: Scalable Infra, Observability & Automation
Senior SRE: Scalable Infra, Observability & Automation

Early Warning • Scottsdale (AZ)

Hybrid
USD 106,000 - 156,000
Healthcare Coverage
401(k) Company Match
Paid Time Off
+2
Senior SRE & Software Engineer: Reliability Lead
Senior SRE & Software Engineer: Reliability Lead

CVS Health • Woonsocket (RI)

On-site
USD 93,000 - 204,000
Global SRE Leader: Reliability, AI Ops & Excellence
Global SRE Leader: Reliability, AI Ops & Excellence

Jobtailor • Minnesota

On-site
USD 170,000 - 210,000
Associate Site Reliability Engineer: Automation & Growth
Associate Site Reliability Engineer: Automation & Growth

Verint • United States

On-site
USD 90,000 - 140,000