Site Reliability Engineer (SRE) – Production Services

TechDigital Group

Pittsburgh (Allegheny County)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

TechDigital Group is seeking an experienced Reliability Engineer to automate high-volume workflows, implement self-healing capabilities, and define SLOs for critical services. You will drive observability, runbooks, and self-service tooling to reduce manual intervention.

The role focuses on automating top support types, improving batch reliability, and tightening incident response with data-driven decision making in a growth-focused environment.

Qualifications

  • 8–10 years of experience in reliability engineering or DevOps roles.
  • Proficient in Java Spring Boot, Kafka, and automation of CI/CD pipelines.
  • Experience building self-healing and observable systems.
  • Ability to define SLOs and manage incident response.

Responsibilities

  • Automate the top 5 high-volume support and request types.
  • Build self-service and agent-driven solutions to reduce manual work.
  • Harden operational workflows for consistency, auditability, and resilience.
  • Implement auto-retry and backoff for recurring failure patterns.
  • Define and manage SLOs for critical services and batch processes.
  • Apply error budget concepts to guide reliability and release decisions.
  • Improve batch reliability through standardized recovery patterns and monitoring.
  • Build reliability dashboards tracking incidents, repeat issues, failure rates, and automation coverage.
  • Improve operational reporting and visibility across incidents, problems, and changes.
  • Develop and expand runbooks for key production scenarios.
  • Convert runbooks into automated remediation workflows.
  • Enable self-service for repeat operational requests.
  • Drive conversion of repeat incidents into permanent fixes and known problems.
  • Implement self-healing capabilities to minimize manual intervention.
  • Optimize alerting systems (Moogsoft) to reduce noise and improve signal quality.
  • Leverage automation and AI to resolve recurring issues with minimal human involvement.

Skills

Java Spring Boot
Apache Kafka
DevOps
CI/CD automation

Job description

Mandatory Skills
  • Java Spring boot
  • Apache Kafka
  • Dev Ops
  • CI/CD automation

Years of experience required: 8-10

Job Description
Automation & Efficiency
  • Automate the top 5 high-volume support and request types
  • Build self-service and agent-driven solutions to reduce manual work
  • Harden operational workflows for consistency, auditability, and resilience
  • Implement auto-retry and backoff for recurring failure patterns
Reliability Engineering
  • Define and manage Service Level Objectives (SLOs) for critical services and batch processes
  • Apply error budget concepts to guide reliability and release decisions
  • Improve batch reliability through standardized recovery patterns and monitoring
Observability & Metrics
  • Build reliability dashboards tracking incidents, repeat issues, failure rates, and automation coverage
  • Improve operational reporting and visibility across incidents, problems, and changes
Runbooks & Self-Service
  • Develop and expand runbooks for key production scenarios
  • Convert runbooks into automated remediation workflows
  • Enable self-service for repeat operational requests
  • Drive conversion of repeat incidents into permanent fixes and known problems
Self-Healing & Intelligent Operations
  • Implement self-healing capabilities to minimize manual intervention
  • Optimize alerting systems (Moogsoft) to reduce noise and improve signal quality
  • Leverage automation and AI to resolve recurring issues with minimal human involvement
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE Contract C2C jobs Pittsburg, PA
SRE Contract C2C jobs Pittsburg, PA

Tech Mirrors • Pittsburgh

On-site
USD 110,000 - 150,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Virtual Tech Gurus • Puerto Rico

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

SRE • Puerto Rico

Hybrid
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Jobtailor • New Hampshire

On-site
USD 110,000 - 160,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Knack Solutions • Reston (VA)

On-site
USD 120,000 - 160,000
Site Reliability engineering (SRE)
Site Reliability engineering (SRE)

TechDigital Group • San Leandro (CA)

On-site
USD 100,000 - 150,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Methodic • San Francisco (CA)

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

System One • Pittsburgh

On-site
USD 140,000 - 190,000
Site Reliability Engineer – Lead
Site Reliability Engineer – Lead

Jobtailor • Arizona

On-site
USD 140,000 - 230,000
Software Engineering Manager – Site Reliability Center
Software Engineering Manager – Site Reliability Center

Jobtailor • Alabama

On-site
USD 120,000 - 160,000