Site Reliability Engineer in Pittsburgh

Energy Jobline ZR

Pittsburgh (Allegheny County)

On-site

USD 130,000 - 180,000

Full time

10 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Energy Jobline ZR seeks a Senior Site Reliability Engineer to bolster production operations and application reliability across multiple sites, including Pittsburgh, Cleveland, and Dallas. You will lead incident management, perform log analysis, iterate on automation, and mentor teams across distributed locations.

Ideal candidates bring 5+ years in IT with hands-on SRE, production support, and strong leadership skills.

Qualifications

  • 5+ years of overall IT experience.
  • 2–3 years of business analytics and technical leadership experience.
  • Strong experience with production/application support
  • Strong understanding of Site Reliability Engineering and production operations.
  • Hands-on experience troubleshooting production applications and analyzing log files.
  • Strong knowledge of system-management, monitoring, and support analytics tools.
  • Experience with incident management, root cause analysis, and problem resolution.
  • Strong understanding of AIOps and NLP concepts.
  • Experience identifying opportunities for automation and process improvement.
  • Strong problem-solving and analytical capabilities.
  • Ability to recommend efficient and cost-effective technical solutions.
  • Experience working with geographically distributed/onshore-offshore teams.
  • Excellent client-facing verbal and written communication skills.

Responsibilities

  • Monitor distributed systems and proactively identify potential production issues.
  • Support troubleshooting and participate in on-call activities.
  • Manage, track, and coordinate production incidents and application outages.
  • Lead incident-analysis and problem-management meetings.
  • Identify opportunities for operational and production-support automation.
  • Monitor applications and related infrastructure to maintain system reliability.
  • Coordinate follow-up activities through incident resolution and closure.
  • Troubleshoot complex application issues using system and application logs.
  • Participate in critical incident calls and contribute technical expertise toward resolution.
  • Perform root cause analysis and recommend corrective actions.
  • Research and reproduce user issues to validate solutions.
  • Resolve technical problems that cannot be handled by junior team members.
  • Provide technical guidance and solutions to the production-support team.
  • Introduce process improvements and innovative solutions for operational challenges.
  • Develop and maintain SOPs, operational procedures, and knowledge documentation.
  • Collaborate with offshore and geographically distributed teams.
  • Work with client technical teams, SMEs, and leadership.
  • Support extended or weekend hours when required during critical production events.
  • Participate in overlapping business-hour shifts for critical meetings and activities.

Skills

Site Reliability Engineering
Production Support
Incident Management
Log Analysis
Automation
Root Cause Analysis
Agile
Technical Leadership
Client-Facing Production Support

Tools

Dynatrace
DT Managed
GlassBox
ITCAM
ITCAMs
TrueSight
Oracle Enterprise Manager (OEM)
Tomcat
Apache
WebSphere (WAS)
IIS

Job description

Job DescriptionJob Description
Senior Site Reliability Engineer (SRE)

Location: Pittsburgh, PA / Cleveland, OH / Dallas, TX

FTE

Position Overview

We are seeking an experienced Senior Site Reliability Engineer (SRE) to support production operations, application reliability, performance management, and continuous improvement initiatives. The selected candidate will work closely with production support and engineering teams to ensure critical internal and external applications maintain appropriate levels of availability, reliability, and uptime. This role requires strong experience in production support, incident management, monitoring, troubleshooting, log analysis, automation identification, infrastructure technologies, databases, and application servers. The SRE will also provide technical leadership and collaborate with geographically distributed teams.

Key Skills
  • Site Reliability Engineering (SRE)
  • Production Support / Application Support
  • Incident & Problem Management
  • Linux
  • Windows Server
  • Oracle / PL/SQL / DB2
  • Dynatrace / DT Managed
  • GlassBox / ITCAM / TrueSight / OEM
  • Tomcat / Apache / WebSphere (WAS) / IIS
  • REST & SOAP Web Services
  • Log Analysis & Troubleshooting
  • AIOps / NLP
  • Monitoring & Performance Management
  • Automation
  • Root Cause Analysis
  • Business Analytics
  • Agile
  • Technical Leadership
  • Client-Facing Production Support
Responsibilities
  • Monitor distributed systems and proactively identify potential production issues.
  • Support troubleshooting and participate in on-call activities.
  • Manage, track, and coordinate production incidents and application outages.
  • Lead incident-analysis and problem-management meetings.
  • Identify opportunities for operational and production-support automation.
  • Monitor applications and related infrastructure to maintain system reliability.
  • Coordinate follow-up activities through incident resolution and closure.
  • Troubleshoot complex application issues using system and application logs.
  • Participate in critical incident calls and contribute technical expertise toward resolution.
  • Perform root cause analysis and recommend corrective actions.
  • Research and reproduce user issues to validate solutions.
  • Resolve technical problems that cannot be handled by junior team members.
  • Provide technical guidance and solutions to the production-support team.
  • Introduce process improvements and innovative solutions for operational challenges.
  • Develop and maintain SOPs, operational procedures, and knowledge documentation.
  • Collaborate with offshore and geographically distributed teams.
  • Work with client technical teams, SMEs, and leadership.
  • Support extended or weekend hours when required during critical production events.
  • Participate in overlapping business-hour shifts for critical meetings and activities.
Required Qualifications
  • 5+ years of overall IT experience.
  • 2–3 years of business analytics and technical leadership experience.
  • Strong experience with production/application support
  • Strong understanding of Site Reliability Engineering and production operations.
  • Hands-on experience troubleshooting production applications and analyzing log files.
  • Strong knowledge of system-management, monitoring, and support analytics tools.
  • Experience with incident management, root cause analysis, and problem resolution.
  • Strong understanding of AIOps and NLP concepts.
  • Experience identifying opportunities for automation and process improvement.
  • Strong problem-solving and analytical capabilities.
  • Ability to recommend efficient and cost-effective technical solutions.
  • Experience working with geographically distributed/onshore-offshore teams.
  • Excellent client-facing verbal and written communication skills.
Database Technologies

Strong knowledge of:

  • Oracle
  • PL/SQL
  • DB2
Web Services

Experience developing and consuming:

  • REST APIs
  • SOAP Web Services

Experience should preferably be within an operational/production environment.

Application Servers / Web Servers

Strong knowledge of:

  • Tomcat
  • Apache
  • WebSphere (WAS)
  • IIS
Operating Systems
  • Extensive experience with Linux
  • Good understanding of Windows Server
  • Linux and Windows server configuration and troubleshooting
Monitoring & Support Tools

Experience with monitoring tools such as:

  • Dynatrace
  • Dynatrace Managed / DT Managed
  • GlassBox
  • ITCAM / ITCAMS
  • TrueSight
  • Oracle Enterprise Manager (OEM)
Additional Skills
  • Agile methodology
  • SOP and technical documentation
  • Performance management
  • System reliability and availability
  • Production incident coordination
  • Technical research and solution evaluation
  • Process improvement
  • Automation opportunity identification
  • Strong stakeholder and client communication

#M1 #DI-CB2 #L1 - KB1

Ref: #404-IT Pittsburgh

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

System One • Dallas (TX)

On-site
USD 130,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

Jobtailor • New Jersey

On-site
USD 120,000 - 180,000
Site Reliability Engineer -Jersey City, NJ & Dallas, TX
Site Reliability Engineer -Jersey City, NJ & Dallas, TX

StradIT • Jersey City (NJ)

Hybrid
USD 120,000 - 160,000
Software Engineering Manager-Site Reliability
Software Engineering Manager-Site Reliability

Fairygodboss • Farmers Branch (TX)

Hybrid
USD 150,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Brooksource • San Antonio (TX)

On-site
USD 80,000 - 120,000
Site Reliability Engineer -Jersey City, NJ & Dallas, TX
Site Reliability Engineer -Jersey City, NJ & Dallas, TX

StradIT • Dallas (TX)

Hybrid
USD 140,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Senior Vice President, Site Reliability Engineer
Senior Vice President, Site Reliability Engineer

BNY Mellon • Town of Florida (NY)

On-site
USD 170,000 - 230,000
Site Reliability Engineer -- SINDC5717546
Site Reliability Engineer -- SINDC5717546

Compunnel Inc. • Denton (TX)

Hybrid
USD 120,000 - 150,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Knack Solutions • Reston (VA)

On-site
USD 120,000 - 160,000