Site Reliability Engineer

System One

Pittsburgh (Allegheny County)

On-site

USD 140,000 - 190,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

System One is seeking a Senior Site Reliability Engineer (SRE) to support production operations, reliability, and performance management across multiple environments. The role requires hands-on incident management, log analysis, automation, and collaboration with distributed teams to maintain uptime for critical applications.

Ideal candidates will have strong Linux/Windows experience, Oracle/PL/SQL/DB2 knowledge, and experience with Dynatrace, ITCAM, OEM in an operational production setting.

Qualifications

  • 5+ years in IT with a focus on production operations and site reliability.
  • 2–3 years in business analytics and technical leadership.
  • Experience with production/application support in client-facing environments.
  • Strong understanding of SRE and production operations.
  • Hands-on troubleshooting and log analysis in production.
  • Familiarity with monitoring/analytics tools and incident management.
  • Experience with automation and process improvement.

Responsibilities

  • Monitor distributed systems and proactively identify issues.
  • Support troubleshooting and participate in on-call activities.
  • Manage production incidents and outages; coordinate response.
  • Lead incident-analysis and problem-management meetings.
  • Identify opportunities for automation in operations.
  • Monitor applications and infrastructure to ensure reliability.
  • Coordinate follow-up through incident resolution and closure.
  • Troubleshoot using system and application logs; reproduce issues.

Skills

SRE
Production Support
Incident Management
Log Analysis
Automation
Root Cause Analysis
Agile
Technical Leadership
Client-Facing Production Support
Performance Management
Monitoring
Distributed Teams

Tools

Dynatrace
GlassBox
ITCAM
TrueSight
Oracle Enterprise Manager (OEM)
Tomcat
Apache
WebSphere (WAS)
IIS
Oracle
PL/SQL
DB2

Job description

Senior Site Reliability Engineer (SRE)
Location: Pittsburgh, PA / Cleveland, OH / Dallas, TX
FTE

Position Overview

We are seeking an experienced Senior Site Reliability Engineer (SRE) to support production operations, application reliability, performance management, and continuous improvement initiatives. The selected candidate will work closely with production support and engineering teams to ensure critical internal and external applications maintain appropriate levels of availability, reliability, and uptime. This role requires strong experience in production support, incident management, monitoring, troubleshooting, log analysis, automation identification, infrastructure technologies, databases, and application servers. The SRE will also provide technical leadership and collaborate with geographically distributed teams.

Key Skills
  • Site Reliability Engineering (SRE)
  • Production Support / Application Support
  • Incident & Problem Management
  • Linux
  • Windows Server
  • Oracle / PL/SQL / DB2
  • Dynatrace / DT Managed
  • GlassBox / ITCAM / TrueSight / OEM
  • Tomcat / Apache / WebSphere (WAS) / IIS
  • REST & SOAP Web Services
  • Log Analysis & Troubleshooting
  • AIOps / NLP
  • Monitoring & Performance Management
  • Automation
  • Root Cause Analysis
  • Business Analytics
  • Agile
  • Technical Leadership
  • Client-Facing Production Support
Responsibilities
  • Monitor distributed systems and proactively identify potential production issues.
  • Support troubleshooting and participate in on-call activities.
  • Manage, track, and coordinate production incidents and application outages.
  • Lead incident-analysis and problem-management meetings.
  • Identify opportunities for operational and production-support automation.
  • Monitor applications and related infrastructure to maintain system reliability.
  • Coordinate follow-up activities through incident resolution and closure.
  • Troubleshoot complex application issues using system and application logs.
  • Participate in critical incident calls and contribute technical expertise toward resolution.
  • Perform root cause analysis and recommend corrective actions.
  • Research and reproduce user issues to validate solutions.
  • Resolve technical problems that cannot be handled by junior team members.
  • Provide technical guidance and solutions to the production-support team.
  • Introduce process improvements and innovative solutions for operational challenges.
  • Develop and maintain SOPs, operational procedures, and knowledge documentation.
  • Collaborate with offshore and geographically distributed teams.
  • Work with client technical teams, SMEs, and leadership.
  • Support extended or weekend hours when required during critical production events.
  • Participate in overlapping business-hour shifts for critical meetings and activities.
Required Qualifications
  • 5+ years of overall IT experience.
  • 2–3 years of business analytics and technical leadership experience.
  • Strong experience with production/application support in a client-facing environment.
  • Strong understanding of Site Reliability Engineering and production operations.
  • Hands-on experience troubleshooting production applications and analyzing log files.
  • Strong knowledge of system-management, monitoring, and support analytics tools.
  • Experience with incident management, root cause analysis, and problem resolution.
  • Strong understanding of AIOps and NLP concepts.
  • Experience identifying opportunities for automation and process improvement.
  • Strong problem-solving and analytical capabilities.
  • Ability to recommend efficient and cost-effective technical solutions.
  • Experience working with geographically distributed/onshore-offshore teams.
  • Excellent client-facing verbal and written communication skills.
Database Technologies

Strong knowledge of:

  • Oracle
  • PL/SQL
  • DB2
Web Services

Experience developing and consuming:

  • REST APIs
  • SOAP Web Services

Experience should preferably be within an operational/production environment.

Application Servers / Web Servers

Strong knowledge of:

  • Tomcat
  • Apache
  • WebSphere (WAS)
  • IIS
Operating Systems
  • Extensive experience with Linux
  • Good understanding of Windows Server
  • Linux and Windows server configuration and troubleshooting
Monitoring & Support Tools

Experience with monitoring tools such as:

  • Dynatrace
  • Dynatrace Managed / DT Managed
  • GlassBox
  • ITCAM / ITCAMS
  • TrueSight
  • Oracle Enterprise Manager (OEM)
Additional Skills
  • Agile methodology
  • SOP and technical documentation
  • Performance management
  • System reliability and availability
  • Production incident coordination
  • Technical research and solution evaluation
  • Process improvement
  • Automation opportunity identification
  • Strong stakeholder and client communication
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

System One • Dallas (TX)

On-site
USD 130,000 - 170,000
Software Engineering Manager – Site Reliability Center
Software Engineering Manager – Site Reliability Center

Jobtailor • Alabama

On-site
USD 120,000 - 160,000
SRE Production Support
SRE Production Support

SelectMinds LLC • Livonia (MI)

On-site
USD 100,000 - 140,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Luxoft • Buffalo (NY)

On-site
USD 140,000 - 190,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Luxoft • United States

On-site
USD 140,000 - 180,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Luxoft • Wilmington (DE)

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Knack Solutions • Reston (VA)

On-site
USD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

Brooksource • San Antonio (TX)

On-site
USD 80,000 - 120,000
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)

Mindlance • Irving (TX)

Hybrid
USD 120,000 - 160,000