Software Engineering Manager-Site Reliability

Fairygodboss

Farmers Branch (TX)

Hybrid

USD 150,000 - 210,000

Full time

3 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

PNC is seeking a Software Engineering Manager - Site Reliability Engineering (SRE) to lead a team responsible for reliability, scalability, and operational excellence of mission-critical platforms powering digital experiences. The role blends technical leadership, hands-on problem solving, and people management across distributed systems.

You will drive production stability, implement proactive reliability practices, and oversee monitoring, logging, and automation initiatives while collaborating

Qualifications

  • 5+ years of related experience and 3+ years of management experience.
  • Strong experience in Site Reliability Engineering, Production Support, or DevOps.
  • Proven ability to lead teams in high-availability, enterprise environments.
  • Deep understanding of incident, problem, and change management frameworks.
  • Hands-on knowledge of monitoring tools, cloud/infrastructure platforms, and automation.
  • Experience improving system reliability, observability, and operational maturity.
  • Strong communication skills with the ability to lead during high-pressure situations.

Responsibilities

  • Lead SRE teams; coach, develop, set goals, and drive ownership; align tech and business objectives.
  • Provide after-hours leadership and on-call support during critical incidents.
  • Lead incident management and remediation across application, infrastructure, and databases.
  • Provide technical leadership in production support and escalation point for complex issues.
  • Drive root-cause analysis and permanent fixes to reduce recurring incidents.
  • Oversee change management and release execution with risk assessments and CAB representation.

Skills

SRE leadership
Production Support
DevOps
Incident management
Observability
Automation
Communication
High-availability
Team leadership

Tools

MongoDB
Cassandra
Elasticsearch
Redis
Kafka
Oracle
SQL
Linux
Windows

Job description

Position Overview

At PNC, our people are our greatest differentiator and competitive advantage in the markets we serve. We are all united in delivering the best experience for our customers. We work together each day to foster an inclusive workplace culture where all of our employees feel respected, valued and have an opportunity to contribute to the company's success. As a(n) [position title] within PNC's [name of division] organization, you will be based in [city/state location of position].

Job Profile
Position Overview

At PNC, our people are our greatest differentiator and competitive advantage in the markets we serve. We are all united in delivering the best experience for our customers. We work together each day to foster an inclusive workplace culture where all of our employees feel respected, valued and have an opportunity to contribute to the company's success. As a Software Engineering Manager within PNC's Site Reliability organization, you will be based in one of these Technology Hub locations: Pittsburgh, PA, Cleveland, OH, Birmingham, AL, Dallas, TX, Phoenix AZ. Weekly time in the office is needed.

Needed skills/experience:
  • 5 + years of related experience and 3+ years of management experience.
  • Strong experience in Site Reliability Engineering, Production Support, or DevOps.
  • Proven ability to lead teams in high-availability, enterprise environments
  • Deep understanding of incident, problem, and change management frameworks
  • Hands-on knowledge of monitoring tools, cloud/infrastructure platforms, and automation
  • Experience improving system reliability, observability, and operational maturity
  • Strong communication skills with the ability to lead during high-pressure situations.
  • Experience with OCP under infrastructure (Linux/Windows, OCP), MongoDB, Cassandra under databases (Oracle, SQL, MongoDB, Cassandra) and working knowledge of Elasticsearch, Redis, MQ and Kafka is a plus.

The Site Reliability Center (SRC) is focused on establishing a culture of operational excellence by ensuring infrastructure, platforms, and applications adhere to SRC onboarding standards that improve reliability, enable proactive issue resolution, and reduce customer impact. This role supports the vision of building a collaborative technology organization across application, infrastructure, and security teams to deliver a stable, reliable, and secure environment. Key responsibilities include driving customer-centric service improvements, implementing proactive and preventative reliability practices, fostering cross-functional collaboration, enhancing monitoring and observability capabilities, promoting a blameless culture of continuous learning, and reducing operational toil through automation. The ideal candidate will help improve service performance, strengthen operational resiliency, and advance automation and observability initiatives that enhance the overall customer experience.

As a Software Engineering Manager - Site Reliability Engineering (SRE), you will lead a team responsible for ensuring the reliability, scalability, and operational excellence of mission-critical platforms that power PNC's digital experiences. This role blends technical leadership, hands-on problem solving, and people management, driving both production stability and continuous improvement across complex distributed systems. You will... .

  • Manage SRE and related Teams; lead, coach, and develop a team of SRE engineers; set clear goals, drive accountability, and foster a culture of ownership and excellence; partner with cross-functional stakeholders to align technology and business objectives; support talent development, performance management, and succession planning; encourage innovation, continuous learning, and DevOps/SRE best practices.
  • Provide after-hours operational leadership and on-call support. Participate in an on-call leadership rotation supporting critical production services, major incidents, and high-severity customer-impacting events. Availability outside standard business hours, including evenings, weekends, and holidays, may be required to support incident response, change events, escalations, and business continuity needs.
  • Lead incident management & remediation; manage and actively participate in end-to-end incident response for major (P1/P2) incidents; guide real-time triage, diagnostics, and troubleshooting across application, infrastructure, and network layers; ensure rapid execution of remediation actions and service restoration; provide clear, timely communication to stakeholders during incidents; oversee post-incident analysis, reporting, and documentation to drive improvements.
  • Provide technical leadership in production support; serve as an escalation point for complex production issues; guide troubleshooting across: applications, infrastructure (Linux/Windows), databases (Oracle, SQL), middleware and integrations; ensure efficient log, metric, and system analysis; oversee batch/ETL monitoring and recovery processes; foster strong collaboration across engineering, infrastructure, and vendor teams.
  • Drive problem management & root cause resolution; lead root cause analysis (RCA) efforts for major and recurring incidents; ensure ownership and resolution of problem records; drive permanent fixes and systemic improvements to eliminate repeat issues, identify trends and patterns to reduce risk and improve stability; partner with engineering teams to resolve code defects and system gaps and promote knowledge sharing via runbooks, knowledge articles, and error catalogs.
  • Oversee change management & release execution; ensure safe and compliant execution of production changes and releases; validate change readiness, testing, rollback strategies, and risk assessments; represent the team in CAB reviews, providing technical risk evaluation; oversee post-implementation reviews (CPIR) and ensure follow-through and drive improvements in change success rate and reduction in production defects.
  • Advance monitoring, alert
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer in Pittsburgh
Site Reliability Engineer in Pittsburgh

Energy Jobline ZR • Pittsburgh

On-site
USD 130,000 - 180,000
SRE Engineering Manager — Lead 24x7 Reliability (On‑Site)
SRE Engineering Manager — Lead 24x7 Reliability (On‑Site)

PNC Financial Services Group, Inc. • United States

On-site
USD 101,000 - 186,000
Medical/prescription coverage
Dental and vision options
401(k) with company match
+2
SRE Engineering Manager - Lead 24x7 Ops & Reliability
SRE Engineering Manager - Lead 24x7 Ops & Reliability

PNC Financial Services Group, Inc. • Pittsburgh, Northern (KY)

Hybrid
USD 100,000 - 186,000
Software Engineer Manager
Software Engineer Manager

Jobtailor • Alabama

On-site
USD 120,000 - 180,000
SRE Software Engineering Manager: Lead Reliability
SRE Software Engineering Manager: Lead Reliability

Fairygodboss • Birmingham (AL)

On-site
USD 110,000 - 204,000
Site Reliability Engineering Manager — Production Stability
Site Reliability Engineering Manager — Production Stability

Fairygodboss • Farmers Branch (TX)

Hybrid
USD 150,000 - 210,000
SRE Engineering Manager — 24x7 Production Reliability
SRE Engineering Manager — 24x7 Production Reliability

PNC • Pittsburgh

On-site
USD 115,000 - 150,000
Medical and prescription drug coverage
401(k) with PNC match
Paid time off and holiday benefits
+1
Site Reliability Software Engineering Manager
Site Reliability Software Engineering Manager

Fairygodboss • Cleveland (OH)

On-site
USD 100,000 - 204,000
Software Engineer Manager
Software Engineer Manager

Fairygodboss • Birmingham (AL)

On-site
USD 110,000 - 204,000
Site Reliability Engineer Sr.
Site Reliability Engineer Sr.

PNC • Cleveland (OH)

On-site
USD 86,000 - 144,000
Competitive salary
In-office role
Benefits package
+1