Software Engineering Manager-Site Reliability

Fairygodboss

Phoenix (AZ)

Hybrid

USD 140,000 - 190,000

Full time

5 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

PNC is seeking an experienced Software Engineering Manager – Site Reliability Engineering (SRE) to lead a team ensuring reliability, scalability, and operational excellence of mission-critical platforms powering PNC's digital experiences. You will blend technical leadership with hands-on problem solving and people management across distributed systems.

The role requires managing SRE engineers, participating in on-call rotations, driving incident response, and aligning cross-functional teams to

Qualifications

  • 5+ years of related experience and 3+ years of management experience.
  • Strong experience in Site Reliability Engineering, Production Support, or DevOps.
  • Proven ability to lead teams in high-availability, enterprise environments.
  • Deep understanding of incident, problem, and change management frameworks.
  • Hands-on knowledge of monitoring tools, cloud/infrastructure platforms, and automation.
  • Experience improving system reliability, observability, and operational maturity.
  • Strong communication skills with the ability to lead during high-pressure situations.
  • Experience with OCP under infrastructure (Linux/Windows, OCP), MongoDB, Cassandra under databases (Oracle, SQL, MongoDB, Cassandra) and working knowledge of Elasticsearch, Redis, MQ and Kafka is a plus.

Responsibilities

  • Manage SRE and related teams; lead, coach, and develop a team of SRE engineers; set clear goals and foster ownership and excellence.
  • Provide after-hours operational leadership and on-call support.
  • Lead incident management & remediation for major incidents; guide real-time triage and troubleshooting; ensure rapid remediation and communications.
  • Provide technical leadership in production support; escalation point for complex production issues; guide troubleshooting across applications and infrastructure.
  • Drive root cause analysis for major incidents; implement permanent fixes and knowledge sharing via runbooks and articles.
  • Oversee change management and release execution; validate readiness, testing, rollback strategies, and risk assessments.
  • Advance monitoring, alerting, and observability; promote best practices and reduce toil with automation.

Skills

Site Reliability Engineering
Production Support
DevOps
Incident management
Cloud/infrastructure
Automation
Communication under pressure
Leadership

Tools

OCP (Linux/Windows)
MongoDB
Cassandra
Oracle
SQL
Elasticsearch
Redis
MQ
Kafka

Job description

Position Overview

At PNC, our people are our greatest differentiator and competitive advantage in the markets we serve. We are all united in delivering the best experience for our customers. We work together each day to foster an inclusive workplace culture where all of our employees feel respected, valued and have an opportunity to contribute to the company's success. As a(n) [position title] within PNC's [name of division] organization, you will be based in [city/state location of position].

Job Profile
Position Overview

At PNC, our people are our greatest differentiator and competitive advantage in the markets we serve. We are all united in delivering the best experience for our customers. We work together each day to foster an inclusive workplace culture where all of our employees feel respected, valued and have an opportunity to contribute to the company's success. As a Software Engineering Manager within PNC's Site Reliability organization, you will be based in one of these Technology Hub locations: Pittsburgh, PA, Cleveland, OH, Birmingham, AL, Dallas, TX, Phoenix AZ. Weekly time in the office is needed.

Needed skills/experience:
  • 5 + years of related experience and 3+ years of management experience.
  • Strong experience in Site Reliability Engineering, Production Support, or DevOps.
  • Proven ability to lead teams in high-availability, enterprise environments
  • Deep understanding of incident, problem, and change management frameworks
  • Hands-on knowledge of monitoring tools, cloud/infrastructure platforms, and automation
  • Experience improving system reliability, observability, and operational maturity
  • Strong communication skills with the ability to lead during high-pressure situations.
  • Experience with OCP under infrastructure (Linux/Windows, OCP), MongoDB, Cassandra under databases (Oracle, SQL, MongoDB, Cassandra) and working knowledge of Elasticsearch, Redis, MQ and Kafka is a plus.

The Site Reliability Center (SRC) is focused on establishing a culture of operational excellence by ensuring infrastructure, platforms, and applications adhere to SRC onboarding standards that improve reliability, enable proactive issue resolution, and reduce customer impact. This role supports the vision of building a collaborative technology organization across application, infrastructure, and security teams to deliver a stable, reliable, and secure environment. Key responsibilities include driving customer-centric service improvements, implementing proactive and preventative reliability practices, fostering cross-functional collaboration, enhancing monitoring and observability capabilities, promoting a blameless culture of continuous learning, and reducing operational toil through automation. The ideal candidate will help improve service performance, strengthen operational resiliency, and advance automation and observability initiatives that enhance the overall customer experience.

As a Software Engineering Manager - Site Reliability Engineering (SRE), you will lead a team responsible for ensuring the reliability, scalability, and operational excellence of mission-critical platforms that power PNC's digital experiences. This role blends technical leadership, hands-on problem solving, and people management, driving both production stability and continuous improvement across complex distributed systems. You will... .

  • Manage SRE and related Teams; lead, coach, and develop a team of SRE engineers; set clear goals, drive accountability, and foster a culture of ownership and excellence; partner with cross-functional stakeholders to align technology and business objectives; support talent development, performance management, and succession planning; encourage innovation, continuous learning, and DevOps/SRE best practices.
  • Provide after-hours operational leadership and on-call support. Participate in an on-call leadership rotation supporting critical production services, major incidents, and high-severity customer-impacting events. Availability outside standard business hours, including evenings, weekends, and holidays, may be required to support incident response, change events, escalations, and business continuity needs.
  • Lead incident management & remediation; manage and actively participate in end-to-end incident response for major (P1/P2) incidents; guide real-time triage, diagnostics, and troubleshooting across application, infrastructure, and network layers; ensure rapid execution of remediation actions and service restoration; provide clear, timely communication to stakeholders during incidents; oversee post-incident analysis, reporting, and documentation to drive improvements.
  • Provide technical leadership in production support; serve as an escalation point for complex production issues; guide troubleshooting across: applications, infrastructure (Linux/Windows), databases (Oracle, SQL), middleware and integrations; ensure efficient log, metric, and system analysis; oversee batch/ETL monitoring and recovery processes; foster strong collaboration across engineering, infrastructure, and vendor teams.
  • Drive problem management & root cause resolution; lead root cause analysis (RCA) efforts for major and recurring incidents; ensure ownership and resolution of problem records; drive permanent fixes and systemic improvements to eliminate repeat issues, identify trends and patterns to reduce risk and improve stability; partner with engineering teams to resolve code defects and system gaps and promote knowledge sharing via runbooks, knowledge articles, and error catalogs.
  • Oversee change management & release execution; ensure safe and compliant execution of production changes and releases; validate change readiness, testing, rollback strategies, and risk assessments; represent the team in CAB reviews, providing technical risk evaluation; oversee post-implementation reviews (CPIR) and ensure follow-through and drive improvements in change success rate and reduction in production defects.
  • Advance monitoring, alert
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineering Manager-Site Reliability
Software Engineering Manager-Site Reliability

Fairygodboss • Farmers Branch (TX)

Hybrid
USD 150,000 - 210,000
Site Reliability Engineer in Pittsburgh
Site Reliability Engineer in Pittsburgh

Energy Jobline ZR • Pittsburgh

On-site
USD 130,000 - 180,000
SRE Engineering Manager - Lead 24x7 Ops & Reliability
SRE Engineering Manager - Lead 24x7 Ops & Reliability

PNC Financial Services Group, Inc. • Pittsburgh, Northern (KY)

Hybrid
USD 100,000 - 186,000
Software Engineer Manager
Software Engineer Manager

Jobtailor • Alabama

On-site
USD 120,000 - 180,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

Fairygodboss • Phoenix (AZ)

Hybrid
USD 140,000 - 190,000
Site Reliability Engineering Manager — Production Stability
Site Reliability Engineering Manager — Production Stability

Fairygodboss • Farmers Branch (TX)

Hybrid
USD 150,000 - 210,000
SRE Engineering Manager — 24x7 Production Reliability
SRE Engineering Manager — 24x7 Production Reliability

PNC • Pittsburgh

On-site
USD 115,000 - 150,000
Medical and prescription drug coverage
401(k) with PNC match
Paid time off and holiday benefits
+1
Site Reliability Software Engineering Manager
Site Reliability Software Engineering Manager

Fairygodboss • Cleveland (OH)

On-site
USD 100,000 - 204,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 240,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Hobbsnews • Jersey City (NJ)

On-site
USD 153,000 - 192,000
Benefits eligible
Discretionary incentive plan