Software Engineering Group Manager – Site Reliability

Jobtailor

Alabama

On-site

USD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor is seeking a senior Production Operations leader to own major incident response for high-impact events, guiding troubleshooting across applications, infrastructure, and databases. You will drive root cause analysis, champion systemic fixes, and shape reliability strategies to improve uptime and system health.

This role requires 8+ years of experience and 5+ years in management, with proven leadership in SRE/Infrastructure Engineering.

Qualifications

  • 8+ years of related experience.
  • 5+ years of management experience.
  • Proven leadership in Production Operations, SRE, or Infrastructure Engineering.
  • Deep expertise in incident, problem, and change management.

Responsibilities

  • Lead during moments that matter; own major incident response for high-impact (P1/P2) events, ensuring rapid resolution and clear communication.
  • Provide technical leadership in production support; serve as an escalation point for complex production issues; guide troubleshooting across: applications, infrastructure (Linux/Windows), databases (Oracle, SQL), middleware and integrations; ensure efficient log, metric, and system analysis; oversee batch/ETL monitoring and recovery processes; foster strong collaboration across engineering, infrastructure, and vendor teams.
  • Drive root cause and real fixes; champion a culture of accountability through deep root cause analysis; eliminate repeat issues by driving permanent, systemic solutions and turn data and trends into actionable improvements.
  • Shape reliability at scale; Define and evolve reliability strategy across availability, resiliency, and performance; lead improvements in uptime, MTTR, and overall system health and partner with engineering to embed reliability into system design.
  • Modernize operations; advance observability with best-in-class monitoring, alerting, and event management; leverage tools like Dynatrace, BigPanda, and Logscale to enable proactive detection and drive automation to reduce manual effort and create self-healing systems.
  • Ensure safe and reliable change; oversee change and release governance to enable fast and safe deployments; improve change success rates, reduce production defects, and lead post-release reviews that fuel continuous improvement.
  • Lead a Global 24x7 Operation; manage distributed teams supporting critical systems around the clock; create seamless handoffs and strong operational discipline across regions and elevate team performance, engagement, and growth. Build a trusted, compliant environment; ensure alignment with enterprise governance, audit, and regulatory standards and strengthen risk management, controls, and operational documentation.

Skills

Leadership
Production Ops
SRE
Infrastructure Eng
Incident Management
Problem Management
Change Management
Executive Presence
Communication
OCP/Linux/Windows
MongoDB
Cassandra
Elasticsearch
Redis
MQ/Kafka

Job description

Responsibilities
  • Lead during moments that matter; own major incident response for high-impact (P1/P2) events, ensuring rapid resolution and clear communication.
  • Provide technical leadership in production support; serve as an escalation point for complex production issues; guide troubleshooting across: applications, infrastructure (Linux/Windows), databases (Oracle, SQL), middleware and integrations; ensure efficient log, metric, and system analysis; oversee batch/ETL monitoring and recovery processes; foster strong collaboration across engineering, infrastructure, and vendor teams.
  • Drive root cause and real fixes; champion a culture of accountability through deep root cause analysis; eliminate repeat issues by driving permanent, systemic solutions and turn data and trends into actionable improvements.
  • Shape reliability at scale; Define and evolve reliability strategy across availability, resiliency, and performance; lead improvements in uptime, MTTR, and overall system health and partner with engineering to embed reliability into system design.
  • Modernize operations; advance observability with best-in-class monitoring, alerting, and event management; leverage tools like Dynatrace, BigPanda, and Logscale to enable proactive detection and drive automation to reduce manual effort and create self-healing systems.
  • Ensure safe and reliable change; oversee change and release governance to enable fast and safe deployments; improve change success rates, reduce production defects, and lead post-release reviews that fuel continuous improvement.
  • Lead a Global 24x7 Operation; manage distributed teams supporting critical systems around the clock; create seamless handoffs and strong operational discipline across regions and elevate team performance, engagement, and growth. Build a trusted, compliant environment; ensure alignment with enterprise governance, audit, and regulatory standards and strengthen risk management, controls, and operational documentation.
Requirements
  • 8+ years of related experience and 5+ years of management experience.
  • Proven leadership experience in Production Operations, SRE, or Infrastructure Engineering.
  • Deep expertise in incident, problem, and change management within complex environments.
  • Passion for building reliable, scalable, customer-centric platforms.
  • Track record of improving operational metrics and leading high-performing teams.
  • Strong executive presence and communication skills.
  • Experience with OCP under infrastructure (Linux/Windows, OCP), MongoDB, Cassandra under databases (Oracle, SQL, MongoDB, Cassandra) and working knowledge of Elasticsearch, Redis, MQ and Kafka is a plus.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineering Group Manager – Site Reliability Center
Software Engineering Group Manager – Site Reliability Center

Jobtailor • Alabama

On-site
USD 140,000 - 200,000
Software Engineering Manager – Site Reliability Center
Software Engineering Manager – Site Reliability Center

Jobtailor • Alabama

On-site
USD 120,000 - 160,000
Director, Site Reliability Engineering
Director, Site Reliability Engineering

Jobtailor • California (MO)

On-site
USD 180,000 - 260,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 260,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Luxoft • Buffalo (NY)

On-site
USD 140,000 - 190,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Luxoft • Wilmington (DE)

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

System One • Pittsburgh

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Veloc Inc • Coppell (TX)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobtailor • Arlington (VA)

On-site
USD 140,000 - 200,000
Sr. Director, Site Reliability and Platform Engineering
Sr. Director, Site Reliability and Platform Engineering

Optomi • Tacoma (WA)

On-site
USD 150,000 - 200,000