Software Engineering Group Manager – Site Reliability Center

Jobtailor

Alabama

On-site

USD 140,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor is seeking a senior Production Operations leader to own incident response, drive reliability across large-scale systems, and guide cross‑functional teams. You will shape observability, governance, and continuous improvement across availability, resiliency, and performance.

Strong executive communication and proven leadership are essential. You will manage distributed teams across regions, oversee change governance, and partner with engineering to embed reliability into system design,

Qualifications

  • 8+ years of related experience with 5+ years in a management role.
  • Proven leadership in Production Operations, SRE, or Infrastructure Engineering.
  • Expertise in incident, problem, and change management within complex environments.
  • Track record of improving operational metrics and leading high-performing teams.

Responsibilities

  • Lead during critical incidents (P1/P2) with rapid resolution and clear communication.
  • Provide technical leadership in production support; serve as escalation point for complex issues across apps, infra, databases, middleware, and integrations.
  • Drive root cause analysis and implement permanent systemic fixes to reduce repeats.

Skills

Incident management
Problem management
Change management
Production support
Root cause analysis
Reliability engineering
Monitoring
Automation
System design
Leadership
Executive communication
Troubleshooting

Tools

Dynatrace
BigPanda
Logscale
MongoDB
Cassandra
Oracle
SQL
Elasticsearch
Redis
MQ
Kafka

Job description

Responsibilities
  • Lead during moments that matter; own major incident response for high-impact (P1/P2) events, ensuring rapid resolution and clear communication.
  • Provide technical leadership in production support; serve as an escalation point for complex production issues; guide troubleshooting across: applications, infrastructure (Linux/Windows), databases (Oracle, SQL), middleware and integrations.
  • Drive root cause and real fixes; champion a culture of accountability through deep root cause analysis; eliminate repeat issues by driving permanent, systemic solutions and turn data and trends into actionable improvements.
  • Shape reliability at scale; define and evolve reliability strategy across availability, resiliency, and performance; lead improvements in uptime, MTTR, and overall system health and partner with engineering to embed reliability into system design.
  • Modernize operations; advance observability with best-in-class monitoring, alerting, and event management; leverage tools like Dynatrace, BigPanda, and Logscale to enable proactive detection and drive automation to reduce manual effort and create self‑healing systems.
  • Ensure safe and reliable change; oversee change and release governance to enable fast and safe deployments; improve change success rates, reduce production defects, and lead post‑release reviews that fuel continuous improvement.
  • Lead a Global 24x7 Operation; manage distributed teams supporting critical systems around the clock; create seamless handoffs and strong operational discipline across regions and elevate team performance, engagement, and growth.
Requirements
  • 8+ years of related experience and 5+ years of management experience.
  • Proven leadership experience in Production Operations, SRE, or Infrastructure Engineering.
  • Deep expertise in incident, problem, and change management within complex environments.
  • Passion for building reliable, scalable, customer‑centric platforms.
  • Track record of improving operational metrics and leading high‑performing teams.
  • Strong executive presence and communication skills.
  • Experience with OCP under infrastructure (Linux/Windows, OCP), MongoDB, Cassandra under databases (Oracle, SQL, MongoDB, Cassandra) and working knowledge of Elasticsearch, Redis, MQ and Kafka is a plus.
Hard Skills
  • incident management
  • problem management
  • change management
  • production support
  • root cause analysis
  • reliability engineering
  • monitoring
  • alerting
  • automation
  • system design
Soft Skills
  • leadership
  • communication
  • operational discipline
  • team performance
  • engagement
  • growth
  • accountability
  • customer‑centric focus
  • executive presence
  • troubleshooting
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineering Group Manager – Site Reliability
Software Engineering Group Manager – Site Reliability

Jobtailor • Alabama

On-site
USD 120,000 - 180,000
Software Engineering Manager – Site Reliability Center
Software Engineering Manager – Site Reliability Center

Jobtailor • Alabama

On-site
USD 120,000 - 160,000
Director, Site Reliability Engineering
Director, Site Reliability Engineering

Jobtailor • California (MO)

On-site
USD 180,000 - 260,000
Site Reliability Engineer
Site Reliability Engineer

System One • Pittsburgh

On-site
USD 140,000 - 190,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Luxoft • Buffalo (NY)

On-site
USD 140,000 - 190,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Luxoft • Wilmington (DE)

On-site
USD 140,000 - 190,000
Site Reliability Engineer – Lead
Site Reliability Engineer – Lead

Jobtailor • Arizona

On-site
USD 140,000 - 230,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 260,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Luxoft • United States

On-site
USD 140,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Veloc Inc • Coppell (TX)

On-site
USD 140,000 - 190,000