Software Engineering Manager – Site Reliability Center

Jobtailor

Alabama

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor is seeking an experienced SRE/Production Support leader in Alabama to guide a high‑availability environment. You will mentor SRE engineers, govern incident response, and drive continuous improvement across monitoring, automation, and release processes.

You will coordinate across engineering, infrastructure, and vendor teams, ensuring rapid incident remediation and robust disaster recovery capabilities. Remote work is not indicated; on‑site leadership is expected.

Qualifications

  • 5+ years in SRE/Production Support/DevOps with leadership
  • Experience in high‑availability, enterprise environments
  • Strong incident, problem and change management knowledge
  • Hands‑on with monitoring tools, cloud/infrastructure and automation
  • Excellent communication for high‑pressure incident situations
  • Familiarity with databases (Oracle/SQL/MongoDB/Cassandra) and MQ/Kafka is a plus

Responsibilities

  • Lead SRE and related teams; coach and develop engineers; align technology with business goals
  • Handle major incident management end-to-end; guide triage, diagnostics and remediation
  • Provide technical leadership in production support across applications and infra
  • Drive root‑cause analysis and permanent fixes; create runbooks and knowledge articles
  • Oversee change management and release execution; CAB participation and post‑implementation reviews
  • Advance monitoring, alerting and observability; optimize tooling like Dynatrace, BigPanda, Logscale
  • Champion resiliency, scalability and availability; oversee DR and failover testing
  • Guide capacity planning and performance tuning; work with development teams
  • Lead 24x7 production support and on-call rotations; manage incident bridges and escalations
  • Drive automation to reduce toil; implement standardized runbooks and automation
  • Ensure governance, risk and compliance; support audits and risk controls

Skills

Site Reliability Engineering
Production Support
DevOps
Team Leadership
Incident Management
Monitoring & Observability
Automation
Capacity Planning
Change Management
Disaster Recovery

Tools

Dynatrace
BigPanda
Logscale

Job description

Responsibilities
  • Manage SRE and related teams; lead, coach, and develop a team of SRE engineers; set clear goals, drive accountability, and foster a culture of ownership and excellence; partner with cross‑functional stakeholders to align technology and business objectives; support talent development, performance management, and succession planning; encourage innovation, continuous learning, and DevOps/SRE best practices.
  • Lead incident management & remediation; manage and actively participate in end‑to‑end incident response for major (P1/P2) incidents; guide real‑time triage, diagnostics, and troubleshooting across application, infrastructure, and network layers; ensure rapid execution of remediation actions and service restoration; provide clear, timely communication to stakeholders during incidents; oversee post‑incident analysis, reporting, and documentation to drive improvements.
  • Provide technical leadership in production support; serve as an escalation point for complex production issues; guide troubleshooting across applications, infrastructure (Linux/Windows), databases (Oracle, SQL), middleware and integrations; ensure efficient log, metric, and system analysis; oversee batch/ETL monitoring and recovery processes; foster strong collaboration across engineering, infrastructure, and vendor teams.
  • Drive problem management & root‑cause resolution; lead root‑cause analysis (RCA) efforts for major and recurring incidents; ensure ownership and resolution of problem records; drive permanent fixes and systemic improvements to eliminate repeat issues, identify trends and patterns to reduce risk and improve stability; partner with engineering teams to resolve code defects and system gaps and promote knowledge sharing via runbooks, knowledge articles, and error catalogs.
  • Oversee change management & release execution; ensure safe and compliant execution of production changes and releases; validate change readiness, testing, rollback strategies, and risk assessments; represent the team in CAB reviews, providing technical risk evaluation; oversee post‑implementation reviews (CPIR) and ensure follow‑through and drive improvements in change success rate and reduction in production defects.
  • Advance monitoring, alerting & observability; lead efforts to build and optimize monitoring, dashboards, and alerting frameworks, champion use of tools such as Dynatrace, BigPanda, Logscale, and enterprise platforms, improve signal‑to‑noise ratio through alert tuning; enable proactive issue detection before customer impact; strengthen event management and observability practices.
  • Champion resiliency, stability & availability; lead efforts to ensure high availability of critical systems; oversee disaster recovery, failover, and continuity testing; identify and eliminate single points of failure and drive improvements in MTTR, uptime, and service reliability.
  • Enable scalability & performance optimization; guide capacity planning and performance tuning strategies; ensure systems scale effectively under peak demand; partner with development teams for performance‑driven design improvements; optimize system configurations to improve efficiency and throughput.
  • Lead a 24x7 production support model; manage team participation in a 24x7 on‑call rotation; oversee engagement in incident bridges, war rooms, and escalations; support pod‑based operating models aligned to key applications; ensure seamless handoffs and global support continuity.
  • Drive Automation & Operational Efficiency; identify and prioritize opportunities to reduce manual effort through automation; implement automation across incident remediation, monitoring and alerting, deployment and validation, promote standardized runbooks and automation frameworks and improve operational metrics and reduce toil.
  • Ensure Governance, Risk & Compliance; maintain adherence to enterprise policies and regulatory standards; support audits, vulnerability remediation, and risk controls; ensure accurate documentation and operational procedures and champion security, access management, and data governance practices.
Requirements
  • 5 + years of related experience and 3+ years of management experience.
  • Strong experience in Site Reliability Engineering, Production Support, or DevOps.
  • Proven ability to lead teams in high‑availability, enterprise environments.
  • Deep understanding of incident, problem, and change management frameworks.
  • Hands‑on knowledge of monitoring tools, cloud/infrastructure platforms, and automation.
  • Experience improving system reliability, observability, and operational maturity.
  • Strong communication skills with the ability to lead during high‑pressure situations.
  • Experience with OCP under infrastructure (Linux/Windows, OCP), MongoDB, Cassandra under databases (Oracle, SQL, MongoDB, Cassandra) and working knowledge of Elasticsearch, Redis, MQ and Kafka is a plus.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineering Group Manager – Site Reliability Center
Software Engineering Group Manager – Site Reliability Center

Jobtailor • Alabama

On-site
USD 140,000 - 200,000
Software Engineering Group Manager – Site Reliability
Software Engineering Group Manager – Site Reliability

Jobtailor • Alabama

On-site
USD 120,000 - 180,000
Director, Site Reliability Engineering
Director, Site Reliability Engineering

Jobtailor • California (MO)

On-site
USD 180,000 - 260,000
Site Reliability Engineer – Lead
Site Reliability Engineer – Lead

Jobtailor • Arizona

On-site
USD 140,000 - 230,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Luxoft • Buffalo (NY)

On-site
USD 140,000 - 190,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Luxoft • Wilmington (DE)

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

System One • Pittsburgh

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 260,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Luxoft • United States

On-site
USD 140,000 - 180,000