SRE Manager: Lead Cloud Reliability & Automation

Hader Solutions

O’Fallon (MO)

On-site

USD 120,000 - 150,000

Full time

15 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Hader Solutions is seeking a Manager of Site Reliability Engineering to lead a team responsible for secure, scalable, and highly available platforms. You will design and implement SRE best practices for reliability and performance while developing automation for deployment and operations.

You will enhance observability across complex cloud-native systems and collaborate with development and security teams to improve incident response and overall system reliability.

Qualifications

  • Experience leading SRE or similar teams managing cloud-native environments.
  • Strong understanding of cloud platforms, automation, and system reliability principles.
  • Excellent collaboration, communication, and leadership skills.
  • Proficiency in implementing monitoring, alerting, and incident management strategies.
  • Ability to drive technological innovation and foster a culture of continuous improvement.

Responsibilities

  • Lead a team to ensure secure, scalable, and highly available platforms.
  • Design and implement SRE best practices for reliability and performance.
  • Develop automation for deployment and operational processes.
  • Enhance observability across complex cloud-native systems.
  • Collaborate with development and security teams to improve system reliability and incident response.
  • Mentor engineers, foster continuous improvement, and promote a culture of innovation and learning.

Skills

SRE leadership
Cloud platforms
Automation
Observability
Incident management

Job description

Hader Solutions is seeking a Manager of Site Reliability Engineering to lead a team responsible for secure, scalable, and highly available platforms. You will design and implement SRE best practices for reliability and performance while developing automation for deployment and operations.

You will enhance observability across complex cloud-native systems and collaborate with development and security teams to improve incident response and overall system reliability.

Get your free, confidential resume review.

or drag and drop your file here.