Site Reliability Engineer - Full Time

Hader Solutions

O’Fallon (MO)

On-site

USD 120,000 - 150,000

Full time

14 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Hader Solutions is seeking a Manager of Site Reliability Engineering to lead a team responsible for secure, scalable, and highly available platforms. You will design and implement SRE best practices for reliability and performance while developing automation for deployment and operations.

You will enhance observability across complex cloud-native systems and collaborate with development and security teams to improve incident response and overall system reliability.

Qualifications

  • Experience leading SRE or similar teams managing cloud-native environments.
  • Strong understanding of cloud platforms, automation, and system reliability principles.
  • Excellent collaboration, communication, and leadership skills.
  • Proficiency in implementing monitoring, alerting, and incident management strategies.
  • Ability to drive technological innovation and foster a culture of continuous improvement.

Responsibilities

  • Lead a team to ensure secure, scalable, and highly available platforms.
  • Design and implement SRE best practices for reliability and performance.
  • Develop automation for deployment and operational processes.
  • Enhance observability across complex cloud-native systems.
  • Collaborate with development and security teams to improve system reliability and incident response.
  • Mentor engineers, foster continuous improvement, and promote a culture of innovation and learning.

Skills

SRE leadership
Cloud platforms
Automation
Observability
Incident management

Job description

Role: Manager, Site Reliability Engineering

Responsibilities
  • Lead a team to ensure secure, scalable, and highly available platforms.
  • Design and implement SRE best practices for reliability and performance.
  • Develop automation for deployment and operational processes.
  • Enhance observability across complex cloud-native systems.
  • Collaborate with development and security teams to improve system reliability and incident response.
  • Mentor engineers, foster continuous improvement, and promote a culture of innovation and learning.
Qualifications & Skills
  • Experience leading SRE or similar teams managing cloud-native environments.
  • Strong understanding of cloud platforms, automation, and system reliability principles.
  • Excellent collaboration, communication, and leadership skills.
  • Proficiency in implementing monitoring, alerting, and incident management strategies.
  • Ability to drive technological innovation and foster a culture of continuous improvement.
Get your free, confidential resume review.

or drag and drop your file here.