Site Reliability Engineering Manager

Nationsbenefits

Plantation (FL)

Remote

USD 140,000 - 190,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Unlimited PTO
Fully remote (US-based)
Competitive compensation
Growth opportunity
Cloud-native tech exposure

Job summary

NationsBenefits, a Healthcare FinTech provider delivering supplemental benefits and member engagement solutions, seeks a Manager, Site Reliability Engineering (SRE) to lead the US-based SRE team and drive reliability across production platforms.

This player-coach role blends people leadership with hands-on technical guidance, mentoring engineers, managing incidents, and coordinating with India for global follow-the-sun coverage.

Qualifications

  • 5–8 years of experience in Site Reliability Engineering, DevOps, Production Support, or Platform Engineering.
  • 1–2+ years of experience leading, mentoring, or managing engineers.
  • Leading in a player‑coach model and hands‑on incident management.
  • Production Kubernetes and Docker experience.
  • Datadog or similar observability platforms.
  • Helm, CI/CD pipelines, deployment automation.
  • ITIL and Agile methodologies knowledge.
  • SQL, MySQL, or NoSQL databases.
  • Excellent communication and stakeholder management.
  • Willingness to participate in PagerDuty on‑call and follow‑the‑sun model.

Responsibilities

  • Lead US-based SRE team and mentor engineers.
  • Conduct 1:1s, performance reviews, and career development.
  • Own hiring, onboarding, and retention for the team.
  • Manage on-call incidents and act as escalation point.
  • Drive incident triage, resolution, and RCA.
  • Track operational KPIs, SLAs, and SLOs.
  • Improve reliability, observability, and resilience with Datadog.
  • Embed reliability into the SDLC with DevOps teams.
  • Collaborate with India for follow-the-sun coverage.
  • Maintain documentation and compliance (HIPAA, PCI DSS, SOC 2, ISO 27001, HITRUST).

Skills

Team leadership
Incident management
Datadog
Kubernetes
Docker
SRE tooling
PowerShell
Bash
Python
Java
C#
CI/CD pipelines
Deployment automation
ITIL
Agile methodologies
SQL/NoSQL
Cloud platforms
On-call management
Communication

Tools

Datadog
PagerDuty
Helm
Jenkins

Job description

NationsBenefits is recognized as one of the fastest-growing companies in America and a Healthcare Fintech provider of supplemental benefits, flex cards, and member engagement solutions. We partner with managed care organizations to provide innovative healthcare solutions that drive growth, improve outcomes, reduce costs, and bring value to their members.Through our comprehensive suite of innovative supplemental benefits, fintech payment platforms, and member engagement solutions, we help health plans deliver high-quality benefits to their members that address the social determinants of health and improve member health outcomes and satisfaction.Our compliance-focused infrastructure, proprietary technology systems, and premier service delivery model allow our health plan partners to deliver high-quality, value-based care to millions of members.We offer a fulfilling work environment that attracts top talent and encourages all associates to contribute to delivering premier service to internal and external customers alike. Our goal is to transform the healthcare industry for the better! We provide career advancement opportunities from within the organization across multiple locations in the US, South America, and India.
Location: Remote (US-based candidates only)
Manager, Site Reliability Engineering (SRE)
Position Overview

We are seeking a Manager, Site Reliability Engineering (SRE) to lead our US-based SRE team and drive operational excellence across our production platforms.

This is a player-coach leadership role that combines people management with hands‑on technical leadership. You will mentor and grow a team of Site Reliability Engineers while actively participating in major incident response, reliability initiatives, and operational reviews. The role is a key part of our global follow‑the‑sun support model and requires close collaboration with SRE leadership in India.

Key Responsibilities
Team Leadership & Development
  • Lead, mentor, and develop a US-based team of Site Reliability Engineers.
  • Conduct regular 1:1s, performance reviews, and career development discussions.
  • Own hiring, onboarding, and retention efforts as the team scales.
  • Foster a culture of ownership, blameless postmortems, and continuous improvement.
Operational Excellence & Incident Management
  • Lead day‑to‑day production operations and ensure timely incident triage, resolution, and escalation.
  • Serve as an escalation point and incident commander for major production incidents.
  • Drive problem management and root cause analysis processes.
  • Carry PagerDuty on‑call escalation responsibilities for critical issues.
  • Track and report operational KPIs, SLAs, and SLOs, including availability, MTTR, and incident trends.
Reliability & Automation
  • Improve system reliability, observability, and resilience using Datadog and related tooling.
  • Drive automation, self‑healing capabilities, and runbook maturity.
  • Partner with Development, DevOps, DevSecOps, and Engineering teams to embed reliability into the SDLC.
  • Contribute hands‑on to tooling, automation, and technical reviews as needed.
Collaboration & Global Alignment
  • Coordinate closely with SRE leadership in India to ensure seamless follow‑the‑sun coverage.
  • Represent the US SRE organization in cross‑functional planning and operational reviews.
  • Communicate effectively with both technical and non‑technical stakeholders.
Documentation & Compliance
  • Maintain high-quality documentation for incidents, postmortems, runbooks, and operational procedures.
  • Ensure adherence to healthcare and fintech compliance standards, including HIPAA, PCI DSS, SOC 2, ISO 27001, and HITRUST.
Required Qualifications
  • 5–8 years of experience in Site Reliability Engineering, DevOps, Production Support, or Platform Engineering.
  • 1–2+ years of experience leading, mentoring, or managing engineers.
  • Demonstrated success operating in a player‑coach leadership model.
  • Strong hands‑on experience with production incident management and escalation processes.
  • Proficiency with Datadog or similar observability platforms.
  • Hands‑on experience with Kubernetes and Docker in production environments.
  • Strong scripting or programming skills in PowerShell, Bash, Python, Java, or C#.
  • Experience with Helm, CI/CD pipelines, and deployment automation.
  • Working knowledge of ITIL processes and Agile methodologies.
  • Experience working with SQL, MySQL, or NoSQL databases.
  • Excellent communication and stakeholder management skills.
  • Willingness to participate in PagerDuty on‑call escalation and work within a global follow‑the‑sun operating model.
Preferred Qualifications
  • Experience with cloud platforms such as AWS, Azure, or GCP.
  • Experience building or scaling SRE teams and on‑call programs.
  • Experience defining and managing SLOs, SLIs, and error budgets.
  • Prior experience in the healthcare or fintech industry.
  • Knowledge of security and compliance frameworks relevant to regulated environments.
Why Join NationsBenefits?
  • Competitive compensation and comprehensive benefits.
  • Unlimited PTO.
  • Fully remote work environment (US-based).
  • Opportunity to lead and grow a high‑impact SRE organization.
  • Exposure to modern cloud‑native technologies and large‑scale reliability challenges.
  • Collaborative culture focused on innovation, learning, and continuous improvement.
  • Meaningful work that directly impacts healthcare technology and millions of members.
Ideal Candidate

We are looking for a technically strong SRE leader who enjoys building teams, improving operational maturity, and remaining hands‑on during critical production events. The ideal candidate combines leadership, systems thinking, and automation expertise to help scale reliability practices across a fast‑growing Healthcare FinTech organization.
NationsBenefits is an Equal Opportunity Employer.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineering Manager
Site Reliability Engineering Manager

NationsBenefits, LLC • Plantation (FL)

On-site
USD 140,000 - 180,000
Fully remote
Unlimited PTO
Competitive compensation
+1
Site Reliability Engineer II
Site Reliability Engineer II

Nationsbenefits • Plantation (FL)

Remote
USD 110,000 - 150,000
Unlimited PTO
Competitive compensation
Remote work
Site Reliability Engineer II
Site Reliability Engineer II

NationsBenefits, LLC • United States

On-site
USD 110,000 - 160,000
Unlimited PTO
Competitive benefits
Career growth
Remote SRE Manager: Lead Reliability & Automation
Remote SRE Manager: Lead Reliability & Automation

NationsBenefits, LLC • Plantation (FL)

Remote
USD 140,000 - 180,000
Fully remote
Unlimited PTO
Competitive compensation
+1
Remote SRE Leader: Build Scalable Reliability
Remote SRE Leader: Build Scalable Reliability

Nationsbenefits • Plantation (FL)

Remote
USD 140,000 - 190,000
Unlimited PTO
Fully remote (US-based)
Competitive compensation
+2
Site Reliability Engineering Manager
Site Reliability Engineering Manager

Jobot • Erie

On-site
USD 165,000 - 190,000
Flexible paid time off
Affordable health, dental, and vision insurance
Monthly fitness reimbursement
+4
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Practice by Numbers • United States

On-site
USD 120,000 - 160,000
High ownership and autonomy
Strong engineering culture
Impactful work on healthcare infrastructure
Remote SRE Manager: Reliability & Scale for Fintech Health
Remote SRE Manager: Reliability & Scale for Fintech Health

NationsBenefits • Plantation (FL)

On-site
USD 140,000 - 190,000
Unlimited PTO
Fully remote work (US-based)
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

OutSolve • Mission (KS)

On-site
USD 90,000 - 130,000
100% remote work environment
Competitive compensation
Professional development opportunities
+1
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Practice By Numbers, Inc. • Bellevue (WA), Northern (KY)

On-site
USD 140,000 - 200,000