Site Reliability Engineer II

Nationsbenefits

Plantation (FL)

Remote

USD 110,000 - 150,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Unlimited PTO
Competitive compensation
Remote work

Job summary

NationsBenefits is seeking a Site Reliability Engineer II (SRE) to join our growing SRE team. In this role, you will help ensure the availability, reliability, and performance of our production platforms by monitoring systems, responding to incidents, troubleshooting infrastructure issues, and driving automation initiatives.

You will collaborate with Development, DevSecOps, and Engineering teams to maintain cloud-native applications while supporting healthcare and fintech services.

Qualifications

  • 3–5 years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or Production Support.
  • Hands-on experience with production incident response, troubleshooting, and escalation.
  • Experience with Datadog or similar monitoring and observability platforms.
  • Strong experience with Kubernetes, including monitoring, troubleshooting, and workload management.
  • Experience with Docker or other container technologies.
  • Working knowledge of SQL, MySQL, or NoSQL databases.
  • Ability to work effectively in high-volume, mission-critical production environments.
  • Strong analytical, troubleshooting, and problem-solving skills.
  • Excellent written and verbal communication skills.
  • Willingness to work weekday shifts as part of a global Follow-the-Sun support model.

Responsibilities

  • Serve as first responder for production incidents, triage and resolve.
  • Monitor application health, infrastructure performance, and system availability; optimize dashboards.
  • Troubleshoot Kubernetes pods and deployments; support containerized applications in Kubernetes and Docker.
  • Participate in Follow-the-Sun production support and on-call rotation.
  • Develop automation scripts and operational tools (Python, PowerShell, Bash, C#, Java).
  • Collaborate with Software Engineers, DevSecOps, Infrastructure and Platform teams.
  • Maintain incident documentation and post-incident reviews; ensure HIPAA/PCI SOC 2 ISO 27001 HITRUST alignment.

Skills

SRE Experience
Incident Response
Kubernetes
Docker
Datadog
CI/CD
Cloud Knowledge

Tools

Kubernetes
Docker
CI/CD Tools

Job description

NationsBenefits is recognized as one of the fastest-growing companies in America and a Healthcare Fintech provider of supplemental benefits, flex cards, and member engagement solutions. We partner with managed care organizations to provide innovative healthcare solutions that drive growth, improve outcomes, reduce costs, and bring value to their members.

Through our comprehensive suite of innovative supplemental benefits, fintech payment platforms, and member engagement solutions, we help health plans deliver high-quality benefits to their members that address the social determinants of health and improve member health outcomes and satisfaction.

Our compliance-focused infrastructure, proprietary technology systems, and premier service delivery model allow our health plan partners to deliver high-quality, value-based care to millions of members.

We offer a fulfilling work environment that attracts top talent and encourages all associates to contribute to delivering premier service to internal and external customers alike. Our goal is to transform the healthcare industry for the better! We provide career advancement opportunities from within the organization across multiple locations in the US, South America, and India.

Location: Remote (US-Based Candidates Only)
Site Reliability Engineer II (SRE)

Position Overview

We are seeking a Site Reliability Engineer II (SRE) to join our growing Site Reliability Engineering team. In this role, you will help ensure the availability, reliability, and performance of our production platforms by monitoring systems, responding to incidents, troubleshooting infrastructure issues, and driving automation initiatives.

You will collaborate closely with Development, DevSecOps, and Engineering teams to maintain highly available cloud-native applications while supporting mission-critical healthcare and fintech services.

This position is ideal for someone who enjoys solving production challenges, improving operational efficiency, and working in a fast-paced environment.

Key Responsibilities
Incident Management
  • Serve as the first responder for production incidents by identifying, triaging, and resolving issues.
  • Monitor and respond to alerts generated by Datadog and other monitoring platforms.
  • Perform initial root cause analysis and elevate incidents according to defined SLAs.
  • Communicate incident status and resolution updates to internal stakeholders.
  • Partner with senior engineers to resolve complex production issues.
Monitoring & Platform Reliability
  • Continuously monitor application health, infrastructure performance, and system availability.
  • Configure and optimize monitoring dashboards and alert thresholds.
  • Troubleshoot Kubernetes environments, including pod failures, deployment rollbacks, and log analysis.
  • Support containerized applications running in Kubernetes and Docker environments.
Production Support
  • Participate in a weekday "Follow-the-Sun" production support model with global engineering teams.
  • Participate in an on-call rotation for critical production systems as needed.
  • Help maintain high availability and system uptime.
Automation & Continuous Improvement
  • Develop automation scripts and operational tools using one or more of the following:
    • Python
    • PowerShell
    • Bash
    • C#
    • Java
  • Support CI/CD pipeline monitoring and deployment reliability.
  • Contribute to self-healing solutions and automation initiatives to reduce manual operational tasks.
Collaboration
  • Work closely with Software Engineers, DevSecOps, Infrastructure, and Platform teams.
  • Recommend improvements to monitoring, tooling, and operational processes.
  • Collaborate effectively with globally distributed engineering teams.
Documentation & Compliance
  • Maintain accurate documentation for incidents, troubleshooting procedures, and post-incident reviews.
  • Ensure operational processes align with industry security and compliance standards, including HIPAA, PCI DSS, SOC 2, ISO 27001, and HITRUST.
Required Qualifications
  • 3–5 years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or Production Support.
  • Hands-on experience with production incident response, troubleshooting, and escalation.
  • Experience with Datadog or similar monitoring and observability platforms.
  • Strong experience with Kubernetes, including monitoring, troubleshooting, and workload management.
  • Experience with Docker or other container technologies.
  • Working knowledge of SQL, MySQL, or NoSQL databases.
  • Ability to work effectively in high-volume, mission-critical production environments.
  • Strong analytical, troubleshooting, and problem-solving skills.
  • Excellent written and verbal communication skills.
  • Willingness to work weekday shifts as part of a global Follow-the-Sun support model.
Preferred Qualifications
  • Experience with cloud platforms such as Microsoft Azure, AWS, or Google Cloud Platform (GCP).
  • Familiarity with CI/CD pipelines and deployment automation.
  • Experience with Helm Charts and Kubernetes deployments.
  • Knowledge of ITIL principles and Agile methodologies.
  • Experience supporting regulated environments such as healthcare or fintech.
  • Scripting or programming experience in Python, PowerShell, Bash, Java, or C#.
Why Join NationsBenefits?
  • Work on technology that positively impacts millions of healthcare members.
  • Join a collaborative, innovative, and supportive engineering culture.
  • Exposure to modern cloud-native technologies and enterprise-scale infrastructure.
  • Competitive compensation and comprehensive benefits.
  • Unlimited Paid Time Off (PTO).
  • Opportunities for career growth and professional development.
  • Work with talented global engineering teams on challenging, high-impact projects.
  • Maintain a healthy work-life balance while contributing to mission-critical platforms.

NationsBenefits is an Equal Opportunity Employer.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer II
Site Reliability Engineer II

NationsBenefits, LLC • United States

On-site
USD 110,000 - 160,000
Unlimited PTO
Competitive benefits
Career growth
Site Reliability Engineering Manager
Site Reliability Engineering Manager

Nationsbenefits • Plantation (FL)

Remote
USD 140,000 - 190,000
Unlimited PTO
Fully remote (US-based)
Competitive compensation
+2
Site Reliability Engineering Manager
Site Reliability Engineering Manager

NationsBenefits, LLC • Plantation (FL)

On-site
USD 140,000 - 180,000
Fully remote
Unlimited PTO
Competitive compensation
+1
Remote SRE II — Cloud Reliability & Automation
Remote SRE II — Cloud Reliability & Automation

Nationsbenefits • Plantation (FL)

Remote
USD 110,000 - 150,000
Unlimited PTO
Competitive compensation
Remote work
Remote SRE II - Cloud-Native Healthcare
Remote SRE II - Cloud-Native Healthcare

NationsBenefits • Plantation (FL)

Remote
USD 110,000 - 150,000
Competitive compensation
Unlimited PTO
Career development opportunities
+2
Senior Site Reliability Engineer II
Senior Site Reliability Engineer II

LexisNexis Risk Solutions • San Jose (CA), Northern (KY)

Hybrid
USD 105,000 - 175,000
401(k) with match
Wellbeing programs
Life Insurance
+1
Remote SRE II: Cloud-Native Reliability & Automation
Remote SRE II: Cloud-Native Reliability & Automation

NationsBenefits, LLC • Plantation (FL)

On-site
USD 110,000 - 160,000
Unlimited PTO
Competitive compensation & benefits
Career growth opportunities
+1
Remote SRE II: Kubernetes & Automation
Remote SRE II: Kubernetes & Automation

NationsBenefits, LLC • United States

On-site
USD 110,000 - 160,000
Unlimited PTO
Competitive benefits
Career growth
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Storm2 • Scottsdale (AZ)

On-site
USD 140,000 - 150,000
Competitive healthcare, dental, and vision coverage
401(k) with company match
Generous PTO and paid holidays
+1
Site Reliability Engineer (SRE) – II
Site Reliability Engineer (SRE) – II

Huntington National Bank • Columbus (OH)

On-site
USD 90,000 - 120,000