Site Reliability Engineer — AI-Driven Cloud Resilience

fabrichealth

New York (NY)

On-site

USD 135,000 - 160,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Medical
Dental
Vision
Unlimited PTO
401(k)
Stock options
Bonus

Job summary

FabricHealth seeks a Site Reliability Engineer to architect and safeguard our AWS/EKS platform, delivering resilient, scalable infrastructure for healthcare experiences at scale. You will design clusters, optimize cost, and define guardrails that keep cognitive load low for the engineering community.

Responsibilities span automation and AI-assisted operations, observability, incident management, and HIPAA compliance.

Qualifications

  • 5+ years of experience in SRE or Platform Engineering managing production environments at scale.
  • Expert depth in AWS (EKS, EC2, RDS, S3) and Kubernetes management.
  • Proficiency with Terraform, Datadog, Helm, and GitHub Actions.
  • Strong coding/scripting in Python, Bash, or Go.
  • Preferred experience building agentic workflows or AI-assisted tooling.
  • Rigor-first mindset with HIPAA-compliant, high-availability architecture.

Responsibilities

  • Designing, deploying, and maintaining Kubernetes (EKS) clusters for enterprise-grade availability.
  • Optimizing the footprint of infrastructure objects across core AWS services for performance, cost, and reliability.
  • Evolving a scalable infrastructure management platform with guardrails to maximize engineering agency.
  • Defining golden paths for workload orchestration and integration with dependencies.
  • Providing robust, reusable GitHub Actions components to streamline the software delivery lifecycle.
  • Developing internal tools that replace manual operations with autonomous systems.
  • Exploring and deploying agentic workflows for AI-assisted runbooks.
  • Driving observability by collecting metrics, traces, and logs to meet SLOs.
  • Leading incident response and blameless postmortems to reduce MTTR.
  • Defining and monitoring platform-level SLIs/SLOs for healthcare performance standards.
  • Ensuring HIPAA compliance across all infrastructure.
  • Reviewing architectural decisions and surfacing risks early.
  • Mentoring engineers on reliability and contributing a clinical-safety perspective.

Skills

SRE
Platform Engineering
AWS
Kubernetes
Python
Bash
Go

Tools

Terraform
Datadog
Helm
GitHub Actions

Job description

FabricHealth seeks a Site Reliability Engineer to architect and safeguard our AWS/EKS platform, delivering resilient, scalable infrastructure for healthcare experiences at scale. You will design clusters, optimize cost, and define guardrails that keep cognitive load low for the engineering community.

Responsibilities span automation and AI-assisted operations, observability, incident management, and HIPAA compliance.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer - Cloud & Kubernetes
Senior Site Reliability Engineer - Cloud & Kubernetes

Fabric • New York (NY)

On-site
USD 140,000 - 190,000
Senior SRE: AWS/Kubernetes Reliability + AI-Driven Ops
Senior SRE: AWS/Kubernetes Reliability + AI-Driven Ops

Fabric Labs, Inc. • New York (NY), Northern (KY)

Hybrid
USD 135,000 - 160,000
Medical, dental, vision
Unlimited PTO
401(k) plan
+2
Site Reliability Engineer
Site Reliability Engineer

Fabric • New York (NY)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer — AI Platform Scale
Senior Site Reliability Engineer — AI Platform Scale

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Kontakt.io • United States

On-site
USD 150,000 - 190,000
Senior Site Reliability Engineer - Cloud, Kubernetes & Automation
Senior Site Reliability Engineer - Cloud, Kubernetes & Automation

Socure • Town of Concord (NY)

On-site
USD 140,000 - 210,000
Senior DevOps/SRE for Healthcare AI Platform
Senior DevOps/SRE for Healthcare AI Platform

eSolutionsFirst • Palo Alto (CA), Northern (KY)

Hybrid
USD 170,000 - 220,000
Equity
Senior Site Reliability Engineer - Healthcare Infra Equity
Senior Site Reliability Engineer - Healthcare Infra Equity

Enzo Health • Lehi (UT)

On-site
USD 120,000 - 180,000
Competitive salary
Equity
401k & Insurance
+2
Senior Site Reliability Engineer — Cloud, Resilience & Automation
Senior Site Reliability Engineer — Cloud, Resilience & Automation

Compunnel Inc. • Denton (TX)

Hybrid
USD 120,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

Fabric Labs, Inc. • New York (NY), Northern (KY)

On-site
USD 135,000 - 160,000
Medical, dental, vision
Unlimited PTO
401(k) plan
+2