Senior SRE: AWS/Kubernetes Reliability + AI-Driven Ops

Fabric Labs, Inc.

New York, Northern (NY, KY)

Hybrid

USD 135,000 - 160,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Medical, dental, vision
Unlimited PTO
401(k) plan
Stock options
Bonuses

Job summary

Fabric Labs, Inc. is seeking a Site Reliability Engineer to evolve and safeguard the infrastructure powering healthcare experiences for millions of patients, while reducing toil for the Fabric tech community.

You will architect and operate our AWS/EKS platform, design for resilience, and enable automation with AI-assisted runbooks. The role emphasizes observability, incident response, and HIPAA-compliant, high-availability systems across teams.

Qualifications

  • 5+ years in SRE or Platform Engineering for large-scale production.
  • Deep AWS (EKS, EC2, RDS, S3) and Kubernetes management.
  • Proficient with Terraform, Datadog, Helm, and GitHub Actions.
  • Strong coding/scripting in Python, Bash, or Go.
  • Experience with agentic workflows or AI-assisted tooling.
  • Rigor-first mindset with HIPAA-compliant, high-availability design.

Responsibilities

  • Architect, deploy, and maintain Kubernetes (EKS) clusters for enterprise availability.
  • Optimize AWS services for performance, cost, and reliability.
  • Build automation and AI-assisted runbooks to reduce manual toil.
  • Advance observability with metrics, traces, and logs to meet SLOs.
  • Lead incident response and blameless postmortems to cut MTTR.
  • Ensure HIPAA compliance and collaborate on cross‑functional reviews.

Skills

AWS
EKS
Kubernetes management
Terraform
Datadog
Helm
GitHub Actions
Python
Bash
Go
Agentic workflows
HIPAA compliance

Tools

Terraform
Datadog
Helm
GitHub Actions

Job description

Fabric Labs, Inc. is seeking a Site Reliability Engineer to evolve and safeguard the infrastructure powering healthcare experiences for millions of patients, while reducing toil for the Fabric tech community.

You will architect and operate our AWS/EKS platform, design for resilience, and enable automation with AI-assisted runbooks. The role emphasizes observability, incident response, and HIPAA-compliant, high-availability systems across teams.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer - Cloud & Kubernetes
Senior Site Reliability Engineer - Cloud & Kubernetes

Fabric • New York (NY)

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Fabric • New York (NY)

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Fabric Labs, Inc. • New York (NY), Northern (KY)

Hybrid
USD 135,000 - 160,000
Medical, dental, vision
Unlimited PTO
401(k) plan
+2
Senior SRE: Cloud, Kubernetes & Automation
Senior SRE: Cloud, Kubernetes & Automation

Socure • Carson City (NV)

On-site
USD 150,000 - 190,000
Senior SRE – AI Cloud Platform, Kubernetes Expert
Senior SRE – AI Cloud Platform, Kubernetes Expert

Socket.dev • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health, dental, vision coverage for in
Wellness and commuter stipends
401k with 2% company match
+1
Senior Site Reliability Engineer — AI Platform Scale
Senior Site Reliability Engineer — AI Platform Scale

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000
Senior SRE - AWS, HIPAA Compliance & Automation
Senior SRE - AWS, HIPAA Compliance & Automation

USA-Medtronic MiniMed, Inc 1017 • Los Angeles (CA)

On-site
USD 124,000 - 212,000
Health insurance
401(k) plan with company match
Employee Stock Purchase Plan
+1
Remote Site Reliability Engineer — AWS, Kubernetes, IaC
Remote Site Reliability Engineer — AWS, Kubernetes, IaC

b.well Connected Health • United States

On-site
USD 153,000 - 190,000
Senior SRE: AI-Driven Infra, Observability, Security
Senior SRE: AI-Driven Infra, Observability, Security

Socket.dev • United States

On-site
USD 150,000 - 180,000
Remote-friendly culture
Hybrid work in Nashville
Medical, dental, vision insurance
+3
Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

JPMorgan Chase & Co. • Jersey City (NJ)

On-site
USD 150,000 - 210,000