Senior Site Reliability Engineer | Cloud & AI/ML Ops

Artha Nexgen

Austin, Northern (TX, KY)

Hybrid

USD 241,000 - 270,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity incentive
Flexible PTO
Health/Dental/Vision plans
401(k)

Job summary

Garner Health is seeking a Staff Site Reliability Engineer to own the reliability strategy for its cloud infrastructure and AI/ML workloads. You will define SLOs, drive incident response, and build automated, scalable observability across AWS, Kubernetes, and supporting tools.

You will mentor engineers, shape technical direction, and ensure security and HIPAA compliance across the platform, delivering high uptime and cost-efficient operations in a remote-friendly role.

Qualifications

  • 7+ years of hands-on experience operating production cloud infrastructure at scale in an SRE, DevOps, or platform engineering role.
  • Deep expertise with Kubernetes and Terraform in a cloud-first environment (AWS preferred), with a track record of architecting reliability for systems at scale
  • Experience designing an organization's reliability practice (SLO frameworks, observability platforms, incident response programs, and blameless post-incident reviews) and the judgment to know when to build vs. buy
  • Strong Python or Go skills applied to infrastructure automation (Kubernetes API experience a plus)
  • Track record driving cloud cost-efficiency and performance optimization across compute, storage, and networking
  • Mentorship experience and the ability to set technical direction as the senior reliability voice
  • Excellent communication skills—able to make complex reliability concepts land with both technical and non-technical stakeholders
  • Fluency with AI tools (e.g., Claude) applied to real engineering and operations workflows, or strong motivation to build it fast
  • Experience supporting AI/ML or data-intensive workloads in production is a plus
  • Experience operating in a security-conscious or regulated environment (HIPAA, SOC 2) is a plus
  • A desire to be a part of a high-performing, mission-driven team that operates with intense urgency, a strong sense of individual accountability, and a commitment to authentic feedback

Responsibilities

  • Own the Reliability Strategy: Architect and own the end-to-end reliability, performance, and resilience of Garner's cloud environments (AWS, Kubernetes), including those powering AI/ML workloads; design the SLO framework our critical services are measured against and lead the technical decision-making that keeps us ahead of scale
  • Lead the Incident Response Program: Set the standard for how Garner responds to incidents: serve in and level up the on-call rotation, lead response for the most complex escalations, drive deep-dive root cause analysis, and build the review culture that sees corrective actions through to resolution
  • Own Observability: Architect the monitoring, alerting, and observability platform that lets us detect and resolve issues before users feel them, and that lets stakeholders quickly identify the health of every team's products
  • Translate Ambiguity: Take high-level, ambiguous scaling and reliability requirements and transform them into well-defined, automated, and composable infrastructure-as-code deliverables (Terraform); proactively identify and implement cost-efficiency and performance gains across the stack to maximize cloud ROI
  • Course-Correct Technical Direction: Proactively identify when infrastructure workflows or technical paths are inefficient or fragile and redirect efforts to ensure the highest ROI for the engineering function, paying down impactful tech debt and using AI tools and automation to convert repetitive operational work into hands-free, monitored processes
  • Multiply the Engineering Team: Build and own the deployment and observability standards that empower the broader engineering team to ship AI features faster and more reliably; mentor engineers across the organization and provide high-quality feedback that raises the bar for operational rigor and discipline
  • Uphold Security & Compliance: Ensure our infrastructure and operations meet Garner's security and HIPAA compliance obligations, and lead rigorous review of infrastructure changes so platform work meets the same standards as our customer-facing products

Skills

Kubernetes
Terraform
Python
Go
Cloud infrastructure
Incident response
Mentorship

Tools

AWS
Datadog
GitLab
Istio
PostgreSQL
NATS

Job description

Garner Health is seeking a Staff Site Reliability Engineer to own the reliability strategy for its cloud infrastructure and AI/ML workloads. You will define SLOs, drive incident response, and build automated, scalable observability across AWS, Kubernetes, and supporting tools.

You will mentor engineers, shape technical direction, and ensure security and HIPAA compliance across the platform, delivering high uptime and cost-efficient operations in a remote-friendly role.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer: Scalable Health Platform
Senior Site Reliability Engineer: Scalable Health Platform

Socket.dev • Williamsburg (VA)

Hybrid
USD 180,000 - 240,000
Competitive compensation including equ
Health coverage
401(k) retirement plan
+2
Senior Site Reliability Engineer — AI Platform Scale
Senior Site Reliability Engineer — AI Platform Scale

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000
Senior DevOps Engineer & SRE - Healthcare AI (Hybrid)
Senior DevOps Engineer & SRE - Healthcare AI (Hybrid)

Qualified Health • Palo Alto (CA)

Hybrid
USD 170,000 - 220,000
Equity
Medical insurance
Dental insurance
+3
Staff Site Reliability Engineer Visa Austin, Texas, US
Staff Site Reliability Engineer Visa Austin, Texas, US

Artha Nexgen • Austin (TX), Northern (KY)

Hybrid
USD 241,000 - 270,000
Equity incentive
Flexible PTO
Health/Dental/Vision plans
+1
Senior Staff Site Reliability Engineer
Senior Staff Site Reliability Engineer

Pivotal Health • Santa Monica (CA)

Hybrid
USD 180,000 - 250,000
Competitive compensation
Health, dental, and vision coverage
401(k) retirement plan
+2
Senior DevOps Engineer for Healthcare AI Platform SRE
Senior DevOps Engineer for Healthcare AI Platform SRE

Transformcap • Palo Alto (CA)

Hybrid
USD 170,000 - 220,000
Equity
Medical insurance
Flexible hours
+1
Senior Cloud Reliability Engineer — AWS & Automation
Senior Cloud Reliability Engineer — AWS & Automation

LeanData Inc. • Santa Clara (CA), Northern (KY)

Hybrid
USD 140,000 - 180,000
Senior SRE – AI Cloud Platform, Kubernetes Expert
Senior SRE – AI Cloud Platform, Kubernetes Expert

Socket.dev • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health, dental, vision coverage for in
Wellness and commuter stipends
401k with 2% company match
+1
Remote AI Enablement Site Reliability Engineer
Remote AI Enablement Site Reliability Engineer

Health Catalyst • United States

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer — Kubernetes & AI-Driven Ops
Senior Site Reliability Engineer — Kubernetes & AI-Driven Ops

Fal • San Francisco (CA)

On-site
USD 180,000 - 250,000
Relocation assistance to San Francisco
Visa sponsorship
Competitive salary and equity
+1