Site Reliability Engineer

Enzo Health

Lehi (UT)

On-site

USD 120,000 - 180,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salary
Equity
401k & Insurance
High ownership
Direct collaboration with founders

Job summary

Enzo Health is hiring a Senior Site Reliability Engineer to join our Security and Site Reliability team. You will focus on a stable, scalable AWS and Kubernetes platform, reliable Postgres operations, and safer delivery. You will mature observability, on‑call, and incident response practices across engineering.

This is a hands‑on role. You will diagnose production problems, improve infrastructure, write automation, and help product engineers operate their services with confidence.

Qualifications

  • At least 5 years of experience in site reliability, platform, infrastructure, or production engineering.
  • Strong hands‑on experience with AWS and production Kubernetes.
  • Strong experience with Terraform and infrastructure as code.
  • Experience operating Postgres in production, including backup and restore, performance, and safe migrations.
  • Experience building and operating CI/CD and deployment systems.
  • Experience with modern observability tools across metrics, logs, and traces.
  • Strong programming or scripting skills for automation and operational tools.
  • A proven track record diagnosing production failures and leading incidents through recovery.
  • Ability to write clear automation, runbooks, and technical standards.
  • Ability to work independently and partner well with application engineers.

Responsibilities

  • Operate and improve our AWS and Kubernetes environments.
  • Own infrastructure changes through Terraform, including modules, state, review standards, and drift control.
  • Improve environment management, cluster practices, and deployment reliability.
  • Make CI/CD and releases safer through clear checks, reliable rollbacks, environment consistency, and release observability.
  • Improve capacity planning, resource controls, resilience, and cost visibility.
  • Build automation that removes repetitive operational work and reduces avoidable failures.
  • Own the operational health of Postgres in production.
  • Verify backups and regularly test restore procedures against agreed recovery targets.
  • Improve monitoring for connections, storage, slow queries, locks, and other important failure signals.
  • Guide safe schema migrations, database access, and production change procedures.
  • Find performance and capacity risks before they affect customers.
  • Maintain clear database runbooks for common failures and emergency work.
  • Mature dashboards, logs, traces, and alert routing.
  • Validate service‑level indicators and objectives; close coverage gaps.
  • Participate in on‑call rotation and write runbooks for engineers.
  • Lead or support incident response and blameless post‑incident reviews.
  • Use reliability data to set priorities and measure improvement.

Skills

AWS
Kubernetes
Terraform
Postgres
CI/CD
Observability
Automation
Scripting
Incident response

Tools

GitHub Actions

Job description

Enzo Health is a healthcare technology company transforming home health operations through purpose-built artificial intelligence. We deliver a secure, HIPAA-compliant AI platform that unifies intake, clinical documentation, coding, and quality assurance—enabling agencies to reclaim time and revenue while elevating patient care.

Enzo addresses the critical challenges facing home health agencies today: rising operational costs, clinician burnout, shrinking reimbursement margins, and increasing compliance demands. Our integrated AI solution automates documentation workflows from referral to final QA, allowing clinical staff to focus on delivering exceptional patient care.

Our Solutions
  • Enzo Intake: Delivers intake decisions in seconds, automatically extracting key data from referrals to increase admissions and reduce processing delays.
  • Enzo Scribe: Auto-generates OASIS documentation, clinical narratives, and care plans, reducing documentation time by up to 75%.
  • Enzo QA: Ensures documentation meets the highest clinical standards with approximately 95% coding accuracy, reducing compliance risk while increasing reimbursement by an average of $185 per episode.
Our Impact

Trusted by top-performing home health agencies nationwide, Enzo delivers measurable results: documentation time under 25 minutes per visit, referral intake under 5 minutes, and 30-50% savings per episode of care. These efficiencies effectively double staff capacity while maintaining exceptional quality and compliance.

As reimbursement pressures intensify, Enzo Health empowers agencies to navigate cost-cutting measures without compromising care quality, positioning AI as the essential strategy for sustainable growth in home health.

About the role

Enzo builds software for home health care teams. Our systems support important clinical and business workflows, so they must be available, secure, and easy to operate.

We are hiring a Senior Site Reliability Engineer to join our Security and Site Reliability team. You will focus on a stable, scalable AWS and Kubernetes platform, reliable Postgres operations, and safer delivery. You will also help mature observability, on‑call, and incident response practices across engineering.

This is a hands‑on role. You will diagnose production problems, improve infrastructure, write automation, and help product engineers operate their services with confidence. You will make our systems safer without slowing product delivery.

What you'll do

Strengthen the production platform

  • Operate and improve our AWS and Kubernetes environments.
  • Own infrastructure changes through Terraform, including modules, state, review standards, and drift control.
  • Improve environment management, cluster practices, and deployment reliability.
  • Make CI/CD and releases safer through clear checks, reliable rollbacks, environment consistency, and release observability.
  • -mprove capacity planning, resource controls, resilience, and cost visibility.
  • Build automation that removes repetitive operational work and reduces avoidable failures.
  • Own the operational health of Postgres in production.
  • Verify backups and regularly test restore procedures against agreed recovery targets.
  • Improve monitoring for connections, storage, slow queries, locks, and other important failure signals.
  • Guide safe schema migrations, database access, and production change procedures.
  • Find performance and capacity risks before they affect customers.
  • Maintain clear database runbooks for common failures and emergency work.

Mature production operations

  • Improve dashboards, monitors, logs, traces, and alert routing.
  • Validate existing service‑level indicators and objectives, then close important coverage gaps.
  • Participate in and improve the company‑wide on‑call rotation.
  • Write practical runbooks and make escalation paths clear for engineers across the company.
  • Lead or support incident response, recovery, and blameless post‑incident reviews.
  • Use incident and reliability data to set priorities and measure improvement.

Work across Security and engineering

  • Work with Security to keep infrastructure changes consistent with SOC 2 and health care security requirements.
  • Apply least‑privilege access, secure defaults, secrets management, encryption, and auditable change practices.
  • Support vulnerability remediation, disaster‑recovery exercises, and secure production access.
  • Help product teams include reliability and operational risk in technical decisions.

What success looks like in the first six months

  • Kubernetes, AWS, and Terraform have clear operating standards and fewer manual failure points.
  • Deployments are observable, repeatable, and easy to roll back.
  • Postgres backups and restores are tested, and database health and capacity risks are visible.
  • Database migrations and emergency access follow safe, documented procedures.
  • Production alerts are useful and actionable, with clear ownership and less noise.
  • On‑call responders have the runbooks, access, and escalation paths that they need.
  • Production incidents result in tracked corrective work and measurable reliability gains.
Qualifications
  • At least 5 years of experience in site reliability, platform, infrastructure, or production engineering.
  • Strong hands‑on experience with AWS and production Kubernetes.
  • Strong experience with Terraform and infrastructure as code.
  • Experience operating Postgres in production, including backup and restore, performance, and safe migrations.
  • Experience building and operating CI/CD and deployment systems.
  • Experience with modern observability tools across metrics, logs, and traces.
  • Strong programming or scripting skills for automation and operational tools.
  • A strong record of diagnosing production failures and leading incidents through recovery.
  • Ability to write clear automation, runbooks, and technical standards.
  • Ability to work independently and partner well with application engineers.
Nice to have
  • Experience with Kubernetes security controls, policy as code, and software supply‑chain security.
  • Experience in health care or another regulated environment.
  • Experience supporting SOC 2 controls or audit evidence.
  • Experience improving cloud cost and capacity efficiency.
Our stack
  • AWS, Kubernetes, and Terraform
  • GitHub‑based CI/CD
Working model

This is an in‑office role in Lehi, Utah. You will be part of the Security and Site Reliability team and work closely with engineering leads and product engineers. The role includes participation in the engineering on‑call rotation.

What We Offer
  • Competitive salary and meaningful equity
  • 401k & insurance (medical, dental, vision, HSA, and more)
  • High ownership and the ability to shape the security function from day one
  • Direct collaboration with founders and engineering leadership
  • The opportunity to defend technology that meaningfully improves healthcare operations
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Tandem Inc. • Lehi (UT), Northern (KY)

Hybrid
USD 120,000 - 180,000
Competitive salary
Meaningful equity
401k & insurance
+1
Senior Full Stack Engineer
Senior Full Stack Engineer

Tandem Inc. • Lehi (UT)

On-site
USD 100,000 - 130,000
Competitive salary
Meaningful equity
Remote-friendly work environment
Senior Site Reliability Engineer - Healthcare Infra Equity
Senior Site Reliability Engineer - Healthcare Infra Equity

Enzo Health • Lehi (UT)

On-site
USD 120,000 - 180,000
Competitive salary
Equity
401k & Insurance
+2
Security Architect
Security Architect

Tandem Inc. • Lehi (UT)

On-site
USD 130,000 - 160,000
Competitive salary
Meaningful equity
Direct collaboration with leadership
Senior Site Reliability Engineer II
Senior Site Reliability Engineer II

Juniper Square • United States

On-site
USD 165,000 - 195,000
Health, dental, and vision care
Life insurance
Mental wellness coverage
+3
Senior Systems Engineer Onsite or Remote San Francisco, California
Senior Systems Engineer Onsite or Remote San Francisco, California

Evidently Ltd. • San Francisco (CA), Northern (KY)

Hybrid
USD 200,000 - 250,000
Site Reliability Engineer
Site Reliability Engineer

ProdataKey • Draper (UT)

On-site
USD 75,000 - 125,000
Comprehensive medical coverage
Dental and vision coverage
401(k) with company match
+2
Senior SRE - HealthTech Reliability & Security
Senior SRE - HealthTech Reliability & Security

Tandem Inc. • Lehi (UT), Northern (KY)

Hybrid
USD 120,000 - 180,000
Competitive salary
Meaningful equity
401k & insurance
+1
3510- Site Reliability Engineer II
3510- Site Reliability Engineer II

Innovaccer • Dallas (TX)

On-site
USD 110,000 - 140,000
Generous Paid Time Off: 22 days per year plus company holidays
Best-in-Class Parental Leave
Comprehensive insurance coverage
Platform Site Reliability Engineer
Platform Site Reliability Engineer

Specter • San Francisco (CA)

On-site
USD 180,000 - 230,000