Cloud Site Reliability Engineer - AWS, Observability & DR

Cadwell

United States

Remote

USD 120,000 - 130,000

Full time

8 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Cadwell is seeking a Cloud Site Reliability Engineer to own the reliability, performance, and security of its cloud-hosted infrastructure supporting healthcare facilities using Cadwell’s software suite. You will implement IaC, deploy automation, and establish robust monitoring for AWS environments.

You will lead incident response, design backup/DR strategies, and collaborate with software engineering to improve application reliability and cost efficiency within a HIPAA/GDPR compliant framework.

Qualifications

  • 8+ years in site reliability engineering, cloud infrastructure, or DevOps.
  • Hands-on experience operating production AWS environments.
  • Experience in regulated environments (medical device, healthcare, or financial services) preferred.
  • AWS certification preferred.
  • Familiarity with virtualization technologies preferred.
  • Familiarity with HL7 integration tooling preferred.

Responsibilities

  • Build, maintain, and continuously improve Cadwell’s AWS cloud infrastructure supporting hosted customer environments, applying infrastructure-as-code practices using Terraform, YAML, and JSON/Jinja.
  • Automate build, test, and deployment pipelines for cloud-hosted Cadwell applications to reduce manual effort and eliminate configuration drift across customer environments.
  • Build and maintain log ingestion, monitoring, alerting, and observability systems that provide early warning of degradation in hosted clinical environments.
  • Lead incident response for cloud-hosted environments, serving as the escalation point for availability and performance events and conducting post-mortems with corrective actions tracked to closure.
  • Design, implement, and routinely test backup and disaster recovery strategies for hosted customer data to meet the recovery time and recovery point objectives committed to customers.
  • Optimize cloud compute, storage, and lifecycle policies to balance performance, clinical data retention requirements, and cost across the hosted footprint.
  • Implement and maintain cybersecurity best practices across cloud infrastructure, including identity and access management, network segmentation, encryption, patching, and vulnerability remediation.
  • Partner with software engineering to improve application reliability, scalability, and performance, contributing to architecture decisions for new and migrating hosted deployments.
  • Partner with Cadwell’s enterprise support and project delivery teams on hosted environment onboarding, migrations, and upgrades, providing cloud infrastructure expertise throughout the customer lifecycle.
  • Ensure cloud infrastructure and processes comply with applicable healthcare data privacy and security requirements (e.g., HIPAA, GDPR) and with Cadwell’s quality system, including IEC 62304 software lifecycle processes.
  • Document infrastructure architecture, runbooks, and escalation procedures so that support and on-call staff can operate hosted environments consistently.
  • Other activities as directed, assigned, or requested.

Skills

AWS cloud
Terraform
CI/CD automation
Security best practices
Incident response
Technical communication
Mentoring engineers

Education

Bachelor’s degree in Computer Science or related field
AWS certification (Solutions Architect or SysOps)

Tools

AWS ECS
VMware
Hyper‑V
Mirth Connect

Job description

Cadwell is seeking a Cloud Site Reliability Engineer to own the reliability, performance, and security of its cloud-hosted infrastructure supporting healthcare facilities using Cadwell’s software suite. You will implement IaC, deploy automation, and establish robust monitoring for AWS environments.

You will lead incident response, design backup/DR strategies, and collaborate with software engineering to improve application reliability and cost efficiency within a HIPAA/GDPR compliant framework.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Cloud Site Reliability Engineer
Cloud Site Reliability Engineer

Cadwell • United States

Remote
USD 120,000 - 130,000
Site Reliability Engineer — AI-Driven Cloud Resilience
Site Reliability Engineer — AI-Driven Cloud Resilience

fabrichealth • New York (NY)

On-site
USD 135,000 - 160,000
Medical
Dental
Vision
+4
Senior Site Reliability Engineer - Cloud Platform & Growth
Senior Site Reliability Engineer - Cloud Platform & Growth

RXinsider LTD. • Cranberry Township

Hybrid
USD 120,000 - 180,000
Senior Site Reliability Engineer — Cloud, Observability & Automation
Senior Site Reliability Engineer — Cloud, Observability & Automation

Nocd- • Chicago (IL)

Hybrid
USD 160,000 - 200,000
Downtown Chicago office with on-site =
Hybrid work model (3x a week in-office
Competitive pay with incentives
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Kontakt.io • United States

On-site
USD 150,000 - 190,000
Senior Cloud Platform Engineer — Remote & DR Resilience
Senior Cloud Platform Engineer — Remote & DR Resilience

HealthEdge Software, Inc. • Northern (KY)

Hybrid
USD 110,000 - 118,000
Site Reliability Engineer - Build Resilient Cloud Systems
Site Reliability Engineer - Build Resilient Cloud Systems

Avalore • Virginia (MN)

On-site
USD 107,000 - 220,000
Health care plan
401k/IRA with matching
Life Insurance
+4
Datacenter Reliability Engineer
Datacenter Reliability Engineer

Amazon Web Services (AWS) • Herndon (VA)

On-site
USD 117,000 - 160,000
Remote Senior Site Reliability Engineer — AI‑Driven Infra
Remote Senior Site Reliability Engineer — AI‑Driven Infra

Precisely • Atlanta (GA)

On-site
USD 150,000 - 190,000
Senior Site Reliability Engineer - Cloud & Automation Lead
Senior Site Reliability Engineer - Cloud & Automation Lead

Ss • Waltham (MA)

Hybrid
USD 130,000 - 140,000