Remote SRE II: Cloud, Data Ops & Incident Response

Cohere Health, Inc.

Boston (MA)

Hybrid

USD 100,000 - 110,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Fully remote
5% travel
Medical insurance
Dental insurance
Vision insurance
Life insurance
Disability insurance
Employee assistance program
401(k) plan with company match
Parental leave (up to 14 weeks)

Job summary

Cohere Health is seeking an operational-focused Site Reliability Engineer to maximize the availability, performance, and resilience of our production healthcare systems. You will bridge AWS cloud infrastructure, MERN stack applications, and large-scale data workflows.

You’ll spend ~60% on live incident remediation, data pipeline operations, and infrastructure tuning, and ~40% on building automated solutions to reduce toil and improve reliability across the platform.

Qualifications

  • Minimum of 3+ years operating multi-tenant cloud SaaS platforms at scale.
  • Deep AWS experience including Lambda, ECS/EKS, EMR/Glue, EC2, VPC, IAM, and CloudWatch.
  • Proficiency in Python (including PySpark) and Node.js for automation.
  • Experience managing distributed data orchestration pipelines and ETL tools.
  • Production MySQL and Athena performance tuning.
  • Infrastructure as code with Terraform or OpenTofu.
  • HIPAA-regulated environment experience.
  • 4+ years software/systems with 1–2 years cloud ops and data workflows.
  • Crisis management and strong communication during outages.

Responsibilities

  • Maintain uptime, scalability, and security of AWS-hosted MERN apps and data architectures.
  • Optimize serverless architectures in AWS Lambda, addressing cold starts and timeouts.
  • Oversee PySpark data workflows and SOPs for large-scale ingestion; triage failures.
  • Participate in on-call rotation to quickly triage outages and data bottlenecks.
  • Ensure HIPAA, SOC2, HITRUST compliance across runtimes and data pipelines.
  • Automate toil elimination: seed data, provision infrastructure, recover pipelines.
  • Build dashboards and alerts for Node.js loops, PySpark stages, memory leaks.
  • Lead blameless post-mortems and implement permanent fixes.

Skills

SaaS platform operations
AWS Cloud Engineering
Automation & data scripting
Data pipelines
Database administration
IaC (Terraform/OpenTofu)
HIPAA compliance
Distributed systems
Live incident response

Job description

Cohere Health is seeking an operational-focused Site Reliability Engineer to maximize the availability, performance, and resilience of our production healthcare systems. You will bridge AWS cloud infrastructure, MERN stack applications, and large-scale data workflows.

You’ll spend ~60% on live incident remediation, data pipeline operations, and infrastructure tuning, and ~40% on building automated solutions to reduce toil and improve reliability across the platform.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote SRE II — Cloud Ops & Data Pipelines
Remote SRE II — Cloud Ops & Data Pipelines

Cohere Health • United States

On-site
USD 100,000 - 110,000
Medical insurance
Dental & Vision
401K with company match
+3
Remote SRE II - Cloud-Native Healthcare
Remote SRE II - Cloud-Native Healthcare

NationsBenefits • Plantation (FL)

Remote
USD 110,000 - 150,000
Competitive compensation
Unlimited PTO
Career development opportunities
+2
Remote SRE II: Cloud-Native Reliability & Automation
Remote SRE II: Cloud-Native Reliability & Automation

NationsBenefits, LLC • Plantation (FL)

On-site
USD 110,000 - 160,000
Unlimited PTO
Competitive compensation & benefits
Career growth opportunities
+1
Remote SRE II: Kubernetes & Automation
Remote SRE II: Kubernetes & Automation

NationsBenefits, LLC • United States

On-site
USD 110,000 - 160,000
Unlimited PTO
Competitive benefits
Career growth
Senior Cloud SRE: AWS, Serverless & Incident Response
Senior Cloud SRE: AWS, Serverless & Incident Response

Apply • Northern (KY)

Hybrid
USD 120,000 - 150,000
Remote SRE Manager: Lead Reliability & Automation
Remote SRE Manager: Lead Reliability & Automation

NationsBenefits, LLC • Plantation (FL)

Remote
USD 140,000 - 180,000
Fully remote
Unlimited PTO
Competitive compensation
+1
Remote SRE Lead: Multi-Cloud Incident Command & CI/CD
Remote SRE Lead: Multi-Cloud Incident Command & CI/CD

AgileEngine, LLC • Brazil (IN), Northern (KY)

Hybrid
USD 120,000 - 160,000
Remote work 100%
Flexible hours
Annual learning budget
Remote Senior SRE — Cloud Ownership & Automation
Remote Senior SRE — Cloud Ownership & Automation

Synthesia • United States

On-site
USD 120,000 - 160,000
Senior Cloud SRE & Automation Engineer
Senior Cloud SRE & Automation Engineer

Orion Health group • Dallas (TX), Northern (KY)

Hybrid
USD 120,000 - 165,000
Principal Site Reliability Engineer (SRE)
Principal Site Reliability Engineer (SRE)

Symmetrio • United States

Hybrid
USD 120,000 - 160,000
Health Care Plan (Medical, Dental & Vision)
Retirement Plan (401k, IRA)
Paid Time Off (Vacation, Sick & Public Holidays)