Site Reliability Engineer ll

Cohere Health, Inc.

Boston (MA)

On-site

USD 100,000 - 110,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Fully remote
5% travel
Medical insurance
Dental insurance
Vision insurance
Life insurance
Disability insurance
Employee assistance program
401(k) plan with company match
Parental leave (up to 14 weeks)

Job summary

Cohere Health is seeking an operational-focused Site Reliability Engineer to maximize the availability, performance, and resilience of our production healthcare systems. You will bridge AWS cloud infrastructure, MERN stack applications, and large-scale data workflows.

You’ll spend ~60% on live incident remediation, data pipeline operations, and infrastructure tuning, and ~40% on building automated solutions to reduce toil and improve reliability across the platform.

Qualifications

  • Minimum of 3+ years operating multi-tenant cloud SaaS platforms at scale.
  • Deep AWS experience including Lambda, ECS/EKS, EMR/Glue, EC2, VPC, IAM, and CloudWatch.
  • Proficiency in Python (including PySpark) and Node.js for automation.
  • Experience managing distributed data orchestration pipelines and ETL tools.
  • Production MySQL and Athena performance tuning.
  • Infrastructure as code with Terraform or OpenTofu.
  • HIPAA-regulated environment experience.
  • 4+ years software/systems with 1–2 years cloud ops and data workflows.
  • Crisis management and strong communication during outages.

Responsibilities

  • Maintain uptime, scalability, and security of AWS-hosted MERN apps and data architectures.
  • Optimize serverless architectures in AWS Lambda, addressing cold starts and timeouts.
  • Oversee PySpark data workflows and SOPs for large-scale ingestion; triage failures.
  • Participate in on-call rotation to quickly triage outages and data bottlenecks.
  • Ensure HIPAA, SOC2, HITRUST compliance across runtimes and data pipelines.
  • Automate toil elimination: seed data, provision infrastructure, recover pipelines.
  • Build dashboards and alerts for Node.js loops, PySpark stages, memory leaks.
  • Lead blameless post-mortems and implement permanent fixes.

Skills

SaaS platform operations
AWS Cloud Engineering
Automation & data scripting
Data pipelines
Database administration
IaC (Terraform/OpenTofu)
HIPAA compliance
Distributed systems
Live incident response

Job description

This is a remote-first role that may require travel to Boston, MA for new hire onboarding and occasional in-person team meetings and company events.

We are seeking an operational-focused Site Reliability Engineer (SRE) to maximize the availability, performance, and resilience of our production healthcare systems. In this role, you will bridge the gap between AWS cloud infrastructure, MERN stack applications, and large-scale data workflows. You will spend roughly 60% of your time on live incident remediation, data pipeline operations, and Node.js/Python infrastructure tuning, and 40% on engineering automated solutions to eliminate operational toil.

What you’ll do:
  • Production Operations: Maintain the continuous uptime, scalability, and security of our AWS-hosted MERN applications and backend data architectures.
  • Serverless Execution: Manage, optimize, and troubleshoot event-driven architectures running on AWS Lambda, focusing on cold-start mitigation, memory allocation, and execution timeouts.
  • Data Pipeline Execution: Monitor scheduled PySpark data workflows, execute standard operating procedures (SOPs) for large-scale data ingestion, and rapidly triage, rerun, or patch failed data processing jobs.
  • Incident Management: Participate in a collaborative on-call rotation to rapidly triage, debug, and mitigate live application outages and data flow bottlenecks.
  • Healthcare Compliance: Maintain strict HIPAA, SOC2, and HITRUST compliance profiles across all runtime environments, storage systems, and data pipelines handling Protected Health Information (PHI).
  • Toil Elimination: Engineer automated workflows to eliminate repetitive tasks like manual data seeding, infrastructure provisioning, and routine PySpark pipeline recovery steps.
  • Observability Engineering: Build specialized dashboards and alerts to monitor Node.js event loops, PySpark job execution stages, driver/worker memory leaks, and data pipeline throughput anomalies.
  • Post-Mortem Culture: Lead blameless post-mortems for operational and data processing failures, translating system crashes into permanent structural fixes.
What you’ll need:
  • SaaS Platform Experience: Minimum of 3+ years of hands‑on experience operating multi‑tenant, cloud‑hosted, or cloud‑native SaaS platforms at scale.
  • AWS Cloud Engineering: Deep expertise operating AWS core services, specifically AWS Lambda, Amazon ECS/EKS, Amazon EMR or AWS Glue (for Spark), EC2, VPC networking, IAM permissions, and CloudWatch.
  • Automation & Data Languages: Professional competency in writing, debugging, and maintaining automation scripts and data tools using Python (including PySpark APIs) and Node.js.
  • Data Operations: Experience managing and troubleshooting distributed data orchestration pipelines, ETL tools, message queues (e.g., AWS SQS/SNS, RabbitMQ), or stream processing frameworks.
  • Database Administration: Practical experience managing, sharding, indexing, and optimizing production‑grade MySQL DB & Athena (RDS or self‑hosted).
  • Infrastructure as Code: Proven ability to deploy and maintain immutable infrastructure utilizing Terraform or OpenTofu.
  • Healthcare Experience: Minimum 1 year working within HIPAA‑regulated environments. Direct experience securing data‑at‑rest and data‑in‑transit containing sensitive patient records is preferred.
  • Education & Experience: Minimum of 4 years of software/systems experience, with at least 1–2 years focused on live cloud operations and distributed data workflow management is preferred.
  • Crisis Management: Calm under pressure with a methodical approach to identifying and isolating PySpark driver OOM (Out of Memory) errors or data corruption during high‑stress outages. Attention to detail and effective communication skills will be critical in working with clients and internal stakeholders is preferred.
Benefits

Fully remote opportunity with about 5% travel.

Medical, dental, vision, life, disability insurance, and Employee Assistance Program.

401K retirement plan with company match; flexible spending and health savings account.

Up to 14 weeks of paid parental leave.

The salary range for this position is $100,000 to $110,000 annually; as part of a total benefits package which includes health insurance, 401k and bonus. In accordance with state applicable laws, Cohere is required to provide a reasonable estimate of the compensation range for this role. Individual pay decisions are ultimately based on a number of factors, including but not limited to qualifications for the role, experience level, skillset, and internal alignment.

This role is not eligible for hire in: CA

Equal Opportunity Statement

Cohere Health is an Equal Opportunity Employer. We are committed to fostering an environment of mutual respect where equal employment opportunities are available to all. To us, it’s personal.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer, Client Implementations
Senior Software Engineer, Client Implementations

Cohere Health, Inc. • Boston (MA)

Hybrid
USD 128,000 - 145,000
Medical and dental insurance
401(k) retirement plan with match
Up to 14 weeks of paid parental leave
Lead Software Engineer
Lead Software Engineer

Cohere Health, Inc. • Boston (MA)

Hybrid
USD 155,000 - 175,000
Fully remote work
Travel opportunities (~5%)
Health insurance
+9
Lead Software Engineer - Integrations New
Lead Software Engineer - Integrations New

Cohere Health • Northern (KY)

Hybrid
USD 150,000 - 185,000
Medical, dental, vision insurance
401(k) with company match
Parental leave up to 14 weeks
+1
Manager, Client Analytics
Manager, Client Analytics

Socket.dev • United States

On-site
USD 115,000 - 130,000
Fully remote
Health insurance
401K and bonus
+3
Manager, Forward Deployed Automation Strategy
Manager, Forward Deployed Automation Strategy

Define Ventures • United States

On-site
USD 115,000 - 130,000
Remote with travel
Health insurance
401K plan
+3
Lead Forward Deployed Automation Strategist
Lead Forward Deployed Automation Strategist

Define Ventures • United States

On-site
USD 115,000 - 130,000
Medical, dental, vision, life, and ADP
401K with company match
Flexible spending and health savings
+3
Manager, Client Analytics New
Manager, Client Analytics New

Cohere Health, Inc. • Boston (MA)

Hybrid
USD 115,000 - 130,000
Health insurance
401k plan
Bonus
+1
Senior Clinical Informaticist New
Senior Clinical Informaticist New

Cohere Health, Inc. • Boston (MA)

Hybrid
USD 75,000 - 100,000
Remote-first
Health insurance
401K with match
+3
Remote SRE II — Cloud Ops & Data Pipelines
Remote SRE II — Cloud Ops & Data Pipelines

Cohere Health • United States

On-site
USD 100,000 - 110,000
Medical insurance
Dental & Vision
401K with company match
+3
Senior Integration Analyst New
Senior Integration Analyst New

Cohere Health, Inc. • Boston (MA)

Hybrid
USD 105,000 - 120,000
Medical insurance
401K with company match
Paid parental leave