Senior Site Reliability Engineer (Systems Engineer III)

ophelia

United States

On-site

USD 140,000 - 210,000

Full time

5 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Ophelia is seeking a Senior Site Reliability Engineer (SE III) to own reliability, availability, and performance of our Google Cloud Platform-based systems. You will define SLOs/SLIs, build observability, and drive incident response with a focus on MTTR, while expanding our IaC footprint and IAM hardening practices.

You will mentor a team of full‑stack engineers, promote best practices, and collaborate with IT on HIPAA/SOC 2 controls, security, and regulatory compliance.

Qualifications

  • Experience as a Senior Site Reliability Engineer or equivalent in a cloud environment.
  • Familiar with GCP, Cloud Run, Firestore, and Firebase.
  • Knowledge of IAM, secrets management, and HIPAA/SOC 2 controls is beneficial.

Responsibilities

  • Own reliability, availability, and performance of the GCP platform.
  • Define and operationalize SLOs and SLIs with the team.
  • Build observability, incident response, and runbooks.
  • Mature incident management with clear categorization, escalation, and AI-assisted tooling.

Skills

SRE experience
GCP platform
Incident response
Team coaching

Tools

Terraform
PagerDuty
Firestore
Firebase
Cloud Run
Claude
Gemini
Node/TypeScript

Job description

Are you looking for a role in a company that's solving one of the greatest challenges of our lifetime? Ophelia helps people end their opioid use and restore their quality of life with respect for their time and dignity. Our mission is to make evidence-based treatments for opioid use disorder (OUD) accessible to everyone... and we're looking to bring more people onto our team to help us achieve it.

Ophelia is a venture-backed, healthcare startup that helps individuals with OUD by providing FDA-approved medication and clinical care through a telehealth platform. Our approach is discreet, convenient, and affordable. We've been successfully operating in 16 states for almost six years and we're excited to continue our growth. We are a team of physicians, scientists, entrepreneurs, researchers and White House advisors, backed by leading technology and healthcare investors working to re-imagine and re-build OUD treatment in America.

About the Role

As a Senior Site Reliability Engineer (SE III) at Ophelia, you will be our dedicated reliability, availability, and performance hire, playing a key role in making the systems that support our mission of treating opioid use disorder through telehealth stable, observable, and fast. Our stack is TypeScript on Node, React, and Firebase (e.g., Firestore, Authentication, Hosting) running on Cloud Run in Google Cloud Platform (GCP). Your work will have a direct impact on patients, clinicians, and our ability to scale.

You will own and drive our system stability and availability initiative, which is already underway: Terraform-driven uptime checks feed our engineering key performance indicators (KPIs), we have an initial service level objective (SLO) plan, and we've recently revamped our on-call rotation and runbook to keep the on-call engineer focused on severity incidents (SEVs). You will take this from a good start to a mature practice: defining and operationalizing SLOs, building out logging, monitoring, and alerting grounded in sound measurement (e.g., percentile-based latency indicators rather than averages), moving incident response into PagerDuty with clear categorization and escalation paths, expanding our infrastructure as code (IaC) footprint, and working down known pain points such as long-latency endpoints and Cloud Run cold starts. You will also partner closely with our head of IT on IAM hardening, secrets management, audit logging, and HIPAA/SOC 2 infrastructure controls. You will be the directly responsible individual (DRI) for our mean time to recovery (MTTR) KPI and for the measurement and reporting of our availability KPIs.

As a senior engineer, you will be a force multiplier for our team of fullstack engineers who are focused on product work. That means leveling up the team's GCP and DevOps skills through documentation, runbooks, pairing, and code review, and bringing reliability recommendations to new features as they are designed, while primarily owning the infrastructure work yourself. This role is entirely reliability-focused for the first six months; after that, you may contribute to product work as needed, though product work will be at most about 25% of your time.

Ophelia's Technology team actively encourages and invests in AI-augmented ways of working: using tools like Claude and Gemini to accelerate infrastructure code, runbooks, log analysis, and incident triage, and staying current as the AI landscape evolves. Consistent with Ophelia's AI Position Statement, we treat AI as a force multiplier for our team, not a replacement for human judgment: it's there to help you move faster from alert to root cause and spend more time on the reliability and architecture decisions that actually require a person. You'll be expected to use AI fluently in your own workflow, to build it into our operational tooling, and to help shape how the larger technology team uses it responsibly and effectively as it evolves.

Together, we will help hard-to-reach individuals treat their opioid dependence. While direct experience in this treatment area is not mandatory, knowledge of the healthcare space, including understanding health outcomes, benchmarks, systems, and regulatory compliance (HIPAA), is highly beneficial.

Responsibilities

Own reliability, availability, and performance of our GCP platform. Define and operationalize SLOs and service level indicators (SLIs) with the team, and drive the work that moves them (e.g., endpoint latency, Cloud Run cold starts, time to recovery), and own our disaster recovery posture (e.g., Firestore backup and restore testing, recovery point and recovery time objectives, and multi-region resilience).

Build observability and incident response. Extend our GCP Cloud Operations setup (log-based metrics, alerting policies, dashboards), move alerting from Slack into PagerDuty, participate in our on-call rotation alongside the rest of the team, and mature that process with clear incident categorization, escalation paths, and an AI-assiste

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Analyst, Strategic Finance New
Senior Analyst, Strategic Finance New

Ophelia • Northern (KY)

Hybrid
USD 125,000 - 130,000
Health insurance
PTO 20 days
Company holidays
+2
Senior Analyst, Strategic Finance Ophelia · Full-time · On-site $120,000–130,000 53 minutes ago
Senior Analyst, Strategic Finance Ophelia · Full-time · On-site $120,000–130,000 53 minutes ago

Emploive • Northern (KY)

Hybrid
USD 120,000 - 130,000
Medical Insurance
PTO 20 days + 5 weeks after years 2/5
Company Holidays 10
+2
Product Analyst III
Product Analyst III

Ophelia • United States

Remote
USD 124,000 - 134,000
Territory Manager (Capital District) Albany, New York
Territory Manager (Capital District) Albany, New York

Ophelia Health, Inc. • City of Albany (NY), Northern (KY)

Hybrid
USD 75,000 - 90,000
Health insurance
PTO 20 days
Holidays
+2
Associate, Clinical Operations (Data and Systems)
Associate, Clinical Operations (Data and Systems)

Ophelia • United States

On-site
USD 60,000 - 65,000
Health insurance
PTO 20 days (4 weeks)
Work From Home stipend
+1
Senior Associate, Community Partnership Operations
Senior Associate, Community Partnership Operations

Ophelia • Northern (KY)

Hybrid
USD 85,000 - 95,000
Field Partnerships Growth Manager
Field Partnerships Growth Manager

Ophelia • City of Albany (NY)

On-site
USD 80,000 - 100,000
Competitive medical, vision, and health insurance
20 days PTO, increasing with tenure
Work From Home stipend
+1
Territory Manager (Capital District)
Territory Manager (Capital District)

Ophelia • City of Albany (NY)

On-site
USD 80,000 - 100,000
Competitive medical, vision, and health insurance
20 days PTO, increasing with tenure
Work From Home stipend
+1
Territory Manager (NYC Region)
Territory Manager (NYC Region)

Ophelia • New York (NY)

On-site
USD 75,000 - 90,000
Health insurance
PTO (20 days)
Remote work stipend
+2
Telehealth Nurse Practitioner or Physician Assistant (Remote) - New York License
Telehealth Nurse Practitioner or Physician Assistant (Remote) - New York License

Ophelia • United States

Remote
USD 110,000 - 125,000
Remote work anywhere in the United St
Competitive insurance options