Site Reliability Engineer

Stability Technology

United States

On-site

USD 120,000 - 170,000

Full time

13 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Stability Technology is seeking a Site Reliability / Observability Engineer to build and own monitoring, logging, tracing, and alerting across our nationwide healthcare platform. You will define healthy system criteria, develop dashboards, and reduce alert noise while ensuring reliability for telehealth and RCM services.

The role requires strong technical judgement, collaboration with Enterprise Tech, and hands-on work with Datadog, Grafana/Prometheus, CloudWatch, and AWS services.

Qualifications

  • 3+ years of experience building and operating observability infrastructure in production.
  • Experience defining, tracking, and communicating SLIs/SLOs.
  • Strong AWS fundamentals are required, particularly CloudWatch, Lambda, and EventBridge.
  • Scripting or automation experience using Python, Node.js, or similar.

Responsibilities

  • Design and implement monitoring, logging, and distributed tracing across internal platforms and telephony.
  • Define and track SLIs and SLOs for critical systems.
  • Build actionable, low-noise alerting that identifies meaningful operational issues.
  • Own the observability toolchain across metrics, logs, traces, dashboards, and on-call tooling.
  • Support incident response with real-time visibility and lead post-incident analysis.
  • Instrument new systems and integrations, including AWS-based telephony and EHR.
  • Track and report system reliability trends to technology leadership.
  • Tune thresholds and retire alerts that reflect meaningful risk.

Skills

Observability
AWS fundamentals
Scripting
Incident response
Alerting
Monitoring

Tools

Datadog
New Relic
Grafana
Prometheus
CloudWatch
AWS Lambda
EventBridge
PagerDuty
GitHub
Confluence
Jira
Slack

Job description

Our client is a growing healthcare technology organization that provides clinical and operational infrastructure supporting virtual care delivery nationwide. Its technology ecosystem spans telehealth, revenue cycle management, practice management, provider staffing, payer coverage, clinician licensing, and other mission-critical healthcare services.

Because these systems directly support patient access to care, system availability, reliability, and proactive monitoring are critical.

Role Summary

The Site Reliability / Observability Engineer will build and own the monitoring, logging, tracing, and alerting environment that provides Enterprise Technology and Engineering teams with real-time visibility into system health.

This person will define what “healthy” looks like across critical systems, develop dashboards and alerts that surface issues early, and continuously reduce unnecessary alert noise. The role requires strong technical expertise as well as the judgment to distinguish meaningful signals from routine system activity.

Key Responsibilities
  • Design and implement monitoring, logging, and distributed tracing across internal platforms, integrations, and telephony infrastructure.
  • Define and track SLIs and SLOs for critical clinical, telephony, and business systems.
  • Build actionable, low-noise alerting that identifies meaningful operational issues.
  • Own the observability toolchain across metrics, logs, traces, dashboards, incident response, and on-call tooling.
  • Support incident response with real-time system visibility and lead post-incident data analysis.
  • Instrument new systems and integrations, including AWS-based telephony, EHR, and business systems.
  • Track and report system reliability trends to technology leadership.
  • Continuously tune thresholds and retire alerts that no longer reflect meaningful risk.
Required Qualifications
  • 3+ years of experience building and operating observability or monitoring infrastructure in a production environment.
  • Hands-on experience with modern observability platforms such as Datadog, New Relic, Grafana/Prometheus, CloudWatch, or similar technologies.
  • Experience defining, tracking, and communicating SLIs/SLOs.
  • Strong AWS fundamentals, particularly CloudWatch, Lambda, and EventBridge.
  • Scripting or automation experience using Python, Node.js, or similar technologies.
  • Experience participating in or leading incident response and post-incident reviews.
  • Strong ability to distinguish signal from noise and build alerts based on meaningful operational risk.
Preferred Qualifications
  • Experience monitoring or instrumenting contact center/telephony infrastructure, including Amazon Connect or similar platforms.
  • Experience supporting healthcare or other highly regulated/compliance-focused environments.
  • Experience with incident management and on-call platforms such as PagerDuty or Opsgenie.
  • Understanding of cost-aware observability, including balancing data retention and granularity against platform costs.
Tools & Technologies

Datadog | CloudWatch | Grafana | Prometheus | AWS Lambda | EventBridge | PagerDuty | GitHub | Confluence | Jira | Slack

What Success Looks Like
  • Monitoring identifies incidents before users or clients report them.
  • On-call engineers trust that alerts represent legitimate issues.
  • Leadership has clear visibility into system health and reliability trends.
  • New systems launch with observability built in from the beginning.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability & Observability Engineer
Site Reliability & Observability Engineer

Stability Technology • United States

On-site
USD 120,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

Infosys • Richardson (TX)

On-site
USD 80,000 - 120,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobgether • United States

Hybrid
USD 120,000 - 160,000
Competitive compensation package
Flexible work arrangements
Professional development opportunities
+2
Senior Site Reliability Engineer, Observability
Senior Site Reliability Engineer, Observability

blockchaincapital.com • New York (NY)

On-site
USD 130,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

O.C. Tanner • Salt Lake City (UT)

On-site
USD 130,000 - 180,000
SRE/Observability Engineer
SRE/Observability Engineer

BlueSky Resource Solutions • United States

Remote
USD 100,000 - 130,000
Observability & Monitoring Engineer
Observability & Monitoring Engineer

Compunnel, Inc. • Des Moines (IA)

On-site
USD 110,000 - 150,000
Site Reliability Engineer Lead
Site Reliability Engineer Lead

TechDigital Group • Tyson (AZ)

On-site
USD 120,000 - 160,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink • United States

Remote
USD 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Oracle • Frankfort (KY)

On-site
USD 120,000 - 160,000