Our client is a growing healthcare technology organization that provides clinical and operational infrastructure supporting virtual care delivery nationwide. Its technology ecosystem spans telehealth, revenue cycle management, practice management, provider staffing, payer coverage, clinician licensing, and other mission-critical healthcare services.
Because these systems directly support patient access to care, system availability, reliability, and proactive monitoring are critical.
Role Summary
The Site Reliability / Observability Engineer will build and own the monitoring, logging, tracing, and alerting environment that provides Enterprise Technology and Engineering teams with real-time visibility into system health.
This person will define what “healthy” looks like across critical systems, develop dashboards and alerts that surface issues early, and continuously reduce unnecessary alert noise. The role requires strong technical expertise as well as the judgment to distinguish meaningful signals from routine system activity.
Key Responsibilities
- Design and implement monitoring, logging, and distributed tracing across internal platforms, integrations, and telephony infrastructure.
- Define and track SLIs and SLOs for critical clinical, telephony, and business systems.
- Build actionable, low-noise alerting that identifies meaningful operational issues.
- Own the observability toolchain across metrics, logs, traces, dashboards, incident response, and on-call tooling.
- Support incident response with real-time system visibility and lead post-incident data analysis.
- Instrument new systems and integrations, including AWS-based telephony, EHR, and business systems.
- Track and report system reliability trends to technology leadership.
- Continuously tune thresholds and retire alerts that no longer reflect meaningful risk.
Required Qualifications
- 3+ years of experience building and operating observability or monitoring infrastructure in a production environment.
- Hands-on experience with modern observability platforms such as Datadog, New Relic, Grafana/Prometheus, CloudWatch, or similar technologies.
- Experience defining, tracking, and communicating SLIs/SLOs.
- Strong AWS fundamentals, particularly CloudWatch, Lambda, and EventBridge.
- Scripting or automation experience using Python, Node.js, or similar technologies.
- Experience participating in or leading incident response and post-incident reviews.
- Strong ability to distinguish signal from noise and build alerts based on meaningful operational risk.
Preferred Qualifications
- Experience monitoring or instrumenting contact center/telephony infrastructure, including Amazon Connect or similar platforms.
- Experience supporting healthcare or other highly regulated/compliance-focused environments.
- Experience with incident management and on-call platforms such as PagerDuty or Opsgenie.
- Understanding of cost-aware observability, including balancing data retention and granularity against platform costs.
Tools & Technologies
Datadog | CloudWatch | Grafana | Prometheus | AWS Lambda | EventBridge | PagerDuty | GitHub | Confluence | Jira | Slack
What Success Looks Like
- Monitoring identifies incidents before users or clients report them.
- On-call engineers trust that alerts represent legitimate issues.
- Leadership has clear visibility into system health and reliability trends.
- New systems launch with observability built in from the beginning.