Cloud Systems Engineer - Site Reliability

TherapyNotes.com

United States

Remote

USD 110,000 - 150,000

Full time

11 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Health insurance
Retirement plan
Profit sharing
Mentorship program
Onboarding plan
Open work environment

Job summary

TherapyNotes.com is seeking a Site Reliability Engineer to improve reliability and operability across a growing 24x7 SaaS environment. You will partner with software, infrastructure, and security teams to raise availability, performance, and resilience while reducing toil.

You will define reliable SLIs/SLOs, own incident management, and drive automation using Bash/PowerShell/Python, Terraform/OpenTofu, and Ansible.

Qualifications

  • BS degree in Information Systems, Engineering, or equivalent.
  • 5+ years in Systems Engineering, Cloud/Platform/DevOps/SRE.
  • Azure and Kubernetes experience preferred.
  • Strong Linux networking and troubleshooting of distributed systems.
  • Datadog experience preferred; Prometheus, Grafana or New Relic helpful.
  • Scripting and automation with Bash/PowerShell/Python; IaC and configuration management.
  • On-call rotations, incident response, RCA, and post-incident improvements.
  • Agile/DevOps experience; ITSM practices.
  • Software development experience or code-level investigation is a plus.

Responsibilities

  • Own and improve Datadog visibility across metrics, logs, traces, dashboards, and alerts.
  • Design and maintain high-availability, high-throughput systems for a 24x7 SaaS platform.
  • Define and improve reliability via SLIs, SLOs, and error budgets.
  • Lead incident management as commander or responder; triage and RCA.
  • Investigate issues across infra and app layers with metrics and traces.
  • Improve deployment safety with automated validation and rollback capabilities.
  • Ensure systems are maintainable by development and operations.
  • Provide on-call coverage and guidance to tech teams.
  • Comply with security and HIPAA policies.

Skills

5+ years engineering experience
Cloud/Platform Engineering
DevOps
SRE
Linux fundamentals
Scripting (Bash/PowerShell/Python)

Education

BS in Information Systems/Engineering or equivalent

Tools

Datadog
Prometheus
Grafana
New Relic
Terraform/OpenTofu
Ansible

Job description

About Us

TherapyNotes is the go-to superhero for behavioral health Practice Management and EHR software! Our top-notch SaaS solution handles scheduling, billing, documenting, telehealth, and more so clinicians can focus on awesome patient care.

We’re a dynamic team of pros who love to innovate and push the envelope, keeping our software cutting-edge. Join us, and let’s revolutionize behavioral health software together while making a real difference!

About The Position

We are seeking a Site Reliability Engineer to improve the reliability and operability of the production services and shared platforms supporting our growing 24x7 SaaS environment. In this role, you will apply software and systems engineering practices to improve availability, performance, scalability, resilience, observability, incident response, and operational automation. You will partner with software development, infrastructure, database, security, and other technology teams to establish measurable reliability goals, reduce operational toil, and ensure services are supportable throughout their lifecycle. If you are passionate about building reliable systems, solving complex production problems, and driving continuous improvement, we want to hear from you.

  • BS degree in Information Systems, Engineering, or equivalent experience.
  • 5+ years of engineering experience in Systems Engineering, Cloud or Platform Engineering, DevOps, Software Engineering, and/or SRE.
  • Experience designing and operating production systems using cloud-based compute, storage, networking, and containerization technologies; Azure and Kubernetes preferred.
  • Strong Linux systems and networking fundamentals, with experience troubleshooting complex distributed systems in production.
  • Expertise with an observability platform; Datadog experience strongly preferred. Experience with Prometheus, Grafana, New Relic, or equivalent platforms is also valuable.
  • Experience with scripting and operational automation using tools such as Bash, PowerShell, or Python, along with infrastructure-as-code and configuration-management practices.
  • Experience participating in production on-call rotations, incident response, root cause analysis, and post-incident improvement.
  • Experience working in Agile/DevOps environments and operating production services using ITSM practices where applicable.
  • Prior software development experience—or experience investigating application behavior through code, logs, and distributed traces is a plus
Responsibilities
  • Own and continuously improve how we use Datadog to make reliability visible and actionable across metrics, logs, traces, dashboards, monitors, alerts, and service-level views.
  • Design, implement, and maintain high-availability, high-throughput, data- and compute-intensive critical systems supporting a growing 24x7 SaaS platform.
  • Partner with service owners to define and improve reliability through meaningful SLIs, SLOs, error budgets, actionable alerting, and operational-readiness practices.
  • Participate in and help drive incident management for production events, serving as an incident commander or technical responder as needed. Coordinate triage, service restoration, escalation, communication, incident documentation, root cause analysis, and completion of corrective actions.
  • Partner with development teams to investigate issues across the infrastructure and application layers using metrics, logs, distributed traces, and code-level context.
  • Improve deployment safety and service resilience through automated validation, recovery and rollback capabilities, reliability testing, and analysis of system failure modes.
  • Partner with other technical leaders to ensure all newly introduced systems are supportable and maintainable by both development and operations.
  • Provide escalated technical guidance and support to other technology teams throughout the organization.
  • Provide on-call coverage for production support and other duties as required.
  • Ensure supported systems and operational activities comply with organizational security, HIPAA, and operating policies.
  • Identify and eliminate repetitive operational toil using Bash, PowerShell, Python, or Ansible. Manage infrastructure as code using Terraform/OpenTofu and configuration automation using Ansible.
  • Competitive salary - $110,000-$150,000
  • Employer sponsored health, dental, vision, life, and disability insurance
  • Retirement plan with company contribution
  • Annual company profit sharing
  • Personal development/training budget
  • Open, collaborative work environment
  • Extensive 2-week onboarding plan
  • Comprehensive mentorship program
Equal Opportunity Employer Statement & Applicant Rights

TherapyNotes LLC is an Equal Opportunity Employer and does not discriminate based on race, color, religion, sex, national origin, age, disability, genetic information, or any other protected status under federal, state, or local law. We are committed to providing a workplace free of discrimination and harassment.For more information about your rights under federal employment laws, please review the following:

  • Know Your Rights: Workplace Discrimination is Illegal
  • Family and Medical Leave Act (FMLA): Employee Rights Under FMLA

If you require a reasonable accommodation during the application process, please contact humanresources@therapynotes.com.

9/17/2026

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Software Developer
Senior Software Developer

TherapyNotes.com • Philadelphia

Remote
USD 120,000 - 140,000
Competitive salary
Health, dental, vision, life insurance
Retirement plan with company match
+4
Cloud Systems Engineer
Cloud Systems Engineer

TherapyNotes, LLC • United States

On-site
USD 100,000 - 120,000
Employer sponsored health insurance
Dental insurance
Vision insurance
+7
Cloud Security Engineer
Cloud Security Engineer

TherapyNotes.com • Philadelphia

On-site
USD 110,000 - 150,000
Competitive salary
Health insurance
Retirement plan
+5
Cloud Security Engineer
Cloud Security Engineer

TherapyNotes.com • Pennsylvania

On-site
USD 110,000 - 150,000
Health, dental, vision, life, and STD
401(k) with company contributions
Profit sharing
+2
Cyber Security Engineer (Application Security)
Cyber Security Engineer (Application Security)

TherapyNotes.com • United States

Remote
USD 110,000 - 150,000
Health insurance
Dental insurance
Vision insurance
+3
Cyber Security Engineer (Application Security)
Cyber Security Engineer (Application Security)

TherapyNotes.com • Philadelphia

On-site
USD 110,000 - 150,000
Salary up to $150k
Insurance
Retirement plan
+5
Senior Site Reliability Engineer
Senior Site Reliability Engineer

NeuroFlow • Philadelphia

On-site
USD 140,000 - 210,000
Flexible work schedule
Unlimited PTO
Medical coverage
+6
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

United States Digital Space LLC • United States

Hybrid
USD 160,000 - 180,000
Hybrid work model
401(k) matching
Parental leave
+1
Sr Platform DevOps Engr, SRE - Remote
Sr Platform DevOps Engr, SRE - Remote

UnitedHealth Group • Eden Prairie (MN)

Hybrid
USD 92,000 - 164,000
Comprehensive benefits package
Equity stock purchase plan
401(k) contribution
Software Developer
Software Developer

TherapyNotes, LLC • United States

On-site
USD 65,000 - 115,000
Employer sponsored health, dental, and vision insurance
Retirement plan with company contribution
Annual profit sharing
+2