Site Reliability Engineer - Cloud, Kubernetes & Observability

Socket.dev

Philadelphia (Philadelphia County)

On-site

USD 110,000 - 150,000

Full time

4 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Health insurance
Retirement plan
Profit sharing
Training budget
Onboarding program
Mentorship
Open work environment

Job summary

TherapyNotes LLC is seeking a Site Reliability Engineer to strengthen the reliability and operability of our 24×7 SaaS environment, collaborating with software, infrastructure, security, and database teams. You will improve observability, automate deployments, and drive incident response while aligning with SLIs/SLOs and ITSM processes.

You will design high-availability systems, own critical metrics, and participate in on-call rotations, incident management, and post-incident reviews to prevent

Qualifications

  • BS degree in Information Systems, Engineering, or equivalent experience.
  • 5+ years of engineering experience in Systems Engineering, Cloud or Platform Engineering, DevOps, Software Engineering, and/or SRE.
  • Experience designing and operating production systems using cloud-based compute, storage, networking, and containerization technologies; Azure and Kubernetes preferred.
  • Strong Linux systems and networking fundamentals, with experience troubleshooting complex distributed systems in production.
  • Datadog experience strongly preferred; Prometheus, Grafana, New Relic, or equivalent platforms is also valuable.
  • Scripting and operational automation using Bash, PowerShell, or Python, with infrastructure-as-code and configuration-management practices.
  • Experience participating in production on-call rotations, incident response, root cause analysis, and post-incident improvement.
  • Experience working in Agile/DevOps environments and operating production services using ITSM practices where applicable.
  • Prior software development experience—or experience investigating application behavior through code, logs, and distributed traces—is a plus

Responsibilities

  • Own and continuously improve how we use Datadog to make reliability visible and actionable across metrics, logs, traces, dashboards, monitors, alerts, and service-level views.
  • Design, implement, and maintain high-availability, high-throughput, data- and compute-intensive critical systems supporting a growing 24×7 SaaS platform.
  • Partner with service owners to define and improve reliability through meaningful SLIs, SLOs, error budgets, actionable alerting, and operational-readiness practices.
  • Participate in and help drive incident management for production events, serving as an incident commander or technical responder as needed.
  • Coordinate triage, service restoration, escalation, communication, incident documentation, root cause analysis, and completion of corrective actions.
  • Partner with development teams to investigate issues across the infrastructure and application layers using metrics, logs, distributed traces, and code-level context.
  • Improve deployment safety and service resilience through automated validation, recovery and rollback capabilities, reliability testing, and analysis of system failure modes.
  • Partner with other technical leaders to ensure all newly introduced systems are supportable and maintainable by both development and operations.
  • Provide escalated technical guidance and support to other technology teams throughout the organization.
  • Provide on-call coverage for production support and other duties as required.
  • Ensure supported systems and operational activities comply with organizational security, HIPAA, and operating policies.
  • Identify and eliminate repetitive operational toil using Bash, PowerShell, Python, or Ansible. Manage infrastructure as code using Terraform/OpenTofu and configuration automation using Ansible.

Skills

Linux fundamentals
Incident response
On-call rotations
Agile/DevOps
Distributed systems

Education

BS degree in IS/Engineering

Tools

Datadog
Prometheus
Grafana
New Relic
Terraform/OpenTofu
Ansible
Bash
PowerShell
Python

Job description

TherapyNotes LLC is seeking a Site Reliability Engineer to strengthen the reliability and operability of our 24×7 SaaS environment, collaborating with software, infrastructure, security, and database teams. You will improve observability, automate deployments, and drive incident response while aligning with SLIs/SLOs and ITSM processes.

You will design high-availability systems, own critical metrics, and participate in on-call rotations, incident management, and post-incident reviews to prevent

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer: Observability, Security & Growth
Site Reliability Engineer: Observability, Security & Growth

tennr • New York (NY)

On-site
USD 120,000 - 180,000
Unlimited PTO
Health benefits
401(k) match
+3
Senior Site Reliability Engineer - Cloud & Observability
Senior Site Reliability Engineer - Cloud & Observability

Innovaccer • Dallas (TX)

On-site
USD 110,000 - 140,000
Senior Site Reliability Engineer — Cloud Observability
Senior Site Reliability Engineer — Cloud Observability

Guidehouse • San Antonio (TX)

On-site
USD 106,000 - 176,000
Medical, Rx, Dental & Vision Insurance
401(k) Retirement Plan
Tuition Reimbursement
+2
Site Reliability Engineer: Cloud Platform & Resilience
Site Reliability Engineer: Cloud Platform & Resilience

New York Technology Partners • Chicago (IL)

On-site
USD 120,000 - 160,000
Senior Site Reliability Engineer – Cloud, Kubernetes & Automation
Senior Site Reliability Engineer – Cloud, Kubernetes & Automation

Socure • United States

On-site
USD 160,000 - 180,000
Senior Site Reliability Engineer: Observability & Cloud
Senior Site Reliability Engineer: Observability & Cloud

VBeyond Corporation • Jersey City (NJ)

On-site
USD 100,000 - 260,000
Site Reliability Engineer
Site Reliability Engineer

Evlo AI • Denver (CO)

On-site
USD 120,000 - 180,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

New York Technology Partners • Chicago (IL)

On-site
USD 120,000 - 160,000
Senior Site Reliability Engineer – Scale & Observability
Senior Site Reliability Engineer – Scale & Observability

Inspire Brands, Inc. • Atlanta (GA)

On-site
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

RTS RTech Solutions • Austin (TX), Northern (KY)

Hybrid
USD 110,000 - 160,000