Senior SRE - Scale & Reliability for Cloud Platforms

CentralReach

Holmdel Township (NJ)

Hybrid

USD 160,000 - 180,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Health benefits
PTO
401(k) matching
Parental leave
Hybrid work schedules

Job summary

CentralReach, a leading ABA and IDD care software provider, seeks a Sr. SRE to own production reliability across our global platforms. You will drive SLOs, error budgets, dashboards, and end-to-end observability alongside Software Engineering stakeholders in a hybrid work environment.

Responsibilities include incident response, capacity planning, release readiness, and automation to reduce toil. Proficiency in AWS, Kubernetes, and CI/CD is required; base salary ranges from $160,000 to $180,000

Qualifications

  • Experience with monitoring, APM, and observability tools such as Splunk, Prometheus, Datadog, and OpenTelemetry.
  • Experience implementing observability strategies for logs, metrics, and traces.
  • Strong understanding of CI/CD practices and tools such as Jenkins, GitHub Actions, GitLab, Argo, and Kargo.
  • Strong understanding of major cloud providers, preferably AWS, and cloud-native infrastructure concepts.
  • Strong understanding of containerization technologies, including Kubernetes and Helm.
  • Experience with one or more programming languages, such as Java, Python, or Go, and familiarity with .NET application development.
  • Strong understanding of Linux, Windows, software development, systems, networking, and cloud concepts.
  • Experience using AI to improve productivity and amplify technical skills.

Responsibilities

  • Own production reliability, including availability, latency, performance, capacity planning, monitoring, emergency response, and uptime for production environments.
  • Define, maintain, and improve SLOs, SLIs, error budgets, actionable dashboards, and observability practices.
  • Analyze, troubleshoot, and resolve operational issues that affect service reliability and SLO performance.
  • Build and automate multi-environment observability capabilities, including capacity forecasting based on usage patterns.
  • Reduce toil and increase development velocity through automation and continuous improvement.
  • Provide production support, including incident, change, and problem management; root cause analysis; service restoration; runbooks; and standard operating procedures.
  • Identify data-driven opportunities to improve system architecture, availability, performance, and reliability.
  • Collaborate with software engineering teams on release management, roadmap planning, and operational readiness.
  • Implement and manage reliability and observability tools such as Datadog, Prometheus, and Grafana.

Skills

Monitoring & observability
CI/CD
Cloud platforms (AWS)
Kubernetes
Programming languages (Java, Python,Go

Tools

Splunk
Prometheus
Datadog
OpenTelemetry
Jenkins
GitHub Actions
GitLab
Argo
Kargo
Kubernetes
Helm

Job description

CentralReach, a leading ABA and IDD care software provider, seeks a Sr. SRE to own production reliability across our global platforms. You will drive SLOs, error budgets, dashboards, and end-to-end observability alongside Software Engineering stakeholders in a hybrid work environment.

Responsibilities include incident response, capacity planning, release readiness, and automation to reduce toil. Proficiency in AWS, Kubernetes, and CI/CD is required; base salary ranges from $160,000 to $180,000

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: Observability, Cloud & Reliability Leader
Senior SRE: Observability, Cloud & Reliability Leader

CentralReach, LLC • Holmdel Township (NJ), Northern (KY)

Hybrid
USD 160,000 - 180,000
Health benefits
PTO and holidays
401(k) matching
+2
Senior SRE: Cloud Reliability, Observability Lead
Senior SRE: Cloud Reliability, Observability Lead

CentralReach • Fort Lauderdale (FL)

Hybrid
USD 160,000 - 180,000
Hybrid work model
Health benefits
PTO & 401(k) matching
+1
Data Platforms SRE — Cloud Reliability & Automation
Data Platforms SRE — Cloud Reliability & Automation

Centralreach-8 • Holmdel Township (NJ)

On-site
USD 135,000 - 160,000
Hybrid work model
Comprehensive health benefits
Generous PTO
+3
Senior SRE: Scale Reliability, Observability & Resilience
Senior SRE: Scale Reliability, Observability & Resilience

Early Warning Services LLC • Scottsdale (AZ)

Hybrid
USD 106,000 - 130,000
Healthcare Coverage
401(k) Retirement Plan
Paid Time Off
+2
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

CentralReach • Holmdel Township (NJ)

Hybrid
USD 160,000 - 180,000
Health benefits
PTO
401(k) matching
+2
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

CentralReach • Fort Lauderdale (FL)

Hybrid
USD 160,000 - 180,000
Hybrid work model
Health benefits
PTO & 401(k) matching
+1
Senior SRE: Cloud, Observability & Resilience
Senior SRE: Cloud, Observability & Resilience

Supernova Technology™ • Chicago (IL)

On-site
USD 130,000 - 170,000
Senior SRE Leader: Cloud, Reliability & Scale
Senior SRE Leader: Cloud, Reliability & Scale

AVG • Tempe (AZ), Northern (KY)

Hybrid
USD 140,000 - 190,000
Senior Production SRE: Cloud & On-Prem Reliability
Senior Production SRE: Cloud & On-Prem Reliability

Weights & Biases • New York (NY)

On-site
USD 140,000 - 180,000
Medical Insurance
Dental Insurance
Vision Insurance
+15
SRE: Scale, Uptime & Observability (Remote)
SRE: Scale, Uptime & Observability (Remote)

WorkOS • Denver (NC)

Remote
USD 175,000 - 250,000