Sr. Site Reliability Engineer New Holmdel, New Jersey

CentralReach, LLC

Holmdel Township, Northern (NJ, KY)

Hybrid

USD 160,000 - 180,000

Full time

2 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Health benefits
PTO and holidays
401(k) matching
Parental leave
Hybrid work model

Job summary

CentralReach is seeking a Sr. SRE to own production reliability and drive modern reliability practices across our platform. You will collaborate with software engineering to define SLOs, build dashboards, and automate observability across multi-environment deployments.

The role requires strong experience with cloud AWS, Kubernetes, and leading tools like Datadog, Prometheus, and Grafana, plus solid CI/CD expertise and scripting in Java, Python, or Go.

Qualifications

  • Experience with monitoring, APM, and observability tools such as Splunk, Prometheus, Datadog, and OpenTelemetry.
  • Strong understanding of CI/CD practices and tools (Jenkins, GitHub Actions, GitLab, Argo, and Kargo).
  • Experience with major cloud providers, preferably AWS, and cloud-native infrastructure concepts.
  • Knowledge of containerization technologies, including Kubernetes and Helm.
  • Proficiency in programming languages such as Java, Python, or Go, and familiarity with .NET development.
  • Solid grasp of Linux, Windows, networking, and cloud concepts.
  • Experience using AI to improve productivity and amplify technical skills.

Responsibilities

  • Own production reliability, including availability, latency, performance, capacity planning, monitoring, and uptime for production environments.
  • Define, maintain, and improve SLOs, SLIs, error budgets, dashboards, and observability practices.
  • Troubleshoot and resolve operational issues affecting reliability and SLOs.
  • Build and automate multi-environment observability capabilities and capacity forecasting.
  • Reduce toil and increase development velocity through automation and continuous improvement.
  • Provide production support including incident, change, and problem management; RCA; runbooks; SOPs.
  • Collaborate with software teams on release management, roadmap planning, and operational readiness.
  • Implement and manage reliability tools such as Datadog, Prometheus, and Grafana.

Skills

Monitoring & Observability
SRE practices
CI/CD
Cloud AWS
Kubernetes
Prometheus
Datadog
Grafana
Jenkins
GitHub Actions

Tools

Prometheus
Datadog
Grafana
Jenkins
GitHub Actions
GitLab
Argo
Kubernetes
Helm

Job description

CentralReach is a leading provider of autism and IDD care software for Applied Behavior Analysis (ABA), multidisciplinary therapy, and special education. Trusted by more than 200,000 users, we enable therapy providers, educators, and employers to scale the way they deliver ABA and related therapies with innovative technology, market-leading industry expertise, and world-class customer satisfaction.

ThePlatform Engineeringgroup atCentralReachbuilds the underlying technologies that power our Public and Private Cloud Platforms worldwide. The group is responsible for storage, data infrastructure, IT, observability systems, DevOps, SRE, provisioning, compute, orchestration platform, internal tools, internal platforms (laptops, networks, systems etc.) and services - all the components that make up theCentralReachPlatform.

If you have a passion for the future, enjoy and thrive in an agile, fast-moving, ever-changing startup environment, welcome and take on technical challenges of all shapes and sizes, have excellent interpersonal skill and sense of humor and enjoy rolling up your sleeves and jumping in, then read on!

As a Sr. SRE, you will work closely with the key stakeholders in Software Engineering to drive adoption of modern reliability practices like SLOs, error budget policies, actionable alerts, incident retrospectives, chaos testing, and end-to-end ownership.

Key Accountabilities:

  • Own production reliability, including availability, latency, performance, capacity planning, monitoring, emergency response, and uptime for production environments.
  • Define,maintain, and improve SLOs, SLIs, error budgets, actionable dashboards, and observability practices.
  • Analyze, troubleshoot, and resolve operational issues that affect service reliability and SLO performance.
  • Build and automate multi-environment observability capabilities, including capacity forecasting based on usage patterns.
  • Reduce toil and increase development velocity through automation and continuous improvement.
  • Provide production support, including incident, change, and problem management; root cause analysis; service restoration; runbooks; and standard operating procedures.
  • Identifydata-driven opportunities to improve system architecture, availability, performance, and reliability.
  • Collaborate with software engineering teams on release management, roadmap planning, and operational readiness.
  • Implement and manage reliability and observability tools such as Datadog, Prometheus, and Grafana.

Desired Skills and Experience:

  • Experience with monitoring, APM, and observability tools such as Splunk, Prometheus, Datadog, andOpenTelemetry.
  • Experience implementing observability strategies for logs, metrics, and traces.
  • Strong understanding of CI/CD practices and tools such as Jenkins, GitHub Actions, GitLab, Argo, and Kargo.
  • Strong understanding of major cloud providers, preferably AWS, and cloud-native infrastructure concepts.
  • Strong understanding of containerization technologies, including Kubernetes and Helm.
  • Experience with one or more programming languages, such as Java, Python, or Go, and familiarity with .NET application development.
  • Strong understanding of Linux, Windows, software development, systems, networking, and cloud concepts.
  • Experience using AI to improve productivity and amplify technical skills.

Base Salary Range

$160,000 - $180,000 USD

Backed by Roper Technologies, Inc. (Nasdaq: ROP), CentralReach is entering an exciting phase of growth, innovation, and scale.

Recognized as one of the best places to work over 10 times by organizations such as Inc, Built In, and NJBIZ, our culture is centered around impact, inclusion, and flexibility. As a hybrid company with collaborative offices in Ft. Lauderdale, FL; Holmdel, NJ; and Verona, Italy, we foster a workplace where top talent can thrive and make a real difference in the lives of those we serve.
We offer competitive compensation, comprehensive health benefits, generous PTO, 401(k) matching, and paid parental leave to our full-time employees. Our team members also enjoy hybrid work schedules, career development support, wellness programs, and opportunities to give back through CR Cares, our community engagement initiative.

Protecting your information is important to us. Please take a moment to review our Notice of Privacy Practices for Job Applicants. Applicant-Privacy-Notice-110725 to understand how we collect, use, store, and protect your personal information during the recruitment process.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

CentralReach • Fort Lauderdale (FL)

Hybrid
USD 160,000 - 180,000
Hybrid work model
Health benefits
PTO & 401(k) matching
+1
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

CentralReach • Holmdel Township (NJ)

Hybrid
USD 160,000 - 180,000
Health benefits
PTO
401(k) matching
+2
Site Reliability Engineer (SRE), Data Products New Holmdel, New Jersey
Site Reliability Engineer (SRE), Data Products New Holmdel, New Jersey

CentralReach, LLC • Holmdel Township (NJ)

On-site
USD 135,000 - 160,000
Competitive compensation
Health benefits
Generous PTO
+6
Site Reliability Engineer (SRE), Data Products
Site Reliability Engineer (SRE), Data Products

Centralreach-8 • Holmdel Township (NJ)

On-site
USD 135,000 - 160,000
Hybrid work model
Comprehensive health benefits
Generous PTO
+3
Sr. Software Engineer, Ruby on Rails New Remote - US
Sr. Software Engineer, Ruby on Rails New Remote - US

CentralReach, LLC • Northern (KY)

Remote
USD 120,000 - 180,000
Health benefits
401(k) matching
Paid parental leave
+2
Sr. Software Engineer, Ruby on Rails
Sr. Software Engineer, Ruby on Rails

Far Coder • Northern (KY)

Hybrid
USD 120,000 - 180,000
Hybrid work model
Health benefits
PTO & 401(k) matching
Sr. Software Engineer, Ruby on Rails
Sr. Software Engineer, Ruby on Rails

Centralreach-8 • Holmdel Township (NJ)

Hybrid
USD 120,000 - 180,000
Hybrid work model
Health benefits
401(k) matching
+1
Sr. Quality Engineer New Remote - US
Sr. Quality Engineer New Remote - US

CentralReach, LLC • Northern (KY)

Hybrid
USD 115,000 - 140,000
Health benefits
Generous PTO
401(k) matching
+1
Sr. Software Engineer, Ruby on Rails
Sr. Software Engineer, Ruby on Rails

CentralReach • United States

Hybrid
USD 120,000 - 180,000
competitive compensation
comprehensive health benefits
generous PTO
+6
Sr. Quality Engineer
Sr. Quality Engineer

Centralreach-8 • Holmdel Township (NJ)

Hybrid
USD 115,000 - 140,000
Health benefits
Generous PTO
401(k) matching
+2