Senior SRE: Observability, Cloud & Reliability Leader

CentralReach, LLC

Holmdel Township, Northern (NJ, KY)

Hybrid

USD 160,000 - 180,000

Full time

9 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Health benefits
PTO and holidays
401(k) matching
Parental leave
Hybrid work model

Job summary

CentralReach is seeking a Sr. SRE to own production reliability and drive modern reliability practices across our platform. You will collaborate with software engineering to define SLOs, build dashboards, and automate observability across multi-environment deployments.

The role requires strong experience with cloud AWS, Kubernetes, and leading tools like Datadog, Prometheus, and Grafana, plus solid CI/CD expertise and scripting in Java, Python, or Go.

Qualifications

  • Experience with monitoring, APM, and observability tools such as Splunk, Prometheus, Datadog, and OpenTelemetry.
  • Strong understanding of CI/CD practices and tools (Jenkins, GitHub Actions, GitLab, Argo, and Kargo).
  • Experience with major cloud providers, preferably AWS, and cloud-native infrastructure concepts.
  • Knowledge of containerization technologies, including Kubernetes and Helm.
  • Proficiency in programming languages such as Java, Python, or Go, and familiarity with .NET development.
  • Solid grasp of Linux, Windows, networking, and cloud concepts.
  • Experience using AI to improve productivity and amplify technical skills.

Responsibilities

  • Own production reliability, including availability, latency, performance, capacity planning, monitoring, and uptime for production environments.
  • Define, maintain, and improve SLOs, SLIs, error budgets, dashboards, and observability practices.
  • Troubleshoot and resolve operational issues affecting reliability and SLOs.
  • Build and automate multi-environment observability capabilities and capacity forecasting.
  • Reduce toil and increase development velocity through automation and continuous improvement.
  • Provide production support including incident, change, and problem management; RCA; runbooks; SOPs.
  • Collaborate with software teams on release management, roadmap planning, and operational readiness.
  • Implement and manage reliability tools such as Datadog, Prometheus, and Grafana.

Skills

Monitoring & Observability
SRE practices
CI/CD
Cloud AWS
Kubernetes
Prometheus
Datadog
Grafana
Jenkins
GitHub Actions

Tools

Prometheus
Datadog
Grafana
Jenkins
GitHub Actions
GitLab
Argo
Kubernetes
Helm

Job description

CentralReach is seeking a Sr. SRE to own production reliability and drive modern reliability practices across our platform. You will collaborate with software engineering to define SLOs, build dashboards, and automate observability across multi-environment deployments.

The role requires strong experience with cloud AWS, Kubernetes, and leading tools like Datadog, Prometheus, and Grafana, plus solid CI/CD expertise and scripting in Java, Python, or Go.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE: Cloud Reliability, Observability Lead
Senior SRE: Cloud Reliability, Observability Lead

CentralReach • Fort Lauderdale (FL)

Hybrid
USD 160,000 - 180,000
Hybrid work model
Health benefits
PTO & 401(k) matching
+1
Senior SRE - Scale & Reliability for Cloud Platforms
Senior SRE - Scale & Reliability for Cloud Platforms

CentralReach • Holmdel Township (NJ)

Hybrid
USD 160,000 - 180,000
Health benefits
PTO
401(k) matching
+2
Senior SRE Leader: Cloud, Reliability & Scale
Senior SRE Leader: Cloud, Reliability & Scale

AVG • Tempe (AZ), Northern (KY)

Hybrid
USD 140,000 - 190,000
Senior SRE Lead: Reliability, Observability & Cloud
Senior SRE Lead: Reliability, Observability & Cloud

Shieldai • San Mateo (CA)

On-site
USD 180,000 - 240,000
SRE Lead: Production Reliability & Observability Architect
SRE Lead: Production Reliability & Observability Architect

TechDigital Group • Woonsocket (RI)

On-site
USD 140,000 - 190,000
Senior SRE: Observability, Automation & Hybrid Cloud
Senior SRE: Observability, Automation & Hybrid Cloud

Colorado-Public-Employees • Denver (CO)

Hybrid
USD 140,000 - 165,000
Hybrid work option
On-call rotation
Work from home eligibility
Senior SRE: Observability & Cloud Automation Lead
Senior SRE: Observability & Cloud Automation Lead

Cosm • El Segundo (CA)

On-site
USD 110,000 - 145,000
Senior SRE & Reliability Architect (Cloud & Observability)
Senior SRE & Reliability Architect (Cloud & Observability)

Gen • Tempe (AZ)

On-site
USD 150,000 - 160,000
Senior SRE: Cloud Reliability & AI-Driven Incident Triage
Senior SRE: Cloud Reliability & AI-Driven Incident Triage

Clearwater Analytics • Boise (ID)

On-site
USD 130,000 - 170,000
SRE Lead: Reliability & Cloud Observability Architect
SRE Lead: Reliability & Cloud Observability Architect

BlackCube Labs • San Diego (CA)

On-site
USD 190,000 - 280,000