Site Reliability Engineer: Observability & Incidents

Zoolatech

Buffalo (NY)

Hybrid

USD 110,000 - 150,000

Full time

5 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Zoolatech is seeking a detail-oriented Site Reliability Engineer to join our Client Technology sub-organization. You will monitor critical systems, participate in on-call incidents, and analyze root causes to improve reliability and observability.

You will collaborate across teams to enhance CI/CD and prevent downtime, contributing to a robust and scalable platform. Ideal candidates will have 3+ years in SRE, strong scripting skills, cloud familiarity, and experience with Docker and Kubernetes.

Qualifications

  • 3+ years of professional experience in Site Reliability Engineering.
  • Understanding of SRE principles with focus on monitoring, alerting, and incident management.
  • Exposure to observability tools such as Prometheus, Datadog, New Relic, or Grafana, and logging platforms like Splunk or Elasticsearch.
  • Proficiency in programming or scripting languages (Python, Go, Bash, or Java) for automation and troubleshooting.
  • Familiarity with cloud platforms (AWS, GCP, or Azure) and services.
  • Understanding of Docker and Kubernetes.

Responsibilities

  • Maintain real-time 'eyes on glass' monitoring dashboards to detect anomalies in performance and availability.
  • Participate in on-call rotations to respond to incidents, troubleshoot, mitigate outages, and restore service quickly.
  • Perform deep root cause analyses and document findings to prevent recurrence.
  • Collaborate to refine monitoring, logging, and alerting systems for actionable insights.
  • Write and maintain scripts to automate routine operational tasks and reporting.
  • Support the definition and tracking of SLOs and SLIs for reliability measurement.
  • Assist in improving CI/CD pipelines and workflows to minimize downtime.
  • Create and maintain documentation for monitoring configurations, incident handling, and RCAs.
  • Collaborate with software engineering and infrastructure teams to boost fault tolerance and scalability.

Skills

Monitoring
Incident response
Root cause analysis
Scripting
Cloud platforms
Docker & Kubernetes
Observability tools
Communication

Education

Bachelor’s degree in CS or related field

Tools

Prometheus
Datadog
New Relic
Grafana
Splunk
Elasticsearch
Python
Go
Bash
Java
Docker
Kubernetes
AWS
GCP
Azure

Job description

Zoolatech is seeking a detail-oriented Site Reliability Engineer to join our Client Technology sub-organization. You will monitor critical systems, participate in on-call incidents, and analyze root causes to improve reliability and observability.

You will collaborate across teams to enhance CI/CD and prevent downtime, contributing to a robust and scalable platform. Ideal candidates will have 3+ years in SRE, strong scripting skills, cloud familiarity, and experience with Docker and Kubernetes.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer Central Europe
Site Reliability Engineer Central Europe

Zoolatech • Buffalo (NY)

Hybrid
USD 110,000 - 150,000
Site Reliability Engineer - Incident & Observability Lead
Site Reliability Engineer - Incident & Observability Lead

Worky • Atlanta (GA)

Hybrid
USD 100,000 - 120,000
Medical & Dental
Hybrid/Remote work
Vision Insurance
+4
Site Reliability Engineer
Site Reliability Engineer

RTS RTech Solutions • Austin (TX), Northern (KY)

On-site
USD 110,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

BlueSky Resource Solutions • Duluth (GA)

On-site
USD 120,000 - 180,000
Site Reliability Engineer II: Reliability & Observability
Site Reliability Engineer II: Reliability & Observability

Early Warning Services LLC • Scottsdale (AZ)

Hybrid
USD 83,000 - 132,000
Healthcare Coverage
401(k) Retirement Plan
Paid Time Off
+1
Site Reliability Engineer: Distributed Systems Observability
Site Reliability Engineer: Distributed Systems Observability

engineeringjobs.net, Inc. • Sunnyvale (CA)

On-site
USD 140,000 - 210,000
Site Reliability Engineer — Global Scale & Observability
Site Reliability Engineer — Global Scale & Observability

Human Ventures, LLC. • United States

On-site
USD 90,000 - 100,000
Site Reliability Engineer – Observability & Automation Lead
Site Reliability Engineer – Observability & Automation Lead

Talentify • Houston (TX)

On-site
USD 140,000 - 180,000
Senior Site Reliability Engineer: Observability & Cloud
Senior Site Reliability Engineer: Observability & Cloud

VBeyond Corporation • Jersey City (NJ)

On-site
USD 100,000 - 260,000
Senior Site Reliability Engineer – Observability & Cloud
Senior Site Reliability Engineer – Observability & Cloud

Cosm Inc. • El Segundo (CA), Northern (KY)

Hybrid
USD 110,000 - 145,000