Site Reliability Engineer Central Europe

Zoolatech

Buffalo (NY)

Hybrid

USD 110,000 - 150,000

Full time

28 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Zoolatech is seeking a detail-oriented Site Reliability Engineer to join our Client Technology sub-organization. You will monitor critical systems, participate in on-call incidents, and analyze root causes to improve reliability and observability.

You will collaborate across teams to enhance CI/CD and prevent downtime, contributing to a robust and scalable platform. Ideal candidates will have 3+ years in SRE, strong scripting skills, cloud familiarity, and experience with Docker and Kubernetes.

Qualifications

  • 3+ years of professional experience in Site Reliability Engineering.
  • Understanding of SRE principles with focus on monitoring, alerting, and incident management.
  • Exposure to observability tools such as Prometheus, Datadog, New Relic, or Grafana, and logging platforms like Splunk or Elasticsearch.
  • Proficiency in programming or scripting languages (Python, Go, Bash, or Java) for automation and troubleshooting.
  • Familiarity with cloud platforms (AWS, GCP, or Azure) and services.
  • Understanding of Docker and Kubernetes.

Responsibilities

  • Maintain real-time 'eyes on glass' monitoring dashboards to detect anomalies in performance and availability.
  • Participate in on-call rotations to respond to incidents, troubleshoot, mitigate outages, and restore service quickly.
  • Perform deep root cause analyses and document findings to prevent recurrence.
  • Collaborate to refine monitoring, logging, and alerting systems for actionable insights.
  • Write and maintain scripts to automate routine operational tasks and reporting.
  • Support the definition and tracking of SLOs and SLIs for reliability measurement.
  • Assist in improving CI/CD pipelines and workflows to minimize downtime.
  • Create and maintain documentation for monitoring configurations, incident handling, and RCAs.
  • Collaborate with software engineering and infrastructure teams to boost fault tolerance and scalability.

Skills

Monitoring
Incident response
Root cause analysis
Scripting
Cloud platforms
Docker & Kubernetes
Observability tools
Communication

Education

Bachelor’s degree in CS or related field

Tools

Prometheus
Datadog
New Relic
Grafana
Splunk
Elasticsearch
Python
Go
Bash
Java
Docker
Kubernetes
AWS
GCP
Azure

Job description

Our clientis a leading U.S. fashion retailer, offering apparel, footwear, beauty, and home goods. It operates 350+ stores and robust online platforms, combining in-store and digital experiences.

Client Technology sub-organization is committed to delivering reliable and scalable systems that power critical services for our customers. We are seeking a motivated and detail-oriented Site Reliability Engineer (SRE) to join our team with a strong focus on proactive monitoring, incident response, and root cause analysis. This role is ideal for someone passionate about ensuring system stability and performance, while diving deep technically to understand and resolve issues when incidents occur.

As an SRE, you will play a key role in maintaining "eyes on glass" monitoring to detect and respond to system anomalies, ensuring the health and reliability of our services. You will also collaborate with teams to address root causes of incidents and continuously improve observability and reliability processes.

Monitor critical systems: Maintain real-time "eyes on glass" monitoring dashboards to proactively identify and respond to anomalies in system performance and availability.

Incident response: Participate in on-call rotations to respond to incidents, troubleshoot issues, mitigate outages, and restore service as quickly as possible.

Root cause analysis: Dive deep into technical investigations to identify the underlying causes of incidents, documenting findings and working with teams to prevent recurrence.

Observability enhancement: Collaborate with teams to refine monitoring, logging, and alerting systems to provide actionable insights and reduce time-to-detection and resolution.

Automation: Write and maintain scripts to automate routine operational tasks, incident remediation, and reporting.

SLOs and SLIs: Support the definition and tracking of Service Level Objectives (SLOs) and Service Level Indicators (SLIs) to measure and improve system reliability.

System optimization: Assist in improving CI/CD pipelines and workflows to ensure seamless deployments and minimize downtime.

Documentation: Create and maintain detailed documentation for monitoring configurations, incident handling procedures, and root cause analysis findings.

Collaboration: Work closely with software engineering and infrastructure teams to improve fault tolerance, scalability, and operational readiness.

3+ years of professional experience in Site Reliability Engineering

Bachelor’s degree in computer science, engineering, or a related field, or equivalent practical experience.

Understanding of site reliability engineering principles, with a strong focus on monitoring, alerting, and incident management.

Exposure to observability tools such as Prometheus, Datadog, New Relic, or Grafana, and logging platforms like Splunk or Elasticsearch.

Proficiency in one or more programming or scripting languages (e.g., Python, Go, Bash, or Java) to assist with automation and troubleshooting.

Familiarity with cloud platforms (AWS, Google Cloud Platform (GCP), or Azure) and their services.

Understanding of containerization and orchestration technologies like Docker and Kubernetes.

Strong analytical skills with the ability to dive deep into technical issues to identify and resolve root causes.

Excellent communication skills for incident reporting, documentation, and collaboration with cross-functional teams.

A proactive mindset and attention to detail, with a willingness to learn and grow in a fast-paced, collaborative environment.

Explore similar open positions that match your experience.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Brooksource • San Antonio (TX)

On-site
USD 80,000 - 120,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 240,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Methodic • San Francisco (CA)

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Harvey Nash • United States

Remote
USD 120,000 - 150,000
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

On-site
USD 130,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mission Staffing • New York (NY)

On-site
USD 140,000 - 200,000
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

Talentify • Houston (TX)

On-site
USD 140,000 - 180,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

State of Wisconsin Investment Board • Madison (WI)

On-site
USD 140,000 - 180,000