Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
Zoolatech is seeking a detail-oriented Site Reliability Engineer to join our Client Technology sub-organization. You will monitor critical systems, participate in on-call incidents, and analyze root causes to improve reliability and observability.
You will collaborate across teams to enhance CI/CD and prevent downtime, contributing to a robust and scalable platform. Ideal candidates will have 3+ years in SRE, strong scripting skills, cloud familiarity, and experience with Docker and Kubernetes.
Our clientis a leading U.S. fashion retailer, offering apparel, footwear, beauty, and home goods. It operates 350+ stores and robust online platforms, combining in-store and digital experiences.
Client Technology sub-organization is committed to delivering reliable and scalable systems that power critical services for our customers. We are seeking a motivated and detail-oriented Site Reliability Engineer (SRE) to join our team with a strong focus on proactive monitoring, incident response, and root cause analysis. This role is ideal for someone passionate about ensuring system stability and performance, while diving deep technically to understand and resolve issues when incidents occur.
As an SRE, you will play a key role in maintaining "eyes on glass" monitoring to detect and respond to system anomalies, ensuring the health and reliability of our services. You will also collaborate with teams to address root causes of incidents and continuously improve observability and reliability processes.
Monitor critical systems: Maintain real-time "eyes on glass" monitoring dashboards to proactively identify and respond to anomalies in system performance and availability.
Incident response: Participate in on-call rotations to respond to incidents, troubleshoot issues, mitigate outages, and restore service as quickly as possible.
Root cause analysis: Dive deep into technical investigations to identify the underlying causes of incidents, documenting findings and working with teams to prevent recurrence.
Observability enhancement: Collaborate with teams to refine monitoring, logging, and alerting systems to provide actionable insights and reduce time-to-detection and resolution.
Automation: Write and maintain scripts to automate routine operational tasks, incident remediation, and reporting.
SLOs and SLIs: Support the definition and tracking of Service Level Objectives (SLOs) and Service Level Indicators (SLIs) to measure and improve system reliability.
System optimization: Assist in improving CI/CD pipelines and workflows to ensure seamless deployments and minimize downtime.
Documentation: Create and maintain detailed documentation for monitoring configurations, incident handling procedures, and root cause analysis findings.
Collaboration: Work closely with software engineering and infrastructure teams to improve fault tolerance, scalability, and operational readiness.
3+ years of professional experience in Site Reliability Engineering
Bachelor’s degree in computer science, engineering, or a related field, or equivalent practical experience.
Understanding of site reliability engineering principles, with a strong focus on monitoring, alerting, and incident management.
Exposure to observability tools such as Prometheus, Datadog, New Relic, or Grafana, and logging platforms like Splunk or Elasticsearch.
Proficiency in one or more programming or scripting languages (e.g., Python, Go, Bash, or Java) to assist with automation and troubleshooting.
Familiarity with cloud platforms (AWS, Google Cloud Platform (GCP), or Azure) and their services.
Understanding of containerization and orchestration technologies like Docker and Kubernetes.
Strong analytical skills with the ability to dive deep into technical issues to identify and resolve root causes.
Excellent communication skills for incident reporting, documentation, and collaboration with cross-functional teams.
A proactive mindset and attention to detail, with a willingness to learn and grow in a fast-paced, collaborative environment.
Explore similar open positions that match your experience.