Site Reliability Engineer Poland

Zoolatech

Wrocław

Ibrido

PLN 170.000 - 250.000

Tempo pieno

4 giorni fa
Candidati tra i primi
Generatore di candidature

Ricevi una risposta da questo datore di lavoro — un curriculum e una lettera di presentazione personalizzati, che corrispondono esattamente a ciò che sta cercando.

Supera i filtri ATS

Descrizione del lavoro

Zoolatech in Wrocław, Poland seeks a driven Site Reliability Engineer to ensure system stability and performance. You will maintain real-time monitoring, respond to incidents, and perform in-depth root cause analysis while collaborating with software and infra teams to raise observability and reliability.

The role emphasizes automation, CI/CD improvements, and comprehensive documentation to support fault-tolerant deployments in a fast-paced environment.

Competenze

  • 3+ years of professional experience in Site Reliability Engineering.
  • Understanding of SRE principles with focus on monitoring, alerting, and incident management.
  • Exposure to observability tools such as Prometheus, Datadog, New Relic, Grafana and logging platforms like Splunk or Elasticsearch.
  • Proficiency in programming or scripting languages (e.g., Python, Go, Bash, or Java).
  • Familiarity with cloud platforms (AWS, GCP, or Azure) and containerization like Docker/Kubernetes.

Mansioni

  • Monitor critical systems with real-time dashboards and proactive alerting.
  • Participate in on-call rotations and incident response to mitigate outages.
  • Perform root cause analyses and document findings for prevention.
  • Improve observability, logging, and alerting to reduce MTTR.
  • Automate routine operational tasks and remediation.
  • Assist in defining and tracking SLOs/SLIs.
  • Contribute to CI/CD improvements and deployment workflows.
  • Provide clear incident reporting and cross-team collaboration.

Conoscenze

Monitoring
Incident response
Root cause analysis
Observability
Automation
SLOs/SLIs
CI/CD
Documentation
Collaboration
Programming/scripting

Formazione

Bachelor's degree in CS/Engineering

Strumenti

Prometheus
Datadog
New Relic
Grafana
Splunk/Elasticsearch

Descrizione del lavoro

Our clientis a leading U.S. fashion retailer, offering apparel, footwear, beauty, and home goods. It operates 350+ stores and robust online platforms, combining in-store and digital experiences.

Client Technology sub-organization is committed to delivering reliable and scalable systems that power critical services for our customers. We are seeking a motivated and detail-oriented Site Reliability Engineer (SRE) to join our team with a strong focus on proactive monitoring, incident response, and root cause analysis. This role is ideal for someone passionate about ensuring system stability and performance, while diving deep technically to understand and resolve issues when incidents occur.

As an SRE, you will play a key role in maintaining "eyes on glass" monitoring to detect and respond to system anomalies, ensuring the health and reliability of our services. You will also collaborate with teams to address root causes of incidents and continuously improve observability and reliability processes.

Monitor critical systems: Maintain real-time "eyes on glass" monitoring dashboards to proactively identify and respond to anomalies in system performance and availability.

Incident response: Participate in on-call rotations to respond to incidents, troubleshoot issues, mitigate outages, and restore service as quickly as possible.

Root cause analysis: Dive deep into technical investigations to identify the underlying causes of incidents, documenting findings and working with teams to prevent recurrence.

Observability enhancement: Collaborate with teams to refine monitoring, logging, and alerting systems to provide actionable insights and reduce time-to-detection and resolution.

Automation: Write and maintain scripts to automate routine operational tasks, incident remediation, and reporting.

SLOs and SLIs: Support the definition and tracking of Service Level Objectives (SLOs) and Service Level Indicators (SLIs) to measure and improve system reliability.

System optimization: Assist in improving CI/CD pipelines and workflows to ensure seamless deployments and minimize downtime.

Documentation: Create and maintain detailed documentation for monitoring configurations, incident handling procedures, and root cause analysis findings.

Collaboration: Work closely with software engineering and infrastructure teams to improve fault tolerance, scalability, and operational readiness.

3+ years of professional experience in Site Reliability Engineering

Bachelor’s degree in computer science, engineering, or a related field, or equivalent practical experience.

Understanding of site reliability engineering principles, with a strong focus on monitoring, alerting, and incident management.

Exposure to observability tools such as Prometheus, Datadog, New Relic, or Grafana, and logging platforms like Splunk or Elasticsearch.

Proficiency in one or more programming or scripting languages (e.g., Python, Go, Bash, or Java) to assist with automation and troubleshooting.

Familiarity with cloud platforms (AWS, Google Cloud Platform (GCP), or Azure) and their services.

Understanding of containerization and orchestration technologies like Docker and Kubernetes.

Strong analytical skills with the ability to dive deep into technical issues to identify and resolve root causes.

Excellent communication skills for incident reporting, documentation, and collaboration with cross-functional teams.

A proactive mindset and attention to detail, with a willingness to learn and grow in a fast-paced, collaborative environment.

Explore similar open positions that match your experience.

Ottieni la revisione del curriculum gratis e riservata.

o trascina qui il file.

Similar jobs

Offerte di lavoro simili che vale la pena confrontare

Site Leader
Site Leader

Weekday AI • Polska

Ibrido
PLN 480.000 - 900.000
Site Reliability Engineer
Site Reliability Engineer

Cavendish Professionals • Wrocław

In loco
PLN 150.000 - 190.000
Site Reliability Engineer
Site Reliability Engineer

EPAM Systems • Polonia

In loco
PLN 180.000 - 300.000
Health insurance
Multisport
Relocation support
+2
Site Reliability Engineer
Site Reliability Engineer

SIX • Warszawa

In loco
PLN 240.000 - 360.000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Luxoft • Polonia

In loco
PLN 180.000 - 260.000
Private Medical & Dental care
Life Insurance covered
Internal Mobility program
Poland-Based SRE: Reliable Systems, Observability & Automation
Poland-Based SRE: Reliable Systems, Observability & Automation

Zoolatech • Wrocław

Ibrido
PLN 170.000 - 250.000
Senior Fullstack Engineer Poland
Senior Fullstack Engineer Poland

Zoolatech • Wrocław

Ibrido
PLN 160.000 - 260.000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Cytiva • Kraków

In loco
PLN 260.000 - 380.000
Remote work arrangement
Senior Site Reliability Engineer - Platform Reliability (Resilience)
Senior Site Reliability Engineer - Platform Reliability (Resilience)

Elasticsearch B.V. • Polska

In loco
PLN 359.000 - 465.000
Health coverage
Flexible locations
Vacation days
+3
SRE Java/Python
SRE Java/Python

Infotree Global Solutions Inc. • Warszawa

In loco
PLN 120.000 - 180.000