Service Reliability Engineer (K/M)

HeadHR

Wrocław

Hybrid

PLN 180,000 - 240,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

HeadHR in Wrocław, Poland, is seeking a Service Reliability Engineer to define SLIs/SLOs and improve reliability across Java/Spring Boot, Python and Node.js microservices. You will optimize latency and memory, perform JVM diagnostics, and implement circuit breakers and retries.

You will build Datadog observability and operate services on AWS with EKS/ECS, API Gateway, RDS and more. Strong Linux knowledge and CI/CD experience are essential.

Qualifications

  • At least 3 years of experience as a Service Reliability Engineer.
  • Strong hands-on experience with Java and Spring Boot.
  • Good proficiency in Python or another backend programming language.
  • Knowledge of AWS, including EKS, ECS, IAM, VPC, RDS, S3 and API Gateway.
  • Hands-on Datadog experience with APM, Logs, dashboards, SLOs, RUM and Synthetic Monitoring.
  • Docker, Kubernetes, SQL Server or another SQL dialect, and MongoDB indexes, replica sets and performance tuning.
  • CI/CD with GitHub Actions, Jenkins and GitLab CI; queues and messaging with SQS, SNS or Kafka.
  • Basic Linux and networking fundamentals, including HTTP, TLS, DNS and load balancing.

Responsibilities

  • Define SLIs and SLOs and improve reliability across Java/Spring Boot, Python and Node.js microservices.
  • Optimize latency, throughput, memory and GC; perform JVM diagnostics; implement circuit breakers, retries and graceful degradation.
  • Build Datadog observability covering metrics, logs, traces, synthetics, RUM, APM and dashboards.
  • Triage incidents, restore services, communicate proactively, lead post-incident reviews, perform root cause analysis and track remediation.
  • Operate services on EKS and ECS and manage AWS API Gateway, Lambda, RDS, IAM, S3, VPC, CloudWatch, ALB and NLB.
  • Maintain CI/CD pipelines using GitHub Actions, Jenkins and GitLab CI; build Python and shell automation.
  • Troubleshoot MSSQL/SQL and MongoDB, including replica sets, indexes and PITR, and ensure reliable backup, restore, failover and performance tuning.

Skills

Java
Spring Boot
Python

Tools

Datadog
Docker
Kubernetes
GitHub Actions
Jenkins
GitLab CI
SQS
SNS
Kafka
SQL Server
MongoDB
AWS

Job description

Responsibilities:
  • Define SLIs and SLOs and improve reliability across Java/Spring Boot, Python and Node.js microservices.
  • Optimize latency, throughput, memory and GC, perform JVM diagnostics, and implement circuit breakers, retries and graceful degradation.
  • Build Datadog observability covering metrics, logs, traces, synthetics, RUM, APM and dashboards.
  • Triage incidents, restore services, communicate proactively, lead post-incident reviews, perform root cause analysis and track remediation.
  • Operate services on EKS and ECS and manage AWS API Gateway, Lambda, RDS, IAM, S3, VPC, CloudWatch, ALB and NLB.
  • Maintain CI/CD pipelines using GitHub Actions, Jenkins and GitLab CI; build Python and shell automation.
  • Troubleshoot MSSQL/SQL and MongoDB, including replica sets, indexes and PITR, and ensure reliable backup, restore, failover and performance tuning.
Formal details:
  • Employment contract; start date to be agreed.
  • Position available since October 2026 in Wrocław at Silver Tower Office Centre.
Kogo poszukujemy?
Key requirements:
  • At least 3 years of experience as a Service Reliability Engineer.
  • Strong hands-on experience with Java and Spring Boot.
  • Good proficiency in Python or another backend programming language.
  • Knowledge of AWS, including EKS, ECS, IAM, VPC, RDS, S3 and API Gateway.
  • Hands-on Datadog experience with APM, Logs, dashboards, SLOs, RUM and Synthetic Monitoring.
  • Docker, Kubernetes, SQL Server or another SQL dialect, and MongoDB indexes, replica sets and performance tuning.
  • CI/CD with GitHub Actions, Jenkins and GitLab CI; queues and messaging with SQS, SNS or Kafka.
  • Basic Linux and networking fundamentals, including HTTP, TLS, DNS and load balancing.
Additional advantage:
  • Previous experience in the automotive industry.
Toyota DNA:
  • Courage to pursue challenging targets.
  • Creativity and innovative thinking.
  • Coaching through knowledge and feedback sharing.
  • Curiosity and fact-based observation.
  • Respectful, inclusive collaboration and customer orientation.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

EV Backend Product Lead (K/M)
EV Backend Product Lead (K/M)

HeadHR • Wrocław

Hybrid
PLN 180,000 - 240,000
Health insurance
Sport card
Lunch subsidy
+3
Support Engineer
Support Engineer

HeadHR • Wrocław

Hybrid
PLN 169,000 - 330,000
Hybrid work in Wroclaw
Equity incentives
Annual bonus possibility
+4
Senior Frontend Developer - Android (K/M)
Senior Frontend Developer - Android (K/M)

HeadHR • Wrocław

Hybrid
PLN 120,000 - 180,000
Office in Wroclaw
Software Tester (K/M)
Software Tester (K/M)

HeadHR • Wrocław

Hybrid
PLN 120,000 - 180,000
Ubezpieczenie zdrowotne
Karta sportowa
Dofinansowanie posiłków
+3
System & Data Engineer (Linux / SQL Area)
System & Data Engineer (Linux / SQL Area)

HeadHR • Łódź

On-site
PLN 120,000 - 180,000
Attractive salary
Exposure to new technologies
Internal and external training access
+2
Senior Java Developer
Senior Java Developer

HeadHR • Kraków

Hybrid
PLN 180,000 - 280,000
Fullstack Software Engineer
Fullstack Software Engineer

HeadHR • Kraków

Hybrid
PLN 180,000 - 300,000
Senior Fullstack Engineer
Senior Fullstack Engineer

HeadHR • Kraków

On-site
PLN 120,000 - 160,000
Technical Lead
Technical Lead

HeadHR • Kraków

On-site
PLN 200,000 - 320,000
Senior Fullstack Developer (Java)
Senior Fullstack Developer (Java)

HeadHR • Poland

Hybrid
PLN 180,000 - 300,000
Hybrid work (1x weekly in Warsaw)