Senior Site Reliability Engineer

Allegro.eu S.A.

Warszawa

Hybrid

PLN 260,000 - 460,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Flexible hours
RSUs
Annual bonus
Well-located offices
MacBook Pro
Fringe benefits cafeteria
English classes
Training budget
Volunteer day
Social events

Job summary

Allegro.eu S.A. is seeking a senior reliability engineer to lead large-scale system stability initiatives across a vast distributed architecture. You will drive the design of complex reliability platforms, mentor engineers, and own critical CI/CD and IaC processes.

A focus on Chaos Engineering and AI-driven incident response will be central to this role. You will work in a hybrid model with flexible hours, participate in ownership-driven development, and help scale Allegro's infrastructure and

Qualifications

  • Solid engineering background with distributed systems and microservices architecture.
  • Proficiency in at least one programming language used for infrastructure automation, tooling, and backend services (Kotlin, Java, Python or Go).
  • Hands-on experience with containerized environments and orchestration tools (Kubernetes and Docker).
  • Experience with Chaos Engineering and performance testing frameworks (Gatling or similar).
  • Deep experience configuring monitoring and observability stacks (Prometheus, Grafana, ELK).
  • Strong knowledge of CI/CD pipelines and infrastructure as code (IaC).
  • Data-driven mindset, mentoring ability, and sharing code quality standards.

Responsibilities

  • Lead and coordinate team-level projects to improve system reliability across a massive distributed architecture.
  • Drive design and implementation of complex reliability solutions and internal platforms.
  • Design and build an AI agent to enhance incident response and act as senior Incident Commander during outages.
  • Develop and scale self-service performance testing with real-time observability.
  • Build automated resiliency platform for fault injection and quarterly DC failovers.
  • Own reliability of team systems, manage technical debt, automate infrastructure.

Skills

Distributed systems
Kubernetes
Docker
Kotlin/Java/Python/Go
Chaos Engineering
Gatling
Prometheus/Grafana/ELK
CI/CD
IaC
Mentoring

Tools

Kubernetes
Docker

Job description

In your daily work, you will handle the following tasks:
  • Leading and coordinating team-level projects to improve system reliability across our massive distributed architecture.
  • Driving the design and implementation of complex reliability solutions and internal platforms, heavily utilizing our core ecosystem.
  • Designing and building an AI agent to revolutionize our incident response workflow, alongside acting as a senior Incident Commander during major outages and leading blameless post-mortems.
  • Developing and scaling our self-service performance testing platform, empowering engineering teams to run large-scale, distributed load tests with real-time observability.
  • Building an automated resiliency platform for continuous fault injection and driving high-stakes, quarterly full Data Center failover experiments.
  • Owning the reliability of the team's systems, managing technical debt, and automating infrastructure.
  • Mentoring less experienced engineers and setting the standard for team code quality.
  • Influencing how the team communicates and makes decisions on difficult system architecture problems at the Allegro scale.
We are looking for people with:
  • Solid Engineering Background with a deep understanding of distributed systems and microservices architecture.
  • Proficiency in at least one programming language used for infrastructure automation, tooling, and backend services (Kotlin, Java, Python or Go).
  • Hands-on, advanced experience managing containerized environments and orchestration tools, specifically Kubernetes and Docker, alongside networking and service mesh concepts.
  • Hands-on experience with Chaos Engineering practices (e.g., fault injection, failure scenarios) and performance testing frameworks (practical knowledge of Gatling or similar tooling is a plus).
  • Deep experience configuring monitoring and observability stacks (e.g., Prometheus, Grafana, ELK), paired with a proven ability to identify single points of failure in complex systems and act as an Incident Commander during major outages.
  • Strong practical knowledge of building and maintaining CI/CD pipelines and managing infrastructure using IaC tools.
  • A strong sense of ownership and a data-driven mindset, with a track record of mentoring less experienced engineers, setting code quality standards, and shaping team architecture decisions.
What's in it for you:
  • Flexible working hours in the hybrid model (4/1) - working hours start between 7:00 a.m. and 10:00 a.m. We also have 30 days of occasional remote work.
  • Long term discretionary incentive plan based on Allegro.eu shares (restricted stock units).
  • Annual bonus based on your annual performance and company results.
  • Well-located offices (with e.g. fully equipped kitchens, bicycle parking, terraces full of greenery) and excellent work tools (e.g., raised desks, ergonomic chairs, interactive conference rooms).
  • A 16\" or 14\" MacBook Pro or corresponding Dell with Windows (if you don't like Macs) and all the necessary accessories.
  • A wide selection of fringe benefits in a cafeteria plan - you choose what you like (e.g., medical, sports or lunch packages, insurance, purchase vouchers).
  • English classes that we pay for related to the specific nature of your job.
  • A training budget, inter-team tourism, hackathons, and an internal learning platform where you will find multiple trainings.
  • An additional day off for volunteering, which you can use alone, with a team, or with a larger group of people connected by a common goal.
  • Social events for Allegro people - Spin Kilometers, Family Day, Fat Thursday, Advent of Code, and many other occasions we enjoy.

And that's just the beginning! You can read more about the benefits here .

#goodtobehere means that:
  • You will join a team you can count on - we work with top-class specialists who have knowledge- and experience-sharing in their DNA.
  • You will love our level of autonomy in team organization, the space for continuous development, and the opportunity to try new things.
  • You get to choose which technology solves the problem and you are responsible for what you create.
  • You will value our Developer Experience and the full platform of tools and technologies that make creating software easier. We rely on an internal ecosystem based on self-service and widely used tools such as Kubernetes, Docker, Consul, GitHub, and GitHub Actions. Thanks to this, you can contribute to Allegro from your very first days on the job.
  • You will be equipped with modern AI tools to automate repetitive tasks, allowing you to focus on developing new services and refining existing ones (also leveraging AI support).
  • You will create solutions that will be used (and loved!) by your friends, family and millions of our customers.
  • You will meet the Allegro Scale, which starts with over 1000 microservices, an open-source data bus (Hermes) with 300K+ rps, a Service Mesh with 1M+ rps, tens of petabytes of data, and production-used machine learning.
  • You will become part of Allegro Tech - We speak at industry conferences, cooperate with tech communities, run our own blog (it's been over 10 years!), record podcasts, lead guilds, and we organize our own internal conference - the Allegro Tech Meeting. We create solutions we love (and can) to talk about!
Don’t wait until you join us! Let's meet online!

Get to know our team, take a peek at our office life and check out what else we do at Allegro.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Allegro • Warszawa

Hybrid
PLN 260,000 - 380,000
Flexible hybrid model (4/1)
Share-based incentive plan
Annual bonus
Senior Technology Portfolio Manager
Senior Technology Portfolio Manager

Allegro.eu S.A. • Poland

Hybrid
PLN 320,000 - 520,000
Flexible hours
Remote work 30 days/year
Annual bonus
+8
Senior Software Engineer (Java / Kotlin) - Product Page
Senior Software Engineer (Java / Kotlin) - Product Page

Allegro • Toruń

Hybrid
PLN 210,000 - 3,469,000
Annual bonus
RSUs (Allegro shares)
MacBook Pro or Dell laptop
+4
Senior Data Engineer (Scala)
Senior Data Engineer (Scala)

Allegro • Warszawa

Hybrid
PLN 240,000 - 420,000
Well-located offices
MacBook Pro / Dell laptop
Cafeteria plan benefits
+4
Engineering Manager
Engineering Manager

Allegro • Warszawa

Hybrid
PLN 320,000 - 480,000
Annual bonus
RSUs / long-term incentive plan
Hybrid work model (4/1) with 30 days遠
+2
Product Manager - Communication Automation
Product Manager - Communication Automation

Allegro • Warszawa

Hybrid
PLN 180,000 - 280,000
Hybrid work model
Annual bonus
RSUs
+1
Product Manager - Communication Automation
Product Manager - Communication Automation

Allegro • Kraków

Hybrid
PLN 240,000 - 320,000
Hybrid work model (4/1)
Discretionary incentive plan (Allegro)
Annual bonus
+7
Senior Manager, Data
Senior Manager, Data

Allegro.eu S.A. • Warszawa

On-site
PLN 450,000 - 750,000
Flexible hybrid model
RSUs discretionary incentive plan
Annual bonus
+7
Mid Cybersecurity Engineer - CSIRT
Mid Cybersecurity Engineer - CSIRT

Allegro.eu S.A. • Warszawa

Hybrid
PLN 260,000 - 380,000
MacBook Pro
Fringe benefits (cafeteria plan)
English classes
+4
Senior Scala Engineer
Senior Scala Engineer

JobCubby • Poland

Hybrid
PLN 205,000 - 284,000