Engineering Team Leader (Site Reliability Engineering)

XTB online investing

Warszawa

Hybrid

PLN 304,000 - 386,000

Full time

2 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Training budget
Birthday off
Parental leave
Equipment provided
Private medical care

Job summary

XTB online investing, a global investment company, is seeking an Engineering Team Leader to drive the Site Reliability Engineering (SRE) team. You will shape reliability practices and implement scalable observability across millions of users.

Collaborating with infrastructure and development teams, you will mentor engineers, steer end-to-end reliability initiatives, and advance data-driven operations. This hybrid role is based in Warsaw with flexible office options.

Qualifications

  • Several years of experience in SRE, Infrastructure, or DevOps roles managing high-scale, distributed environments.
  • Proven leadership in formal management roles, mentoring and developing high-performing SRE/DevOps teams.
  • Extensive experience building scalable, observable infrastructure systems (Azure, Kubernetes, On-prem).
  • Ability to deliver end-to-end reliability strategies and drive architectural improvements across large projects.
  • Strong cross-functional collaboration with product engineering teams and distributed/remote teams.

Responsibilities

  • Shape and grow a high-performing Site Reliability Engineering team, fostering technical excellence and ownership.
  • Define and drive the SRE platform strategy, ensuring scalability and architectural alignment with business objectives.
  • Own and oversee 24/7 on-call and incident management, tooling, procedures, and post-mortems.
  • Define and track KPIs for team performance; use data to drive reliability improvements.
  • Oversee observability ecosystem design and evolution, telemetry, logging, tracing, and sampling.

Skills

Python
Kubernetes
Ansible
Azure
On-prem
Observability
AIOps
Distributed systems

Tools

Prometheus
Grafana
OTEL
ELK
Tempo
Thanos
Datadog
Splunk
New Relic

Job description

We are building XTB - a global investment company offering innovative technological solutions that allow our clients to effectively manage their finances in multiple ways. All of this within a single, intuitive XTB app already used by over one million users worldwide!
We are a certified Great Place to Work company.
We are looking for an Engineering Team Leader to drive the development and growth of the Site Reliability Engineering team. In this role, you will have the opportunity to shape the technical and operational direction of SRE practices, lead a resilience strategy, and play a key role in defining and delivering solutions that ensure the reliability and scalability of XTB systems for millions of clients across a growing organization.

Responsibilities
  • Team Leadership: Shape and grow a high-performing Site Reliability Engineering team, fostering an environment of technical excellence, ownership, and continuous improvement.
  • Reliability Strategy: Define and drive the SRE platform strategy in close collaboration with infrastructure and development teams, ensuring alignment with organizational business objectives. Oversee reliability strategy across the organization, ensuring consistent architectural alignment, scalability, and the adoption of industry-standard SRE best practices.
  • Incident Management: Own and oversee organization-wide 24/7 on-call and incident management processes. Manage incident tooling, establish and maintain operational procedures, ensure compliance with regulations, oversee reporting, and drive continuous improvement of incident response strategy.
  • Data-Driven Management: Define and track measurable objectives (KPIs) for team performance. Leverage data to drive improvements and build management metrics that provide clear visibility into operational health and team productivity.
  • Observability Engineering: Oversee the design, development, and evolution of the organization-wide observability ecosystem. Lead the strategy for implementing standardized telemetry, including structured logging, distributed tracing, and intelligent sampling.
Requirements
  • Professional Background: Several years of experience in SRE, Infrastructure, or DevOps roles managing high-scale, distributed environments.
  • Leadership Experience: Proven track record in a formal management role, leading, mentoring, and developing high-performing SRE or DevOps engineering teams.
  • Infrastructure & Reliability: Extensive experience building and maintaining scalable, reliable, and observable infrastructure systems (Azure, Kubernetes, On-prem).
  • Reliability Strategy: Demonstrated ability to deliver end-to-end reliability strategies, drive architectural improvements, and manage large-scale technical projects from design to production.
  • Cross-Functional Collaboration: Proven ability to drive cultural change, act as a strategic partner to product engineering teams, and work effectively with distributed/remote teams.
Leadership & Management
  • Execution: Break down complex projects into actionable tasks; drive incremental value.
  • Growth: Mentor and support the professional growth of team members.
  • Collaboration: Facilitate workshops; build technical community; resolve conflicts effectively.
  • Operational Excellence: Drive operational excellence and reliability culture within the team; lead incident management and champion post-mortems.
  • Strategy: Proactively manage technical debt; align team output with organizational goals.
Technical Skills We Expect
  • Programming & Scripting: Strong Python skills for building scalable automation, internal tools, and scripts.
  • Cloud & Orchestration: Expertise in managing Kubernetes, configuration management with Ansible, and designing resilient infrastructure on Azure and on-prem.
  • Observability Engineering: Deep proficiency in building standardized telemetry systems. Mastery of tools like Prometheus, Grafana, OTEL, ELK, Tempo, Thanos, and similar.
  • AI & Automation: Using AI/ML for AIOps, anomaly detection, log analysis, and optimizing reliability workflows.
Nice to have
  • Experience with commercial APM platforms (e.g., Datadog, Splunk, New Relic) and chaos engineering tooling.
  • Proficiency with cloud cost management and FinOps principles.
What We Offer
  • Real influence on the development of the company and the product.
  • Work in an experienced team that is happy to share its knowledge.
  • A clear vision of development thanks to regular feedback and clear career paths.
  • Regular team-building meetings.
Benefits
  • A training budget for courses and conferences that interest you.
  • An extra day off on your birthday.
  • An extra day off for parents.
  • Equipment tailored to your needs.
  • Private medical care and group insurance.
  • Access to an e-learning platform for learning English and a benefits platform.
  • Access to a wellbeing platform and the opportunity to take advantage of workshops and private therapy sessions.
  • Remote work, from the office in Warsaw or from a coworking space in your city.

27,200 zł - 34,600 zł a month
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Engineering Team Leader (Site Reliability Engineering)
Engineering Team Leader (Site Reliability Engineering)

XTB online investing • Poland

Hybrid
PLN 320,000 - 520,000
Remote work
Private medical care
Group insurance
+5
Site Reliability Engineer
Site Reliability Engineer

XTB online investing • Warszawa

Hybrid
PLN 199,000 - 253,000
Training budget for courses
Birthday day off
Parental leave day off
+5
Site Reliability Engineer
Site Reliability Engineer

XTB online investing • Poland

Hybrid
PLN 180,000 - 300,000
Remote work
Office in Warsaw
Birthday leave
+4
Platform Engineering Team Leader
Platform Engineering Team Leader

XTB online investing • Warszawa

Hybrid
PLN 304,000 - 386,000
Private medical care
Group insurance
Remote work flexibility
+1
Senior Java Software Engineer
Senior Java Software Engineer

XTB online investing • Warszawa

Hybrid
PLN 228,000 - 301,000
Training budget for courses
Birthday off
Parental leave
+4
Senior Project Manager
Senior Project Manager

XTB online investing • Warszawa

Hybrid
PLN 199,000 - 253,000
Training budget
Birthday leave
Parental leave
+5
Data Engineer
Data Engineer

XTB online investing • Warszawa

On-site
PLN 156,000 - 199,000
Training budget
Birthday off
Parental leave
+4
Product Manager
Product Manager

XTB online investing • Warszawa

Hybrid
PLN 161,000 - 205,000
Real influence on product development
Knowledge-sharing team
Regular feedback and clear career path
+9
AI Engineer
AI Engineer

XTB online investing • Warszawa

Hybrid
PLN 171,000 - 217,000
Private medical care
Group insurance
Remote work from Warsaw
+4
Software Engineer (DevOps) - Site Reliability Engineer
Software Engineer (DevOps) - Site Reliability Engineer

Revolut • Kraków

On-site
GBP 70,000 - 110,000