Kafka Reliability Engineer

Be | Shaping the Future Poland

Warszawa

Remote

PLN 240,000 - 360,000

Full time

47 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Mindgram access
Free gym at Q22

Job summary

Be | Shaping the Future Poland is seeking a Senior SRE / Platform Engineer with Kafka to support a production readiness initiative for a critical booking platform. You will improve reliability, observability, resilience, and operational excellence across distributed systems and event-driven architectures.

The ideal candidate will have hands-on experience with Kafka-based systems, observability tooling, cloud-native deployments, and SRE practices.

Qualifications

  • Strong background as an SRE, Platform/DevOps engineer or reliability engineer.
  • Hands-on experience with Apache Kafka in production environments.
  • Experience with monitoring, alerting, logging, and observability platforms.
  • Knowledge of distributed systems and event-driven architectures.
  • Experience with Kubernetes and containerized environments.
  • Experience with CI/CD pipelines and deployment automation.
  • Understanding of incident management and operational excellence practices.
  • Experience with MongoDB replication and backup concepts.
  • Strong troubleshooting and problem-solving skills.

Responsibilities

  • Define and implement monitoring strategies based on Golden Signals.
  • Design and maintain dashboards, metrics, alerts, and reporting.
  • Improve centralized logging and distributed tracing capabilities.
  • Develop alerting rules, thresholds, and runbooks.
  • Support on-call processes and incident response activities.
  • Design and optimize Kafka topics, partitioning, and consumer groups.
  • Implement retry mechanisms, DLQ, and idempotent processing.
  • Define schema governance and messaging standards.
  • Monitor Kafka performance, lag, and throughput.
  • Improve reliability of asynchronous workflows and backpressure handling.
  • Define and manage SLIs, SLOs, and error budgets.
  • Participate in incident management and post-mortem activities.
  • Drive reliability-by-design across services and platforms.
  • Identify and implement improvements reducing operational overhead.
  • Support CI/CD pipelines and release automation.
  • Work with Canary/Blue-Green/rollback strategies.
  • Collaborate on load testing and production readiness assessments.
  • Work with Kubernetes and container orchestration platforms.

Skills

SRE/Platform Engineer
Kafka in prod
Observability tooling
Kubernetes
CI/CD pipelines
Incident management

Tools

MongoDB
Docker

Job description

Be | Shaping the Future Poland has a proven position of being a reliable partner for financial services organisations to analyse complex requirements, find solutions and implement them in their entirety, regardless of their complexity. Since the foundation of Be Poland in 2013, we have been continually expanding and customising our spectrum of services. Today, we are privileged to have in our team the best individuals in each sector we operate within the financial services industry.

Role: Senior SRE / Platform Engineer with Kafka

Location: fully remote from Poland

Contract Type: B2B

We are looking for experienced Senior SRE / Platform Engineers with Kafka to support a production readiness initiative for a critical booking processing platform. The assignment focuses on improving reliability, observability, resilience, and operational excellence across distributed systems and event-driven architectures. The ideal candidate combines strong hands-on experience with Kafka-based systems, observability tooling, cloud-native deployments, and Site Reliability Engineering practices.

Key Responsibilities

Observability & Monitoring

  • Define and implement monitoring strategies based on Golden Signals
  • Design and maintain dashboards, metrics, alerts, and reporting
  • Improve centralized logging and distributed tracing capabilities
  • Develop alerting rules, thresholds, and operational runbooks
  • Support on-call processes and incident response activities

Kafka Reliability & Messaging

  • Design and optimize Kafka topics, partitioning strategies, and consumer groups
  • Implement retry mechanisms, dead-letter queues (DLQ), and idempotent processing
  • Define schema governance and messaging standards
  • Monitor Kafka performance, consumer lag, and throughput
  • Improve reliability of asynchronous workflows, including backpressure handling and failure recovery

Reliability Engineering & SRE Practices

  • Define and manage SLIs, SLOs, and error budgets
  • Participate in incident management and post-mortem activities
  • Drive reliability-by-design principles across services and platforms
  • Identify and implement improvements that reduce operational overhead

Release & Deployment

  • Support CI/CD pipelines and deployment automation
  • Implement quality gates and release management processes
  • Work with Canary, Blue-Green, and rollback strategies
  • Support feature flag frameworks and version management
  • Collaborate on load testing and production readiness assessments
  • Work with container orchestration platforms such as Kubernetes

Data Protection & Disaster Recovery

  • Define backup and restore strategies
  • Support disaster recovery planning and testing
  • Contribute to RPO/RTO definitions and operational procedures
  • Participate in DR exercises and resilience testing

Required Skills & Experience

  • Strong experience as an SRE, Platform Engineer, DevOps Engineer, or Reliability Engineer
  • Hands-on experience with Apache Kafka in production environments
  • Experience with monitoring, alerting, logging, and observability platforms
  • Knowledge of distributed systems and event-driven architectures
  • Experience with Kubernetes and containerized environments
  • Experience with CI/CD pipelines and deployment automation
  • Understanding of incident management and operational excellence practices
  • Experience with MongoDB, including replication and backup concepts
  • Strong troubleshooting and problem-solving skills

Nice to Have

  • Experience in financial services or other highly regulated environments
  • Experience with distributed tracing solutions
  • Knowledge of cloud platforms (AWS, Azure, or GCP)
  • Experience implementing SLO/SLI frameworks
  • Certifications related to Kubernetes, cloud technologies, or SRE practices

Our offer:

  • Competitive remuneration on B2B contract
  • Access to Mindgram - mental health & well-being platform
  • Free gym at Q22
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote Senior Kafka Reliability Engineer
Remote Senior Kafka Reliability Engineer

Be | Shaping the Future Poland • Warszawa

Remote
PLN 240,000 - 360,000
Mindgram access
Free gym at Q22
Kafka SRE: Build & Maintain Scalable Streaming Platform
Kafka SRE: Build & Maintain Scalable Streaming Platform

Solid Company • Kraków

Hybrid
PLN 180,000 - 240,000
Kafka Engineer
Kafka Engineer

Mindbox Sp. z o.o. • Wrocław

On-site
PLN 180,000 - 280,000
Flexible cooperation model
Hybrid work setup
InterPolska Health Care
Senior Kafka Engineer / Solution Designer
Senior Kafka Engineer / Solution Designer

Euroclear • Polska

Hybrid
PLN 180,000 - 320,000
Kafka SRE
Kafka SRE

Solid Company • Kraków

Hybrid
PLN 180,000 - 240,000
Senior Backend / Fullstack Developer
Senior Backend / Fullstack Developer

Winged IT • Polska

Hybrid
PLN 220,000 - 248,000
Fully remote cooperation
Occasional business trips to Germany
International project experience
Senior Backend Engineer – Kafka & Cloud (Remote)
Senior Backend Engineer – Kafka & Cloud (Remote)

Winged IT • Polska

Hybrid
PLN 220,000 - 248,000
Fully remote cooperation
Occasional business trips to Germany
International project experience
Senior Staff Software Engineer
Senior Staff Software Engineer

Kontakt.io • Kraków

On-site
PLN 212,404 - 297,366
Equity in Series C company
Private healthcare
Multisport card
+1
Kafka Systems Engineer
Kafka Systems Engineer

BEC Financial Technologies • Warszawa

On-site
PLN 180,000 - 240,000
Mental health support
Free lunch at the office
Professional development budget
+3
Kafka Engineer
Kafka Engineer

B3 Consulting Poland • Wrocław

On-site
PLN 180,000 - 260,000