Be | Shaping the Future Poland has a proven position of being a reliable partner for financial services organisations to analyse complex requirements, find solutions and implement them in their entirety, regardless of their complexity. Since the foundation of Be Poland in 2013, we have been continually expanding and customising our spectrum of services. Today, we are privileged to have in our team the best individuals in each sector we operate within the financial services industry.
Role: Senior SRE / Platform Engineer with Kafka
Location: fully remote from Poland
Contract Type: B2B
We are looking for experienced Senior SRE / Platform Engineers with Kafka to support a production readiness initiative for a critical booking processing platform. The assignment focuses on improving reliability, observability, resilience, and operational excellence across distributed systems and event-driven architectures. The ideal candidate combines strong hands-on experience with Kafka-based systems, observability tooling, cloud-native deployments, and Site Reliability Engineering practices.
Key Responsibilities
Observability & Monitoring
- Define and implement monitoring strategies based on Golden Signals
- Design and maintain dashboards, metrics, alerts, and reporting
- Improve centralized logging and distributed tracing capabilities
- Develop alerting rules, thresholds, and operational runbooks
- Support on-call processes and incident response activities
Kafka Reliability & Messaging
- Design and optimize Kafka topics, partitioning strategies, and consumer groups
- Implement retry mechanisms, dead-letter queues (DLQ), and idempotent processing
- Define schema governance and messaging standards
- Monitor Kafka performance, consumer lag, and throughput
- Improve reliability of asynchronous workflows, including backpressure handling and failure recovery
Reliability Engineering & SRE Practices
- Define and manage SLIs, SLOs, and error budgets
- Participate in incident management and post-mortem activities
- Drive reliability-by-design principles across services and platforms
- Identify and implement improvements that reduce operational overhead
Release & Deployment
- Support CI/CD pipelines and deployment automation
- Implement quality gates and release management processes
- Work with Canary, Blue-Green, and rollback strategies
- Support feature flag frameworks and version management
- Collaborate on load testing and production readiness assessments
- Work with container orchestration platforms such as Kubernetes
Data Protection & Disaster Recovery
- Define backup and restore strategies
- Support disaster recovery planning and testing
- Contribute to RPO/RTO definitions and operational procedures
- Participate in DR exercises and resilience testing
Required Skills & Experience
- Strong experience as an SRE, Platform Engineer, DevOps Engineer, or Reliability Engineer
- Hands-on experience with Apache Kafka in production environments
- Experience with monitoring, alerting, logging, and observability platforms
- Knowledge of distributed systems and event-driven architectures
- Experience with Kubernetes and containerized environments
- Experience with CI/CD pipelines and deployment automation
- Understanding of incident management and operational excellence practices
- Experience with MongoDB, including replication and backup concepts
- Strong troubleshooting and problem-solving skills
Nice to Have
- Experience in financial services or other highly regulated environments
- Experience with distributed tracing solutions
- Knowledge of cloud platforms (AWS, Azure, or GCP)
- Experience implementing SLO/SLI frameworks
- Certifications related to Kubernetes, cloud technologies, or SRE practices
Our offer:
- Competitive remuneration on B2B contract
- Access to Mindgram - mental health & well-being platform
- Free gym at Q22