Data Engineer (Spark)

Addepto

Białystok

Hybrid

PLN 180,000 - 210,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Remote work opportunities
Flexible working hours
Training & conference budget
Medical care and well-being support
Career development

Job summary

Addepto is a leading AI consulting and data engineering company that builds scalable, ROI-focused AI solutions for global enterprises and startups. The Data Engineer (Spark) role offers a chance to work on large automotive data platforms and cutting-edge data technologies.

You will contribute to streaming and batch data processing, leverage Spark, Cloudera, Airflow, Iceberg, Python, and AWS, and collaborate with cross-functional teams in a dynamic, supportive environment.

Qualifications

  • At least 3 years of commercial experience implementing, developing, or maintaining Big Data systems.
  • Strong programming skills in Python (or Java/Scala): writing clean, OOP code.
  • Hands-on with Spark, Cloudera, Kafka, Data Platform, Airflow, NiFi, Docker, and Iceberg.
  • Excellent understanding of dimensional data and data modeling techniques.
  • Experience deploying solutions in cloud environments.
  • Consulting experience with client management and effective communication.
  • Fluent English (C1 level).
  • Bachelor’s degree in technical or mathematical studies.

Responsibilities

  • Develop and maintain a high-performance data processing platform for automotive data, ensuring scalability and reliability.
  • Design and implement data pipelines that process large volumes of data in both streaming and batch modes.
  • Optimize data workflows using Spark, Cloudera, and Airflow.
  • Work with data lake technologies (e.g., Iceberg) to manage structured and unstructured data.
  • Collaborate with cross-functional teams to understand data requirements and integration needs.
  • Monitor and troubleshoot the platform to ensure high availability, performance, and accuracy.
  • Leverage AWS for infrastructure management and scaling.
  • Write and maintain high-quality Python (or Java/Scala) code for data processing tasks and automation.

Skills

Python
SQL
Spark
Airflow
Cloudera

Education

Bachelor's degree in technical or mathematical studies

Tools

Docker
Databricks
Kubernetes
Kafka
NiFi
Trino
Hudi
Java
Scala
Iceberg
AWS

Job description

Data Engineer (Spark)

Miejsce pracy: Białystok

Technologies we use
  • Python
  • SQL
  • Spark
  • Airflow
  • Cloudera
Optional
  • Kubernetes
  • Kafka
  • NiFi
  • Trino
  • Hudi
  • Java
  • Scala
  • Docker
  • Databricks
  • Iceberg
  • AWS
Operating system
  • Windows
  • macOS
About the project

As a Data Engineer, you will have the exciting opportunity to work with a team of technology experts on challenging projects across various industries, leveraging cutting-edge technologies. Here are some of the projects we are seeking talented individuals to join:

Development and maintenance of a large platform for processing automotive data. A significant amount of data is processed in both streaming and batch modes. The technology stack includes Spark, Cloudera, Airflow, Iceberg, Python, and AWS.

Design and development of a universal data platform for global aerospace companies. This Azure and Databricks powered initiative combines diverse enterprise and public data sources. The data platform is at the early stages of the development, covering design of architecture and processes as well as giving freedom for technology selection.

Centralized reporting platform for a growing US telecommunications company. This project involves implementing BigQuery and Looker as the central platform for data reporting. It focuses on centralizing data, integrating various CRMs, and building executive reporting solutions to support decision-making and business growth.

This is how we organize our work

This is how we work

  • in house
  • at the client's site
  • you can change the project
  • you have influence on the choice of tools and technologies
  • you have influence on the technological solutions applied
  • you have influence on the product
  • you develop the code "from scratch"
  • you focus on product development
  • agile
  • scrum
Team members
  • technical leader
  • big data developer
  • data scientist
  • project manager
Your responsibilities
  • Develop and maintain a high-performance data processing platform for automotive data, ensuring scalability and reliability.
  • Design and implement data pipelines that process large volumes of data in both streaming and batch modes.
  • Optimize data workflows to ensure efficient data ingestion, processing, and storage using technologies such as Spark, Cloudera, and Airflow.
  • Work with data lake technologies (e.g., Iceberg) to manage structured and unstructured data efficiently.
  • Collaborate with cross-functional teams to understand data requirements and ensure seamless integration of data sources.
  • Monitor and troubleshoot the platform, ensuring high availability, performance, and accuracy of data processing.
  • Leverage cloud services (AWS) for infrastructure management and scaling of processing workloads.
  • Write and maintain high-quality Python (or Java/Scala) code for data processing tasks and automation.
Our requirements
  • At least 3 years of commercial experience implementing, developing, or maintaining Big Data systems, data governance and data management processes.
  • Strong programming skills in Python (or Java/Scala): writing a clean code, OOP design.
  • Hands-on with Big Data technologies like Spark, Cloudera, Kafka, Data Platform, Airflow, NiFi, Docker, and Iceberg,.
  • Excellent understanding of dimensional data and data modeling techniques.
  • Experience implementing and deploying solutions in cloud environments.
  • Consulting experience with excellent communication and client management skills, including prior experience directly interacting with clients as a consultant.
  • Ability to work independently and take ownership of project deliverables.
  • Fluent English (at least C1 level).
  • Bachelor’s degree in technical or mathematical studies.
Optional
  • Experience with an MLOps framework such as Kubeflow or MLFlow.
  • Familiarity with Databricks and/or dbt.
Division of working time
  • 80% New features
  • 5% Maintenance / Bug fixing
  • 5% Client support
  • 5% Document writing
  • 5% Meetings
What we offer
  • Work in a supportive team of passionate enthusiasts of AI & Big Data.
  • Engage with top-tier global enterprises and cutting-edge startups on international projects.
  • Enjoy flexible work arrangements, allowing you to work remotely or from modern offices and coworking spaces.
  • Accelerate your professional growth through career paths, knowledge-sharing initiatives, language classes, and sponsored training or conferences, including a partnership with Databricks, which offers industry-leading training materials and certifications.
  • Choose from various employment options: B2B, employment contracts, or contracts of mandate.
  • Make use of 20 fully paid days off available for B2B contractors and individuals under contracts of mandate.
  • Participate in team-building events and utilize the integration budget.
  • Celebrate work anniversaries, birthdays, and milestones.
  • Access medical and sports packages, eye care, and well-being support services, including psychotherapy and coaching.
  • Get full work equipment for optimal productivity, including a laptop and other necessary devices.
  • With our backing, you can boost your personal brand by speaking at conferences, writing for our blog, or participating in meetups.
  • Experience a smooth onboarding with a dedicated buddy, and start your journey in our friendly, supportive, and autonomous culture.
Development opportunities we offer
  • assistance in preparation to public speeches
  • conferences abroad
  • conferences in Poland
  • development budget
  • external training
  • industry-specific e-learning platforms
  • intracompany training
  • soft skills training
  • substantive support from technological leaders
  • support of IT events
  • technical knowledge exchange within the company
  • time for development of your ideas
Benefits
  • sharing the costs of sports activities
  • private medical care
  • sharing the costs of foreign language classes
  • sharing the costs of professional training & courses
  • remote work opportunities
  • flexible working time
  • fruits
  • integration events
  • no dress code
  • video games at work
  • birthday celebration
  • employee referral program
  • Reimbursement for prescription glasses or contact lenses
Recruitment stages
  • Let's get to know each other (45 min)
  • Technical Interview (1 h)
  • Management Interview (1 h)
Addepto

Addepto is a leading AI consulting (https://addepto.com/ai-consulting/) and data engineering (https://addepto.com/data-engineering-services/) company that builds scalable, ROI-focused AI solutions for some of the world's largest enterprises and pioneering startups, including Rolls Royce, Continental, Porsche, ABB, and WGU. With an exclusive focus on Artificial Intelligence and Big Data, Addepto helps organizations unlock the full potential of their data through systems designed for measurable business impact and long-term growth.

The company's work extends beyond client engagements. Drawing from real-world challenges and insights, Addepto has developed its own product - ContextClue - and actively contributes open-source solutions to the AI community. This commitment to transforming practical experience into scalable innovation has earned Addepto recognition by Forbes as one of the top 10 AI consulting companies worldwide.

As part of KMS Technology, a US-based global technology group, Addepto combines deep AI specialization with enterprise-scale delivery capabilities—enabling the partnership to move clients from AI experimentation to production impact, securely and at scale.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer (Spark)
Data Engineer (Spark)

Addepto • Województwo pomorskie

On-site
PLN 120,000 - 180,000
Flexible work arrangements
Remote or office options
Professional development opportunities
Data Engineer (Spark)
Data Engineer (Spark)

Addepto • Kraków

Hybrid
PLN 180,000 - 240,000
Elastyczny czas pracy
Praca zdalna lub stacjonarna
Szkolenia i konferencje
+1
Senior Data Engineer (Spark)
Senior Data Engineer (Spark)

Addepto • Poznań

Hybrid
PLN 210,000 - 330,000
Flexible work arrangements
Training & conferences
Private medical care
+2
AI Tech Lead
AI Tech Lead

Addepto • Wrocław

Hybrid
PLN 240,000 - 420,000
Remote work flexibility
Career development opportunities
Team-building events
+7
Data Engineer (Databricks)
Data Engineer (Databricks)

Addepto • Poland

Remote
PLN 80,000 - 100,000
Flexible work arrangements
20 fully paid days off
Medical and sports packages
+1
Data Scientist / AI Engineer
Data Scientist / AI Engineer

Addepto • Warszawa

On-site
PLN 295,000 - 464,000
Flexible work arrangements
Career paths and knowledge-sharing initiatives
20 fully paid days off
+2
Remote Cloud-Native Solution Architect | Azure, Kafka & DDD
Remote Cloud-Native Solution Architect | Azure, Kafka & DDD

Addepto • Wrocław

On-site
PLN 190,000 - 270,000
AI Engineer
AI Engineer

Addepto • Warszawa

Hybrid
PLN 120,000 - 160,000
Career paths and knowledge-sharing initiatives
Flexible work arrangements
20 fully paid days off
+2
Junior Data Engineer (Power BI)
Junior Data Engineer (Power BI)

Addepto • Poland

Hybrid
PLN 90,000 - 130,000
Remote or hybrid work
Career development and training
Databricks training & certificates
+2
Senior/Lead Software Engineer
Senior/Lead Software Engineer

Addepto • Wrocław

On-site
Flexible work arrangements
Career paths and knowledge-sharing initiatives
20 fully paid days off
+3