Data Engineer (Spark)

Addepto

Town of Poland (NY)

Hybrid

USD 110,000 - 170,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Flexible remote or office work
Professional training and conferences
Paid time off
Formal onboarding buddy
Medical and wellness benefits

Job summary

Addepto is seeking a data platform engineer to build and maintain a high‑performance data processing platform for automotive data. You will design scalable pipelines for streaming and batch workloads using Spark, Cloudera, Airflow, and Iceberg, and collaborate with cross‑functional teams to meet data requirements.

You will also optimize data workflows on AWS, write production-grade Python or Java/Scala code, and help ensure reliability, availability, and accurate analytics for enterprise clients.

Qualifications

  • 4+ years of commercial experience with Big Data systems and data management.
  • Strong programming in Python or Java/Scala with clean OO design.
  • Hands-on with Spark, Cloudera, Airflow, Iceberg and related data platforms.
  • Cloud experience in AWS and data processing in production environments.
  • Consulting experience and excellent client-facing communication.

Responsibilities

  • Develop and maintain a high-performance data processing platform for automotive data.
  • Design and implement data pipelines for streaming and batch processing.
  • Optimize workflows for ingestion, processing, and storage using Spark, Cloudera, Airflow.
  • Work with data lake technologies (e.g., Iceberg) for structured/unstructured data.
  • Collaborate with cross-functional teams to integrate data sources.
  • Monitor and troubleshoot to ensure high availability and accuracy.
  • Leverage AWS for infrastructure and scaling of workloads.
  • Write high-quality Python or Java/Scala code for data tasks and automation.

Skills

Python
Java/Scala
Big Data
Cloud AWS
Spark
Cloudera
Airflow
Iceberg
Kafka

Education

Bachelor's degree in technical or mathematical studies

Tools

Databricks
dbt

Job description

Addepto is a leading AI consulting (https://addepto.com/ai-consulting/) and data engineering (https://addepto.com/data-engineering-services/) company that builds scalable, ROI-focused AI solutions for some of the world's largest enterprises and pioneering startups, including Rolls Royce, Continental, Porsche, ABB, and WGU. With an exclusive focus on Artificial Intelligence and Big Data, Addepto helps organizations unlock the full potential of their data through systems designed for measurable business impact and long-term growth.

The company's work extends beyond client engagements. Drawing from real-world challenges and insights, Addepto has developed its own product – ContextClue – and actively contributes open-source solutions to the AI community. This commitment to transforming practical experience into scalable innovation has earned Addepto recognition by Forbes as one of the top 10 AI consulting companies worldwide.

As part of KMS Technology, a US-based global technology group, Addepto combines deep AI specialization with enterprise-scale delivery capabilities—enabling the partnership to move clients from AI experimentation to production impact, securely and at scale.

Projects we are seeking talented individuals to join:
  • Development and maintenance of a large platform for processing automotive data. A significant amount of data is processed in both streaming and batch modes. The technology stack includes Spark, Cloudera, Airflow, Iceberg, Python, and AWS.
  • Design and development of a universal data platform for global aerospace companies. This Azure and Databricks powered initiative combines diverse enterprise and public data sources. The data platform is at the early stages of the development, covering design of architecture and processes as well as giving freedom for technology selection.
  • Centralized reporting platform for a growing US telecommunications company. This project involves implementing BigQuery and Looker as the central platform for data reporting. It focuses on centralizing data, integrating various CRMs, and building executive reporting solutions to support decision‑making and business growth.
Your main responsibilities:
  • Develop and maintain a high-performance data processing platform for automotive data, ensuring scalability and reliability.
  • Design and implement data pipelines that process large volumes of data in both streaming and batch modes.
  • Optimize data workflows to ensure efficient data ingestion, processing, and storage using technologies such as Spark, Cloudera, and Airflow.
  • Work with data lake technologies (e.g., Iceberg) to manage structured and unstructured data efficiently.
  • Collaborate with cross-functional teams to understand data requirements and ensure seamless integration of data sources.
  • Monitor and troubleshoot the platform, ensuring high availability, performance, and accuracy of data processing.
  • Leverage cloud services (AWS) for infrastructure management and scaling of processing workloads.
  • Write and maintain high-quality Python (or Java/Scala) code for data processing tasks and automation.
What you’ll need to succeed in this role:
  • At least 4 years of commercial experience implementing, developing, or maintaining Big Data systems, data governance and data management processes.
  • Strong programming skills in Python (or Java/Scala): writing a clean code, OOP design.
  • Hands‑on with Big Data technologies like Spark, Cloudera, Kafka, Data Platform, Airflow, NiFi, Docker, and Iceberg.
  • Excellent understanding of dimensional data and data modeling techniques.
  • Experience implementing and deploying solutions in cloud environments.
  • Consulting experience with excellent communication and client management skills, including prior experience directly interacting with clients as a consultant.
  • Ability to work independently and take ownership of project deliverables.
  • Fluent English (at least C1 level).
  • Bachelor’s degree in technical or mathematical studies.
Nice to have:
  • Experience with an MLOps framework such as Kubeflow or MLFlow.
  • Familiarity with Databricks and/or dbt.
Perks & Benefits:
  • Work in a supportive team of passionate enthusiasts of AI & Big Data.
  • Engage with top-tier global enterprises and cutting-edge startups on international projects.
  • Enjoy flexible work arrangements, allowing you to work remotely or from modern offices and coworking spaces.
  • Accelerate your professional growth through career paths, knowledge-sharing initiatives, language classes, and sponsored training or conferences, including a partnership with Databricks, which offers industry-leading training materials and certifications.
  • Choose your preferred form of cooperation: B2B or a contract of mandate, and make use of 20 fully paid days off.
  • Participate in team-building events and utilize the integration budget.
  • Celebrate work anniversaries, birthdays, and milestones.
  • Access medical and sports packages, eye care, and well-being support services, including psychotherapy and coaching.
  • Get full work equipment for optimal productivity, including a laptop and other necessary devices.
  • With our backing, you can boost your personal brand by speaking at conferences, writing for our blog, or participating in meetups.
  • Experience a smooth onboarding with a dedicated buddy, and start your journey in our friendly, supportive, and autonomous culture.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior/Lead Software Engineer / Solution Architect
Senior/Lead Software Engineer / Solution Architect

Addepto • Town of Poland (NY)

Hybrid
USD 110,000 - 150,000
Flexible work arrangements
Career development opportunities
Medical and sports packages
DataBricks Developer (IoT sphere)
DataBricks Developer (IoT sphere)

Coherent Solutions • Town of Poland (NY)

Hybrid
USD 110,000 - 155,000
Health insurance
Flexible remote/hybrid options
Referral program
+3
Strategic Partnerships Manager
Strategic Partnerships Manager

Addepto • Town of Poland (NY)

Hybrid
USD 90,000 - 120,000
Flexible work arrangements
Career growth and training
Medical and sports packages
+1
AI Engineer
AI Engineer

CloudDevs • Town of Poland (NY)

Remote
USD 90,000 - 120,000
20 fully paid days off
Career paths and training opportunities
Access to medical and sports packages
Data Engineer Remote Latin America
Data Engineer Remote Latin America

Fractal River • United States

Hybrid
USD 70,000 - 120,000
Personal development plan
Access to a reference library
Unlimited access to AI tools
+3
Senior Agentic Data Management Advisor
Senior Agentic Data Management Advisor

Lingaro • Town of Poland (NY)

Hybrid
USD 120,000 - 160,000
Stable employment
Flexibility regarding working hours
Unlimited access to the Udemy learning platform
Data Engineer
Data Engineer

AgileEngine • United States

Hybrid
USD 110,000 - 140,000
Professional growth
Competitive compensation
Exciting project selection
+1
Senior Data Engineer ID75059
Senior Data Engineer ID75059

AgileEngine • New York (NY)

Hybrid
USD 140,000 - 210,000
Professional growth
Competitive compensation
Exciting projects
+1
Senior Data Engineer ID75059
Senior Data Engineer ID75059

AgileEngine • Richmond (VT)

Hybrid
USD 120,000 - 180,000
Professional growth
Competitive USD-based pay
Exciting projects
+1
Senior Data Engineer
Senior Data Engineer

AgileEngine, LLC • United States

On-site
USD 140,000 - 200,000
100% remote work with flexible hours