Remote Spark Big Data Engineer for AI Training

ixolabs.ai

France

À distance

EUR 82 656 - 165 312

Temps partiel

14 jours+

Recevez plus de réponses des employeurs

Envoyez un CV adapté au poste en quelques minutes.

Résumé du poste

ixolabs.ai is seeking a Big Data Engineer (Spark) to help design and optimize Spark-based AI training pipelines. You will write and review Spark code across Scala, PySpark, and Spark SQL, focusing on scalable data processing, partitioning, and efficient joins.

Ideal candidates have extensive Spark production experience, strong knowledge of Spark internals, and hands-on with modern table formats and cloud platforms. Flexible, remote, contractor engagement with 10–25 hours per week is offered.

Qualifications

  • 6 years in big data engineering, including 4 years writing production Spark.
  • Strong familiarity with both Scala Spark and PySpark.
  • Deep knowledge of Spark internals (Catalyst, Tungsten, AQE).
  • Experience with at least one modern table format (Delta, Iceberg, Hudi).
  • Comfort with cloud data platforms (Databricks, EMR, Dataproc, Synapse).
  • Familiarity with Kafka, Airflow, or dbt is a plus.

Responsabilités

  • Generate and evaluate Spark instruction-response pairs covering DataFrame, SQL, and Structured Streaming APIs.
  • Review AI-generated code in Scala Spark, PySpark, and Spark SQL.
  • Provide feedback on partitioning strategies, broadcast joins, and skew handling.
  • Validate AI handling of Delta Lake, Iceberg, and Hudi table formats.
  • Evaluate cluster sizing, dynamic allocation, and Spark-on-Kubernetes patterns.
  • Identify subtle issues in shuffle behavior, serialization, and AQE-related regressions.

Connaissances

Big data engineering
Spark production experience
Scala Spark
PySpark
Cloud data platforms

Outils

Spark (Core)
Delta Lake
Iceberg
Hudi
Databricks
EMR
Dataproc
Synapse
Kafka
Airflow
dbt
Spark on Kubernetes

Description du poste

ixolabs.ai is seeking a Big Data Engineer (Spark) to help design and optimize Spark-based AI training pipelines. You will write and review Spark code across Scala, PySpark, and Spark SQL, focusing on scalable data processing, partitioning, and efficient joins.

Ideal candidates have extensive Spark production experience, strong knowledge of Spark internals, and hands-on with modern table formats and cloud platforms. Flexible, remote, contractor engagement with 10–25 hours per week is offered.

Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

Remote AWS Cloud Engineer for AI Training
Remote AWS Cloud Engineer for AI Training

ixolabs.ai • France

À distance
Remote Streaming Data Engineer for Real-time AI Training
Remote Streaming Data Engineer for Real-time AI Training

ixolabs.ai • France

À distance
EUR 35 000 - 52 000
Remote work
Flexible hours
Weekly payments via Stripe
AI Training MLOps Engineer — Remote & Flexible Hours
AI Training MLOps Engineer — Remote & Flexible Hours

ixolabs.ai • France

À distance
Data Engineering & Big Data – Spark / Databricks / ETL Engineer
Data Engineering & Big Data – Spark / Databricks / ETL Engineer

HCLTech • Paris

Hybride
EUR 65 000 - 100 000
Senior Data Engineer
Senior Data Engineer

Hashlist • Poissy

Sur place
EUR 85 000 - 125 000
Remote Freelance Data Architect & AI Trainer
Remote Freelance Data Architect & AI Trainer

10x.Team • Paris

Sur place
Remote work
Flexible hours
Long-term collaboration
Scala Engineer
Scala Engineer

ixolabs.ai • France

À distance
Remote work
Flexible hours
Weekly payments via Stripe
AI-Driven AppSec & DevSecOps Engineer (Remote)
AI-Driven AppSec & DevSecOps Engineer (Remote)

ixolabs.ai • France

À distance
Remote AI Engineer — Build Scalable ML Pipelines
Remote AI Engineer — Build Scalable ML Pipelines

YO IT Consulting • Paris

Sur place
EUR 50 000 - 80 000
Data Engineer
Data Engineer

Move2cloud • Marseille

Sur place
EUR 45 000 - 70 000