N’envoyez pas un CV générique — générez un CV et une lettre de motivation adaptés à ce poste précis.
Genesis in Paris is seeking a senior data engineer to design, build, and maintain large-scale data pipelines for robotics foundation model training and evaluation. You will own core data infrastructure including data models, storage, ingestion, and orchestration, collaborating with a team to standardize processing across real-world teleoperation and synthetic simulation datasets.
The ideal candidate has 8+ years in data engineering, deep distributed systems knowledge (Spark, Kafka), and hands-on
Design, build, and maintain large-scale data pipelines (batch and streaming) for robotics foundation model training and evaluation at petabyte scale
Own core data infrastructure: data model, storage systems, ingestion pipelines, transformation frameworks, and orchestration layers
Standardize data models and unify processing pipelines across real-world teleoperation and synthetic simulation datasets
Collaborate with a team of driven individuals committed to building general-purpose Physical AI
Excellent software engineering skills (Python, Go, or similar)
Extensive experience designing, building, and maintaining large-scale data pipelines (8+ years)
Deep understanding of distributed systems (Spark, Kafka, or similar)
Extensive experience with data storage technologies (data lakes, warehouses, object stores like S3)
Experience running and maintaining production-grade infrastructure (Kubernetes, Terraform)
Bonus: Experience supporting AI systems, in particular embodied AI like self-driving