Data & ML Infrastructure Lead

S27a

Paris

Sur place

EUR 120 000 - 160 000

Plein temps

14 jours+

Recevez plus de réponses des employeurs

Envoyez un CV adapté au poste en quelques minutes.

Résumé du poste

UMA is seeking a Data & ML Infrastructure Lead to own and scale the data backbone of our robotics platform. You will design and operate the end-to-end data platform, from on-robot capture to training-ready datasets, with a focus on reliability, scale, and reproducibility.

You will build pipelines, optimize data loading for training, and grow the data infrastructure team as the system evolves. A strong background in Python and systems programming is essential, with experience across big data

Qualifications

  • 8+ years in data infrastructure, ML infrastructure, or large-scale data-platform engineering at senior/lead level
  • Proven track record building data-intensive infrastructure in production and at scale
  • Deep experience with large-scale multimodal and time-series data and the storage systems behind it
  • Hands-on experience optimizing training data pipelines and data loading throughput
  • Treat versioning, lineage, observability and reproducibility as core engineering concerns
  • Strong Python, plus experience in a high-performance compiled language (C++, Rust)
  • Experience with compute/training infrastructure (GPU clusters, distributed training, Slurm, cloud) or potential to grow into it
  • Ability to reason end-to-end about performance, scalability, reliability, and cost
  • Thrives in a fast-paced startup, autonomous, execution-driven, curious about AI and robotics
  • Bonus: robotics/AV data, RL/continuous-learning loops, fleet-scale data collection
  • Bonus: public projects or open-source contributions

Responsabilités

  • Lead the design of the data platform end to end from robot data capture to training-ready datasets
  • Build versioned, orchestrated pipelines (Airflow) for post-processing and dataset statistics
  • Make training data loading fast with efficient decoding, prefetching and optimized formats
  • Produce and manage datasets from RL and continuous-learning loops
  • Design visualization, exploration, and metadata tooling to inspect and debug data at scale
  • Design storage and transfer with cost-efficient formats for exploration and high-throughput training
  • Grow compute and training infrastructure as needed and set production-grade practices

Connaissances

Python
C++
Rust
Airflow
Spark/Ray
Multimodal data
Time-series data
Data platforms
Observability
End-to-end systems

Formation

Bachelor's degree in Computer Science or related field

Outils

S3
Docker
Kubernetes

Description du poste

Your Mission

As Data & ML Infrastructure Lead, you will own and scale the data backbone of UMA: the systems that record, store, version, serve, and visualize the data our robots produce, and the infrastructure that turns that data into trained policies. Data is one of the central challenges in robotics AI, and we're looking for someone who has built data infrastructure at scale to own it with us.

We already have a working data and training stack and a strong team behind it, so you won't be starting from zero but you'll have the mandate to shape the best possible architecture, redesigning from the ground up where that's what it takes. You'll take ownership of the data platform end to end: recording, storage, versioning, high-throughput loading for training, and the tooling to explore and debug our data at scale, and quickly grow into leading the Data & ML Infrastructure team as it scales. The goal is a tight, scalable loop — a data engine where our fleet continuously produces data, we learn from it in near real time, and we ship improvements back fast — much of which doesn't exist off the shelf. The architectural calls you make now will determine how smoothly we get there. Compute and training infrastructure (GPU clusters, distributed training, scheduling) is a secondary but growing part of the role; the data platform is the core today

Key responsibilities :

  • Lead the design of our data platform end to end — from on-robot capture of multimodal data (video, depth, proprioception, actions, sensors) to training-ready datasets — built for reliability, scale, and reproducibility

  • Build versioned, orchestrated pipelines (e.g. Airflow) for post-processing and dataset statistics, so every run is fully reproducible

  • Make training data loading fast with efficient video decoding, prefetching, and a training-optimized format that keeps GPUs fed

  • Produce and manage datasets from RL and continuous-learning loops, closing the loop between deployment and training

  • Design and build visualization, exploration, and metadata tooling to inspect, curate, and debug our data at scale — central to our data-centric strategy

  • Design storage and transfer with formats suited to both exploration and high-throughput training, scaling cost-effectively as data and fleet grow

  • Grow our compute and training infrastructure (GPU clusters, distributed training, scheduling) as that need arises, and help set data standards and production-grade practices as the team grows

What You Bring to the Table
  • 8+ years in data infrastructure, ML infrastructure, or large-scale data-platform engineering, at a senior, lead, or staff level

  • Proven track record building data-intensive infrastructure in production and at scale — distributed data processing (e.g. Spark/Ray), workflow orchestration, cloud — ideally a data flywheel powering a continuously improving ML system

  • Deep experience with large-scale multimodal and time-series data (video, sensor streams, high-frequency signals) and the storage systems behind it — object storage (e.g. S3) for media, databases for structured data

  • Hands-on experience optimizing training data pipelines — loading throughput, video decoding, prefetching, keeping GPUs fed

  • Treat versioning, lineage, observability and reproducibility as core engineering concerns

  • Strong Python, with solid experience in a high-performance compiled language (C++, Rust) with the taste to build tooling that is reliable, maintainable, and pleasant to use

  • Good grasp of compute/training infrastructure (GPU clusters, distributed training, Slurm, cloud) — or the clear ability to grow into it

  • Ability to reason about systems end-to-end — performance, scalability, reliability, cost — and make and defend the right trade-offs

  • Thrive in a hands-on, fast-paced startup, building from scratch as the company grows: autonomous, rigorous, execution-driven, easy to work with, and broadly curious about AI, robotics, and systems

  • Bonus : robotics, autonomous vehicles, or other embodied/physical-AI data (adjacent large-scale multimodal — AV, video, geospatial/sensor — counts strongly), RL/continuous-learning loops, or fleet-scale data collection

  • Bonus : public projects, open-source contributions, maintained tools, or technical writing

  • We value exceptional builders over perfect resumes. If you have a world-class data-infrastructure track record and the drive to build the backbone that lets a robotics company scale, we strongly encourage you to apply — even if you don't tick every box

Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

Data & ML Infrastructure Lead
Data & ML Infrastructure Lead

UMA • Paris

Sur place
EUR 70 000 - 100 000
Robotics Data & ML Platform Lead
Robotics Data & ML Platform Lead

S27a • Paris

Sur place
EUR 120 000 - 160 000
Research Scientist / Engineer
Research Scientist / Engineer

UMA • Paris

Sur place
EUR 50 000 - 70 000
Staff Software Engineer, Infrastructure (Cloud)
Staff Software Engineer, Infrastructure (Cloud)

AeroVect • Paris

Sur place
EUR 110 000 - 150 000
Head of Data & ML Infrastructure
Head of Data & ML Infrastructure

UMA • Paris

Sur place
EUR 70 000 - 100 000
Data Engineer - Foundational
Data Engineer - Foundational

Harmattan AI • Paris

Sur place
EUR 70 000 - 90 000
Research Engineer — Robot Learning
Research Engineer — Robot Learning

Bleu Robotics • Paris

Sur place
EUR 90 000 - 130 000
Ambitious research mission
Advanced robots fleet
Research that ships to production
+5
Applied AI Engineer
Applied AI Engineer

Norbert Health • Paris

Sur place
EUR 70 000 - 90 000
Competitive salary and equity
High autonomy and technical ownership
Transparent, mission-driven culture
Data Engineer (Detect & Track Distillation)
Data Engineer (Detect & Track Distillation)

AI Chopping Block • Paris

Hybride
EUR 70 000 - 110 000
Data Engineer (Detect & Track Distillation)
Data Engineer (Detect & Track Distillation)

Harmattan AI • Paris

Sur place
EUR 70 000 - 95 000