Founding Data Engineer

Meetdavis

Paris

Sur place

EUR 90 000 - 130 000

Plein temps

Il y a 3 jours
Soyez parmi les premiers à postuler
Générateur de candidature

Obtenez une réponse de cet employeur — un CV et une lettre de motivation adaptés exactement à ce qu’on recherche pour ce poste.

Passez les filtres ATS

Résumé du poste

Meetdavis is hiring a Founding Data Engineer to build the data behind our foundation model, trained from scratch to generate buildings as geometric graphs. You will own the data end to end, from raw floorplans and synthetic generation to a canonical representation, curation and large-scale pretraining.

You will architect the data stack, write pipelines, run experiments, and perform data ablations yourself, combining real and synthetic sources.

Qualifications

  • 5+ years of strong data engineering experience.
  • Experience building datasets for large-scale pretraining, with real or synthetic sources.
  • Senior, hands-on coder able to architect data stack and set strategy.

Responsabilités

  • Own data end-to-end: from raw floorplans to canonical graphs and validation.
  • Architect the data stack and set data strategy; write pipelines and run experiments.
  • Perform ablations to determine data impact and support large-scale pretraining.

Connaissances

Python
Data pipelines
Distributed systems
Senior hands-on

Description du poste

TL;DR Davis is hiring a Founding Data Engineer to build the data behind our foundation model, trained from scratch to generate buildings as geometric graphs. You will own the data end to end, from raw floorplans and synthetic generation to a canonical representation, curation, large scale pretraining and the ablations that tell us which data actually moves the model. Here the dataset is part of the algorithm.

About Davis

Davis is an AI-native real estate company accelerating early-stage development and architectural design. Today developers coordinate four to five fragmented stakeholders over weeks or months. Soon they will need only one: Davis. We turn every input that shapes a development decision into decision-ready outputs: investor-grade feasibility studies, investment analysis, and architect-certified designs, delivered in days. Every stage pairs our proprietary AI systems with expert review, so velocity never comes at the cost of reliability. We closed a $5.5M preseed co-led by Heartcore Capital and Balderton Capital, with Yellow, Evantic and Entrepreneur First, alongside angels from the founding teams of Spacemaker, Black Forest Labs, Hugging Face, Supabase, Cleo and Spore Bio. We already work with leading developers and expect to support hundreds of projects over the coming year, deepening our research, our hiring, and our coverage of the development process end to end.

The Role

You will own the data our foundation model learns from, a model we train from scratch to generate buildings as geometric graphs. Part of the corpus comes from real floorplans as images and PDFs that have to become clean, standardized graphs. A large part will be synthetic, procedurally generated building graphs, geometry and rendered floorplans, with controlled variation in style, scan noise, annotations and furniture, each kept with its ground truth graph automatically. You will think about the whole loop, from raw and synthetic data to a canonical structured representation, curation and validation, the training dataset, large scale pretraining, evaluation, data ablations, and back to improving the generator and the data mixture. You are senior enough to architect the data stack and set the data strategy, and hands on enough to write the pipelines, run the experiments, train the models and do the ablations yourself.

What you will be working on
  • Corpus from raw sources. Turn real floorplans (images, PDFs, scans) into clean, standardized building graphs, with the geometry and semantics that make them trainable.
  • Synthetic data generation. Explore strategies to expand the dataset with synthetic data.
  • Canonical representation and curation. Define the standardized representation, then filter, deduplicate, quality score and validate at scale, with versioning, provenance and lineage.
  • Pretraining data and mixtures. Assemble the training datasets, design the mixture and the curriculum, and blend synthetic and real data for large scale pretraining.
  • Data ablations. Train models to learn which data actually helps, read the results, and feed them back into the generator and the mixture.
  • Evaluation. Build the eval harness (datasets, metrics, regression tests, monitoring) that tracks data and model quality over time.
What We Are Looking For

You have built the dataset, not just trained on it. You have personally built or generated the data used for a large pretraining run, from raw or synthetic sources, rather than only training on a dataset someone handed you. Senior and deeply hands on. At least 5 years of strong experience, senior enough to architect the data stack and set strategy, but still coding the pipelines, running the experiments and doing the ablations yourself.

Data as a first class problem

A track record where the data itself is the object: curation, filtering, deduplication, quality scoring, mixtures, synthetic generation.

Strong engineering

Deep Python, clean and typed code, async and concurrency, distributed data pipelines, TDD culture.

Nice to have

Computer vision and document understanding, images to structured output, OCR, layout extraction, segmentation, geometry extraction, vectorization, raster to vector, image to scene graph, 3D or CAD. Graph and structured scientific data, molecular, protein or scene graphs, meshes, CAD, BIM, 3D geometry, or relational and structured world data. Synthetic worlds and simulation, a structured state to a simulator to a renderer to synthetic images with perfect labels, then a perception model, with an eye on the sim to real gap. Foundation models from scratch, real involvement in a large pretraining run, not only fine tuning. Public evidence of data ownership, lead on a dataset, a Hugging Face release, a dataset card, a technical blog on your pipeline, or a talk on data curation or synthetic data. GIS and geometry, parcels, zoning layers, projections, computation.

Obtenez votre examen gratuit et confidentiel de votre CV.

ou faites glisser et déposez votre fichier ici.

Similar jobs

Postes similaires à comparer

Founding Data Engineer
Founding Data Engineer

Carbon Data Solutions • Paris

Sur place
EUR 120 000 - 180 000
Founding Data Engineer - Architect AI Data Stack
Founding Data Engineer - Architect AI Data Stack

Carbon Data Solutions • Paris

Sur place
EUR 120 000 - 180 000
Founding Data Engineer: Build Data for Graphs
Founding Data Engineer: Build Data for Graphs

Meetdavis • Paris

Sur place
EUR 90 000 - 130 000
AI Researcher - Full time
AI Researcher - Full time

Meetdavis • Paris

Sur place
EUR 90 000 - 150 000
Competitive salary
Meaningful equity
Equity incentives
AI Research Intern
AI Research Intern

Meetdavis • Paris

Sur place
EUR 13 000 - 20 000
AI Engineer - Full time
AI Engineer - Full time

Meetdavis • Paris

Sur place
EUR 70 000 - 110 000
AI Engineer Intern
AI Engineer Intern

Meetdavis • Paris

Sur place
EUR 10 000 - 15 000
Residential Architect
Residential Architect

Meetdavis • Paris

Sur place
EUR 42 000 - 65 000
Generalist Architect
Generalist Architect

Meetdavis • Paris

Sur place
EUR 55 000 - 75 000
Office in central Paris
Health insurance
Meal vouchers
+2
Data Engineer - Foundational
Data Engineer - Foundational

Harmattan AI • Paris

Sur place
EUR 70 000 - 90 000