Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Causal, a San Francisco startup, is building a Large Physics foundation Model to learn physics from sensory data. We seek data engineers to own datasets end-to-end—from source discovery to training-ready pipelines and access controls—ensuring data quality for multimodal physical data.
You will design scalable pipelines using Spark, Ray, and Beam, develop QA checks, and collaborate with researchers to verify that datasets translate into model performance.
Our mission is general causal intelligence; AI that is capable of (1) predicting the future and (2) identifying the actions to alter it.
To achieve this breakthrough, we are building a Large Physics foundation Model (LPM) because physical systems, unlike text or images, are governed by verifiable cause and effect. We believe that scaling on physics will enable an understanding of causality required to predict and control physical systems, starting with weather.
Our founding team has built and deployed AI against the physical world in robotics, drug discovery, and particle physics at institutions like DeepMind, Waymo, Cruise, Insitro, Nabla Bio, and CERN.
We look for data engineers who are excited to tackle unsolved problems. Data is critical to any ML model but is especially consequential for our thesis to learn physics from sensory observations. The vast majority of meaningful progress in AI comes not from new architectures, but from training on data that is carefully curated with specific characteristics, quality, and scale.
Your mission is to own every dataset end to end — from discovering the source and securing access, to writing the pipelines that ingest it, to guaranteeing it enters training clean, standardized, and correct.
We value a relentless approach to problem-solving, rapid execution, and the ability to quickly learn in unfamiliar domains.