Human-centered physical intelligence starts withmodels that learn from how people actually do things in the real world. Our systems run in homes and workplaces with people doing real tasks, and every deployment returns hours of continuous, time-synchronized multimodal data. Nothing comparable exists publicly, and the models and benchmarks for this kind of data have yet to be built.
We're looking for an experienced researcher to be a driving force behind this work: shaping what these models should be capable of and building the evals that measure it, pretraining and post-training multimodal foundation models on our data, and shipping them into systems operating in the real world, where telemetry returns within days and shapes what we collect next. Research is what moves the company forward. The models we train determine what our systems can do, and the benchmarks we build define how we judge progress.
What You'll Do
Design and run the eval suite
- Design the eval suite that measures the capabilities we care about over real-world input, and stand it up as a repeatable, versioned system the whole team runs against.
- Build novel evals for tasks without an established benchmark, with metrics tied to real-world usefulness.
- Extend evals beyond offline benchmarks to online evaluation of how models actually behave in deployment, where static metrics stop being predictive.
- Analyze failure modes, red-team behavior, and turn results into concrete priorities for data and training.
Train, deploy, and study models
- Post-train multimodal foundation models on our data — SFT, RL, and preference optimization — and design experiments that isolate what improves performance.
- Take models from research into production — export, optimization, and inference across edge and cloud targets under real-time constraints — and build the telemetry that feeds failures and edge cases back into the data engine.
- Prototype agentic systems that perceive, reason, and act over continuous multimodal input, both in real time and offline over recorded data.
- Run scaling-law studies — how capability moves with data volume, diversity, model size, and compute — and use them to shape the research and data roadmap.
Scale the data engine
- Work across the full data lifecycle for large volumes of multimodal, time-synchronized data captured from real hardware on real-world tasks.
- Grow the engine along two axes at once: raw volume and domain coverage, so our models generalize across the environments, tasks, embodiments, and people they'll encounter.
- Build the tooling that tightens the loop between deployment, data, and training.
- Turn signal from real-world use into supervision that existing datasets don't offer.
- Work cross-functionally to translate model needs into human data collection priorities, partnering with engineering on the pipeline infrastructure that stores, versions, and serves it.
What We're Looking For
- PhD in machine learning, computer vision, robotics, or a related field, with a record of training models that reached production.
- Deep experience training and post-training multimodal models (vision, video, audio, language) — SFT, RLHF, DPO or GRPO, and building RL environments.
- A track record of designing evals and benchmarks — ideally for tasks without an established one — and strong opinions about what makes an eval trustworthy.
- Hands-on experience building data pipelines for model training at scale: you've owned curation, annotation, and quality on real datasets, and you know where data problems hide.
- Experience taking models into production — on-device or real-time inference on compute-constrained hardware, quantization, and the tradeoffs that come with tight latency and compute budgets.
- Strong engineering fundamentals: you write clean, reproducible code and are comfortable building infrastructure when the work calls for it.
- Practical fluency with AI coding tools and agents as part of your daily workflow, with good judgment about when to lean on them.
- Startup DNA: high ownership, comfort with ambiguity, and the judgment to make pragmatic calls on a small, flat, collaborative team.
Stack & Skills
Fluent
- PyTorch and the modern training stack — distributed training, mixed precision, experiment tracking, and reproducible pipelines
- Multimodal and vision-language models (VLMs) — architectures, post-training recipes (SFT, RLHF, DPO/GRPO), and RL environments
- Eval methodology — benchmark design, held-out set construction, statistical rigor, and human preference data
- Python, plus the ability to move into systems code when a pipeline needs it
Working knowledge
- Video understanding — temporal segmentation, action recognition, and long-horizon reasoning
- Data engineering for ML — large-scale multimodal data processing, dataset versioning, annotation tooling, and quality/coverage measurement
- Speech and audio models — streaming ASR and diarization in real-world conditions
- Retrieval — multimodal embeddings and vector search
- Model deployment and optimization — ONNX, TensorRT, vLLM, or similar inference stacks; quantization, distillation, and profiling on edge and real-time targets
- Cloud infrastructure (AWS or GCP), containers, and orchestration for training and data jobs
- Streaming and real-time inference over continuous sensor input
Bonus
- Experience building physical AI data engines — the pipelines, annotation systems, and quality loops behind large real-world datasets
- Experience building or running human data-collection programs, including consent and privacy handling
- Published work at top venues (NeurIPS, CVPR, ICCV, ICML, CoRL, RSS, ICRA) grounded in real-world deployment
- Experience with active learning, data selection, or scaling-law studies
- Research in embodied perception, human-robot or human-agent interaction, or robot learning
- Cross-embodiment learning, world models, or sim-to-real transfer
- Learning from demonstration or policy learning over multimodal sensor streams — vision-language-action models, diffusion policies, or behavior cloning from teleoperation data
- Simulation environments — Isaac, MuJoCo, or similar — for training, evaluation, or data generation
- Robot perception stacks — SLAM, 3D reconstruction, depth and point-cloud processing, and grasp or affordance prediction
- Real-world robot evaluation on physical hardware, or contributions to open robot learning ecosystems (Open X-Embodiment, LeRobot, RoboCasa)
Who You'll Work With
You’ll be working with a tightly knit team that has pioneered frontier AI research, shipped tens of millions of devices, scaled infrastructure used by millions every day, and defined how people and machines interact in the physical world. They’ve done it at Meta, Amazon, Microsoft, and Valve.
Why Join
- Early role at a funded startup founded by the team that built one of the most advanced physical AI data engines in the industry.
- Work on an open research problem with real-world multimodal data that no lab has, and build the benchmarks the field will use to measure progress.
- Your research directly shapes our system capabilities and informs what we build and collect next.
- Work across the whole loop: data, evals, training, and inference.
- Small team, exceptional peers, no bureaucracy.
Compensation:Competitive salary plus meaningful early-stage equity and benefits.
All applicants must be legally authorized to work in the United States. Noösphere will consider sponsorship of qualified candidates for employment visa status where required. Noösphere is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.