Get more replies from employers
Send a job-specific resume in minutes.
Humyn Labs is building a two-sided, open data platform for physical AI, integrating sound, sight, motion, and touch data from field operations across 20+ countries. As Platform Architect, you will own the capture, processing, and evaluation stacks, working closely with ML researchers to define benchmarks and ensure high-quality training signals.
This role is hands-on with hardware specs and live contracts, with phase-oriented delivery.
Humyn Labs converts real human action - sound, sight, motion, and touch - into training signal for physical AI. We run verified field capture across 20+ countries in India, Southeast Asia, Latin America, and the Middle East: the real-world environments where physical AI deploys, not the labs where it is built.
Sound is proven: 100,000 hours across 33 languages, revenue-generating today, with BRIDGE - our published benchmark evaluating 15 ASR models across 22 languages and 7 voice parameters.
Sight is the current bet: 500,000 hours across 10+ countries, residential and commercial, with a live labeling stack (6DoF pose, hand-skeletal tracking, IMU/SLAM) and a discard rate below 15%, down from 40–60%. Motion is integrating: joint physics and navigation, captured through the sight pipeline. Touch is the frontier: teleoperated grippers with force and tactile sensors, built as a byproduct of motion collection - not a separate build.
The physical AI data problem has four stages of increasing depth: foundation-building (scale, format, diversity), task generalisation (domain-aware labeling), deployment readiness (environment-specific evaluation), and continuous refinement (correction-episode capture from deployed robots). The technology leader we hire will build a platform that serves all four - simultaneously, at scale.
This is not a workflow tool. It is not a B2C application. It is not a systems integration project. It is an open data platform with two sides that must compound each other - and a live delivery engine that pays for it today.
Phase 1 (now) - Humyn is the first supplier, proving the pipelines directly: sound done, sight in production, motion integrating. Phase 2 (next) - Humyn publishes the fusion spec (timestamps, hardware, SOPs) and certified external networks join the supply side. Phase 3 (destination) - the full two-sided platform: any certified network's pipelines fuse into one signal; any lab buys it. You are hired in Phase 1 to build the architecture Phases 2 and 3 run on.
An open architecture that allows multiple verified collectors, domain experts, and operational partners across 20+ countries to contribute multi-modal sensory data through Humyn's verification and labeling pipeline. Every modality runs the same four stages - collection, validation, processing, labelling - with human-in-the-loop QC across all of it. The platform must scale without Humyn owning every collection operation: contributors bring data; Humyn's pipeline processes, verifies, labels, and routes it. In Phase 2, the fusion spec you harden becomes the certification standard every external network must meet to plug in.
An open interface that allows all four buyer cohorts - foundation model builders, humanoid builders, world model companies, voice AI platforms - to benchmark their models against Humyn's held-out evaluation sets: to identify which tasks fail, in which environments, under which conditions. The gap between what a model can do and what the evaluation reveals it cannot becomes a structured data brief that Humyn fulfils. BRIDGE for sound is the working proof of this model. The physical AI equivalent - deployment-readiness benchmarking for robots - is what you will build.
Quality is decided at capture time, not in post-processing. You will own the hardware and software discipline that field operations execute across 20+ countries: sensor selection and evaluation (stereo depth, global-shutter RGB, IMU specifications), hardware time-sync and calibration workflows, on-device validation, and the per-unit economics of running capture fleets at scale. You will make the build‑vs‑buy calls on rigs and drive the hardware roadmap as specs from frontier labs tighten - including the touch path: teleoperated grippers with force/tactile sensing, designed as an extension of motion collection.
This is not future platform work. Live contracts with frontier labs carry hard hardware specifications and delivery deadlines today. You own throughput, spec compliance, and discard rate on those deliveries from day one.
The pipeline that converts raw multi-modal capture into embodiment-aligned training signal: 6DoF head pose, 21-keypoint hand-skeletal tracking, IMU/SLAM spatial reconstruction, dense natural-language action labeling, verifiable provenance records, and format-agnostic delivery in MCAP, RLDS, and LeRobot v3. Fusion is the core conversion: many synchronized streams become one multi-modal signal per domain - the same spec across every modality, geography, and environment. This stack is what separates a raw-footage commodity supplier from a verified-signal platform priced an order of magnitude higher. You will own it, extend it across modalities, and make it the processing standard researchers depend on.
Feedback that arrives in post-processing is useless - the session is over. You will maintain and extend the real‑time quality validation layer that operates at collection time across 20+ countries: frame‑drift detection, signal‑to‑noise monitoring, domain compliance checking, and instant feedback to field contributors in their own language. This is how Humyn got from 40–60% discard rates to below 15%. You will drive it further.