Stand out for this role — generate a tailored resume and cover letter in about a minute.
TechTree seeks an ML engineer to own the end-to-end video data pipeline from raw capture to validated episodes. You will benchmark and distill vision-language models, calibrate against human labels, and own eval sets and thresholds.
Build scalable QC/QA to keep every clip within spec and drive egocentric video tasks across languages and venues. You will deploy CV/robotics models in production, manage multi-camera data, and collaborate directly with founders and clients on performance and data
Own the video data pipeline that frontier robotics labs train on, from raw capture to validated episodes.
Hub is one of the fastest-growing data companies, providing real-world training data to the largest AI and robotics companies. Based in San Francisco and backed by Y Combinator and top VCs, our mission is to advance embodied AGI through real-world data and research.
As a founding ML engineer you own Hub's video data pipeline end to end, from raw capture in the field to a dataset a frontier robotics lab trains on. It is the layer that decides what we are allowed to ship.
Petabytes of video from multi-camera rigs we build ourselves: RGB-D, RGB and IMU, metric depth, global and rolling shutter, hardware sync. You turn raw bundles into validated episodes.
Egocentric video with narration, across languages and environments: speech recognition, translation, chapter segmentation, QA against each customer's taxonomy.
Quality control as an ML problem. You benchmark vision-language models, decide which ones we trust to judge our data, fine-tune and distill our own, and own the eval sets and thresholds behind that call.
Scalable QC and QA pipelines that trim, cut and quarantine, so no clip ever ships out of spec.
Annotation at scale against demanding customer taxonomies, plus hand tracking and fine-grained manipulation. Every human verdict becomes a training label.
The next modalities: tactile and teleoperation sit in the same problem space.
Live production for several of the top 5 AI companies, at high volume, with hard deadlines and specs.
A top engineering school, 3+ years of applied ML in computer vision or robotics. Less experience is fine for outliers: the bar is what you've built.
You've run VLMs as judges in production: panel agreement, calibration against human labels, fine-tuning and distillation.
You've trained and deployed CV or robotics models in production: detection and tracking, pose and hand estimation, depth, visual-inertial odometry, action recognition.
You know the video stack deeply: codecs and frame timing, multi-view geometry, intrinsics and extrinsics, distortion, temporal alignment across sensors.
Exceptional individual achievement: elite rankings at competitions or concours, hackathons won, projects at real scale.
An active GitHub and Hugging Face: recent contributions, open weights and datasets, reproduced results.
You follow the literature and can tell what's worth implementing from what's noise.
Agentic engineering as a craft: a custom harness, and a loop for your agents to verify their own work through tests, training runs and evals.
You can come to our Paris office once it opens.
Egocentric vision, IMUs, MCAP, ROS.
World models, video generation, VLAs or robot foundation models.
PyTorch, CUDA, multi-GPU training and distributed inference. Quantisation, batching and throughput matter as much as accuracy.
VLMs as judges: vLLM serving open-weight models on H100s, hosted models behind one provider interface, and our own fine-tuned and distilled judges.
Multi-stream HEVC video plus high-rate IMU per recording, hardware-synchronised and calibrated, delivered as MCAP.
Postgres for state, S3 for bytes, Kubernetes for compute, GPU inference at petabyte scale. Cost per hour processed is an engineering target.
Every threshold is a named constant tied to the customer requirement it comes from. Every quarantine carries a code, evidence and an owner.
$90,000 to $120,000 yearly salary.
Stock options between 0.25% and 0.5%, granted at signature.
Your own GPU budget.
Being able to come to our Paris office once it opens is a big plus.
Direct work with the founders and with the biggest AI labs.