Founding Machine Learning Engineer

TechTree

San Francisco (CA)

On-site

USD 90,000 - 120,000

Full time

22 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Stock options 0.25–0.5%
GPU budget
Paris office access
Direct work with founders

Job summary

TechTree seeks an ML engineer to own the end-to-end video data pipeline from raw capture to validated episodes. You will benchmark and distill vision-language models, calibrate against human labels, and own eval sets and thresholds.

Build scalable QC/QA to keep every clip within spec and drive egocentric video tasks across languages and venues. You will deploy CV/robotics models in production, manage multi-camera data, and collaborate directly with founders and clients on performance and data

Qualifications

  • 3+ years of applied ML in computer vision or robotics
  • Experience running VLMs in production and calibrating against human labels
  • Experience deploying CV/robotics models at scale (detection, tracking, depth, odometry)
  • Deep video stack knowledge: codecs, multi-view geometry, sensors timing

Responsibilities

  • Own the video data pipeline end to end from raw capture to validated episodes.
  • Benchmark and distill vision-language models, set evals and thresholds.
  • Build scalable QC/QA pipelines to ensure no non-conforming clips ship.
  • Annotate at scale against customer taxonomies and handle hand tracking tasks.
  • Collaborate with founders and clients for production-grade data pipelines.

Skills

Applied ML in CV/robotics
Video data pipelines
Production model deployment
VLMs as judges

Education

Top engineering school

Tools

PyTorch
CUDA
Kubernetes
PostgreSQL
S3

Job description

Own the video data pipeline that frontier robotics labs train on, from raw capture to validated episodes.

Hub is one of the fastest-growing data companies, providing real-world training data to the largest AI and robotics companies. Based in San Francisco and backed by Y Combinator and top VCs, our mission is to advance embodied AGI through real-world data and research.

The Role

As a founding ML engineer you own Hub's video data pipeline end to end, from raw capture in the field to a dataset a frontier robotics lab trains on. It is the layer that decides what we are allowed to ship.

What You'll Own

Petabytes of video from multi-camera rigs we build ourselves: RGB-D, RGB and IMU, metric depth, global and rolling shutter, hardware sync. You turn raw bundles into validated episodes.

Egocentric video with narration, across languages and environments: speech recognition, translation, chapter segmentation, QA against each customer's taxonomy.

Quality control as an ML problem. You benchmark vision-language models, decide which ones we trust to judge our data, fine-tune and distill our own, and own the eval sets and thresholds behind that call.

Scalable QC and QA pipelines that trim, cut and quarantine, so no clip ever ships out of spec.

Annotation at scale against demanding customer taxonomies, plus hand tracking and fine-grained manipulation. Every human verdict becomes a training label.

The next modalities: tactile and teleoperation sit in the same problem space.

Live production for several of the top 5 AI companies, at high volume, with hard deadlines and specs.

The Profile

A top engineering school, 3+ years of applied ML in computer vision or robotics. Less experience is fine for outliers: the bar is what you've built.

You've run VLMs as judges in production: panel agreement, calibration against human labels, fine-tuning and distillation.

You've trained and deployed CV or robotics models in production: detection and tracking, pose and hand estimation, depth, visual-inertial odometry, action recognition.

You know the video stack deeply: codecs and frame timing, multi-view geometry, intrinsics and extrinsics, distortion, temporal alignment across sensors.

Exceptional individual achievement: elite rankings at competitions or concours, hackathons won, projects at real scale.

An active GitHub and Hugging Face: recent contributions, open weights and datasets, reproduced results.

You follow the literature and can tell what's worth implementing from what's noise.

Agentic engineering as a craft: a custom harness, and a loop for your agents to verify their own work through tests, training runs and evals.

Nice to Have

You can come to our Paris office once it opens.

Egocentric vision, IMUs, MCAP, ROS.

World models, video generation, VLAs or robot foundation models.

Stack

PyTorch, CUDA, multi-GPU training and distributed inference. Quantisation, batching and throughput matter as much as accuracy.

VLMs as judges: vLLM serving open-weight models on H100s, hosted models behind one provider interface, and our own fine-tuned and distilled judges.

Multi-stream HEVC video plus high-rate IMU per recording, hardware-synchronised and calibrated, delivered as MCAP.

Postgres for state, S3 for bytes, Kubernetes for compute, GPU inference at petabyte scale. Cost per hour processed is an engineering target.

Every threshold is a named constant tied to the customer requirement it comes from. Every quarantine carries a code, evidence and an owner.

What We Offer

$90,000 to $120,000 yearly salary.

Stock options between 0.25% and 0.5%, granted at signature.

Your own GPU budget.

Being able to come to our Paris office once it opens is a big plus.

Direct work with the founders and with the biggest AI labs.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Research Engineer - Robotics
ML Research Engineer - Robotics

D24 Search Limited • San Francisco (CA)

On-site
USD 270,000 - 330,000
Equity
Machine Learning Engineer
Machine Learning Engineer

Human Archive • San Francisco (CA)

On-site
USD 120,000 - 160,000
Staff, Robotics ML/Data Engineer
Staff, Robotics ML/Data Engineer

Persona AI Inc • Houston (TX)

On-site
USD 140,000 - 180,000
Competitive compensation
Bonus program
Medical benefits (99% employer covered
+4
Founding ML Engineer
Founding ML Engineer

a16z-speedrun • San Francisco (CA)

On-site
USD 150,000 - 230,000
Member of Technical Staff: Training Infrastructure
Member of Technical Staff: Training Infrastructure

Wintermeyer Ventures • San Francisco (CA)

On-site
USD 200,000 - 375,000
Senior Software Engineer, ML Systems
Senior Software Engineer, ML Systems

Voxel • San Francisco (CA)

On-site
USD 200,000 - 240,000
Health insurance
Dental insurance
Vision insurance
+6
ML Software Engineer
ML Software Engineer

Humble Robotics • San Francisco (CA)

On-site
USD 100,000 - 300,000
Staff Machine Learning Engineer
Staff Machine Learning Engineer

Unity Technologies SF • Mountain View (CA)

On-site
USD 172,000 - 284,000
Equity awards
Annual discretionary bonuses
Sales commissions
+4
Member of Technical Staff, MLE
Member of Technical Staff, MLE

AIC • San Francisco (CA)

On-site
USD 180,000 - 250,000
Medical benefits
Unlimited PTO
Equity
Member of Technical Staff - Training Platform
Member of Technical Staff - Training Platform

Prime Intellect AI • San Francisco (CA)

Hybrid
USD 150,000 - 300,000
Remote or SF office
Visa sponsorship
Relocation support
+3