ML Research Engineer

Origin Lab

Los Angeles (CA)

On-site

USD 180,000 - 280,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Origin Lab is seeking a senior, hands-on research hybrid role combining ML engineering and scientific inquiry. You will work with a multimodal data corpus, test models across architectures, reproduce findings, and publish the strongest results with well-maintained checkpoints and evaluation harnesses.

You will liaise with academic labs and private model makers to run studies our data enables, and guide cross-functional teams from dataset design to strategic outcomes.

Qualifications

  • Hands-on experience training, testing, and using video and multimodal models.
  • Generalist across ML engineering, research science, and data science.
  • Rigorous experimental design with preregistered endpoints and kill criteria.
  • Fluency with open-weights ecosystem and Hugging Face tooling.

Responsibilities

  • Choose what experiments, what models, and what data to run.
  • Demonstrate knowledge across many models and quantify data performance.
  • Reproduce published results on our data and publish the strongest findings with clean checkpoints and released evaluation harnesses.
  • Own and liaise with research partnerships with academic labs and private model makers.
  • Work cross-functionally: translate business needs into datasets and strategy.
  • Publish: design short and long-term publication strategies.

Skills

Video and multimodal models
ML engineering
Data wrangling
Experiment design
Open-weights ecosystem
PyTorch
Distributed training
Reproducibility

Tools

Hugging Face
safetensors
Experiment tracking
Open model landscape

Job description

The role

We have strong hypotheses about why our data is so unique and worth more than what the internet produces. Our multimodal data corpus will reach more than 500,000 hours and more than 10 petabytes before the end of this year. Your job is to explore it, improve it and prove out our hypothesis as models and research needs grow and change. You are also the bridge to the gap of what’s possible from the research and how our data can be used to improve and solve data needs.

You use publicly available models, open weights, and private partnerships with model makers to test, validate, and stress our data across different architectures, use cases, and training scenarios. You reproduce published findings on our corpus, run the experiments that show where our data wins and where it does not, and take the strongest results toward publication and academic collaboration.

This is a senior, hands-on, catch-all research seat: part ML engineer, part research scientist. You have tinkered with a lot of models and you like it that way.

What you\'ll do
  • Choose what experiments, what models, and what data to run.
  • Demonstrate knowledge and expertise across many models, staying up on the latest findings, and newly quantifying the performance of our data.
  • Reproduce published results on our data, and publish the strongest findings: clean checkpoints, honest model cards, released evaluation harnesses.
  • Own and liaise with research partnerships: work with academic labs and private model makers to run studies our data uniquely enables.
  • Work cross-functionally: take a business or customer need, define the dataset that meets it, and advise on the best strategy to solve it.
  • Publish: design short and long-term publication strategies from the work you lead
What we\'re looking for

Senior IC, 5+ years. We care about depth over breadth in one place: hands-on experience training, testing, and using video and multimodal models, and organizing the data behind them.

  • Deep, hands-on experience with video and multimodal models: you have trained, fine-tuned, evaluated, and just plain tinkered with many of them, and you organize and wrangle the data they run on.
  • A generalist who spans ML engineering, research science, and data science, and is comfortable owning a question from dataset to result.
  • Rigorous experimental design: matched-compute, dose-response studies, preregistered endpoints and kill criteria, paired statistics, reported nulls. If a result depends on a choice made after seeing the data, you know it does not count.
  • Fluency in the open-weights ecosystem: Hugging Face stack, safetensors, experiment tracking, and today\'s open model landscape (Depth Anything, VGGT, SAM 2, Wan, Cosmos, LTX). You pin exact checkpoints and read licenses as carefully as code.
  • Applied deep learning in PyTorch, including distributed fine-tuning and evaluation, and the instinct to diagnose a run that is quietly wrong, not just one that crashes.
  • Reproducibility discipline: you validate a harness against a published number before trusting a figure on our data.
Nice to have
  • Major plus: experience with video game data or content, whether capture, engine telemetry, or game-derived datasets.
  • Depth and camera geometry: metric vs. affine-invariant depth, intrinsics and extrinsics, OpenCV/COLMAP/TUM conventions, pose metrics.
  • Controllable video generation and world models: camera-trajectory and depth conditioning, action conditioning, recoverability metrics, world-model evaluation (WorldScore, VBench-class).
  • Prior public research output, academic collaborations, or published benchmarks.
  • Large-scale data handling: WebDataset streaming over terabyte-scale video, decoding depth blobs to metric float32.
  • Agentic coding tools (Claude, Cursor, Codex) used to move fast on harness and analysis code.
About Origin Lab

Origin Lab delivers AI-enriched catalogs across video game capture, 3D environments, TV/film, and animation, licensed at the source with audit-ready provenance. We work with AI researchers from Oxford, Google Research, and others to drive breakthroughs in Artificial World Intelligence.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Full Stack Engineer
Full Stack Engineer

Origin Lab • San Francisco (CA)

On-site
USD 120,000 - 150,000
Research Scientist, Video Understanding & World Models
Research Scientist, Video Understanding & World Models

Kindredventures • New York (NY)

On-site
USD 120,000 - 150,000
Member of Technical Staff, Machine Learning_CNTR
Member of Technical Staff, Machine Learning_CNTR

PulseRise Technologies LTD • San Francisco (CA)

On-site
USD 150,000 - 210,000
Head of AI Data Sales
Head of AI Data Sales

Origin Lab • San Francisco (CA)

On-site
USD 200,000 - 580,000
Research Scientist, SLAM & VIO
Research Scientist, SLAM & VIO

Mecka • New York (NY)

On-site
USD 100,000 - 140,000
Access to proprietary data
Cutting-edge research environment
Opportunity for high impact
Machine Learning Engineer
Machine Learning Engineer

Jack • San Francisco (CA)

On-site
USD 190,000 - 240,000
Health insurance
Series A funding
Small, high-impact team
+1
Member of Technical Staff, Data Engineering
Member of Technical Staff, Data Engineering

Odyssey • Palo Alto (CA)

On-site
USD 140,000 - 190,000
Research Engineer, Data
Research Engineer, Data

Harnham • California (MO)

On-site
USD 100,000 - 130,000
Machine Learning Engineer
Machine Learning Engineer

Human Archive • San Francisco (CA)

On-site
USD 120,000 - 160,000
Product Manager
Product Manager

Origin Lab • Los Angeles (CA)

On-site
USD 120,000 - 160,000