Data Engineer - Video & Multimodal Training Data (World Models)

International Association of Plumbing and Mechanical Officials (IAPMO)

Alameda, Northern (CA, KY)

Hybrid

USD 150,000 - 210,000

Full time

8 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Unknown organization in the United States is seeking a data engineer to curate large-scale multimodal datasets, including video, audio, and robotics content, to train flagship models.

You will work with researchers to optimize captioning, quality scoring, deduplication, and dataset versioning, while tuning infrastructure and storage at petabyte scale.

Qualifications

  • Experience building training corpora for visual or multimodal ML at scale.
  • Hands-on with video/audio codecs, frame-level work, and syncing modalities.
  • Production experience with Python, Kubernetes, Ray or Flyte, and large storage.

Responsibilities

  • Curate large-scale multimodal datasets across video, robotics and audio.
  • Implement filtering, deduplication, quality scoring, captioning and shot detection.
  • Manage platforms processing millions of hours of content at scale.
  • Versioning and reproducibility of datasets, with clear lineage.
  • Collaborate with researchers on data quality and data mix.

Skills

Multimodal datasets
Data quality assessment
Python production
Kubernetes in production
Ray / Flyte
Petabyte-scale storage

Tools

Python
Kubernetes
Ray
Flyte

Job description

Our client builds world model systems that learn how the world actually behaves over time and generate it back, interactively, in real time. Seven model releases since December 2024, including the first real-time synchronised audio-and-video world model and a four-player shared world.

Data is fast becoming one of the biggest bottlenecks in building world models. Models are only as good as the data behind them, and getting that data right is one of the hardest and most important problems any research lab has.

What you'd own
  • Curating large-scale multimodal datasets across video, robotics and audio to train flagship models
  • The selection logic: filtering, deduplication, quality scoring, captioning and shot detection. The decisions that set the ceiling on model quality
  • Platforms processing millions of hours of content, reliably and at a cost that isn't embarrassing
  • Dataset versioning, lineage and reproducibility, which show which data was used and whether you can rebuild it
  • Working directly with researchers on caption quality and data mix, and closing the loop on which data moved which eval

Some days you're tuning infrastructure. Some days you're sitting with a researcher, working out why the model learned the wrong thing.

What we're looking for
  • You've built training corpora for visual or multimodal ML at a scale where you had to write code to judge quality because you couldn't review them yourself
  • Hands-on with video and/or audio - codecs, decoding, frame-level work, keeping modalities in sync. Not media as opaque blobs in object storage
  • Python, Kubernetes in production, Ray, Flyte or equivalent, petabyte-scale object storage, sharded dataset formats
  • You've deliberately discarded data and can explain exactly how you decided

We care more about how you think about data than about your years of experience or your publication record. Data engineering or research background - both work.

Also worth a conversation if:

you've never touched video, but you've run petabyte-to-exabyte data infrastructure where selection, layout, throughput and cost were the whole problem. That skill transfers. They've hired against it before.

Why this seat:

There is no incumbent data platform and no VP of Engineering in post. You'd define versioning, lineage and the quality bar rather than inherit someone else's from three years ago. The scope doesn't usually fit in one role: video, audio, robotics, synthetic data from our adversarial RL work, and licensed corpora. At a larger lab, that's five teams. You'd own a slice of one. And we name every contributor on our model release pages, including infrastructure and data people.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer - Video & Multimodal Training Data (World Models)
Data Engineer - Video & Multimodal Training Data (World Models)

This is Growth • San Francisco (CA)

On-site
USD 120,000 - 210,000
Member of Technical Staff, Data Engineering
Member of Technical Staff, Data Engineering

Odyssey • Palo Alto (CA)

On-site
USD 140,000 - 190,000
Member of Technical Staff, Data Engineering
Member of Technical Staff, Data Engineering

Odyssey • Palo Alto (CA)

On-site
USD 120,000 - 180,000
Multimodal Data Engineer — Video & Audio Pipelines
Multimodal Data Engineer — Video & Audio Pipelines

This is Growth • San Francisco (CA)

On-site
USD 120,000 - 210,000
Research Member of Technical Staff- Data Infrastructure
Research Member of Technical Staff- Data Infrastructure

Rhoda AI • Palo Alto (CA)

On-site
USD 180,000 - 280,000
Research Engineer, Physical AI (Robotics, World Models)
Research Engineer, Physical AI (Robotics, World Models)

Orbifold AI • Palo Alto (CA)

On-site
USD 180,000 - 240,000
ML Research Engineer, Data San Francisco, CA · On-site
ML Research Engineer, Data San Francisco, CA · On-site

Weave Robotics, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 200,000
ML / Data Platform Engineer - World Models
ML / Data Platform Engineer - World Models

Deca Talent • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff (Data Intelligence)
Member of Technical Staff (Data Intelligence)

Reka • United States

On-site
USD 100,000 - 130,000
ML Research Engineer
ML Research Engineer

Origin Lab • Los Angeles (CA)

On-site
USD 180,000 - 280,000