Research Engineer, Multimodal Data

Eventual

United States

On-site

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

In-person, tight-knit SF team
Competitive compensation
Startup equity
Catered lunches and dinners
Commuter benefit
Team-building events
Health, vision, dental
Flexible PTO
Latest Apple equipment
401(k) plan with match

Job summary

Eventual is seeking a Research Engineer on the Visual Understanding team in the United States. You will own the layer that makes petabytes of video queryable by content, enabling corpus-scale annotation with cost-effective economics.

You will define the roadmap, train models, and build datasets that power customer training runs. This is a research engineering role focused on production impact and practical deployment, not just papers.

Qualifications

  • Strong familiarity with modern vision and multimodal models and deployable SOTA.
  • Experience running models at scale on video and sensor data for perception tasks.
  • Background from perception teams in robotics or visual-data companies or related research labs.

Responsibilities

  • Own the visual understanding roadmap end-to-end and production-inference at corpus scale.
  • Train, fine-tune, and evaluate VLMs, VQA, embeddings, and perception models against datasets.
  • Reduce per-clip annotation cost via model selection, distillation, batching, and pipelining.
  • Build rich, queryable datasets and design robust taxonomies with researchers.
  • Coordinate with dataloader and storage teams so outputs flow into the index and GPU pipelines.
  • Collaborate with researchers at partner labs for tight feedback loops.

Skills

Vision models
Multimodal models
Scale training
Embeddings
Cloud infra
Big data
Experimentation

Job description

About Eventual

Every breakthrough Physical AI system — humanoid robots, autonomous vehicles, video generation models — is trained on petabytes of video, lidar, radar, and sensor data. But today's data platforms (Databricks, Snowflake) were built for spreadsheet-like analytics, not the multimodal corpora that power AI. As a result, robotics and video-AI teams iterate on model improvement about once a week. Most of that week isn't training — it's finding the right data: writing CV heuristics over raw footage, paying annotators for edge cases, hand-curating clips before a cluster ever spins up. GPU bandwidth has grown 2-3× per generation. Storage and pipelines haven't. The gap widens every year.

Eventual was founded in 2022 to close it. Our open-source engine, Daft, is the distributed data engine purpose-built for multimodal AI — already running 2 PB/day at Amazon, 60-100 PB at another FAANG company, and in production at Mobileye, TogetherAI, and CloudKitchens. We are building a video-native index on top of our engine for Physical AI that collapses the data iteration loop. Describe the dataset you want, get a curated table in minutes, feed it to your GPUs at line rate. One iteration per day becomes the norm.

We're building this in partnership with the top PhysicalAI labs and public AI infrastructure companies today. We have raised $30M from Felicis, CRV, Microsoft M12, Citi, Essence, Y Combinator, Caffeinated Capital, Array.vc, and angels from the co-founders of Databricks and Perplexity. We've assembled a world-class team from AWS, Render, Pinecone and Tesla. We have spent our careers powering the last generation of PhysicalAI in self-driving, and are excited to now do this for the next.

Join our small (but powerful!) team working together 4 days/week in our SF Mission district office.

Your Role

As a Research Engineer on the Visual Understanding team, you'll own the layer that makes petabytes of video queryable by content. Physical AI teams have raw footage, lidar, radar, and sim outputs scattered across object stores with no way to find what they need without weeks of human annotation. We change that economics: we run vision-language models over every clip in a corpus along axes the customer cares about (gripper type, failure mode, object class, scene, motion density), so a researcher can ask "left-arm grasp failures on deformable objects" and get a curated dataset in minutes.

You'll define the roadmap for our visual understanding capabilities, train and select the models that make corpus-scale annotation tractable at single-digit cents per hour of video, and build the rich datasets that go on to train customer models. This is a research engineering role — meaning you'll read papers and run experiments, but you ship to production and your work is judged by what it does for customer training runs.

Key Responsibilities
  • Own the visual understanding roadmap end-to-end: from picking the model family for a customer's taxonomy to landing it in production inference at corpus scale.

  • Train, fine-tune, and evaluate VLMs, VQA models, embedding models, and convolutional perception models against customer datasets and benchmarks.

  • Drive down per-clip annotation cost — model selection, distillation, batching, decode pipelining — so "annotate every clip in a 10K-hour corpus" stays economical.

  • Build the rich, queryable datasets that customers train on: design taxonomies with researchers, instrument quality, version the outputs.

  • Partner with the dataloading and storage teams so visual understanding outputs flow into the index and on to the GPU without re-engineering.

  • Work directly with researchers at our partner labs — your shortest feedback loop is their next training iteration.

What we look for
  • Strong familiarity with modern vision and multimodal models — convolution nets, VLMs, VQA, embeddings — and a sense for the SOTA that's actually deployable today vs. on a leaderboard.

  • Experience running these models at scale on real video and sensor data, ideally for perception tasks (detection, tracking, segmentation, retrieval, captioning).

  • Background from a perception team at a self-driving, robotics, or visual-data company — or equivalent depth from a research lab.

  • Comfortable with cloud infrastructure and large-scale data processing — you don't need to be a distributed-systems engineer, but you've shipped jobs that ran on thousands of GPU-hours of video.

  • Bias toward data and infrastructure: you reach for "annotate the whole corpus" before "fine-tune another model."

Nice to have
  • Experience training vision or multimodal models from scratch (not just calling APIs).

  • ML/AI research background — papers, citations, or a research org on your resume.

  • Hands-on time with big-data frameworks like Spark, Ray, or Daft.

  • Worked on embeddings, retrieval, or content-aware search at scale.

  • Experience designing labeling taxonomies or running annotation programs.

Perks & Benefits
  • In-person, tight-knit team — 4 days/week in our SF Mission office.

  • Competitive comp and meaningful startup equity.

  • Catered lunches and dinners for SF employees.

  • Commuter benefit.

  • Team-building events and poker nights.

  • Health, vision, and dental coverage.

  • Flexible PTO.

  • Latest Apple equipment.

  • 401(k) plan with match.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer, Multimodal Data
Research Engineer, Multimodal Data

Eventual • San Francisco (CA)

On-site
USD 120,000 - 150,000
Catered lunches and dinners
Commuter benefit
Health, vision, and dental coverage
+2
Software Engineer, Product
Software Engineer, Product

Alumni Ventures • San Francisco (CA)

On-site
USD 150,000 - 250,000
Competitive compensation and startup equity
Catered lunches and dinners
Commuter benefit
+3
Software Engineer, High Performance Computing
Software Engineer, High Performance Computing

Eventual • United States

On-site
USD 100,000 - 130,000
Catered lunches and dinners
Flexible PTO
Health, vision, and dental coverage
+1
Dataloading Systems Engineer — GPU-Scale Data Pipelines
Dataloading Systems Engineer — GPU-Scale Data Pipelines

Eventual • United States

On-site
Software Engineer, High Performance Computing
Software Engineer, High Performance Computing

Eventual Inc. • San Francisco (CA)

On-site
USD 180,000 - 240,000
In-person, tight-knit team — 4 days a‑
Software Engineer, Large-Scale Data Query Systems
Software Engineer, Large-Scale Data Query Systems

Doist • San Francisco (CA)

On-site
USD 180,000 - 240,000
In-person 4 days/week in office
Competitive comp and startup equity
Catered lunches and dinners for SF
+6
Software Engineer, Large-Scale Data Query Systems
Software Engineer, Large-Scale Data Query Systems

Eventual • San Francisco (CA)

On-site
USD 150,000 - 250,000
Competitive pay
Startup equity
Catered meals and dinners
+6
Research, Vision Expertise
Research, Vision Expertise

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Research Engineer / Research Scientist, Vision
Research Engineer / Research Scientist, Vision

United States Digital Space LLC • New York (NY)

Hybrid
USD 350,000 - 850,000
Competitive compensation
Generous vacation and parental leave
Flexible working hours
Research Vision Expertise
Research Vision Expertise

Thinking Machines Lab • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1