Staff Data Engineer, Multimodal AI Pipelines

Veeda

California (MO)

On-site

USD 130,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We seek a Member of Technical Staff - Data to design and operate large-scale data pipelines ingesting video, lidar, and robot trajectories using Ray Data and Daft, with GPU-accelerated decoding and sequential streaming layouts.

The role requires strong Python skills, experience with distributed pipelines over hundreds of terabytes, and a focus on provenance, licensing constraints, and data quality.

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, or equivalent hands-on data engineering experience.
  • Built and operated distributed data pipelines (e.g., Ray Data, Daft, Spark) over hundreds of terabytes, with rigor in idempotency, backfills, and schema evolution.
  • Strong Python skills and comfortable in the video stack (codecs, containers, ffmpeg, GPU decode) and in columnar and object storage formats.
  • Able to design and defend a data mixture empirically, running curation ablations that measure downstream model quality.
  • Experience working inside real licensing constraints on what may and may not be trained on, with provenance treated as a hard requirement.

Responsibilities

  • Multimodal Ingest: Build ingest for video, lidar, and robot trajectories on Ray Data and Daft, with GPU decode (NVDEC, DALI) and resharding into WebDataset and Lance layouts that stream sequentially rather than seeking per sample.
  • Curation & Filtering: Decide what earns a slot using blur, exposure, and camera-trajectory scoring plus embedding deduplication over cuVS indexes, and prove each filter with a downstream ablation, not a dataset-size delta.
  • Annotation & Auto-Labeling: Produce the labels the models need, such as VLM captions, camera pose from feed-forward reconstruction (VGGT, MASt3R), and depth and segmentation pseudo-labels, and hold each to a measured error rate against human review.
  • Real & Synthetic Interop: Normalize episodic data across formats such as LeRobotDataset v3, Open X-Embodiment, and RLDS, reconciling action spaces, control rates, and frame timing, and account for the simulated share of every training mixture.
  • Provenance, Licensing & Governance: Track license terms, restricted-source flags, and C2PA content credentials at source granularity, and version datasets as immutable manifests so any checkpoint traces back to the exact bytes that trained it.

Skills

Python
Data pipelines
Video processing

Education

Bachelor's degree in Computer Science or related

Tools

Ray Data
Daft
Spark

Job description

Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We seek a Member of Technical Staff - Data to design and operate large-scale data pipelines ingesting video, lidar, and robot trajectories using Ray Data and Daft, with GPU-accelerated decoding and sequential streaming layouts.

The role requires strong Python skills, experience with distributed pipelines over hundreds of terabytes, and a focus on provenance, licensing constraints, and data quality.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Technical Staff Engineer - World Models & Multimodal AI
Technical Staff Engineer - World Models & Multimodal AI

Veeda • California (MO)

On-site
USD 130,000 - 200,000
Member of Technical Staff - Data
Member of Technical Staff - Data

Veeda • California (MO)

On-site
USD 130,000 - 180,000
AI Infrastructure Engineer: HPC GPU Clusters
AI Infrastructure Engineer: HPC GPU Clusters

Veeda AI • Seattle (WA)

On-site
USD 180,000 - 240,000
AI Infrastructure Architect for HPC & GPU Clusters
AI Infrastructure Architect for HPC & GPU Clusters

Veeda • California (MO)

On-site
USD 150,000 - 200,000
Staff Engineer, World Models for Embodied AI
Staff Engineer, World Models for Embodied AI

Veeda AI • Seattle (WA)

On-site
USD 150,000 - 230,000
ML Ops Engineer: Reproducible Runs & Scalable Tools
ML Ops Engineer: Reproducible Runs & Scalable Tools

Veeda • California (MO)

On-site
USD 120,000 - 190,000
Staff Engineer, Physical AI Simulation & World Models
Staff Engineer, Physical AI Simulation & World Models

Veeda Innovation • California (MO), Northern (KY)

Hybrid
USD 140,000 - 190,000
ML Systems Engineer — Performance & Scale
ML Systems Engineer — Performance & Scale

Veeda Innovation • Northern (KY)

Hybrid
USD 150,000 - 230,000
Remote AI Data Engineer: Scalable Pipelines & ML Workloads
Remote AI Data Engineer: Scalable Pipelines & ML Workloads

Bright Vision Technologies • Charlotte (NC)

On-site
USD 100,000 - 150,000
Member of Technical Staff — Data Ingestion & Quality
Member of Technical Staff — Data Ingestion & Quality

Causal Labs • San Francisco (CA)

On-site
USD 120,000 - 170,000