Member of Technical Staff - Data

Veeda AI

Toronto

On-site

CAD 120,000 - 180,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Veeda AI in Toronto is seeking a Member of Technical Staff - Data to help build and manage large-scale multimodal datasets for physical AI research.

You will design and operate distributed data pipelines, ingest video, lidar, and robot trajectories, ensure data quality with rigorous curation and provenance, and collaborate with researchers to align datasets with model development needs.

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, or equivalent hands-on experience in large-scale data engineering.
  • Built and operated distributed data pipelines (e.g., Ray Data, Daft, Spark) over hundreds of terabytes, with rigor in idempotency, backfills, and schema evolution.
  • Strong Python skills and comfortable in the video stack (codecs, containers, ffmpeg, GPU decode) and in columnar and object storage formats.
  • Able to design and defend a data mixture empirically, running curation ablations that measure downstream model quality.
  • Experience working inside real licensing constraints on what may and may not be trained on, with provenance treated as a hard requirement.

Responsibilities

  • Multimodal Ingest: Build ingest for video, lidar, and robot trajectories on Ray Data and Daft, with GPU decode (NVDEC, DALI) and resharding into WebDataset and Lance layouts that stream sequentially rather than seeking per sample.
  • Curation & Filtering: Decide what earns a slot using blur, exposure, and camera-trajectory scoring plus embedding deduplication over cuVS indexes, and prove each filter with a downstream ablation, not a dataset-size delta.
  • Annotation & Auto-Labeling: Produce the labels the models need, such as VLM captions, camera pose from feed-forward reconstruction (VGGT, MASt3R), and depth and segmentation pseudo-labels, and hold each to a measured error rate against human review.
  • Real & Synthetic Interop: Normalize episodic data across formats such as LeRobotDataset v3, Open X-Embodiment, and RLDS, reconciling action spaces, control rates, and frame timing, and account for the simulated share of every training mixture.
  • Provenance, Licensing & Governance: Track license terms, restricted-source flags, and C2PA content credentials at source granularity, and version datasets as immutable manifests so any checkpoint traces back to the exact bytes that trained it.

Skills

Python programming
Data engineering
Data governance
Experimentation & ablations

Education

Bachelor's degree in Computer Science or Computer Engineering or equivalent hands-on data engineering experience

Tools

Ray Data
Daft
Spark
ffmpeg
GPU decode
Open X-Embodiment

Job description

Member of Technical Staff - Data
About Us

Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence. If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the opportunity to make an outsized impact from day one.

Responsibilities
  • Multimodal Ingest: Build ingest for video, lidar, and robot trajectories on Ray Data and Daft, with GPU decode (NVDEC, DALI) and resharding into WebDataset and Lance layouts that stream sequentially rather than seeking per sample.

  • Curation & Filtering: Decide what earns a slot using blur, exposure, and camera-trajectory scoring plus embedding deduplication over cuVS indexes, and prove each filter with a downstream ablation, not a dataset-size delta.

  • Annotation & Auto-Labeling: Produce the labels the models need, such as VLM captions, camera pose from feed-forward reconstruction (VGGT, MASt3R), and depth and segmentation pseudo-labels, and hold each to a measured error rate against human review.

  • Real & Synthetic Interop: Normalize episodic data across formats such as LeRobotDataset v3, Open X-Embodiment, and RLDS, reconciling action spaces, control rates, and frame timing, and account for the simulated share of every training mixture.

  • Provenance, Licensing & Governance: Track license terms, restricted-source flags, and C2PA content credentials at source granularity, and version datasets as immutable manifests so any checkpoint traces back to the exact bytes that trained it.

Requirements
  • Bachelor's degree in Computer Science, Computer Engineering, or equivalent hands-on experience in large-scale data engineering.

  • Built and operated distributed data pipelines (e.g., Ray Data, Daft, Spark) over hundreds of terabytes, with rigor in idempotency, backfills, and schema evolution.

  • Strong Python skills and comfortable in the video stack (codecs, containers, ffmpeg, GPU decode) and in columnar and object storage formats.

  • Able to design and defend a data mixture empirically, running curation ablations that measure downstream model quality.

  • Experience working inside real licensing constraints on what may and may not be trained on, with provenance treated as a hard requirement.

Nice to Have
  • Experience with robot trajectory formats and tooling such as LeRobot, RLDS, ROS 2 bags, or MCAP.

  • Experience with sensor calibration, hardware time synchronization, and non-pinhole camera models such as fisheye or ftheta.

  • Experience running GPU-accelerated curation with RAPIDS or NeMo Curator.

  • Experience building PII, face, and plate redaction into a video pipeline at scale.

  • Managed annotation vendors and built the QA statistics that keep them honest.

  • Built lakehouse storage on Iceberg or Delta and reduced object-storage cost without losing read throughput.

  • Published on data curation or contributed to open-source data tooling such as DataTrove or video2dataset.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff - World Models
Member of Technical Staff - World Models

Veeda Innovation • Toronto

On-site
CAD 120,000 - 180,000
Member of Technical Staff - World Models
Member of Technical Staff - World Models

Veeda AI • Toronto

On-site
CAD 120,000 - 180,000
Member of Technical Staff - Robotics
Member of Technical Staff - Robotics

Veeda AI • Toronto

On-site
CAD 110,000 - 150,000
Member of Technical Staff - ML Operations
Member of Technical Staff - ML Operations

Veeda AI • Toronto

On-site
CAD 120,000 - 150,000
Member of Technical Staff - ML Performance
Member of Technical Staff - ML Performance

Veeda AI • Toronto

On-site
CAD 120,000 - 180,000
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda AI • Toronto

On-site
CAD 120,000 - 165,000
Member of Technical Staff - Simulation
Member of Technical Staff - Simulation

Veeda AI • Toronto

On-site
CAD 120,000 - 180,000
Senior Data Developer
Senior Data Developer

Visier • Vancouver

On-site
CAD 120,000 - 160,000
AI Engineer
AI Engineer

Valsoft Corporation • Canada

On-site
CAD 100,000 - 140,000
Software Engineer, ML Ops
Software Engineer, ML Ops

AeroVect Technologies Inc. • Toronto

On-site
CAD 110,000 - 150,000