Research Scientist, Video & Multimodal

Innodata

Ridgefield Park (NJ)

On-site

USD 160,000 - 185,000

Full time

15 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Innodata is hiring a Research Scientist to own the science behind video and multimodal evaluations. You will define data designs, evaluation criteria, and experimental protocols to advance video understanding and generation.

The role emphasizes rigorous experiments, cross‑functional collaboration, and publishing benchmarks and methodologies. Required are 5+ years in video or multimodal ML, a CS/EE degree (MS/PhD preferred), and strong PyTorch experience.

Qualifications

  • 5+ years of hands-on industry experience in video understanding or multimodal ML.
  • Bachelor's in CS/EE or related technical field required; MS/PhD preferred.
  • Proficient with PyTorch and model fine-tuning in video/vision-language tasks.

Responsibilities

  • Define data design, schemas, and evaluation criteria for video and multimodal models.
  • Create evaluation methodologies for temporal grounding, long-horizon reasoning, and cross-modal retrieval.
  • Develop data collection and labeling plans with annotation teams and synthetic-data pipelines.
  • Run experiments linking data choices to model performance and publish results.

Skills

Video understanding
Multimodal ML
PyTorch fundamentals
HuggingFace transformers
Datasets & benchmarks

Education

Bachelor's in CS/EE or related
MS or PhD preferred

Tools

ffmpeg
decord
WebDataset
Parquet/Arrow
HuggingFace datasets

Job description

Innodata(Nasdaq: INOD) is a global data engineering company. We believe that data and Artificial Intelligence (AI) are inextricably linked.Our mission is to enable the responsible advancement of artificial intelligence by providing the data, evaluation frameworks, and human expertise required to build AI systems that can be trusted at scale.We provide a range of transferable solutions, platforms, and services for Generative AI / AI builders and adopters. In every relationship, we honor our 36+ year legacy delivering the highest quality data and outstanding outcomes for our customers.

Scope of the Role

Video is where multimodal models are weakest and hardest to grade. Temporal reasoning, long-form understanding, grounding events in time, and holding audio, video, and text together do not fall out of image benchmarks — and the evaluations for them are still immature. Closing that gap is gated as much by how we design data and evaluation as by architecture. Innodata builds that data and those evaluations for the customers and frontier labs advancing video and multimodal models, and we are hiring a Research Scientist to own the science behind it.

What You’ll Own

You will define how Innodata designs, structures, and evaluates video data for video and multimodal models, and you will validate those choices experimentally. Concretely, you will:

  • Translate the requirements of video and multimodal models — video understanding, temporal and event localization, action recognition, long-form video, video-language models, video generation, cross-modal reasoning, and multimodal retrieval and grounding — into concrete data specifications: modalities, annotation schemas, sampling, and evaluation criteria.
  • Build evaluation methodology for video understanding — temporal grounding accuracy, long-context and long-horizon reasoning, and dynamic multi-turn, cross-modal, and retrieval-and-grounding evaluation — clear about when model-based scoring is trustworthy and when a human is needed.
  • Build evaluation methodology for video generation — fidelity, temporal coherence, and physical plausibility, including generative video used as a world model — the regime where automatic metrics are weakest and human judgment matters most.
  • Decide how existing and incoming video should be structured, enriched, and sampled to extract the most model value from it, including from messy, domain-specific footage.
  • Run experiments that prove data decisions matter: fine-tune and evaluate models on Innodata data, with ablations tying specific data choices to measurable improvement.
  • Design adversarial and stumping evaluations that surface where video and multimodal systems fail, and turn those failures into better data.
  • Publish. Turn what you learn into benchmarks, methodology, and papers that advance the field and earn the trust of the customers and frontier labs we partner with.
  • Work with annotation teams, subject-matter experts, and the synthetic-data pipeline to turn specifications into operational collection and labeling plans.
You'll Thrive in This Role If You Have
  • Roughly 5+ years of hands-on industry experience in video understanding or multimodal ML. We weight practical experience over formal credentials; a PhD with a compelling, current research agenda can offset the lower end.
  • A Bachelor's degree in computer science, electrical engineering, or a related technical or quantitative field is required; an advanced degree (MS or PhD) in a relevant field is preferred.
  • Trained and evaluated video or multimodal models yourself, with strong PyTorch fundamentals.
  • Fluency in the formats and tooling video work runs on: ffmpeg and decord pipelines, temporal and COCO-style annotation, WebDataset, Parquet and Arrow, and HuggingFace datasets.
  • Experience fine-tuning large video or vision-language models with the modern toolchain (HuggingFace transformers, PEFT, efficient inference), and with long-form video, streaming, temporal segmentation, or synthetic video generation.
  • A way of thinking in datasets and benchmarks: you have built evaluation sets, calibrated difficulty, and argued about what makes video data good for a given objective.
  • A track record the field recognizes: first-author publications or strong open-source contributions at venues such as CVPR, ICCV, ECCV, NeurIPS, or ICLR.
  • The ability to work directly with the research scientists at the customers and frontier labs we partner with, and to explain data and modeling decisions clearly to both expert and non-expert audiences, backed by a rigorous, reproducible approach to experiments and documentation.
  • Bonus: interest or hands-on experience in responsible-AI evaluation and red-teaming — safety and robustness testing for video and multimodal systems.

The expected salary range for this position is $160,000 - $185,000 p/year, based on experience, skills, and qualifications.

As set forth in Innodata Inc.’s Equal Employment Opportunity policy,we do not discriminate on the basis of any protected group status under any applicable law.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Video & Multimodal AI Research Scientist
Video & Multimodal AI Research Scientist

Innodata • Ridgefield Park (NJ)

On-site
USD 160,000 - 185,000
Research Scientist, Speech & Audio
Research Scientist, Speech & Audio

Innodata Inc. • United States

On-site
USD 160,000 - 185,000
Research Scientist, Robotics & World Models
Research Scientist, Robotics & World Models

Innodata Inc. • United States

Remote
USD 160,000 - 185,000
Robotics / Physical AI Research Engineer
Robotics / Physical AI Research Engineer

Innodata Inc. • New Jersey

Hybrid
USD 180,000 - 300,000
ML Research Engineer
ML Research Engineer

Origin Lab • Los Angeles (CA)

On-site
USD 180,000 - 280,000
AI Research Scientist
AI Research Scientist

Noösphere • Seattle (WA)

On-site
USD 150,000 - 210,000
Early-stage equity
Benefits package
Research Data Scientist
Research Data Scientist

Innodata • Austin (TX)

On-site
USD 160,000 - 185,000
Research, Vision Expertise
Research, Vision Expertise

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

On-site
USD 350,000 - 475,000
Health benefits
Dental benefits
Vision benefits
+3
Applied AI Analyst - Flexible Hours
Applied AI Analyst - Flexible Hours

Innodata Inc. • United States

Hybrid
USD 18,000 - 23,000
Member of Technical Staff, Data Engineering
Member of Technical Staff, Data Engineering

Odyssey • Palo Alto (CA)

On-site
USD 140,000 - 190,000