Research Scientist, Video Understanding & World Models

Mecka AI

New York (NY)

On-site

USD 100,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mecka AI in New York is seeking a Research Scientist specializing in video understanding to lead the development of video representation models. The role involves training large-scale models and fine-tuning video-language models with an emphasis on real-world robotics data.

The ideal candidate will have extensive experience with PyTorch and modern video representation techniques. This position presents a unique opportunity to influence production systems and contribute to cutting-edge research in egocentric embodied video.

Qualifications

  • Deep experience in multi-GPU or distributed training.
  • Strong understanding of multimodal modeling.
  • Experience with egocentric/embodied datasets.

Responsibilities

  • Own model architecture and training strategy.
  • Train and fine-tune video encoders and models.
  • Turn checkpoints into usable artifacts.

Skills

Deep experience training large models in PyTorch
Understanding of modern video representation learning
Ability to run rigorous experiments
Experience with video VLMs
Strong software engineering discipline

Job description

About the Role

We are looking for a Research Scientist, Video Understanding to own Mecka’s video understanding agenda end-to-end: train large-scale video representation and video-language models on our egocentric + stereo corpus, and turn the resulting checkpoints into production signals the rest of the stack ships on.

This role is focused on large model training, video encoders, video-language models, VLMs/VLAs, and temporal representation learning on real-world robotics data.

What You'll Work On
Large-Scale Training & Architecture
  • Own model architecture and training strategy across Mecka’s task families (manipulation, locomotion, daily activity, long-horizon behavior).
  • Run self-supervised and multimodal pretraining (VideoMAE / VJEPA / VideoPrism / InternVideo-class) with rigorous evals and clean ablations.
Video-Language & Multimodal Modeling
  • Train and fine-tune video encoders and video-language models (temporal transformers, joint-embedding models, contrastive objectives, masked modeling, instruction/video alignment).
  • Incorporate useful priors (pose, depth, camera motion, optical flow) when it improves representation quality.
Research - Production Signals
  • Turn checkpoints into usable artifacts: embeddings and model outputs that downstream systems can reliably consume (retrieval, labeling, QA, analytics).
  • Build a disciplined training + eval workflow with regression tracking and reproducible runs.
Who You Are
Required Background
  • Deep experience training large models in PyTorch (or equivalent), including multi-GPU or distributed training.
  • Strong understanding of modern video representation learning and/or multimodal modeling.
  • Ability to run rigorous experiments and communicate results clearly.
Strong Signals:
  • Experience with video VLMs / VLA-adjacent systems (VideoCLIP, InstructBLIP-Video, LLaVA-Video-class).
  • Experience with egocentric / embodied datasets (Ego4D, EgoExo4D, EPIC-Kitchens, Something-Something).
  • Strong software engineering discipline: you write research code that can be shipped.
Why This Role
  • Work on a domain - egocentric embodied video - where data is scarce everywhere except here.
  • Own a research agenda that directly feeds production systems and product outcomes.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Scientist, Video Understanding & World Models
Research Scientist, Video Understanding & World Models

Kindredventures • New York (NY)

On-site
USD 120,000 - 150,000
Research Scientist, Video Understanding & World Models
Research Scientist, Video Understanding & World Models

Mecka AI • New York (NY)

On-site
USD 100,000 - 130,000
Lead Research Scientist, Video Understanding & World Models
Lead Research Scientist, Video Understanding & World Models

Kindredventures • New York (NY)

On-site
USD 120,000 - 150,000
Research Scientist, SLAM & VIO
Research Scientist, SLAM & VIO

Mecka • New York (NY)

On-site
USD 100,000 - 140,000
Access to proprietary data
Cutting-edge research environment
Opportunity for high impact
Machine Learning Engineer (Video Understanding & Segmentation)
Machine Learning Engineer (Video Understanding & Segmentation)

Glint Tech Solutions • Santa Clara (CA)

On-site
USD 140,000 - 210,000
Research Scientist (Spatial AI & Neural Reconstruction)
Research Scientist (Spatial AI & Neural Reconstruction)

Mecka AI • New York (NY)

On-site
USD 120,000 - 150,000
Member of Technical Staff, Vision/Language
Member of Technical Staff, Vision/Language

XDOF • San Mateo (CA)

On-site
USD 130,000 - 210,000
Research Scientist (Spatial AI & Neural Reconstruction)
Research Scientist (Spatial AI & Neural Reconstruction)

Kindredventures • New York (NY)

On-site
USD 120,000 - 160,000
Research Member of Technical Staff - Data & Evaluation
Research Member of Technical Staff - Data & Evaluation

Rhoda AI • Mountain View (CA)

On-site
USD 120,000 - 160,000
Machine Learning Engineer
Machine Learning Engineer

Human Archive • San Francisco (CA)

On-site
USD 120,000 - 160,000