Research Intern: Proficiency Estimation in Procedural Videos

Honda Research Institute USA

San Jose (CA)

On-site

USD 23,000 - 37,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Honda Research Institute USA in San Jose, CA invites a PhD research intern to advance proficiency estimation in long procedural videos. The project explores action recognition, error detection, and skill assessment under narrated guidance, with emphasis on multimodal video-language alignment and model fine-tuning.

The role targets independent researchers with strong publication records, capable of driving ideation to experimentation and contributing to high-impact studies and potential patents.

Qualifications

  • PhD student in Computer Vision, Machine Learning, AI, or related field.
  • Excellent publication record in top-tier conferences.
  • Experience with multimodal language models and video understanding.
  • Strong programming skills and PyTorch proficiency.
  • Strong written and verbal communication skills.
  • Ability to independently drive research from ideation to publication.

Responsibilities

  • Conduct cutting-edge research in learning from narration to evaluate proficiency in procedural videos.
  • Design and implement algorithms for aligning video and text descriptions and training multi-modal language models.
  • Perform literature review, run experiments, and analyze results.
  • Lead or contribute to research paper writing and potential conference submissions.
  • Write reproducible research code using deep learning frameworks like PyTorch.

Job description

Research Intern: Proficiency Estimation in Procedural Videos - Honda Research Institute USA

Job Number: P25INT-62

Honda Research Institute USA (HRI-US) is seeking a highly motivated and independent PhD research intern to join our team in advancing the frontiers of human action understanding and computer vision in long procedural videos. The project focuses on human proficiency estimation in procedural tasks. The research associated with this position involves recognizing actions and errors, and evaluating human skill supervised by narration in videos. This role is ideal for a researcher with a strong background in video understanding and vision-language models. The intern will work on real-world challenges involving long-horizon human activity videos, and contribute to high-impact publications and patents.

San Jose, CA

Key Responsibilities

  • Conduct cutting-edge research in learning from narration to learn and evaluate proficiency in procedural videos.
  • Design and implement novel algorithms for aligning video and text descriptions and training/finetuning multi-modal language models.
  • Perform literature review, formulate hypotheses, run experiments, and analyze results
  • Lead or contribute to research paper writing, including potential submission to top-tier computer vision or machine learning conferences (e.g., CVPR, ICCV, NeurIPS, ECCV).
  • Write well-structured, efficient code using deep learning frameworks such as PyTorch.

Minimum Qualifications

  • Currently enrolled in a PhD program in Computer Vision, Machine Learning, Artificial Intelligence, or a closely related field.
  • Publication record in top-tier conferences (e.g., CVPR, ICCV, ECCV, WACV, NeurIPS, ICLR).
  • Prior experience with multimodal language models (i.e, Q formers, LoRA, and LLMs) in video understanding, OR video-language representation alignment (e.g., CLIP).
  • Previous publication in a problem involving procedural videos.
  • Excellent programming skills, ability to write reproducible research code, and proficiency in deep learning frameworks, especially PyTorch.
  • Strong written and verbal communication skills.
  • Ability to independently drive research, from ideation to experimentation and publication.

Bonus Qualifications

  • Previous publication experience in any of the following areas:
    • Previous experience in (hand/body) pose estimation or its application in videos.
    • Diffusion model
    • Video or video-text alignment
    • Sequence modeling
    • Human-object interaction
    • Temporal action segmentation, error detection or skill assessment in videos.

Years of Work Experience Required 0

Desired Start Date 1/11/2027

Internship Duration 3 Months

Position Keywords Learning from Narration, Action Understanding, Long Video Understanding, Skill Assessment, Proficiency Estimation, Error Detection and Recognition

Candidates must have the legal right to work in the U.S.A.

For California applicants: Please review the Honda California Job Applicant Privacy Notice before submitting your application.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Intern: Vision-Based Assessment of Human Proficiency and Skill Through Force Estimation
Research Intern: Vision-Based Assessment of Human Proficiency and Skill Through Force Estimation

Honda Research Institute USA • San Jose (CA)

On-site
USD 15,000 - 25,000
PhD Research Intern: Proficiency in Procedural Video ML
PhD Research Intern: Proficiency in Procedural Video ML

Honda Research Institute USA • San Jose (CA)

On-site
USD 23,000 - 37,000
Research Scientist: Human-Centric Visual Intelligence ...
Research Scientist: Human-Centric Visual Intelligence ...

Honda Research Institute USA, Inc. • San Jose (CA)

On-site
USD 100,000 - 130,000
Research Intern: Semantic-Aware Interaction and Behavior Planning for Autonomous Driving
Research Intern: Semantic-Aware Interaction and Behavior Planning for Autonomous Driving

Honda Research Institute USA • Mountain View (CA)

On-site
USD 34,000 - 52,000
Research Intern: Closed-loop E2E Driving
Research Intern: Closed-loop E2E Driving

Honda Research Institute USA • Mountain View (CA)

On-site
USD 35,000 - 55,000
Vision-Based Proficiency Research Intern — Multimodal Data
Vision-Based Proficiency Research Intern — Multimodal Data

Honda Research Institute USA • San Jose (CA)

On-site
USD 15,000 - 25,000
Research Intern: Memory Representation for Dexterous Manipulation
Research Intern: Memory Representation for Dexterous Manipulation

Honda Research Institute USA • San Jose (CA)

On-site
USD 20,000 - 27,000
Research Intern: Meta-Cognition and Internal Mechanisms for Multi-Modal Foundation Model-based Agent
Research Intern: Meta-Cognition and Internal Mechanisms for Multi-Modal Foundation Model-based Agent

Honda Research Institute USA • San Jose (CA)

On-site
USD 13,225,000 - 19,837,000
Scene Understanding for Autonomous Mobility
Scene Understanding for Autonomous Mobility

Honda Research Institute USA • Mountain View (CA)

On-site
USD 80,000 - 120,000
Scene Understanding for Autonomous Mobility
Scene Understanding for Autonomous Mobility

Honda Research Institute USA, Inc. • San Jose (CA)

On-site
USD 100,000 - 130,000