Research Intern: Vision-Based Assessment of Human Proficiency and Skill Through Force Estimation

Honda Research Institute USA

San Jose (CA)

On-site

USD 15,000 - 25,000

Part time

8 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Honda Research Institute USA in San Jose, CA seeks a Research Intern to develop vision-based models for assessing human proficiency and skill via force estimation. You will design data collection protocols, build multimodal capture setups, and create baselines using egocentric and exocentric video streams with force data.

The role emphasizes hands-on data collection and evaluation of vision-based approaches, with potential for publications.

Qualifications

  • Currently enrolled in an MS or PhD program in Robotics, Computer Vision, Machine Learning, AI, or a closely related field.
  • Hands-on experience building or operating multimodal data collection setups with calibration and time synchronization across streams.
  • Strong programming skills in Python with PyTorch and writing clean, reproducible code.
  • Comfort with hardware-in-the-loop debugging: device SDKs, drivers, serial/USB interfaces, data logging, and end-to-end troubleshooting of recording pipelines.
  • Willingness to participate directly in data collection and in recording sessions with participants.
  • Ability to drive a project from data collection protocol design through dataset capture to baseline models and evaluation.
  • Strong written and verbal communication skills.

Responsibilities

  • Design the data collection protocol and determine task variations to span proficiency levels.
  • Build and validate capture setup (egocentric glasses, force sensors) and data pipelines.
  • Synchronize egocentric streams with exocentric views and force measurements.
  • Conduct data collection sessions with participants and/or researchers.
  • Develop vision-based baselines for force estimation and evaluate against ground-truth forces.
  • Design and validate metrics to capture task proficiency from signals.
  • Optionally contribute to publications and dataset releases.

Skills

Python
PyTorch
Multimodal data collection
Sensor calibration
Time synchronization
Strong communication
Project management

Education

MS or PhD in Robotics, Computer Vision, ML or related field

Tools

Device SDKs

Job description

Research Intern: Vision-Based Assessment of Human Proficiency and Skill Through Force Estimation - Honda Research Institute USA
Research Intern: Vision-Based Assessment of Human Proficiency and Skill Through Force Estimation

Job Number: P25INT-72

Honda Research Institute USA (HRI-US) is seeking a highly motivated and independent intern to work on human action understanding and proficiency estimation in procedural videos. This role centers on recovering physical interaction signals while performing manual tasks directly from multicamera video. The intern will build the multimodal capture setup, collect a dataset with synchronized ground-truth measurements, and use it to develop vision-based models and evaluation metrics.

San Jose, CA

Key Responsibilities

  • Design the data collection protocol: which procedural tasks to capture, and what variations to introduce so the resulting dataset spans a meaningful range of proficiency levels.
  • Build and validate the capture setup, Egocentric glasses and force sensors, and their data pipelines (egocentric video, gaze, hand pose, IMU, forces).
  • Synchronize egocentric streams with exocentric camera views and with the force sensor measurements.
  • Run the data collection: both as a participant yourself, wearing the glasses and force sensors while performing procedural tasks, and by recording sessions with other participants. This is human data collection, not robot demonstration.
  • Develop vision-based baselines for force estimation, optionally incorporating gaze, IMU, and pose, and evaluate them against the measured ground-truth forces.
  • Design and validate metrics that capture task proficiency from the estimated and measured signals.
  • Optionally contribute to research publications and dataset releases (e.g., CVPR, ICCV, ECCV, NeurIPS).

This is an engineering-heavy role, but the data collection is a means to understanding proficiency, not the goal in itself. It calls for research judgment throughout: deciding which variations matter, what meaningfully captures proficiency, and which baselines are worth building.

Minimum Qualifications

  • Currently enrolled in an MS or PhD program in Robotics, Computer Vision, Machine Learning, Artificial Intelligence, or a closely related field.
  • Hands-on experience building or operating multimodal data collection setups (e.g., wearable or head-mounted cameras, multi-camera rigs, IMUs, motion capture, or force/tactile sensors), including sensor calibration and time synchronization across streams.
  • Strong programming skills in Python with proficiency in PyTorch, and the ability to write clean, efficient, and reproducible code.
  • Comfort with hardware-in-the-loop debugging: device SDKs, drivers, serial/USB interfaces, data logging, and end-to-end troubleshooting of recording pipelines.
  • Willingness to participate directly in data collection, using egocentric glasses and force sensors and performing procedural tasks, as well as running recording sessions with other participants.
  • Ability to independently drive a project from data collection protocol design through dataset capture to baseline models and evaluation.
  • Strong written and verbal communication skills.

Bonus Qualifications

  • Experience with egocentric glasses or wearable head-mounted cameras, including Machine Perception Services outputs such as eye gaze, hand tracking, and SLAM/device pose.
  • Experience with force, pressure, or tactile sensing hardware (e.g., force sensors, force-sensitive resistors, load cells), including calibration, drift compensation, and noise handling.
  • Experience with multi-camera setups: intrinsic and extrinsic calibration, hardware or software synchronization, and egocentric–exocentric spatial and temporal alignment.
  • Prior work with egocentric or procedural human activity datasets (e.g., Ego4D, Ego-Exo4D, Assembly101, EPIC-KITCHENS) and their annotation formats and toolchains.
  • Experience inferring physical quantities from vision (e.g., force, contact, or pressure estimation), or in hand–object interaction and contact modeling.
  • Familiarity with skill and proficiency assessment, action quality assessment, or temporal action segmentation in procedural videos.
  • Experience with video understanding, including long-form video or temporal modeling, and familiarity with video encoders (e.g., VideoMAE, V-JEPA, TimeSformer) and/or fine-tuning approaches.
  • Experience designing human-subject data collection protocols, including participant instructions, consent, and IRB or equivalent review processes.
  • Publications at top-tier venues (e.g., CVPR, ICCV, ECCV, NeurIPS, ICRA, IROS).
Years of Work Experience Required 0

Desired Start Date 1/11/2027

Internship Duration 3 Months

Position Keywords Video Action Segmentation, Human Action Understanding

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Intern: Proficiency Estimation in Procedural Videos
Research Intern: Proficiency Estimation in Procedural Videos

Honda Research Institute USA • San Jose (CA)

On-site
USD 23,000 - 37,000
Vision-Based Proficiency Research Intern — Multimodal Data
Vision-Based Proficiency Research Intern — Multimodal Data

Honda Research Institute USA • San Jose (CA)

On-site
USD 15,000 - 25,000
Research Scientist: Human-Centric Visual Intelligence ...
Research Scientist: Human-Centric Visual Intelligence ...

Honda Research Institute USA, Inc. • San Jose (CA)

On-site
USD 100,000 - 130,000
Research Intern: Test-Time Adaptation For Embodied Agents
Research Intern: Test-Time Adaptation For Embodied Agents

Honda Research Institute USA • San Jose (CA)

On-site
USD 33,000 - 50,000
Research Intern: Memory Representation for Dexterous Manipulation
Research Intern: Memory Representation for Dexterous Manipulation

Honda Research Institute USA • San Jose (CA)

On-site
USD 20,000 - 27,000
Research Intern: Semantic-Aware Interaction and Behavior Planning for Autonomous Driving
Research Intern: Semantic-Aware Interaction and Behavior Planning for Autonomous Driving

Honda Research Institute USA • Mountain View (CA)

On-site
USD 34,000 - 52,000
Research Intern: Meta-Cognition and Internal Mechanisms for Multi-Modal Foundation Model-based Agent
Research Intern: Meta-Cognition and Internal Mechanisms for Multi-Modal Foundation Model-based Agent

Honda Research Institute USA • San Jose (CA)

On-site
USD 13,225,000 - 19,837,000
Scene Understanding for Autonomous Mobility
Scene Understanding for Autonomous Mobility

Honda Research Institute USA, Inc. • San Jose (CA)

On-site
USD 100,000 - 130,000
Research Scientist: Multimodal Learning for Embodied ...
Research Scientist: Multimodal Learning for Embodied ...

Honda Research Institute USA, Inc. • San Jose (CA)

On-site
USD 90,000 - 130,000
Scene Understanding for Autonomous Mobility
Scene Understanding for Autonomous Mobility

Honda Research Institute USA • Mountain View (CA)

On-site
USD 80,000 - 120,000