Research Intern: Vision-Based Assessment of Human Proficiency and Skill Through Force Estimation - Honda Research Institute USA

Research Intern: Vision-Based Assessment of Human Proficiency and Skill Through Force Estimation

Your application is being processed

Research Intern: Vision-Based Assessment of Human Proficiency and Skill Through Force Estimation

Job Number: P25INT-72
​Honda Research Institute USA (HRI-US) is seeking a highly motivated and independent intern to work on human action understanding and proficiency estimation in procedural videos. This role centers on recovering physical interaction signals while performing manual tasks directly from multicamera video. The intern will build the multimodal capture setup, collect a dataset with synchronized ground-truth measurements, and use it to develop vision-based models and evaluation metrics.
San Jose, CA

 

Key Responsibilities

 

  • ​Design the data collection protocol: which procedural tasks to capture, and what variations to introduce so the resulting dataset spans a meaningful range of proficiency levels.
  • Build and validate the capture setup, Egocentric glasses and force sensors, and their data pipelines (egocentric video, gaze, hand pose, IMU, forces).
  • Synchronize egocentric streams with exocentric camera views and with the force sensor measurements.
  • Run the data collection: both as a participant yourself, wearing the glasses and force sensors while performing procedural tasks, and by recording sessions with other participants. This is human data collection, not robot demonstration.
  • Develop vision-based baselines for force estimation, optionally incorporating gaze, IMU, and pose, and evaluate them against the measured ground-truth forces.
  • Design and validate metrics that capture task proficiency from the estimated and measured signals.
  • Optionally contribute to research publications and dataset releases (e.g., CVPR, ICCV, ECCV, NeurIPS).

  This is an engineering-heavy role, but the data collection is a means to understanding proficiency, not the goal in itself. It calls for research judgment throughout: deciding which variations matter, what meaningfully captures proficiency, and which baselines are worth building.

 

Minimum Qualifications

 

  • ​Currently enrolled in an MS or PhD program in Robotics, Computer Vision, Machine Learning, Artificial Intelligence, or a closely related field.
  • Hands-on experience building or operating multimodal data collection setups (e.g., wearable or head-mounted cameras, multi-camera rigs, IMUs, motion capture, or force/tactile sensors), including sensor calibration and time synchronization across streams.
  • Strong programming skills in Python with proficiency in PyTorch, and the ability to write clean, efficient, and reproducible code.
  • Comfort with hardware-in-the-loop debugging: device SDKs, drivers, serial/USB interfaces, data logging, and end-to-end troubleshooting of recording pipelines.
  • Willingness to participate directly in data collection, using egocentric glasses and force sensors and performing procedural tasks, as well as running recording sessions with other participants.
  • Ability to independently drive a project from data collection protocol design through dataset capture to baseline models and evaluation.
  • Strong written and verbal communication skills.

 

Bonus Qualifications

  • Experience with egocentric glasses or wearable head-mounted cameras, including Machine Perception Services outputs such as eye gaze, hand tracking, and SLAM/device pose.
  • Experience with force, pressure, or tactile sensing hardware (e.g., force sensors, force-sensitive resistors, load cells), including calibration, drift compensation, and noise handling.
  • Experience with multi-camera setups: intrinsic and extrinsic calibration, hardware or software synchronization, and egocentric–exocentric spatial and temporal alignment.
  • Prior work with egocentric or procedural human activity datasets (e.g., Ego4D, Ego-Exo4D, Assembly101, EPIC-KITCHENS) and their annotation formats and toolchains.
  • Experience inferring physical quantities from vision (e.g., force, contact, or pressure estimation), or in hand–object interaction and contact modeling.
  • Familiarity with skill and proficiency assessment, action quality assessment, or temporal action segmentation in procedural videos.
  • Experience with video understanding, including long-form video or temporal modeling, and familiarity with video encoders (e.g., VideoMAE, V-JEPA, TimeSformer) and/or fine-tuning approaches.
  • Experience designing human-subject data collection protocols, including participant instructions, consent, and IRB or equivalent review processes.
  • Publications at top-tier venues (e.g., CVPR, ICCV, ECCV, NeurIPS, ICRA, IROS).

 

Years of Work Experience Required   0
Desired Start Date  1/11/2027
Internship Duration  3 Months
Position Keywords  Video Action Segmentation, Human Action Understanding 

Alternate Way to Apply

Send an e-mail to careers@honda-ri.com with the following:
- Subject line including the job number(s) you are applying for 
- Recent CV 
- A cover letter highlighting relevant background (Optional)

Please, do not contact our office to inquiry about your application status.