Research Intern: Proficiency Estimation in Procedural Videos - Honda Research Institute USA

Research Intern: Proficiency Estimation in Procedural Videos

Your application is being processed

Research Intern: Proficiency Estimation in Procedural Videos

Job Number: P25INT-62
Honda Research Institute USA (HRI-US) is seeking a highly motivated and independent PhD research intern to join our team in advancing the frontiers of human action understanding and computer vision in long procedural videos. The project focuses on human proficiency estimation in procedural tasks. The research associated with this position involves recognizing actions and errors, and evaluating human skill supervised by narration in videos. This role is ideal for a researcher with a strong background in video understanding and vision-language models. The intern will work on real-world challenges involving long-horizon human activity videos, and contribute to high-impact publications and patents.
San Jose, CA

 

Key Responsibilities

 

  • ​Conduct cutting-edge research in learning from narration to learn and evaluate proficiency in procedural videos.
  • Design and implement novel algorithms for aligning video and text descriptions and training/finetuning multi-modal language models.
  • Perform literature review, formulate hypotheses, run experiments, and analyze results
  • Lead or contribute to research paper writing, including potential submission to top-tier computer vision or machine learning conferences (e.g., CVPR, ICCV, NeurIPS, ECCV).
  • Write well-structured, efficient code using deep learning frameworks such as PyTorch.

 

Minimum Qualifications

 

  • ​Currently enrolled in a PhD program in Computer Vision, Machine Learning, Artificial Intelligence, or a closely related field.
  • Publication record in top-tier conferences (e.g., CVPR, ICCV, ECCV, WACV, NeurIPS, ICLR).
  • Prior experience with multimodal language models (i.e, Q formers, LoRA, and LLMs) in video understanding, OR video-language representation alignment (e.g., CLIP).
  • Previous publication in a problem involving procedural videos.
  • Excellent programming skills, ability to write reproducible research code, and proficiency in deep learning frameworks, especially PyTorch.
  • Strong written and verbal communication skills.
  • Ability to independently drive research, from ideation to experimentation and publication.

 

Bonus Qualifications

  • Previous publication experience in any of the following areas:
    • Previous experience in (hand/body) pose estimation or its application in videos.
    • Diffusion model
    • Video or video-text alignment
    • Sequence modeling
    • Human-object interaction
    • Temporal action segmentation, error detection or skill assessment in videos.

 

Years of Work Experience Required   0
Desired Start Date  1/11/2027
Internship Duration  3 Months
Position Keywords  Learning from Narration, Action Understanding, Long Video Understanding, Skill Assessment, Proficiency Estimation, Error Detection and Recognition 

Alternate Way to Apply

Send an e-mail to careers@honda-ri.com with the following:
- Subject line including the job number(s) you are applying for 
- Recent CV 
- A cover letter highlighting relevant background (Optional)

Please, do not contact our office to inquiry about your application status.