Your application is being processed
Research Intern: Proficiency Estimation in Procedural Videos
Job Number: P25INT-62
Honda Research Institute USA (HRI-US) is seeking a highly motivated and independent PhD research intern to join our team in advancing the frontiers of human action understanding and computer vision in long procedural videos. The project focuses on human proficiency estimation in procedural tasks. The research associated with this position involves recognizing actions and errors, and evaluating human skill supervised by narration in videos. This role is ideal for a researcher with a strong background in video understanding and vision-language models. The intern will work on real-world challenges involving long-horizon human activity videos, and contribute to high-impact publications and patents.
San Jose, CA
|
Key Responsibilities
|
|
- Conduct cutting-edge research in learning from narration to learn and evaluate proficiency in procedural videos.
- Design and implement novel algorithms for aligning video and text descriptions and training/finetuning multi-modal language models.
- Perform literature review, formulate hypotheses, run experiments, and analyze results
- Lead or contribute to research paper writing, including potential submission to top-tier computer vision or machine learning conferences (e.g., CVPR, ICCV, NeurIPS, ECCV).
- Write well-structured, efficient code using deep learning frameworks such as PyTorch.
Minimum Qualifications
|
|
- Currently enrolled in a PhD program in Computer Vision, Machine Learning, Artificial Intelligence, or a closely related field.
- Publication record in top-tier conferences (e.g., CVPR, ICCV, ECCV, WACV, NeurIPS, ICLR).
- Prior experience with multimodal language models (i.e, Q formers, LoRA, and LLMs) in video understanding, OR video-language representation alignment (e.g., CLIP).
- Previous publication in a problem involving procedural videos.
- Excellent programming skills, ability to write reproducible research code, and proficiency in deep learning frameworks, especially PyTorch.
- Strong written and verbal communication skills.
- Ability to independently drive research, from ideation to experimentation and publication.
Bonus Qualifications
- Previous publication experience in any of the following areas:
- Previous experience in (hand/body) pose estimation or its application in videos.
- Diffusion model
- Video or video-text alignment
- Sequence modeling
- Human-object interaction
- Temporal action segmentation, error detection or skill assessment in videos.
|
| Years of Work Experience Required |
0 |
| Desired Start Date |
1/11/2027 |
| Internship Duration |
3 Months |
| Position Keywords |
Learning from Narration, Action Understanding, Long Video Understanding, Skill Assessment, Proficiency Estimation, Error Detection and Recognition |
|
|
|
Alternate Way to Apply
Send an e-mail to careers@honda-ri.com with the following:
- Subject line including the job number(s) you are applying for
- Recent CV
- A cover letter highlighting relevant background (Optional)
Please, do not contact our office to inquiry about your application status.