Research Intern: Human Action Understanding from Long Videos - Honda Research Institute USA

Research Intern: Human Action Understanding from Long Videos

Your application is being processed

Research Intern: Human Action Understanding from Long Videos

Job Number: P25INT-58
Honda Research Institute USA (HRI-US) is seeking a highly motivated and independent intern to join our team in advancing the frontiers of human action understanding. This role focuses on temporal video understanding in long, untrimmed procedural videos and sits at the intersection of applied research and engineering, with an emphasis on building efficient models for real-world systems. The intern will work on challenging problems involving egocentric and exocentric video streams. Potential project topics include, but are not limited to, developing methods for real-time/online action segmentation, as well as multi-view action understanding, including cross-view transfer and learning view-invariant representations. Additional directions include multimodal alignment (e.g., hand/body pose and video) for action segmentation and proficiency estimation in procedural videos. Depending on project outcomes, the intern may also contribute to high-impact publications and patents.
San Jose, CA

 

Key Responsibilities

 

  • Design and implement methods for human action understanding in long, untrimmed procedural videos, including temporal action segmentation and recognition, across egocentric and exocentric video. Directions include multi-view and cross-view learning, view-invariant representation learning, multimodal alignment (e.g., hand and body pose with video), and online/streaming inference under latency constraints.
  • Develop scalable video models suited to real-world deployment: knowledge distillation across views, modalities, and architectures, lightweight architecture design, and analysis of accuracy, latency, and memory trade-offs.
  • Benchmark models rigorously, conduct error analysis, and build tooling for reproducible experimentation and visualization.
  • Review the literature, formulate hypotheses, design experiments, and analyze results.
  • Write clean, reproducible code in modern deep learning frameworks (e.g., PyTorch).
  • Collaborate with the team, and contribute to research publications (e.g., CVPR, ICCV, ECCV, NeurIPS).   

 

Minimum Qualifications

 

  • Currently enrolled in an MS or PhD program in Computer Vision, Machine Learning, Robotics, Artificial Intelligence, or a closely related field.
  • Experience with multimodal LLMs, and their applications to video understanding tasks, including long-form video or temporal modeling, and familiarity with video encoders (e.g., VideoMAE, LaViLa, TimeSformer) and/or fine-tuning approaches.
  • Strong programming skills, with the ability to write clean, efficient, and reproducible code, and proficiency in deep learning frameworks (e.g., PyTorch).
  • Familiarity with model efficiency considerations (e.g., latency, memory, or throughput).
  • Solid problem-solving skills and the ability to independently drive projects from ideation to experimentation (and optionally publication).
  • Strong written and verbal communication skills.

 

Bonus Qualifications

  • Familiarity with procedural untrimmed video understanding and downstream tasks such as temporal action segmentation, detection or anticipation.
  • Familiarity with self-supervised or contrastive representation learning.
  • Experience working with procedural human action datasets (e.g., long-form, untrimmed videos) and associated annotations.
  • Experience in multiview and multimodal learning (e.g., ego, exo views, hand pose, body pose).
  • Background in learning video representations, modern video understanding methods, and efficient fine-tuning techniques.

 

Years of Work Experience Required   0
Desired Start Date  1/11/2027
Internship Duration  3 Months
Position Keywords

 Video Action Segmentation, Human Action Understanding 

Alternate Way to Apply

Send an e-mail to careers@honda-ri.com with the following:
- Subject line including the job number(s) you are applying for 
- Recent CV 
- A cover letter highlighting relevant background (Optional)

Please, do not contact our office to inquiry about your application status.