Research Intern: Multimodal AI for Collective State Understanding via Audio-Visual Fusion - Honda Research Institute USA

Research Intern: Multimodal AI for Collective State Understanding via Audio-Visual Fusion

Your application is being processed

Research Intern: Multimodal AI for Collective State Understanding via Audio-Visual Fusion

Job Number: P25INT-63
​Honda Research Institute USA (HRI-US) is seeking a research intern to develop theory-informed multimodal learning methods for understanding collective states and emerging norms in human and human-AI teams. The intern will connect multimodal behavioral signals, including verbal, nonverbal, and interaction dynamics, to constructs such as team cohesion, trust, and shared mental models. Building on these insights, they will develop data-driven machine learning methods that incorporate established behavioral and organizational theory, with an emphasis on causal and counterfactual modeling. The goal is to produce rigorous, interpretable research suitable for publication at a leading multimodal interaction venue.
San Jose, CA

 

Key Responsibilities

 

  • Translate sociology and psychology constructs into measurable group states, or loss functions in models.
  • Build multimodal baseline and theory-fused group representations.
  • Design actor, modality, and behavior counterfactuals in the model.
  • Conduct machine learning experiments and validate the model performance across datasets.

 

Minimum Qualifications

 

  • ​Current PhD student in Computer Science, Robotics, Cognitive Science or related fields.
  • Strong machine learning and large language model tuning skills.
  • Experience with temporal or multimodal human data.
  • Proven publication records in machine learning, natural language processing, computer vision, or human-computer interaction venues

 

Bonus Qualifications

  • ​Experience in social signal processing and affective computing with publications in top conferences.
  • Research experiences with group interaction or human-AI teaming.
  • Familiarity with causal inference or counterfactual modeling such as graph or sequence models.
  • Experience with multimodal foundation models.

 

Years of Work Experience Required   0
Desired Start Date  1/11/2027
Internship Duration  3 Months
Position Keywords  Multimodal machine learning, affective computing, human-AI teaming, causal ML, collective intelligence 

Alternate Way to Apply

Send an e-mail to careers@honda-ri.com with the following:
- Subject line including the job number(s) you are applying for 
- Recent CV 
- A cover letter highlighting relevant background (Optional)

Please, do not contact our office to inquiry about your application status.