Multimodal transformer and directed relational graph learning for emotion recognition in conversations with potential applications in college student mental health research

Multimodal emotion recognition in conversations aims to identify the emotion expressed in each utterance by jointly analyzing textual, acoustic, visual, and contextual information. However, existing methods still face challenges in coordinating intra-modal dependency modeling, cross-modal interaction, and graph-based conversational reasoning. To address these challenges, this paper proposes MTG-ERC, a multimodal conversational emotion recognition framework that integrates transformer-based attention mechanisms with directed relational graph learning. First, textual, acoustic, and visual features are projected into a unified representation space. Intra-modal self-attention and cross-modal attention are then employed to capture modality-specific contextual dependencies and complementary interactions among modalities. Subsequently, a directed multi-relational graph module combines relational graph convolution and Graph Transformer operations to model temporal context, speaker-related interactions, and multimodal dependencies among utterances. Finally, global attention-based representations and local graph-based contextual representations are integrated for utterance-level emotion classification. Experiments on the IEMOCAP and MELD datasets show that MTG-ERC achieves competitive performance compared with the evaluated conversational emotion recognition baselines. Ablation results further indicate that intra-modal attention, cross-modal interaction, and directed relational modeling make complementary contributions to the final performance. The proposed framework also has potential applications in college student mental health research, where its multimodal affective modeling capability could support the analysis of emotional patterns in conversational data.

Authors

Institutions

Publication Details

Journal
Frontiers in Psychology
Published
2026-09-14
DOI
https://doi.org/10.3389/fpsyg.2026.1867034
Primary Topic
Emotion and Mood Recognition
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Multimodal transformer and directed relational graph learning for emotion recognition in conversations with potential applications in college student mental health research

Jing Zhao, Feng Liu
Frontiers in Psychology
Emotion and Mood Recognition
article

Multimodal transformer and directed relational graph learning for emotion recognition in conversations with potential applications in college student mental health research

Jing Zhao, Feng Liu
article en

Abstract

Multimodal emotion recognition in conversations aims to identify the emotion expressed in each utterance by jointly analyzing textual, acoustic, visual, and contextual information. However, existing methods still face challenges in coordinating intra-modal dependency modeling, cross-modal interaction, and graph-based conversational reasoning. To address these challenges, this paper proposes MTG-ERC, a multimodal conversational emotion recognition framework that integrates transformer-based attention mechanisms with directed relational graph learning. First, textual, acoustic, and visual features are projected into a unified representation space. Intra-modal self-attention and cross-modal attention are then employed to capture modality-specific contextual dependencies and complementary interactions among modalities. Subsequently, a directed multi-relational graph module combines relational graph convolution and Graph Transformer operations to model temporal context, speaker-related interactions, and multimodal dependencies among utterances. Finally, global attention-based representations and local graph-based contextual representations are integrated for utterance-level emotion classification. Experiments on the IEMOCAP and MELD datasets show that MTG-ERC achieves competitive performance compared with the evaluated conversational emotion recognition baselines. Ablation results further indicate that intra-modal attention, cross-modal interaction, and directed relational modeling make complementary contributions to the final performance. The proposed framework also has potential applications in college student mental health research, where its multimodal affective modeling capability could support the analysis of emotional patterns in conversational data.

Frontiers in PsychologyVol. 17
Anhui Business College (CN), QuantumCTek (China) (CN)
Quality Education
Openalex Percentile: Top 7%
Emotion and Mood Recognition
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Multimodal transformer and directed relational graph learning for emotion recognition in conversations with potential applications in college student mental health research — Jing Zhao, Feng Liu · Frontiers in Psychology (2026) | TGRS Research Map | TGRS