Adaptive Learning-based Semantic Enhancement for Multimodal Dialogue Discourse Parsing

Dialogue Discourse Parsing (DDP) aims to identify dependency structures and relation types among utterances in conversations. Existing multimodal DDP methods often directly fuse low-level audio-visual features, making it difficult to establish explicit semantic links between nonverbal behaviors and discourse relations and potentially introducing irrelevant modality noise. To address these limitations, we propose ALSE-MDDP , an A daptive L earning-based S emantic E nhancement framework for M ultimodal D ialogue D iscourse P arsing. ALSE-MDDP includes three modules: T ext P re- P arsing (TPP), which performs text-only parsing and identifies low-confidence predictions; E xplicit S emantic A ugmentation (ESA), which converts discourse-relevant audio-visual behaviors into textual cues; and A daptive F eedback P reference O ptimization (AFPO), which uses parser feedback to optimize cue generation. Experiments on DraDDP and MODDP show that ALSE-MDDP achieves 56.79 and 57.50 Link&Rel F1, outperforming the strongest multimodal baseline, LLaMIPa* (Qwen2.5-Omni-7B), by 3.45 and 5.23 points, respectively.

Authors

Institutions

Publication Details

Journal
Information Processing & Management
Published
2026-10-04
DOI
https://doi.org/10.1016/j.ipm.2026.105209
Primary Topic
Speech and dialogue systems
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Adaptive Learning-based Semantic Enhancement for Multimodal Dialogue Discourse Parsing

Peifeng Li, Qiaoming Zhu, Shannan Liu
Information Processing & Management
Speech and dialogue systems
article

Adaptive Learning-based Semantic Enhancement for Multimodal Dialogue Discourse Parsing

Peifeng Li, Qiaoming Zhu, Shannan Liu
article en

Abstract

Dialogue Discourse Parsing (DDP) aims to identify dependency structures and relation types among utterances in conversations. Existing multimodal DDP methods often directly fuse low-level audio-visual features, making it difficult to establish explicit semantic links between nonverbal behaviors and discourse relations and potentially introducing irrelevant modality noise. To address these limitations, we propose ALSE-MDDP , an A daptive L earning-based S emantic E nhancement framework for M ultimodal D ialogue D iscourse P arsing. ALSE-MDDP includes three modules: T ext P re- P arsing (TPP), which performs text-only parsing and identifies low-confidence predictions; E xplicit S emantic A ugmentation (ESA), which converts discourse-relevant audio-visual behaviors into textual cues; and A daptive F eedback P reference O ptimization (AFPO), which uses parser feedback to optimize cue generation. Experiments on DraDDP and MODDP show that ALSE-MDDP achieves 56.79 and 57.50 Link&Rel F1, outperforming the strongest multimodal baseline, LLaMIPa* (Qwen2.5-Omni-7B), by 3.45 and 5.23 points, respectively.

Information Processing & ManagementVol. 64(2)
Soochow University (CN)
National Natural Science Foundation of China
Openalex Percentile: Top 11%
Speech and dialogue systems
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.