Adaptive Learning-based Semantic Enhancement for Multimodal Dialogue Discourse Parsing
Dialogue Discourse Parsing (DDP) aims to identify dependency structures and relation types among utterances in conversations. Existing multimodal DDP methods often directly fuse low-level audio-visual features, making it difficult to establish explicit semantic links between nonverbal behaviors and discourse relations and potentially introducing irrelevant modality noise. To address these limitations, we propose ALSE-MDDP , an A daptive L earning-based S emantic E nhancement framework for M ultimodal D ialogue D iscourse P arsing. ALSE-MDDP includes three modules: T ext P re- P arsing (TPP), which performs text-only parsing and identifies low-confidence predictions; E xplicit S emantic A ugmentation (ESA), which converts discourse-relevant audio-visual behaviors into textual cues; and A daptive F eedback P reference O ptimization (AFPO), which uses parser feedback to optimize cue generation. Experiments on DraDDP and MODDP show that ALSE-MDDP achieves 56.79 and 57.50 Link&Rel F1, outperforming the strongest multimodal baseline, LLaMIPa* (Qwen2.5-Omni-7B), by 3.45 and 5.23 points, respectively.
Authors
- Peifeng Li (ORCID: https://orcid.org/0000-0003-4850-3128)
- Qiaoming Zhu (ORCID: https://orcid.org/0000-0002-2708-8976)
- Shannan Liu (ORCID: https://orcid.org/0009-0001-1098-7406)
Institutions
- Soochow University (CN)
Publication Details
- Journal
- Information Processing & Management
- Published
- 2026-10-04
- DOI
- https://doi.org/10.1016/j.ipm.2026.105209
- Primary Topic
- Speech and dialogue systems
- Type
- article
- Field-Weighted Citation Impact
- 0.00
Funders
- National Natural Science Foundation of China