Adaptive learning behavior recognition using multimodal transformer based dynamic modality regulation

Abstract With multimodal interaction becoming increasingly common in intelligent learning systems, behavior streams from clicks, speech, facial expressions, and gestures often show time-varying reliability, short high-saliency events, and correlated or nested labels. These characteristics make static fusion and label-agnostic modeling unstable under imbalance and distribution shift. This paper proposes MFT-Net for an 18-label learning behavior recognition task in which multimodal sequences are mapped to behavior categories and optional nested sub-labels. MFT-Net integrates modal response filtering, dynamic modality weight regulation, and structured label embedding into a Transformer-based sequence model. The response filtering module preserves salient temporal segments while suppressing low-activity noise; the dynamic regulation module adjusts modality contributions according to cross-modal consistency; and the structured label embedding module introduces label semantics into attention-based sequence matching. Experiments on 5240 multimodal learning-behavior sequences show that MFT-Net achieves 96.2% accuracy and 95.5% F1 score, converges in 8 epochs, and obtains 34 ms average inference latency with an 18.6 MB model footprint under the workstation GPU setting. Robustness tests under skewed sample distributions and frame-length perturbations further indicate that the proposed model improves recognition stability while keeping computational cost within a deployable range.

Authors

Institutions

Publication Details

Journal
Discover Artificial Intelligence
Published
2026-09-18
DOI
https://doi.org/10.1007/s44163-026-02112-3
Primary Topic
Emotion and Mood Recognition
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Adaptive learning behavior recognition using multimodal transformer based dynamic modality regulation

Lei Zujun, Faze Liang, Xueyao Du
Discover Artificial Intelligence
Emotion and Mood Recognition
article

Adaptive learning behavior recognition using multimodal transformer based dynamic modality regulation

Lei Zujun, Faze Liang, Xueyao Du
article en

Abstract

Abstract With multimodal interaction becoming increasingly common in intelligent learning systems, behavior streams from clicks, speech, facial expressions, and gestures often show time-varying reliability, short high-saliency events, and correlated or nested labels. These characteristics make static fusion and label-agnostic modeling unstable under imbalance and distribution shift. This paper proposes MFT-Net for an 18-label learning behavior recognition task in which multimodal sequences are mapped to behavior categories and optional nested sub-labels. MFT-Net integrates modal response filtering, dynamic modality weight regulation, and structured label embedding into a Transformer-based sequence model. The response filtering module preserves salient temporal segments while suppressing low-activity noise; the dynamic regulation module adjusts modality contributions according to cross-modal consistency; and the structured label embedding module introduces label semantics into attention-based sequence matching. Experiments on 5240 multimodal learning-behavior sequences show that MFT-Net achieves 96.2% accuracy and 95.5% F1 score, converges in 8 epochs, and obtains 34 ms average inference latency with an 18.6 MB model footprint under the workstation GPU setting. Robustness tests under skewed sample distributions and frame-length perturbations further indicate that the proposed model improves recognition stability while keeping computational cost within a deployable range.

Discover Artificial IntelligenceVol. 6(1)
Chongqing Academy of Environmental Science (CN), Yango University (CN)
Openalex Percentile: Top 7%
Emotion and Mood Recognition
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Adaptive learning behavior recognition using multimodal transformer based dynamic modality regulation — Lei Zujun, Faze Liang, et al. · Discover Artificial Intelligence (2026) | TGRS Research Map | TGRS