Adaptive learning behavior recognition using multimodal transformer based dynamic modality regulation
Abstract With multimodal interaction becoming increasingly common in intelligent learning systems, behavior streams from clicks, speech, facial expressions, and gestures often show time-varying reliability, short high-saliency events, and correlated or nested labels. These characteristics make static fusion and label-agnostic modeling unstable under imbalance and distribution shift. This paper proposes MFT-Net for an 18-label learning behavior recognition task in which multimodal sequences are mapped to behavior categories and optional nested sub-labels. MFT-Net integrates modal response filtering, dynamic modality weight regulation, and structured label embedding into a Transformer-based sequence model. The response filtering module preserves salient temporal segments while suppressing low-activity noise; the dynamic regulation module adjusts modality contributions according to cross-modal consistency; and the structured label embedding module introduces label semantics into attention-based sequence matching. Experiments on 5240 multimodal learning-behavior sequences show that MFT-Net achieves 96.2% accuracy and 95.5% F1 score, converges in 8 epochs, and obtains 34 ms average inference latency with an 18.6 MB model footprint under the workstation GPU setting. Robustness tests under skewed sample distributions and frame-length perturbations further indicate that the proposed model improves recognition stability while keeping computational cost within a deployable range.
Authors
- Lei Zujun
- Faze Liang
- Xueyao Du (ORCID: https://orcid.org/0000-0002-5852-5825)
Institutions
- Chongqing Academy of Environmental Science (CN)
- Yango University (CN)
Publication Details
- Journal
- Discover Artificial Intelligence
- Published
- 2026-09-18
- DOI
- https://doi.org/10.1007/s44163-026-02112-3
- Primary Topic
- Emotion and Mood Recognition
- Type
- article
- Field-Weighted Citation Impact
- 0.00