A multimodal framework of continuous music emotion recognition for adaptive cockpit lighting
Continuous music emotion recognition remains sensitive to dataset shift and imperfect lyric alignment. We propose Cockpit-EmoNet, which combines a frozen MERT acoustic encoder and a partially fine-tuned GTE text encoder through bidirectional cross-attention, feature-wise gating, and joint regression–alignment learning. On the pooled PMEmo–DEAM benchmark, Cockpit-EmoNet achieves valence PCC/CCC of 0.69/0.67 and arousal PCC/CCC of 0.81/0.79. Across five runs, it significantly reduces song-level errors relative to Cross-Attention Only ( \\(p_{\\textrm{adj}}<0.001\\) ) and improves all within- and cross-dataset evaluation settings. An offline cockpit-lighting analysis further shows that lower VA prediction errors are associated with smaller color and transition deviations. Controlled cockpit validation further provides initial physical and perceptual evidence for the downstream lighting application.
Authors
- Xingang Mou
- Lvlong Chen
- Dongming Wang
- Wei Shen
Institutions
- Wuhan University of Technology (CN)
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-09-08
- DOI
- https://doi.org/10.1038/s41598-026-70305-y
- Primary Topic
- Emotion and Mood Recognition
- Type
- article
- Field-Weighted Citation Impact
- 0.00