A multimodal framework of continuous music emotion recognition for adaptive cockpit lighting

Continuous music emotion recognition remains sensitive to dataset shift and imperfect lyric alignment. We propose Cockpit-EmoNet, which combines a frozen MERT acoustic encoder and a partially fine-tuned GTE text encoder through bidirectional cross-attention, feature-wise gating, and joint regression–alignment learning. On the pooled PMEmo–DEAM benchmark, Cockpit-EmoNet achieves valence PCC/CCC of 0.69/0.67 and arousal PCC/CCC of 0.81/0.79. Across five runs, it significantly reduces song-level errors relative to Cross-Attention Only ( \\(p_{\\textrm{adj}}<0.001\\) ) and improves all within- and cross-dataset evaluation settings. An offline cockpit-lighting analysis further shows that lower VA prediction errors are associated with smaller color and transition deviations. Controlled cockpit validation further provides initial physical and perceptual evidence for the downstream lighting application.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-08
DOI
https://doi.org/10.1038/s41598-026-70305-y
Primary Topic
Emotion and Mood Recognition
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A multimodal framework of continuous music emotion recognition for adaptive cockpit lighting

Xingang Mou, Lvlong Chen, Dongming Wang, Wei Shen
Scientific Reports
Emotion and Mood Recognition
article

A multimodal framework of continuous music emotion recognition for adaptive cockpit lighting

Xingang Mou, Lvlong Chen, Dongming Wang, Wei Shen
article en

Abstract

Continuous music emotion recognition remains sensitive to dataset shift and imperfect lyric alignment. We propose Cockpit-EmoNet, which combines a frozen MERT acoustic encoder and a partially fine-tuned GTE text encoder through bidirectional cross-attention, feature-wise gating, and joint regression–alignment learning. On the pooled PMEmo–DEAM benchmark, Cockpit-EmoNet achieves valence PCC/CCC of 0.69/0.67 and arousal PCC/CCC of 0.81/0.79. Across five runs, it significantly reduces song-level errors relative to Cross-Attention Only ( \(p_{\textrm{adj}}<0.001\) ) and improves all within- and cross-dataset evaluation settings. An offline cockpit-lighting analysis further shows that lower VA prediction errors are associated with smaller color and transition deviations. Controlled cockpit validation further provides initial physical and perceptual evidence for the downstream lighting application.

Scientific Reports
Wuhan University of Technology (CN)
Openalex Percentile: Top 7%
Emotion and Mood Recognition
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

A multimodal framework of continuous music emotion recognition for adaptive cockpit lighting — Xingang Mou, Lvlong Chen, et al. · Scientific Reports (2026) | TGRS Research Map | TGRS