Emotion Classification in Music: Leveraging Machine Learning for Music Therapy and Emotional Response Analysis
Emotion classification in music is a central problem in affective computing with critical applications in music therapy, personalized media, and emotionally adaptive systems. This study proposes a novel multi-modal stacking ensemble framework for multi-label music emotion recognition using the Emotify dataset. The proposed architecture integrates acoustic features, listener metadata (age, gender, mood), and genre information to capture both perceptual and contextual determinants of emotional response. Six modeling phases were evaluated using Random Forest, Multi-Layer Perceptron (MLP), XGBoost, and ensemble learning strategies. The final Stacking Ensemble, combining Random Forest, MLP, and XGBoost through a logistic regression meta-classifier, achieved a subset accuracy of 0.41, Hamming loss of 0.25, and macro-averaged F1 score of 0.67, outperforming all individual models as well as recent state-of-the-art approaches evaluated on the same Emotify dataset. Compared with existing CNN-LSTM and LLM-based methods, which report macro-F1 values below 0.51 or rely on shortsegment evaluation, the proposed framework delivers substantially higher performance on full-length music clips with strict multi-label evaluation. The results demonstrate that stacked ensemble learning combined with multi-modal feature fusion provides a robust and generalizable solution for modeling complex emotional landscapes in music, enabling accurate emotionaware music therapy systems, adaptive media platforms, and personalized music recommendation engines.
Authors
- Le Yang
- Jing Wu (ORCID: https://orcid.org/0009-0008-6657-6893)
- Yonggang Pan
Publication Details
- Journal
- International Journal of Pattern Recognition and Artificial Intelligence
- Published
- 2026-09-30
- DOI
- https://doi.org/10.1142/s0218001426500497
- Primary Topic
- Emotion and Mood Recognition
- Type
- article
- Field-Weighted Citation Impact
- 0.00