Music Genre Classification Algorithm Based on Multi-Scale Fusion
In the era of music streaming, the vast volume of digital music resources renders manual classification impractical. Although deep learning has achieved significant advances in music genre classification tasks, existing methods still face challenges such as insufficient feature extraction, low classification accuracy, loss of temporal sequence information, and limited dataset sizes. This study addresses these challenges in several ways: first, to tackle the problem of varying music durations affecting classification accuracy, log-Mel spectrograms are generated from segmented audio data; second, to improve classification accuracy on small datasets, we propose HybridNet, a deep learning model that employs a deep multi-scale fusion module to extract richer feature information. Subsequently, to further mitigate the loss of temporal information, we introduce a bidirectional long short-term memory (BiLSTM) layer resulting in the Bi-HybridNet model, which captures key sequential information in both forward and backward directions. Experimental results demonstrate that the proposed method achieves 86.92% accuracy on the FMA-small dataset, indicating improved classification performance over classical models. This study provides a practical technical solution for music genre classification, offering certain reference values in feature extraction and temporal modeling.
Authors
- Tingting CHEN (ORCID: https://orcid.org/0000-0002-1869-1909)
- Yonggang Tang
Institutions
- Suzhou University of Science and Technology (CN)
- Soochow University (CN)
Publication Details
- Journal
- Journal of Advanced Computational Intelligence and Intelligent Informatics
- Published
- 2026-09-19
- DOI
- https://doi.org/10.20965/jaciii.2026.p1640
- Primary Topic
- Music and Audio Processing
- Type
- article
- Field-Weighted Citation Impact
- 0.00