An intelligent vocal singing quality assessment model based on adaptive convolution and temporal graph modeling
Automatic assessment of vocal singing quality is an important research topic in intelligent music education. The core challenges lie in extracting multi-scale features from complex acoustic signals, modeling hierarchical temporal dependencies during the singing process, and jointly evaluating multi-dimensional quality attributes. Existing methods remain insufficient in terms of feature extraction adaptability, cross-scale structural modeling capability, and interpretability of assessment results. To this end, an intelligent vocal singing quality assessment model was presented based on adaptive convolution and temporal graph modeling. The model comprises three core components: a feature extraction mechanism that uses fixed multi-scale convolutional branches and learns adaptive fusion weights through a channel attention mechanism to capture acoustic features at different time–frequency scales; a cross-granularity graph attention mechanism that constructs frame-level and segment-level dual-layer graph structures and models hierarchical temporal dependencies of singing signals through cross-layer attention propagation; and a multi-task vocal quality assessment model that achieves joint assessment of sub-dimensions including pitch accuracy, rhythm, timbre, and expressiveness along with overall quality through dimension-decoupled scoring branches and hierarchical consistency constraints. Experimental results on three public datasets demonstrate that our method outperforms existing comparison methods across all four metrics of R 2 , Mean Absolute Error (MAE), Root Mean Square Error (RMSE), and Pearson Correlation Coefficient (PCC), with the overall score R 2 reaching 0.921, representing a 4.3% improvement over the best comparison method.
Authors
- Adeel Ashraf Cheema (ORCID: https://orcid.org/0000-0002-8310-6251)
- Li Sun (ORCID: https://orcid.org/0000-0001-7950-0953)
Institutions
- National University of Computer and Emerging Sciences (PK)
- Guangxi Normal University (CN)
Publication Details
- Journal
- PeerJ Computer Science
- Published
- 2026-09-24
- DOI
- https://doi.org/10.7717/peerj-cs.4090
- Primary Topic
- Music and Audio Processing
- Type
- article
- Field-Weighted Citation Impact
- 0.00