An intelligent vocal singing quality assessment model based on adaptive convolution and temporal graph modeling

Automatic assessment of vocal singing quality is an important research topic in intelligent music education. The core challenges lie in extracting multi-scale features from complex acoustic signals, modeling hierarchical temporal dependencies during the singing process, and jointly evaluating multi-dimensional quality attributes. Existing methods remain insufficient in terms of feature extraction adaptability, cross-scale structural modeling capability, and interpretability of assessment results. To this end, an intelligent vocal singing quality assessment model was presented based on adaptive convolution and temporal graph modeling. The model comprises three core components: a feature extraction mechanism that uses fixed multi-scale convolutional branches and learns adaptive fusion weights through a channel attention mechanism to capture acoustic features at different time–frequency scales; a cross-granularity graph attention mechanism that constructs frame-level and segment-level dual-layer graph structures and models hierarchical temporal dependencies of singing signals through cross-layer attention propagation; and a multi-task vocal quality assessment model that achieves joint assessment of sub-dimensions including pitch accuracy, rhythm, timbre, and expressiveness along with overall quality through dimension-decoupled scoring branches and hierarchical consistency constraints. Experimental results on three public datasets demonstrate that our method outperforms existing comparison methods across all four metrics of R 2 , Mean Absolute Error (MAE), Root Mean Square Error (RMSE), and Pearson Correlation Coefficient (PCC), with the overall score R 2 reaching 0.921, representing a 4.3% improvement over the best comparison method.

Authors

Institutions

Publication Details

Journal
PeerJ Computer Science
Published
2026-09-24
DOI
https://doi.org/10.7717/peerj-cs.4090
Primary Topic
Music and Audio Processing
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

An intelligent vocal singing quality assessment model based on adaptive convolution and temporal graph modeling

Adeel Ashraf Cheema, Li Sun
PeerJ Computer Science
Music and Audio Processing
article

An intelligent vocal singing quality assessment model based on adaptive convolution and temporal graph modeling

Adeel Ashraf Cheema, Li Sun
article en

Abstract

Automatic assessment of vocal singing quality is an important research topic in intelligent music education. The core challenges lie in extracting multi-scale features from complex acoustic signals, modeling hierarchical temporal dependencies during the singing process, and jointly evaluating multi-dimensional quality attributes. Existing methods remain insufficient in terms of feature extraction adaptability, cross-scale structural modeling capability, and interpretability of assessment results. To this end, an intelligent vocal singing quality assessment model was presented based on adaptive convolution and temporal graph modeling. The model comprises three core components: a feature extraction mechanism that uses fixed multi-scale convolutional branches and learns adaptive fusion weights through a channel attention mechanism to capture acoustic features at different time–frequency scales; a cross-granularity graph attention mechanism that constructs frame-level and segment-level dual-layer graph structures and models hierarchical temporal dependencies of singing signals through cross-layer attention propagation; and a multi-task vocal quality assessment model that achieves joint assessment of sub-dimensions including pitch accuracy, rhythm, timbre, and expressiveness along with overall quality through dimension-decoupled scoring branches and hierarchical consistency constraints. Experimental results on three public datasets demonstrate that our method outperforms existing comparison methods across all four metrics of R 2 , Mean Absolute Error (MAE), Root Mean Square Error (RMSE), and Pearson Correlation Coefficient (PCC), with the overall score R 2 reaching 0.921, representing a 4.3% improvement over the best comparison method.

PeerJ Computer ScienceVol. 12
National University of Computer and Emerging Sciences (PK), Guangxi Normal University (CN)
Quality Education
Openalex Percentile: Top 10%
Music and Audio Processing
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

An intelligent vocal singing quality assessment model based on adaptive convolution and temporal graph modeling — Adeel Ashraf Cheema, Li Sun · PeerJ Computer Science (2026) | TGRS Research Map | TGRS