Environmental Sound Classification for Audio-Based Action Recognition: A Comparative Study of Classical and Deep Models
BackgroundEnvironmental Sound Classification helps smart systems understand sounds happening in real-life situations.It is used in areas like watching over places, making cities smarter, keeping an eye on health, and better understanding media content. ObjectiveThis study compares three traditional machine learning classifiers -K-Nearest Neighbours (KNN), Support Vector Machine (SVM), and Random Forest (RF) -with a Convolutional Neural Network (CNN) for classifying environmental sounds using the UrbanSound8K benchmark. MethodsFor the classical models, four groups of handcrafted features -Mel-Frequency Cepstral Coefficients, Chroma, spectral features, and Tonnetz -were extracted from the audio recordings.The CNN model used log-Mel spectrograms as its input.All models were tested using the 10-fold cross-validation method provided by UrbanSound8K.In each round of testing, preprocessing and scaling were done only on the training data and then applied to the test part.The results were measured using accuracy, precision, recall, macro F1-score, confidence intervals, and paired statistical tests. ResultsAmong the classical approaches, RF achieved the highest macro F1-score (65.72%), followed by linear SVM (62.04%) and KNN (56.26%).The proposed CNN obtained a macro F1-score of approximately 70.42% with an accuracy of approximately 69.35%, outperforming all classical baselines.Both the paired Student's t-test and Wilcoxon signed-rank test indicated statistically significant differences between the CNN and each classical baseline (p < 0.05). ConclusionThe results demonstrate that CNNs provide superior environmental sound classification performance by learning discriminative spectral-temporal representations directly from log-Mel spectrograms, whereas RF remains a computationally efficient alternative for resource-constrained deployments.A graphical user interface-based prototype is presented as a proof-of-concept deployment framework illustrating practical inference rather than as a separately evaluated scientific contribution.
Authors
- Shreya Jare
- Rushali Rajaram Katkar
- Sushant Sharma
- Sushama A Shirke
Institutions
- United States Department of the Army (US)
Publication Details
- Journal
- Cureus Journal of Computer Science.
- Published
- 2026-09-16
- DOI
- https://doi.org/10.7759/s44389-026-00281-x
- Primary Topic
- Music and Audio Processing
- Type
- article
- Field-Weighted Citation Impact
- 0.00