Environmental Sound Classification for Audio-Based Action Recognition: A Comparative Study of Classical and Deep Models

BackgroundEnvironmental Sound Classification helps smart systems understand sounds happening in real-life situations.It is used in areas like watching over places, making cities smarter, keeping an eye on health, and better understanding media content. ObjectiveThis study compares three traditional machine learning classifiers -K-Nearest Neighbours (KNN), Support Vector Machine (SVM), and Random Forest (RF) -with a Convolutional Neural Network (CNN) for classifying environmental sounds using the UrbanSound8K benchmark. MethodsFor the classical models, four groups of handcrafted features -Mel-Frequency Cepstral Coefficients, Chroma, spectral features, and Tonnetz -were extracted from the audio recordings.The CNN model used log-Mel spectrograms as its input.All models were tested using the 10-fold cross-validation method provided by UrbanSound8K.In each round of testing, preprocessing and scaling were done only on the training data and then applied to the test part.The results were measured using accuracy, precision, recall, macro F1-score, confidence intervals, and paired statistical tests. ResultsAmong the classical approaches, RF achieved the highest macro F1-score (65.72%), followed by linear SVM (62.04%) and KNN (56.26%).The proposed CNN obtained a macro F1-score of approximately 70.42% with an accuracy of approximately 69.35%, outperforming all classical baselines.Both the paired Student's t-test and Wilcoxon signed-rank test indicated statistically significant differences between the CNN and each classical baseline (p < 0.05). ConclusionThe results demonstrate that CNNs provide superior environmental sound classification performance by learning discriminative spectral-temporal representations directly from log-Mel spectrograms, whereas RF remains a computationally efficient alternative for resource-constrained deployments.A graphical user interface-based prototype is presented as a proof-of-concept deployment framework illustrating practical inference rather than as a separately evaluated scientific contribution.

Authors

Institutions

Publication Details

Journal
Cureus Journal of Computer Science.
Published
2026-09-16
DOI
https://doi.org/10.7759/s44389-026-00281-x
Primary Topic
Music and Audio Processing
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Environmental Sound Classification for Audio-Based Action Recognition: A Comparative Study of Classical and Deep Models

Shreya Jare, Rushali Rajaram Katkar, Sushant Sharma, Sushama A Shirke
Cureus Journal of Computer Science.
Music and Audio Processing
article

Environmental Sound Classification for Audio-Based Action Recognition: A Comparative Study of Classical and Deep Models

Shreya Jare, Rushali Rajaram Katkar, Sushant Sharma, Sushama A Shirke
article en

Abstract

BackgroundEnvironmental Sound Classification helps smart systems understand sounds happening in real-life situations.It is used in areas like watching over places, making cities smarter, keeping an eye on health, and better understanding media content. ObjectiveThis study compares three traditional machine learning classifiers -K-Nearest Neighbours (KNN), Support Vector Machine (SVM), and Random Forest (RF) -with a Convolutional Neural Network (CNN) for classifying environmental sounds using the UrbanSound8K benchmark. MethodsFor the classical models, four groups of handcrafted features -Mel-Frequency Cepstral Coefficients, Chroma, spectral features, and Tonnetz -were extracted from the audio recordings.The CNN model used log-Mel spectrograms as its input.All models were tested using the 10-fold cross-validation method provided by UrbanSound8K.In each round of testing, preprocessing and scaling were done only on the training data and then applied to the test part.The results were measured using accuracy, precision, recall, macro F1-score, confidence intervals, and paired statistical tests. ResultsAmong the classical approaches, RF achieved the highest macro F1-score (65.72%), followed by linear SVM (62.04%) and KNN (56.26%).The proposed CNN obtained a macro F1-score of approximately 70.42% with an accuracy of approximately 69.35%, outperforming all classical baselines.Both the paired Student's t-test and Wilcoxon signed-rank test indicated statistically significant differences between the CNN and each classical baseline (p < 0.05). ConclusionThe results demonstrate that CNNs provide superior environmental sound classification performance by learning discriminative spectral-temporal representations directly from log-Mel spectrograms, whereas RF remains a computationally efficient alternative for resource-constrained deployments.A graphical user interface-based prototype is presented as a proof-of-concept deployment framework illustrating practical inference rather than as a separately evaluated scientific contribution.

Cureus Journal of Computer Science.
United States Department of the Army (US)
Climate action
Openalex Percentile: Top 10%
Music and Audio Processing
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Environmental Sound Classification for Audio-Based Action Recognition: A Comparative Study of Classical and Deep Models — Shreya Jare, Rushali Rajaram Katkar, et al. · Cureus Journal of Computer Science. (2026) | TGRS Research Map | TGRS