SPEAKER ACCENT RECOGNITION USING MACHINE LEARNING

Speaker Accent Recognition using Machine Learning classifies speech into six accent groups: Spanish, French, German, Italian, British English, and American English. The executed notebook uses 329 records containing twelve Mel-frequency cepstral coefficient (MFCC) features and no missing values, together with ten WAV files for demonstration. A pipeline standardises the twelve features and trains a class-balanced Random Forest classifier with 500 trees on a stratified 80:20 split. The model is evaluated on 66 samples and the notebook also extracts mean MFCC vectors from uploaded WAV files for interactive prediction. The notebook reports 83.33% accuracy, 86.80% weighted precision, 83.33% weighted recall, and 81.66% weighted F1-score. American English dominates the dataset, and French, Italian, and British English have lower recall. The result is therefore presented as an academic prototype, not as evidence of universal accent identification. Keywords: speaker accent recognition, MFCC, audio classification, random forest, librosa, machine learning.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-15
DOI
https://doi.org/10.5281/zenodo.22767637
Primary Topic
Speech Recognition and Synthesis
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

SPEAKER ACCENT RECOGNITION USING MACHINE LEARNING

Gagan Rana, Kaushal Kumar, Md. Rehan ., Md. Arman Alam
Zenodo (CERN European Organization for Nuclear Research)
Speech Recognition and Synthesis
article

SPEAKER ACCENT RECOGNITION USING MACHINE LEARNING

Gagan Rana, Kaushal Kumar, Md. Rehan ., Md. Arman Alam
article en

Abstract

Speaker Accent Recognition using Machine Learning classifies speech into six accent groups: Spanish, French, German, Italian, British English, and American English. The executed notebook uses 329 records containing twelve Mel-frequency cepstral coefficient (MFCC) features and no missing values, together with ten WAV files for demonstration. A pipeline standardises the twelve features and trains a class-balanced Random Forest classifier with 500 trees on a stratified 80:20 split. The model is evaluated on 66 samples and the notebook also extracts mean MFCC vectors from uploaded WAV files for interactive prediction. The notebook reports 83.33% accuracy, 86.80% weighted precision, 83.33% weighted recall, and 81.66% weighted F1-score. American English dominates the dataset, and French, Italian, and British English have lower recall. The result is therefore presented as an academic prototype, not as evidence of universal accent identification. Keywords: speaker accent recognition, MFCC, audio classification, random forest, librosa, machine learning.

Zenodo (CERN European Organization for Nuclear Research)
Munger University (IN)
Quality Education
Openalex Percentile: Top 8%
Speech Recognition and Synthesis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

SPEAKER ACCENT RECOGNITION USING MACHINE LEARNING — Gagan Rana, Kaushal Kumar, et al. · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS