Emergency Call Verification Using Cepstral Features With Neural Networks and Federated Learning
ABSTRACT The reliance on vocal communication services, particularly in environments where the security aspects are highly valued such as emergencies, call for reliable methods for that detect the manipulation of environmental sound. One of the current challenges is the detection of SceneFake audio where only the environment sound is altered but not the speech content. This paper presents a privacy‐preserving scene‐manipulated audio detection framework that combines Linear Frequency Cepstral Coefficients (LFCC) oriented cepstral feature representation with Convolutional Neural Network (CNN) and Dense Neural Network (DNN) classifiers under both centralized and Federated Learning (FL) settings. The LFCC‐oriented cepstral representation is employed as the front‐end feature extraction technique to capture discriminative characteristics of environmental audio, while CNN and DNN models perform the classification of authentic and manipulated audio scenes. The federated learning technique helps in the training of the model in a distributed fashion by various clients without sharing the raw audio data in order to maintain privacy. SMOTE is used in order to handle the problem of class imbalance during the preprocessing stage for the training set. The performance of the suggested models is validated using the SceneFake dataset. Among the evaluated models, the LFCC‐CNN with Federated Learning achieves the lowest EER, outperforming the corresponding centralized CNN and DNN implementations. This work shows the efficiency of the suggested framework for privacy‐preserving detection of altered audio from the environment and reveals its possible uses in the field of emergency communication and multimedia forensics.
Authors
- Mohit Dua (ORCID: https://orcid.org/0000-0001-7071-8323)
- Sakshi
Institutions
- National Institute of Technology Kurukshetra (IN)
Publication Details
- Journal
- Fire and Materials
- Published
- 2026-10-06
- DOI
- https://doi.org/10.1002/fam.70106
- Primary Topic
- Speech Recognition and Synthesis
- Type
- article
- Field-Weighted Citation Impact
- 0.00