Vauxguard: Zero-Shot LLM-Mediated Interpretable Acoustic Forensics for Synthetic Speech Detection
Vauxguard is a web-based synthetic speech detection system that combines deterministic acoustic feature extraction with large-language-model-mediated classification. The system extracts seven interpretable signal-processing features from 16 kHz mono audio and uses an LLM to classify recordings as genuine or synthetic. A pilot evaluation on 50 balanced samples, including 25 human recordings and 25 synthetic clips generated using ElevenLabs and OpenAI TTS, achieved 86.0% accuracy, 82.1% precision, 92.0% recall, and an F1 score of 86.8%. This preprint describes the system architecture, evaluation methodology, limitations, and reproducibility materials.
Authors
- Eshaan Revankar
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-19
- DOI
- https://doi.org/10.5281/zenodo.22848991
- Primary Topic
- Speech Recognition and Synthesis
- Type
- preprint