A geometry-aware deep learning framework for detecting diffusion-era synthetic speech in digital security applications
Synthetic speech detection is an important AI security task for mitigating audio misinformation, voice impersonation, and trust violations in digital systems. Recent diffusion- and flow-matching-based generators produce highly realistic speech, making conventional detectors less reliable under unseen-generator and cross-paradigm conditions. This study investigates diffusion-era synthetic speech detection using pretrained audio representations and proposes GeoOT-Former, a geometry-aware optimal transport framework. The model constructs complementary hyperbolic and spherical token views from frozen pretrained embeddings and aligns them through entropic optimal transport before gated fusion and attention-based aggregation. This design is intended to improve cross-view correspondence and enhance detection robustness. Comprehensive experiments on DiffSSD show that GeoOT-Former consistently improves performance over single-encoder baselines, while evaluation on DFADD demonstrates cross-paradigm transfer capability. Among the evaluated encoders, Audio-MAMBA achieves the strongest performance, attaining 0.13% EER on DiffSSD-Dtest and 0.03% EER on DFADD. These results suggest that geometry-aware alignment is effective for detecting diffusion-era synthetic speech in digital security applications.
Authors
- Vikrant Bhateja (ORCID: https://orcid.org/0000-0002-3259-8874)
- Mohd Mujtaba Akhtar (ORCID: https://orcid.org/0009-0000-1982-7110)
Institutions
- Veer Bahadur Singh Purvanchal University (IN)
Publication Details
- Journal
- International Journal of Computers and Applications
- Published
- 2026-09-17
- DOI
- https://doi.org/10.1080/1206212x.2026.2732228
- Primary Topic
- Speech and Audio Processing
- Type
- article
- Field-Weighted Citation Impact
- 0.00