A geometry-aware deep learning framework for detecting diffusion-era synthetic speech in digital security applications

Synthetic speech detection is an important AI security task for mitigating audio misinformation, voice impersonation, and trust violations in digital systems. Recent diffusion- and flow-matching-based generators produce highly realistic speech, making conventional detectors less reliable under unseen-generator and cross-paradigm conditions. This study investigates diffusion-era synthetic speech detection using pretrained audio representations and proposes GeoOT-Former, a geometry-aware optimal transport framework. The model constructs complementary hyperbolic and spherical token views from frozen pretrained embeddings and aligns them through entropic optimal transport before gated fusion and attention-based aggregation. This design is intended to improve cross-view correspondence and enhance detection robustness. Comprehensive experiments on DiffSSD show that GeoOT-Former consistently improves performance over single-encoder baselines, while evaluation on DFADD demonstrates cross-paradigm transfer capability. Among the evaluated encoders, Audio-MAMBA achieves the strongest performance, attaining 0.13% EER on DiffSSD-Dtest and 0.03% EER on DFADD. These results suggest that geometry-aware alignment is effective for detecting diffusion-era synthetic speech in digital security applications.

Authors

Institutions

Publication Details

Journal
International Journal of Computers and Applications
Published
2026-09-17
DOI
https://doi.org/10.1080/1206212x.2026.2732228
Primary Topic
Speech and Audio Processing
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A geometry-aware deep learning framework for detecting diffusion-era synthetic speech in digital security applications

Vikrant Bhateja, Mohd Mujtaba Akhtar
International Journal of Computers and Applications
Speech and Audio Processing
article

A geometry-aware deep learning framework for detecting diffusion-era synthetic speech in digital security applications

Vikrant Bhateja, Mohd Mujtaba Akhtar
article en

Abstract

Synthetic speech detection is an important AI security task for mitigating audio misinformation, voice impersonation, and trust violations in digital systems. Recent diffusion- and flow-matching-based generators produce highly realistic speech, making conventional detectors less reliable under unseen-generator and cross-paradigm conditions. This study investigates diffusion-era synthetic speech detection using pretrained audio representations and proposes GeoOT-Former, a geometry-aware optimal transport framework. The model constructs complementary hyperbolic and spherical token views from frozen pretrained embeddings and aligns them through entropic optimal transport before gated fusion and attention-based aggregation. This design is intended to improve cross-view correspondence and enhance detection robustness. Comprehensive experiments on DiffSSD show that GeoOT-Former consistently improves performance over single-encoder baselines, while evaluation on DFADD demonstrates cross-paradigm transfer capability. Among the evaluated encoders, Audio-MAMBA achieves the strongest performance, attaining 0.13% EER on DiffSSD-Dtest and 0.03% EER on DFADD. These results suggest that geometry-aware alignment is effective for detecting diffusion-era synthetic speech in digital security applications.

International Journal of Computers and Applications
Veer Bahadur Singh Purvanchal University (IN)
Openalex Percentile: Top 10%
Speech and Audio Processing
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

A geometry-aware deep learning framework for detecting diffusion-era synthetic speech in digital security applications — Vikrant Bhateja, Mohd Mujtaba Akhtar · International Journal of Computers and Applications (2026) | TGRS Research Map | TGRS