LGF-Net: a spatial-frequency deepfake detection network with learnable gabor filters and gated cross-modal Fusion
Abstract Deepfakes have become increasingly realistic due to recent advances in face manipulation techniques, making reliable detection in unconstrained environments more challenging. Existing spatial-frequency deepfake detection methods often rely on fixed hand-crafted frequency transforms and simple fusion strategies, which may limit their adaptability and cross-dataset generalization. To address these limitations, we propose LGF-Net, a unified framework that jointly models spatial semantics and adaptive spectral cues for deepfake detection. The Frequency Representation Module employs learnable Gabor filters and a frequency-aware attention mechanism to capture manipulation-specific spectral patterns. Moreover, the Spatial Representation Module uses multi-rate dilated convolutions to model both subtle local artifacts and long-range structural inconsistencies, while a gated cross-modal fusion module integrates the two representations into a compact forensic descriptor. Experimental results on FF++ (HQ), Celeb-DF (V2), DPDC, and DFD show that LGF-Net achieves competitive intra-dataset and cross-dataset performance compared with several state-of-the-art deepfake detection methods.
Authors
- Deepika Koundal (ORCID: https://orcid.org/0000-0003-1688-8772)
- Sanjeev Kumar (ORCID: https://orcid.org/0000-0001-7728-3668)
- Mukesh Pandey (ORCID: https://orcid.org/0009-0000-3343-1294)
Institutions
- University of Eastern Finland (FI)
- University of Petroleum and Energy Studies (IN)
Publication Details
- Journal
- Discover Computing
- Published
- 2026-09-15
- DOI
- https://doi.org/10.1007/s10791-026-10548-5
- Primary Topic
- Face recognition and analysis
- Type
- article
- Field-Weighted Citation Impact
- 0.00