Explainable Spatio-Temporal Attention Framework for GAN-Based Deepfake Detection
The rapid growth of Generative Adversarial Networks (GANs) has made synthetic media incredibly realistic and led to the emergence of deepfakes that can harm cybersecurity, trust in the digital world, and information integrity. The current state of the art in deepfake detection faces several issues, such as limited generalization across several manipulation methods and compressed media, and poor spatiotemporal learning capabilities. To solve these problems, this research develops a novel method that incorporates a robust spatiotemporal learning model that uses an Extreme Inception Network (XceptionNet) architecture to extract spatial features, a Bidirectional Long Short-Term Memory (BiLSTM) network for temporal sequence learning, and a temporal attention mechanism for adaptive frame analysis. The proposed solution was tested on a popular dataset, FaceForensics++, which includes multiple types of GAN-based manipulation, such as DeepFakes, Face2Face, FaceSwap, FaceShifter, and NeuralTextures. The results show improved accuracy in both binary and multiclass classification, along with Explainable Artificial Intelligence tools such as Gradient-weighted Class Activation Mapping (Grad-CAM) and SHapley Additive exPlanations (SHAP).
Authors
- Muhammad Hatim Binsawad (ORCID: https://orcid.org/0000-0003-0915-7058)
- Abdullah Alhejaili (ORCID: https://orcid.org/0000-0003-3234-1937)
Institutions
- King Abdulaziz University (SA)
Publication Details
- Journal
- Electronics
- Published
- 2026-09-21
- DOI
- https://doi.org/10.3390/electronics15184343
- Primary Topic
- Generative Adversarial Networks and Image Synthesis
- Type
- article
- Field-Weighted Citation Impact
- 0.00