An Improved Deepfake Detection Approach Using Hybrid Architecture Based on EfficientNet and Vision Transformer
The rapid spread of deepfake content poses a significant threat to the credibility of digital multimedia, creating an urgent need for accurate and robust detection methods. This paper proposes a hybrid deep learning framework that combines EfficientNet-B0 and Vision Transformer (ViT-B/16) through feature concatenation to exploit both local spatial representations and global contextual dependencies for image-level deepfake detection. The proposed model was trained and evaluated on the deepfake and real images dataset. Experimental results demonstrate that the proposed framework achieves an accuracy of 0.9870, an F1-score of 0.9871, and an AUC of 0.9990. Additional cross-dataset evaluation on the CelebDF-v2 image dataset and robustness experiments under common image degradations further demonstrates the strong generalization capability and practical applicability of the proposed approach. These results confirm that integrating CNN-based and Transformer-based feature extraction provides an effective and reliable solution for deepfake image detection.
Authors
- Omar Banimelhem (ORCID: https://orcid.org/0000-0002-2649-7161)
- Abeer O. Alsharu
Institutions
- Jordan University of Science and Technology (JO)
Publication Details
- Journal
- Journal of Imaging
- Published
- 2026-09-09
- DOI
- https://doi.org/10.3390/jimaging12090427
- Primary Topic
- Generative Adversarial Networks and Image Synthesis
- Type
- article
- Field-Weighted Citation Impact
- 0.00