An Improved Deepfake Detection Approach Using Hybrid Architecture Based on EfficientNet and Vision Transformer

The rapid spread of deepfake content poses a significant threat to the credibility of digital multimedia, creating an urgent need for accurate and robust detection methods. This paper proposes a hybrid deep learning framework that combines EfficientNet-B0 and Vision Transformer (ViT-B/16) through feature concatenation to exploit both local spatial representations and global contextual dependencies for image-level deepfake detection. The proposed model was trained and evaluated on the deepfake and real images dataset. Experimental results demonstrate that the proposed framework achieves an accuracy of 0.9870, an F1-score of 0.9871, and an AUC of 0.9990. Additional cross-dataset evaluation on the CelebDF-v2 image dataset and robustness experiments under common image degradations further demonstrates the strong generalization capability and practical applicability of the proposed approach. These results confirm that integrating CNN-based and Transformer-based feature extraction provides an effective and reliable solution for deepfake image detection.

Authors

Institutions

Publication Details

Journal
Journal of Imaging
Published
2026-09-09
DOI
https://doi.org/10.3390/jimaging12090427
Primary Topic
Generative Adversarial Networks and Image Synthesis
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

An Improved Deepfake Detection Approach Using Hybrid Architecture Based on EfficientNet and Vision Transformer

Omar Banimelhem, Abeer O. Alsharu
Journal of Imaging
Generative Adversarial Networks and Image Synthesis
article

An Improved Deepfake Detection Approach Using Hybrid Architecture Based on EfficientNet and Vision Transformer

Omar Banimelhem, Abeer O. Alsharu
article en

Abstract

The rapid spread of deepfake content poses a significant threat to the credibility of digital multimedia, creating an urgent need for accurate and robust detection methods. This paper proposes a hybrid deep learning framework that combines EfficientNet-B0 and Vision Transformer (ViT-B/16) through feature concatenation to exploit both local spatial representations and global contextual dependencies for image-level deepfake detection. The proposed model was trained and evaluated on the deepfake and real images dataset. Experimental results demonstrate that the proposed framework achieves an accuracy of 0.9870, an F1-score of 0.9871, and an AUC of 0.9990. Additional cross-dataset evaluation on the CelebDF-v2 image dataset and robustness experiments under common image degradations further demonstrates the strong generalization capability and practical applicability of the proposed approach. These results confirm that integrating CNN-based and Transformer-based feature extraction provides an effective and reliable solution for deepfake image detection.

Journal of ImagingVol. 12(9)
Jordan University of Science and Technology (JO)
Openalex Percentile: Top 13%
Generative Adversarial Networks and Image Synthesis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

An Improved Deepfake Detection Approach Using Hybrid Architecture Based on EfficientNet and Vision Transformer — Omar Banimelhem, Abeer O. Alsharu · Journal of Imaging (2026) | TGRS Research Map | TGRS