Dual-domain self-supervised feature alignment via spectral–spatial representation learning for deepfake anomaly detection

Deepfake anomaly detection has become increasingly critical in visual security applications. However, many existing methods rely heavily on large-scale supervised forgery annotations and often struggle to capture subtle manipulation traces distributed across both spatial and frequency domains. To address these limitations, we propose a Self-Supervised Dual-Domain Alignment (SDDA) framework that jointly learns spatial and spectral representations without requiring labeled forged samples. Specifically, SDDA employs two parameter-sharing encoders to extract complementary features from RGB images and their frequency-domain counterparts, while a cross-domain alignment module enforces consistency between spatial textures and spectral signatures. To provide explicit and reproducible self-supervised supervision, four complementary descriptors are further derived from FAN feature maps: Local Structural Deviation (LSD) for local gradient inconsistency, Global Pattern Discrepancy (GPD) for holistic channel-correlation differences, Local Residual Difference (LRD) for residual-level inconsistency, and Total Consistency Deviation (TCD) for multi-layer feature deviation. These descriptors are predicted from the fused spatial-frequency embedding, encouraging the model to learn manipulation-sensitive representations in the absence of fake samples. In addition, a dual-domain contrastive objective enhances the discrimination of subtle forgery-related anomalies while improving robustness against real-world degradations such as compression and illumination variations. Experimental results demonstrate that SDDA consistently outperforms state-of-the-art baselines, particularly under challenging unseen-manipulation scenarios. Overall, the proposed framework provides an interpretable and label-efficient solution for dual-domain deepfake anomaly detection.

Authors

Institutions

Publication Details

Journal
PLoS ONE
Published
2026-09-11
DOI
https://doi.org/10.1371/journal.pone.0358049
Primary Topic
Digital Media Forensic Detection
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Dual-domain self-supervised feature alignment via spectral–spatial representation learning for deepfake anomaly detection

Yi He
PLoS ONE
Digital Media Forensic Detection
article

Dual-domain self-supervised feature alignment via spectral–spatial representation learning for deepfake anomaly detection

Yi He
article en

Abstract

Deepfake anomaly detection has become increasingly critical in visual security applications. However, many existing methods rely heavily on large-scale supervised forgery annotations and often struggle to capture subtle manipulation traces distributed across both spatial and frequency domains. To address these limitations, we propose a Self-Supervised Dual-Domain Alignment (SDDA) framework that jointly learns spatial and spectral representations without requiring labeled forged samples. Specifically, SDDA employs two parameter-sharing encoders to extract complementary features from RGB images and their frequency-domain counterparts, while a cross-domain alignment module enforces consistency between spatial textures and spectral signatures. To provide explicit and reproducible self-supervised supervision, four complementary descriptors are further derived from FAN feature maps: Local Structural Deviation (LSD) for local gradient inconsistency, Global Pattern Discrepancy (GPD) for holistic channel-correlation differences, Local Residual Difference (LRD) for residual-level inconsistency, and Total Consistency Deviation (TCD) for multi-layer feature deviation. These descriptors are predicted from the fused spatial-frequency embedding, encouraging the model to learn manipulation-sensitive representations in the absence of fake samples. In addition, a dual-domain contrastive objective enhances the discrimination of subtle forgery-related anomalies while improving robustness against real-world degradations such as compression and illumination variations. Experimental results demonstrate that SDDA consistently outperforms state-of-the-art baselines, particularly under challenging unseen-manipulation scenarios. Overall, the proposed framework provides an interpretable and label-efficient solution for dual-domain deepfake anomaly detection.

PLoS ONEVol. 21(9)
Batangas State University (PH), Henan University of Urban Construction (CN)
Reduced inequalities
Openalex Percentile: Top 13%
Digital Media Forensic Detection
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.