End-to-End Reconstruction-Classification Learning for Face Forgery Detection
Junyi CaoChao MaTaiping YaoShen ChenShouhong DingXiaokang Yang
Proposes an end-to-end framework that models the distribution of genuine faces through reconstruction learning and pairs encoder-decoder features via multi-scale bipartite graphs to improve generalizability against unseen deepfake manipulation methods.
Rapid advances in digital manipulation have made creating highly realistic fake facial images and videos straightforward. Malicious use of these technologies threatens identity authentication, public information integrity, and security. Most existing detection systems look for specific manipulation artifacts present in their training data, such as local noise or texture irregularities. Consequently, they experience significant performance drops when confronted with unfamiliar forgery techniques or compressed media in real-world settings.
The article demonstrates a novel face forgery detection framework called RECCE (Reconstruction-Classification Learning). The objective is to establish a detection model that reliably identifies manipulated faces—including previously unseen manipulation types—by learning the common, compact representations of genuine human faces rather than overfitting to specific manipulation patterns.
The approach combines reconstruction and classification tasks into a unified, end-to-end framework. Instead of training a reconstruction network on all images, the model trains an encoder-decoder architecture solely to reconstruct authentic faces, using metric learning to compact real face representations and separate real from fake embeddings. To reason through forgery traces across different spatial resolutions, the framework connects encoder and decoder features using a multi-scale graph module. A reconstruction-guided attention module then uses pixel-level discrepancies between the input image and the reconstructed image to highlight suspected forgery areas for the final classifier. The methodology was evaluated against existing benchmark datasets containing diverse manipulation techniques and real-world internet videos, including FaceForensics++, Celeb-DF, WildDeepfake, and the Deepfake Detection Challenge dataset.
The experimental findings show that the proposed framework consistently outperforms existing methods, particularly in real-world and cross-dataset testing. In low-quality and heavily compressed video environments, the framework achieved an area under the curve of 95.02%, exceeding competing frequency-based detectors by 1.72%. When tested across unseen datasets to measure generalization, the model attained a 64.31% score on real-world internet videos where standard approaches dropped near 60%, and it outperformed existing methods by 4.57%. On the large-scale Deepfake Detection Challenge dataset, the framework surpassed existing state-of-the-art tools across accuracy, ranking metrics, and overall error loss. Furthermore, robustness evaluations demonstrated superior resilience against common digital perturbations, outperforming competing methods by 6.31% against blur and 4.44% against pixelation.
These results demonstrate that modeling the common structure of authentic faces allows security systems to treat unknown manipulation methods as outliers. This shifts deepfake detection from reactive pattern matching to proactive anomaly identification. For organizations managing media verification, content moderation, or biometric security, this approach mitigates the operational risk of detection failures caused by evolving forgery algorithms and platform compression.
Based on these findings, development teams should adopt reconstruction-classification frameworks that prioritize modeling authentic data distributions over specific artifact profiles. Next steps include implementing real-world pilot deployments in automated media review pipelines to evaluate end-to-end processing speeds and resource consumption. The primary boundary condition is that the system relies on an underlying image backbone and standard image-level supervision; confidence in its core performance gains is high across major industry benchmarks, though ongoing evaluation against emerging generative video techniques remains necessary.
- Paper: FaceForensics++: Learning to Detect Manipulated Facial Images, Andreas Rössler et al. (2019). It introduces the foundational benchmark and baseline architectures for facial manipulation detection upon which the source's evaluation and problem framing rely.
- Paper: Celeb-DF: A Large-Scale Challenging Dataset for DeepFake Forensics, Yuezun Li et al. (2019). It provides the challenging, high-quality deepfake benchmark dataset that highlights generalization failures on unseen manipulations addressed by the source.
- Paper: MesoNet: a Compact Facial Video Forgery Detection Network, Darius Afchar et al. (2018). It establishes early convolutional network approaches for detecting facial video tampering under internet compression, motivating the need for more generalizable feature representations.
- Paper: CNN-Generated Images Are Surprisingly Easy to Spot… for Now, Sheng-Yu Wang et al. (2019). It examines how classifiers generalize across unseen generative models, providing crucial context for the source's objective of detecting unknown forgery patterns.
- Paper: GANomaly: Semi-Supervised Anomaly Detection via Adversarial Training, Samet Akcay et al. (2018). It introduces reconstruction-based anomaly detection via autoencoders trained exclusively on normal data, serving as a conceptual foundation for reconstruction learning on genuine faces.
- Paper: Memorizing Normality to Detect Anomaly: Memory-Augmented Deep Autoencoder for Unsupervised Anomaly Detection, Dong Gong et al. (2019). It models normal patterns through constrained autoencoder reconstruction to detect anomalies, directly informing the source's strategy of using reconstruction differences to expose forgeries.
- Paper: Implicit Identity Driven Deepfake Face Swapping Detection, Baojin Huang et al. (2023). It extends the quest for generalizable deepfake detection by contrasting explicit appearance with implicit identity residuals to overcome unseen manipulations.
- Paper: Hierarchical Fine-Grained Image Forgery Detection and Localization, Xiao Guo et al. (2023). It expands forgery detection beyond binary classification by presenting a hierarchical framework that performs simultaneous manipulation localization and fine-grained method attribution.
- Paper: Frequency-Aware Deepfake Detection: Improving Generalizability through Frequency Space Domain Learning, Chuangchuang Tan et al. (2024). It advances generalizability against unseen generators by shifting representation learning directly into the frequency domain.
- Paper: Rethinking the Up-Sampling Operations in CNN-Based Generative Network for Generalizable Deepfake Detection, Chuangchuang Tan et al. (2024). It further develops generalization capabilities across diverse generative models by identifying universal spatial artifacts from up-sampling operations.
- Paper: AVFF: Audio-Visual Feature Fusion for Video Deepfake Detection, Trevine Oorloff et al. (2024). It broadens video deepfake detection beyond visual reconstruction by fusing multimodal audio-visual correspondences.
