End-to-End Reconstruction-Classification Learning for Face Forgery Detection

Junyi CaoChao MaTaiping YaoShen ChenShouhong DingXiaokang Yang

article2022CVPR373 citations

Proposes an end-to-end framework that models the distribution of genuine faces through reconstruction learning and pairs encoder-decoder features via multi-scale bipartite graphs to improve generalizability against unseen deepfake manipulation methods.

Listen

Rapid advances in digital manipulation have made creating highly realistic fake facial images and videos straightforward. Malicious use of these technologies threatens identity authentication, public information integrity, and security. Most existing detection systems look for specific manipulation artifacts present in their training data, such as local noise or texture irregularities. Consequently, they experience significant performance drops when confronted with unfamiliar forgery techniques or compressed media in real-world settings.

The article demonstrates a novel face forgery detection framework called RECCE (Reconstruction-Classification Learning). The objective is to establish a detection model that reliably identifies manipulated faces—including previously unseen manipulation types—by learning the common, compact representations of genuine human faces rather than overfitting to specific manipulation patterns.

The approach combines reconstruction and classification tasks into a unified, end-to-end framework. Instead of training a reconstruction network on all images, the model trains an encoder-decoder architecture solely to reconstruct authentic faces, using metric learning to compact real face representations and separate real from fake embeddings. To reason through forgery traces across different spatial resolutions, the framework connects encoder and decoder features using a multi-scale graph module. A reconstruction-guided attention module then uses pixel-level discrepancies between the input image and the reconstructed image to highlight suspected forgery areas for the final classifier. The methodology was evaluated against existing benchmark datasets containing diverse manipulation techniques and real-world internet videos, including FaceForensics++, Celeb-DF, WildDeepfake, and the Deepfake Detection Challenge dataset.

The experimental findings show that the proposed framework consistently outperforms existing methods, particularly in real-world and cross-dataset testing. In low-quality and heavily compressed video environments, the framework achieved an area under the curve of 95.02%, exceeding competing frequency-based detectors by 1.72%. When tested across unseen datasets to measure generalization, the model attained a 64.31% score on real-world internet videos where standard approaches dropped near 60%, and it outperformed existing methods by 4.57%. On the large-scale Deepfake Detection Challenge dataset, the framework surpassed existing state-of-the-art tools across accuracy, ranking metrics, and overall error loss. Furthermore, robustness evaluations demonstrated superior resilience against common digital perturbations, outperforming competing methods by 6.31% against blur and 4.44% against pixelation.

These results demonstrate that modeling the common structure of authentic faces allows security systems to treat unknown manipulation methods as outliers. This shifts deepfake detection from reactive pattern matching to proactive anomaly identification. For organizations managing media verification, content moderation, or biometric security, this approach mitigates the operational risk of detection failures caused by evolving forgery algorithms and platform compression.

Based on these findings, development teams should adopt reconstruction-classification frameworks that prioritize modeling authentic data distributions over specific artifact profiles. Next steps include implementing real-world pilot deployments in automated media review pipelines to evaluate end-to-end processing speeds and resource consumption. The primary boundary condition is that the system relies on an underlying image backbone and standard image-level supervision; confidence in its core performance gains is high across major industry benchmarks, though ongoing evaluation against emerging generative video techniques remains necessary.

Cover for End-to-End Reconstruction-Classification Learning for Face Forgery Detection

Abstract

Existing face forgery detectors mainly focus on specific forgery patterns like noise characteristics, local textures, or frequency statistics for forgery detection. This causes specialization of learned representations to known forgery patterns presented in the training set, and makes it difficult to detect forgeries with unknown patterns. In this paper, from a new perspective, we propose a forgery detection framework emphasizing the common compact representations of genuine faces based on reconstruction-classification learning. Reconstruction learning over real images enhances the learned representations to be aware of forgery patterns that are even unknown, while classification learning takes the charge of mining the essential discrepancy between real and fake images, facilitating the understanding of forgeries. To achieve better representations, instead of only using the encoder in reconstruction learning, we build bipartite graphs over the encoder and decoder features in a multi-scale fashion. We further exploit the reconstruction difference as guidance of forgery traces on the graph output as the final representation, which is fed into the classifier for forgery detection. The reconstruction and classification learning is optimized end-to-end. Extensive experiments on large-scale benchmark datasets demonstrate the superiority of the proposed method over state of the arts.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Proposed Method
  • 3.1. Reconstruction Learning
  • 3.2. Multi-scale Graph Reasoning
  • 3.3. Reconstruction Guided Attention
  • 3.4. Loss Function
  • 4. Experiments
  • 4.1. Experimental Setup
  • 4.2. Experimental Results
  • 4.3. Ablation Study
  • 4.4. Experimental Analysis
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — RECCE Architecture for Face Forgery Detection

    model/method

    The Reconstruction-Classification Learning (RECCE) framework detects facial forgeries by learning compact representations of genuine faces rather than overfitting to specific manipulation artifacts. It comprises three core interacting schemes optimized end-to-end:

    1. Reconstruction Learning Network: An encoder-decoder autoencoder F\mathcal{F} trained exclusively to reconstruct real facial images with injected white noise, forcing the model to capture the compact distribution of authentic human faces. When applied to fake facial images, the network fails to accurately reconstruct forged regions, exposing distributional discrepancies.
    2. Multi-Scale Graph Reasoning (MGR): Constructs bipartite graphs linking the feature vectors of the encoder output with latent features across multiple intermediate decoder blocks. This propagates and aggregates discrepancy information captured during decoding back to enrich the encoder representation.
    3. Reconstruction Guided Attention (RGA): Calculates a pixel-level difference mask m=∣X^−X∣m = |\hat{X} - X| between the reconstructed image X^\hat{X} and the input image XX. This mask is passed through convolutional layers to produce an attention map that spatially weights the enhanced feature maps from the MGR module, directing the subsequent binary classifier to focus on suspected manipulated regions.
  2. Knowl 2 — Genuine-Face Reconstruction and Embedding Metric Learning Formulations

    model/method

    In the RECCE framework, representation learning is constrained by two specialized losses applied to the encoder-decoder reconstruction network F\mathcal{F}:

    1. Denoising Reconstruction Loss (Lr\mathcal{L}_r): To prevent the autoencoder from learning a trivial identity mapping, zero-mean white noise is added to the input image X∈Rh×w×3X \in \mathbb{R}^{h \times w \times 3} to form X~\tilde{X}, yielding reconstruction X^=F(X~)\hat{X} = \mathcal{F}(\tilde{X}). The reconstruction loss is computed exclusively over genuine (real) samples in a mini-batch RR: Lr=1∣R∣∑i∈R∥X^i−Xi∥1\mathcal{L}_r = \frac{1}{|R|} \sum_{i \in R} \|\hat{X}_i - X_i\|_1 where ∣R∣|R| denotes the cardinality of real samples in the mini-batch.

    2. Metric-Learning Loss (Lm\mathcal{L}_m): Applied to global average pooled feature vectors Fˉ∈Rc\bar{F} \in \mathbb{R}^c extracted from the output of the final encoder block and each intermediate decoder block. It minimizes the distance between pairs of real samples while maximizing the distance between real and fake samples: Lm=1NRR∑i∈R,j∈Rd(Fˉi,Fˉj)−1NRF∑i∈R,j∈Fd(Fˉi,Fˉj)\mathcal{L}_m = \frac{1}{N_{RR}} \sum_{i \in R, j \in R} d(\bar{F}_i, \bar{F}_j) - \frac{1}{N_{RF}} \sum_{i \in R, j \in F} d(\bar{F}_i, \bar{F}_j) where RR and FF denote sets of real and fake samples in the batch; NRRN_{RR} and NRFN_{RF} are the total number of (real, real) and (real, fake) pairs, respectively; and d(a,b)d(a, b) is the cosine distance normalized to [0,1][0, 1]: d(a,b)=1−a∥a∥2⋅b∥b∥22d(a, b) = \frac{1 - \frac{a}{\|a\|_2} \cdot \frac{b}{\|b\|_2}}{2} Crucially, no compactness constraint is enforced on fake samples because forgery artifacts vary substantially across generation methods.

  3. Knowl 3 — Multi-Scale Graph Reasoning (MGR) Formulation

    model/method

    The Multi-Scale Graph Reasoning (MGR) module combines encoder features and multi-scale decoder features into bipartite graphs to capture forgery cues at varying spatial granularities.

    Let the encoder output feature map and a decoder block feature map at a specific scale be represented as vertex sets Venc={venci}i=1h1×w1V_{enc} = \{v^i_{enc}\}_{i=1}^{h_1 \times w_1} and Vdec={vdeci}i=1h2×w2V_{dec} = \{v^i_{dec}\}_{i=1}^{h_2 \times w_2}, respectively. For an encoder node venciv^i_{enc}, let N(venci)={vdeci,j}j=1N\mathcal{N}(v^i_{enc}) = \{v^{i,j}_{dec}\}_{j=1}^N denote its spatially corresponding local neighborhood in VdecV_{dec}.

    1. Node features are projected into a shared latent space via two neural networks g1(⋅)g_1(\cdot) and g2(⋅)g_2(\cdot): v~enci=g1(venci),v~deci,j=g2(vdeci,j)\tilde{v}^i_{enc} = g_1(v^i_{enc}), \quad \tilde{v}^{i,j}_{dec} = g_2(v^{i,j}_{dec})

    2. Bipartite attention coefficients aja_j measuring the importance of decoder node vdeci,jv^{i,j}_{dec} to encoder node venciv^i_{enc} are computed via concatenation [⋅∥⋅][\cdot \| \cdot] and a single-layer network ϕ\phi: aj=exp⁡(ϕ(v~enci∥v~deci,j))∑vdeci,l∈N(venci)exp⁡(ϕ(v~enci∥v~deci,l))a_j = \frac{\exp\left(\phi\left(\tilde{v}^i_{enc} \| \tilde{v}^{i,j}_{dec}\right)\right)}{\sum_{v^{i,l}_{dec} \in \mathcal{N}(v^i_{enc})} \exp\left(\phi\left(\tilde{v}^i_{enc} \| \tilde{v}^{i,l}_{dec}\right)\right)}

    3. Decoder features are aggregated with channel-wise modulation using a non-linear channel richness transform ψ(⋅)∈[0,1]c\psi(\cdot) \in [0, 1]^c: vaggi=∑j=1Najv~deci,j⊗(1−ψ(venci))v^i_{agg} = \sum_{j=1}^N a_j \tilde{v}^{i,j}_{dec} \otimes \left(1 - \psi(v^i_{enc})\right) where ⊗\otimes is element-wise multiplication, selectively boosting decoder channels where encoder channel activations are weak.

    4. Aggregated vectors {vaggi}\{v^i_{agg}\} across all decoder scales are concatenated with venciv^i_{enc}, passed through a sigmoid function followed by two fully-connected layers to yield the enhanced feature vector venhiv^i_{enh}, and assembled spatially into feature maps FenhF_{enh}.

  4. Knowl 4 — Reconstruction Guided Attention (RGA) Formulation

    model/method

    The Reconstruction Guided Attention (RGA) module leverages pixel-level reconstruction discrepancies to highlight suspected forgery traces on the enhanced feature representations.

    Given the input facial image X∈Rh×w×3X \in \mathbb{R}^{h \times w \times 3} and the reconstructed image X^=F(X~)\hat{X} = \mathcal{F}(\tilde{X}), the pixel-wise difference mask mm is defined as: m=∣X^−X∣m = |\hat{X} - X| where ∣⋅∣|\cdot| denotes the absolute value operation.

    Using the enhanced feature maps FenhF_{enh} produced by the Multi-Scale Graph Reasoning (MGR) module, spatial attention is generated and applied via: Fenh′=σ(f1(m))⊗f2(Fenh)F'_{enh} = \sigma(f_1(m)) \otimes f_2(F_{enh}) Fatt=Fenh′+FenhF_{att} = F'_{enh} + F_{enh} where f1f_1 and f2f_2 denote convolutional operations, σ\sigma is the sigmoid activation function, and ⊗\otimes represents element-wise multiplication. Spatial resizing of intermediate tensors is performed via bilinear interpolation where needed. The resulting attended representation FattF_{att} is fed into the classification head for binary real/fake prediction.

  5. Knowl 5 — RECCE End-to-End Joint Loss Function

    equation

    The overall objective function L\mathcal{L} of the RECCE framework jointly optimizes binary classification and genuine-face reconstruction in an end-to-end manner:

    L=Lcls+λ1Lr+λ2Lm\mathcal{L} = \mathcal{L}_{cls} + \lambda_1 \mathcal{L}_r + \lambda_2 \mathcal{L}_m

    where:

    • Lcls\mathcal{L}_{cls} is the standard binary cross-entropy loss between predicted labels y^\hat{y} and ground-truth authenticity labels y∈{0,1}y \in \{0, 1\}.
    • Lr\mathcal{L}_r is the L1L_1 reconstruction loss computed solely over genuine facial samples RR in the mini-batch: Lr=1∣R∣∑i∈R∥X^i−Xi∥1\mathcal{L}_r = \frac{1}{|R|} \sum_{i \in R} \|\hat{X}_i - X_i\|_1.
    • Lm\mathcal{L}_m is the metric-learning loss applied to global pooled features across encoder and decoder blocks, enforcing compactness among real samples and separation between real and fake samples.
    • λ1\lambda_1 and λ2\lambda_2 are balancing hyperparameters empirically set to λ1=0.1\lambda_1 = 0.1 and λ2=0.1\lambda_2 = 0.1.
  6. Knowl 6 — RECCE Implementation and Training Setup

    experimental setup

    The RECCE framework uses the Xception architecture as its feature extraction and encoding backbone. Key training hyperparameters and environment specifications include:

    • Optimizer: Adam with an initial learning rate of 2×10−42 \times 10^{-4} and a weight decay of 1×10−51 \times 10^{-5}.
    • Learning Rate Schedule: Step learning rate scheduler.
    • Batch Size: 32.
    • Loss Weights: λ1=0.1\lambda_1 = 0.1 and λ2=0.1\lambda_2 = 0.1 in the total loss function L=Lcls+λ1Lr+λ2Lm\mathcal{L} = \mathcal{L}_{cls} + \lambda_1 \mathcal{L}_r + \lambda_2 \mathcal{L}_m.
    • Data Augmentation: Random horizontal flipping only.
    • Benchmark Datasets Evaluated:
      1. FaceForensics++ (FF++): Includes 4 manipulation algorithms (Deepfakes [DF], Face2Face [F2F], FaceSwap [FS], NeuralTextures [NT]) under high-quality (c23) and low-quality/heavily compressed (c40) settings.
      2. Celeb-DF: 590 real and 5,639 high-quality DeepFake videos.
      3. WildDeepfake (WDF): In-the-wild dataset consisting of 3,805 real and 3,509 fake video sequences.
      4. Deepfake Detection Challenge (DFDC): 128,154 facial videos spanning 960 subjects with diverse manipulations and perturbations.
    • Metrics Reported: Accuracy (Acc, %), Area Under the ROC Curve (AUC, %), Equal Error Rate (EER, %), and LogLoss (for DFDC).
  7. Knowl 7 — Intra-Dataset Performance Comparison on Standard Benchmarks

    data/table

    Under intra-dataset evaluation (training and testing on splits of the same dataset), RECCE consistently outperforms prior face forgery detection approaches across multiple quality levels, particularly under heavy video compression (FF++ c40) and complex in-the-wild datasets (DFDC, WildDeepfake).

    Methods FF++ (c23) FF++ (c40) Celeb-DF WildDeepfake
    Acc (%) AUC (%) Acc (%) AUC (%) Acc (%) AUC (%) Acc (%) AUC (%)
    MesoNet 83.10 – 70.47 – – – 64.47 –
    Multi-task 85.65 85.43 81.30 75.59 – – – –
    Xception 95.73 96.30 86.86 89.30 97.90 99.73 77.25 86.76
    Face X-ray – 87.40 – 61.60 – – – –
    Two-branch 96.43 98.70 86.34 86.59 – – – –
    SPSL 91.50 95.32 81.57 82.82 – – – –
    RFM 95.69 98.79 87.06 89.83 97.96 99.94 77.38 83.92
    Freq-SCL 96.69 99.28 89.00 92.39 – – – –
    Add-Net 96.78 97.74 87.50 91.01 96.93 99.55 76.25 86.17
    F3^3-Net 97.52 98.10 90.43 93.30 95.95 98.93 80.66 87.53
    MultiAtt 97.60 99.29 88.69 90.40 97.92 99.94 82.86 90.71
    RECCE (Ours) 97.06 99.32 91.03 95.02 98.59 99.94 83.25 92.02

    On the challenging DFDC benchmark, RECCE achieves 81.20%81.20\% Acc, 91.33%91.33\% AUC, and 0.43410.4341 LogLoss, outperforming Xception (79.35%79.35\% Acc, 89.50%89.50\% AUC, 0.49160.4916 LogLoss), RFM (80.83%80.83\% Acc, 89.75%89.75\% AUC, 0.58100.5810 LogLoss), and MultiAtt (76.81%76.81\% Acc, 90.32%90.32\% AUC, 0.52910.5291 LogLoss).

  8. Knowl 8 — Cross-Dataset and Cross-Manipulation Generalization Performance

    data/table

    To evaluate generalization to unseen manipulation techniques and datasets, models trained on FaceForensics++ (c40) were tested directly on WildDeepfake, Celeb-DF, and DFDC without fine-tuning, and cross-manipulation tests were conducted across individual manipulation techniques within FF++ c40.

    Methods WildDeepfake Celeb-DF DFDC
    AUC (%) ↑\uparrow EER (%) ↓\downarrow AUC (%) ↑\uparrow EER (%) ↓\downarrow AUC (%) ↑\uparrow EER (%) ↓\downarrow
    Xception 62.72 40.65 61.80 41.73 63.61 40.58
    RFM 57.75 45.45 65.63 38.54 66.01 39.05
    Add-Net 62.35 41.42 65.29 38.90 64.78 40.23
    F3^3-Net 57.10 45.12 61.51 42.03 64.60 39.84
    MultiAtt 59.74 43.73 67.02 37.90 68.01 37.17
    RECCE (Ours) 64.31 40.53 68.71 35.73 69.06 36.08

    Under fine-grained cross-manipulation testing across Deepfakes (DF), Face2Face (F2F), FaceSwap (FS), and NeuralTextures (NT) on FF++ c40, RECCE achieves higher cross-technique average AUC scores:

    • Trained on DF: RECCE achieves 70.76%70.76\% cross average AUC vs. 63.13%63.13\% (Freq-SCL) and 66.58%66.58\% (MultiAtt).
    • Trained on F2F: RECCE achieves 70.95%70.95\% cross average AUC vs. 63.19%63.19\% (Freq-SCL) and 70.01%70.01\% (MultiAtt).
    • Trained on FS: RECCE achieves 67.84%67.84\% cross average AUC vs. 60.09%60.09\% (Freq-SCL) and 66.26%66.26\% (MultiAtt).
    • Trained on NT: RECCE achieves 74.47%74.47\% cross average AUC vs. 69.10%69.10\% (Freq-SCL) and 72.02%72.02\% (MultiAtt).
  9. Knowl 9 — Ablation Analysis of RECCE Architectural Components and Loss Constraints

    empirical result

    Ablation experiments conducted on the WildDeepfake dataset isolate the performance contribution of each architectural module and loss constraint in RECCE:

    1. Component Ablation:

      • Baseline Xception: 77.25%77.25\% Acc, 86.76%86.76\% AUC.
      • Baseline + Reconstruction Learning (Rec): 81.19%81.19\% Acc, 89.61%89.61\% AUC (+3.94%+3.94\% Acc, +2.85%+2.85\% AUC over baseline).
      • Baseline + Rec + MGR (no RGA): 81.48%81.48\% Acc, 91.10%91.10\% AUC (+1.49%+1.49\% AUC over Rec alone).
      • Baseline + Rec + RGA (no MGR): 82.15%82.15\% Acc, 89.71%89.71\% AUC.
      • Full RECCE (Rec + MGR + RGA): 83.25%83.25\% Acc, 92.02%92.02\% AUC.
    2. Reconstruction and Metric Constraint Ablation:

      • Reconstruction loss Lr\mathcal{L}_r applied to both real and fake faces (with Lm\mathcal{L}_m): 80.62%80.62\% Acc, 88.92%88.92\% AUC. Restricting Lr\mathcal{L}_r strictly to genuine faces yields an improvement of +2.63%+2.63\% Acc and +3.10%+3.10\% AUC, showing that reconstructing fake faces impairs the discovery of a compact real-face distribution.
      • RECCE without Metric Loss Lm\mathcal{L}_m: 81.36%81.36\% Acc, 90.49%90.49\% AUC. Adding Lm\mathcal{L}_m yields +1.89%+1.89\% Acc and +1.53%+1.53\% AUC by tightening the genuine-face cluster and pushing fake samples away in feature space.
  10. Knowl 10 — Robustness of RECCE Under Visual Perturbations

    data/table

    To evaluate resilience to common social media transformations, models were tested on the WildDeepfake dataset subjected to five perturbation types: image compression, Gaussian blur, contrast jitter, saturation jitter, and pixelation.

    Methods Compress Blur Contrast Saturate Pixelate Avg.
    Xception 86.01 78.29 81.90 84.96 66.24 79.48
    RFM 83.74 75.34 79.77 82.59 71.25 78.54
    Add-Net 83.34 79.66 84.46 85.13 64.33 79.38
    F3^3-Net 86.71 78.99 86.53 87.67 73.23 82.63
    MultiAtt 89.64 80.98 89.30 90.37 79.44 85.95
    RECCE (Ours) 89.65 87.29 91.19 91.74 83.88 88.75

    While frequency- and texture-dependent models undergo sharp performance drops under Gaussian blur (destroying frequency patterns) and pixelation (destroying texture details), RECCE maintains high AUC (87.29%87.29\% on blur, 83.88%83.88\% on pixelation), achieving an average AUC of 88.75%88.75\% (+2.80%+2.80\% higher than MultiAtt).

Coverage note — None was omitted; all primary methodological components, loss formulations, architectural modules, experimental results, generalization tests, ablations, and robustness evaluations are covered.

References

  1. 1.Darius Afchar, Vincent Nozick, Junichi Yamagishi, and Isao Echizen. Mesonet: A compact facial video forgery detection network. In WIFS, 2018.
  2. 2.Dmitri Bitouk, Neeraj Kumar, Samreen Dhillon, Peter N. Belhumeur, and Shree K. Nayar. Face swapping: Automatically replacing faces in photographs. ACM Trans. Graph., 27(3):39, 2008.
  3. 3.Shenhao Cao, Qin Zou, Xiuqing Mao, Dengpan Ye, and Zhongyuan Wang. Metric learning for anti-compression facial forgery detection. In ACM MM, 2021.
  4. 4.Caroline Chan, Shiry Ginosar, Tinghui Zhou, and Alexei A. Efros. Everybody dance now. In ICCV, 2019.
  5. 5.Olivier Chapelle, Bernhard Schölkopf, and Alexander Zien. Semi-supervised learning. IEEE Trans. Neural Networks, 20(3):542, 2009.
  6. 6.Shen Chen, Taiping Yao, Yang Chen, Shouhong Ding, Jilin Li, and Rongrong Ji. Local relation learning for face forgery detection. In AAAI, 2021.
  7. 7.François Chollet. Xception: Deep learning with depthwise separable convolutions. In CVPR, 2017.
  8. 8.Hao Dang, Feng Liu, Joel Stehouwer, Xiaoming Liu, and Anil K. Jain. On the detection of digital face manipulation. In CVPR, 2020.
  9. 9.Brian Dolhansky, Joanna Bitton, Ben Pflaum, Jikuo Lu, Russ Howes, Menglin Wang, and Cristian Canton Ferrer. The deepfake detection challenge (dfdc) dataset. arXiv preprint arXiv:2006.07397, 2020.
  10. 10.Mengnan Du, Shiva K. Pentyala, Yuening Li, and Xia Hu. Towards generalizable deepfake detection with locality-aware autoencoder. In CIKM, 2020.
  11. 11.Yue Gao, Fangyun Wei, Jianmin Bao, Shuyang Gu, Dong Chen, Fang Wen, and Zhouhui Lian. High-fidelity and arbitrary face editing. In CVPR, 2021.
  12. 12.Qiqi Gu, Shen Chen, Taiping Yao, Yang Chen, Shouhong Ding, and Ran Yi. Exploiting fine-grained face forgery clues via progressive enhancement learning. In AAAI, 2022.
  13. 13.Zhihao Gu, Yang Chen, Taiping Yao, Shouhong Ding, Jilin Li, Feiyue Huang, and Lizhuang Ma. Spatiotemporal inconsistency learning for deepfake video detection. In ACM MM, 2021.
  14. 14.Zhihao Gu, Yang Chen, Taiping Yao, Shouhong Ding, Jilin Li, and Lizhuang Ma. Delving into the local: Dynamic inconsistency learning for deepfake video detection. In AAAI, 2022.
  15. 15.Alexandros Haliassos, Konstantinos Vougioukas, Stavros Petridis, and Maja Pantic. Lips don't lie: A generalisable and robust approach to face forgery detection. In CVPR, 2021.
  16. 16.Zhizhong Han, Xiyang Wang, Yu-Shen Liu, and Matthias Zwicker. Multi-angle point cloud-vae: Unsupervised feature learning for 3d point clouds from multiple angles by joint self-reconstruction and half-to-half prediction. In ICCV, 2019.
  17. 17.Liming Jiang, Ren Li, Wayne Wu, Chen Qian, and Chen Change Loy. Deeperforensics-1.0: A large-scale dataset for real-world face forgery detection. In CVPR, 2020.
  18. 18.Yuming Jiang, Ziqi Huang, Xingang Pan, Chen Change Loy, and Ziwei Liu. Talk-to-edit: Fine-grained facial editing via dialog. In ICCV, 2021.
  19. 19.Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In ICLR, 2015.
  20. 20.Iryna Korshunova, Wenzhe Shi, Joni Dambre, and Lucas Theis. Fast face-swap using convolutional neural networks. In ICCV, 2017.
  21. 21.Akash Kumar, Arnav Bhavsar, and Rajesh Verma. Detecting deepfakes with metric learning. In International Workshop on Biometrics and Forensics, 2020.
  22. 22.Jiaming Li, Hongtao Xie, Jiahong Li, Zhongyuan Wang, and Yongdong Zhang. Frequency-aware discriminative feature learning supervised by single-center loss for face forgery detection. In CVPR, 2021.
  23. 23.Lingzhi Li, Jianmin Bao, Ting Zhang, Hao Yang, Dong Chen, Fang Wen, and Baining Guo. Face X-ray for more general face forgery detection. In CVPR, 2020.
  24. 24.Yuezun Li, Xin Yang, Pu Sun, Honggang Qi, and Siwei Lyu. Celeb-DF: A large-scale challenging dataset for deepfake forensics. In CVPR, 2020.
  25. 25.Honggu Liu, Xiaodan Li, Wenbo Zhou, Yuefeng Chen, Yuan He, Hui Xue, Weiming Zhang, and Nenghai Yu. Spatial-phase shallow learning: Rethinking face forgery detection in frequency domain. In CVPR, 2021.
  26. 26.Xinhai Liu, Xinchen Liu, Zhizhong Han, and Yu-Shen Liu. Spu-net: Self-supervised point cloud upsampling by coarse-to-fine reconstruction with self-projection optimization. arXiv preprint arXiv:2012.04439, 2020.
  27. 27.Siwei Lyu. Deepfake detection: Current challenges and next steps. In ICME Workshop, 2020.
  28. 28.Lars Maaløe, Casper Kaae Sønderby, Søren Kaae Sønderby, and Ole Winther. Auxiliary deep generative models. In ICML, 2016.
  29. 29.Iacopo Masi, Aditya Killekar, Royston Marian Mascarenhas, Shenoy Pratik Gurudatt, and Wael AbdAlmageed. Two-branch recurrent network for isolating deepfakes in videos. In ECCV, 2020.
  30. 30.Huy H. Nguyen, Fuming Fang, Junichi Yamagishi, and Isao Echizen. Multi-task learning for detecting and segmenting manipulated facial images and videos. In BTAS, 2019.
  31. 31.Huy H. Nguyen, Junichi Yamagishi, and Isao Echizen. Capsule-forensics: Using capsule networks to detect forged images and videos. In ICASSP, 2019.
  32. 32.Deepak Pathak, Philipp Krähenbóhl, Jeff Donahue, Trevor Darrell, and Alexei A. Efros. Context encoders: Feature learning by inpainting. In CVPR, 2016.
  33. 33.Yuyang Qian, Guojun Yin, Lu Sheng, Zixuan Chen, and Jing Shao. Thinking in frequency: Face forgery detection by mining frequency-aware clues. In ECCV, 2020.
  34. 34.Antti Rasmus, Mathias Berglund, Mikko Honkala, Harri Valpola, and Tapani Raiko. Semi-supervised learning with ladder networks. In NeurIPS, 2015.
  35. 35.Andreas Rössler, Davide Cozzolino, Luisa Verdoliva, Christian Riess, Justus Thies, and Matthias Nießner. Faceforensics++: Learning to detect manipulated facial images. In ICCV, 2019.
  36. 36.Lukas Ruff, Robert A. Vandermeulen, Nico Görnitz, Alexander Binder, Emmanuel Müller, Klaus-Robert Müller, and Marius Kloft. Deep semi-supervised anomaly detection. In ICLR, 2020.
  37. 37.Florian Schroff, Dmitry Kalenichenko, and James Philbin. Facenet: A unified embedding for face recognition and clustering. In CVPR, 2015.
  38. 38.Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. Grad-CAM: Visual explanations from deep networks via gradient-based localization. In CVPR, 2017.
  39. 39.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. In ICLR, 2015.
  40. 40.Ke Sun, Taiping Yao, Shen Chen, Shouhong Ding, Jilin Li, and Rongrong Ji. Dual contrastive learning for general face forgery detection. In AAAI, 2022.
  41. 41.Supasorn Suwajanakorn, Steven M. Seitz, and Ira Kemelmacher-Shlizerman. Synthesizing Obama: Learning lip sync from audio. ACM Trans. Graph., 36(4):95:1–95:13, 2017.
  42. 42.Justus Thies, Michael Zollhöfer, and Matthias Nießner. Deferred neural rendering: Image synthesis using neural textures. ACM Trans. Graph., 38(4):66:1–66:12, 2019.
  43. 43.Laurens van der Maaten and Geoffrey Hinton. Visualizing data using t-sne. Journal of Machine Learning Research, 9(86):2579–2605, 2008.
  44. 44.Petar Velickovic, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio. Graph attention networks. In ICLR, 2018.
  45. 45.Chengrui Wang and Weihong Deng. Representative forgery mining for fake face detection. In CVPR, 2021.
  46. 46.Xinyao Wang, Taiping Yao, Shouhong Ding, and Lizhuang Ma. Face manipulation detection via auxiliary supervision. In ICONIP, 2020.
  47. 47.Zhuhui Wang, Shijie Wang, Haojie Li, Zhi Dou, and Jianjun Li. Graph-propagation based correlation learning for weakly supervised fine-grained image classification. In AAAI, 2020.
  48. 48.Yandong Wen, Kaipeng Zhang, Zhifeng Li, and Yu Qiao. A discriminative feature learning approach for deep face recognition. In ECCV, 2016.
  49. 49.Davis Wertheimer, Luming Tang, and Bharath Hariharan. Few-shot classification with feature map reconstruction networks. In CVPR, 2021.
  50. 50.Wayne Wu, Yunxuan Zhang, Cheng Li, Chen Qian, and Chen Change Loy. ReenactGAN: Learning to reenact faces via boundary transfer. In ECCV, 2018.
  51. 51.Burhaneddin Yaman, Chetan Shenoy, Zilin Deng, Steen Moeller, Hossam El-Rewaidy, Reza Nezafat, and Mehmet Akçakaya. Self-supervised physics-guided deep learning reconstruction for high-resolution 3D LGE CMR. In International Symposium on Biomedical Imaging, 2021.
  52. 52.Jie Yang, Yong Shi, and Zhiquan Qi. DFR: Deep feature reconstruction for unsupervised anomaly segmentation. arXiv preprint arXiv:2012.07122, 2020.
  53. 53.Guangming Yao, Yi Yuan, Tianjia Shao, Shuang Li, Shanqi Liu, Yong Liu, Mengmeng Wang, and Kun Zhou. One-shot face reenactment using appearance adaptive normalization. In AAAI, 2021.
  54. 54.Ryota Yoshihashi, Wen Shao, Rei Kawakami, Shaodi You, Makoto Iida, and Takeshi Naemura. Classification-reconstruction learning for open-set recognition. In CVPR, 2019.
  55. 55.Hanqing Zhao, Wenbo Zhou, Dongdong Chen, Tianyi Wei, Weiming Zhang, and Nenghai Yu. Multi-attentional deepfake detection. In CVPR, 2021.
  56. 56.Yifan Zhao, Ke Yan, Feiyue Huang, and Jia Li. Graph-based high-order relation discovery for fine-grained recognition. In CVPR, 2021.
  57. 57.Hong-Yu Zhou, Chixiang Lu, Sibei Yang, Xiaoguang Han, and Yizhou Yu. Preservational learning improves self-supervised medical image models by reconstructing diverse contexts. In ICCV, 2021.
  58. 58.Peng Zhou, Xintong Han, Vlad I. Morariu, and Larry S. Davis. Two-stream neural networks for tampered face detection. In CVPR Workshops, 2017.
  59. 59.Xiangyu Zhu, Hao Wang, Hongyan Fei, Zhen Lei, and Stan Z. Li. Face forgery detection by 3D decomposition. In CVPR, 2021.
  60. 60.Bojia Zi, Minghao Chang, Jingjing Chen, Xingjun Ma, and Yu-Gang Jiang. WildDeepfake: A challenging real-world dataset for deepfake detection. In ACM MM, 2020.

Citation

MLA
Cao, J., et al. “End-to-End Reconstruction-Classification Learning for Face Forgery Detection”. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 4103–12, https://doi.org/10.1109/CVPR52688.2022.00408.
APA
Cao, J., Ma, C., Yao, T., Chen, S., Ding, S., & Yang, X. (2022). End-to-End Reconstruction-Classification Learning for Face Forgery Detection. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 4103–4112. https://doi.org/10.1109/CVPR52688.2022.00408
Chicago
Cao, J., C. Ma, T. Yao, S. Chen, S. Ding, and X. Yang. 2022. “End-to-End Reconstruction-Classification Learning for Face Forgery Detection”. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 4103–12. https://doi.org/10.1109/CVPR52688.2022.00408.
Harvard
Cao, J. et al. (2022) “End-to-End Reconstruction-Classification Learning for Face Forgery Detection”, 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 4103–4112. Available at: https://doi.org/10.1109/CVPR52688.2022.00408.
Vancouver
1. Cao J, Ma C, Yao T, Chen S, Ding S, Yang X (2022) End-to-End Reconstruction-Classification Learning for Face Forgery Detection. In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 4103–4112

BibTeX

@inproceedings{Cao_2022, title={End-to-End Reconstruction-Classification Learning for Face Forgery Detection}, url={http://dx.doi.org/10.1109/CVPR52688.2022.00408}, DOI={10.1109/cvpr52688.2022.00408}, booktitle={2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Cao, Junyi and Ma, Chao and Yao, Taiping and Chen, Shen and Ding, Shouhong and Yang, Xiaokang}, year={2022}, month=June, pages={4103–4112} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE