Hierarchical Fine-Grained Image Forgery Detection and Localization

Xiao GuoXiaohong LiuZhiyuan RenSteven GroszIacopo MasiXiaoming Liu

article2023CVPR230 citations

Proposes a hierarchical multi-branch framework and a structured dataset that unify image forgery detection and pixel-level localization across both deep generative synthesis and conventional image editing by explicitly modeling multi-level manipulation attributes.

Listen

Rapid advancements in artificial intelligence and accessible photo-editing software have made generating convincing synthetic imagery and manipulating visual content easier than ever. This widespread proliferation of manipulated media fuels misinformation campaigns, posing severe reputational, legal, and operational risks across modern digital ecosystems. While prior forensic tools generally specialized in either identifying conventional image editing (such as splicing and inpainting) or detecting entirely synthesized images, real-world deployment requires a unified framework capable of handling both manipulation types simultaneously.

The article demonstrates a unified framework designed to perform simultaneous image forgery detection, pixel-level manipulation localization, and fine-grained classification of the specific generation method used. To achieve this, the authors introduce a hierarchical multi-branch architecture alongside a comprehensive forensic dataset spanning modern diffusion models, generative adversarial networks, and standard editing techniques.

To bridge the gap across manipulation types, the authors formulated a four-level hierarchical structure that classifies manipulated imagery from broad categories down to specific source tools. They built a benchmark containing over 1.8 million images spanning thirteen distinct manipulation methods paired with high-resolution pixel masks. The proposed network processes both color and frequency domains using multiple specialized branches corresponding to different hierarchical levels, a self-attention localization module optimized via metric learning, and masked partial convolutions that feed localized spatial cues directly into attribute classification.

Across extensive testing on seven distinct benchmarks, the unified framework consistently matched or outperformed existing state-of-the-art approaches. On the newly introduced dataset, it achieved an overall image detection performance of 96.8% Area Under the Curve (AUC) and a localization performance of 95.3% AUC, outperforming specialized baselines by 2.6% to 9.3%. In fine-grained attribute classification, the model attained an overall score of 87.6%, compared to under 50% for previous generative model attribution techniques. Ablation experiments confirmed that removing the hierarchical path prediction decreased detection accuracy by 3.6% AUC, demonstrating the necessity of structured dependency learning.

These findings indicate that integrating pixel-level localization with hierarchical attribute attribution significantly bolsters general forensic accuracy. Rather than treating manipulation types as mutually exclusive, leveraging coarse-to-fine relational dependencies enables organizations to better track manipulation lineage and assess media tampering risks in open environments. Practically, this approach provides actionable provenance data essential for compliance, legal verification, and automated content moderation.

Organizations seeking to implement visual forensics should transition away from single-domain detectors toward multi-level unified models capable of simultaneous detection and localization. Further engineering work should focus on scaling forensic datasets and evaluating defenses against rapid developments in diffusion-based inpainting. While the framework demonstrates strong robustness, confidence is moderated by observed classification challenges in edge cases, including small inpainting patches, images with extreme lighting, and closely related generative model architectures.

Cover for Hierarchical Fine-Grained Image Forgery Detection and Localization

Abstract

Differences in forgery attributes of images generated in CNN-synthesized and image-editing domains are large, and such differences make a unified image forgery detection and localization (IFDL) challenging. To this end, we present a hierarchical fine-grained formulation for IFDL representation learning. Specifically, we first represent forgery attributes of a manipulated image with multiple labels at different levels. Then we perform fine-grained classification at these levels using the hierarchical dependency between them. As a result, the algorithm is encouraged to learn both comprehensive features and inherent hierarchical nature of different forgery attributes, thereby improving the IFDL representation. Our proposed IFDL framework contains three components: multi-branch feature extractor, localization and classification modules. Each branch of the feature extractor learns to classify forgery attributes at one level, while localization and classification modules segment the pixel-level forgery region and detect image-level forgery, respectively. Lastly, we construct a hierarchical fine-grained dataset to facilitate our study. We demonstrate the effectiveness of our method on 7 different benchmarks, for both tasks of IFDL and forgery attribute classification. Our source code and dataset can be found: github.com/CHELSEA234/HiFi-IFDL.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. HiFi-Net
  • 3.1. Multi-Branch Feature Extractor
  • 3.2. Localization Module
  • 3.3. Classification Module
  • 3.4. Training and Inference
  • 4. Hierarchical Fine-grained IFDL dataset
  • 5. Experiments
  • 5.1. Image Forgery Detection and Localization
  • 5.1.2 Image Editing Datasets
  • 5.1.3 Diverse Fake Face Dataset
  • 5.2. Ablation Study
  • 5.3. Forgery Attribute performance
  • 6. Conclusion
  • References

Knowls

  1. Knowl 1 — HiFi-Net Architecture for Image Forgery Detection and Localization

    model/method

    The Hierarchical Fine-grained Network (HiFi-Net) is designed to perform unified image forgery detection and pixel-level localization across both CNN-synthesized and traditional image editing domains for an input image X∈R3×W×HX \in \mathbb{R}^{3 \times W \times H}. The architecture contains three core components:

    1. Multi-Branch Feature Extractor: Features are extracted using dual color and frequency blocks, where the frequency block applies a Laplacian of Gaussian (LoG) operator to capture artifact patterns across both RGB and spectral domains. Four hierarchical branches θb\theta_b (b∈{1,2,3,4}b \in \{1, 2, 3, 4\}) generate multi-resolution feature maps. Branch θ1\theta_1 operates on low-resolution representations for coarse binary classification, while branch θ4\theta_4 processes the highest-resolution feature representation F∈R512×W×HF \in \mathbb{R}^{512 \times W \times H} to capture fine-grained spatial manipulation cues.

    2. Localization Module: Connected to the output FF of the highest-resolution branch θ4\theta_4, this module applies 1×11 \times 1 convolutions to construct projection maps Fg=g(F)F_g = g(F), Fϕ=ϕ(F)F_\phi = \phi(F), and Fψ=ψ(F)F_\psi = \psi(F). A spatial self-attention matrix is computed as: As=softmax(FϕTFψ)A_s = \text{softmax}\left(F_\phi^T F_\psi\right) This attention map projects FgF_g into a global feature representation F′=AsFg∈R512×W×HF' = A_s F_g \in \mathbb{R}^{512 \times W \times H}, which is used to output the predicted binary localization mask M^∈RW×H\hat{M} \in \mathbb{R}^{W \times H}.

    3. Classification Module: Reuses the localization mask M^\hat{M} to construct a masked image Xmask=X⊙M^X_{\text{mask}} = X \odot \hat{M} containing only the predicted manipulated regions. Partial convolutions (PConv) process XmaskX_{\text{mask}} to extract prior representations that feed into the hierarchical attribute classification branches.

  2. Knowl 2 — Hierarchical Forgery Attribute Taxonomy and Path Conditioning

    model/method

    Rather than treating forgery types as flat and mutually exclusive classes, HiFi-Net organizes forgery attributes into a 4-level directed hierarchical tree:

    • Level 1 (2 classes): Fully-synthesized vs. Partially-manipulated.
    • Level 2 (4 classes): Diffusion model, GAN-based method, Conventional image editing, CNN-based partial manipulation.
    • Level 3: Conditional vs. Unconditional generation.
    • Level 4 (14 classes): Real class, and 13 specific forgery methods (DDPM, DDIM, GDM, LDM, StarGANv2, HiSD, StyleGANv2-ada, StyleGAN3, STGAN, FaceShifter, Splicing, Inpainting, and Copy-Move).

    To exploit parent-child dependencies across tree levels, predictions are conditioned hierarchically. For an input image XX, let θb(X)\theta_b(X) denote the logits at branch b∈{1,2,3,4}b \in \{1, 2, 3, 4\} and p(yb−1∣X)p(y_{b-1} \mid X) denote the predicted probabilities at level b−1b-1. The conditional probability distribution p(yb∣X)p(y_b \mid X) at level bb is computed as: p(yb∣X)=softmax(θb(X)⊙(1+p(yb−1∣X)))p(y_b \mid X) = \text{softmax}\left(\theta_b(X) \odot (1 + p(y_{b-1} \mid X))\right) where ⊙\odot represents broadcasting and element-wise multiplication by the probability of the corresponding parent category in level b−1b-1. For the root level (b=1b=1), p(y0∣X)=0p(y_0 \mid X) = 0.

  3. Knowl 3 — Hyperspherical Metric Learning Loss for Pixel-Level Forgery Localization

    equation

    To enhance spatial boundary separation between authentic and manipulated pixels, HiFi-Net employs a pixel-level hyperspherical metric learning localization loss. Let Fij′∈RDF'_{ij} \in \mathbb{R}^D denote the DD-dimensional feature embedding at pixel coordinate (i,j)(i, j) in the final mask prediction layer, Mij∈{0,1}M_{ij} \in \{0, 1\} be the binary ground truth manipulation indicator (Mij=0M_{ij}=0 for real, Mij=1M_{ij}=1 for forged), c∈RDc \in \mathbb{R}^D be the fixed reference center computed by averaging feature embeddings over all real pixels across the training dataset, and τ>0\tau > 0 be a predefined Euclidean margin.

    The localization loss Lloc\mathcal{L}_{\text{loc}} over an image of size H×WH \times W is defined as: Lloc=1HW∑i=1H∑j=1WL(Fij′,Mij;c,τ)\mathcal{L}_{\text{loc}} = \frac{1}{HW} \sum_{i=1}^H \sum_{j=1}^W \mathcal{L}\left(F'_{ij}, M_{ij}; c, \tau\right) where the per-pixel loss term is: L(Fij′,Mij;c,τ)={∥Fij′−c∥2if Mij=0 (real)max⁡(0,τ−∥Fij′−c∥2)if Mij=1 (forged)\mathcal{L}\left(F'_{ij}, M_{ij}; c, \tau\right) = \begin{cases} \|F'_{ij} - c\|_2 & \text{if } M_{ij} = 0 \text{ (real)} \\ \max\left(0, \tau - \|F'_{ij} - c\|_2\right) & \text{if } M_{ij} = 1 \text{ (forged)} \end{cases}

    The objective forces feature vectors of authentic pixels to collapse tightly toward center cc, while pushing embeddings of forged pixels outside the hyper-sphere radius τ\tau.

  4. Knowl 4 — Mask-Guided Feature Extraction via Partial Convolutions

    model/method

    HiFi-Net reuses the predicted binary localization mask M^∈RW0×H0\hat{M} \in \mathbb{R}^{W_0 \times H_0} to isolate the forged image content Xmask=X⊙M^X_{\text{mask}} = X \odot \hat{M}. To prevent zero-padded authentic regions from distorting the manipulation representations, feature extraction on XmaskX_{\text{mask}} is conducted using Partial Convolution (PConv) layers.

    Let WparW_{\text{par}} denote the convolution kernel weights. The partial convolution operation is defined as: Xmask′=WparT(X⊙M^)X'_{\text{mask}} = W_{\text{par}}^T (X \odot \hat{M}) where the kernel dot product is renormalized exclusively by the count of active mask pixels within its receptive field. Across successive layers, the mask M^\hat{M} is dynamically updated as: M^′={1if ∥M^∥1>00otherwise\hat{M}' = \begin{cases} 1 & \text{if } \|\hat{M}\|_1 > 0 \\ 0 & \text{otherwise} \end{cases}

    Two lightweight partial convolutional layers extract localization-conditioned priors (e.g., distinguishing sparse attribute modifications by STGAN from holistic facial swaps by FaceShifter) to inform downstream fine-grained attribute classification.

  5. Knowl 5 — HiFi-Net Joint Training Objective and Inference Protocol

    equation

    HiFi-Net is trained end-to-end using a joint objective function that integrates branch-specific cross-entropy classification losses Lclsb\mathcal{L}_{\text{cls}}^b for b∈{1,2,3,4}b \in \{1, 2, 3, 4\} and the metric localization loss Lloc\mathcal{L}_{\text{loc}}: Ltot={λLloc+Lcls1+Lcls2+Lcls3+Lcls4if X is forgedλLloc+Lcls4if X is real\mathcal{L}_{\text{tot}} = \begin{cases} \lambda \mathcal{L}_{\text{loc}} + \mathcal{L}_{\text{cls}}^1 + \mathcal{L}_{\text{cls}}^2 + \mathcal{L}_{\text{cls}}^3 + \mathcal{L}_{\text{cls}}^4 & \text{if } X \text{ is forged} \\ \lambda \mathcal{L}_{\text{loc}} + \mathcal{L}_{\text{cls}}^4 & \text{if } X \text{ is real} \end{cases} where Lclsb=−∑ky^b,klog⁡p(yb,k∣X)\mathcal{L}_{\text{cls}}^b = -\sum_k \hat{y}_{b,k} \log p(y_{b,k} \mid X) measures cross-entropy between the predicted conditional probabilities and ground-truth categorical label y^b\hat{y}_b, and λ\lambda is a balancing hyperparameter.

    During inference:

    1. Pixel-level localization is generated directly from the localization module output M^\hat{M}.
    2. Forgery attribute classification uses the level-4 probability vector p(y4∣X)p(y_4 \mid X).
    3. Binary image-level detection classifies XX as forged if the maximum predicted class probability at level 4 corresponds to any of the 13 manipulation techniques rather than the real category.
  6. Knowl 6 — HiFi-IFDL Hierarchical Benchmark Dataset

    experimental setup

    The Hierarchical Fine-grained Image Forgery Detection and Localization (HiFi-IFDL) dataset covers 13 distinct manipulation methods spanning CNN generation and conventional editing, paired with authentic images from 6 source datasets (FFHQ, AFHQ, CelebA-HQ, YouTube Face, MSCOCO, and LSUN):

    • Unconditional Diffusion: DDPM, DDIM (source: LSUN)
    • Conditional Diffusion: GDM, LDM (source: LSUN)
    • Unconditional GAN: StyleGANv2-ada, StyleGAN3 (source: FFHQ, AFHQ)
    • Conditional GAN: StarGANv2, HiSD (source: CelebA-HQ)
    • CNN-based Partial Manipulation: STGAN (source: CelebA-HQ), FaceShifter (source: YouTube Face)
    • Conventional Image Editing: Splicing, Inpainting, Copy-Move

    Each of the 13 forgery methods contains 100,000 images, and authentic datasets contribute either 100,000 images each or their entire volume. The overall dataset comprises:

    • Training set: 1,710,000 images
    • Validation set: 15,000 images
    • Test set: 174,000 images

    Every manipulated image is annotated with a high-resolution pixel-level binary ground truth mask indicating the precise manipulation boundary.

  7. Knowl 7 — Image Forgery Detection and Localization Performance on HiFi-IFDL

    data/table

    Detection and localization performance of HiFi-Net compared with baseline models on the HiFi-IFDL test benchmark. Metrics reported are Area Under the ROC Curve (AUC, %) and F1 score (%). Pre-trained models are marked with ∗*; unmarked models are trained on the HiFi-IFDL training set.

    Forgery Detection CNN-syn. Image Edit. Overall
    AUC F1 AUC F1 AUC F1
    CNN-detector∗^* 76.5 60.5 54.8 33.5 56.5 40.5
    CNN-detector 92.3 90.0 87.0 74.7 90.1 83.7
    Two-branch 93.3 89.2 83.3 66.7 86.7 80.2
    Att. Xception 93.8 91.2 90.8 82.1 87.3 90.0
    PSCC 94.6 93.2 90.7 82.3 93.2 91.3
    HiFi-Net (Ours) 97.0 96.1 91.5 85.9 96.8 94.1
    Forgery Localization CNN-syn. Image Edit. Overall
    AUC F1 AUC F1 AUC F1
    OSN-detector∗^* 51.4 38.8 83.2 70.1 79.4 56.5
    CatNet∗^* 48.6 31.9 86.1 79.4 78.3 65.1
    CatNet 92.5 81.5 92.0 88.2 92.4 86.8
    Att. Xception 89.1 87.7 83.3 79.3 87.1 86.5
    PSCC 94.3 96.8 91.1 86.5 92.7 94.9
    HiFi-Net (Ours) 98.4 97.0 93.0 90.1 95.3 96.9

    HiFi-Net outperforms prior methods across both detection (96.8% overall AUC, 94.1% overall F1) and localization (95.3% overall AUC, 96.9% overall F1). Models specialized strictly for editing or GANs (e.g., CatNet, CNN-detector) fail to generalize across both domains without joint multi-level hierarchical training.

  8. Knowl 8 — Cross-Domain Evaluation on Image Editing and Face Manipulation Benchmarks

    data/table

    HiFi-Net localization and detection performance evaluated on standard image editing benchmarks (Columbia, Coverage, CASIA, NIST16, IMD20) and the Diverse Fake Face Dataset (DFFD).

    Pre-trained Localization AUC (%) Columbia Coverage NIST16 CASIA IMD20 Average
    ManTra-Net 82.4 81.9 79.5 81.7 74.8 80.0
    SPAN 93.6 92.2 84.0 79.7 75.0 84.9
    PSCC 98.2 84.7 85.5 82.9 80.6 86.3
    ObjectFormer 95.5 92.8 87.2 84.3 82.1 88.3
    HiFi-Net (pre-trained as in PSCC) 98.3 93.2 87.0 85.8 82.9 89.4
    HiFi-Net (pre-trained on HiFi-IFDL) 98.4 92.4 86.9 86.6 83.4 89.6
    Fine-tuned Localization (AUC / F1 %) Coverage CASIA NIST16
    SPAN 93.7 / 55.8 83.8 / 40.8 96.1 / 58.2
    PSCC 94.1 / 72.3 87.5 / 55.4 99.6 / 81.9
    ObjectFormer 95.7 / 75.8 88.2 / 57.9 99.6 / 82.4
    HiFi-Net (Ours) 96.1 / 80.1 88.5 / 61.6 98.9 / 85.0

    On the CASIA image-level detection benchmark, HiFi-Net obtains 99.5% AUC and 97.4% F1 (compared to ObjectFormer's 99.7% AUC and 97.3% F1). On DFFD face manipulations, HiFi-Net attains an IoU of 0.411 and Pixel-wise Binary Classification Accuracy (PBCA) of 0.801 for localization (surpassing Attention Xception's 0.401 / 0.786), while achieving 99.45% AUC and 88.50% PBCA on image detection.

  9. Knowl 9 — Ablation Analysis of HiFi-Net Components

    data/table

    Ablation study evaluating the individual contributions of the multi-branch classification module (MM), localization module (LL), partial convolution (PP), metric localization loss (Lloc\mathcal{L}_{\text{loc}}), multi-level classification losses (Lcls\mathcal{L}_{\text{cls}}), and hierarchical path prediction on the HiFi-IFDL benchmark.

    Configuration Modules Loss Function Detection Localization
    AUC (%) F1 (%) AUC (%) F1 (%)
    Full HiFi-Net M,L,PM, L, P Lcls,Lloc\mathcal{L}_{\text{cls}}, \mathcal{L}_{\text{loc}} 96.8 94.1 95.3 96.9
    No Lcls\mathcal{L}_{\text{cls}} M,L,PM, L, P Lloc\mathcal{L}_{\text{loc}} 65.0 70.0 93.4 95.0
    No Lloc\mathcal{L}_{\text{loc}} M,L,PM, L, P Lcls\mathcal{L}_{\text{cls}} 95.8 92.4 66.0 58.0
    Level-4 Classification Only M,L,PM, L, P Lcls4,Lloc\mathcal{L}_{\text{cls}}^4, \mathcal{L}_{\text{loc}} 93.1 91.7 92.5 93.9
    No Path Conditioning M,L,PM, L, P Lclsind,Lloc\mathcal{L}_{\text{cls}}^{\text{ind}}, \mathcal{L}_{\text{loc}} 93.2 92.8 93.2 94.8
    No Partial Convolution M,LM, L Lcls,Lloc\mathcal{L}_{\text{cls}}, \mathcal{L}_{\text{loc}} 96.6 93.0 94.8 96.0

    Key takeaways:

    1. Removing Lcls\mathcal{L}_{\text{cls}} degrades detection by 24.1% F1 and localization by 1.9% AUC, indicating that fine-grained classification features provide strong priors for localization.
    2. Omitting intermediate hierarchical levels (Level-4 only) drops detection AUC by 3.7% and localization AUC by 2.8%.
    3. Disabling hierarchical path conditioning (independent prediction Lclsind\mathcal{L}_{\text{cls}}^{\text{ind}}) reduces detection AUC by 3.6%.
  10. Knowl 10 — Fine-Grained Forgery Attribute Classification and Forensic Failure Modes

    empirical result

    HiFi-Net achieves an F1 score of 87.63% in 14-way fine-grained forgery attribute classification at level 4 on the HiFi-IFDL benchmark. It substantially outperforms flat generative attribution baselines (FEN at 48.00%, Attr at 31.11%), as well as isolated single-level training (84.62%) and unconditioned multi-scale training (85.37%).

    Empirical error analysis reveals four main failure modes:

    1. Authentic Image Distortions: Real images featuring prominent watermarks, compression artifacts, or extreme lighting conditions are prone to false-positive classification as forged.
    2. Small Inpainting Regions: Tampered images with tiny inpainting areas provide insufficient spatial and spectral cues, resulting in misclassification as authentic.
    3. Intra-Architecture Generator Ambiguity: Generated images from closely related GAN architectures (e.g., StyleGANv2-ada vs. StyleGAN3) produce highly similar frequency patterns and are frequently confused with one another.
    4. Editing-to-Diffusion Generalization Gap: Models trained strictly on conventional image editing techniques generalize poorly to diffusion-based inpainting manipulations.

Coverage note — None was omitted; all contributed aspects of the problem formulation, network architecture, loss formulations, dataset construction, benchmark evaluations, ablation experiments, and limitations were incorporated.

References

  1. 1.Survey: More americans get news from internet than news-papers or radio. http://www.cnn.com/2010/TECH/03/01/social.network.news/index.html.
  2. 2.Nist: Nist nimble 2016 datasets., May 2016.
  3. 3.Infodemic - world Health Organization. https://www.who.int/health-topics/infodemic, 2022.
  4. 4.Wael AbdAlmageed, Hengameh Mirzaalian, Xiao Guo, Linda M Randolph, Veeraya K Tanawattanacharoen, Mitchell E Geffner, Heather M Ross, and Mimi S Kim. As-sessment of facial morphologic features in patients with con-genital adrenal hyperplasia using deep learning. JAMA network open, 2020.
  5. 5.Vishal Asnani, Xi Yin, Tal Hassner, Sijia Liu, and Xiaom-ing Liu. Proactive image manipulation detection. In CVPR, 2022.
  6. 6.Vishal Asnani, Xi Yin, Tal Hassner, and Xiaoming Liu. Re-verse engineering of generative models: Inferring model hyperparameters from generated images. arXiv preprint arXiv:2106.07873, 2021.
  7. 7.Vishal Asnani, Xi Yin, Tal Hassner, and Xiaoming Liu. Malp: Manipulation localization using a proactive scheme. In CVPR, 2023.
  8. 8.Tu Bui, Ning Yu, and John Collomosse. Repmix: Represen-tation mixing for robust attribution of synthesized images. In ECCV, 2022.
  9. 9.Peter J Burt and Edward H Adelson. The laplacian pyramid as a compact image code. In Readings in computer vision. Elsevier, 1987.
  10. 10.Lucy Chai, David Bau, Ser-Nam Lim, and Phillip Isola. What makes fake images detectable? understanding prop-erties that generalize. In ECCV, 2020.
  11. 11.Xinru Chen, Chengbo Dong, Jiaqi Ji, Juan Cao, and Xirong Li. Image manipulation detection by multi-view multi-scale supervision. In ICCV, 2021.
  12. 12.Yunjey Choi, Youngjung Uh, Jaejun Yoo, and Jung-Woo Ha. Stargan v2: Diverse image synthesis for multiple domains. In CVPR, 2020.
  13. 13.Riccardo Corvi, Davide Cozzolino, Giada Zingarini, Gio-vanni Poggi, Koki Nagano, and Luisa Verdoliva. On the de-tection of synthetic images generated by diffusion models. arXiv preprint arXiv:2211.00680, 2022.
  14. 14.Davide Cozzolino, Justus Thies, Andreas Rossler, Christian Riess, Matthias Nießner, and Luisa Verdoliva. Forensictrans-fer: Weakly-supervised domain adaptation for forgery detec-tion. arXiv preprint arXiv:1812.02510, 2018.
  15. 15.Debayan Deb, Xiaoming Liu, and Anil Jain. Unified detec-tion of digital and physical face attacks. In FG, 2023.
  16. 16.Brian Dolhansky, Russ Howes, Ben Pflaum, Nicole Baram, and Cristian Canton Ferrer. The Deepfake Detection Chal-lenge (DFDC) Preview Dataset. arXiv preprint arXiv: 1910.08854, 2019.
  17. 17.Chengbo Dong, Xinru Chen, Ruohan Hu, Juan Cao, and Xirong Li. Mvss-net: Multi-view multi-scale supervised net-works for image manipulation detection. TPAMI, 2022.
  18. 18.Jing Dong, Wei Wang, and Tieniu Tan. Casia image tamper-ing detection evaluation database. 2013 IEEE China Summit and ICSIP, 2013.
  19. 19.Nicholas Dufour, Andrew Gully, Per Karlsson, Alexey Vic-tor Vorbyov, Thomas Leung, Jeremiah Childs, and Christoph Bregler. Deepfakes detection dataset by Google & Jigsaw, 2019.
  20. 20.Xiao Guo and Jongmoo Choi. Human motion prediction via learning local structure representations and temporal depen-dencies. In AAAI, 2019.
  21. 21.Xiao Guo, Yaojie Liu, Anil Jain, and Xiaoming Liu. Multi-domain learning for updating face anti-spoofing models. In ECCV, 2022.
  22. 22.Xiao Guo, Hengameh Mirzaalian, Ekraam Sabir, Ayush Jaiswal, and Wael Abd-Almageed. Cord19sts: Covid-19 semantic textual similarity dataset. arXiv preprint arXiv:2007.02461, 2020.
  23. 23.Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P Kingma, Ben Poole, Mohammad Norouzi, David J Fleet, et al. Imagen video: High definition video generation with diffusion mod-els. arXiv preprint arXiv:2210.02303, 2022.
  24. 24.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffu-sion probabilistic models. In NeurIPS, 2020.
  25. 25.I Hsu, Xiao Guo, Premkumar Natarajan, Nanyun Peng, et al. Discourse-level relation extraction via graph pooling. In AAAI DLG Wrokshop, 2021.
  26. 26.Xuefeng Hu, Zhihan Zhang, Zhenye Jiang, Syomantak Chaudhuri, Zhenheng Yang, and Ram Nevatia. Span: spatial pyramid attention network for image manipulation localiza-tion. In ECCV, 2020.
  27. 27.Yihao Huang, Felix Juefei-Xu, Qing Guo, Yang Liu, and Geguang Pu. Fakelocator: Robust localization of gan-based face manipulations. TIFS, 2022.
  28. 28.Liming Jiang, Ren Li, Wayne Wu, Chen Qian, and Chen Change Loy. Deeperforensics-1.0: A large-scale dataset for real-world face forgery detection. In CVPR, 2020.
  29. 29.Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen. Progressive growing of gans for improved quality, stability, and variation. In ICLR, 2018.
  30. 30.Tero Karras, Miika Aittala, Janne Hellsten, Samuli Laine, Jaakko Lehtinen, and Timo Aila. Training generative adver-sarial networks with limited data. In NeurIPS, 2020.
  31. 31.Tero Karras, Miika Aittala, Samuli Laine, Erik Hark ¨ onen, ¨ Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Alias-free generative adversarial networks. In NeurIPS, 2021.
  32. 32.Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. In CVPR, 2019.
  33. 33.Diederik P Kingma and Max Welling. Auto-encoding varia-tional bayes. In ICLR, 2014.
  34. 34.Pavel Korshunov and Sebastien Marcel. Vulnerability as-sessment and detection of deepfake videos. In ICB, 2019.
  35. 35.Myung-Joon Kwon, Seung-Hun Nam, In-Jae Yu, Heung-Kyu Lee, and Changick Kim. Learning jpeg compression artifacts for image manipulation detection and localization. IJCV, 2022.
  36. 36.Cheng-Han Lee, Ziwei Liu, Lingyun Wu, and Ping Luo. Maskgan: Towards diverse and interactive facial image ma-nipulation. In CVPR, 2020.
  37. 37.Lingzhi Li, Jianmin Bao, Hao Yang, Dong Chen, and Fang Wen. Faceshifter: Towards high fidelity and occlusion aware face swapping. In CVPR, 2020.
  38. 38.Lingzhi Li, Jianmin Bao, Ting Zhang, Hao Yang, Dong Chen, Fang Wen, and Baining Guo. Face x-ray for more general face forgery detection. In CVPR, 2020.
  39. 39.Xinyang Li, Shengchuan Zhang, Jie Hu, Liujuan Cao, Xi-aopeng Hong, Xudong Mao, Feiyue Huang, Yongjian Wu, and Rongrong Ji. Image-to-image translation via hierarchi-cal style disentanglement. In CVPR, 2022.
  40. 40.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence ´ Zitnick. Microsoft coco: Common objects in context. In ECCV, 2014.
  41. 41.Guilin Liu, Fitsum A Reda, Kevin J Shih, Ting-Chun Wang, Andrew Tao, and Bryan Catanzaro. Image inpainting for ir-regular holes using partial convolutions. In ECCV, 2018.
  42. 42.Ming Liu, Yukang Ding, Min Xia, Xiao Liu, Errui Ding, Wangmeng Zuo, and Shilei Wen. Stgan: A unified selec-tive transfer network for arbitrary image attribute editing. In CVPR, 2019.
  43. 43.Xiaohong Liu, Yaojie Liu, Jun Chen, and Xiaoming Liu. Pscc-net: Progressive spatio-channel correlation network for image manipulation detection and localization. TCSVT, 2022.
  44. 44.Francesco Marra, Diego Gragnaniello, Davide Cozzolino, and Luisa Verdoliva. Detection of gan-generated fake im-ages over social networks. In MIPR, 2018.
  45. 45.Francesco Marra, Diego Gragnaniello, Luisa Verdoliva, and Giovanni Poggi. Do gans leave artificial fingerprints? In MIPR, 2019.
  46. 46.Iacopo Masi, Aditya Killekar, Royston Marian Mascaren-has, Shenoy Pratik Gurudatt, and Wael AbdAlmageed. Two-branch recurrent network for isolating deepfakes in videos. In ECCV, 2020.
  47. 47.Owen Mayer and Matthew C Stamm. Learned forensic source similarity for unknown camera models. In ICASSP, 2018.
  48. 48.Tian-Tsong Ng, Jessie Hsu, and Shih-Fu Chang. Columbia image splicing detection evaluation dataset. DVMM lab. Columbia Univ CalPhotos Digit Libr, 2009.
  49. 49.Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. Glide: Towards photorealistic image genera-tion and editing with text-guided diffusion models. In ICML, 2021.
  50. 50.Adam Novozamsky, Babak Mahdian, and Stanislav Saic. Imd2020: A large-scale annotated dataset tailored for detect-ing manipulated images. In WACV Workshop, 2020.
  51. 51.Patrick Perez, Michel Gangnet, and Andrew Blake. Poisson ´ image editing. ACM SIGGRAPH, 2003.
  52. 52.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image gen-eration with clip latents. arXiv preprint arXiv:2204.06125, 2022.
  53. 53.Jonas Ricker, Simon Damm, Thorsten Holz, and Asja Fis-cher. Towards the detection of diffusion model deepfakes. arXiv preprint arXiv:2210.14571, 2022.
  54. 54.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. High-resolution image syn- ¨ thesis with latent diffusion models. In CVPR, 2022.
  55. 55.Andreas Rossler, Davide Cozzolino, Luisa Verdoliva, Chris- ¨ tian Riess, Justus Thies, and Matthias Nießner. Faceforen-sics++: Learning to detect manipulated facial images. In ICCV, 2019.
  56. 56.Lukas Ruff, Robert Vandermeulen, Nico Goernitz, Lucas Deecke, Shoaib Ahmed Siddiqui, Alexander Binder, Em-manuel Muller, and Marius Kloft. Deep one-class classifi- ¨ cation. In ICML, 2018.
  57. 57.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S Sara Mahdavi, Rapha Gontijo Lopes, et al. Photorealistic text-to-image diffusion models with deep language understanding. arXiv preprint arXiv:2205.11487, 2022.
  58. 58.Ulrich Scherhag, Christian Rathgeb, Johannes Merkle, Ralph Breithaupt, and Christoph Busch. Face recognition systems under morphing attacks: A survey. IEEE Access, 2019.
  59. 59.Husrev Taha Sencar, Luisa Verdoliva, and Nasir Memon. Multimedia Forensics. Springer Nature, 2022.
  60. 60.Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, et al. Make-a-video: Text-to-video generation without text-video data. arXiv preprint arXiv:2209.14792, 2022.
  61. 61.Jiaming Song, Chenlin Meng, and Stefano Ermon. Denois-ing diffusion implicit models. 2021.
  62. 62.Joel Stehouwer, Hao Dang, Feng Liu, Xiaoming Liu, and Anil Jain. On the detection of digital face manipulation. In CVPR, 2020.
  63. 63.Junke Wang, Zuxuan Wu, Jingjing Chen, Xintong Han, Ab-hinav Shrivastava, Ser-Nam Lim, and Yu-Gang Jiang. Ob-jectformer for image manipulation detection and localiza-tion. In CVPR, 2022.
  64. 64.Sheng-Yu Wang, Oliver Wang, Richard Zhang, Andrew Owens, and Alexei A Efros. Cnn-generated images are sur-prisingly easy to spot... for now. In CVPR, 2020.
  65. 65.Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaim-ing He. Non-local neural networks. In CVPR, 2018.
  66. 66.Bihan Wen, Ye Zhu, Ramanathan Subramanian, Tian-Tsong Ng, Xuanjing Shen, and Stefan Winkler. Coverage—a novel database for copy-move forgery detection. In ICIP, 2016.
  67. 67.Yue Wu, Wael Abd-Almageed, and Prem Natarajan. Buster-net: Detecting copy-move image forgery with source/target localization. In ECCV, 2018.
  68. 68.Yue Wu, Wael AbdAlmageed, and Premkumar Natarajan. Mantra-net: Manipulation tracing network for detection and localization of image forgeries with anomalous features. In CVPR, 2019.
  69. 69.Haiwei Wu et al. Robust image forgery detection over online social network shared images. In CVPR, 2022.
  70. 70.Fisher Yu, Ari Seff, Yinda Zhang, Shuran Song, Thomas Funkhouser, and Jianxiong Xiao. Lsun: Construction of a large-scale image dataset using deep learning with humans in the loop. arXiv preprint arXiv:1506.03365, 2015.
  71. 71.Ning Yu, Larry S Davis, and Mario Fritz. Attributing fake images to gans: Learning and analyzing gan fingerprints. In ICCV, 2019.
  72. 72.Han Zhang, Ian Goodfellow, Dimitris Metaxas, and Augus-tus Odena. Self-attention generative adversarial networks. In ICML, 2019.
  73. 73.Xu Zhang, Svebor Karaman, and Shih-Fu Chang. Detecting and simulating artifacts in gan fake images. WIFS, 2019.
  74. 74.Tianchen Zhao, Xiang Xu, Mingze Xu, Hui Ding, Yuanjun Xiong, and Wei Xia. Learning self-consistency for deepfake detection. In CVPR, 2021.
  75. 75.Peng Zhou, Bor-Chun Chen, Xintong Han, Mahyar Najibi, Abhinav Shrivastava, Ser-Nam Lim, and Larry Davis. Gen-erate, segment, and refine: Towards generic manipulation segmentation. In AAAI, 2020.
  76. 76.Peng Zhou, Xintong Han, Vlad I Morariu, and Larry S Davis. Learning rich features for image manipulation detection. In CVPR, 2018.

Citation

MLA
Guo, X., et al. “Hierarchical Fine-Grained Image Forgery Detection and Localization”. arXiv, 2023, http://arxiv.org/abs/2303.17111v1.
APA
Guo, X., Liu, X., Ren, Z., Grosz, S., Masi, I., & Liu, X. (2023). Hierarchical Fine-Grained Image Forgery Detection and Localization. arXiv. http://arxiv.org/abs/2303.17111v1
Chicago
Guo, X., X. Liu, Z. Ren, S. Grosz, I. Masi, and X. Liu. 2023. “Hierarchical Fine-Grained Image Forgery Detection and Localization”. arXiv. http://arxiv.org/abs/2303.17111v1.
Harvard
Guo, X. et al. (2023) “Hierarchical Fine-Grained Image Forgery Detection and Localization”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2303.17111v1.
Vancouver
1. Guo X, Liu X, Ren Z, Grosz S, Masi I, Liu X (2023) Hierarchical Fine-Grained Image Forgery Detection and Localization. arXiv

BibTeX

@article{guo2023hierarchical,
  title = {Hierarchical Fine-Grained Image Forgery Detection and Localization},
  author = {Guo, Xiao and Liu, Xiaohong and Ren, Zhiyuan and Grosz, Steven and Masi, Iacopo and Liu, Xiaoming},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2303.17111v1},
  eprint = {2303.17111}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE