Hierarchical Fine-Grained Image Forgery Detection and Localization
Xiao GuoXiaohong LiuZhiyuan RenSteven GroszIacopo MasiXiaoming Liu
Proposes a hierarchical multi-branch framework and a structured dataset that unify image forgery detection and pixel-level localization across both deep generative synthesis and conventional image editing by explicitly modeling multi-level manipulation attributes.
Rapid advancements in artificial intelligence and accessible photo-editing software have made generating convincing synthetic imagery and manipulating visual content easier than ever. This widespread proliferation of manipulated media fuels misinformation campaigns, posing severe reputational, legal, and operational risks across modern digital ecosystems. While prior forensic tools generally specialized in either identifying conventional image editing (such as splicing and inpainting) or detecting entirely synthesized images, real-world deployment requires a unified framework capable of handling both manipulation types simultaneously.
The article demonstrates a unified framework designed to perform simultaneous image forgery detection, pixel-level manipulation localization, and fine-grained classification of the specific generation method used. To achieve this, the authors introduce a hierarchical multi-branch architecture alongside a comprehensive forensic dataset spanning modern diffusion models, generative adversarial networks, and standard editing techniques.
To bridge the gap across manipulation types, the authors formulated a four-level hierarchical structure that classifies manipulated imagery from broad categories down to specific source tools. They built a benchmark containing over 1.8 million images spanning thirteen distinct manipulation methods paired with high-resolution pixel masks. The proposed network processes both color and frequency domains using multiple specialized branches corresponding to different hierarchical levels, a self-attention localization module optimized via metric learning, and masked partial convolutions that feed localized spatial cues directly into attribute classification.
Across extensive testing on seven distinct benchmarks, the unified framework consistently matched or outperformed existing state-of-the-art approaches. On the newly introduced dataset, it achieved an overall image detection performance of 96.8% Area Under the Curve (AUC) and a localization performance of 95.3% AUC, outperforming specialized baselines by 2.6% to 9.3%. In fine-grained attribute classification, the model attained an overall score of 87.6%, compared to under 50% for previous generative model attribution techniques. Ablation experiments confirmed that removing the hierarchical path prediction decreased detection accuracy by 3.6% AUC, demonstrating the necessity of structured dependency learning.
These findings indicate that integrating pixel-level localization with hierarchical attribute attribution significantly bolsters general forensic accuracy. Rather than treating manipulation types as mutually exclusive, leveraging coarse-to-fine relational dependencies enables organizations to better track manipulation lineage and assess media tampering risks in open environments. Practically, this approach provides actionable provenance data essential for compliance, legal verification, and automated content moderation.
Organizations seeking to implement visual forensics should transition away from single-domain detectors toward multi-level unified models capable of simultaneous detection and localization. Further engineering work should focus on scaling forensic datasets and evaluating defenses against rapid developments in diffusion-based inpainting. While the framework demonstrates strong robustness, confidence is moderated by observed classification challenges in edge cases, including small inpainting patches, images with extreme lighting, and closely related generative model architectures.
- Paper: CNN-Generated Images Are Surprisingly Easy to Spot… for Now, Sheng-Yu Wang et al. (2019). It establishes the benchmark framework for analyzing synthetic image forgery across diverse CNN generators, providing essential context for unified image forgery representation learning.
- Paper: FaceForensics++: Learning to Detect Manipulated Facial Images, Andreas Rössler et al. (2019). It provides foundational facial manipulation benchmarks and standard forgery detection methodologies that underpin multi-domain forgery analysis.
- Paper: Celeb-DF: A Large-Scale Challenging Dataset for DeepFake Forensics, Yuezun Li et al. (2019). It introduces a challenging standard dataset and evaluation framework for high-quality synthetic manipulation detection that contextualizes generalized forgery detection.
- Paper: MesoNet: a Compact Facial Video Forgery Detection Network, Darius Afchar et al. (2018). It establishes intermediate feature-level representations for detecting compressed facial forgeries, serving as an early baseline for fine-grained forensic analysis.
- Paper: Deep Layer Aggregation, F. Yu et al. (2017). It introduces hierarchical deep layer aggregation techniques that directly inform multi-branch, multi-level feature extraction in visual tasks.
- Paper: Rethinking the Up-Sampling Operations in CNN-Based Generative Network for Generalizable Deepfake Detection, Chuangchuang Tan et al. (2024). It advances generalizable forgery detection by isolating specific spatial upsampling artifacts across diverse generative models, building beyond hierarchical attribute classification.
- Paper: Frequency-Aware Deepfake Detection: Improving Generalizability through Frequency Space Domain Learning, Chuangchuang Tan et al. (2024). It extends deepfake detection generalization across unseen generators by focusing on frequency-domain representation learning.
