AVFF: Audio-Visual Feature Fusion for Video Deepfake Detection
Trevine OorloffSurya KoppisettiNicolò BonettiniDivyaraj SolankiBen ColmanYaser YacoobAli ShahriyariGaurav Bharaj
Proposes a two-stage deepfake detection framework that learns intrinsic cross-modal correspondences from real videos via complementary masking and feature fusion, achieving 98.6% accuracy on the FakeAVCeleb benchmark.
The rapid advancement of generative artificial intelligence has made creating deceptive deepfake videos increasingly accessible, posing severe risks including fraud, defamation, and disinformation. Many current automated detection systems rely either on visual-only clues or on single-stage supervised training. These traditional systems frequently overlook the subtle, natural synchronization between speech audio and facial movements—such as mouth articulation and emotional expression—which is difficult for generative models to recreate accurately and is essential for detecting novel, unseen deepfakes.
The article demonstrates a novel deepfake detection framework named Audio-Visual Feature Fusion (AVFF). The primary objective is to capture the natural correspondence between audio and visual streams to significantly improve the detection of synthetic videos where either or both modalities have been manipulated.
To achieve this, the article establishes a two-stage approach. In the first stage, the system undergoes self-supervised pre-training exclusively on real videos to learn the intrinsic relationship between human speech and facial dynamics. The model divides audio and video into temporal slices and applies complementary masking, obscuring 50% of the data in each modality such that masked sections in one stream correspond to visible sections in the other. It then uses cross-modal networks to predict the hidden segments across modalities and reconstruct the video. In the second stage, the model uses these rich representations to train a classifier on labeled real and fake videos to detect multi-modal dissonance.
The findings confirm that this cross-modal representation learning substantially outperforms existing methods. On the benchmark FakeAVCeleb evaluation, the proposed method achieves 98.6% accuracy and a 99.1% Area Under the ROC Curve (AUC)—a standard metric measuring classification capability where 100% indicates perfect discrimination. This represents an improvement of 14.9 percentage points in accuracy and 9.9 percentage points in AUC over current top-performing audio-visual systems, as well as an 8.7 percentage point accuracy improvement over top visual-only systems. Furthermore, the model maintains high resilience when tested against previously unseen deepfake generation tools (achieving over 92% AUC across diverse synthetic categories) and adapts successfully when evaluated on entirely separate datasets.
These results demonstrate that learning the baseline rules of authentic human speech and facial behavior provides a far stronger defensive capability than merely searching for known manipulation artifacts. For organizations managing digital media integrity, security operations, and compliance, deploying multi-modal verification models can substantially reduce the operational and reputational risks associated with sophisticated generative AI scams.
Organizations evaluating or deploying deepfake detection defenses should prioritize solutions that incorporate cross-modal audio-visual verification rather than unimodal visual scanners. Future implementation roadmaps should focus on extending this framework to handle non-humanoid video synthesis and refining cross-modal representations for other tasks, such as emotion recognition.
Decision-makers should note that the system requires inputs containing both audio and visual streams with a single coherent speaker. Performance may degrade in scenarios with asynchronous audio lag, background voice overlaps, multi-speaker dialogue, or severe facial occlusions such as masks or hands. Nevertheless, within standard single-speaker portrait video conditions, the article provides high confidence in the framework's superior accuracy and generalization capabilities.
- Paper: Contrastive Audio-Visual Masked Autoencoder, Yuan Gong et al. (2023). Provides the foundational self-supervised paradigm uniting contrastive matching and masked autoencoding across audio and visual modalities that AVFF adapts for deepfake detection.
- Paper: Multimodal Deep Learning, Jiquan Ngiam et al. (2011). Establishes seminal principles for cross-modal feature learning and autoencoder-based audio-visual reconstruction underlying cross-stream deep learning representations.
- Paper: FaceForensics++: Learning to Detect Manipulated Facial Images, Andreas Rössler et al. (2019). Introduces foundational facial manipulation benchmarks and baseline forensic evaluation protocols for video forgery detection.
- Paper: Celeb-DF: A Large-Scale Challenging Dataset for DeepFake Forensics, Yuezun Li et al. (2019). Presents a challenging large-scale video deepfake dataset and standardized cross-manipulation benchmarking used to evaluate generalized video forensics.
- Paper: DeepfakeBench: A Comprehensive Benchmark of Deepfake Detection, Zhiyuan Yan et al. (2023). Provides a comprehensive unified framework and standard metrics for evaluating generalizable video deepfake detection methods.
- Paper: Frequency-Aware Deepfake Detection: Improving Generalizability through Frequency Space Domain Learning, Chuangchuang Tan et al. (2024). Extends generalizable deepfake detection to the frequency domain to identify generative synthesis artifacts across unseen architectures.
