AVFF: Audio-Visual Feature Fusion for Video Deepfake Detection

Trevine OorloffSurya KoppisettiNicolò BonettiniDivyaraj SolankiBen ColmanYaser YacoobAli ShahriyariGaurav Bharaj

article2024CVPR134 citations

Proposes a two-stage deepfake detection framework that learns intrinsic cross-modal correspondences from real videos via complementary masking and feature fusion, achieving 98.6% accuracy on the FakeAVCeleb benchmark.

Listen

The rapid advancement of generative artificial intelligence has made creating deceptive deepfake videos increasingly accessible, posing severe risks including fraud, defamation, and disinformation. Many current automated detection systems rely either on visual-only clues or on single-stage supervised training. These traditional systems frequently overlook the subtle, natural synchronization between speech audio and facial movements—such as mouth articulation and emotional expression—which is difficult for generative models to recreate accurately and is essential for detecting novel, unseen deepfakes.

The article demonstrates a novel deepfake detection framework named Audio-Visual Feature Fusion (AVFF). The primary objective is to capture the natural correspondence between audio and visual streams to significantly improve the detection of synthetic videos where either or both modalities have been manipulated.

To achieve this, the article establishes a two-stage approach. In the first stage, the system undergoes self-supervised pre-training exclusively on real videos to learn the intrinsic relationship between human speech and facial dynamics. The model divides audio and video into temporal slices and applies complementary masking, obscuring 50% of the data in each modality such that masked sections in one stream correspond to visible sections in the other. It then uses cross-modal networks to predict the hidden segments across modalities and reconstruct the video. In the second stage, the model uses these rich representations to train a classifier on labeled real and fake videos to detect multi-modal dissonance.

The findings confirm that this cross-modal representation learning substantially outperforms existing methods. On the benchmark FakeAVCeleb evaluation, the proposed method achieves 98.6% accuracy and a 99.1% Area Under the ROC Curve (AUC)—a standard metric measuring classification capability where 100% indicates perfect discrimination. This represents an improvement of 14.9 percentage points in accuracy and 9.9 percentage points in AUC over current top-performing audio-visual systems, as well as an 8.7 percentage point accuracy improvement over top visual-only systems. Furthermore, the model maintains high resilience when tested against previously unseen deepfake generation tools (achieving over 92% AUC across diverse synthetic categories) and adapts successfully when evaluated on entirely separate datasets.

These results demonstrate that learning the baseline rules of authentic human speech and facial behavior provides a far stronger defensive capability than merely searching for known manipulation artifacts. For organizations managing digital media integrity, security operations, and compliance, deploying multi-modal verification models can substantially reduce the operational and reputational risks associated with sophisticated generative AI scams.

Organizations evaluating or deploying deepfake detection defenses should prioritize solutions that incorporate cross-modal audio-visual verification rather than unimodal visual scanners. Future implementation roadmaps should focus on extending this framework to handle non-humanoid video synthesis and refining cross-modal representations for other tasks, such as emotion recognition.

Decision-makers should note that the system requires inputs containing both audio and visual streams with a single coherent speaker. Performance may degrade in scenarios with asynchronous audio lag, background voice overlaps, multi-speaker dialogue, or severe facial occlusions such as masks or hands. Nevertheless, within standard single-speaker portrait video conditions, the article provides high confidence in the framework's superior accuracy and generalization capabilities.

arXiv: 2406.02951
Cover for AVFF: Audio-Visual Feature Fusion for Video Deepfake Detection

Abstract

With the rapid growth in deepfake video content, we require improved and generalizable methods to detect them. Most existing detection methods either use uni-modal cues or rely on supervised training to capture the dissonance between the audio and visual modalities. While the former disregards the audio-visual correspondences entirely, the latter predominantly focuses on discerning audio-visual cues within the training corpus, thereby potentially overlooking correspondences that can help detect unseen deepfakes. We present Audio-Visual Feature Fusion (AVFF), a two-stage cross-modal learning method that explicitly captures the correspondence between the audio and visual modalities for improved deepfake detection. The first stage pursues representation learning via self-supervision on real videos to capture the intrinsic audio-visual correspondences. To extract rich cross-modal representations, we use contrastive learning and autoencoding objectives, and introduce a novel audio-visual complementary masking and feature fusion strategy. The learned representations are tuned in the second stage, where deepfake classification is pursued via supervised learning on both real and fake videos. Extensive experiments and analysis suggest that our novel representation learning paradigm is highly discriminative in nature. We report 98.6% accuracy and 99.1% AUC on the FakeAVCeleb dataset, outperforming the current audio-visual state-of-the-art by 14.9% and 9.9%, respectively.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 2.1. Multi-Modal Representation Learning
  • 2.2. Deepfake Detection
  • 3. Method
  • 3.1. Preprocessing
  • 3.2. Representation Learning Stage
  • 3.3. Deepfake Classification Stage
  • 4. Experiments and Results
  • 4.1. Implementation
  • 4.2. Evaluation and Discussion
  • 5. Ablation Study
  • 6. Conclusion, Limitations, and Future Work
  • References

Knowls

  1. Knowl 1 — AVFF two-stage audio-visual deepfake detector

    model/method

    Audio-Visual Feature Fusion (AVFF) detects videos in which either the audio, the visual content, or both modalities have been generated or manipulated. It uses two sequential stages. First, a self-supervised representation-learning stage is trained only on real talking-face videos so that audio and facial-video embeddings capture the temporal correspondence between speech and facial motion. This stage combines audio-visual contrastive learning with masked autoencoding, using complementary masking and cross-modal token prediction rather than independent modality reconstruction. Second, a supervised classifier uses the learned unimodal and cross-modal representations on labeled real and fake videos. The intended detection signal is the reduced audio-visual cohesion of synthesized samples relative to real videos.

  2. Knowl 2 — Audio-visual preprocessing and temporally aligned tokenization

    model/method

    Each video is preprocessed by sampling visual frames at 5 frames per second and audio at 16 kHz. Face regions are cropped and backgrounds are removed with FaceX-Zoo, while the waveform is converted to a log-mel spectrogram with LL frequency bins. For a video of duration TT, the audio representation is xa∈RTa×Lx_a\in\mathbb{R}^{T_a\times L}, where TaT_a is the number of audio frames, and the visual representation is xv∈RTv×C×H×Wx_v\in\mathbb{R}^{T_v\times C\times H\times W}, where TvT_v is the number of visual frames, CC is the number of image channels, and HH and WW are frame height and width. Audio is tokenized into non-overlapping 16×1616\times16 two-dimensional patches; visual data is tokenized into non-overlapping 2×16×162\times16\times16 three-dimensional spatio-temporal patches. The resulting audio and visual token sequences are divided into K=8K=8 equal temporal slices, with audio slice ii and visual slice ii covering the same time interval. Separate transformer encoders EaE_a and EvE_v produce audio features aa and visual features vv with learnable positional embeddings.

  3. Knowl 3 — Complementary masking and cross-modal feature fusion

    model/method

    AVFF masks half of the eight temporal feature slices in each modality, with complementary binary masks Ma,Mv∈{0,1}8M_a,M_v\in\{0,1\}^{8} satisfying Ma+Mv=1M_a+M_v=\mathbf{1}; therefore, whenever an audio slice is masked, the visual slice at the same time is visible, and vice versa. For modality p∈{a,v}p\in\{a,v\}, the visible and masked features are pvis=Mp⊙pp_{\mathrm{vis}}=M_p\odot p and pmsk=(1−Mp)⊙pp_{\mathrm{msk}}=(\mathbf{1}-M_p)\odot p, where ⊙\odot is elementwise multiplication. A learnable audio-to-visual network converts visible audio slices into predicted visual slices, v~=A2V(avis)\tilde v=A2V(a_{\mathrm{vis}}), and a visual-to-audio network converts visible visual slices into predicted audio slices, a~=V2A(vvis)\tilde a=V2A(v_{\mathrm{vis}}). Each conversion network contains a single-layer MLP for token-count matching followed by one transformer block. The masked audio slices in aa are replaced by the corresponding temporal slices of a~\tilde a, and the masked visual slices in vv are replaced by the corresponding slices of v~\tilde v, producing fused embeddings a′a' and v′v'. Thus, masked-token reconstruction cannot rely only on a modality-specific learned mask token: it must use information from the temporally corresponding other modality.

  4. Knowl 4 — Dual contrastive, masked-reconstruction, and adversarial training objective

    equation

    The representation-learning stage optimizes a contrastive objective together with masked reconstruction and Wasserstein adversarial objectives. Let a minibatch contain NN paired audio-visual samples, let p,q∈{a,v}p,q\in\{a,v\} denote audio and visual modalities, and let P={(a,v),(v,a)}\mathcal{P}=\{(a,v),(v,a)\} be the two ordered modality pairs. For sample ii, let zp,iz_{p,i} be the mean of the encoder feature over its patch dimension and let zˉp,i=zp,i/∥zp,i∥2\bar z_{p,i}=z_{p,i}/\lVert z_{p,i}\rVert_2 be its normalized vector. With temperature τ>0\tau>0, the bidirectional contrastive loss is

    Lc=−12N∑(p,q)∈P∑i=1Nlog⁡exp⁡(zˉp,iTzˉq,i/τ)∑j=1Nexp⁡(zˉp,iTzˉq,j/τ).\mathcal{L}_c=-\frac{1}{2N}\sum_{(p,q)\in\mathcal{P}}\sum_{i=1}^{N}\log\frac{\exp\left(\bar z_{p,i}^{\mathsf T}\bar z_{q,i}/\tau\right)}{\sum_{j=1}^{N}\exp\left(\bar z_{p,i}^{\mathsf T}\bar z_{q,j}/\tau\right)}.

    The reconstruction loss is evaluated only on tokens masked in the corresponding modality. If xp,msk(i)x_{p,\mathrm{msk}}^{(i)} and x^p,msk(i)\hat x_{p,\mathrm{msk}}^{(i)} are the original and reconstructed masked audio or visual tokens for sample ii, then

    Lrec=1N∑i=1N∑p∈{a,v}∥xp,msk(i)−x^p,msk(i)∥22.\mathcal{L}_{\mathrm{rec}}=\frac{1}{N}\sum_{i=1}^{N}\sum_{p\in\{a,v\}}\left\lVert x_{p,\mathrm{msk}}^{(i)}-\hat x_{p,\mathrm{msk}}^{(i)}\right\rVert_2^2.

    The decoders reconstruct audio and visual inputs from a′a' and v′v'. Each modality also has a discriminator DpD_p. The generator-side Wasserstein loss and discriminator-side loss used on masked tokens are

    Ladv(G)=−1N∑i=1N∑p∈{a,v}Dp(x^p,msk(i)),\mathcal{L}_{\mathrm{adv}}^{(G)}=-\frac{1}{N}\sum_{i=1}^{N}\sum_{p\in\{a,v\}}D_p\left(\hat x_{p,\mathrm{msk}}^{(i)}\right), Ladv(D)=1N∑i=1N∑p∈{a,v}[Dp(x^p,msk(i))−Dp(xp,msk(i))].\mathcal{L}_{\mathrm{adv}}^{(D)}=\frac{1}{N}\sum_{i=1}^{N}\sum_{p\in\{a,v\}}\left[D_p\left(\hat x_{p,\mathrm{msk}}^{(i)}\right)-D_p\left(x_{p,\mathrm{msk}}^{(i)}\right)\right].

    The generator objective is

    L(G)=λcLc+λrecLrec+λadvLadv(G),\mathcal{L}^{(G)}=\lambda_c\mathcal{L}_c+\lambda_{\mathrm{rec}}\mathcal{L}_{\mathrm{rec}}+\lambda_{\mathrm{adv}}\mathcal{L}_{\mathrm{adv}}^{(G)},

    where λc\lambda_c, λrec\lambda_{\mathrm{rec}}, and λadv\lambda_{\mathrm{adv}} are loss weights. The contrastive term aligns paired audio and visual embeddings, while masked reconstruction forces the decoder to use cross-modal predictions.

  5. Knowl 5 — Cross-modal classifier and sliding-window inference

    algorithm

    After representation learning, AVFF processes labeled real and fake videos without masking. The audio and visual encoders produce unimodal features aa and vv for every temporal slice. The A2V and V2A networks produce cross-modal features for every slice, yielding an audio-side combined embedding fa=a⊕a~f_a=a\oplus\tilde a and a visual-side combined embedding fv=v⊕v~f_v=v\oplus\tilde v, where ⊕\oplus concatenates along the feature dimension. The classifier consists of modality-specific patch-reduction networks Ψa\Psi_a and Ψv\Psi_v, followed by a classification head Γ\Gamma:

    l=Q(fa,fv)=Γ(Ψa(fa)⊕Ψv(fv)),l=Q(f_a,f_v)=\Gamma\left(\Psi_a(f_a)\oplus\Psi_v(f_v)\right),

    where ll is the real/fake logit vector. The classifier is trained with standard cross-entropy loss against the real/fake label. At inference, a video is divided into blocks of the training duration TT with stride T/8T/8, equal to one temporal slice. The classifier produces logits for every block, and the final real/fake decision is obtained from the mean logits across all blocks.

  6. Knowl 6 — Training and evaluation protocol

    experimental setup

    The self-supervised representation stage is trained on LRS3, which contains real talking-face videos and no deepfake samples. The supervised classifier is trained using FakeAVCeleb, whose manipulated samples include visual synthesis with FaceSwap, FSGAN, or Wav2Lip and audio synthesis with SV2TTS; samples may contain manipulation in one or both modalities. In the intra-dataset experiment, 70% of FakeAVCeleb is used for training and the remaining 30% is an unseen test set. Cross-manipulation evaluation leaves one of five manipulation categories out for testing and trains on the other four. Cross-dataset evaluation trains on FakeAVCeleb and tests on a subset of KoDF. Accuracy, average precision (AP), and ROC area under the curve (AUC) are averaged over runs with different random seeds. For audio-visual methods, a video is labeled fake when either modality is manipulated; for visual-only methods, it is labeled fake only when the visual modality is manipulated.

  7. Knowl 7 — Intra-dataset FakeAVCeleb performance

    data/table

    On a 70%-30% train-test split of FakeAVCeleb, AVFF substantially outperforms the evaluated visual-only and audio-visual baselines. The comparison uses accuracy (ACC) and ROC area under the curve (AUC); the audio-visual methods are evaluated for videos with either or both modalities manipulated.

    Method Modality ACC AUC
    Xception V 67.9 70.5
    LipForensics V 80.1 82.4
    FTCN V 64.9 84.0
    CViT V 69.7 71.8
    RealForensics V 89.9 94.6
    Emotions Don't Lie AV 78.1 79.8
    MDS AV 82.8 86.5
    AVFakeNet AV 78.4 83.4
    VFD AV 81.5 86.1
    AVoiD-DF AV 83.7 89.2
    AVFF AV 98.6 99.1

    AVFF reaches 98.6% accuracy and 99.1% AUC, exceeding the strongest listed audio-visual baseline, AVoiD-DF, by 14.9 percentage points in accuracy and 9.9 percentage points in AUC. It also exceeds the strongest listed visual-only baseline, RealForensics, by 8.7 accuracy points and 4.5 AUC points.

  8. Knowl 8 — Generalization to unseen manipulation methods

    data/table

    Cross-manipulation evaluation trains on four FakeAVCeleb categories and tests on the held-out fifth category. The categories are RVFA (real visual, fake audio generated with SV2TTS), FVRA-WL (fake visual generated with Wav2Lip, real audio), FVFA-FS (FaceSwap + Wav2Lip + SV2TTS), FVFA-GAN (FaceSwapGAN + Wav2Lip + SV2TTS), and FVFA-WL (Wav2Lip + SV2TTS). AVG-FV averages categories containing fake visuals. Values are AP and AUC percentages.

    Method RVFA FVRA-WL FVFA-FS FVFA-GAN FVFA-WL AVG-FV
    AP AUC AP AUC AP AUC AP AUC AP AUC AP AUC
    Xception - - 88.2 88.3 92.3 93.5 67.6 68.5 91.0 91.0 84.8 85.3
    LipForensics - - 97.8 97.7 99.9 99.9 61.5 68.1 98.6 98.7 89.4 91.1
    FTCN - - 96.2 97.4 100. 100. 77.4 78.3 95.6 96.5 92.3 93.1
    RealForensics - - 88.8 93.0 99.3 99.1 99.8 99.8 93.4 96.7 95.3 97.1
    AV-DFD 74.9 73.3 97.0 97.4 99.6 99.7 58.4 55.4 100. 100. 88.8 88.1
    AVAD (LRS2) 62.4 71.6 93.6 93.7 95.3 95.8 94.1 94.3 93.8 94.1 94.2 94.5
    AVAD (LRS3) 70.7 80.5 91.1 93.0 91.0 92.3 91.6 92.7 91.4 93.1 91.3 92.8
    AVFF 93.3 92.4 94.8 98.2 100. 100. 99.9 100. 99.4 99.8 98.5 99.5

    AVFF obtains AP above 93% and AUC above 92% in every held-out category, with AVG-FV of 98.5 AP and 99.5 AUC. It is best in most categories and remains competitive in the remaining ones, whereas several baselines degrade strongly on the FVFA-GAN and RVFA categories.

  9. Knowl 9 — Cross-dataset generalization to KoDF

    data/table

    A classifier trained on FakeAVCeleb is evaluated without retraining on a subset of the different KoDF dataset. The cross-dataset results are reported as AP and AUC percentages.

    Method Modality AP AUC
    Xception V 76.9 77.7
    LipForensics V 89.5 86.6
    FTCN V 66.8 68.1
    RealForensics V 95.7 93.6
    AV-DFD AV 79.6 82.1
    AVAD AV 87.6 86.9
    AVFF AV 93.1 95.5

    AVFF achieves 93.1 AP and 95.5 AUC, the best values among the listed methods, demonstrating that the representation learned with real-video audio-visual correspondence transfers to a different dataset distribution.

  10. Knowl 10 — Real-fake separation emerges during self-supervised representation learning

    empirical result

    A t-SNE visualization of embeddings produced at the end of AVFF's representation-learning stage shows clear separation between real and fake FakeAVCeleb samples, even though the representation stage was trained exclusively on real LRS3 videos and never saw deepfake examples. The fake samples also form distinct clusters associated with different manipulation categories. Adjacent clusters correspond to related generation procedures; in particular, the FVRA-WL and FVFA-WL clusters are neighboring because both use Wav2Lip. The authors interpret these patterns as evidence that the self-supervised audio-visual representation captures manipulation-sensitive correspondence cues rather than merely memorizing the supervised deepfake classes. This separation is consistent with downstream AUC values reaching approximately 90% within the first one to three classifier-training epochs.

  11. Knowl 11 — Ablation evidence for each AVFF component

    data/table

    The ablation evaluation compares the full AVFF pipeline with variants that remove or replace individual representation-learning or classification components. AP and AUC are percentages.

    Variant AP AUC
    Only contrastive loss 84.2 90.3
    AVFF without cross-modality fusion 87.2 93.1
    AVFF without complementary masking 78.9 90.7
    Only feature embeddings 89.7 97.6
    Only cross-modal embeddings 94.6 98.0
    Mean pooling over patch dimension 96.5 98.1
    AVFF 96.7 99.1

    Using only contrastive learning removes the autoencoding, masking, fusion, and decoding components and substantially lowers performance. Replacing cross-modal predicted tokens with shared learned mask tokens also reduces performance, while replacing complementary masking with random masking produces the largest AP decrease among the representation-learning ablations. Concatenating unimodal and cross-modal embeddings is better than using either alone, and learned unimodal patch-reduction networks slightly outperform mean pooling, indicating that preserving weighted nonlinear combinations of patch cues is useful.

  12. Knowl 12 — Operational limitations of audio-visual correspondence detection

    limitation

    AVFF depends on meaningful temporal coherence between speech audio and facial visual content. It is therefore expected to be difficult to apply to asynchronous videos with audio lag, videos with mismatched audio, videos containing multiple speakers, or videos in which important facial regions are occluded by masks or hands. The method also requires both audio and visual modalities at inference time and does not support videos containing only one modality.

Coverage note — Related work, acknowledgements, future-work suggestions, and supplementary implementation details were omitted because they are not part of the paper's central contributed method or reported main experiments.

References

  1. 1.Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman. Lrs3-ted: a large-scale dataset for visual speech recognition. arXiv preprint arXiv:1809.00496, 2018. 2, 6
  2. 2.Shruti Agarwal, Hany Farid, Ohad Fried, and Maneesh Agrawala. Detecting deep-fake videos from phoneme-viseme mismatches. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops, pages 660–661, 2020. 1, 3
  3. 3.Martin Arjovsky, Soumith Chintala, and Léon Bottou. Wasserstein generative adversarial networks. In International conference on machine learning, pages 214–223. PMLR, 2017. 5
  4. 4.Weiming Bai, Yufan Liu, Zhipeng Zhang, Bing Li, and Weiming Hu. Aunet: Learning relations between action units for face forgery detection. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 24709–24719, 2023. 1
  5. 5.Tadas Baltrušaitis, Chaitanya Ahuja, and Louis-Philippe Morency. Multimodal machine learning: A survey and taxonomy. IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(2):423–443, 2019. 2
  6. 6.Zhixi Cai, Shreya Ghosh, Kalin Stefanov, Abhinav Dhall, Jianfei Cai, Hamid Rezatofighi, Reza Haffari, and Munawar Hayat. Marlin: Masked autoencoder for facial video representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1493–1504, 2023. 4, 5, 6, 8
  7. 7.Harry Cheng, Yangyang Guo, Tianyi Wang, Qi Li, Xiaojun Chang, and Liqiang Nie. Voice-face homogeneity tells deepfake. arXiv preprint arXiv:2203.02195, 2022. 3, 6
  8. 8.Komal Chugh, Parul Gupta, Abhinav Dhall, and Ramanathan Subramanian. Not made for each other-audio-visual dissonance-based deepfake detection and localization. In Proceedings of the 28th ACM international conference on multimedia, pages 439–447, 2020. 3, 6
  9. 9.J. S. Chung and A. Zisserman. Out of time: automated lip sync in the wild. In Workshop on Multi-view Lip-reading, ACCV, 2016. 1, 2
  10. 10.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. In North American Chapter of the Association for Computational Linguistics, 2019. 2
  11. 11.Brian Dolhansky, Joanna Bitton, Ben Pflaum, Jikuo Lu, Russ Howes, Menglin Wang, and Cristian Canton Ferrer. The deepfake detection challenge (dfdc) dataset. arXiv preprint arXiv:2006.07397, 2020. 2, 7
  12. 12.Shichao Dong, Jin Wang, Renhe Ji, Jiajun Liang, Haoqiang Fan, and Zheng Ge. Implicit identity leakage: The stumbling block to improving deepfake detection generalization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 3994–4004, 2023. 3
  13. 13.Xiaoyi Dong, Jianmin Bao, Dongdong Chen, Ting Zhang, Weiming Zhang, Nenghai Yu, Dong Chen, Fang Wen, and Baining Guo. Protecting celebrities from deepfake with identity consistency transformer. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 9458–9468, 2022. 3
  14. 14.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations, 2020. 3
  15. 15.Chao Feng, Ziyang Chen, and Andrew Owens. Self-supervised video forensics by audio-visual anomaly detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10491–10503, 2023. 6, 7
  16. 16.Mariana-Iuliana Georgescu, Eduardo Fonseca, Radu Tudor Ionescu, Mario Lucic, Cordelia Schmid, and Anurag Arnab. Audiovisual masked autoencoders. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 16144–16154, 2023. 2
  17. 17.Yuan Gong, Andrew Rouditchenko, Alexander H. Liu, David Harwath, Leonid Karlinsky, Hilde Kuehne, and James R. Glass. Contrastive audio-visual masked autoencoder. In ICLR. OpenReview.net, 2023. 2, 4, 5, 6, 8
  18. 18.Andrey Guzhov, Federico Raue, Jörn Hees, and Andreas Dengel. Audioclip: Extending clip to image, text and audio. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 976–980. IEEE, 2022. 1, 2, 4
  19. 19.Alexandros Haliassos, Konstantinos Vougioukas, Stavros Petridis, and Maja Pantic. Lips don’t lie: A generalisable and robust approach to face forgery detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 5039–5049, 2021. 3, 6, 7
  20. 20.Alexandros Haliassos, Rodrigo Mira, Stavros Petridis, and Maja Pantic. Leveraging real talking faces via self-supervision for robust forgery detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14950–14962, 2022. 2, 3, 6, 7
  21. 21.Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. Masked autoencoders are scalable vision learners. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 15979–15988, 2022. 2, 8
  22. 22.Baojin Huang, Zhongyuan Wang, Jifan Yang, Jiaxin Ai, Qin Zou, Qian Wang, and Dengpan Ye. Implicit identity driven deepfake face swapping detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 4490–4499, 2023. 3
  23. 23.Po-Yao Huang, Hu Xu, Juncheng Li, Alexei Baevski, Michael Auli, Wojciech Galuba, Florian Metze, and Christoph Feichtenhofer. Masked autoencoders that listen. Advances in Neural Information Processing Systems, 35:28708–28720, 2022. 4, 5, 6, 8
  24. 24.Hafsa Ilyas, Ali Javed, and Khalid Mahmood Malik. Avfakenet: A unified end-to-end dense swin transformer deep learning model for audio–visual deepfakes detection. Applied Soft Computing, 136:110124, 2023. 3, 6
  25. 25.Yujin Jeong, Wonjeong Ryoo, Seunghyun Lee, Dabin Seo, Wonmin Byeon, Sangpil Kim, and Jinkyu Kim. The power of sound (tpos): Audio reactive video generation with stable diffusion. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7822–7832, 2023. 8
  26. 26.Ye Jia, Yu Zhang, Ron Weiss, Quan Wang, Jonathan Shen, Fei Ren, Patrick Nguyen, Ruoming Pang, Ignacio Lopez Moreno, Yonghui Wu, et al. Transfer learning from speaker verification to multispeaker text-to-speech synthesis. Advances in neural information processing systems, 31, 2018. 6
  27. 27.Tackhyun Jung, Sangwon Kim, and Keecheon Kim. Deepvision: Deepfakes detection using human eye blinking pattern. IEEE Access, 8:83144–83154, 2020. 3
  28. 28.Bachir Kaddar, Sid Ahmed Fezza, Wassim Hamidouche, Zahid Akhtar, and Abdenour Hadid. Hcit: Deepfake video detection using a hybrid model of cnn features and vision transformer. In 2021 International Conference on Visual Communications and Image Processing (VCIP), pages 1–5, 2021. 3
  29. 29.Hasam Khalid, Shahroz Tariq, Minha Kim, and Simon S Woo. Fakeavceleb: A novel audio-video multimodal deepfake dataset. arXiv preprint arXiv:2108.05080, 2021. 2, 6
  30. 30.Pavel Korshunov and Sébastien Marcel. Deepfakes: a new threat to face recognition? assessment and detection. arXiv preprint arXiv:1812.08685, 2018. 7
  31. 31.Iryna Korshunova, Wenzhe Shi, Joni Dambre, and Lucas Theis. Fast face-swap using convolutional neural networks. In Proceedings of the IEEE international conference on computer vision, pages 3677–3685, 2017. 6
  32. 32.Patrick Kwon, Jaeseong You, Gyuhyeon Nam, Sungwoo Park, and Gyeongsu Chae. Kodf: A large-scale korean deepfake detection dataset. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10744–10753, 2021. 7
  33. 33.Yuezun Li, Ming-Ching Chang, and Siwei Lyu. In ictu oculi: Exposing ai created fake videos by detecting eye blinking. In 2018 IEEE International Workshop on Information Forensics and Security (WIFS), pages 1–7, 2018. 3
  34. 34.Kevin Lutz and Robert Bassett. Deepfake detection with inconsistent head poses: Reproducibility and analysis. CoRR, abs/2108.12715, 2021. 3
  35. 35.Shugao Ma, Tomas Simon, Jason Saragih, Dawei Wang, Yuecheng Li, Fernando De la Torre, and Yaser Sheikh. Pixel codec avatars. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 64–73, 2021. 1
  36. 36.Trisha Mittal, Uttaran Bhattacharya, Rohan Chandra, Aniket Bera, and Dinesh Manocha. Emotions don’t lie: An audio-visual deepfake detection method using affective cues. In Proceedings of the 28th ACM international conference on multimedia, pages 2823–2832, 2020. 1, 2, 3, 6
  37. 37.Trisha Mittal, Uttaran Bhattacharya, Rohan Chandra, Aniket Bera, and Dinesh Manocha. M3er: Multiplicative multimodal emotion recognition using facial, textual, and speech cues. In Proceedings of the AAAI conference on artificial intelligence, pages 1359–1367, 2020. 1
  38. 38.Pedro Morgado, Ishan Misra, and Nuno Vasconcelos. Robust audio-visual instance discrimination. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12934–12945, 2021. 4
  39. 39.Yuval Nirkin, Yosi Keller, and Tal Hassner. Fsgan: Subject agnostic face swapping and reenactment. In Proceedings of the IEEE/CVF international conference on computer vision, pages 7184–7193, 2019. 6
  40. 40.Utkarsh Ojha, Yuheng Li, and Yong Jae Lee. Towards universal fake image detectors that generalize across generative models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 24480–24489, 2023. 3
  41. 41.KR Prajwal, Rudrabha Mukhopadhyay, Vinay P Namboodiri, and CV Jawahar. A lip sync expert is all you need for speech to lip generation in the wild. In Proceedings of the 28th ACM international conference on multimedia, pages 484–492, 2020. 1, 6
  42. 42.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PMLR, 2021. 1, 2, 4
  43. 43.Andreas Rossler, Davide Cozzolino, Luisa Verdoliva, Christian Riess, Justus Thies, and Matthias Nießner. Faceforensics++: Learning to detect manipulated facial images. In Proceedings of the IEEE/CVF international conference on computer vision, pages 1–11, 2019. 2, 6, 7
  44. 44.Andrew Rouditchenko, Angie Boggust, David Harwath, Brian Chen, Dhiraj Joshi, Samuel Thomas, Kartik Audhkhasi, Hilde Kuehne, Rameswar Panda, Rogerio Feris, et al. Avlnet: Learning audio-visual language representations from instructional videos. In Annual Conference of the International Speech Communication Association. International Speech Communication Association, 2021. 4
  45. 45.Ludan Ruan, Yiyang Ma, Huan Yang, Huiguo He, Bei Liu, Jianlong Fu, Nicholas Jing Yuan, Qin Jin, and Baining Guo. Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10219–10228, 2023. 8
  46. 46.Laurens van der Maaten and Geoffrey E. Hinton. Visualizing data using t-sne. Journal of Machine Learning Research, 9:2579–2605, 2008. 2
  47. 47.Jun Wang, Yinglu Liu, Yibo Hu, Hailin Shi, and Tao Mei. Facex-zoo: A pytorch toolbox for face recognition. In Proceedings of the 29th ACM International Conference on Multimedia, pages 3779–3782, 2021. 4
  48. 48.Ting-Chun Wang, Arun Mallya, and Ming-Yu Liu. One-shot free-view neural talking-head synthesis for video conferencing. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10039–10049, 2021. 1
  49. 49.Zhendong Wang, Jianmin Bao, Wengang Zhou, Weilun Wang, and Houqiang Li. Altfreezing for more general video face forgery detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 4129–4138, 2023. 3
  50. 50.Deressa Wodajo and Solomon Atnafu. Deepfake video detection using convolutional vision transformer. arXiv preprint arXiv:2102.11126, 2021. 3, 6
  51. 51.Lior Wolf, Tal Hassner, and Itay Maoz. Face recognition in unconstrained videos with matched background similarity. In CVPR 2011, pages 529–534. IEEE, 2011. 2
  52. 52.Haojie Wu, Pan Hui, and Pengyuan Zhou. Deepfake in the metaverse: An outlook survey. arXiv preprint arXiv:2306.07011, 2023. 1
  53. 53.Wenyuan Yang, Xiaoyu Zhou, Zhikai Chen, Bofei Guo, Zhongjie Ba, Zhihua Xia, Xiaochun Cao, and Kui Ren. Avoid-df: Audio-visual joint learning for detecting deepfake. IEEE Transactions on Information Forensics and Security, 18:2015–2029, 2023. 2, 3, 6
  54. 54.Xin Yang, Yuezun Li, and Siwei Lyu. Exposing deep fakes using inconsistent head poses. In ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 8261–8265, 2019. 3
  55. 55.Xiaohui Zhao, Yang Yu, Rongrong Ni, and Yao Zhao. Exploring complementarity of global and local spatiotemporal information for fake face video detection. In ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 2884–2888, 2022. 3
  56. 56.Yinglin Zheng, Jianmin Bao, Dong Chen, Ming Zeng, and Fang Wen. Exploring temporal coherence for more general video face forgery detection. In Proceedings of the IEEE/CVF international conference on computer vision, pages 15044–15054, 2021. 3, 6, 7
  57. 57.Yipin Zhou and Ser-Nam Lim. Joint audio-visual deepfake detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 14800–14809, 2021. 3, 7
  58. 58.Yang Zhou, Zhan Xu, Chris Landreth, Evangelos Kalogerakis, Subhransu Maji, and Karan Singh. Visemenet: Audio-driven animator-centric speech animation. ACM Transactions on Graphics (TOG), 37(4):1–10, 2018. 1
  59. 59.Hao Zhu, Wayne Wu, Wentao Zhu, Liming Jiang, Siwei Tang, Li Zhang, Ziwei Liu, and Chen Change Loy. Celebv-hq: A large-scale video facial attributes dataset. In European conference on computer vision, pages 650–667. Springer, 2022. 2
  60. 60.Wanyi Zhuang, Qi Chu, Zhentao Tan, Qiankun Liu, Haojie Yuan, Changtao Miao, Zixiang Luo, and Nenghai Yu. Uia-vit: Unsupervised inconsistency-aware method based on vision transformer for face forgery detection. In Computer Vision – ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part V, page 391–407, Berlin, Heidelberg, 2022. Springer-Verlag. 3

Citation

MLA
Oorloff, T., et al. “AVFF: Audio-Visual Feature Fusion for Video Deepfake Detection”. arXiv, 2024, http://arxiv.org/abs/2406.02951v1.
APA
Oorloff, T., Koppisetti, S., Bonettini, N., Solanki, D., Colman, B., Yacoob, Y., Shahriyari, A., & Bharaj, G. (2024). AVFF: Audio-Visual Feature Fusion for Video Deepfake Detection. arXiv. http://arxiv.org/abs/2406.02951v1
Chicago
Oorloff, T., S. Koppisetti, N. Bonettini, et al. 2024. “AVFF: Audio-Visual Feature Fusion for Video Deepfake Detection”. arXiv. http://arxiv.org/abs/2406.02951v1.
Harvard
Oorloff, T. et al. (2024) “AVFF: Audio-Visual Feature Fusion for Video Deepfake Detection”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2406.02951v1.
Vancouver
1. Oorloff T, Koppisetti S, Bonettini N, Solanki D, Colman B, Yacoob Y, Shahriyari A, Bharaj G (2024) AVFF: Audio-Visual Feature Fusion for Video Deepfake Detection. arXiv

BibTeX

@article{oorloff2024avff,
  title = {AVFF: Audio-Visual Feature Fusion for Video Deepfake Detection},
  author = {Oorloff, Trevine and Koppisetti, Surya and Bonettini, Nicolò and Solanki, Divyaraj and Colman, Ben and Yacoob, Yaser and Shahriyari, Ali and Bharaj, Gaurav},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2406.02951v1},
  eprint = {2406.02951}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE