Revealing Single Frame Bias for Video-and-Language Learning

Jie LeiTamara L. BergMohit Bansal

article2023ACL167 citations

Reveals a pervasive static appearance bias in standard video-and-language benchmarks by showing that single-frame training paired with inference-time frame ensembling outperforms multi-frame methods, while proposing two new action-focused retrieval tasks to properly evaluate temporal reasoning.

Listen

Standard artificial intelligence systems for video and language understanding typically process multiple video frames during training to capture changes over time. However, training multi-frame models demands substantial computational power, memory, and time, which significantly increases development and infrastructure costs. This raises an important question for technology leaders: does processing full multi-frame sequences during training provide enough value to justify these high computational expenses, or can simpler, more efficient methods achieve comparable performance?

The article evaluates whether a model trained using only a single randomly selected frame per video can match or exceed the performance of complex multi-frame systems across standard video-language tasks, such as video retrieval and question answering. In doing so, it investigates whether widely used benchmark datasets rely on genuine temporal understanding or are instead dominated by static visual cues like background scenes and still objects.

To conduct this evaluation, the researchers developed an architecture called SINGULARITY. The model pairs standard vision and language encoders with a cross-modal fusion encoder. During training, the system samples only a single frame per video, drastically cutting computational overhead. During testing and deployment, the model samples multiple frames and applies an early fusion strategy, combining all visual features before making a video-level prediction. The approach was evaluated across six established video retrieval and question answering benchmarks using pre-training datasets ranging from 5.4 million to 17.3 million image-text and video-text pairs, and compared against baseline systems trained on up to 400 million examples.

The findings show that single-frame training achieves state-of-the-art results across standard benchmarks while dramatically lowering resource requirements. Pre-trained on just 5.4 million examples, the single-frame model matched or outperformed existing multi-frame models trained on dozens to hundreds of millions of samples. Expanding pre-training to 17.3 million examples pushed performance even higher, achieving top retrieval scores on benchmarks such as DiDeMo and ActivityNet Captions. In efficiency terms, the single-frame approach required up to 16 times less pre-training computation than comparable models and trained between 2.8 and 8.5 times faster during task adaptation, allowing substantially larger batch sizes on standard hardware. Furthermore, early fusion consistently outperformed traditional late fusion methods (such as score averaging), which frequently suffer from noisy individual frame predictions.

These results demonstrate a substantial static appearance bias in popular video-language benchmarks: standard tests largely evaluate whether a model recognizes objects and environments rather than actions occurring over time. For enterprise applications where static visual matching is sufficient (such as tagging or general video search), organizations can deploy lightweight single-frame architectures to achieve major cost savings and faster deployment cycles without sacrificing accuracy. However, because standard benchmarks mask shortcomings in temporal reasoning, the researchers adapted the action-focused Something-Something v2 dataset into two new retrieval benchmarks. On these motion-critical tasks, the single-frame model underperformed multi-frame baselines by significant margins (such as a 10.9-point deficit in template retrieval), confirming that single-frame training cannot replace temporal modeling when tracking true physical actions.

Organizations should adopt a split strategy based on use-case requirements. Teams focused on general video search, cataloging, and high-volume question answering should leverage single-frame training with early fusion to maximize throughput and minimize cloud compute costs. Conversely, teams building applications that depend strictly on sequence and motion (such as safety monitoring, gesture control, or fine-grained action classification) must incorporate temporal modeling modules, such as the multi-frame temporal variant introduced in the article. Researchers and evaluation teams should immediately incorporate fine-grained action benchmarks to ensure systems are tested on genuine temporal understanding rather than static visual shortcuts.

Cover for Revealing Single Frame Bias for Video-and-Language Learning

Abstract

Training an effective video-and-language model intuitively requires multiple frames as model inputs. However, it is unclear whether using multiple frames is beneficial to downstream tasks, and if yes, whether the performance gain is worth the drastically-increased computation and memory costs resulting from using more frames. In this work, we explore single-frame models for video-and-language learning. On a diverse set of video-and-language tasks (including text-to-video retrieval and video question answering), we show the surprising result that, with large-scale pre-training and a proper frame ensemble strategy at inference time, a single-frame trained model that does not consider temporal information can achieve better performance than existing methods that use multiple frames for training. This result reveals the existence of a strong “static appearance bias” in popular video-and-language datasets. Therefore, to allow for a more comprehensive evaluation of video-and-language models, we propose two new retrieval tasks based on existing fine-grained action recognition datasets that encourage temporal modeling. Our code is available at https://github.com/jayleicn/singularity.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methods
  • 4 Experiments
  • 4.1 Downstream Task Setup
  • 4.2 Comparison on Existing Datasets
  • 4.3 New Temporal Tasks
  • 5 Analysis
  • 6 Conclusion
  • Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • Appendix
  • A.1 Additional Modeling Details
  • A.2 Additional Experiments
  • A.3 Additional Data Details

Knowls

  1. Knowl 1 — Single-frame training with early-fusion video inference

    model/method

    SINGULARITY uses a vision encoder FvF_v, a language encoder FlF_l, and a multimodal transformer encoder HH containing self-attention, cross-attention, and feed-forward layers. For a video V=[f1,…,fT]V=[f_1,\ldots,f_T] with TT frames and paired text SS, training randomly selects one frame ftf_t and predicts from its visual representation and the text representation:

    ptrain=H(Fl(S),Fv(ft)).p_{\text{train}}=H(F_l(S),F_v(f_t)).

    At inference, the model uniformly samples TtestT_{\text{test}} frames fτ1,…,fτTtestf_{\tau_1},\ldots,f_{\tau_{T_{\text{test}}}}, encodes them separately, concatenates their visual representations, and makes one video-level prediction:

    pvideo=H(Fl(S),[Fv(fτ1);…;Fv(fτTtest)]).p_{\text{video}}=H\left(F_l(S),[F_v(f_{\tau_1});\ldots;F_v(f_{\tau_{T_{\text{test}}}})]\right).

    Here ptrainp_{\text{train}} and pvideop_{\text{video}} are task predictions or scores, tt is a randomly selected training-frame index, each τi\tau_i is an inference-frame index, and [;][;] denotes concatenation along the visual-token sequence. The approach therefore uses one frame per training example but lets the multimodal encoder jointly assess multiple frames at inference; the base model has no explicit temporal encoder.

  2. Knowl 2 — Pretraining architecture, data, and objectives

    model/method

    SINGULARITY is initialized with a BEiT({\text{BASE}}) vision encoder pretrained on ImageNet-21K, the first nine layers of BERT({\text{BASE}}) as its language encoder, and the last three layers of the same BERT model as its multimodal encoder; the multimodal cross-attention layers are initialized randomly. Pretraining combines image-text data from COCO, Visual Genome, SBU Captions, CC3M, and CC12M with WebVid video-text data. For WebVid, the model samples one frame per video during training. The 5M corpus contains 5.44 million image/video-text examples from CC3M and WebVid; the 17M corpus contains 17.28 million images and videos and 18.41 million text examples from all listed datasets.

    The model is pretrained with three objectives: vision-text contrastive learning aligns pooled vision and text representations; masked language modeling predicts masked tokens using textual and visual context, with a 50% masking ratio; and vision-text matching predicts whether a vision-text pair matches, using hard negative sampling. Pretraining runs for 10 epochs with AdamW, an initial learning rate of 10−410^{-4}, and batch size 128 per GPU on three NVIDIA A100 GPUs.

  3. Knowl 3 — Text-to-video retrieval on established benchmarks

    empirical result

    On MSRVTT, DiDeMo, and ActivityNet Captions, SINGULARITY is fine-tuned using one frame per training example. Evaluation uses recall at rank KK (R@KK); DiDeMo and ActivityNet Captions use paragraph-to-video retrieval, while MSRVTT uses standard text-to-video retrieval. Inference uses 12 frames for MSRVTT and DiDeMo and 32 for the longer ActivityNet Captions videos. The table compares the paper's 5M- and 17M-pretrained models with CLIP4Clip, a method pretrained on 400M examples and fine-tuned with more frames. SINGULARITY performs especially strongly on DiDeMo and ActivityNet Captions, although its MSRVTT results do not exceed CLIP4Clip at every rank.

    Method Pretraining Train frames MSRVTT DiDeMo ActivityNet Captions
    R@1 R@5 R@10 R@1 R@5 R@10 R@1 R@5 R@10
    CLIP4Clip 400M 12/64/64 42.0 68.6 78.7 42.8 68.5 79.2 40.5 72.4 98.2
    SINGULARITY 5M 1 36.8 65.9 75.5 47.4 75.2 84.0 43.0 70.6 81.3
    SINGULARITY 17M 1 41.5 68.7 77.0 53.9 79.4 86.9 47.1 75.5 85.5
  4. Knowl 4 — Video question answering with one-frame training

    empirical result

    SINGULARITY is trained with one video frame per example for open-ended video question answering and uses 12 frames at inference; MSRVTT-MC uses the MSRVTT retrieval model to select among five candidate captions. The metric is accuracy. The table compares results with representative systems that use substantially more pretraining data and/or training frames. The 17M SINGULARITY model obtains the highest listed ActivityNet-QA accuracy and ties the listed best MSRVTT-MC result, while its MSRVTT-QA score is below All-in-one's result.

    Method Pretraining Train frames MSRVTT-QA ActivityNet-QA MSRVTT-MC
    JustAsk 69M 640 41.5 38.9 –
    MERLOT 180M 5 43.1 41.4 90.9
    All-in-one 138M 9 44.3 – 92.0
    SINGULARITY 5M 1 42.7 41.8 92.0
    SINGULARITY 17M 1 43.5 43.1 92.1
  5. Knowl 5 — SSv2 retrieval tasks designed to test temporal understanding

    definition

    The paper defines two text-to-video retrieval tasks from Something-Something v2 (SSv2), whose action videos often require fine-grained temporal reasoning. In SSv2-Template Retrieval, the text queries are the dataset's 174 action templates, such as “Throwing [something] in the air and catching it”; the object is omitted so the task emphasizes action and temporal structure. In SSv2-Label Retrieval, queries are the annotated labels, such as “Throwing keys in the air and catching it,” so matching can rely on both objects and their actions. Each task uses 168,913 training videos. Because test annotations are unavailable, evaluation uses 2,088 validation videos, sampled as 12 videos for each template. The template task is intended to focus more narrowly on temporal understanding, whereas the label task evaluates both static appearance and motion.

  6. Knowl 6 — Explicit temporal modeling improves SSv2 retrieval

    empirical result

    The paper's SINGULARITY-temporal variant adds a two-layer temporal transformer after the vision encoder. It adds learned temporal position encodings to frame representations and passes the transformer outputs to the multimodal encoder. Starting from a single-frame-pretrained checkpoint, the authors perform a second pretraining stage on WebVid with four input frames for five epochs. On SSv2, this variant substantially improves over single-frame SINGULARITY and reaches or exceeds the listed retrieval baselines on the reported metrics. In particular, the template task's R@1 rises from 42.0 to 77.0 with 5M pretraining, consistent with the need for temporal information when query objects are removed. R@KK denotes recall at rank KK; the Frozen label-task training run did not converge.

    Method Pretraining SSv2-Label SSv2-Template
    R@1 R@5 R@10 R@1 R@5 R@10
    Frozen (4 frames) 5M – – – 52.9 94.8 99.4
    CLIP4Clip (12 frames) 400M 43.1 71.4 80.7 77.0 96.6 98.3
    SINGULARITY (1 frame) 5M 36.4 64.9 75.4 42.0 86.2 94.3
    SINGULARITY-temporal (4 frames) 5M 44.1 73.5 82.2 77.0 98.9 99.4
    SINGULARITY-temporal (4 frames) 17M 47.4 75.9 84.0 77.6 96.0 98.9
  7. Knowl 7 — Joint frame fusion outperforms score-level fusion

    empirical result

    The authors compare early fusion, which concatenates encoded frame representations before the multimodal encoder produces a video-level prediction, against late fusion, which independently predicts from each frame and aggregates the scores using LogSumExp, max, or mean. Across the tested inference counts (1, 2, 4, 8, 12, and 16 frames), early fusion performs better on MSRVTT retrieval and ActivityNet-QA than the three late-fusion strategies. Increasing the number of frames generally improves performance, and early fusion improves consistently as more frames are added. Late fusion can instead lose accuracy with additional frames—for example, ActivityNet-QA with max aggregation performs worse beyond four frames than with four. The paper attributes the difference to late fusion's reliance on potentially inaccurate, unstable predictions made from individual frames without their surrounding context.

  8. Knowl 8 — More cross-modal pretraining narrows the single- versus multi-frame gap

    empirical result

    The paper compares one-frame SINGULARITY with a four-frame model under four cross-modal pretraining conditions: no cross-modal pretraining, WebVid (2.49 million videos), the 5.44-million-example corpus, and the 17.28-million-example corpus. Both models improve substantially as pretraining data increases. Across the evaluated downstream tasks, the performance difference between the one-frame and four-frame models generally shrinks nearly monotonically with more pretraining, supporting the authors' hypothesis that large-scale pretraining helps compensate for the noisier, less-contextual single-frame training signal. This trend is not universal: tasks requiring fine-grained temporal modeling, such as SSv2-Label Retrieval, still benefit from explicit multi-frame modeling.

  9. Knowl 9 — Single-frame training reduces training cost

    empirical result

    On one RTX A6000 GPU with 48 GB of memory, the authors measured training time over 8,394 DiDeMo training examples. The 17M-pretrained one-frame SINGULARITY model trained 2.8 times faster than the four-frame Frozen model and 8.5 times faster than the 64-frame CLIP4Clip model, while the paper reports better downstream performance than those baselines in this comparison. Its maximum allowed batch size on one GPU was 190, compared with 50 for Frozen. For pretraining, SINGULARITY took about one day on the 5M corpus and four days on the 17M corpus using three A100 GPUs; the paper reports that its 5M pretraining required one-sixteenth the computation of AlignPrompt's 5M pretraining.

  10. Knowl 10 — Single-frame modeling is limited on temporal tasks

    limitation

    The authors state that single-frame training does not work well on genuinely temporal tasks such as the proposed SSv2 tasks, where recognizing fine-grained actions can require multiple ordered frames. They also identify a greater dependence on pretraining data than for multi-frame models. More broadly, because the system learns from dataset distributions and single-frame inputs can omit necessary context, its predictions may be biased or unreliable on tasks requiring deeper temporal understanding.

Coverage note — Supplementary zero-shot retrieval, image-text and image-question-answering results, and image-size and training-objective ablations are omitted because they are secondary to the paper's core single-frame method, temporal-task proposal, and principal video-task analyses.

References

  1. 1.Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018. Bottom-up and top-down attention for image captioning and visual question answering. In CVPR.
  2. 2.Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell. 2017. Localizing moments in video with natural language. In ICCV.
  3. 3.Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015. Vqa: Visual question answering. In ICCV.
  4. 4.Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman. 2021. Frozen in time: A joint video and image encoder for end-to-end retrieval. In ICCV.
  5. 5.Hangbo Bao, Li Dong, and Furu Wei. 2021. Beit: Bert pre-training of image transformers. In ICLR.
  6. 6.Gedas Bertasius, Heng Wang, and Lorenzo Torresani. 2021. Is space-time attention all you need for video understanding. In ICML.
  7. 7.Shyamal Buch, Cristobal Eyzaguirre, Adrien Gaidon, Jiajun Wu, Li Fei-Fei, and Juan Carlos Niebles. 2022. Revisiting the “video” in video-language understanding. arXiv preprint arXiv:2206.01720.
  8. 8.Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. 2015. Activitynet: A large-scale video benchmark for human activity understanding. In CVPR.
  9. 9.Soravit Changpinyo, Piyush Sharma, Nan Ding, and Radu Soricut. 2021. Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. In CVPR.
  10. 10.Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick. 2015. Microsoft coco captions: Data collection and evaluation server. arXiv.
  11. 11.Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2020. Uniter: Learning universal image-text representations. In ECCV.
  12. 12.Jaemin Cho, Jie Lei, Hao Tan, and Mohit Bansal. 2021. Unifying vision-and-language tasks via text generation. arXiv.
  13. 13.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009. Imagenet: A large-scale hierarchical image database. In CVPR.
  14. 14.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In NAACL.
  15. 15.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR.
  16. 16.Victor Escorcia, Mattia Soldan, Josef Sivic, Bernard Ghanem, and Bryan Russell. 2019. Temporal localization of moments in video collections with natural language. arXiv.
  17. 17.Valentin Gabeur, Chen Sun, Karteek Alahari, and Cordelia Schmid. 2020. Multi-modal transformer for video retrieval. In ECCV.
  18. 18.Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al. 2017a. The" something something" video database for learning and evaluating visual common sense. In ICCV.
  19. 19.Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2017b. Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In CVPR.
  20. 20.Dan Hendrycks, Kimin Lee, and Mantas Mazeika. 2019. Using pre-training can improve model robustness and uncertainty. In ICML.
  21. 21.Andrew Jaegle, Sebastian Borgeaud, Jean-Baptiste Alayrac, Carl Doersch, Catalin Ionescu, David Ding, Skanda Koppula, Daniel Zoran, Andrew Brock, Evan Shelhamer, et al. 2022. Perceiver io: A general architecture for structured inputs & outputs. In ICLR.
  22. 22.Andrew Jaegle, Felix Gimeno, Andrew Brock, Andrew Zisserman, Oriol Vinyals, and Joao Carreira. 2021. Perceiver: General perception with iterative attention. In ICML.
  23. 23.Yunseok Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim. 2017. Tgif-qa: Toward spatio-temporal reasoning in visual question answering. In CVPR.
  24. 24.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V Le, Yunhsuan Sung, Zhen Li, and Tom Duerig. 2021. Scaling up visual and vision-language representation learning with noisy text supervision. arXiv.
  25. 25.Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al. 2017. The kinetics human action video dataset. arXiv.
  26. 26.Wonjae Kim, Bokyung Son, and Ildoo Kim. 2021. Vilt: Vision-and-language transformer without convolution or region supervision. In ICML.
  27. 27.Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles. 2017a. Dense-captioning events in videos. In ICCV.
  28. 28.Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al. 2017b. Visual genome: Connecting language and vision using crowdsourced dense image annotations. IJCV.
  29. 29.Jie Lei, Linjie Li, Luowei Zhou, Zhe Gan, Tamara L Berg, Mohit Bansal, and Jingjing Liu. 2021. Less is more: Clipbert for video-and-language learning via sparse sampling. In CVPR.
  30. 30.Jie Lei, Licheng Yu, Mohit Bansal, and Tamara L Berg. 2018. Tvqa: Localized, compositional video question answering. In EMNLP.
  31. 31.Jie Lei, Licheng Yu, Tamara L Berg, and Mohit Bansal. 2020. What is more likely to happen next? video-and-language future event prediction. In EMNLP.
  32. 32.Dongxu Li, Junnan Li, Hongdong Li, Juan Carlos Niebles, and Steven CH Hoi. 2021a. Align and prompt: Video-and-language pre-training with entity prompts. arXiv.
  33. 33.Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. arXiv preprint arXiv:2201.12086.
  34. 34.Junnan Li, Ramprasaath R Selvaraju, Akhilesh Deepak Gotmare, Shafiq Joty, Caiming Xiong, and Steven Hoi. 2021b. Align before fuse: Vision and language representation learning with momentum distillation. In NeurIPS.
  35. 35.Linjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan, Licheng Yu, and Jingjing Liu. 2020a. Hero: Hierarchical encoder for video+ language omni-representation pre-training. In EMNLP.
  36. 36.Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. 2019. Visualbert: A simple and performant baseline for vision and language. arXiv.
  37. 37.Wei Li, Can Gao, Guocheng Niu, Xinyan Xiao, Hao Liu, Jiachen Liu, Hua Wu, and Haifeng Wang. 2021c. Unimo: Towards unified-modal understanding and generation via cross-modal contrastive learning. In ACL.
  38. 38.Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, et al. 2020b. Oscar: Object-semantics aligned pre-training for vision-language tasks. In ECCV.
  39. 39.Yingwei Li, Yi Li, and Nuno Vasconcelos. 2018. Resound: Towards action recognition without representation bias. In ECCV.
  40. 40.Xudong Lin, Gedas Bertasius, Jue Wang, Shih-Fu Chang, Devi Parikh, and Lorenzo Torresani. 2021. Vx2text: End-to-end learning of video-based text generation from multimodal inputs. In CVPR.
  41. 41.Yan-Bo Lin, Jie Lei, Mohit Bansal, and Gedas Bertasius. 2022. Eclipse: Efficient long-range video retrieval using sight and sound. In ECCV.
  42. 42.Yang Liu, Samuel Albanie, Arsha Nagrani, and Andrew Zisserman. 2020. Use what you have: Video retrieval using representations from collaborative experts. In BMVC.
  43. 43.Ilya Loshchilov and Frank Hutter. 2017. Sgdr: Stochastic gradient descent with warm restarts. In ICLR.
  44. 44.Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In ICLR.
  45. 45.Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In NeurIPS.
  46. 46.Huaishao Luo, Lei Ji, Ming Zhong, Yang Chen, Wen Lei, Nan Duan, and Tianrui Li. 2021. Clip4clip: An empirical study of clip for end to end video clip retrieval. arXiv preprint arXiv:2104.08860.
  47. 47.Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic. 2019. Howto100m: Learning a text-video embedding by watching hundred million narrated video clips. In ICCV.
  48. 48.Vicente Ordonez, Girish Kulkarni, and Tamara Berg. 2011. Im2text: Describing images using 1 million captioned photographs. NeurIPS.
  49. 49.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. Pytorch: An imperative style, high-performance deep learning library. In NeurIPS.
  50. 50.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. arXiv.
  51. 51.Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In ACL.
  52. 52.Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. 2012. Ucf101: A dataset of 101 human actions classes from videos in the wild. arXiv.
  53. 53.Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid. 2019. Videobert: A joint model for video and language representation learning. In ICCV.
  54. 54.Hao Tan and Mohit Bansal. 2019. Lxmert: Learning cross-modality encoder representations from transformers. In EMNLP.
  55. 55.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou. 2021. Training data-efficient image transformers & distillation through attention. In ICML.
  56. 56.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In NeurIPS.
  57. 57.Alex Jinpeng Wang, Yixiao Ge, Rui Yan, Yuying Ge, Xudong Lin, Guanyu Cai, Jianping Wu, Ying Shan, Xiaohu Qie, and Mike Zheng Shou. 2022. All in one: Exploring unified video-language pre-training. arXiv.
  58. 58.Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool. 2016. Temporal segment networks: Towards good practices for deep action recognition. In ECCV.
  59. 59.Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang. 2017. Video question answering via gradually refined attention over appearance and motion. In ACM MM.
  60. 60.Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko, Armen Aghajanyan, Florian Metze, Luke Zettlemoyer, and Christoph Feichtenhofer. 2021. Videoclip: Contrastive pre-training for zero-shot video-text understanding. In EMNLP.
  61. 61.Jun Xu, Tao Mei, Ting Yao, and Yong Rui. 2016. Msr-vtt: A large video description dataset for bridging video and language. In CVPR.
  62. 62.Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, and Cordelia Schmid. 2021. Just ask: Learning to answer questions from millions of narrated videos. In ICCV.
  63. 63.Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. 2014. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. TACL.
  64. 64.Youngjae Yu, Jongseok Kim, and Gunhee Kim. 2018. A joint sequence fusion model for video question answering and retrieval. In ECCV.
  65. 65.Zhou Yu, Dejing Xu, Jun Yu, Ting Yu, Zhou Zhao, Yueting Zhuang, and Dacheng Tao. 2019. Activitynet-qa: A dataset for understanding complex web videos via question answering. In AAAI.
  66. 66.Rowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu, Jae Sung Park, Jize Cao, Ali Farhadi, and Yejin Choi. 2021. Merlot: Multimodal neural script knowledge models. NeurIPS.
  67. 67.Peng Zhang, Yash Goyal, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2016. Yin and yang: Balancing and answering binary visual questions. In CVPR.
  68. 68.Linchao Zhu and Yi Yang. 2020. Actbert: Learning global-local video-text representations. In CVPR.

Citation

MLA
Lei, J., et al. “Revealing Single Frame Bias for Video-and-Language Learning”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 487–507, https://doi.org/10.18653/v1/2023.acl-long.29.
APA
Lei, J., Berg, T., & Bansal, M. (2023). Revealing Single Frame Bias for Video-and-Language Learning. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 487–507. https://doi.org/10.18653/v1/2023.acl-long.29
Chicago
Lei, J., T. Berg, and M. Bansal. 2023. “Revealing Single Frame Bias for Video-and-Language Learning”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 487–507. https://doi.org/10.18653/v1/2023.acl-long.29.
Harvard
Lei, J., Berg, T. and Bansal, M. (2023) “Revealing Single Frame Bias for Video-and-Language Learning”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 487–507. Available at: https://doi.org/10.18653/v1/2023.acl-long.29.
Vancouver
1. Lei J, Berg T, Bansal M (2023) Revealing Single Frame Bias for Video-and-Language Learning. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 487–507

BibTeX

@inproceedings{lei-etal-2023-revealing,
    title = "Revealing Single Frame Bias for Video-and-Language Learning",
    author = "Lei, Jie  and
      Berg, Tamara  and
      Bansal, Mohit",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.29/",
    doi = "10.18653/v1/2023.acl-long.29",
    pages = "487--507"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/