Language Models with Image Descriptors are Strong Few-Shot Video-Language Learners

Zhenhailong WangManling LiRuochen XuLuowei ZhouJie LeiXudong LinShuohang WangZiyi YangChenguang ZhuDerek Hoiem

article2022NeurIPS175 citations

Proposes VidIL, a framework that decomposes video content into multi-level textual descriptions via frozen image-language models, enabling large language models to perform generative video tasks with few-shot prompting without requiring any video pretraining or finetuning.

Listen

Artificial intelligence systems often struggle to generalize to new video understanding tasks when provided with only a few annotated examples. Current video-language models typically focus solely on encoding visual data without generating text, or they rely heavily on computationally expensive pretraining and finetuning over millions of video-text pairs. Moreover, conventional video models struggle to bridge the gap between noisy speech transcripts and rich visual scenes across time.

The article evaluates a modular framework, named VidIL, to demonstrate that pretrained large language models can perform diverse video-to-text generative tasks using only a few examples, completely eliminating the need for video-specific pretraining or finetuning.

The approach converts video content into a structured, unified textual format across three hierarchical tiers: visual tokens (identifying objects, events, and attributes via image encoders), frame-level captions, and video-level summaries. These elements are arranged chronologically using temporal transition markers (such as "First," "Then," and "Finally") and combined with optional speech transcripts. A frozen language model receives this structured text alongside a small set of dynamically selected in-context examples to generate outputs across multiple benchmarks, including open-domain and instructional video datasets.

Key findings demonstrate the effectiveness and efficiency of this approach:

  • In video future event prediction, the framework achieved 72.0% accuracy with only 10 labeled examples, outperforming fully supervised models trained on over 20,000 video instances (which achieved 68.4%).
  • In few-shot video question answering, the 5-shot model achieved 21.2% accuracy on MSR-VTT and 39.1% on MSVD, surpassing zero-shot baselines by large margins and outperforming heavy models like Flamingo-3B across multiple shot settings.
  • In video captioning, incorporating speech transcripts lifted instructional captioning quality from a 27.0 baseline score to 111.6, demonstrating robust cross-domain flexibility.
  • In semi-supervised text-video retrieval, using the model to generate synthetic labels on unlabeled videos improved retrieval recall across open-domain benchmarks, achieving results comparable to models trained on fully human-annotated data.

These results show that organizations can bypass expensive, specialized video pretraining pipelines by chaining off-the-shelf image models with general-purpose language models. This substantially cuts training compute costs, accelerates deployment timelines, and simplifies the integration of multi-modal streams like audio and text transcripts into existing workflows.

Organizations developing video analytics should adopt modular prompting architectures when tackling low-resource or rapid-adaptation video tasks. Teams should also implement dynamic in-context example selection rather than static prompting to maximize output accuracy. Where large pools of unlabeled video exist, teams can employ this framework as an automated pseudo-labeling tool to bootstrap downstream retrieval models.

Confidence in these findings is strong across generative, question answering, and retrieval benchmarks. However, leaders should note that converting visual data into pure text inevitably discards fine-grained spatial and low-level visual details, making this approach less suitable for specialized spatial localization tasks. Furthermore, because the framework relies on large language models trained on massive internet data, organizations must implement safeguards to monitor and mitigate potential social biases in generated outputs.

arXiv: 2205.10747
Cover for Language Models with Image Descriptors are Strong Few-Shot Video-Language Learners

Abstract

The goal of this work is to build flexible video-language models that can generalize to various video-to-text tasks from few examples. Existing few-shot video-language learners focus exclusively on the encoder, resulting in the absence of a video-to-text decoder to handle generative tasks. Video captioners have been pretrained on large-scale video-language datasets, but they rely heavily on finetuning and lack the ability to generate text for unseen tasks in a few-shot setting. We propose VidIL, a few-shot Video-language Learner via Image and Language models, which demonstrates strong performance on few-shot video-to-text tasks without the necessity of pretraining or finetuning on any video datasets. We use image-language models to translate the video content into frame captions, object, attribute, and event phrases, and compose them into a temporal-aware template. We then instruct a language model, with a prompt containing a few in-context examples, to generate a target output from the composed content. The flexibility of prompting allows the model to capture any form of text input, such as automatic speech recognition (ASR) transcripts. Our experiments demonstrate the power of language models in understanding videos on a wide variety of video-language tasks, including video captioning, video question answering, video caption retrieval, and video future event prediction. Especially, on video future event prediction, our few-shot model significantly outperforms state-of-the-art supervised models trained on large-scale video datasets. Code and processed data are publicly available for research purposes at https://github.com/MikeWangWZHL/VidIL.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Image-Language Models and Their Applications on Video-Language Tasks
  • 2.2 Unifying MultiModal Tasks with Language Models
  • 3 Method
  • 3.1 Frame Level: Image Captioning
  • 3.2 Visual Token Level: Structure-Aware Visual Tokenization
  • 3.3 Video Level: Temporal-Aware Few-shot Prompting
  • 4 Experiments
  • 4.1 Experimental Setup
  • 4.2 Few-shot Video Captioning
  • 4.3 Few-shot Video Question Answering
  • 4.4 Few-shot Video-Language Event Prediction
  • 4.5 Semi-supervised Text-Video Retrieval
  • 4.6 Ablation Studies
  • 5 Conclusions, Limitations and Future Work
  • 6 Broader Impact
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — VidIL Hierarchical Video Representation Framework

    model/method

    VidIL (Video-language Learner via Image and Language models) is a few-shot video-to-text learning framework that operates without pretraining or finetuning on video datasets. The method decomposes a video into three hierarchical levels of unified textual representations:

    1. Visual Token Level: Salient fine-grained semantic entities—comprising objects, events, and visual attributes—are extracted for each sampled video frame using a contrastive multimodal retrieval model (such as CLIP-L/14) against structured text vocabularies.
    2. Frame Level: High-level visual semantics are captured by generating and filtering natural language frame descriptions using a pretrained image captioning model (such as BLIP fine-tuned on COCO).
    3. Video Level: The visual tokens, frame captions, and optional auxiliary textual modalities (such as Automatic Speech Recognition transcripts) are composed into a structured, temporal-aware text prompt. The prompt is prefixed with natural language transition markers (such as "First... Then... Finally...") and passed along with in-context demonstration examples to a frozen large language model (such as InstructGPT) to generate the target video-level text output.
  2. Knowl 2 — Retrieval-Based Structure-Aware Visual Tokenization

    model/method

    Rather than relying on closed classification vocabularies, VidIL tokenizes sparsely sampled video frames into textual descriptors via multimodal retrieval across three vocabularies:

    • Object Vocabulary: Comprises approximately 20,000 full classes from OpenImages.
    • Event Vocabulary: Constructed by parsing Visual Genome object synsets using Semantic Role Labeling (SRL) to select phrases containing at least one verb and one argument, followed by Sentence-BERT semantic deduplication to eliminate near-identical event phrases.
    • Attribute Vocabulary: Derived from Visual Genome attribute synsets.

    For each sampled frame ff, visual embedding vf\mathbf{v}_f is computed alongside textual embeddings et\mathbf{e}_t for candidate tokens tt in a vocabulary using a pretrained CLIP-L/14 encoder. The top 5 visual tokens per frame are selected according to cosine similarity:

    sim(vf,et)=vf⋅et∥vf∥∥et∥\text{sim}(\mathbf{v}_f, \mathbf{e}_t) = \frac{\mathbf{v}_f \cdot \mathbf{e}_t}{\|\mathbf{v}_f\| \|\mathbf{e}_t\|}

  3. Knowl 3 — Temporal-Aware Few-Shot Prompting and Generation Formulation

    equation

    Given a task instruction ss, a few-shot context cc containing demonstration examples and the structured video representation of the test instance, and a task query suffix qq designating the target format, the target textual sequence y=(y1,y2,…,yL)y = (y_1, y_2, \dots, y_L) of length LL is generated auto-regressively by a frozen language model according to:

    yl=arg⁡max⁡yp(y∣s,c,q,y<l)y_l = \arg\max_{y} p(y \mid s, c, q, y_{<l})

    where y<l=(y1,…,yl−1)y_{<l} = (y_1, \dots, y_{l-1}) denotes the sequence of tokens generated prior to step ll.

    To encode the temporal evolution of visual states, each visual token set and frame caption in context cc is prefixed with temporal transition markers (such as "First,", "Then,", "After that,", "Finally,") ordered by frame timestamp index. This natural language ordering prompts the language model to distinguish temporal dynamics (e.g., distinguishing a sunrise from a sunset based on event sequencing).

  4. Knowl 4 — Dynamic In-Context Example Selection and Ordering

    algorithm

    To select prompt demonstrations that are semantically relevant to a query and mitigate recency bias in frozen autoregressive language models, VidIL dynamically filters and sorts in-context examples from a support set:

    Input: Candidate support set S of M labeled video-text pairs {(V_i, Y_i)},
           Test instance representation V_test,
           Test query text Q_test (question text for QA, or frame captions for captioning),
           Number of demonstration shots N (where N <= M),
           Pretrained sentence encoder SBERT
    Output: Ordered in-context prompt context C
    for each (V_i, Y_i) in S do
        Extract reference text Q_i from (V_i, Y_i) (question text or frame captions)
        sim_i = CosineSimilarity(SBERT(Q_i), SBERT(Q_test))
    end for
    Select subset S_sub of size N from S containing the pairs with top N highest sim_i values
    Sort S_sub in ascending order of sim_i: S_ordered = [(V_(1), Y_(1)), ..., (V_(N), Y_(N))] such that sim_(1) <= ... <= sim_(N)
    C = EmptyPrompt()
    for each (V_(j), Y_(j)) in S_ordered do
        C = AppendExample(C, V_(j), Y_(j))
    end for
    C = AppendTestInstance(C, V_test)
    return C
  5. Knowl 5 — Few-Shot Video Captioning Benchmark Results

    data/table

    In 10-shot video captioning evaluated across open-domain datasets (MSR-VTT, VaTeX) and instructional videos (YouCook2), VidIL outperforms both video-pretrained (UniVL) and image-pretrained (BLIP) baselines on average CIDEr score without video pretraining or fine-tuning.

    Method #VideoPT ASR MSR-VTT YouCook2 VaTex Avg C
    B-4 M C B-4 M C B-4 M C
    UniVL 1.2M No 2.1 9.5 3.6 3.3 11.6 34.1 1.7 8.0 2.1 13.3
    BLIP 0 No 27.7 23.0 39.5 0.7 3.4 11.5 13.5 15.4 20.7 23.9
    BLIPcap 0 No 21.6 22.7 30.2 3.7 3.8 9.4 20.7 17.4 28.9 22.8
    VidIL (ours) 0 No 26.0 24.7 36.3 2.6 9.5 27.0 22.2 20.0 36.7 33.3
    UniVL 1.2M Yes - - - 4.3 12.2 48.6 2.7 10.2 3.4 26.0
    VidIL (ours) 0 Yes - - - 10.7 19.4 111.6 23.2 20.6 38.9 75.3

    Metrics reported are BLEU-4 (B-4), METEOR (M), and CIDEr (C). Avg C is the average CIDEr across benchmarks. #VideoPT indicates the number of videos used in multimodal pretraining. Integrating ASR subtitles into the textual prompt raises VidIL's CIDEr score on YouCook2 from 27.0 to 111.6.

  6. Knowl 6 — Few-Shot Video-Language Future Event Prediction on VLEP

    empirical result

    On the Video-Language Future Event Prediction (VLEP) benchmark, the task requires predicting which of two candidate future events is more likely to follow a given video and dialogue context. VidIL formulates this as conditional generation: generating free-form continuation text and mapping it to the candidate event choices using Sentence-BERT embedding similarity.

    On the hidden VLEP test set:

    • VLEP Baseline (supervised with 20,142 labeled videos): 67.5% accuracy.
    • MERLOT (supervised with 20,142 labeled videos): 68.4% accuracy.
    • VidIL (10-shot) (frozen image/language models, 0 video pretraining): 72.0% accuracy.
    • Human Performance: 90.5% accuracy.

    With only 10 labeled examples, 10-shot VidIL outperforms fully supervised state-of-the-art models trained on over 20,000 video instances by ~3.6-4.5% absolute accuracy.

  7. Knowl 7 — Few-Shot Video Question Answering Evaluation

    data/table

    VidIL evaluated with 5 in-context demonstration shots on video question answering benchmarks MSR-VTT_QA and MSVD_QA outperforms zero-shot/few-shot image-language baselines and achieves performance competitive with large-scale video-pretrained models.

    Method #videoPT #videoFT MSR-VTT Acc (%) MSVD Acc (%)
    BLIP 0 0-shot 0.55 0.45
    BLIP 0 5-shot 0.84 0.53
    BLIP_VQA 0 0-shot 19.2 35.2
    VidIL (ours) 0 5-shot 21.2 39.1
    Flamingo-3B 27M 4-shot 14.9 33.0
    Flamingo-3B 27M 8-shot 19.6 37.0
    Flamingo-80B 27M 4-shot 23.9 41.7
    Flamingo-80B 27M 8-shot 27.6 45.5
    ALPRO 2M full-shot 42.1 45.9

    VidIL with 5 shots (using InstructGPT and 0 video pretraining) achieves 21.2% on MSR-VTT and 39.1% on MSVD, outperforming 8-shot Flamingo-3B (pretrained on 27 million video-text pairs) and zero-shot BLIP_VQA.

  8. Knowl 8 — Semi-Supervised Text-Video Retrieval via Few-Shot Pseudo-Labeling

    data/table

    In a semi-supervised retrieval scenario with 10 labeled training videos and thousands of unlabeled videos, VidIL serves as a few-shot video captioner to generate pseudo-captions for unlabeled training videos. A standard dual-encoder base model (BLIP) is then finetuned on the pseudo-labeled corpora.

    Pseudo Label Source MSR-VTT (10 / 7010) VaTex (10 / 22685)
    t_R1 t_R5 v_R1 v_R5 t_R1 t_R5 v_R1 v_R5
    Zero-shot BLIP (no pseudo) 33.2 57.2 40.5 62.8 28.2 53.4 34.0 58.6
    UniVL pseudo labels 33.1 57.3 33.6 57.7 25.5 47.7 26.1 49.1
    BLIP pseudo labels 35.6 60.8 39.8 60.4 26.3 50.5 29.3 53.6
    BLIPcap pseudo labels 35.3 58.0 39.1 63.3 23.9 46.8 27.5 49.7
    VidIL pseudo labels (ours) 39.6 64.5 40.8 65.2 33.3 59.1 33.7 59.5
    Ground Truth (Full data) 43.6 66.2 43.1 67.2 40.1 66.4 40.1 66.6

    Columns t_R1/t_R5 denote video-to-text Recall@1 and Recall@5; v_R1/v_R5 denote text-to-video Recall@1 and Recall@5. Finetuning on VidIL-generated pseudo-labels yields substantial gains over zero-shot BLIP and other pseudo-labeling sources, approaching the performance of full ground-truth fine-tuning on MSR-VTT Recall@5 (65.2 vs. 67.2).

  9. Knowl 9 — Ablation of Video Representation Components and In-Context Shots

    data/table

    Ablation experiments conducted on the MSVD_QA validation set illustrate the contributions of visual token granularities, temporal sampling, and in-context selection strategies (mean accuracy and standard deviation across 3 random shot samples):

    Video Representation Setting Accuracy (% Avg ±\pm Std)
    Frame captions only 39.6 ±\pm 3.7
    Frame + Objects 40.3 ±\pm 2.9
    Frame + Objects + Events 39.9 ±\pm 2.8
    Frame + Objects + Attributes 40.9 ±\pm 2.9
    Frame + Objects + Events + Attributes 40.8 ±\pm 2.4
    Reduced to single frame 38.5 ±\pm 2.4
    Reversed temporal order 40.7 ±\pm 1.7

    Key observations:

    1. Incorporating multi-level visual tokens (objects, events, attributes) increases accuracy from 39.6% to 40.8% and reduces variance from 3.7 to 2.4.
    2. Reducing temporal coverage to a single middle frame degrades accuracy to 38.5%.
    3. Dynamic in-context selection with N=5N=5 examples chosen from larger pools stabilizes performance across shots (at 30 shots: 41.1 ±\pm 1.9 with selection vs. 40.0 ±\pm 2.9 without selection).
  10. Knowl 10 — Limitations of Symbolic Textual Representations and Benchmark Temporal Sensitivity

    limitation

    The VidIL framework exhibits two primary limitations:

    1. Loss of Low-Level Spatial Features: Converting video frames exclusively into symbolic textual descriptions (captions, object labels, event phrases, and attributes) discards continuous, pixel-level spatial information. This loss of fine-grained spatial coordinates can impair performance on tasks requiring spatial visual reasoning (such as localized spatial visual question answering).
    2. Benchmark Insensitivity to Temporal Dynamics: Reversing the temporal ordering of visual tokens and frame captions in the prompt resulted in only a marginal decrease in MSVD_QA performance (from 40.8% to 40.7%), indicating that current standard video QA benchmarks rely predominantly on static visual semantics rather than fine-grained temporal tracking across frames.

Coverage note — None was omitted; all contributed methods, algorithmic formulations, prompt design principles, empirical benchmark evaluations across captioning, QA, VLEP, semi-supervised retrieval, ablations, and identified limitations are fully covered.

References

  1. 1.Hassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang, Shih-Fu Chang, Yin Cui, and Boqing Gong. Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text. Advances in Neural Information Processing Systems, 34, 2021.
  2. 2.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. ArXiv preprint, abs/2204.14198, 2022.
  3. 3.Jean-Baptiste Alayrac, Adrià Recasens, Rosalia Schneider, Relja Arandjelovic, Jason Ramapuram, Jeffrey De Fauw, Lucas Smaira, Sander Dieleman, and Andrew Zisserman. Self-supervised multimodal versatile networks. In Hugo Larochelle, Marc'Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin, editors, Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, 2020.
  4. 4.Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. VQA: visual question answering. In 2015 IEEE International Conference on Computer Vision, ICCV 2015, Santiago, Chile, December 7-13, 2015, pages 2425–2433. IEEE Computer Society, 2015.
  5. 5.Satanjeev Banerjee and Alon Lavie. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, pages 65–72, Ann Arbor, Michigan, 2005. Association for Computational Linguistics.
  6. 6.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. In Hugo Larochelle, Marc'Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin, editors, Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, 2020.
  7. 7.David Chen and William Dolan. Collecting highly parallel data for paraphrase evaluation. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, pages 190–200, Portland, Oregon, USA, 2011. Association for Computational Linguistics.
  8. 8.Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. Uniter: Universal image-text representation learning. In European conference on computer vision, pages 104–120. Springer, 2020.
  9. 9.Jaemin Cho, Jie Lei, Hao Tan, and Mohit Bansal. Unifying vision-and-language tasks via text generation. In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research, pages 1931–1942. PMLR, 2021.
  10. 10.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Fei-Fei Li. Imagenet: A large-scale hierarchical image database. In 2009 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR 2009), 20-25 June 2009, Miami, Florida, USA, pages 248–255. IEEE Computer Society, 2009.
  11. 11.Han Fang, Pengfei Xiong, Luhui Xu, and Yu Chen. Clip2video: Mastering video-text retrieval via image clip. ArXiv preprint, abs/2106.11097, 2021.
  12. 12.Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He. Slowfast networks for video recognition. In 2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019, Seoul, Korea (South), October 27 - November 2, 2019, pages 6201–6210. IEEE, 2019.
  13. 13.Tianyu Gao, Adam Fisch, and Danqi Chen. Making pre-trained language models better few-shot learners. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3816–3830, Online, 2021. Association for Computational Linguistics.
  14. 14.Matt Gardner, Joel Grus, Mark Neumann, Oyvind Tafjord, Pradeep Dasigi, Nelson F. Liu, Matthew Peters, Michael Schmitz, and Luke Zettlemoyer. AllenNLP: A deep semantic natural language processing platform. In Proceedings of Workshop for NLP Open Source Software (NLP-OSS), pages 1–6, Melbourne, Australia, 2018. Association for Computational Linguistics.
  15. 15.Ting-Hao Kenneth Huang, Francis Ferraro, Nasrin Mostafazadeh, Ishan Misra, Aishwarya Agrawal, Jacob Devlin, Ross Girshick, Xiaodong He, Pushmeet Kohli, Dhruv Batra, C. Lawrence Zitnick, Devi Parikh, Lucy Vanderwende, Michel Galley, and Margaret Mitchell. Visual storytelling. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1233–1239, San Diego, California, 2016. Association for Computational Linguistics.
  16. 16.Zhicheng Huang, Zhaoyang Zeng, Yupan Huang, Bei Liu, Dongmei Fu, and Jianlong Fu. Seeing out of the box: End-to-end pre-training for vision-language representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12976–12985, 2021.
  17. 17.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V. Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research, pages 4904–4916. PMLR, 2021.
  18. 18.Wonjae Kim, Bokyung Son, and Ildoo Kim. Vilt: Vision-and-language transformer without convolution or region supervision. In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research, pages 5583–5594. PMLR, 2021.
  19. 19.Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations. International journal of computer vision, 123(1):32–73, 2017.
  20. 20.Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper Uijlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci, Alexander Kolesnikov, et al. The open images dataset v4. International Journal of Computer Vision, 128(7):1956–1981, 2020.
  21. 21.Jie Lei, Linjie Li, Luowei Zhou, Zhe Gan, Tamara L. Berg, Mohit Bansal, and Jingjing Liu. Less is more: Clipbert for video-and-language learningvia sparse sampling. In CVPR, 2021.
  22. 22.Jie Lei, Licheng Yu, Mohit Bansal, and Tamara Berg. TVQA: Localized, compositional video question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 1369–1379, Brussels, Belgium, 2018. Association for Computational Linguistics.
  23. 23.Jie Lei, Licheng Yu, Tamara Berg, and Mohit Bansal. What is more likely to happen next? video-and-language future event prediction. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 8769–8784, Online, 2020. Association for Computational Linguistics.
  24. 24.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online, 2020. Association for Computational Linguistics.
  25. 25.Dongxu Li, Junnan Li, Hongdong Li, Juan Carlos Niebles, and Steven C.H. Hoi. Align and prompt: Video-and-language pre-training with entity prompts. In arxiv, 2021.
  26. 26.Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. ArXiv preprint, abs/2201.12086, 2022.
  27. 27.Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi. Align before fuse: Vision and language representation learning with momentum distillation. Advances in Neural Information Processing Systems, 34, 2021.
  28. 28.Linjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan, Licheng Yu, and Jingjing Liu. HERO: Hierarchical encoder for Video+Language omni-representation pre-training. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2046–2065, Online, 2020. Association for Computational Linguistics.
  29. 29.Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, et al. Oscar: Object-semantics aligned pre-training for vision-language tasks. In European Conference on Computer Vision, pages 121–137. Springer, 2020.
  30. 30.Chin-Yew Lin and Franz Josef Och. Automatic evaluation of machine translation quality using longest common subsequence and skip-bigram statistics. In Proceedings of the 42nd Annual Meeting of the Association for Computational Linguistics (ACL-04), pages 605–612, Barcelona, Spain, 2004.
  31. 31.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755. Springer, 2014.
  32. 32.Xudong Lin, Gedas Bertasius, Jue Wang, Shih-Fu Chang, Devi Parikh, and Lorenzo Torresani. Vx2text: End-to-end learning of video-based text generation from multimodal inputs. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 7001–7011. IEEE, 2021.
  33. 33.Xudong Lin, Fabio Petroni, Gedas Bertasius, Marcus Rohrbach, Shih-Fu Chang, and Lorenzo Torresani. Learning to recognize procedural activities with distant supervision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13853–13863, 2022.
  34. 34.Xudong Lin, Simran Tiwari, Shiyuan Huang, Manling Li, Mike Zheng Shou, Heng Ji, and Shih-Fu Chang. Towards fast adaptation of pretrained contrastive models for multi-channel video-language retrieval. arXiv preprint arXiv:2206.02082, 2022.
  35. 35.Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In Hanna M. Wallach, Hugo Larochelle, Alina Beygelzimer, Florence d'Alché-Buc, Emily B. Fox, and Roman Garnett, editors, Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pages 13–23, 2019.
  36. 36.Huaishao Luo, Lei Ji, Botian Shi, Haoyang Huang, Nan Duan, Tianrui Li, Jason Li, Taroon Bharti, and Ming Zhou. Univl: A unified video and language pre-training model for multimodal understanding and generation. ArXiv preprint, abs/2002.06353, 2020.
  37. 37.Huaishao Luo, Lei Ji, Ming Zhong, Yang Chen, Wen Lei, Nan Duan, and Tianrui Li. CLIP4Clip: An empirical study of clip for end to end video clip retrieval. ArXiv preprint, abs/2104.08860, 2021.
  38. 38.Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic, and Andrew Zisserman. End-to-end learning of visual representations from uncurated instructional videos. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, pages 9876–9886. IEEE, 2020.
  39. 39.Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic. Howto100m: Learning a text-video embedding by watching hundred million narrated video clips. In 2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019, Seoul, Korea (South), October 27 - November 2, 2019, pages 2630–2640. IEEE, 2019.
  40. 40.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. ArXiv preprint, abs/2203.02155, 2022.
  41. 41.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311–318, Philadelphia, Pennsylvania, USA, 2002. Association for Computational Linguistics.
  42. 42.Mandela Patrick, Po-Yao Huang, Yuki Markus Asano, Florian Metze, Alexander G. Hauptmann, João F. Henriques, and Andrea Vedaldi. Support-set bottlenecks for video-text representation learning. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net, 2021.
  43. 43.Baolin Peng, Chenguang Zhu, Chunyuan Li, Xiujun Li, Jinchao Li, Michael Zeng, and Jianfeng Gao. Few-shot natural language generation for task-oriented dialog. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 172–182, Online, 2020. Association for Computational Linguistics.
  44. 44.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research, pages 8748–8763. PMLR, 2021.
  45. 45.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. ArXiv preprint, abs/1910.10683, 2019.
  46. 46.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21:1–67, 2020.
  47. 47.Nils Reimers and Iryna Gurevych. Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3982–3992, Hong Kong, China, 2019. Association for Computational Linguistics.
  48. 48.Timo Schick and Hinrich Schütze. Exploiting cloze-questions for few-shot text classification and natural language inference. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 255–269, Online, 2021. Association for Computational Linguistics.
  49. 49.Paul Hongsuck Seo, Arsha Nagrani, Anurag Arnab, and Cordelia Schmid. End-to-end generative pretraining for multimodal video captioning. ArXiv preprint, abs/2201.08264, 2022.
  50. 50.Yixuan Su, Tian Lan, Yahui Liu, Fangyu Liu, Dani Yogatama, Yan Wang, Lingpeng Kong, and Nigel Collier. Language models can see: Plugging visual controls in text generation. ArXiv preprint, abs/2205.02655, 2022.
  51. 51.Derek Tam, Rakesh R. Menon, Mohit Bansal, Shashank Srivastava, and Colin Raffel. Improving and simplifying pattern exploiting training. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 4980–4991, Online and Punta Cana, Dominican Republic, 2021. Association for Computational Linguistics.
  52. 52.Hao Tan and Mohit Bansal. LXMERT: Learning cross-modality encoder representations from transformers. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 5100–5111, Hong Kong, China, 2019. Association for Computational Linguistics.
  53. 53.Maria Tsimpoukelli, Jacob L Menick, Serkan Cabi, SM Eslami, Oriol Vinyals, and Felix Hill. Multimodal few-shot learning with frozen language models. Advances in Neural Information Processing Systems, 34:200–212, 2021.
  54. 54.Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. Cider: Consensus-based image description evaluation. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2015, Boston, MA, USA, June 7-12, 2015, pages 4566–4575. IEEE Computer Society, 2015.
  55. 55.Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang. Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework. ArXiv preprint, abs/2202.03052, 2022.
  56. 56.Qiang Wang, Yanhao Zhang, Yun Zheng, Pan Pan, and Xian-Sheng Hua. Disentangled representation learning for text-video retrieval. ArXiv preprint, abs/2203.07111, 2022.
  57. 57.Xin Wang, Jiawei Wu, Junkun Chen, Lei Li, Yuan-Fang Wang, and William Yang Wang. Vatex: A large-scale, high-quality multilingual dataset for video-and-language research. In 2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019, Seoul, Korea (South), October 27 - November 2, 2019, pages 4580–4590. IEEE, 2019.
  58. 58.Zirui Wang, Jiahui Yu, Adams Wei Yu, Zihang Dai, Yulia Tsvetkov, and Yuan Cao. Simvlm: Simple visual language model pretraining with weak supervision. ArXiv preprint, abs/2108.10904, 2021.
  59. 59.Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. Finetuned language models are zero-shot learners. ArXiv preprint, abs/2109.01652, 2021.
  60. 60.Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang. Video question answering via gradually refined attention over appearance and motion. In Proceedings of the 2017 ACM on Multimedia Conference, MM 2017, Mountain View, CA, USA, October 23-27, 2017, pages 1645–1653, 2017.
  61. 61.Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko, Armen Aghajanyan, Florian Metze, Luke Zettlemoyer, and Christoph Feichtenhofer. VideoCLIP: Contrastive pre-training for zero-shot video-text understanding. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6787–6800, Online and Punta Cana, Dominican Republic, 2021. Association for Computational Linguistics.
  62. 62.Jun Xu, Tao Mei, Ting Yao, and Yong Rui. MSR-VTT: A large video description dataset for bridging video and language. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016, pages 5288–5296. IEEE Computer Society, 2016.
  63. 63.Zhengyuan Yang, Zhe Gan, Jianfeng Wang, Xiaowei Hu, Yumao Lu, Zicheng Liu, and Lijuan Wang. An empirical study of gpt-3 for few-shot knowledge-based vqa. ArXiv preprint, abs/2109.05014, 2021.
  64. 64.Ziyi Yang, Yuwei Fang, Chenguang Zhu, Reid Pryzant, Dongdong Chen, Yu Shi, Yichong Xu, Yao Qian, Mei Gao, Yi-Ling Chen, et al. i-code: An integrative and composable multimodal learning framework. ArXiv preprint, abs/2205.01818, 2022.
  65. 65.Fei Yu, Jiji Tang, Weichong Yin, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang. Ernie-vil: Knowledge enhanced vision-language representations through scene graphs. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 3208–3216, 2021.
  66. 66.Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, et al. Florence: A new foundation model for computer vision. ArXiv preprint, abs/2111.11432, 2021.
  67. 67.Rowan Zellers, Jiasen Lu, Ximing Lu, Youngjae Yu, Yanpeng Zhao, Mohammadreza Salehi, Aditya Kusupati, Jack Hessel, Ali Farhadi, and Yejin Choi. Merlot reserve: Neural script knowledge through vision and language and sound. ArXiv preprint, abs/2201.02639, 2022.
  68. 68.Rowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu, Jae Sung Park, Jize Cao, Ali Farhadi, and Yejin Choi. Merlot: Multimodal neural script knowledge models. Advances in Neural Information Processing Systems, 34, 2021.
  69. 69.Andy Zeng, Adrian Wong, Stefan Welker, Krzysztof Choromanski, Federico Tombari, Aveek Purohit, Michael Ryoo, Vikas Sindhwani, Johnny Lee, Vincent Vanhoucke, et al. Socratic models: Composing zero-shot multimodal reasoning with language. ArXiv preprint, abs/2204.00598, 2022.
  70. 70.Pengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang, Lei Zhang, Lijuan Wang, Yejin Choi, and Jianfeng Gao. Vinvl: Revisiting visual representations in vision-language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5579–5588, 2021.
  71. 71.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. Opt: Open pre-trained transformer language models. ArXiv preprint, abs/2205.01068, 2022.
  72. 72.Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. Calibrate before use: Improving few-shot performance of language models. In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research, pages 12697–12706. PMLR, 2021.
  73. 73.Luowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu, Jason J. Corso, and Jianfeng Gao. Unified vision-language pre-training for image captioning and VQA. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2020, New York, NY, USA, February 7-12, 2020, pages 13041–13049. AAAI Press, 2020.
  74. 74.Luowei Zhou, Chenliang Xu, and Jason J. Corso. Towards automatic learning of procedures from web instructional videos. In Sheila A. McIlraith and Kilian Q. Weinberger, editors, Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, (AAAI-18), the 30th innovative Applications of Artificial Intelligence (IAAI-18), and the 8th AAAI Symposium on Educational Advances in Artificial Intelligence (EAAI-18), New Orleans, Louisiana, USA, February 2-7, 2018, pages 7590–7598. AAAI Press, 2018.
  75. 75.Xizhou Zhu, Jinguo Zhu, Hao Li, Xiaoshi Wu, Xiaogang Wang, Hongsheng Li, Xiaohua Wang, and Jifeng Dai. Uni-perceiver: Pre-training unified architecture for generic perception for zero-shot and few-shot tasks. ArXiv preprint, abs/2112.01522, 2021.

Citation

MLA
Wang, Z., et al. “Language Models with Image Descriptors Are Strong Few-Shot Video-Language Learners”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 8483–97, https://proceedings.neurips.cc/paper_files/paper/2022/file/381ceeae4a1feb1abc59c773f7e61839-Paper-Conference.pdf.
APA
Wang, Z., Li, M., Xu, R., Zhou, L., Lei, J., Lin, X., Wang, S., Yang, Z., Zhu, C., Hoiem, D., Chang, S.-F., Bansal, M., & Ji, H. (2022). Language Models with Image Descriptors are Strong Few-Shot Video-Language Learners. Advances in Neural Information Processing Systems, 35, 8483–8497. https://proceedings.neurips.cc/paper_files/paper/2022/file/381ceeae4a1feb1abc59c773f7e61839-Paper-Conference.pdf
Chicago
Wang, Z., M. Li, R. Xu, et al. 2022. “Language Models with Image Descriptors Are Strong Few-Shot Video-Language Learners”. Advances in Neural Information Processing Systems 35: 8483–97. https://proceedings.neurips.cc/paper_files/paper/2022/file/381ceeae4a1feb1abc59c773f7e61839-Paper-Conference.pdf.
Harvard
Wang, Z. et al. (2022) “Language Models with Image Descriptors are Strong Few-Shot Video-Language Learners”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 8483–8497. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/381ceeae4a1feb1abc59c773f7e61839-Paper-Conference.pdf.
Vancouver
1. Wang Z, Li M, Xu R, et al (2022) Language Models with Image Descriptors are Strong Few-Shot Video-Language Learners. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 8483–8497

BibTeX

@inproceedings{wang2022language,
  title = {Language Models with Image Descriptors are Strong Few-Shot Video-Language Learners},
  author = {Wang, Zhenhailong and Li, Manling and Xu, Ruochen and Zhou, Luowei and Lei, Jie and Lin, Xudong and Wang, Shuohang and Yang, Ziyi and Zhu, Chenguang and Hoiem, Derek and Chang, Shih-Fu and Bansal, Mohit and Ji, Heng},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {8483-8497},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/381ceeae4a1feb1abc59c773f7e61839-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission