Built independently by an author, for readers. Read the story and support ChapterPal

keyword

text-video retrieval

Text-video retrieval is a multimodal information retrieval task that involves finding relevant video content based on a natural language text query, or conversely, matching relevant text descriptions to a given video query. It operates by mapping both video data and textual inputs into a shared semantic embedding space, enabling algorithms to measure the similarity between visual features and language semantics. Unlike static image-text retrieval, this task requires modeling complex spatio-temporal dynamics in videos, such as motions, temporal context, and evolving scene interactions, to align them accurately with descriptive text. By bridging visual and linguistic modalities, text-video retrieval enables content-based search, filtering, and indexing across video collections without relying strictly on manual tags or metadata.

6 items

SwinBERT: End-to-End Transformers with Sparse Attention for Video Captioning

SwinBERT: End-to-End Transformers with Sparse Attention for Video Captioning

Kevin Lin, Linjie Li, Chung-Ching Lin, Faisal Ahmed, Zhe Gan, Zicheng Liu, Yumao Lu, Lijuan Wang

OrganizationsMicrosoft

Why you should read this

Presents the first pure end-to-end transformer framework for video captioning that directly processes raw video frames and utilizes a learnable sparse attention mask to reduce temporal redundancy across densely sampled inputs.

The canonical approach to video captioning dictates a caption generation model to learn from offline-extracted dense video features. These feature extractors usually operate on video frames sampled at a fixed frame rate and are often trained on image/video understanding tasks, without adaption to video captioning data. In this work, we present SWINBERT, an end-to-end transformer-based model for video captioning, which takes video frame patches directly as inputs, and outputs a natural language description. Instead of leveraging multiple 2D/3D feature extractors, our method adopts a video transformer to encode spatial-temporal representations that can adapt to variable lengths of video input without dedicated design for different frame rates. Based on this model architecture, we show that video captioning can benefit significantly from more densely sampled video frames as opposed to previous successes with sparsely sampled video frames for video-and-language understanding tasks (e.g., video question answering). Moreover, to avoid the inherent redundancy in consecutive video frames, we propose adaptively learning a sparse attention mask and optimizing it for task-specific performance improvement through better long-range video sequence modeling. Through extensive experiments on 5 video captioning datasets, we show that SWINBERT achieves across-the-board performance improvements over previous methods, often by a large margin. The learned sparse attention masks in addition push the limit to new state of the arts, and can be transferred between different video lengths and between different datasets. Code is available at https://github.com/microsoft/SwinBERT.

Added

2026-09-26

VoP: Text-Video Co-Operative Prompt Tuning for Cross-Modal Retrieval

VoP: Text-Video Co-Operative Prompt Tuning for Cross-Modal Retrieval

Siteng Huang, Biao Gong, Yulin Pan, Jianwen Jiang, Yiliang Lv, Yuyuan Li, Donglin Wang

OrganizationsAlibaba GroupWestlake UniversityZhejiang University

Why you should read this

Proposes an efficient text-video co-operative prompt tuning framework that incorporates spatio-temporal video prompts into CLIP to outperform full fine-tuning on text-video retrieval benchmarks with six times fewer parameters.

Many recent studies leverage the pre-trained CLIP for text-video cross-modal retrieval by tuning the backbone with additional heavy modules, which not only brings huge computational burdens with much more parameters, but also leads to the knowledge forgetting from upstream models. In this work, we propose the VoP: Text-Video Co-operative Prompt Tuning for efficient tuning on the text-video retrieval task. The proposed VoP is an end-to-end framework with both video & text prompts introducing, which can be regarded as a powerful baseline with only 0.1% trainable parameters. Further, based on the spatio-temporal characteristics of videos, we develop three novel video prompt mechanisms to improve the performance with different scales of trainable parameters. The basic idea of the VoP enhancement is to model the frame position, frame context, and layer function with specific trainable prompts, respectively. Extensive experiments show that compared to full fine-tuning, the enhanced VoP achieves a 1.4% average R@1 gain across five text-video retrieval benchmarks with 6× less parameter overhead. The code will be available at https://github.com/bighuang624/VoP.

Added

2026-09-26

Language Models with Image Descriptors are Strong Few-Shot Video-Language Learners

Language Models with Image Descriptors are Strong Few-Shot Video-Language Learners

Zhenhailong Wang, Manling Li, Ruochen Xu, Luowei Zhou, Jie Lei, Xudong Lin, Shuohang Wang, Ziyi Yang, Chenguang Zhu, Derek Hoiem, Shih-Fu Chang, Mohit Bansal, Heng Ji

OrganizationsColumbia UniversityGoogleMicrosoftUniversity of Illinois Urbana-ChampaignUniversity of North Carolina at Chapel Hill

Why you should read this

Proposes VidIL, a framework that decomposes video content into multi-level textual descriptions via frozen image-language models, enabling large language models to perform generative video tasks with few-shot prompting without requiring any video pretraining or finetuning.

The goal of this work is to build flexible video-language models that can generalize to various video-to-text tasks from few examples. Existing few-shot video-language learners focus exclusively on the encoder, resulting in the absence of a video-to-text decoder to handle generative tasks. Video captioners have been pretrained on large-scale video-language datasets, but they rely heavily on finetuning and lack the ability to generate text for unseen tasks in a few-shot setting. We propose VidIL, a few-shot Video-language Learner via Image and Language models, which demonstrates strong performance on few-shot video-to-text tasks without the necessity of pretraining or finetuning on any video datasets. We use image-language models to translate the video content into frame captions, object, attribute, and event phrases, and compose them into a temporal-aware template. We then instruct a language model, with a prompt containing a few in-context examples, to generate a target output from the composed content. The flexibility of prompting allows the model to capture any form of text input, such as automatic speech recognition (ASR) transcripts. Our experiments demonstrate the power of language models in understanding videos on a wide variety of video-language tasks, including video captioning, video question answering, video caption retrieval, and video future event prediction. Especially, on video future event prediction, our few-shot model significantly outperforms state-of-the-art supervised models trained on large-scale video datasets. Code and processed data are publicly available for research purposes at https://github.com/MikeWangWZHL/VidIL.

Added

2026-09-26

Bidirectional Cross-Modal Knowledge Exploration for Video Recognition with Pre-trained Vision-Language Models

Bidirectional Cross-Modal Knowledge Exploration for Video Recognition with Pre-trained Vision-Language Models

Wenhao Wu, Xiaohan Wang, Haipeng Luo, Jingdong Wang, Yi Yang, Wanli Ouyang

OrganizationsBaiduShanghai Artificial Intelligence LaboratoryUniversity of Chinese Academy of SciencesUniversity of SydneyZhejiang University

Why you should read this

Proposes a bidirectional framework called BIKE that transfers pre-trained vision-language knowledge into video recognition by retrieving complementary textual attributes and using category concepts to capture frame-level temporal saliency.

Vision-language models (VLMs) pre-trained on large-scale image-text pairs have demonstrated impressive transferability on various visual tasks. Transferring knowledge from such powerful VLMs is a promising direction for building effective video recognition models. However, current exploration in this field is still limited. We believe that the greatest value of pre-trained VLMs lies in building a bridge between visual and textual domains. In this paper, we propose a novel framework called BIKE, which utilizes the cross-modal bridge to explore bidirectional knowledge: i) We introduce the Video Attribute Association mechanism, which leverages the Video-to-Text knowledge to generate textual auxiliary attributes for complementing video recognition. ii) We also present a Temporal Concept Spotting mechanism that uses the Text-to-Video expertise to capture temporal saliency in a parameter-free manner, leading to enhanced video representation. Extensive studies on six popular video datasets, including Kinetics-400 & 600, UCF-101, HMDB-51, ActivityNet and Charades, show that our method achieves state-of-the-art performance in various recognition scenarios, such as general, zero-shot, and few-shot video recognition. Our best model achieves a state-of-the-art accuracy of 88.6% on the challenging Kinetics-400 using the released CLIP model. The code is available at https://github.com/whwu95/BIKE.

Added

2026-09-26

Revisiting Classifier: Transferring Vision-Language Models for Video Recognition

Revisiting Classifier: Transferring Vision-Language Models for Video Recognition

Wenhao Wu, Zhun Sun, Wanli Ouyang

OrganizationsBaiduShanghai Artificial Intelligence LaboratoryUniversity of Sydney

Why you should read this

Proposes replacing the traditional randomly initialized visual classifier with frozen text embeddings from pre-trained vision-language models, drastically boosting video recognition accuracy and convergence speed across zero-shot, few-shot, and fully supervised benchmarks.

Transferring knowledge from task-agnostic pre-trained deep models for downstream tasks is an important topic in computer vision research. Along with the growth of computational capacity, we now have open-source vision-language pre-trained models in large scales of the model architecture and amount of data. In this study, we focus on transferring knowledge for video classification tasks. Conventional methods randomly initialize the linear classifier head for vision classification, but they leave the usage of the text encoder for downstream visual recognition tasks undiscovered. In this paper, we revise the role of the linear classifier and replace the classifier with different knowledge from the pre-trained model. We utilize the well-pre-trained language model to generate a good semantic target for efficient transferring learning. The empirical study shows that our method improves both the performance and the training speed of video classification, with a negligible change in the model. Our simple yet effective tuning paradigm achieves state-of-the-art performance and efficient training on various video recognition scenarios, i.e., zero-shot, few-shot, and general recognition. In particular, our paradigm achieves the state-of-the-art accuracy of 87.8% on Kinetics-400, and also surpasses previous methods by 20~50% absolute top-1 accuracy under zero-shot, few-shot settings on five video datasets. Code and models are available at https://github.com/whwu95/Text4Vis.

Added

2026-09-26

Text Is MASS: Modeling as Stochastic Embedding for Text-Video Retrieval

Text Is MASS: Modeling as Stochastic Embedding for Text-Video Retrieval

Jiamian Wang, Pichao Wang, Guohao Sun, Dongfang Liu, Sohail A. Dianat, Raghuveer Rao, Majid Rabbani, Zhiqiang Tao

OrganizationsAmazonRochester Institute of TechnologyUnited States Army Research Laboratory

Why you should read this

Proposes T-MASS, a text-video retrieval framework that models concise text queries as stochastic embeddings with adaptive radii to capture the broad semantic scope of rich video content and achieve state-of-the-art retrieval accuracy across multiple benchmark datasets.

The increasing prevalence of video clips has sparked growing interest in text-video retrieval. Recent advances focus on establishing a joint embedding space for text and video, relying on consistent embedding representations to compute similarity. However, the text content in existing datasets is generally short and concise, making it hard to fully describe the redundant semantics of a video. Correspondingly, a single text embedding may be less expressive to capture the video embedding and empower the retrieval. In this study, we propose a new stochastic text modeling method T-MASS, i.e., text is modeled as a stochastic embedding, to enrich text embedding with a flexible and resilient semantic range, yielding a text mass. To be specific, we introduce a similarity-aware radius module to adapt the scale of the text mass upon the given text-video pairs. Plus, we design and develop a support text regularization to further control the text mass during the training. The inference pipeline is also tailored to fully exploit the text mass for accurate retrieval. Empirical evidence suggests that T-MASS not only effectively attracts relevant text-video pairs while distancing irrelevant ones, but also enables the determination of precise text embeddings for relevant pairs. Our experimental results show a substantial improvement of T-MASS over baseline (3% ∼ 6.3% by R@1). Also, T-MASS achieves state-of-the-art performance on five benchmark datasets, including MSRVTT, LSMDC, DiDeMo, VATEX, and Charades. Code and models are available here.

Added

2026-09-26