VoP: Text-Video Co-Operative Prompt Tuning for Cross-Modal Retrieval
Siteng HuangBiao GongYulin PanJianwen JiangYiliang LvYuyuan LiDonglin Wang
Proposes an efficient text-video co-operative prompt tuning framework that incorporates spatio-temporal video prompts into CLIP to outperform full fine-tuning on text-video retrieval benchmarks with six times fewer parameters.
Adapting large foundation models to cross-modal text-video retrieval typically relies on full fine-tuning or adding heavy architectural components. These standard approaches present significant operational challenges: updating all underlying parameters creates high computational overhead, risks catastrophic forgetting of pre-trained knowledge, and requires storing massive, separate model copies for every downstream deployment.
The article introduces and evaluates "Text-Video Co-operative Prompt Tuning" (VoP), an efficient framework that adapts pre-trained dual-encoder models to video-text retrieval by freezing the backbone and optimizing only lightweight prompt tokens. The primary objective is to demonstrate that prompt tuning across both visual and textual branches—augmented by video-specific temporal modeling—can match or exceed the performance of full fine-tuning while drastically reducing trainable parameter storage.
The research evaluated VoP across five benchmark text-video retrieval datasets, including MSR-VTT, DiDeMo, ActivityNet, and LSMDC. The core framework inserts small sets of learnable continuous vectors into every layer of both the text and visual Transformer encoders. To capture the dynamic nature of video without adding heavy temporal networks, the authors developed three targeted prompt mechanisms: position-specific prompts to encode relative frame order, context-specific prompts generated via a lightweight recurrent module, and function-specific prompts that repurpose deeper visual layers to handle spatio-temporal self-attention across frames.
The evaluation yielded several critical findings. First, the baseline VoP method achieved retrieval accuracy comparable to existing efficient tuning protocols while requiring only 0.1% of the original model's trainable parameters. Second, introducing video-specific prompts consistently improved retrieval accuracy, with function-specific prompting outperforming full fine-tuning by 0.3% on average without adding extra parameters. Third, combining function-specific prompts with contextual prompts achieved an average 1.4% gain in top-1 retrieval recall over full fine-tuning across all benchmarks, while reducing parameter overhead by more than sixfold.
These findings indicate that organizations can achieve superior cross-modal retrieval performance without the substantial financial and computational costs of updating and storing complete model backbones. By retaining frozen weights, systems preserve foundational pre-trained knowledge and avoid overfitting on limited downstream data, lowering the risk and infrastructure footprint required for production multi-modal search systems.
Teams implementing text-video retrieval systems should adopt parameter-efficient prompt tuning over full model fine-tuning. Depending on performance and latency budgets, practitioners can deploy baseline VoP for extreme parameter savings or combine function- and context-specific prompts when top retrieval accuracy is paramount. Future initiatives should evaluate pairing these lightweight internal prompt mechanisms with external cross-modal fusion modules.
Confidence in these findings is supported by rigorous benchmarking across five standard retrieval datasets and extensive ablation studies. However, the study relies on a fixed frame-sampling rate and operates within specific architectural bounds, meaning performance trade-offs should be validated when scaling to significantly longer video sequences or alternate base models.
- Paper: Learning to Prompt for Vision-Language Models, Kaiyang Zhou et al. (2021). It introduces continuous prompt tuning (CoOp) for pre-trained vision-language models like CLIP, establishing the foundational prompt adaptation paradigm that VoP adapts for video.
- Paper: Visual Prompt Tuning, Menglin Jia et al. (2022). It formulates Visual Prompt Tuning (VPT) for vision transformers, providing the core parameter-efficient visual token prompt mechanism that VoP extends into spatio-temporal video prompts.
- Paper: Conditional Prompt Learning for Vision-Language Models, Kaiyang Zhou et al. (2022). It establishes conditional prompt learning across modalities in CLIP, providing direct conceptual context for VoP's text-video co-operative prompt design.
- Paper: Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval, Max Bain et al. (2021). It establishes the standard end-to-end transformer architecture and benchmark setting for cross-modal text-video retrieval.
- Paper: ViViT: A Video Vision Transformer, Anurag Arnab et al. (2021). It presents foundational transformer designs for modeling spatial and temporal frame interactions in video.
- Paper: CLIP-Adapter: Better Vision-Language Models with Feature Adapters, Peng Gao et al. (2021). It explores parameter-efficient adaptation of CLIP representations using lightweight residual adapters instead of full fine-tuning.
- Paper: MmAP: Multi-Modal Alignment Prompt for Cross-Domain Multi-Task Learning, Yi Xin et al. (2024). It generalizes multi-modal prompt tuning across modalities in CLIP to tackle multi-task and cross-domain learning scenarios.
- Paper: Text Is MASS: Modeling as Stochastic Embedding for Text-Video Retrieval, Jiamian Wang et al. (2024). It builds directly on text-video retrieval by modeling text queries as stochastic regions rather than single deterministic points.
- Paper: Video-LLaVA: Learning United Visual Representation by Alignment Before Projection, Bin Lin et al. (2023). It extends cross-modal video-language alignment by unifying image and video representations prior to projection into multimodal language models.
- Paper: Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization, Yang Jin et al. (2024). It advances unified video-language pre-training by decoupling spatial keyframe visual tokens and temporal motion tokens for efficient comprehension.
