VoP: Text-Video Co-Operative Prompt Tuning for Cross-Modal Retrieval

Siteng HuangBiao GongYulin PanJianwen JiangYiliang LvYuyuan LiDonglin Wang

article2023CVPR66 citations

Proposes an efficient text-video co-operative prompt tuning framework that incorporates spatio-temporal video prompts into CLIP to outperform full fine-tuning on text-video retrieval benchmarks with six times fewer parameters.

Listen

Adapting large foundation models to cross-modal text-video retrieval typically relies on full fine-tuning or adding heavy architectural components. These standard approaches present significant operational challenges: updating all underlying parameters creates high computational overhead, risks catastrophic forgetting of pre-trained knowledge, and requires storing massive, separate model copies for every downstream deployment.

The article introduces and evaluates "Text-Video Co-operative Prompt Tuning" (VoP), an efficient framework that adapts pre-trained dual-encoder models to video-text retrieval by freezing the backbone and optimizing only lightweight prompt tokens. The primary objective is to demonstrate that prompt tuning across both visual and textual branches—augmented by video-specific temporal modeling—can match or exceed the performance of full fine-tuning while drastically reducing trainable parameter storage.

The research evaluated VoP across five benchmark text-video retrieval datasets, including MSR-VTT, DiDeMo, ActivityNet, and LSMDC. The core framework inserts small sets of learnable continuous vectors into every layer of both the text and visual Transformer encoders. To capture the dynamic nature of video without adding heavy temporal networks, the authors developed three targeted prompt mechanisms: position-specific prompts to encode relative frame order, context-specific prompts generated via a lightweight recurrent module, and function-specific prompts that repurpose deeper visual layers to handle spatio-temporal self-attention across frames.

The evaluation yielded several critical findings. First, the baseline VoP method achieved retrieval accuracy comparable to existing efficient tuning protocols while requiring only 0.1% of the original model's trainable parameters. Second, introducing video-specific prompts consistently improved retrieval accuracy, with function-specific prompting outperforming full fine-tuning by 0.3% on average without adding extra parameters. Third, combining function-specific prompts with contextual prompts achieved an average 1.4% gain in top-1 retrieval recall over full fine-tuning across all benchmarks, while reducing parameter overhead by more than sixfold.

These findings indicate that organizations can achieve superior cross-modal retrieval performance without the substantial financial and computational costs of updating and storing complete model backbones. By retaining frozen weights, systems preserve foundational pre-trained knowledge and avoid overfitting on limited downstream data, lowering the risk and infrastructure footprint required for production multi-modal search systems.

Teams implementing text-video retrieval systems should adopt parameter-efficient prompt tuning over full model fine-tuning. Depending on performance and latency budgets, practitioners can deploy baseline VoP for extreme parameter savings or combine function- and context-specific prompts when top retrieval accuracy is paramount. Future initiatives should evaluate pairing these lightweight internal prompt mechanisms with external cross-modal fusion modules.

Confidence in these findings is supported by rigorous benchmarking across five standard retrieval datasets and extensive ablation studies. However, the study relies on a fixed frame-sampling rate and operates within specific architectural bounds, meaning performance trade-offs should be validated when scaling to significantly longer video sequences or alternate base models.

Cover for VoP: Text-Video Co-Operative Prompt Tuning for Cross-Modal Retrieval

Abstract

Many recent studies leverage the pre-trained CLIP for text-video cross-modal retrieval by tuning the backbone with additional heavy modules, which not only brings huge computational burdens with much more parameters, but also leads to the knowledge forgetting from upstream models. In this work, we propose the VoP: Text-Video Co-operative Prompt Tuning for efficient tuning on the text-video retrieval task. The proposed VoP is an end-to-end framework with both video & text prompts introducing, which can be regarded as a powerful baseline with only 0.1% trainable parameters. Further, based on the spatio-temporal characteristics of videos, we develop three novel video prompt mechanisms to improve the performance with different scales of trainable parameters. The basic idea of the VoP enhancement is to model the frame position, frame context, and layer function with specific trainable prompts, respectively. Extensive experiments show that compared to full fine-tuning, the enhanced VoP achieves a 1.4% average R@1 gain across five text-video retrieval benchmarks with 6× less parameter overhead. The code will be available at https://github.com/bighuang624/VoP.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Methodology
  • 3.1. Preliminary
  • 3.2. Text-Video Co-operative Prompt Tuning (VoP)
  • 3.3. Equipping with Video Prompts
  • 4. Experiments
  • 4.1. Experimental Setup
  • 4.2. Main Results
  • 4.3. Ablation Study
  • 4.4. Qualitative Results
  • 5. Conclusion
  • Acknowledgement
  • References

Knowls

  1. Knowl 1 — Text-Video Co-operative Prompt Tuning Framework

    model/method

    Text-Video Co-operative Prompt Tuning (VoP) adapts a pre-trained dual-encoder vision-language foundation model (such as CLIP) to text-video cross-modal retrieval by prepending continuous learnable prompt vectors to every layer of both the text encoder and the visual encoder, while freezing all original backbone parameters.

    Let the pre-trained text encoder have KK Transformer layers. For the ii-th layer Lit\mathcal{L}^t_i (i∈{1,…,K}i \in \{1, \dots, K\}), a set of learnable textual prompt embeddings Pi−1t∈RPt×dt\mathbf{P}^t_{i-1} \in \mathbb{R}^{P^t \times d_t} (where PtP^t is the text prompt length and dtd_t is the text feature dimension) is prepended to the text representation Wi−1∈RN×dt\mathbf{W}_{i-1} \in \mathbb{R}^{N \times d_t} of the NN input tokens:

    [‾,Wi]=Lit([Pi−1t,Wi−1]),[\underline{\quad}, \mathbf{W}_i] = \mathcal{L}^t_i([\mathbf{P}^t_{i-1}, \mathbf{W}_{i-1}]),

    where [⋅,⋅][\cdot, \cdot] denotes concatenation along the sequence length dimension, and the output tokens at the prompt positions (denoted by ‾\underline{\quad}) are discarded. The final text embedding zt∈Rd\mathbf{z}^t \in \mathbb{R}^d is obtained by projecting the last layer's token at the [extEOS][ ext{EOS}] position.

    For the visual encoder receiving MM patch tokens E0∈RM×dv\mathbf{E}_0 \in \mathbb{R}^{M \times d_v} and a learnable [extCLS][ ext{CLS}] token c0∈Rdv\mathbf{c}_0 \in \mathbb{R}^{d_v}, learnable visual prompt tokens Pi−1v∈RPv×dv\mathbf{P}^v_{i-1} \in \mathbb{R}^{P^v \times d_v} of length PvP^v and dimension dvd_v are prepended at each layer i∈{1,…,K}i \in \{1, \dots, K\}:

    [ci,‾,Ei]=Liv([ci−1,Pi−1v,Ei−1]).[\mathbf{c}_i, \underline{\quad}, \mathbf{E}_i] = \mathcal{L}^v_i([\mathbf{c}_{i-1}, \mathbf{P}^v_{i-1}, \mathbf{E}_{i-1}]).

    The final frame embedding zjv∈Rd\mathbf{z}^v_j \in \mathbb{R}^d for frame j∈{1,…,F}j \in \{1, \dots, F\} is obtained by projecting cK\mathbf{c}_K. The overall video embedding zˉv\mathbf{\bar{z}}^v is the average of all frame embeddings 1F∑j=1Fzjv\frac{1}{F} \sum_{j=1}^F \mathbf{z}^v_j. The model is trained end-to-end via symmetric text-to-video and video-to-text cross-entropy contrastive losses.

  2. Knowl 2 — Position-Specific Video Prompts

    model/method

    Position-Specific Video Prompts (extVoPextP ext{VoP}^{ ext{P}}) inject temporal order information into the visual encoder by parameterizing separate visual prompt vectors for each relative frame position index j∈{1,…,F}j \in \{1, \dots, F\} in a video sequence of FF sampled frames.

    A prompt table indexed by frame position is maintained. When processing the jj-th frame at the ii-th visual encoder layer Liv\mathcal{L}^v_i, the layer transformation is computed as:

    [ci;j,‾,Ei;j]=Liv([ci−1;j,Pi−1;jv,Ei−1;j]),[\mathbf{c}_{i; j}, \underline{\quad}, \mathbf{E}_{i; j}] = \mathcal{L}^v_i([\mathbf{c}_{i-1; j}, \mathbf{P}^v_{i-1; j}, \mathbf{E}_{i-1; j}]),

    where ci−1;j\mathbf{c}_{i-1; j} is the [extCLS][ ext{CLS}] token and Ei−1;j\mathbf{E}_{i-1; j} represents the patch tokens of frame jj, while Pi−1;jv∈RPv×dv\mathbf{P}^v_{i-1; j} \in \mathbb{R}^{P^v \times d_v} denotes the visual prompt tokens specific to frame position jj.

    To balance model capacity and parameter efficiency, only a subset of the prompt tokens in each layer are position-specific, while the remaining tokens are position-agnostic prompts shared across all frames.

  3. Knowl 3 — Context-Specific Video Prompts

    model/method

    Context-Specific Video Prompts (extVoPextC ext{VoP}^{ ext{C}}) dynamically generate input-dependent visual prompt tokens for each frame by capturing cross-frame contextual interactions across the entire video sequence.

    At layer ii of the visual encoder, the [extCLS][ ext{CLS}] tokens from all FF frames of a video are collected into a sequence Ci−1={ci−1;1,ci−1;2,…,ci−1;F}∈RF×dv\mathbf{C}_{i-1} = \{\mathbf{c}_{i-1;1}, \mathbf{c}_{i-1;2}, \dots, \mathbf{c}_{i-1;F}\} \in \mathbb{R}^{F \times d_v}. This sequence is fed into a shared Context Modeling Module (CMM\text{CMM}), implemented as a 1-layer Bidirectional LSTM (BiLSTM), to modulate each frame representation with temporal context:

    C~i−1=CMM(Ci−1).\widetilde{\mathbf{C}}_{i-1} = \text{CMM}(\mathbf{C}_{i-1}).

    For each frame j∈{1,…,F}j \in \{1, \dots, F\}, a shared fully-connected (FC) linear layer generates the prompt tokens from the modulated token c~i−1;j\tilde{\mathbf{c}}_{i-1; j} by projecting and splitting the vector into PvP^v prompt tokens:

    P~i−1;jv=FC(c~i−1;j).\widetilde{\mathbf{P}}^v_{i-1; j} = \text{FC}(\tilde{\mathbf{c}}_{i-1; j}).

    The frame representations are then updated in the original visual encoder layer via:

    [ci;j,‾,Ei;j]=Liv([ci−1;j,P~i−1;jv,Ei−1;j]).[\mathbf{c}_{i; j}, \underline{\quad}, \mathbf{E}_{i; j}] = \mathcal{L}^v_i([\mathbf{c}_{i-1; j}, \widetilde{\mathbf{P}}^v_{i-1; j}, \mathbf{E}_{i-1; j}]).

    The modulated tokens C~i−1\widetilde{\mathbf{C}}_{i-1} are not passed directly to the next layer, thereby preserving the original feature representations while injecting context solely via the generated prompts. The parameters of CMM\text{CMM} and the FC\text{FC} layer are shared across all visual layers.

  4. Knowl 4 — Function-Specific Video Prompts

    model/method

    Function-Specific Video Prompts (extVoPextF ext{VoP}^{ ext{F}}) repurpose the frozen layers of the pre-trained visual Transformer into two distinct functional stages to perform spatio-temporal modeling without adding heavy temporal architectures.

    The visual encoder with KK layers is partitioned at layer KsK_s:

    1. Shallow Layers (i≤Ksi \le K_s): Perform intra-frame spatial self-attention independently on tokens of each frame j∈{1,…,F}j \in \{1, \dots, F\}:
    [ci;j,‾,Ei;j]=Liv([ci−1;j,Pi−1;jv,Ei−1;j]).[\mathbf{c}_{i; j}, \underline{\quad}, \mathbf{E}_{i; j}] = \mathcal{L}^v_i([\mathbf{c}_{i-1; j}, \mathbf{P}^v_{i-1; j}, \mathbf{E}_{i-1; j}]).
    1. Deep Layers (i>Ksi > K_s): Perform inter-frame spatio-temporal self-attention across the combined sequence containing tokens from all FF frames simultaneously. Before layer Ks+1K_s + 1, a trainable frame positional embedding is added to all video tokens. For i>Ksi > K_s, the layer computation is:
    [Ci,‾,Ei;1,Ei;2,…,Ei;F]=Liv([Ci−1,Pi−1v,Ei−1;1,Ei−1;2,…,Ei−1;F]),[\mathbf{C}_i, \underline{\quad}, \mathbf{E}_{i;1}, \mathbf{E}_{i;2}, \dots, \mathbf{E}_{i;F}] = \mathcal{L}^v_i([\mathbf{C}_{i-1}, \mathbf{P}^v_{i-1}, \mathbf{E}_{i-1;1}, \mathbf{E}_{i-1;2}, \dots, \mathbf{E}_{i-1;F}]),

    where Ci−1=[ci−1;1,…,ci−1;F]\mathbf{C}_{i-1} = [\mathbf{c}_{i-1;1}, \dots, \mathbf{c}_{i-1;F}] gathers the [extCLS][ ext{CLS}] tokens of all frames, and Pi−1v\mathbf{P}^v_{i-1} acts as video-level prompt tokens.

    Hybrid variants VoPF+P\text{VoP}^{\text{F+P}} and VoPF+C\text{VoP}^{\text{F+C}} apply position-specific (VoPP\text{VoP}^{\text{P}}) and context-specific (VoPC\text{VoP}^{\text{C}}) prompts, respectively, in the shallow KsK_s layers while employing function-specific inter-frame self-attention in the deeper layers.

  5. Knowl 5 — Retrieval Performance on MSR-VTT-9k Across Tuning Strategies

    data/table

    Evaluation on the MSR-VTT-9k text-video retrieval benchmark comparing full fine-tuning, parameter-efficient baselines, and VoP variants using a pre-trained CLIP (ViT-B/32) backbone.

    Methods Params (M) t2v v2t
    R@1 R@5 R@10 MnR↓\downarrow MdR↓\downarrow R@1 R@5 R@10 MnR↓\downarrow MdR↓\downarrow
    Full 119.8 (100%) 41.7 69.2 79.0 16.5 2.0 42.5 70.9 81.4 11.0 2.0
    Bias 0.1 (0.104%) 39.7 66.5 77.3 17.3 2.0 41.1 68.4 79.2 13.6 2.0
    Proj 0.7 (0.547%) 37.1 63.0 76.1 20.5 3.0 37.2 64.6 75.9 16.7 3.0
    Partial 7.7 (6.410%) 39.8 65.3 75.9 19.3 2.0 37.9 66.1 77.4 15.5 3.0
    AdapterATTN^{\text{ATTN}} 2.0 (1.655%) 37.6 63.2 75.8 18.7 3.0 39.6 66.5 76.8 14.7 2.0
    AdapterFFN^{\text{FFN}} 2.0 (1.655%) 38.2 63.5 76.4 17.9 3.0 39.9 66.8 77.7 14.2 2.0
    VoP 0.1 (0.103%) 39.6 66.7 77.8 17.2 2.0 42.1 68.8 80.7 12.4 2.0
    VoPP^{\text{P}} 0.5 (0.441%) 40.1 65.7 77.7 16.9 2.0 42.5 70.0 79.9 12.4 2.0
    VoPC^{\text{C}} 14.3 (11.898%) 40.8 68.1 79.0 15.8 2.0 42.3 70.1 81.1 11.4 2.0
    VoPF^{\text{F}} 0.1 (0.103%) 42.6 68.4 78.7 15.8 2.0 42.4 70.5 81.0 11.0 2.0
    VoPF+P^{\text{F+P}} 0.4 (0.328%) 43.5 69.3 79.3 14.8 2.0 43.6 71.2 81.2 11.0 2.0
    VoPF+C^{\text{F+C}} 14.1 (11.785%) 44.6 69.9 80.3 16.3 2.0 44.5 70.7 80.6 11.5 2.0

    VoP tunes only 0.103%0.103\% of parameters (0.1 M0.1\text{ M}) while competing closely with full fine-tuning. Adding function-specific prompts (VoPF\text{VoP}^{\text{F}}) surpasses full fine-tuning by +0.9%+0.9\% t2v R@1 without adding extra parameters. Combining prompts (VoPF+C\text{VoP}^{\text{F+C}}) attains 44.6%44.6\% t2v R@1 (+2.9%+2.9\% over Full) with over 8×8\times fewer parameters.

  6. Knowl 6 — Ablation on Co-operative Multi-Modal Prompts

    data/table

    Evaluating the contribution of inserting prompts into text-only, visual-only, versus both encoders simultaneously on the MSR-VTT-9k benchmark.

    Textual Visual R@1 R@5 R@10 MnR↓\downarrow MdR↓\downarrow
    31.5 52.8 63.6 42.9 5.0
    ✓ 36.5 62.7 75.1 18.3 3.0
    ✓ 36.3 63.4 75.0 20.3 3.0
    ✓ ✓ 39.6 66.7 77.8 17.2 2.0

    Prompting either modality in isolation improves zero-shot CLIP (31.5%31.5\% R@1) to 36.5%36.5\% (text-only) and 36.3%36.3\% (visual-only). Prompting both encoders collaboratively (VoP) yields 39.6%39.6\% R@1, confirming that co-operative tuning of both modalities produces complementary cross-modal alignment.

  7. Knowl 7 — Context Modeling Module Architecture and Prompt Generation Strategy

    data/table

    Ablation results comparing temporal Context Modeling Module (CMM) architectures and prompt generation vs. direct token replacement in VoPC\text{VoP}^{\text{C}}.

    1. CMM Architecture Selection across Benchmarks (R@1 / R@5 / R@10):
    Choice of CMM MSR-VTT-9k MSR-VTT-7k DiDeMo ActivityNet LSMDC
    R@1 R@5 R@10 R@1 R@5 R@10 R@1 R@5 R@10 R@1 R@5 R@10 R@1 R@5 R@10
    Transformer (4 layers) 40.1 68.2 78.8 39.5 68.2 78.1 40.4 67.3 77.3 32.0 61.5 74.9 20.3 39.5 47.8
    LSTM (1 layer) 40.6 69.5 79.7 39.5 69.3 78.0 38.6 66.7 77.0 32.4 62.0 75.4 19.6 38.2 47.7
    BiLSTM (1 layer) 40.8 68.1 79.0 40.0 67.3 78.2 40.0 68.0 78.5 32.6 62.5 76.5 20.4 40.0 48.1

    A 1-layer BiLSTM provides the most consistent performance across datasets.

    1. Role of CMM on MSR-VTT-9k:
    Role of CMM R@1 R@5 R@10 MnR↓\downarrow MdR↓\downarrow
    Updating [CLS] tokens directly 38.0 63.6 75.3 18.5 2.0
    Generating prompts (extVoPC ext{VoP}^{\text{C}}) 40.8 68.1 79.0 15.8 2.0

    Using CMM to generate prompt tokens outperforms directly replacing the [extCLS][ ext{CLS}] tokens with CMM outputs by +2.8%+2.8\% R@1, because prompt generation allows cross-frame communication without corrupting the original encoder feature representations.

  8. Knowl 8 — Optimal Visual Layer Partitioning for Function-Specific Video Prompts

    empirical result

    In VoPF\text{VoP}^{\text{F}}, the visual Transformer encoder with K=12K = 12 layers is partitioned into KsK_s shallow layers performing intra-frame spatial self-attention and 12−Ks12 - K_s deep layers performing inter-frame spatio-temporal self-attention. Evaluating the split index Ks∈{2,4,6,8,10,12}K_s \in \{2, 4, 6, 8, 10, 12\} on MSR-VTT-9k text-to-video retrieval yields the following R@1 performance:

    • Ks=2K_s = 2: 39.9%39.9\%
    • Ks=4K_s = 4: 40.1%40.1\%
    • Ks=6K_s = 6: 41.9%41.9\%
    • Ks=8K_s = 8: 42.6%42.6\%
    • Ks=10K_s = 10: 42.3%42.3\%
    • Ks=12K_s = 12 (equivalent to base VoP with spatial attention only): 39.6%39.6\%

    When Ks=0K_s = 0 (all layers performing inter-frame spatio-temporal self-attention), the model fails to generalize and performance collapses. The optimal value is Ks=8K_s = 8, providing a trade-off where shallow layers extract intra-frame spatial features and the last 4 deep layers perform inter-frame temporal modeling.

  9. Knowl 9 — Sensitivity to Prompt Depth, Length, and Specific Token Allocation

    empirical result

    Ablation experiments on hyperparameter choices in VoP on MSR-VTT-9k establish the following design principles:

    1. Prompt Depth: Inserting prompts into layers progressively from input to deeper layers monotonically increases text-to-video R@1: layer 1 only achieves 37.8%37.8\%, layers 1–3 achieve 38.2%38.2\%, layers 1–6 achieve 38.8%38.8\%, layers 1–9 achieve 39.0%39.0\%, and layers 1–12 (every layer) achieve 39.6%39.6\%.
    2. Prompt Length: Varying the total prompt length per layer across {4,8,12,16,20}\{4, 8, 12, 16, 20\} yields R@1 scores of 39.4%39.4\%, 39.6%39.6\%, 39.6%39.6\%, 39.0%39.0\%, and 39.5%39.5\%, respectively. A prompt length of 8 tokens achieves the optimal balance of parameter efficiency and performance.
    3. Video-Specific Prompt Token Ratio: When replacing a total of 8 visual prompt tokens with video-specific prompt tokens in VoPP\text{VoP}^{\text{P}}, VoPC\text{VoP}^{\text{C}}, VoPF+P\text{VoP}^{\text{F+P}}, and VoPF+C\text{VoP}^{\text{F+C}}, replacing exactly 4 out of the 8 tokens (i.e., 4 general tokens + 4 specific tokens) yields the highest performance across variants (e.g., reaching 40.8%40.8\% in VoPC\text{VoP}^{\text{C}} and 44.6%44.6\% in VoPF+C\text{VoP}^{\text{F+C}}), as general tokens preserve cross-video shared concepts while specific tokens capture frame-level or video-level temporal patterns.
  10. Knowl 10 — Cross-Benchmark Performance and Generalization Across Five Retrieval Datasets

    empirical result

    Comparing text-to-video retrieval (t2v R@1) gains relative to full fine-tuning (Full) across five standard benchmarks (MSR-VTT-9k, MSR-VTT-7k, DiDeMo, ActivityNet, and LSMDC):

    • Bias Tuning: −3.7%-3.7\% average R@1 relative gain (0.104%0.104\% parameters).
    • Proj (Linear Projection Only): −5.8%-5.8\% average R@1 (0.547%0.547\% parameters).
    • Partial (Last Layer Only): −2.6%-2.6\% average R@1 (6.410%6.410\% parameters).
    • AdapterATTN^{\text{ATTN}}: −3.9%-3.9\% average R@1 (1.655%1.655\% parameters).
    • AdapterFFN^{\text{FFN}}: −3.6%-3.6\% average R@1 (1.655%1.655\% parameters).
    • Base VoP: −2.8%-2.8\% average R@1 (0.103%0.103\% parameters).
    • VoPP\text{VoP}^{\text{P}}: −2.3%-2.3\% average R@1 (0.441%0.441\% parameters).
    • VoPC\text{VoP}^{\text{C}}: −1.8%-1.8\% average R@1 (11.898%11.898\% parameters).
    • VoPF\text{VoP}^{\text{F}}: +0.3%+0.3\% average R@1 over Full fine-tuning (0.103%0.103\% parameters).
    • VoPF+P\text{VoP}^{\text{F+P}}: +1.2%+1.2\% average R@1 over Full fine-tuning (0.328%0.328\% parameters).
    • VoPF+C\text{VoP}^{\text{F+C}}: +1.4%+1.4\% average R@1 over Full fine-tuning (11.785%11.785\% parameters, over 6×6\times fewer parameters than Full).

    On DiDeMo, VoPF+C\text{VoP}^{\text{F+C}} exceeds Full fine-tuning by +4.8%+4.8\% R@1, demonstrating strong generalization in paragraph-to-video retrieval.

Coverage note — None was omitted; all contributed models (VoP, VoP^P, VoP^C, VoP^F, VoP^{F+P}, VoP^{F+C}), mathematical formulations, main retrieval tables, ablation studies, and multi-dataset empirical results are fully covered.

References

  1. 1.Hyojin Bahng, Ali Jahanian, Swami Sankaranarayanan, and Phillip Isola. Visual prompting: Modifying pixel space to adapt pre-trained models. arXiv preprint arXiv:2203.17274, 2022.
  2. 2.Max Bain, Arsha Nagrani, Gõl Varol, and Andrew Zisserman. Frozen in time: A joint video and image encoder for end-to-end retrieval. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1708–1718, 2021.
  3. 3.Gedas Bertasius, Heng Wang, and Lorenzo Torresani. Is space-time attention all you need for video understanding? In Proceedings of the International Conference on Machine Learning, pages 813–824, 2021.
  4. 4.Johan Bjorck, Kilian Q. Weinberger, and Carla P. Gomes. Understanding decoupled and early weight decay. arXiv preprint arXiv:2012.13841, 2020.
  5. 5.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. In Proceedings of the Advances in Neural Information Processing Systems, pages 1877–1901, 2020.
  6. 6.Han Cai, Chuang Gan, Ligeng Zhu, and Song Han. TinyTL: Reduce memory, not parameters for efficient on-device learning. In Proceedings of the Advances in Neural Information Processing Systems, pages 11285–11297, 2020.
  7. 7.Shoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang, Yibing Song, Jue Wang, and Ping Luo. AdaptFormer: Adapting vision transformers for scalable visual recognition. In Proceedings of the Advances in Neural Information Processing Systems, 2022.
  8. 8.Yu Du, Fangyun Wei, Zihe Zhang, Miaojing Shi, Yue Gao, and Guoqi Li. Learning to prompt for open-vocabulary object detection with vision-language model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14064–14073, 2022.
  9. 9.Valentin Gabeur, Chen Sun, Karteek Alahari, and Cordelia Schmid. Multi-modal transformer for video retrieval. In Proceedings of the European Conference on Computer Vision, pages 214–229, 2020.
  10. 10.Satya Krishna Gorti, Noel Vouitsis, Junwei Ma, Keyvan Golestan, Maksims Volkovs, Animesh Garg, and Guang Wei Yu. X-Pool: Cross-modal language-video attention for text-video retrieval. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4996–5005, 2022.
  11. 11.Alex Graves and Jürgen Schmidhuber. Framewise phoneme classification with bidirectional LSTM and other neural network architectures. Neural Networks, pages 602–610, 2005.
  12. 12.Junxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick, and Graham Neubig. Towards a unified view of parameter-efficient transfer learning. In Proceedings of the International Conference on Learning Representations, 2022.
  13. 13.Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. ActivityNet: A large-scale video benchmark for human activity understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 961–970, 2015.
  14. 14.Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan C. Russell. Localizing moments in video with natural language. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5804–5813, 2017.
  15. 15.Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. Parameter-efficient transfer learning for NLP. In Proceedings of the International Conference on Machine Learning, pages 2790–2799, 2019.
  16. 16.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V. Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In Proceedings of the International Conference on Machine Learning, pages 4904–4916, 2021.
  17. 17.Menglin Jia, Luming Tang, Bor-Chun Chen, Claire Cardie, Serge J. Belongie, Bharath Hariharan, and Ser-Nam Lim. Visual prompt tuning. In Proceedings of the European Conference on Computer Vision, pages 709–727, 2022.
  18. 18.Chen Ju, Tengda Han, Kunhao Zheng, Ya Zhang, and Weidi Xie. Prompting visual-language models for efficient video understanding. In Proceedings of the European Conference on Computer Vision, pages 105–124, 2022.
  19. 19.Jie Lei, Linjie Li, Luowei Zhou, Zhe Gan, Tamara L. Berg, Mohit Bansal, and Jingjing Liu. Less is more: ClipBERT for video-and-language learning via sparse sampling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7331–7341, 2021.
  20. 20.Brian Lester, Rami Al-Rfou, and Noah Constant. The power of scale for parameter-efficient prompt tuning. In Proceedings of the Conference on Empirical Methods in Natural Language Processing, pages 3045–3059, 2021.
  21. 21.Junnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty, Caiming Xiong, and Steven Chu-Hong Hoi. Align before fuse: Vision and language representation learning with momentum distillation. In Proceedings of the Advances in Neural Information Processing Systems, pages 9694–9705, 2021.
  22. 22.Yangguang Li, Feng Liang, Lichen Zhao, Yufeng Cui, Wanli Ouyang, Jing Shao, Fengwei Yu, and Junjie Yan. Supervision exists everywhere: A data efficient contrastive language-image pre-training paradigm. In Proceedings of the International Conference on Learning Representations, 2022.
  23. 23.Xiao Liu, Kaixuan Ji, Yicheng Fu, Zhengxiao Du, Zhilin Yang, and Jie Tang. P-Tuning v2: Prompt tuning can be comparable to fine-tuning universally across scales and tasks. arXiv preprint arXiv:2110.07602, 2021.
  24. 24.Xiao Liu, Yanan Zheng, Zhengxiao Du, Ming Ding, Yujie Qian, Zhilin Yang, and Jie Tang. GPT understands, too. arXiv preprint arXiv:2103.10385, 2021.
  25. 25.Yang Liu, Samuel Albanie, Arsha Nagrani, and Andrew Zisserman. Use what you have: Video retrieval using representations from collaborative experts. In Proceedings of the British Machine Vision Conference, page 279, 2019.
  26. 26.Ilya Loshchilov and Frank Hutter. SGDR: Stochastic gradient descent with warm restarts. In Proceedings of the International Conference on Learning Representations, 2017.
  27. 27.Huaishao Luo, Lei Ji, Ming Zhong, Yang Chen, Wen Lei, Nan Duan, and Tianrui Li. CLIP4Clip: An empirical study of CLIP for end to end video clip retrieval and captioning. Neurocomputing, pages 293–304, 2022.
  28. 28.Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic, and Andrew Zisserman. End-to-end learning of visual representations from uncurated instructional videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9876–9886, 2020.
  29. 29.Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic. HowTo100M: Learning a text-video embedding by watching hundred million narrated video clips. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2630–2640, 2019.
  30. 30.Or Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or, and Dani Lischinski. StyleCLIP: Text-driven manipulation of StyleGAN imagery. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2065–2074, 2021.
  31. 31.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning, pages 8748–8763, 2021.
  32. 32.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  33. 33.Anna Rohrbach, Marcus Rohrbach, and Bernt Schiele. The long-short story of movie description. In Proceedings of the German Conference on Pattern Recognition, pages 209–221, 2015.
  34. 34.Alex Jinpeng Wang, Yixiao Ge, Guanyu Cai, Rui Yan, Xudong Lin, Ying Shan, Xiaohu Qie, and Mike Zheng Shou. Object-aware video-language pre-training for retrieval. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3303–3312, 2022.
  35. 35.Mengmeng Wang, Jiazheng Xing, and Yong Liu. ActionCLIP: A new paradigm for video action recognition. arXiv preprint arXiv:2109.08472, 2021.
  36. 36.Peng Wu, Xiangteng He, Mingqian Tang, Yiliang Lv, and Jing Liu. HANet: Hierarchical alignment networks for video-text retrieval. In Proceedings of the ACM International Conference on Multimedia, pages 3518–3527, 2021.
  37. 37.Jun Xu, Tao Mei, Ting Yao, and Yong Rui. MSR-VTT: A large video description dataset for bridging video and language. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5288–5296, 2016.
  38. 38.Youngjae Yu, Jongseok Kim, and Gunhee Kim. A joint sequence fusion model for video question answering and retrieval. In Proceedings of the European Conference on Computer Vision, pages 487–503, 2018.
  39. 39.Elad Ben Zaken, Yoav Goldberg, and Shauli Ravfogel. BitFit: Simple parameter-efficient fine-tuning for transformer-based masked language-models. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, pages 1–9, 2022.
  40. 40.Shuai Zhao, Linchao Zhu, Xiaohan Wang, and Yi Yang. CenterCLIP: Token clustering for efficient text-video retrieval. In Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 970–981, 2022.
  41. 41.Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. Learning to prompt for vision-language models. International Journal of Computer Vision, pages 2337–2348, 2022.

Citation

MLA
Huang, S., et al. “VoP: Text-Video Co-operative Prompt Tuning for Cross-Modal Retrieval”. arXiv, 2022, http://arxiv.org/abs/2211.12764v3.
APA
Huang, S., Gong, B., Pan, Y., Jiang, J., Lv, Y., Li, Y., & Wang, D. (2022). VoP: Text-Video Co-operative Prompt Tuning for Cross-Modal Retrieval. arXiv. http://arxiv.org/abs/2211.12764v3
Chicago
Huang, S., B. Gong, Y. Pan, et al. 2022. “VoP: Text-Video Co-operative Prompt Tuning for Cross-Modal Retrieval”. arXiv. http://arxiv.org/abs/2211.12764v3.
Harvard
Huang, S. et al. (2022) “VoP: Text-Video Co-operative Prompt Tuning for Cross-Modal Retrieval”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2211.12764v3.
Vancouver
1. Huang S, Gong B, Pan Y, Jiang J, Lv Y, Li Y, Wang D (2022) VoP: Text-Video Co-operative Prompt Tuning for Cross-Modal Retrieval. arXiv

BibTeX

@article{huang2022vop,
  title = {VoP: Text-Video Co-operative Prompt Tuning for Cross-Modal Retrieval},
  author = {Huang, Siteng and Gong, Biao and Pan, Yulin and Jiang, Jianwen and Lv, Yiliang and Li, Yuyuan and Wang, Donglin},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2211.12764v3},
  eprint = {2211.12764}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE