Text Is MASS: Modeling as Stochastic Embedding for Text-Video Retrieval

Jiamian WangPichao WangGuohao SunDongfang LiuSohail A. DianatRaghuveer RaoMajid RabbaniZhiqiang Tao

article2024CVPR84 citationsHighlight

Proposes T-MASS, a text-video retrieval framework that models concise text queries as stochastic embeddings with adaptive radii to capture the broad semantic scope of rich video content and achieve state-of-the-art retrieval accuracy across multiple benchmark datasets.

Listen

Text-video retrieval is increasingly vital as video content proliferates across digital platforms, yet the task faces a fundamental information mismatch. While videos contain complex, redundant visual information across multiple frames, search queries and captions are typically short and concise. Existing methods map both text and video to single fixed points in a shared mathematical space, which often fails because a single text point lacks the semantic richness needed to match the full scope of a video.

The article evaluates a novel retrieval method called T-MASS, which treats text not as a single deterministic point, but as an elastic, stochastic region termed a text mass. The main objective is to demonstrate that modeling text as a flexible semantic range improves alignment with rich video representations and significantly boosts retrieval accuracy without requiring expensive video pre-training or complex visual extraction architectures.

To test this concept, the authors developed a framework built upon standard pre-trained image-text models. The approach introduces a similarity-aware radius network that dynamically scales the size of the text mass based on the text-video pair, ensuring that closely related pairs form tight, precise representations while unrelated pairs remain distant. During training, the model uses a specialized learning objective and a support text regularization vector at the boundary of the mass to control its position and volume. The model was evaluated across five diverse benchmark video datasets using standard recall metrics and tested against leading state-of-the-art systems.

The experimental findings show substantial performance improvements across the board. T-MASS outperformed baseline methods by 3.0% to 6.3% in top-rank accuracy (Recall@1), achieving new state-of-the-art results on five standard benchmarks. On challenging datasets such as LSMDC and Charades, T-MASS improved top-1 recall by over 3.7% and 6.0% respectively compared to prior architectures. Furthermore, the analysis confirmed that the text mass mechanism effectively separates irrelevant pairs in the shared space while drawing true pairs closer together, achieving superior alignment even under varying video frame lengths.

These findings indicate that rethinking text modeling offers a highly cost-effective path to improving cross-modal search. Rather than expending heavy computational resources to pre-train large video encoders on millions of additional clips, organizations can achieve superior retrieval accuracy through flexible text representations. The system requires minimal structural modification to existing pipelines and maintains robust performance across different video inputs.

Organizations developing video search, recommendation, or digital asset management platforms should consider adopting stochastic text representations to upgrade legacy retrieval models. Implementing T-MASS requires balancing the number of inference samples, where testing 10 to 20 samples provides an optimal trade-off between search speed and retrieval accuracy. While the approach delivers strong, reliable gains on standard benchmarks, stakeholders should validate deployment latency under real-time production loads and evaluate whether combining text-mass modeling with large-scale pre-training yields further performance benefits.

Cover for Text Is MASS: Modeling as Stochastic Embedding for Text-Video Retrieval

Abstract

The increasing prevalence of video clips has sparked growing interest in text-video retrieval. Recent advances focus on establishing a joint embedding space for text and video, relying on consistent embedding representations to compute similarity. However, the text content in existing datasets is generally short and concise, making it hard to fully describe the redundant semantics of a video. Correspondingly, a single text embedding may be less expressive to capture the video embedding and empower the retrieval. In this study, we propose a new stochastic text modeling method T-MASS, i.e., text is modeled as a stochastic embedding, to enrich text embedding with a flexible and resilient semantic range, yielding a text mass. To be specific, we introduce a similarity-aware radius module to adapt the scale of the text mass upon the given text-video pairs. Plus, we design and develop a support text regularization to further control the text mass during the training. The inference pipeline is also tailored to fully exploit the text mass for accurate retrieval. Empirical evidence suggests that T-MASS not only effectively attracts relevant text-video pairs while distancing irrelevant ones, but also enables the determination of precise text embeddings for relevant pairs. Our experimental results show a substantial improvement of T-MASS over baseline (3% ∼ 6.3% by R@1). Also, T-MASS achieves state-of-the-art performance on five benchmark datasets, including MSRVTT, LSMDC, DiDeMo, VATEX, and Charades. Code and models are available here.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Preliminaries
  • 3.2. Text-Video Representations
  • 3.3. Proposed Method: T-MASS
  • 4. Experiment
  • 4.1. Experimental Settings
  • 4.2. Performance Comparison
  • 4.3. Model Discussion
  • 5. Conclusions
  • 6. Acknowledgement
  • References

Knowls

  1. Knowl 1 — Stochastic Text Representation as a Semantic Mass

    model/method

    In text-video retrieval, concise text descriptions often fail to encompass the redundant semantics and visual variations of corresponding video clips. Rather than projecting text query tt as a single deterministic point in a dd-dimensional joint space, the T-MASS framework represents text as a stochastic semantic region termed a "text mass".

    Given a deterministic text feature t=ϕt(t)∈Rdt = \phi_t(t) \in \mathbb{R}^d obtained from a pre-trained text encoder ϕt\phi_t, the stochastic text embedding ts∈Rdt_s \in \mathbb{R}^d is generated via the reparameterization trick:

    ts=t+R⊙ϵ,ϵ∼N(0,Id)t_s = t + \mathcal{R} \odot \epsilon, \quad \epsilon \sim \mathcal{N}(0, I_d)

    where ϵ\epsilon is an auxiliary random noise vector sampled from a standard Gaussian distribution N(0,Id)\mathcal{N}(0, I_d), ⊙\odot represents element-wise multiplication, and R∈Rd\mathcal{R} \in \mathbb{R}^d is a scale vector representing the underlying radius of the text mass. Any point sampled from this stochastic text mass is treated as a valid semantic embedding of text tt that can be evaluated against a video embedding vv using a similarity metric s(ts,v)s(t_s, v).

  2. Knowl 2 — Similarity-Aware Radius Module

    model/method

    To dynamically adjust the scale of the text mass according to the semantic alignment between a text query and candidate video frames, a similarity-aware radius module computes the radius vector R∈Rd\mathcal{R} \in \mathbb{R}^d prior to video feature fusion.

    Given the text embedding t∈Rdt \in \mathbb{R}^d and T′T' frame embeddings [f1,…,fT′]∈RT′×d[f_1, \dots, f_{T'}] \in \mathbb{R}^{T' \times d} extracted by a visual encoder ϕv\phi_v, frame-level cosine similarities are computed as:

    Si=s(t,fi)=t⊤fi∥t∥2∥fi∥2,i=1,…,T′S_i = s(t, f_i) = \frac{t^\top f_i}{\|t\|_2 \|f_i\|_2}, \quad i = 1, \dots, T'

    Concatenating these similarity scores produces the frame similarity descriptor vector S=[S1,…,ST′]∈RT′S = [S_1, \dots, S_{T'}] \in \mathbb{R}^{T'}. The radius vector R\mathcal{R} is calculated using a learnable linear transformation layer followed by an exponential scaling activation:

    R=exp⁡(SW)\mathcal{R} = \exp(S W)

    where W∈RT′×dW \in \mathbb{R}^{T' \times d} is a learnable weight matrix and exp⁡(⋅)\exp(\cdot) is applied element-wise. This parameterization allows the network to expand or shrink the text mass flexibly across each feature dimension based on fine-grained text-frame correlations.

  3. Knowl 3 — Stochastic Contrastive Objective with Support Text Regularization

    model/method

    The T-MASS framework optimizes the text mass in the joint embedding space using a combination of a stochastic symmetric cross-entropy loss and a support text regularization loss without directly regularizing the deterministic text center tt.

    Given a training batch of NN text-video pairs with video embeddings vi=ψ([f1,…,fT′],t)∈Rdv_i = \psi([f_1, \dots, f_{T'}], t) \in \mathbb{R}^d, random stochastic text samples ts,i∈Rdt_{s,i} \in \mathbb{R}^d are drawn to compute the stochastic contrastive loss Ls\mathcal{L}_s:

    Ls=12(Lts→v+Lv→ts)\mathcal{L}_s = \frac{1}{2} (\mathcal{L}_{t_s \to v} + \mathcal{L}_{v \to t_s})

    Lts→v=−1N∑i=1Nlog⁡exp⁡(s(ts,i,vi)⋅λ)∑j=1Nexp⁡(s(ts,i,vj)⋅λ),Lv→ts=−1N∑i=1Nlog⁡exp⁡(s(ts,i,vi)⋅λ)∑j=1Nexp⁡(s(ts,j,vi)⋅λ)\mathcal{L}_{t_s \to v} = -\frac{1}{N} \sum_{i=1}^N \log \frac{\exp(s(t_{s,i}, v_i) \cdot \lambda)}{\sum_{j=1}^N \exp(s(t_{s,i}, v_j) \cdot \lambda)}, \quad \mathcal{L}_{v \to t_s} = -\frac{1}{N} \sum_{i=1}^N \log \frac{\exp(s(t_{s,i}, v_i) \cdot \lambda)}{\sum_{j=1}^N \exp(s(t_{s,j}, v_i) \cdot \lambda)}

    where s(⋅,⋅)s(\cdot, \cdot) is cosine similarity and λ\lambda is a learnable temperature scaling factor.

    To control the position and volume of the high-dimensional mass without prior divergence penalties (such as KL-divergence), a support text embedding tsup∈Rdt_{\text{sup}} \in \mathbb{R}^d is defined at the boundary surface of the text mass oriented along the direction connecting tt and vv:

    tsup=t+v−t∥v−t∥2⊙Rt_{\text{sup}} = t + \frac{v - t}{\|v - t\|_2} \odot \mathcal{R}

    A symmetric contrastive regularization loss Lsup\mathcal{L}_{\text{sup}} is computed using tsupt_{\text{sup}} in place of tst_s. The overall training objective is:

    Ltotal=Ls+αLsup\mathcal{L}_{\text{total}} = \mathcal{L}_s + \alpha \mathcal{L}_{\text{sup}}

    where α\alpha is the support text regularization weight parameter (set to α=1.2\alpha = 1.2).

  4. Knowl 4 — T-MASS Retrieval Inference via Closest Stochastic Sample Selection

    algorithm

    During inference, text-video similarity is computed by exploring multiple stochastic samples across the learned text mass to find the representation closest to the candidate video representation.

    Input: Text query tt, Video clip vv, Sample size MM, Text encoder ϕt\phi_t, Frame encoder ϕv\phi_v, Feature fusion module ψ\psi, Radius module weights WW
    Output: Cross-modal similarity score $s^*
    Extract text feature t=ϕt(t)∈Rdt = \phi_t(t) \in \mathbb{R}^d
    Sample T′T' frames and extract frame features fi=ϕv(fi)∈Rdf_i = \phi_v(f_i) \in \mathbb{R}^d for i=1,…,T′i = 1, \dots, T'
    Compute video embedding v=ψ([f1,…,fT′],t)∈Rdv = \psi([f_1, \dots, f_{T'}], t) \in \mathbb{R}^d
    for i=1i = 1 to T′T' do
        Si=(t⊤fi)/(∥t∥2∥fi∥2)S_i = (t^\top f_i) / (\|t\|_2 \|f_i\|_2)
    end for
    S=[S1,…,ST′]∈RT′S = [S_1, \dots, S_{T'}] \in \mathbb{R}^{T'}
    R=exp⁡(SW)∈Rd\mathcal{R} = \exp(S W) \in \mathbb{R}^d
    for k=1k = 1 to MM do
        Sample ϵ(k)∼N(0,Id)\epsilon^{(k)} \sim \mathcal{N}(0, I_d)
        ts(k)=t+R⊙ϵ(k)t_s^{(k)} = t + \mathcal{R} \odot \epsilon^{(k)}
        sim(k)=(ts(k)⊤v)/(∥ts(k)∥2∥v∥2)\text{sim}^{(k)} = (t_s^{(k) \top} v) / (\|t_s^{(k)}\|_2 \|v\|_2)
    end for
    k∗=arg⁡max⁡k∈{1,…,M}sim(k)k^* = \arg\max_{k \in \{1, \dots, M\}} \text{sim}^{(k)}
    t^s=ts(k∗)\hat{t}_s = t_s^{(k^*)}
    s∗=sim(k∗)s^* = \text{sim}^{(k^*)}
    return s∗s^*

    In standard evaluations, M=20M=20 sampling trials are used to balance compute cost and retrieval performance.

  5. Knowl 5 — Text-to-Video Retrieval Performance on MSRVTT and LSMDC

    data/table

    T-MASS evaluated against competitive baselines for text-to-video retrieval on MSRVTT (1K-A split) and LSMDC (1000 test clips) across CLIP ViT-B/32 and ViT-B/16 backbones.

    Method MSRVTT Retrieval LSMDC Retrieval
    R@1 ↑\uparrow R@5 ↑\uparrow R@10 ↑\uparrow MdR ↓\downarrow MnR ↓\downarrow R@1 ↑\uparrow R@5 ↑\uparrow R@10 ↑\uparrow MdR ↓\downarrow MnR ↓\downarrow
    CLIP-ViT-B/32
    X-Pool 46.9 72.8 82.2 2.0 14.3 25.2 43.7 53.5 8.0 53.2
    DiffusionRet 49.0 75.2 82.7 2.0 12.1 24.4 43.1 54.3 8.0 40.7
    UATVR 47.5 73.9 83.5 2.0 12.3 – – – – –
    TEFAL 49.4 75.9 83.9 2.0 12.0 26.8 46.1 56.5 7.0 44.4
    CLIP-ViP 50.1 74.8 84.6 1.0 – 25.6 45.3 54.4 8.0 –
    T-MASS (Ours) 50.2 75.3 85.1 1.0 11.9 28.9 48.2 57.6 6.0 43.3
    CLIP-ViT-B/16
    X-Pool 48.2 73.7 82.6 2.0 12.7 26.1 46.8 56.7 7.0 47.3
    UATVR 50.8 76.3 85.5 1.0 12.4 – – – – –
    CLIP-ViP 54.2 77.2 84.8 1.0 – 29.4 50.6 59.0 5.0 –
    T-MASS (Ours) 52.7 77.1 85.6 1.0 10.5 30.3 52.2 61.3 5.0 40.1

    T-MASS consistently outperforms the baseline X-Pool (by +3.3%+3.3\% R@1 on MSRVTT and +3.7%+3.7\% R@1 on LSMDC with ViT-B/32). Unlike CLIP-ViP, T-MASS does not require auxiliary post-pretraining datasets (such as WebVid-2.5M or HD-VILA-100M).

  6. Knowl 6 — Text-to-Video Retrieval Performance on DiDeMo and VATEX

    data/table

    T-MASS performance on DiDeMo and VATEX benchmark datasets compared across CLIP ViT-B/32 and ViT-B/16 backbones.

    Method DiDeMo Retrieval VATEX Retrieval
    R@1 ↑\uparrow R@5 ↑\uparrow R@10 ↑\uparrow MdR ↓\downarrow MnR ↓\downarrow R@1 ↑\uparrow R@5 ↑\uparrow R@10 ↑\uparrow MdR ↓\downarrow MnR ↓\downarrow
    CLIP-ViT-B/32
    X-Pool 44.6 73.2 82.0 2.0 15.4 60.0 90.0 95.0 1.0 3.8
    DiffusionRet 46.7 74.7 82.7 2.0 14.3 – – – – –
    UATVR 43.1 71.8 82.3 2.0 15.1 61.3 91.0 95.6 1.0 3.3
    CLIP-ViP 48.6 77.1 84.4 2.0 – – – – – –
    T-MASS (Ours) 50.9 77.2 85.3 1.0 12.1 63.0 92.3 96.4 1.0 3.2
    CLIP-ViT-B/16
    X-Pool 47.3 74.8 82.8 2.0 14.2 62.6 91.7 96.0 1.0 3.4
    UATVR 45.8 73.7 83.3 2.0 13.5 64.5 92.6 96.8 1.0 2.8
    CLIP-ViP 50.5 78.4 87.1 1.0 – – – – – –
    T-MASS (Ours) 53.3 80.1 87.7 1.0 9.8 65.6 93.9 97.2 1.0 2.7

    T-MASS achieves top performance across both datasets and backbones, surpassing X-Pool by +6.3%+6.3\% R@1 on DiDeMo (ViT-B/32) and +6.0%+6.0\% R@1 (ViT-B/16), and improving R@1 on VATEX by +3.0%+3.0\% across both backbones.

  7. Knowl 7 — Retrieval Performance on Charades and Video-to-Text on MSRVTT

    data/table

    T-MASS results for text-to-video retrieval on the Charades dataset and video-to-text retrieval on the MSRVTT dataset.

    Charades Text-to-Video Retrieval
    Method R@1 ↑\uparrow R@5 ↑\uparrow R@10 ↑\uparrow MdR ↓\downarrow MnR ↓\downarrow
    CLIP-ViT-B/32
    ClipBERT 6.7 17.3 25.2 32.0 149.7
    CLIP4Clip 9.9 27.1 36.8 21.0 85.4
    X-Pool 11.2 28.3 38.8 20.0 82.7
    T-MASS (Ours) 14.2 36.2 48.3 12.0 54.8
    CLIP-ViT-B/16
    CLIP4Clip 16.0 38.2 48.5 12.0 54.1
    X-Pool 20.7 42.5 53.5 9.0 47.4
    T-MASS (Ours) 26.7 51.7 63.9 5.0 30.0
    MSRVTT Video-to-Text Retrieval
    Method R@1 ↑\uparrow R@5 ↑\uparrow R@10 ↑\uparrow MdR ↓\downarrow MnR ↓\downarrow
    CLIP-ViT-B/32
    CLIP4Clip 42.7 70.9 80.6 2.0 11.6
    CenterCLIP 42.8 71.7 82.2 2.0 10.9
    X-Pool 44.4 73.3 84.0 2.0 9.0
    TS2-Net 45.3 74.1 83.7 2.0 9.2
    DiffusionRet 47.7 73.8 84.5 2.0 8.8
    UATVR 46.9 73.8 83.8 2.0 8.6
    T-MASS (Ours) 47.7 78.0 86.3 2.0 8.0
    CLIP-ViT-B/16
    X-Pool 46.4 73.9 84.1 2.0 8.4
    TS2-Net 46.6 75.9 84.9 2.0 8.9
    CenterCLIP 47.7 75.0 83.3 2.0 10.2
    UATVR 48.1 76.3 85.4 2.0 8.0
    T-MASS (Ours) 50.9 80.2 88.0 1.0 7.4

    On Charades, T-MASS improves R@1 over X-Pool by +3.0%+3.0\% (ViT-B/32) and +6.0%+6.0\% (ViT-B/16), while reducing Mean Rank (MnR) from 47.4 to 30.0. On MSRVTT video-to-text retrieval, T-MASS achieves state-of-the-art results across all metrics (R@1 of 50.9%50.9\% with ViT-B/16).

  8. Knowl 8 — Ablation of Radius Formulations and Loss Components

    data/table

    Ablations on MSRVTT and DiDeMo analyzing the architecture of the similarity-aware radius R\mathcal{R}, training loss terms (deterministic cross-entropy Lce\mathcal{L}_{ce}, stochastic cross-entropy Ls\mathcal{L}_s, and support regularization Lsup\mathcal{L}_{\text{sup}}), and stochastic embedding replacement using CLIP ViT-B/32.

    (a) Radius Module Formulation R\mathcal{R}
    Radius R\mathcal{R} MSRVTT DiDeMo
    R@1 R@5 R@10 MdR MnR R@1 R@5 R@10 MdR MnR
    w/o R\mathcal{R} (baseline) 46.9 72.8 82.2 2.0 14.3 44.6 73.2 82.0 2.0 15.4
    exp⁡(1T′∑Si)\exp(\frac{1}{T'} \sum S_i) 48.7 74.7 83.7 2.0 12.7 48.0 75.4 85.0 2.0 13.0
    exp⁡(θT′∑Si)\exp(\frac{\theta}{T'} \sum S_i) 49.2 75.7 84.7 2.0 11.7 49.7 75.8 85.3 2.0 12.6
    exp⁡(SW)\exp(SW) 49.1 75.7 85.7 2.0 11.9 49.8 78.1 86.0 2.0 11.8
    (b) Text Representation and Loss Objective Ablation on MSRVTT
    tst_s Lce\mathcal{L}_{ce} Ls\mathcal{L}_s Lsup\mathcal{L}_{\text{sup}} R@1 R@5 R@10 MdR MnR –
    ✗ ✓ ✗ ✗ 46.9 72.8 82.2 2.0 14.3 –
    ✓ ✓ ✓ ✗ 48.5 74.8 84.3 2.0 12.3 –
    ✓ ✗ ✓ ✗ 49.1 75.7 85.7 2.0 11.9 –
    ✓ ✗ ✓ ✓ 50.2 75.3 85.1 1.0 11.9 –

    Key observations:

    1. Even an unlearnable similarity-based radius yields substantial gains over point embedding (+1.8%+1.8\% R@1 on MSRVTT and +3.4%+3.4\% on DiDeMo).
    2. Retaining the original deterministic contrastive loss Lce\mathcal{L}_{ce} alongside Ls\mathcal{L}_s degrades R@1 from 49.1%49.1\% to 48.5%48.5\%, as it biases mass regularization toward the single point tt.
    3. Introducing support text regularization Lsup\mathcal{L}_{\text{sup}} increases R@1 to 50.2%50.2\% and reduces Median Rank (MdR) to 1.01.0.
  9. Knowl 9 — Hyperparameter Dynamics of Inference Sampling, Support Loss Weight, and Frame Numbers

    empirical result

    Systematic evaluations show how T-MASS behaves across varying inference sampling trials MM, support regularization weights α\alpha, and input frame counts T′T':

    1. Inference sampling trials (MM) on MSRVTT-1K (ViT-B/32):

      • M=0M=0 (deterministic tt without stochastic sampling): R@1 = 44.4%44.4\%, R@5 = 72.4%72.4\%, R@10 = 81.9%81.9\%, MdR = 2.02.0, MnR = 13.113.1
      • M=5M=5: R@1 = 46.8%46.8\%, R@5 = 74.7%74.7\%, R@10 = 84.0%84.0\%, MdR = 2.02.0, MnR = 12.512.5
      • M=10M=10: R@1 = 50.0%50.0\%, R@5 = 75.2%75.2\%, R@10 = 84.1%84.1\%, MdR = 2.02.0, MnR = 12.312.3
      • M=20M=20: R@1 = 50.2%50.2\%, R@5 = 75.3%75.3\%, R@10 = 85.1%85.1\%, MdR = 1.01.0, MnR = 11.911.9 Performance stabilizes between M=10M=10 and M=20M=20, establishing M=20M=20 as the optimal trade-off point.
    2. Support regularization weight (α\alpha) on MSRVTT: Evaluating α∈{0.5,0.8,1.0,1.2,1.5}\alpha \in \{0.5, 0.8, 1.0, 1.2, 1.5\} demonstrates that retrieval performance peak occurs at α=1.2\alpha = 1.2, outperforming all lower and higher weight values.

    3. Input video frame count (T′T') on Charades: Testing T′∈{12,15,18,21,24}T' \in \{12, 15, 18, 21, 24\} shows consistent performance margins over baseline X-Pool across all frame counts, demonstrating that text mass modeling is robust to variations in temporal sampling density.

  10. Knowl 10 — Geometric Separation Dynamics of the Learned Text Mass

    empirical result

    Analysis of the embedding geometry reveals two distinct mechanisms through which T-MASS improves cross-modal alignment:

    1. Adaptive Radius Dynamics: Tracking the L1L_1-norm ∥R∥1\|\mathcal{R}\|_1 across 5 training epochs on MSRVTT-1K reveals that for true positive (relevant) text-video pairs, ∥R∥1\|\mathcal{R}\|_1 contracts to smaller values (around 420420--440440), forming a compact, precise semantic region. For 999 paired irrelevant videos, ∥R∥1\|\mathcal{R}\|_1 expands to significantly higher values (480480--510510).

    2. Suppression of Irrelevant Similarity and Cross-Entropy Reduction: Evaluating cosine similarities between queries and negative video candidates shows that stochastic text embeddings tst_s produce consistently lower maximum cosine similarity across 999 negative candidates than the deterministic embedding tt. For relevant text-video pairs, tst_s reduces the average text-to-video cross-entropy loss from 0.6350.635 (for tt) to 0.5880.588 (for tst_s).

Coverage note — None was omitted. All key methodological formulations, algorithms, ablation studies, primary benchmark comparisons across 5 datasets, and empirical geometry analyses have been covered.

References

  1. 1.Hassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang, Shih-Fu Chang, Yin Cui, and Boqing Gong. Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text. In NeurIPS, 2021.
  2. 2.Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell. Localizing moments in video with natural language. In ICCV, 2017.
  3. 3.Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman. Frozen in time: A joint video and image encoder for end-to-end retrieval. In ICCV, 2021.
  4. 4.Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman. A clip-hitchhiker’s guide to long video retrieval. arXiv, 2022.
  5. 5.Simion-Vlad Bogolin, Ioana Croitoru, Hailin Jin, Yang Liu, and Samuel Albanie. Cross modal retrieval with querybank normalisation. In CVPR, 2022.
  6. 6.Shizhe Chen, Yida Zhao, Qin Jin, and Qi Wu. Fine-grained video-text retrieval with hierarchical graph reasoning. In CVPR, 2020.
  7. 7.Xing Cheng, Hezheng Lin, Xiangyu Wu, Fan Yang, and Dong Shen. Improving video-text retrieval by multi-stream corpus alignment and dual softmax loss. arXiv, 2021.
  8. 8.Ioana Croitoru, Simion-Vlad Bogolin, Marius Leordeanu, Hailin Jin, Andrew Zisserman, Samuel Albanie, and Yang Liu. Teachtext: Crossmodal generalized distillation for text-video retrieval. In CVPR, 2021.
  9. 9.Chaorui Deng, Qi Chen, Pengda Qin, Da Chen, and Qi Wu. Prompt switch: Efficient clip adaptation for text-video retrieval. In ICCV, 2023.
  10. 10.Jianfeng Dong, Xianke Chen, Minsong Zhang, Xun Yang, Shujie Chen, Xirong Li, and Xun Wang. Partially relevant video retrieval. In ACM MM, 2022.
  11. 11.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv, 2020.
  12. 12.Maksim Dzabraev, Maksim Kalashnikov, Stepan Komkov, and Aleksandr Petiushko. Mdmmt: Multidomain multi-modal transformer for video retrieval. In CVPR, 2021.
  13. 13.Bo Fang, Chang Liu, Yu Zhou, Min Yang, Yuxin Song, Fu Li, Weiping Wang, Xiangyang Ji, Wanli Ouyang, et al. Uatvr: Uncertainty-adaptive text-video retrieval. In ICCV, 2023.
  14. 14.Han Fang, Pengfei Xiong, Luhui Xu, and Yu Chen. Clip2video: Mastering video-text retrieval via image clip. arXiv preprint arXiv:2106.11097, 2021.
  15. 15.Valentin Gabeur, Chen Sun, Karteek Alahari, and Cordelia Schmid. Multi-modal transformer for video retrieval. In ECCV, 2020.
  16. 16.Zijian Gao, Jingyu Liu, Sheng Chen, Dedan Chang, Hao Zhang, and Jinwei Yuan. CLIP2TV: an empirical study on transformer-based methods for video-text retrieval. CoRR, 2021.
  17. 17.Satya Krishna Gorti, Noël Vouitsis, Junwei Ma, Keyvan Golestan, Maksims Volkovs, Animesh Garg, and Guangwei Yu. X-pool: Cross-modal language-video attention for text-video retrieval. In CVPR, 2022.
  18. 18.Peiyan Guan, Renjing Pei, Bin Shao, Jianzhuang Liu, Weimian Li, Jiaxi Gu, Hang Xu, Songcen Xu, Youliang Yan, and Edmund Y Lam. Pidro: Parallel isomeric attention with dynamic routing for text-video retrieval. In ICCV, 2023.
  19. 19.Ning Han, Jingjing Chen, Hao Zhang, Huanwen Wang, and Hao Chen. Adversarial multi-grained embedding network for cross-modal text-video retrieval. ACM MM, 2022.
  20. 20.J. Huang, Y. Li, J. Feng, X. Wu, X. Sun, and R. Ji. Clover: Towards a unified video-language alignment and fusion model. In CVPR, 2023.
  21. 21.Sarah Ibrahimi, Xiaohang Sun, Pichao Wang, Amanmeet Garg, Ashutosh Sanan, and Mohamed Omar. Audio-enhanced text-to-video retrieval using text-conditioned feature alignment. In ICCV, 2023.
  22. 22.Y. Ji, R. Tu, J. Jiang, W. Kong, C. Cai, W. Zhao, H. Wang, Y. Yang, and W. Liu. Seeing what you miss: Vision-language pre-training with semantic completion learning. In CVPR, 2023.
  23. 23.Peng Jin, Jinfa Huang, Fenglin Liu, Xian Wu, Shen Ge, Guoli Song, David Clifton, and Jie Chen. Expectation-maximization contrastive learning for compact video-and-language representations. In NeurIPS, 2022.
  24. 24.Peng Jin, Jinfa Huang, Pengfei Xiong, Shangxuan Tian, Chang Liu, Xiangyang Ji, Li Yuan, and Jie Chen. Video-text as game players: Hierarchical banzhaf interaction for cross-modal representation learning. In CVPR, 2023.
  25. 25.Peng Jin, Hao Li, Zesen Cheng, Jinfa Huang, Zhennan Wang, Li Yuan, Chang Liu, and Jie Chen. Text-video retrieval with disentangled conceptualization and set-to-set alignment. In IJCAI, 2023.
  26. 26.Peng Jin, Hao Li, Zesen Cheng, Kehan Li, Xiangyang Ji, Chang Liu, Li Yuan, and Jie Chen. Diffusionret: Generative text-video retrieval with diffusion model. In ICCV, 2023.
  27. 27.Diederik P Kingma and Max Welling. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, 2013.
  28. 28.Jie Lei, Linjie Li, Luowei Zhou, Zhe Gan, Tamara L Berg, Mohit Bansal, and Jingjing Liu. Less is more: Clipbert for video-and-language learning via sparse sampling. In CVPR, 2021.
  29. 29.Linjie Li, Zhe Gan, Kevin Lin, Chung-Ching Lin, Zicheng Liu, Ce Liu, and Lijuan Wang. Lavender: Unifying video-language understanding as masked language modeling. In CVPR, 2023.
  30. 30.Pandeng Li, Chen-Wei Xie, Liming Zhao, Hongtao Xie, Jiannan Ge, Yun Zheng, Deli Zhao, and Yongdong Zhang. Progressive spatio-temporal prototype matching for text-video retrieval. In ICCV, 2023.
  31. 31.Y. Li, K. Min, S. Tripathi, and N. Vasconcelos. Svitt: Temporal learning of sparse video-text transformers. In CVPR, 2023.
  32. 32.Chengzhi Lin, Ancong Wu, Junwei Liang, Jun Zhang, Wenhang Ge, Wei-Shi Zheng, and Chunhua Shen. Text-adaptive multiple visual prototype matching for video-text retrieval. NeurIPS, 2022.
  33. 33.Yan-Bo Lin, Jie Lei, Mohit Bansal, and Gedas Bertasius. Eclipse: Efficient long-range video retrieval using sight and sound. In ECCV, 2022.
  34. 34.Yang Liu, Samuel Albanie, Arsha Nagrani, and Andrew Zisserman. Use what you have: Video retrieval using representations from collaborative experts. In BMVC, 2019.
  35. 35.Yu Liu, Huai Chen, Lianghua Huang, Di Chen, Bin Wang, Pan Pan, and Lisheng Wang. Animating images to transfer clip for video-text retrieval. In ACM SIGIR, 2022.
  36. 36.Yuqi Liu, Pengfei Xiong, Luhui Xu, Shengming Cao, and Qin Jin. Ts2-net: Token shift and selection transformer for text-video retrieval. In ECCV, 2022.
  37. 37.Ilya Loshchilov and Frank Hutter. Sgdr: Stochastic gradient descent with warm restarts. arXiv preprint arXiv:1608.03983, 2016.
  38. 38.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017.
  39. 39.Huaishao Luo, Lei Ji, Ming Zhong, Yang Chen, Wen Lei, Nan Duan, and Tianrui Li. Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning. In Neurocomputing, 2022.
  40. 40.Antoine Miech, Ivan Laptev, and Josef Sivic. Learning a text-video embedding from incomplete and heterogeneous data. arXiv, 2018.
  41. 41.Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2018.
  42. 42.R. Pei, J. Liu, W. Li, B. Shao, S. Xu, P. Dai, J. Lu, and Y. Yan. Clipping: Distilling clip-based models with a student base for video-language retrieval. In CVPR, 2023.
  43. 43.Jesús Andrés Portillo-Quintero, José Carlos Ortiz-Bayliss, and Hugo Terashima-Marín. A straightforward framework for video retrieval using clip. In MCPR, 2021.
  44. 44.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In ICML, 2021.
  45. 45.Anna Rohrbach, Marcus Rohrbach, Niket Tandon, and Bernt Schiele. A dataset for movie description. In CVPR, 2015.
  46. 46.Gunnar A Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta. Hollywood in homes: Crowdsourcing data collection for activity understanding. In ECCV, 2016.
  47. 47.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. NeurIPS, 2017.
  48. 48.Qiang Wang, Yanhao Zhang, Yun Zheng, Pan Pan, and Xian-Sheng Hua. Disentangled representation learning for text-video retrieval. arXiv, 2022.
  49. 49.Wenzhe Wang, Mengdan Zhang, Runnan Chen, Guanyu Cai, Penghao Zhou, Pai Peng, Xiaowei Guo, Jian Wu, and Xing Sun. Dig into multi-modal cues for video retrieval with hierarchical alignment. In IJCAI, 2021.
  50. 50.Xin Wang, Jiawei Wu, Junkun Chen, Lei Li, Yuan-Fang Wang, and William Yang Wang. Vatex: A large-scale, high-quality multilingual dataset for video-and-language research. In ICCV, 2019.
  51. 51.Xiaohan Wang, Linchao Zhu, and Yi Yang. T2vlad: global-local sequence alignment for text-video retrieval. In CVPR, 2021.
  52. 52.Ziyue Wang, Aozhu Chen, Fan Hu, and Xirong Li. Learn to understand negation in video retrieval. In ACM MM, 2022.
  53. 53.Wenhao Wu, Haipeng Luo, Bo Fang, Jingdong Wang, and Wanli Ouyang. Cap4video: What can auxiliary captions do for text-video retrieval? In CVPR, 2023.
  54. 54.Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko, Armen Aghajanyan, Florian Metze, Luke Zettlemoyer, and Christoph Feichtenhofer. VideoCLIP: Contrastive pre-training for zero-shot video-text understanding. In EMNLP, 2021.
  55. 55.Jun Xu, Tao Mei, Ting Yao, and Yong Rui. Msr-vtt: A large video description dataset for bridging video and language. In CVPR, 2016.
  56. 56.Hongwei Xue, Tiankai Hang, Yanhong Zeng, Yuchong Sun, Bei Liu, Huan Yang, Jianlong Fu, and Baining Guo. Advancing high-resolution video-language representation with large-scale video transcriptions. In CVPR, 2022.
  57. 57.Hongwei Xue, Yuchong Sun, Bei Liu, Jianlong Fu, Ruihua Song, Houqiang Li, and Jiebo Luo. Clip-vip: Adapting pre-trained image-text model to video-language representation alignment. In ICLR, 2023.
  58. 58.Youngjae Yu, Jongseok Kim, and Gunhee Kim. A joint sequence fusion model for video question answering and retrieval. In ECCV, 2018.
  59. 59.N. Zhao, J. Jiao, W. Xie, and D. Lin. Cali-nce: Boosting cross-modal video representation learning with calibrated alignment. In CVPRW, 2023.
  60. 60.Shuai Zhao, Linchao Zhu, Xiaohan Wang, and Yi Yang. Centerclip: Token clustering for efficient text-video retrieval. In ACM SIGIR, 2022.
  61. 61.Linchao Zhu and Yi Yang. Actbert: Learning global-local video-text representations. In CVPR, 2020.

Citation

MLA
Wang, J., et al. “Text Is MASS: Modeling as Stochastic Embedding for Text-Video Retrieval”. arXiv, 2024, http://arxiv.org/abs/2403.17998v1.
APA
Wang, J., Sun, G., Wang, P., Liu, D., Dianat, S., Rabbani, M., Rao, R., & Tao, Z. (2024). Text Is MASS: Modeling as Stochastic Embedding for Text-Video Retrieval. arXiv. http://arxiv.org/abs/2403.17998v1
Chicago
Wang, J., G. Sun, P. Wang, et al. 2024. “Text Is MASS: Modeling as Stochastic Embedding for Text-Video Retrieval”. arXiv. http://arxiv.org/abs/2403.17998v1.
Harvard
Wang, J. et al. (2024) “Text Is MASS: Modeling as Stochastic Embedding for Text-Video Retrieval”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2403.17998v1.
Vancouver
1. Wang J, Sun G, Wang P, Liu D, Dianat S, Rabbani M, Rao R, Tao Z (2024) Text Is MASS: Modeling as Stochastic Embedding for Text-Video Retrieval. arXiv

BibTeX

@article{wang2024text,
  title = {Text Is MASS: Modeling as Stochastic Embedding for Text-Video Retrieval},
  author = {Wang, Jiamian and Sun, Guohao and Wang, Pichao and Liu, Dongfang and Dianat, Sohail and Rabbani, Majid and Rao, Raghuveer and Tao, Zhiqiang},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2403.17998v1},
  eprint = {2403.17998}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE