Text Is MASS: Modeling as Stochastic Embedding for Text-Video Retrieval
Jiamian WangPichao WangGuohao SunDongfang LiuSohail A. DianatRaghuveer RaoMajid RabbaniZhiqiang Tao
Proposes T-MASS, a text-video retrieval framework that models concise text queries as stochastic embeddings with adaptive radii to capture the broad semantic scope of rich video content and achieve state-of-the-art retrieval accuracy across multiple benchmark datasets.
Text-video retrieval is increasingly vital as video content proliferates across digital platforms, yet the task faces a fundamental information mismatch. While videos contain complex, redundant visual information across multiple frames, search queries and captions are typically short and concise. Existing methods map both text and video to single fixed points in a shared mathematical space, which often fails because a single text point lacks the semantic richness needed to match the full scope of a video.
The article evaluates a novel retrieval method called T-MASS, which treats text not as a single deterministic point, but as an elastic, stochastic region termed a text mass. The main objective is to demonstrate that modeling text as a flexible semantic range improves alignment with rich video representations and significantly boosts retrieval accuracy without requiring expensive video pre-training or complex visual extraction architectures.
To test this concept, the authors developed a framework built upon standard pre-trained image-text models. The approach introduces a similarity-aware radius network that dynamically scales the size of the text mass based on the text-video pair, ensuring that closely related pairs form tight, precise representations while unrelated pairs remain distant. During training, the model uses a specialized learning objective and a support text regularization vector at the boundary of the mass to control its position and volume. The model was evaluated across five diverse benchmark video datasets using standard recall metrics and tested against leading state-of-the-art systems.
The experimental findings show substantial performance improvements across the board. T-MASS outperformed baseline methods by 3.0% to 6.3% in top-rank accuracy (Recall@1), achieving new state-of-the-art results on five standard benchmarks. On challenging datasets such as LSMDC and Charades, T-MASS improved top-1 recall by over 3.7% and 6.0% respectively compared to prior architectures. Furthermore, the analysis confirmed that the text mass mechanism effectively separates irrelevant pairs in the shared space while drawing true pairs closer together, achieving superior alignment even under varying video frame lengths.
These findings indicate that rethinking text modeling offers a highly cost-effective path to improving cross-modal search. Rather than expending heavy computational resources to pre-train large video encoders on millions of additional clips, organizations can achieve superior retrieval accuracy through flexible text representations. The system requires minimal structural modification to existing pipelines and maintains robust performance across different video inputs.
Organizations developing video search, recommendation, or digital asset management platforms should consider adopting stochastic text representations to upgrade legacy retrieval models. Implementing T-MASS requires balancing the number of inference samples, where testing 10 to 20 samples provides an optimal trade-off between search speed and retrieval accuracy. While the approach delivers strong, reliable gains on standard benchmarks, stakeholders should validate deployment latency under real-time production loads and evaluate whether combining text-mass modeling with large-scale pre-training yields further performance benefits.
- Paper: Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval, Max Bain et al. (2021). This paper establishes the foundational dual-encoder visual transformer architecture for joint video-text embedding spaces upon which stochastic text retrieval models build.
- Paper: MSR-VTT: A Large Video Description Dataset for Bridging Video and Language, Jun Xu et al. (2016). This work introduces MSR-VTT, the primary benchmark dataset used to evaluate and validate video-text retrieval representations and alignment methods.
- Paper: HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips, Antoine Miech et al. (2019). This study introduces large-scale joint text-video semantic embedding learning and provides key insights into cross-modal retrieval mechanics.
- Paper: Generating Sentences from a Continuous Space, Samuel R. Bowman et al. (2016). This work introduces continuous latent space modeling for text, providing the foundational conceptual basis for representing text as distributions rather than single point vectors.
- Paper: NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models, Chankyu Lee et al. (2025). This work advances generalist dense embedding architectures and contrastive instruction-tuning strategies that build upon modern multi-task retrieval paradigms.
- Paper: MMTEB: Massive Multilingual Text Embedding Benchmark, Kenneth C. Enevoldsen et al. (2025). This benchmark establishes a massive multilingual evaluation framework to test and generalize next-generation dense text and retrieval embeddings.
- Paper: Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis, Chaoyou Fu et al. (2025). This paper extends video-text comprehension evaluation to comprehensive multimodal LLM benchmarks across complex, long-form video reasoning tasks.
