Weakly Supervised Video Moment Localization with Contrastive Negative Sample Mining
Minghang ZhengYanjie HuangQingchao ChenYang Liu
Proposes a weakly supervised video moment localization framework that replaces rigid sliding windows with learnable Gaussian masks and mines hard intra-video negatives to better distinguish confusing segments without frame-level annotations.
Identifying specific moments in untrimmed video streams based on natural language queries is critical for technologies such as automated surveillance and robotics. Traditional fully supervised methods require precise, human-annotated start and end timestamps for every sentence, which is prohibitively costly and difficult to scale. Weakly supervised approaches bypass this bottleneck by learning from video-level descriptions alone, but existing frameworks face significant limitations: they rely on inefficient sliding-window proposals and train models against negative samples drawn from entirely different videos, failing to teach the system how to differentiate confusing, highly similar scenes occurring within the same video.
The article introduces and evaluates Contrastive Negative Sample Mining, a novel weakly supervised framework designed to accurately locate video moments using only video-level descriptions during training. The primary objective is to enhance moment retrieval precision and processing efficiency by generating dynamic, content-aware segment proposals and mining both easy and hard negative samples from within the target video itself.
The authors designed a two-part architecture comprising a mask generator and a mask-conditioned reconstructor. Rather than testing hundreds of arbitrary sliding windows, the system predicts a continuous, learnable bell-shaped temporal mask to represent the positive event. It treats the remaining unhighlighted video frames as easy negative samples and the entire unedited video as a hard negative sample. The system is evaluated by testing its ability to reconstruct masked text queries from these visual segments, trained with a specialized contrastive loss function that enforces proper ranking among positive, hard negative, and easy negative visual contexts. Experiments were conducted on two standard benchmarks: ActivityNet Captions, containing over 19,000 longer videos, and Charades-STA, containing over 12,000 shorter video-query pairs.
The findings show that the proposed method establishes state-of-the-art performance across key benchmark metrics. On ActivityNet Captions, the model achieved top localization accuracy, reaching 55.68% at a moderate overlap threshold and 33.33% at a strict overlap threshold, outperforming prior weakly supervised models without requiring detailed paragraph-level sequence annotations. Furthermore, replacing sliding windows with dynamic masks cut inference latency by more than half, processing a video in 55.8 milliseconds compared to 124 milliseconds for baseline methods. Ablation experiments confirmed that mining both easy and hard intra-video negatives provides crucial supervisory signals, improving mean localization overlap from 28.55% up to 37.14%.
These results demonstrate that weakly supervised systems can achieve high localization precision while substantially lowering data annotation costs and computational overhead. Organizations deploying video search and retrieval systems can eliminate expensive frame-by-frame timestamp labeling in favor of simpler video-level tagging, reducing deployment timelines and operational expenses without sacrificing retrieval quality.
Teams implementing video moment localization should transition from static sliding-window architectures to learnable mask generators and incorporate intra-video contrastive training. Future work should focus on addressing the framework's tendency to predict slightly longer boundaries than necessary due to its reconstruction objective, particularly in short-duration video environments. Confidence in the reported results is high for standard video benchmark tasks, though practitioners should account for performance variations across differing video lengths and domain contexts.
- Paper: HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips, Antoine Miech et al. (2019). Introduces large-scale joint video-text embedding learning alongside intra-video negative sampling, which serves as a foundational precursor to contrastive negative sample mining.
- Paper: Momentum Contrast for Unsupervised Visual Representation Learning, Kaiming He et al. (2020). Establishes foundational contrastive representation learning paradigms and negative sampling mechanics that underlie modern intra-video contrastive loss formulations.
- Paper: Dense-Captioning Events in Videos, Ranjay Krishna et al. (2017). Introduces the core problem formulation of temporal event proposal generation and contextual language grounding across untrimmed videos.
- Paper: Stacked Cross Attention for Image-Text Matching, Kuang-Huei Lee et al. (2018). Pioneers cross-modal attention alignment and hard negative sample mining for vision-language matching.
- Paper: Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval, Max Bain et al. (2021). Establishes end-to-end transformer-based dual encoding for video-text retrieval, providing essential architectural context for multimodal video-language alignment.
- Paper: Proposal-Based Multiple Instance Learning for Weakly-Supervised Temporal Action Localization, Huan Ren et al. (2023). Extends weakly supervised video localization by shifting from segment-level multiple instance learning to proposal-level scoring to improve temporal boundary prediction.
- Paper: TR-DETR: Task-Reciprocal Transformer for Joint Moment Retrieval and Highlight Detection, Hao Sun et al. (2024). Builds upon query-guided temporal moment retrieval by jointly coupling moment retrieval with highlight detection in a reciprocal transformer framework.
- Paper: TRACE: Temporal Grounding Video LLM via Causal Event Modeling, Yongxin Guo 0001 et al. (2025). Generalizes temporal grounding to multimodal large language models by formalizing causal event modeling with structured timestamp and saliency score predictions.
- Paper: TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning, Xiangyu Zeng 0004 et al. (2025). Applies explicit temporal grounding and grounded tuning concepts to long-video multimodal models to improve moment localization and reduce hallucinations.
- Paper: Adaptive Keyframe Sampling for Long Video Understanding, Xi Tang et al. (2025). Explores query-relevant temporal frame filtering and keyframe scoring mechanisms to optimize long video understanding without full-sequence processing.
