Weakly Supervised Video Moment Localization with Contrastive Negative Sample Mining

Minghang ZhengYanjie HuangQingchao ChenYang Liu

article2022AAAI106 citations

Proposes a weakly supervised video moment localization framework that replaces rigid sliding windows with learnable Gaussian masks and mines hard intra-video negatives to better distinguish confusing segments without frame-level annotations.

Listen

Identifying specific moments in untrimmed video streams based on natural language queries is critical for technologies such as automated surveillance and robotics. Traditional fully supervised methods require precise, human-annotated start and end timestamps for every sentence, which is prohibitively costly and difficult to scale. Weakly supervised approaches bypass this bottleneck by learning from video-level descriptions alone, but existing frameworks face significant limitations: they rely on inefficient sliding-window proposals and train models against negative samples drawn from entirely different videos, failing to teach the system how to differentiate confusing, highly similar scenes occurring within the same video.

The article introduces and evaluates Contrastive Negative Sample Mining, a novel weakly supervised framework designed to accurately locate video moments using only video-level descriptions during training. The primary objective is to enhance moment retrieval precision and processing efficiency by generating dynamic, content-aware segment proposals and mining both easy and hard negative samples from within the target video itself.

The authors designed a two-part architecture comprising a mask generator and a mask-conditioned reconstructor. Rather than testing hundreds of arbitrary sliding windows, the system predicts a continuous, learnable bell-shaped temporal mask to represent the positive event. It treats the remaining unhighlighted video frames as easy negative samples and the entire unedited video as a hard negative sample. The system is evaluated by testing its ability to reconstruct masked text queries from these visual segments, trained with a specialized contrastive loss function that enforces proper ranking among positive, hard negative, and easy negative visual contexts. Experiments were conducted on two standard benchmarks: ActivityNet Captions, containing over 19,000 longer videos, and Charades-STA, containing over 12,000 shorter video-query pairs.

The findings show that the proposed method establishes state-of-the-art performance across key benchmark metrics. On ActivityNet Captions, the model achieved top localization accuracy, reaching 55.68% at a moderate overlap threshold and 33.33% at a strict overlap threshold, outperforming prior weakly supervised models without requiring detailed paragraph-level sequence annotations. Furthermore, replacing sliding windows with dynamic masks cut inference latency by more than half, processing a video in 55.8 milliseconds compared to 124 milliseconds for baseline methods. Ablation experiments confirmed that mining both easy and hard intra-video negatives provides crucial supervisory signals, improving mean localization overlap from 28.55% up to 37.14%.

These results demonstrate that weakly supervised systems can achieve high localization precision while substantially lowering data annotation costs and computational overhead. Organizations deploying video search and retrieval systems can eliminate expensive frame-by-frame timestamp labeling in favor of simpler video-level tagging, reducing deployment timelines and operational expenses without sacrificing retrieval quality.

Teams implementing video moment localization should transition from static sliding-window architectures to learnable mask generators and incorporate intra-video contrastive training. Future work should focus on addressing the framework's tendency to predict slightly longer boundaries than necessary due to its reconstruction objective, particularly in short-duration video environments. Confidence in the reported results is high for standard video benchmark tasks, though practitioners should account for performance variations across differing video lengths and domain contexts.

Cover for Weakly Supervised Video Moment Localization with Contrastive Negative Sample Mining

Abstract

Video moment localization aims at localizing the video segments which are most related to the given free-form natural language query. The weakly supervised setting, where only video level description is available during training, is getting more and more attention due to its lower annotation cost. Prior weakly supervised methods mainly use sliding windows to generate temporal proposals, which are independent of video content and low quality, and train the model to distinguish matched video-query pairs and unmatched ones collected from different videos, while neglecting what the model needs is to distinguish the unaligned segments within the video. In this work, we propose a novel weakly supervised solution by introducing Contrastive Negative sample Mining (CNM). Specifically, we use a learnable Gaussian mask to generate positive samples, highlighting the video frames most related to the query, and consider other frames of the video and the whole video as easy and hard negative samples respectively. We then train our network with the Intra-Video Contrastive loss to make our positive and negative samples more discriminative. Our method has two advantages: (1) Our proposal generation process with a learnable Gaussian mask is more efficient and makes our positive sample higher quality. (2) The more difficult intra-video negative samples enable our model to distinguish highly confusing scenes. Experiments on two datasets show the effectiveness of our method. Code can be found at https://github.com/minghangz/cnm.

Table of Contents

  • Introduction
  • Related Work
  • Approach
  • Mask Generator
  • Mask Conditioned Reconstructor
  • Model Training and Inference
  • Experiments
  • Datasets
  • Evaluation Metric
  • Implementation Details
  • Comparisons to the State-Of-The-Art
  • Ablation Study
  • Qualitative Results
  • Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Contrastive Negative Sample Mining framework

    model/method

    Contrastive Negative Sample Mining (CNM) is a weakly supervised video moment localization method that trains from an untrimmed video and its natural-language description without temporal start/end annotations. A multimodal mask generator predicts a query-dependent Gaussian mask over video frames as the positive sample. Frames suppressed by this mask form an easy negative sample, while the entire video forms a hard negative sample because it contains the relevant moment together with redundant content.

    A mask-conditioned reconstructor attempts to reconstruct the query from each masked video sample. The reconstruction losses are used as semantic relevance scores: the positive mask should reconstruct the query best, the whole-video hard negative should perform worse, and the outside-mask easy negative should perform worst. CNM trains these components with reconstruction learning and intra-video contrastive learning, thereby focusing supervision on distinguishing temporally misaligned segments within the same video rather than only mismatched videos.

  2. Knowl 2 — Content-dependent Gaussian mask generation

    model/method

    CNM represents a video as frame features V={vi}i=1N∈RN×DVV=\{v_i\}_{i=1}^{N}\in\mathbb{R}^{N\times D_V} and a query with MM word embeddings as W={wj}j=1M∈RM×DWW=\{w_j\}_{j=1}^{M}\in\mathbb{R}^{M\times D_W}, where NN is the number of sampled frames, DVD_V is the visual feature dimension, and DWD_W is the word-embedding dimension. A transformer encoder-decoder fuses the two modalities and produces frame-aligned features:

    H=D(V,E(W))={hi}i=1N∈RN×DH,H=D(V,E(W))=\{h_i\}_{i=1}^{N}\in\mathbb{R}^{N\times D_H},

    where EE is a transformer text encoder, DD is a multimodal transformer decoder, and DHD_H is the hidden dimension. The final fused feature hNh_N predicts a normalized temporal center c∈(0,1)c\in(0,1) and normalized width w∈(0,1)w\in(0,1) through fully connected layers followed by a sigmoid:

    c=Sigmoid⁡(FC⁡(hN)),w=Sigmoid⁡(FC⁡(hN)).c=\operatorname{Sigmoid}(\operatorname{FC}(h_N)),\qquad w=\operatorname{Sigmoid}(\operatorname{FC}(h_N)).

    The positive mask assigns a differentiable Gaussian weight to frame ii:

    mip=exp⁡(−α(i/N−c)2w2),i=1,…,N,m_i^p=\exp\left(-\frac{\alpha(i/N-c)^2}{w^2}\right),\qquad i=1,\ldots,N,

    where α>0\alpha>0 controls the Gaussian variance. The resulting proposal is content- and query-dependent, can express the temporal progression of an event, and is learned end-to-end rather than generated by enumerating sliding windows.

  3. Knowl 3 — Intra-video easy and hard negative construction

    model/method

    For a video with NN frames and a Gaussian positive mask mp∈RNm^p\in\mathbb{R}^{N}, CNM constructs two negative masks from the same video. The easy negative mask is the complement of the positive mask:

    me=1N−mp,m^e=\mathbf{1}_N-m^p,

    where 1N\mathbf{1}_N is the length-NN all-ones vector. It emphasizes frames outside the predicted query-related segment; these frames may still be difficult because they can share visual background or semantics with the positive moment. The hard negative mask is the entire video:

    mh=1N.m^h=\mathbf{1}_N.

    The hard negative contains the relevant moment together with irrelevant frames, so it is more difficult to distinguish from the positive sample and discourages overly long predictions. If R(m,W)R(m,W) denotes the semantic relevance of mask mm to query embeddings WW, CNM imposes the ordering

    R(mp,W)>R(mh,W)>R(me,W).R(m^p,W)>R(m^h,W)>R(m^e,W).

    This negative-mining scheme supplies temporally confusing negatives from the same video instead of relying exclusively on unrelated videos.

  4. Knowl 4 — Mask-conditioned attention for differentiable temporal conditioning

    model/method

    The mask-conditioned reconstructor replaces the standard self-attention in a transformer encoder with attention weighted by an arbitrary continuous mask m∈RNm\in\mathbb{R}^{N}. Given visual features V∈RN×DVV\in\mathbb{R}^{N\times D_V}, learned projections produce queries Qa∈RN×DHQ_a\in\mathbb{R}^{N\times D_H}, keys Ka∈RN×DHK_a\in\mathbb{R}^{N\times D_H}, and values Ua∈RN×DHU_a\in\mathbb{R}^{N\times D_H}. The unnormalized attention scores are

    A=QaKaTDH∈RN×N.A=\frac{Q_aK_a^{\mathsf T}}{\sqrt{D_H}}\in\mathbb{R}^{N\times N}.

    The mask is applied to every row of AA across its key dimension before a row-wise softmax:

    Em(V,m)=Softmax⁡row ⁣(A⊙(1NmT))Ua∈RN×DH,E_m(V,m)=\operatorname{Softmax}_{\mathrm{row}}\!\left(A\odot(\mathbf{1}_N m^{\mathsf T})\right)U_a\in\mathbb{R}^{N\times D_H},

    where ⊙\odot is elementwise multiplication and 1NmT\mathbf{1}_Nm^{\mathsf T} repeats the mask across all attention-query rows. Thus, frames with larger mask values contribute more to the contextual representation while the operation remains differentiable with respect to the predicted Gaussian mask. The decoder uses query-derived attention queries and keys and values derived from Em(V,m)E_m(V,m), allowing query words to collect information from the masked video context.

  5. Knowl 5 — Mask-conditioned semantic completion

    model/method

    CNM measures the relevance of a masked video segment by asking a transformer reconstructor to complete a partially masked version of the original query. During training, approximately one third of the query words are replaced by a special symbol, with nouns, verbs, and adjectives assigned higher replacement probability. Let W^\widehat W be the resulting masked query, and let mm be one of the positive, easy-negative, or hard-negative masks. The mask-conditioned decoder produces word-level hidden states

    Hm=Dm(W^,Em(V,m),m)∈RM×DH,H^m=D_m\bigl(\widehat W,E_m(V,m),m\bigr)\in\mathbb{R}^{M\times D_H},

    where MM is the query length and DHD_H is the hidden dimension. A fully connected layer followed by softmax predicts the next word over a vocabulary of size NwN_w:

    Pm(wi+1∣V,W^1:i)=Softmax⁡(FC⁡(Hm))i,i=1,…,M−1.P^m(w_{i+1}\mid V,\widehat W_{1:i}) =\operatorname{Softmax}(\operatorname{FC}(H^m))_i, \qquad i=1,\ldots,M-1.

    The reconstruction cross-entropy for mask mm is

    Lcem=−∑i=1M−1log⁡Pm(wi+1∣V,W^1:i),L_{\mathrm{ce}}^m=-\sum_{i=1}^{M-1}\log P^m(w_{i+1}\mid V,\widehat W_{1:i}),

    where wi+1w_{i+1} is the original next query token. The reconstructor is optimized with

    Lrec=Lcep+Lceh,L_{\mathrm{rec}}=L_{\mathrm{ce}}^p+L_{\mathrm{ce}}^h,

    using the positive and whole-video masks. The easy-negative reconstruction loss is not used to train the reconstructor because the easy-negative frames should not contain the query's correct moment.

  6. Knowl 6 — Intra-video contrastive loss and alternating optimization

    algorithm

    CNM trains the Gaussian mask generator and mask-conditioned reconstructor with different losses. Let LcepL_{\mathrm{ce}}^p, LcehL_{\mathrm{ce}}^h, and LceeL_{\mathrm{ce}}^e be the query-completion cross-entropies for the positive Gaussian mask, whole-video hard-negative mask, and outside-mask easy-negative mask, respectively. The Intra-Video Contrastive (IVC) loss is

    LIVC=max⁡(Lcep−Lceh+β1,0)+max⁡(Lcep−Lcee+β2,0),L_{\mathrm{IVC}}= \max\left(L_{\mathrm{ce}}^p-L_{\mathrm{ce}}^h+\beta_1,0\right) +\max\left(L_{\mathrm{ce}}^p-L_{\mathrm{ce}}^e+\beta_2,0\right),

    where β1<β2\beta_1<\beta_2. It requires the positive sample's reconstruction loss to be at least β1\beta_1 lower than the hard-negative loss and at least β2\beta_2 lower than the easy-negative loss.

    The training procedure alternates two parameter updates. With mask-generator parameters θ1\theta_1 fixed, the reconstructor parameters θ2\theta_2 are updated using LrecL_{\mathrm{rec}}. With θ2\theta_2 fixed, θ1\theta_1 is updated using LIVCL_{\mathrm{IVC}}. This separation prevents a trivial early-training solution in which the reconstructor simply assigns poor reconstruction scores to negative samples, which would otherwise make mask-generator optimization unreliable.

  7. Knowl 7 — Direct boundary inference without dense proposal post-processing

    algorithm

    At inference time, CNM applies its mask generator once to a video-query pair and obtains a normalized Gaussian center c∈(0,1)c\in(0,1) and width w∈(0,1)w\in(0,1). If the video is represented by NN sampled frames, the predicted start and end coordinates are

    s=max⁡(c−w/2,0)N,e=min⁡(c+w/2,1)N.s=\max(c-w/2,0)N,\qquad e=\min(c+w/2,1)N.

    The interval (s,e)(s,e) is the localized moment, converted to video-time coordinates according to the frame sampling scheme. CNM does not enumerate dense sliding-window proposals and therefore does not require non-maximum suppression or other proposal-selection post-processing.

  8. Knowl 8 — Datasets, metric, and implementation configuration

    experimental setup

    CNM was evaluated under weak supervision on ActivityNet Captions and Charades-STA. ActivityNet Captions contains 19,290 videos and 37,417/17,505/17,031 moments in the train/validation-1/validation-2 splits; the authors used validation-1 for validation and validation-2 for testing. Its average query, target-moment, and untrimmed-video lengths are 14 words, 36.2 seconds, and 117.6 seconds. Charades-STA contains 12,408 training and 3,720 test video-query pairs, with 5,338 training and 1,334 test videos; its average query, target-moment, and untrimmed-video lengths are 7.2 words, 8.1 seconds, and 30.6 seconds.

    Evaluation uses recall at temporal intersection-over-union thresholds: {0.1,0.3,0.5}\{0.1,0.3,0.5\} for ActivityNet Captions and {0.3,0.5,0.7}\{0.3,0.5,0.7\} for Charades-STA. Recall at threshold mm is the percentage of predicted moments whose temporal IoU with the ground truth exceeds mm.

    Visual features were pre-extracted with CLIP for ActivityNet Captions and I3D for Charades-STA; word features used pretrained GloVe embeddings. The maximum query length was 20 words, the maximum video length was 200 frames, and vocabulary sizes were 8,000 and 1,111 for ActivityNet Captions and Charades-STA, respectively. Both transformers used hidden dimension 256, four attention heads, and three layers. Adam optimization used learning rate 0.0004; β1=0.1\beta_1=0.1 and β2=0.15\beta_2=0.15 on both datasets. The Gaussian parameter was α=5\alpha=5 for ActivityNet Captions and α=5.5\alpha=5.5 for Charades-STA, whose predicted width was additionally capped at 0.45. The GloVe and visual backbone parameters were frozen during training.

  9. Knowl 9 — Weakly supervised benchmark performance

    data/table

    The proposed CNM was compared with weakly supervised video moment localization methods using recall percentages at multiple temporal-IoU thresholds. On ActivityNet Captions, CNM achieved the best reported recall at IoU thresholds 0.3 and 0.5, while its IoU-0.1 recall was slightly below CRM. On Charades-STA, CNM tied the best reported IoU-0.3 recall, but was below the strongest competing results at IoU-0.5 and IoU-0.7. CRM uses additional paragraph-level video annotations during training, whereas CNM uses only video-query pairs.

    ActivityNet Captions: Recall (%)
    Method IoU=0.1 IoU=0.3 IoU=0.5
    Random 38.23 18.64 7.63
    WS-DEC 62.71 41.98 23.34
    EC-SL 68.48 44.29 24.16
    MARN - 47.01 29.95
    SCN 71.48 47.23 29.22
    RTBPN 73.73 49.77 29.63
    WSTG 74.2 44.3 23.6
    WSLLN 75.4 42.8 22.7
    LCNet 78.58 48.49 26.33
    WSTAN 79.78 52.45 30.01
    CRM 81.61 55.26 32.19
    CNM 78.13 55.68 33.33
  10. Knowl 10 — Ablation evidence and observed prediction-length limitation

    empirical result

    Ablations on ActivityNet Captions show that both the learned mask generator and intra-video negative mining contribute to CNM. Replacing the mask generator with sliding-window proposals and a policy-gradient procedure substantially reduced recall at IoU 0.3 and 0.5. Removing either the hard or easy negative reduced performance at stricter IoU thresholds, while removing both caused the largest degradation. Separately optimizing the mask generator with the IVC loss and the reconstructor with the reconstruction loss was much better than jointly optimizing the entire model.

    Mask configuration IoU=0.1 IoU=0.3 IoU=0.5 mIoU
    Full model 78.13 55.68 33.33 37.14
    Without mask generator 79.35 47.71 26.98 34.73

    The learned-mask model ran at 55.8 ms per video on an NVIDIA TITAN X, compared with 124 ms per video for the sliding-window SCN-style alternative. The authors also report a limitation: reconstruction favors predictions that include more information, so CNM can output moments that are longer than the ground-truth interval; this is identified as a possible reason for its weaker Charades-STA performance at IoU 0.5 and 0.7.

Coverage note — Only qualitative example visualizations were omitted because they illustrate the quantitative and methodological findings without adding a separate reproducible contribution.

References

  1. 1.Balntas, V.; Riba, E.; Ponsa, D.; and Mikolajczyk, K. 2016. Learning local feature descriptors with triplets and shallow convolutional neural networks. In Bmvc, volume 1, 3.
  2. 2.Caba Heilbron, F.; Escorcia, V.; Ghanem, B.; and Carlos Niebles, J. 2015. Activitynet: A large-scale video benchmark for human activity understanding. In Proceedings of the ieee conference on computer vision and pattern recognition, 961–970.
  3. 3.Carreira, J.; and Zisserman, A. 2017. Quo vadis, action recognition? a new model and the kinetics dataset. In proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 6299–6308.
  4. 4.Chen, S.; and Jiang, Y.-G. 2021. Towards Bridging Event Captioner and Sentence Localizer for Weakly Supervised Dense Event Captioning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 8425–8435.
  5. 5.Chen, Z.; Ma, L.; Luo, W.; Tang, P.; and Wong, K.-Y. K. 2020. Look closer to ground better: Weakly-supervised temporal grounding of sentence in video. arXiv preprint arXiv:2001.09308.
  6. 6.Collins, R. T.; Lipton, A. J.; Kanade, T.; Fujiyoshi, H.; Duggins, D.; Tsin, Y.; Tolliver, D.; Enomoto, N.; Hasegawa, O.; Burt, P.; et al. 2000. A system for video surveillance and monitoring. VSAM final report, 2000(1-68): 1.
  7. 7.Duan, X.; Huang, W.; Gan, C.; Wang, J.; Zhu, W.; and Huang, J. 2018. Weakly supervised dense event captioning in videos. arXiv preprint arXiv:1812.03849.
  8. 8.Ellouz, F.; Adam, A.; Ciorbaru, R.; and Lederer, E. 1974. Minimal structural requirements for adjuvant activity of bacterial peptidoglycan derivatives. Biochemical and biophysical research communications, 59(4): 1317–1325.
  9. 9.Gao, J.; Sun, C.; Yang, Z.; and Nevatia, R. 2017. TALL: Temporal Activity Localization via Language Query. arXiv:1705.02101.
  10. 10.Gao, M.; Davis, L. S.; Socher, R.; and Xiong, C. 2019. Wslln: Weakly supervised natural language localization networks. arXiv preprint arXiv:1909.00239.
  11. 11.Huang, J.; Liu, Y.; Gong, S.; and Jin, H. 2021. Cross-Sentence Temporal and Semantic Relations in Video Activity Localisation. arXiv preprint arXiv:2107.11443.
  12. 12.Kemp, C. C.; Edsinger, A.; and Torres-Jara, E. 2007. Challenges for robot manipulation in human environments [grand challenges of robotics]. IEEE Robotics & Automation Magazine, 14(1): 20–29.
  13. 13.Krishna, R.; Hata, K.; Ren, F.; Fei-Fei, L.; and Niebles, J. C. 2017. Dense-Captioning Events in Videos. In 2017 IEEE International Conference on Computer Vision (ICCV).
  14. 14.Lin, Z.; Zhao, Z.; Zhang, Z.; Wang, Q.; and Liu, H. 2020. Weakly-Supervised Video Moment Retrieval via Semantic Completion Network. arXiv:1911.08199.
  15. 15.Ma, M.; Yoon, S.; Kim, J.; Lee, Y.; Kang, S.; and Yoo, C. D. 2020. Vlanet: Video-language alignment network for weakly-supervised video moment retrieval. In European Conference on Computer Vision, 156–171. Springer.
  16. 16.Mithun, N. C.; Paul, S.; and Roy-Chowdhury, A. K. 2019. Weakly Supervised Video Moment Retrieval From Text Queries. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019, 11592–11601. Computer Vision Foundation / IEEE.
  17. 17.Neubeck, A.; and Van Gool, L. 2006. Efficient non-maximum suppression. In 18th International Conference on Pattern Recognition (ICPR’06), volume 3, 850–855. IEEE.
  18. 18.Pennington, J.; Socher, R.; and Manning, C. 2014. Glove: Global Vectors for Word Representation. In Conference on Empirical Methods in Natural Language Processing.
  19. 19.Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I. 2021. Learning Transferable Visual Models From Natural Language Supervision. arXiv:2103.00020.
  20. 20.Rodriguez-Opazo, C.; Marrese-Taylor, E.; Fernando, B.; Li, H.; and Gould, S. 2020. DORi: Discovering Object Relationship for Moment Localization of a Natural-Language Query in Video. arXiv:2010.06260.
  21. 21.Song, Y.; Wang, J.; Ma, L.; Yu, Z.; and Yu, J. 2020. Weakly-supervised multi-level attentional reconstruction network for grounding textual queries in videos. arXiv preprint arXiv:2003.07048.
  22. 22.Tan, R.; Xu, H.; Saenko, K.; and Plummer, B. A. 2021. Logan: Latent graph co-attention network for weakly-supervised video moment retrieval. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2083–2092.
  23. 23.Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017. Attention Is All You Need. arXiv:1706.03762.
  24. 24.Wang, H.; Zha, Z.-J.; Chen, X.; Xiong, Z.; and Luo, J. 2020. Dual Path Interaction Network for Video Moment Localization. In Proceedings of the 28th ACM International Conference on Multimedia, MM ’20, 4116–4124. New York, NY, USA: Association for Computing Machinery. ISBN 9781450379885.
  25. 25.Wang, H.; Zha, Z.-J.; Li, L.; Liu, D.; and Luo, J. 2021a. Structured Multi-Level Interaction Network for Video Moment Localization via Language Query. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 7026–7035.
  26. 26.Wang, Y.; Deng, J.; Zhou, W.; and Li, H. 2021b. Weakly Supervised Temporal Adjacent Network for Language Grounding. IEEE Transactions on Multimedia.
  27. 27.Xiao, S.; Chen, L.; Zhang, S.; Ji, W.; Shao, J.; Ye, L.; and Xiao, J. 2021. Boundary Proposal Network for Two-Stage Natural Language Video Localization. arXiv:2103.08109.
  28. 28.Yang, W.; Zhang, T.; Zhang, Y.; and Wu, F. 2021. Local correspondence network for weakly supervised temporal sentence grounding. IEEE Transactions on Image Processing, 30: 3252–3262.
  29. 29.Zhang, M.; Yang, Y.; Chen, X.; Ji, Y.; Xu, X.; Li, J.; and Shen, H. T. 2021. Multi-Stage Aggregated Transformer Network for Temporal Language Localization in Videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 12669–12678.
  30. 30.Zhang, S.; Peng, H.; Fu, J.; and Luo, J. 2020a. Learning 2D Temporal Adjacent Networks for Moment Localization with Natural Language. arXiv:1912.03590.
  31. 31.Zhang, Z.; Lin, Z.; Zhao, Z.; Zhu, J.; and He, X. 2020b. Regularized two-branch proposal networks for weakly-supervised moment retrieval in videos. In Proceedings of the 28th ACM International Conference on Multimedia, 4098–4106.
  32. 32.Zhao, Y.; Zhao, Z.; Zhang, Z.; and Lin, Z. 2021. Cascaded Prediction Network via Segment Tree for Temporal Video Grounding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 4197–4206.
  33. 33.Zhou, H.; Zhang, C.; Luo, Y.; Chen, Y.; and Hu, C. 2021. Embracing Uncertainty: Decoupling and De-bias for Robust Temporal Grounding. arXiv:2103.16848.

Citation

MLA
Zheng, M., et al. “Weakly Supervised Video Moment Localization with Contrastive Negative Sample Mining”. Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, no. 3, 2022, pp. 3517–25, https://doi.org/10.1609/AAAI.V36I3.20263.
APA
Zheng, M., Huang, Y., Chen, Q., & Liu, Y. (2022). Weakly Supervised Video Moment Localization with Contrastive Negative Sample Mining. Proceedings of the AAAI Conference on Artificial Intelligence, 36(3), 3517–3525. https://doi.org/10.1609/AAAI.V36I3.20263
Chicago
Zheng, M., Y. Huang, Q. Chen, and Y. Liu. 2022. “Weakly Supervised Video Moment Localization with Contrastive Negative Sample Mining”. Proceedings of the AAAI Conference on Artificial Intelligence 36 (3): 3517–25. https://doi.org/10.1609/AAAI.V36I3.20263.
Harvard
Zheng, M. et al. (2022) “Weakly Supervised Video Moment Localization with Contrastive Negative Sample Mining”, Proceedings of the AAAI Conference on Artificial Intelligence, 36(3), pp. 3517–3525. Available at: https://doi.org/10.1609/AAAI.V36I3.20263.
Vancouver
1. Zheng M, Huang Y, Chen Q, Liu Y (2022) Weakly Supervised Video Moment Localization with Contrastive Negative Sample Mining. Proceedings of the AAAI Conference on Artificial Intelligence 36:3517–3525

BibTeX

@article{Zheng_2022, title={Weakly Supervised Video Moment Localization with Contrastive Negative Sample Mining}, volume={36}, ISSN={2159-5399}, url={http://dx.doi.org/10.1609/AAAI.V36I3.20263}, DOI={10.1609/aaai.v36i3.20263}, number={3}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, publisher={Association for the Advancement of Artificial Intelligence (AAAI)}, author={Zheng, Minghang and Huang, Yanjie and Chen, Qingchao and Liu, Yang}, year={2022}, month=June, pages={3517–3525} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF