MSR-VTT: A Large Video Description Dataset for Bridging Video and Language

Jun XuTao MeiTing YaoYong Rui

article2016CVPR2,574 citations

Introduces MSR-VTT, a large-scale video-to-text benchmark containing 10,000 open-domain video clips and 200,000 natural sentence annotations, paired with extensive evaluations showing that combining 2D spatial and 3D motion features with soft-attention pooling achieves superior video captioning performance.

Listen

The article addresses the challenge of automatically generating natural language descriptions for videos, noting that existing computer vision methods struggle with the variability and complexity of real-world video content. Current benchmarks are limited in scale, diversity, and domain coverage, which hinders progress compared to image captioning datasets.

The work set out to create a large-scale, representative video description dataset and to benchmark state-of-the-art recurrent neural network approaches for translating video to text.

Researchers collected 10,000 web video clips totaling 41.2 hours from 257 popular search queries spanning 20 categories. They obtained roughly 20 human-annotated sentences per clip through Amazon Mechanical Turk, yielding 200,000 clip-sentence pairs. They then evaluated multiple LSTM-based models that combined frame-level features from networks such as VGG and GoogleNet with temporal features from C3D, using both mean pooling and soft-attention strategies.

The resulting MSR-VTT dataset is substantially larger than prior collections in both sentences and vocabulary size, covers far more diverse real-world content, and includes audio channels. Models that fused C3D temporal features with VGG-19 spatial features and applied soft attention achieved the strongest results, reaching 40.5 BLEU@4 and 29.9 METEOR, outperforming mean-pooling baselines by roughly 1–2 points. Performance varied by category, with temporal features helping action-heavy content and attention helping multi-scene videos.

These findings matter because they supply the training data and evaluation framework needed to advance practical video captioning systems for search, accessibility, and summarization. The hybrid representation approach demonstrates that combining motion and appearance cues improves generalization on complex web video.

Next steps include incorporating audio information, developing methods that handle multi-scene or complex videos, and extending the dataset for tasks such as video summarization. The current 10K-clip version leaves room for further scaling, and results on the most challenging categories remain modest, indicating the need for continued algorithmic work.

Cover for MSR-VTT: A Large Video Description Dataset for Bridging Video and Language

Abstract

While there has been increasing interest in the task of describing video with natural language, current computer vision algorithms are still severely limited in terms of the variability and complexity of the videos and their associated language that they can recognize. This is in part due to the simplicity of current benchmarks, which mostly focus on specific fine-grained domains with limited videos and simple descriptions. While researchers have provided several benchmark datasets for image captioning, we are not aware of any large-scale video description dataset with comprehensive categories yet diverse video content.

In this paper we present MSR-VTT (standing for “MSR-Video to Text”) which is a new large-scale video benchmark for video understanding, especially the emerging task of translating video to text. This is achieved by collecting 257 popular queries from a commercial video search engine, with 118 videos for each query. In its current version, MSR-VTT provides 10K web video clips with 41.2 hours and 200K clip-sentence pairs in total, covering the most comprehensive categories and diverse visual content, and representing the largest dataset in terms of sentence and vocabulary. Each clip is annotated with about 20 natural sentences by 1,327 AMT workers. We present a detailed analysis of MSR-VTT in comparison to a complete set of existing datasets, together with a summarization of different state-of-the-art video-to-text approaches. We also provide an extensive evaluation of these approaches on this dataset, showing that the hybrid Recurrent Neural Network-based approach, which combines single-frame and motion representations with soft-attention pooling strategy, yields the best generalization capability on MSR-VTT.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. The MSR-VTT Datataset
  • 3.1. Collection of Representative Videos
  • 3.2. Clip Selection and Sentence Annotation
  • 3.3. Dataset Split
  • 3.4. Data Statistics
  • 4. Approaches to Video Descriptions
  • 5. Evaluations
  • 5.1. Experiment Settings
  • 5.2. Performance Comparison between Different Video Representations
  • 5.3. Performance Comparison between Different Pooling Strategies
  • 5.4. The Size of hidden Layer of LSTM
  • 5.5. Human Evaluations
  • 6. Conclusions
  • References

Knowls

  1. Knowl 1 — MSR-VTT Dataset Specification and Construction Pipeline

    definition

    The MSR-VTT (MSR-Video to Text) dataset is a large-scale video description benchmark designed for video-to-language translation and video understanding. In its 10K release (MSR-VTT-10K), the dataset comprises 10,000 video clips totaling 41.2 hours of footage derived from 7,180 distinct web videos, annotated with 200,000 clip-sentence pairs (20 unique natural language descriptions per clip) spanning 1,856,523 total words and a vocabulary of 29,316 unique words.

    The dataset spans 20 semantic categories: music, people, gaming, sports (actions), news (events/politics), education, TV shows, movie, animation, vehicles, how-to, travel, science (technology), animal, kids (family), documentary, food, cooking, beauty (fashion), and advertisement.

    The dataset curation pipeline proceeds through four main stages:

    1. Video Collection: 257 representative queries are gathered from a commercial video search engine across the 20 categories, and the top 150 search results per query are crawled. Removing duplicates, low-quality videos, and excessively short videos leaves 30,404 high-quality video files with audio intact.
    2. Shot Boundary Detection and Clip Selection: Color histogram-based shot segmentation detects 3,590,688 video shots. Human evaluators select consecutive shot sequences to form coherent clips of duration between 10 and 30 seconds (median: 2 shots per clip), with at most 3 clips selected per source video to guarantee diversity. This yields 30,000 candidate clips, from which 10,000 are sampled for the 10K dataset.
    3. Sentence Annotation: 1,327 Amazon Mechanical Turk (AMT) workers watch the clips and provide natural language sentences describing them. Duplicate and excessively short sentences are filtered out during post-processing to ensure exactly 20 distinct human descriptions per clip.
    4. Data Splitting: Video clips are partitioned based on search queries and source videos to ensure that clips from the same query or video do not cross split boundaries. The split follows a 65%:5%:30% ratio:
      • Training set: 6,513 clips (65%)
      • Validation set: 497 clips (5%)
      • Test set: 2,990 clips (30%)
  2. Knowl 2 — Cross-Dataset Comparison of Video Description Benchmarks

    data/table

    The scale, domain variety, and linguistic statistics of MSR-VTT-10K are compared below against standard video description benchmarks, including YouCook, TACoS, TACoS Multi-Level (TACoS M-L), M-VAD, MPII-MD, and MSVD.

    Dataset Context Sentence Source #Video #Clip #Sentence Vocabulary Duration (hrs)
    YouCook cooking labeled 88 – 2,668 2,711 2.3
    TACoS cooking AMT workers 123 7,206 18,227 – –
    TACoS M-L cooking AMT workers 185 14,105 52,593 – –
    M-VAD movie DVS 92 48,986 55,905 18,269 84.6
    MPII-MD movie DVS+Script 94 68,337 68,375 24,549 73.6
    MSVD multi-category AMT workers – 1,970 70,028 13,010 5.3
    MSR-VTT-10K 20 categories AMT workers 7,180 10,000 200,000 29,316 41.2

    MSR-VTT-10K contains the largest number of clip-sentence pairs (200,000200,000) and the largest distinct word vocabulary (29,31629,316), with 20 distinct human-annotated sentences per video clip across 20 general web categories. In contrast, existing corpora are either constrained to specialized narrow domains (cooking in YouCook and TACoS; movies with Descriptive Video Service [DVS] audio descriptions in M-VAD and MPII-MD) or offer limited clip counts (MSVD contains 1,970 clips without category labels).

  3. Knowl 3 — RNN-Based Video Captioning Framework with Spatial-Temporal Representations and Attention

    model/method

    A unified deep neural framework for translating video to natural language descriptions combines single-frame visual appearance representations, 3D clip-level spatiotemporal motion features, frame aggregation strategies, and Long Short-Term Memory (LSTM) decoders.

    1. Visual Appearance Representation: Single frames are passed through 2D Convolutional Neural Networks (pre-trained on ImageNet). Visual features are extracted from the 4,096-dimensional fc6\text{fc6} layer of AlexNet, VGG-16, or VGG-19, or from the pool5/7×7s1\text{pool5}/7\times 7\text{s1} layer of GoogleNet.
    2. Temporal and Motion Representation: Temporal and motion dynamics across consecutive video frames are extracted using a 3D Convolutional Neural Network (C3D) pre-trained on Sports-1M. Video clips are segmented into continuous 16-frame non-overlapping or overlapping blocks, each yielding a 4,096-dimensional representation from the C3D fc6\text{fc6} layer.
    3. Temporal Aggregation / Pooling Strategies:
      • Mean Pooling (MP-LSTM): Averages the frame-level and clip-level feature vectors uniformly across the temporal dimension into a single static video representation vector.
      • Soft Attention (SA-LSTM): Employs a dynamic soft attention mechanism to compute a weighted combination of frame/clip representations conditioned on the decoding state at each timestep, allowing the model to focus selectively on salient temporal segments.
    4. Feature Fusion: Spatial 2D CNN representations and spatiotemporal C3D representations are concatenated to form hybrid visual-motion feature vectors ((C3D+VGG-16)(\text{C3D} + \text{VGG-16}) or (C3D+VGG-19)(\text{C3D} + \text{VGG-19})).
    5. Sentence Decoder: A two-layer LSTM with a hidden layer size of dh=512d_h = 512 sequentially generates word tokens via softmax over a vocabulary of the ~20,000 most frequent words, conditioned on the visual inputs and previous word embeddings.
  4. Knowl 4 — Evaluation of Visual Representations and Pooling Mechanisms on MSR-VTT

    data/table

    Fourteen configurations comparing mean pooling (MP-LSTM) and soft attention (SA-LSTM) across seven visual representations on the MSR-VTT test set (2,990 clips) are evaluated using BLEU@1, BLEU@2, BLEU@3, BLEU@4, and METEOR (reported as percentages). All models use an LSTM hidden layer dimension of 512.

    Model BLEU@1 BLEU@2 BLEU@3 BLEU@4 METEOR
    MP-LSTM (AlexNet) 75.9 60.6 46.5 35.4 26.3
    MP-LSTM (GoogleNet) 76.8 61.3 47.2 36.7 27.5
    MP-LSTM (VGG-16) 78.0 62.0 48.7 37.2 28.6
    MP-LSTM (VGG-19) 78.2 62.2 48.9 37.3 28.7
    MP-LSTM (C3D) 79.8 64.7 51.7 39.9 29.3
    MP-LSTM (C3D+VGG-16) 79.8 64.7 52.0 40.1 29.4
    MP-LSTM (C3D+VGG-19) 79.9 64.9 52.1 40.1 29.5
    SA-LSTM (AlexNet) 76.9 61.1 46.8 35.8 27.0
    SA-LSTM (GoogleNet) 77.8 62.2 48.1 37.1 28.4
    SA-LSTM (VGG-16) 78.8 63.2 49.0 37.5 28.8
    SA-LSTM (VGG-19) 79.1 63.3 49.3 37.6 28.9
    SA-LSTM (C3D) 80.2 64.6 51.9 40.1 29.4
    SA-LSTM (C3D+VGG-16) 81.2 65.1 52.3 40.3 29.7
    SA-LSTM (C3D+VGG-19) 81.5 65.0 52.5 40.5 29.9

    Key findings include:

    1. 3D spatiotemporal representations (C3D) outperform individual 2D frame-level representations (AlexNet, GoogleNet, VGG-16, VGG-19) under both pooling schemes.
    2. Soft attention pooling consistently outperforms mean pooling across all feature encoders, improving METEOR by up to 1.4% (e.g., from 28.5% in MP-LSTM to 29.9% in SA-LSTM for C3D+VGG-19).
    3. Concatenating temporal C3D features with spatial VGG-19 features achieves the highest overall accuracy (BLEU@4 of 40.5% and METEOR of 29.9%).
  5. Knowl 5 — Category-Specific Performance Dynamics of Spatial, Temporal, and Attention Encodings

    empirical result

    Analysis of per-category METEOR performance across all 20 MSR-VTT categories reveals systematic interactions between video content structure and modeling choices:

    1. Temporal (C3D) vs. Spatial (VGG-19) Representations:

      • In dynamic, action-dense categories with high visual appearance variability, such as sports/actions, C3D temporal representations significantly outperform VGG-19 frame-based representations (54.4% METEOR for C3D vs. 51.1% for VGG-19).
      • In visually static or slow-paced categories with minimal motion, such as documentary, 2D spatial features from VGG-19 outperform C3D (36.9% METEOR for VGG-19 vs. 34.8% for C3D).
    2. Soft Attention (SA-LSTM) vs. Mean Pooling (MP-LSTM) (evaluated using C3D + VGG-19 features):

      • In heterogeneous categories characterized by multi-scene composition and rapid visual transitions, such as news (41.3% METEOR for SA-LSTM vs. 36.2% for MP-LSTM) and travel (36.2% for SA-LSTM vs. 31.2% for MP-LSTM), soft attention provides substantial performance advantages.
      • In homogeneous categories that typically occur within a single stationary setting, such as cooking, uniform mean pooling achieves comparable or superior results (37.4% METEOR for MP-LSTM vs. 37.8% for SA-LSTM).
  6. Knowl 6 — Comparison of Single-Frame and Frame-Pooling Inputs for Video Captioning

    data/table

    Evaluating caption generation quality using a single frame (the middle frame of each video clip) versus temporal aggregation across all frames using VGG-19 features with a 512-dimensional LSTM hidden layer on MSR-VTT yields:

    Feature / Pooling Method BLEU@4 (%) METEOR (%)
    Single frame (middle frame) 32.4 22.6
    Mean pooling 37.3 28.7
    Soft-Attention 37.6 28.9

    Aggregating multi-frame information via mean pooling improves BLEU@4 by +4.9%+4.9\% and METEOR by +6.1%+6.1\% over a single static middle frame. Soft attention further improves over single-frame input by +5.2%+5.2\% BLEU@4 and +6.3%+6.3\% METEOR.

  7. Knowl 7 — Impact of LSTM Hidden Layer Dimensionality on Video Caption Generation

    data/table

    Evaluating the effect of the LSTM hidden state dimension (dhd_h) on generation quality and model parameter size using clip-based C3D features and mean pooling on MSR-VTT yields:

    Hidden Layer Size BLEU@4 (%) METEOR (%) Parameters
    128 32.5 26.6 3.7M
    256 38.0 29.0 7.6M
    512 39.9 29.3 16.3M

    Performance on BLEU@4 and METEOR improves monotonically as the hidden layer size increases from 128 to 512, establishing dh=512d_h = 512 as the optimal operating setting despite the parameter increase to 16.3M.

  8. Knowl 8 — Human Evaluation and Sentence Consistency on MSR-VTT

    data/table

    Human evaluation of caption generation quality across different visual feature backbones was conducted using 5 human judges rating outputs on a 1–10 numerical scale, where lower scores indicate better quality:

    Feature Correctness Grammar Relevance
    AlexNet 7.8 7.0 7.9
    GoogleNet 6.2 6.8 6.4
    VGG-16 5.3 6.9 5.4
    VGG-19 5.4 6.7 5.2
    C3D 5.1 6.4 5.3
    C3D+VGG-16 5.1 6.1 5.0
    C3D+VGG-19 4.9 6.1 5.1

    The hybrid C3D+VGG-19\text{C3D}+\text{VGG-19} model achieves the best human evaluation scores across correctness (4.9), grammar (6.1), and relevance (5.1).

    Additionally, evaluating inter-annotator consistency across the 20 ground-truth human descriptions per clip by measuring Subject-Verb-Object (SVO) triplet overlap yields an average overlap percentage of 62.7%62.7\% across the dataset.

  9. Knowl 9 — Performance of Non-Parametric K-Nearest Neighbor (KNN) Baselines on MSR-VTT

    data/table

    Non-parametric K-Nearest Neighbor (KNN) retrieval baselines using mean-pooled video feature representations on MSR-VTT achieve the following captioning scores:

    Feature BLEU@4 (%) METEOR (%)
    AlexNet 6.3 14.1
    GoogleNet 8.1 15.2
    VGG-16 8.7 15.5
    VGG-19 7.3 14.5
    C3D 7.5 14.5

    Non-parametric retrieval baselines achieve at best 8.7% BLEU@4 and 15.5% METEOR (using VGG-16 features), significantly trailing parametric LSTM-based generation models (which achieve ~35–40% BLEU@4 and ~26–30% METEOR on equivalent representations).

Coverage note — None was omitted; all key contributions—the MSR-VTT dataset construction and statistics, dataset comparisons, video captioning framework architectures, comprehensive empirical results (Tables 1–7), per-category ablation analyses, and human evaluations—have been captured.

References

  1. 1.S. Banerjee and A. Lavie. METEOR: An automatic metric for mt evaluation with improved correlation with human judgments. In Proceedings of ACL Workshop, pages 65–72, 2005. 6
  2. 2.A. Barbu, A. Bridge, Z. Burchill, D. Coroian, S. Dickinson, S. Fidler, A. Michaux, S. Mussman, S. Narayanaswamy, D. Salvi, et al. Video in sentences out. Proceedings of UAI, 2012. 1, 3
  3. 3.D. L. Chen and W. B. Dolan. Collecting highly parallel data for paraphrase evaluation. In Proceedings of ACL, pages 190–200, 2011. 1, 3, 4
  4. 4.X. Chen and C. L. Zitnick. Mind’s Eye: A recurrent visual representation for image caption generation. In Proceedings of CVPR, 2015. 3
  5. 5.P. Das, C. Xu, R. F. Doell, and J. J. Corso. A thousand frames in just a few words: Lingual description of videos through latent topics and sparse object stitching. In Proceedings of CVPR, pages 2634–2641, 2013. 1, 3, 4
  6. 6.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. ImageNet: A large-scale hierarchical image database. In Proceedings of CVPR, pages 248–255, 2009. 5
  7. 7.J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell. Long-term recurrent convolutional networks for visual recognition and description. In Proceedings of CVPR, 2015. 1, 3
  8. 8.H. Fang, S. Gupta, F. Iandola, R. Srivastava, L. Deng, P. Dollár, J. Gao, X. He, M. Mitchell, J. Platt, et al. From captions to visual concepts and back. In Proceedings of CVPR, 2015. 1, 3
  9. 9.A. Farhadi, M. Hejrati, M. A. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. Forsyth. Every picture tells a story: Generating sentences from images. In Proceedings of ECCV, pages 15–29, 2010. 3
  10. 10.S. Ji, W. Xu, M. Yang, and K. Yu. 3D convolutional neural networks for human action recognition. IEEE Trans. on Pattern Analysis and Machine Intelligence, 35(1):221–231, 2013. 3
  11. 11.A. Karpathy and L. Fei-Fei. Deep visual-semantic alignments for generating image descriptions. In Proceedings of CVPR, 2015. 1, 3
  12. 12.A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei. Large-scale video classification with convolutional neural networks. In Proceedings of CVPR, pages 1725–1732, 2014. 5
  13. 13.R. Kiros, R. Salakhutdinov, and R. Zemel. Multimodal neural language models. In Proceedings of ICML, pages 595–603, 2014. 3
  14. 14.R. Kiros, R. Salakhutdinov, and R. S. Zemel. Unifying visual-semantic embeddings with multimodal neural language models. TACL, 2015. 1
  15. 15.A. Kojima, T. Tamura, and K. Fukunaga. Natural language description of human activities from video images based on concept hierarchy of actions. International Journal of Computer Vision, 50(2):171–184, 2002. 1, 3
  16. 16.N. Krishnamoorthy, K. S. Girish Malkarnenkar, Raymond J. Mooney, and S. Guadarrama. Generating natural-language video descriptions using text-mined knowledge. In Proceedings of AAAI, 2013. 1
  17. 17.A. Krizhevsky, I. Sutskever, and G. E. Hinton. Imagenet classification with deep convolutional neural networks. In Proceedings of NIPS, pages 1097–1105, 2012. 5
  18. 18.G. Kulkarni, V. Premraj, V. Ordonez, S. Dhar, S. Li, Y. Choi, A. C. Berg, and T. Berg. Babytalk: Understanding and generating simple image descriptions. IEEE Trans. on Pattern Analysis and Machine Intelligence, 35(12):2891–2903, 2013. 3
  19. 19.R. Lebret, P. O. Pinheiro, and R. Collobert. Phrase-based image captioning. Proceedings of ICML, 2015. 3
  20. 20.S. Li, G. Kulkarni, T. L. Berg, A. C. Berg, and Y. Choi. Composing simple image descriptions using web-scale N-grams. In Proceedings of International Conference on Computational Natural Language Learning, pages 220–228, 2011. 3
  21. 21.T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick. Microsoft COCO: Common objects in context. In Proceedings of ECCV, pages 740–755, 2014. 1
  22. 22.J. Mao, W. Xu, Y. Yang, J. Wang, and A. L. Yuille. Deep captioning with multimodal recurrent neural networks (m-rnn). In Proceedings of ICLR, 2015. 1, 3
  23. 23.T. Mei, Y. Rui, S. Li, and Q. Tian. Multimedia search reranking: A literature survey. ACM Computing Surveys (CSUR), 46(3):38, 2014. 4
  24. 24.K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu. BLEU: a method for automatic evaluation of machine translation. In Proceedings of ACL, pages 311–318, 2002. 6
  25. 25.M. Regneri, M. Rohrbach, D. Wetzel, S. Thater, B. Schiele, and M. Pinkal. Grounding action descriptions in videos. Transactions of the Association for Computational Linguistics, 1:25–36, 2013. 1, 3, 4
  26. 26.A. Rohrbach, M. Rohrbach, W. Qiu, A. Friedrich, M. Pinkal, and B. Schiele. Coherent multi-sentence video description with variable level of detail. Pattern Recognition, pages 184–195, 2014. 3, 4
  27. 27.A. Rohrbach, M. Rohrbach, N. Tandon, and B. Schiele. A dataset for movie description. Proceedings of CVPR, 2015. 1, 3, 4
  28. 28.M. Rohrbach, W. Qiu, I. Titov, S. Thater, M. Pinkal, and B. Schiele. Translating video content to natural language descriptions. In Proceedings of ICCV, pages 433–440, 2013. 1, 3, 4
  29. 29.K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. In Proceedings of ICLR, 2015. 5
  30. 30.C. Sun, C. Gan, and R. Nevatia. Automatic concept discovery from parallel text and visual corpora. In ICCV, pages 2596–2604, 2015. 3
  31. 31.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich. Going deeper with convolutions. In Proceedings of CVPR, 2015. 5
  32. 32.A. Torabi, C. J. Pal, H. Larochelle, and A. C. Courville. Using descriptive video services to create a large data source for video annotation research. arXiv:1503.01070, 2015. 1, 3, 4
  33. 33.D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri. C3D: generic features for video analysis. In Proceedings of ICCV, 2015. 5
  34. 34.S. Venugopalan, M. Rohrbach, J. Donahue, R. Mooney, T. Darrell, and K. Saenko. Sequence to sequence–video to text. In Proceedings of ICCV, 2015. 1, 3
  35. 35.S. Venugopalan, H. Xu, J. Donahue, M. Rohrbach, R. Mooney, and K. Saenko. Translating videos to natural language using deep recurrent neural networks. In Proceedings of ACL, 2015. 1, 3, 5, 7, 8
  36. 36.O. Vinyals, A. Toshev, S. Bengio, and D. Erhan. Show and Tell: A neural image caption generator. In Proceedings of CVPR, 2015. 1, 3
  37. 37.K. Xu, J. Ba, R. Kiros, A. Courville, R. Salakhutdinov, R. Zemel, and Y. Bengio. Show, attend and tell: Neural image caption generation with visual attention. In Proceedings of ICML, 2015. 3
  38. 38.Y. Yang, C. L. Teo, H. Daume III, and Y. Aloimonos. Corpus-guided sentence generation of natural images. In Proceedings of Intl Conference on Empirical Methods in Natural Language Processing, pages 444–454, 2011. 3
  39. 39.L. Yao, A. Torabi, K. Cho, N. Ballas, C. Pal, H. Larochelle, and A. Courville. Describing videos by exploiting temporal structure. In Proceedings of ICCV, 2015. 1, 3, 5, 7, 8
  40. 40.T. Yao, T. Mei, C.-W. Ngo, and S. Li. Annotation for free: Video tagging by mining user search behavior. In Proceedings of the 21st ACM international conference on Multimedia, pages 977–986. ACM, 2013. 3
  41. 41.P. Young, A. Lai, M. Hodosh, and J. Hockenmaier. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. Transactions of the Association for Computational Linguistics, 2:67–78, 2014. 1
  42. 42.Z.-J. Zha, T. Mei, Z. Wang, and X.-S. Hua. Building a comprehensive ontology to refine video concept detection. In Proceedings of the international workshop on Workshop on multimedia information retrieval, pages 227–236. ACM, 2007. 3

Citation

MLA
Xu, J., et al. “MSR-VTT: A Large Video Description Dataset for Bridging Video and Language”. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 5288–96, https://doi.org/10.1109/CVPR.2016.571.
APA
Xu, J., Mei, T., Yao, T., & Rui, Y. (2016). MSR-VTT: A Large Video Description Dataset for Bridging Video and Language. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 5288–5296. https://doi.org/10.1109/CVPR.2016.571
Chicago
Xu, J., T. Mei, T. Yao, and Y. Rui. 2016. “MSR-VTT: A Large Video Description Dataset for Bridging Video and Language”. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 5288–96. https://doi.org/10.1109/CVPR.2016.571.
Harvard
Xu, J. et al. (2016) “MSR-VTT: A Large Video Description Dataset for Bridging Video and Language”, 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 5288–5296. Available at: https://doi.org/10.1109/CVPR.2016.571.
Vancouver
1. Xu J, Mei T, Yao T, Rui Y (2016) MSR-VTT: A Large Video Description Dataset for Bridging Video and Language. In: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 5288–5296

BibTeX

@inproceedings{Xu_2016, title={MSR-VTT: A Large Video Description Dataset for Bridging Video and Language}, url={http://dx.doi.org/10.1109/CVPR.2016.571}, DOI={10.1109/cvpr.2016.571}, booktitle={2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Xu, Jun and Mei, Tao and Yao, Ting and Rui, Yong}, year={2016}, month=June, pages={5288–5296} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF