Sequence to Sequence -- Video to Text

Subhashini VenugopalanMarcus RohrbachJeff DonahueRaymond MooneyTrevor DarrellKate Saenko

article2015ICCV1,502 citations

Proposes an end-to-end sequence-to-sequence framework using LSTMs to generate natural language captions directly from variable-length video frame sequences by modeling temporal structure across both visual inputs and text outputs.

Listen

Generating natural language descriptions for open-domain videos is a critical challenge for applications such as automated video indexing, human-robot interaction, and assistive narration for the visually impaired. Unlike static image captioning, video description must handle inputs of varying durations and understand complex temporal actions occurring across frames. Prior methods often relied on rigid predefined templates, collapsed time by averaging visual features across an entire clip, or used complex multi-stage pipelines that failed to capture rich linguistic nuances.

The main objective of the article is to demonstrate and evaluate a unified sequence-to-sequence framework called S2VT, which directly maps variable-length sequences of video frames to natural language sentences in an end-to-end learning setup. The system aims to capture temporal dynamics and learn an integrated language model without requiring predefined sentence templates or separate attention mechanisms.

The evaluated approach uses a stacked recurrent neural network architecture—specifically Long Short-Term Memory (LSTM) networks—operating in two sequential phases. In the encoding phase, pre-trained image convolutional networks process raw visual frames and optical flow motion images one by one, building a rich internal representation of the video over time. In the decoding phase, the model generates the sentence word by word conditioned on the encoded visual history. The researchers evaluated the system across three large, open-domain video benchmarks: the Microsoft Video Description (MSVD) YouTube corpus, the MPII Movie Description dataset (MPII-MD), and the Montreal Video Annotation Dataset (M-VAD).

The evaluation yielded several key findings regarding description quality and temporal modeling. First, on the MSVD benchmark, the combined model using raw frames and optical flow achieved a top METEOR accuracy score of 29.8%, outperforming prior template-based, frame-averaging, and complex attention-based baselines. Second, when tested with randomly shuffled video frames, performance dropped significantly from 29.2% to 28.2%, proving that the sequential architecture genuinely learns and benefits from temporal ordering. Third, integrating motion-based optical flow features with visual appearance boosted sentence generation accuracy compared to using appearance alone. Finally, on the challenging MPII-MD and M-VAD movie datasets, the model established state-of-the-art results with scores of 7.1% and 6.7% METEOR, respectively, exceeding prior translation-based and attention-guided models.

These findings imply that end-to-end sequential architectures can effectively bypass complex, hand-engineered feature pipelines and explicit attention mechanisms. This reduces system architectural complexity and engineering overhead while generating more natural and contextually appropriate sentences. Because the single-framework design shares learned parameters across visual encoding and language decoding, it provides a scalable, computationally streamlined path for automated video analysis and accessibility services.

Stakeholders developing automated video processing systems should adopt end-to-end sequential architectures rather than rigid multi-stage classifiers or static frame-averaging pipelines. Practitioners should integrate both motion flow and appearance data to maximize captioning fidelity. Furthermore, teams should scale up training data where feasible, as the underlying neural framework exhibits strong capacity gains when exposed to larger, diverse parallel video-sentence datasets.

Readers should note certain limitations and maintain calibrated expectations for deployment. Overall caption quality scores on movie datasets remain modest (around 7% METEOR), reflecting the high difficulty of open-domain video reasoning, diverse vocabularies, and subtle human actions. In addition, optical flow features alone struggled with domain shifts and polysemous verbs. While confidence is high in the model's architectural superiority over baseline methods, fully autonomous deployments in complex, high-risk settings will require ongoing validation and larger domain-specific datasets.

arXiv: 1505.00487examples/s2vt
Cover for Sequence to Sequence -- Video to Text

Abstract

Real-world videos often have complex dynamics; and methods for generating open-domain video descriptions should be sensitive to temporal structure and allow both input (sequence of frames) and output (sequence of words) of variable length. To approach this problem, we propose a novel end-to-end sequence-to-sequence model to generate captions for videos. For this we exploit recurrent neural networks, specifically LSTMs, which have demonstrated state-of-the-art performance in image caption generation. Our LSTM model is trained on video-sentence pairs and learns to associate a sequence of video frames to a sequence of words in order to generate a description of the event in the video clip. Our model naturally is able to learn the temporal structure of the sequence of frames as well as the sequence model of the generated sentences, i.e. a language model. We evaluate several variants of our model that exploit different visual features on a standard set of YouTube videos and two movie description datasets (M-VAD and MPII-MD).

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Approach
  • 3.1 LSTMs for sequence modeling
  • 3.2 Sequence to sequence video to text
  • 3.3 Video and text representation
  • 4 Experimental Setup
  • 4.1 Video description datasets
  • 4.1.1 Microsoft Video Description Corpus (MSVD)
  • 4.1.2 MPII Movie Description Dataset (MPII-MD)
  • 4.1.3 Montreal Video Annotation Dataset (M-VAD)
  • 4.2 Evaluation Metrics
  • 4.3 Experimental details of our models
  • 4.4 Related approaches
  • 5 Results and Discussion
  • 5.1 MSVD dataset
  • 5.2 Movie description datasets
  • 6 Conclusion
  • References

Knowls

  1. Knowl 1 — S2VT (Sequence-to-Sequence Video-to-Text) Architecture

    model/method

    The Sequence-to-Sequence Video-to-Text (S2VT) model maps a variable-length sequence of input video frames (x1,…,xn)(x_1, \dots, x_n) directly to a variable-length sequence of output words (y1,…,ym)(y_1, \dots, y_m) using a single stacked two-layer Long Short-Term Memory (LSTM) network that shares parameters across encoding and decoding phases.

    The architecture operates in two consecutive stages:

    1. Encoding Stage (time steps 11 to nn): The top LSTM layer receives visual feature embeddings representing successive video frames and encodes them into hidden states. At each time step t∈{1,…,n}t \in \{1, \dots, n\}, the second (lower) LSTM layer receives the hidden representation ht(1)h_t^{(1)} from the top LSTM layer concatenated with a null-padded text input vector (zeros, representing <pad>). During this stage, the network updates its internal cell states and hidden states without computing any prediction loss.

    2. Decoding Stage (time steps n+1n+1 to n+mn+m): Once all nn frames are consumed, the top LSTM layer receives zero-padding input vectors (<pad><pad>). The second LSTM layer is fed a start-of-sentence token (<BOS><BOS>) at step n+1n+1, prompting it to generate a sequence of natural language words. At each decoding step n+tn+t, the second LSTM takes as input the concatenated vector of the top layer's hidden state hn+t(1)h_{n+t}^{(1)} and the embedding of the previous word yt−1y_{t-1}. The hidden state output zn+tz_{n+t} of the second LSTM layer is projected via a softmax layer to yield a probability distribution over the vocabulary. The decoding continues until the model emits an end-of-sentence token (<EOS><EOS>).

  2. Knowl 2 — S2VT State Update and Training Objective Equations

    equation

    At each time step tt, an LSTM unit with input vector xtx_t, previous hidden state ht−1h_{t-1}, and previous cell state ct−1c_{t-1} computes its input gate iti_t, forget gate ftf_t, output gate oto_t, candidate state gtg_t, cell memory ctc_t, and hidden state hth_t according to:

    it=σ(Wxixt+Whiht−1+bi)i_t = \sigma(W_{xi} x_t + W_{hi} h_{t-1} + b_i)

    ft=σ(Wxfxt+Whfht−1+bf)f_t = \sigma(W_{xf} x_t + W_{hf} h_{t-1} + b_f)

    ot=σ(Wxoxt+Whoht−1+bo)o_t = \sigma(W_{xo} x_t + W_{ho} h_{t-1} + b_o)

    gt=ϕ(Wxgxt+Whght−1+bg)g_t = \phi(W_{xg} x_t + W_{hg} h_{t-1} + b_g)

    ct=ft⊙ct−1+it⊙gtc_t = f_t \odot c_{t-1} + i_t \odot g_t

    ht=ot⊙ϕ(ct)h_t = o_t \odot \phi(c_t)

    where σ(z)=(1+e−z)−1\sigma(z) = (1 + e^{-z})^{-1} is the sigmoid activation function, ϕ(z)=tanh⁡(z)\phi(z) = \tanh(z) is the hyperbolic tangent activation function, ⊙\odot represents element-wise multiplication, WijW_{ij} are trainable weight matrices, and bjb_j are trainable bias vectors.

    Given an input video frame sequence X=(x1,…,xn)X = (x_1, \dots, x_n) and target word sequence Y=(y1,…,ym)Y = (y_1, \dots, y_m), the probability of the sequence is modeled autoregressively during the decoding phase:

    p(y1,…,ym∣x1,…,xn)=∏t=1mp(yt∣hn+t−1,yt−1)p(y_1, \dots, y_m \mid x_1, \dots, x_n) = \prod_{t=1}^m p(y_t \mid h_{n+t-1}, y_{t-1})

    For a second-layer LSTM output representation zt∈Rdz_t \in \mathbb{R}^{d} and vocabulary VV, word emission probabilities are computed using a linear projection WyW_y and softmax:

    p(yt=y∣zt)=exp⁡(Wyzt)∑y′∈Vexp⁡(Wy′zt)p(y_t = y \mid z_t) = \frac{\exp(W_y z_t)}{\sum_{y' \in V} \exp(W_{y'} z_t)}

    Model parameters θ\theta are trained end-to-end to maximize the conditional log-likelihood of target sentences across training samples:

    θ∗=arg⁡max⁡θ∑t=1mlog⁡p(yt∣hn+t−1,yt−1;θ)\theta^* = \arg\max_\theta \sum_{t=1}^m \log p(y_t \mid h_{n+t-1}, y_{t-1}; \theta)

    Loss is evaluated only during the decoding steps (t=1,…,mt=1, \dots, m) and backpropagated through time across all encoding and decoding time steps.

  3. Knowl 3 — Visual and Text Feature Representations in S2VT

    model/method

    S2VT constructs continuous 500-dimensional input representations for RGB video frames, optical flow motion fields, and natural language words:

    1. RGB Frame Embeddings: Input video frames are scaled to 256×256256 \times 256 pixels and randomly cropped to 227×227227 \times 227 pixels. Spatial appearance features are extracted from the post-ReLU fc7 activations of pre-trained Convolutional Neural Networks (CNNs), specifically AlexNet (CaffeNet) or 16-layer VGGNet (trained on the 1.2M image ILSVRC-2012 ImageNet classification task). A trainable linear transformation projects these activations to a 500-dimensional embedding space, learned jointly with the LSTM layers.

    2. Optical Flow Embeddings: Motion between consecutive frames is captured via variational optical flow. The horizontal (uu) and vertical (vv) displacement components are centered around 128 and scaled to the range [0,255][0, 255]. The flow magnitude u2+v2\sqrt{u^2 + v^2} is computed and appended as a third channel to form a 3-channel optical flow image. A CNN pre-trained on the UCF101 action recognition dataset processes these flow images, and its fc6 layer activations are linearly embedded into the 500-dimensional input space of the LSTM.

    3. Word Embeddings: Target words are represented as 1-of-∣V∣|V| one-hot vectors and mapped to a 500-dimensional continuous embedding via a learned linear projection layer.

  4. Knowl 4 — Shallow Fusion of RGB and Optical Flow Models

    model/method

    To combine appearance and temporal motion cues, S2VT uses a shallow score fusion method during decoding inference rather than early feature concatenation.

    At each time step tt of sentence generation, the independent RGB-trained S2VT model and Optical Flow-trained S2VT model each output a probability distribution over the vocabulary VV. The combined probability score for emitting word y′y' is computed as a convex combination:

    p(yt=y′)=α⋅prgb(yt=y′)+(1−α)⋅pflow(yt=y′)p(y_t = y') = \alpha \cdot p_{\text{rgb}}(y_t = y') + (1 - \alpha) \cdot p_{\text{flow}}(y_t = y')

    where prgb(yt=y′)p_{\text{rgb}}(y_t = y') is the score predicted by the RGB network, pflow(yt=y′)p_{\text{flow}}(y_t = y') is the score predicted by the optical flow network, and α∈[0,1]\alpha \in [0, 1] is a scalar balancing hyperparameter tuned on the validation set.

  5. Knowl 5 — S2VT Training Protocol, Regularization, and Inference Setup

    experimental setup

    The stacked LSTM network uses 1000 hidden units per layer. During training, the LSTM stack is unrolled for a fixed total of 80 time steps.

    • Sampling and Padding: Video frames are sub-sampled at regular intervals: 1 frame every 10 frames for YouTube clips (MSVD) and 1 frame every 5 frames for movie clips (MPII-MD and M-VAD). If the total number of frames and words is under 80 steps, inputs are padded with zero vectors (<pad>). If the total exceeds 80, the input frame sequence is truncated.
    • Batch Size and Parameter Freezing: Mini-batch sizes reach up to 8 videos for AlexNet and up to 3 videos for optical flow models. When utilizing the 16-layer VGG network, all convolutional layers below fc7 are kept frozen to minimize memory consumption.
    • Optimization and Regularization: For MSVD, parameters are optimized using stochastic gradient descent (SGD). For movie description datasets (MPII-MD and M-VAD), dropout is applied to the inputs and outputs of both LSTM layers to prevent overfitting, and models are trained using ADAM with momentum parameters β1=0.9\beta_1 = 0.9 and β2=0.999\beta_2 = 0.999.
    • Inference: At test time, videos are not truncated, and all sampled frames are encoded. Decoding proceeds greedily by selecting arg⁡max⁡y′∈Vp(yt=y′∣zt)\arg\max_{y' \in V} p(y_t = y' \mid z_t) at each step until the <EOS> token is generated.
  6. Knowl 6 — Video Description Performance on Microsoft Video Description Corpus (MSVD)

    data/table

    Performance of S2VT variants and comparison baselines on the Microsoft Video Description (MSVD) corpus evaluated using the METEOR metric (METEOR 1.5, in %, higher is better):

    Model METEOR (%)
    Factor Graph Model (FGM) 23.9
    Mean pool - AlexNet 26.9
    Mean pool - VGG 27.7
    Mean pool - AlexNet COCO pre-trained 29.1
    Mean pool - GoogleNet 28.7
    Temporal attention - GoogleNet 29.0
    Temporal attention - GoogleNet + 3D-CNN 29.6
    S2VT - Flow (AlexNet) 24.3
    S2VT - RGB (AlexNet) 27.9
    S2VT - RGB (VGG) with random frame order 28.2
    S2VT - RGB (VGG) 29.2
    S2VT - RGB (VGG) + Flow (AlexNet) 29.8

    S2VT with RGB (AlexNet) scores 27.9% METEOR, outperforming mean pooling on AlexNet (26.9%) and VGG (27.7%). Shuffling the frame order during S2VT training (random frame order) degrades performance from 29.2% to 28.2%, verifying that S2VT actively leverages temporal sequence ordering. Combining RGB (VGG) and Flow (AlexNet) yields 29.8% METEOR, exceeding temporal soft-attention with 3D-CNNs.

  7. Knowl 7 — Video Description Performance on MPII-MD and M-VAD Movie Datasets

    data/table

    Performance comparison on the MPII Movie Description (MPII-MD) and Montreal Video Annotation Dataset (M-VAD) benchmarks measured in METEOR (%):

    Approach MPII-MD METEOR (%) M-VAD METEOR (%)
    SMT (best variant) 5.6 –
    Visual-Labels 7.0 6.3
    Temporal attention (GoogleNet + 3D-CNN) – 4.3
    Mean pool (VGG) 6.7 6.1
    S2VT: RGB (VGG) 7.1 6.7

    On MPII-MD, S2VT trained with RGB frames on VGG achieves 7.1% METEOR, improving over Statistical Machine Translation (5.6%) and mean-pooled VGG features (6.7%). On M-VAD, S2VT reaches 6.7% METEOR, outperforming temporal attention (4.3%), mean pooling (6.1%), and Visual-Labels (6.3%). On the combined Large Scale Movie Description Challenge (LSMDC) public test set, S2VT achieves 7.0% METEOR.

  8. Knowl 8 — Levenshtein Edit Distance Analysis of Sentence Originality

    data/table

    The percentage of generated test sentences matching at least one reference sentence in the training corpus within a Levenshtein edit distance kk (where kk denotes the number of token insertions, deletions, or substitutions):

    Dataset k=0k = 0 k≤1k \le 1 k≤2k \le 2 k≤3k \le 3
    MSVD 42.9% 81.2% 93.6% 96.6%
    MPII-MD 28.8% 43.5% 56.4% 83.0%
    M-VAD 15.6% 28.7% 37.8% 45.0%

    On the MSVD dataset (web video clips with multiple reference sentences per video), 42.9% of generated sentences replicate a training sentence verbatim (k=0k=0), and 81.2% differ by at most one word edit (k≤1k \le 1). In contrast, for the movie description corpora (MPII-MD and M-VAD, which have higher visual/linguistic diversity and mostly 1 reference per clip), verbatim reproduction drops to 28.8% and 15.6% respectively, showing substantially greater sentence composition novelty.

Coverage note — Deliberately omitted qualitative video captioning figures and general dataset collection statistics from prior benchmark papers as they do not constitute standalone methodological or empirical contributions.

References

  1. 1.H. Aradhye, G. Toderici, and J. Yagnik. Video2text: Learning to annotate video content. In ICDMW, 2009. 2
  2. 2.T. Brox, A. Bruhn, N. Papenberg, and J. Weickert. High accuracy optical flow estimation based on a theory for warping. In ECCV, pages 25–36, 2004. 2, 4
  3. 3.D. L. Chen and W. B. Dolan. Collecting highly parallel data for paraphrase evaluation. In ACL, 2011. 2, 5
  4. 4.X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Dollar, and C. L. Zitnick. Microsoft COCO captions: Data collection and evaluation server. arXiv:1504.00325, 2015. 5
  5. 5.X. Chen and C. L. Zitnick. Learning a recurrent visual representation for image caption generation. CVPR, 2015. 1
  6. 6.K. Cho, B. van Merriënboer, D. Bahdanau, and Y. Bengio. On the properties of neural machine translation: Encoder-decoder approaches. arXiv:1409.1259, 2014. 3
  7. 7.M. Denkowski and A. Lavie. Meteor universal: Language specific translation evaluation for any target language. In EACL, 2014. 5
  8. 8.J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell. Long-term recurrent convolutional networks for visual recognition and description. In CVPR, 2015. 1, 2, 3, 4
  9. 9.G. Gkioxari and J. Malik. Finding action tubes. 2014. 4
  10. 10.A. Graves and N. Jaitly. Towards end-to-end speech recognition with recurrent neural networks. In ICML, 2014. 1
  11. 11.S. Guadarrama, N. Krishnamoorthy, G. Malkarnenkar, S. Venugopalan, R. Mooney, T. Darrell, and K. Saenko. Youtube2text: Recognizing and describing arbitrary activities using semantic hierarchies and zero-shoot recognition. In ICCV, 2013. 1, 2
  12. 12.S. Hochreiter and J. Schmidhuber. Long short-term memory. Neural computation, 9(8), 1997. 1, 3
  13. 13.P. Hodosh, A. Young, M. Lai, and J. Hockenmaier. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. In TACL, 2014. 6
  14. 14.H. Huang, Y. Lu, F. Zhang, and S. Sun. A multi-modal clustering method for web videos. In ISCTCS. 2013. 2
  15. 15.Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell. Caffe: Convolutional architecture for fast feature embedding. ACMMM, 2014. 2
  16. 16.A. Karpathy and L. Fei-Fei. Deep visual-semantic alignments for generating image descriptions. CVPR, 2015. 1
  17. 17.D. Kingma and J. Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014. 7
  18. 18.R. Kiros, R. Salakhutdinov, and R. S. Zemel. Unifying visual-semantic embeddings with multimodal neural language models. arXiv:1411.2539, 2014. 1
  19. 19.N. Krishnamoorthy, G. Malkarnenkar, R. J. Mooney, K. Saenko, and S. Guadarrama. Generating natural-language video descriptions using text-mined knowledge. In AAAI, July 2013. 2
  20. 20.P. Kuznetsova, V. Ordonez, T. L. Berg, U. C. Hill, and Y. Choi. Treetalk: Composition and compression of trees for image descriptions. In TACL, 2014. 1
  21. 21.C.-Y. Lin. Rouge: A package for automatic evaluation of summaries. In Text Summarization Branches Out: Proceedings of the ACL-04 Workshop, pages 74–81, 2004. 5
  22. 22.T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick. Microsoft coco: Common objects in context. In ECCV, 2014. 6
  23. 23.J. Mao, W. Xu, Y. Yang, J. Wang, and A. L. Yuille. Deep captioning with multimodal recurrent neural networks (m-rnn). arXiv:1412.6632, 2014. 1
  24. 24.J. Y. Ng, M. J. Hausknecht, S. Vijayanarasimhan, O. Vinyals, R. Monga, and G. Toderici. Beyond short snippets: Deep networks for video classification. CVPR, 2015. 3, 4
  25. 25.P. Over, G. Awad, M. Michel, J. Fiscus, G. Sanders, B. Shaw, A. F. Smeaton, and G. Quéenot. TRECVID 2012 – an overview of the goals, tasks, data, evaluation mechanisms and metrics. In Proceedings of TRECVID 2012, 2012. 2
  26. 26.K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu. Bleu: a method for automatic evaluation of machine translation. In ACL, 2002. 5
  27. 27.A. Rohrbach, M. Rohrbach, and B. Schiele. The long-short story of movie description. GCPR, 2015. 7
  28. 28.A. Rohrbach, M. Rohrbach, N. Tandon, and B. Schiele. A dataset for movie description. In CVPR, 2015. 1, 2, 5, 7
  29. 29.M. Rohrbach, W. Qiu, I. Titov, S. Thater, M. Pinkal, and B. Schiele. Translating video content to natural language descriptions. In ICCV, 2013. 1
  30. 30.O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei. ILSVRC, 2014. 4
  31. 31.K. Simonyan and A. Zisserman. Two-stream convolutional networks for action recognition in videos. In NIPS, 2014. 2
  32. 32.K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. CoRR, abs/1409.1556, 2014. 4
  33. 33.N. Srivastava, E. Mansimov, and R. Salakhutdinov. Unsupervised learning of video representations using LSTMs. ICML, 2015. 2
  34. 34.I. Sutskever, O. Vinyals, and Q. V. Le. Sequence to sequence learning with neural networks. In NIPS, 2014. 1, 2, 3
  35. 35.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich. Going deeper with convolutions. CVPR, 2015. 6
  36. 36.J. Thomason, S. Venugopalan, S. Guadarrama, K. Saenko, and R. J. Mooney. Integrating language and vision to generate natural language descriptions of videos in the wild. In COLING, 2014. 2, 6
  37. 37.A. Torabi, C. Pal, H. Larochelle, and A. Courville. Using descriptive video services to create a large data source for video annotation research. arXiv:1503.01070v1, 2015. 2, 5
  38. 38.R. Vedantam, C. L. Zitnick, and D. Parikh. CIDEr: Consensus-based image description evaluation. CVPR, 2015. 5
  39. 39.S. Venugopalan, H. Xu, J. Donahue, M. Rohrbach, R. Mooney, and K. Saenko. Translating videos to natural language using deep recurrent neural networks. In NAACL, 2015. 1, 2, 4, 5, 6, 7
  40. 40.O. Vinyals, A. Toshev, S. Bengio, and D. Erhan. Show and tell: A neural image caption generator. CVPR, 2015. 1, 2, 4
  41. 41.H. Wang and C. Schmid. Action recognition with improved trajectories. In ICCV, pages 3551–3558. IEEE, 2013. 2
  42. 42.S. Wei, Y. Zhao, Z. Zhu, and N. Liu. Multimodal fusion for video search reranking. TKDE, 2010. 2
  43. 43.L. Yao, A. Torabi, K. Cho, N. Ballas, C. Pal, H. Larochelle, and A. Courville. Describing videos by exploiting temporal structure. arXiv:1502.08029v4, 2015. 1, 2, 4, 6, 7, 8
  44. 44.W. Zaremba and I. Sutskever. Learning to execute. arXiv:1410.4615, 2014. 3

Citation

MLA
Venugopalan, S., et al. “Sequence to Sequence -- Video to Text”. arXiv, 2015, http://arxiv.org/abs/1505.00487v3.
APA
Venugopalan, S., Rohrbach, M., Donahue, J., Mooney, R., Darrell, T., & Saenko, K. (2015). Sequence to Sequence -- Video to Text. arXiv. http://arxiv.org/abs/1505.00487v3
Chicago
Venugopalan, S., M. Rohrbach, J. Donahue, R. Mooney, T. Darrell, and K. Saenko. 2015. “Sequence to Sequence -- Video to Text”. arXiv. http://arxiv.org/abs/1505.00487v3.
Harvard
Venugopalan, S. et al. (2015) “Sequence to Sequence -- Video to Text”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1505.00487v3.
Vancouver
1. Venugopalan S, Rohrbach M, Donahue J, Mooney R, Darrell T, Saenko K (2015) Sequence to Sequence -- Video to Text. arXiv

BibTeX

@article{venugopalan2015sequence,
  title = {Sequence to Sequence -- Video to Text},
  author = {Venugopalan, Subhashini and Rohrbach, Marcus and Donahue, Jeff and Mooney, Raymond and Darrell, Trevor and Saenko, Kate},
  year = {2015},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1505.00487v3},
  eprint = {1505.00487}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE