Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading Books

Yukun ZhuRyan KirosRichard ZemelRuslan SalakhutdinovRaquel UrtasunAntonio TorralbaSanja Fidler

article2015ICCV2,774 citations

Introduces a context-aware neural framework to align movie scenes with their corresponding book text, generating story-level visual descriptions that capture complex character states and narrative context beyond standard video captioning.

Listen

The article addresses the challenge of generating rich, story-like explanations for visual content in movies and images, which current caption datasets fail to provide due to their limited semantic depth. Books offer detailed descriptions of characters, scenes, emotions, and narrative evolution, but lack direct visual grounding; many books have been adapted into movies, creating an opportunity to align the two for better visual-language understanding.

The work evaluates methods to align movie shots and subtitle dialogs with book sentences, aiming to produce semantically meaningful correspondences that go beyond simple keyword matching. It develops neural sentence embeddings trained unsupervised on millions of sentences from a large book corpus, extends image-text embeddings to video using DVS descriptions, combines multiple similarity measures via a context-aware CNN, and applies a CRF to enforce timeline consistency.

Evaluation on a new dataset of 11 movie-book pairs with 2,070 annotated correspondences shows strong performance: the full model achieves mean recall of 69% at the paragraph level, doubling average precision compared to baselines, with each component (sentence embedding, visual features, context CNN, CRF) contributing measurable gains. The approach also retrieves the correct book for every movie tested and generates relevant story-like passages for movie clips and static CoCo images.

These findings matter because they enable practical applications such as interactive browsing between movies and books, richer video descriptions for accessibility or analysis, and story generation from visuals, advancing AI systems that ground high-level semantics in visual data. The results outperform prior hand-crafted alignment techniques and demonstrate the value of learned embeddings over keyword-based methods.

Next steps include scaling the alignment to larger corpora of books and movies, improving visual matching for cases where descriptions are sparse or verbose, and exploring extensions to other media or tasks like question answering. Main limitations are the non-exhaustive ground truth, subjective annotation of matches, and variability in how faithfully movies follow books; caution is warranted when generalizing beyond the 11 evaluated pairs.

Cover for Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading Books

Abstract

Books are a rich source of both fine-grained information, how a character, an object or a scene looks like, as well as high-level semantics, what someone is thinking, feeling and how these states evolve through a story. This paper aims to align books to their movie releases in order to provide rich descriptive explanations for visual content that go semantically far beyond the captions available in current datasets. To align movies and books we exploit a neural sentence embedding that is trained in an unsupervised way from a large corpus of books, as well as a video-text neural embedding for computing similarities between movie clips and sentences in the book. We propose a context-aware CNN to combine information from multiple sources. We demonstrate good quantitative performance for movie/book alignment and show several qualitative examples that showcase the diversity of tasks our model can be used for.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 The MovieBook and BookCorpus Datasets
  • 4 Aligning Books and Movies
  • 4.1 Skip-Thought Vectors
  • 4.2 Visual-semantic embeddings of clips and DVS
  • 4.3 Context aware similarity
  • 4.4 Global Movie/Book Alignment
  • 5 Experimental Evaluation
  • 5.1 Movie/Book Alignment
  • 5.2 Describing Movies via the Book
  • 5.3 Book “Retrieval”
  • 5.4 The CoCoBook: Writing Stories for CoCo
  • 6 Conclusion
  • A Qualitative Movie-Book Alignment Results
  • B Borrowing “Lines” from Other Books
  • C The CoCoBook
  • References

Knowls

  1. Knowl 1 — Skip-Thought Vectors for Unsupervised Sentence Representation

    model/method

    Skip-Thought Vectors is an unsupervised neural sentence embedding model that extends the word-level skip-gram objective to the sentence level. Given an ordered sequence of sentences, the model encodes a target sentence sis_i into a continuous vector representation and optimizes the probability of reconstructing its preceding sentence si−1s_{i-1} and succeeding sentence si+1s_{i+1}.

    Let sentence sis_i consist of words wi1,…,wiNw_i^1, \dots, w_i^N with word embeddings xi1,…,xiN∈Rd\mathbf{x}_i^1, \dots, \mathbf{x}_i^N \in \mathbb{R}^d.

    Encoder: The encoder is a Gated Recurrent Unit (GRU) RNN. Dropping the sentence index ii, the hidden state ht\mathbf{h}^t at time step tt is computed as: ht=(1−zt)⊙ht−1+zt⊙hˉt\mathbf{h}^t = (1 - \mathbf{z}^t) \odot \mathbf{h}^{t-1} + \mathbf{z}^t \odot \mathbf{\bar{h}}^t \mathbf{z}^t = \sigma(\mathbf{W}_z \mathbf{x}^t + \mathbf{U}_z \mathbf{h}^{t-1})$$$ \mathbf{r}^t = \sigma(\mathbf{W}_r \mathbf{x}^t + \mathbf{U}_r \mathbf{h}^{t-1})$$$ hˉt=tanh⁡(Wxt+U(rt⊙ht−1))\mathbf{\bar{h}}^t = \tanh(\mathbf{W} \mathbf{x}^t + \mathbf{U} (\mathbf{r}^t \odot \mathbf{h}^{t-1})) where zt\mathbf{z}^t is the update gate, rt\mathbf{r}^t is the reset gate, σ(⋅)\sigma(\cdot) is the element-wise sigmoid function, ⊙\odot denotes component-wise multiplication, and Wz,Uz,Wr,Ur,W,U\mathbf{W}_z, \mathbf{U}_z, \mathbf{W}_r, \mathbf{U}_r, \mathbf{W}, \mathbf{U} are trainable matrices. The final state hi=hN\mathbf{h}_i = \mathbf{h}^N serves as the sentence representation.

    Decoder: Two distinct GRU decoders sharing vocabulary projection matrix V\mathbf{V} decode si+1s_{i+1} and si−1s_{i-1}, conditioned on hi\mathbf{h}_i. For the succeeding sentence si+1s_{i+1}, decoder hidden state hi+1t\mathbf{h}_{i+1}^t is updated via: \mathbf{z}^t = \sigma(\mathbf{W}_z^d \mathbf{x}_{i+1}^{t-1} + \mathbf{U}_z^d \mathbf{h}_{i+1}^{t-1} + \mathbf{C}_z \mathbf{h}_i)$$$ \mathbf{r}^t = \sigma(\mathbf{W}r^d \mathbf{x}{i+1}^{t-1} + \mathbf{U}r^d \mathbf{h}{i+1}^{t-1} + \mathbf{C}_r \mathbf{h}_i) $$\mathbf{\bar{h}}^t = \tanh(\mathbf{W}^d \mathbf{x}_{i+1}^{t-1} + \mathbf{U}^d (\mathbf{r}^t \odot \mathbf{h}_{i+1}^{t-1}) + \mathbf{C} \mathbf{h}_i) hi+1t=(1−zt)⊙hi+1t−1+zt⊙hˉt\mathbf{h}_{i+1}^t = (1 - \mathbf{z}^t) \odot \mathbf{h}_{i+1}^{t-1} + \mathbf{z}^t \odot \mathbf{\bar{h}}^t The probability of generating word wi+1tw_{i+1}^t is: P(wi+1t∣wi+1<t,hi)∝exp⁡(vwi+1t⊤hi+1t)P(w_{i+1}^t \mid w_{i+1}^{<t}, \mathbf{h}_i) \propto \exp(\mathbf{v}_{w_{i+1}^t}^\top \mathbf{h}_{i+1}^t) where vw\mathbf{v}_w is the row of V\mathbf{V} for word index ww.

    Objective Function: The model optimizes the sum of conditional log-probabilities over all sentence triplets in the corpus using the Adam optimizer: L=∑tlog⁡P(wi+1t∣wi+1<t,hi)+∑tlog⁡P(wi−1t∣wi−1<t,hi)\mathcal{L} = \sum_t \log P(w_{i+1}^t \mid w_{i+1}^{<t}, \mathbf{h}_i) + \sum_t \log P(w_{i-1}^t \mid w_{i-1}^{<t}, \mathbf{h}_i) Once trained, the semantic similarity between two sentences sas_a and sbs_b is computed as the inner product ha⊤hb\mathbf{h}_a^\top \mathbf{h}_b between their normalized representations.

  2. Knowl 2 — Context-Aware CNN for Local Movie-Book Alignment

    model/method

    To resolve local ambiguities when aligning subtitle sentences and video shots to text in a book (such as repeated conversational phrases), a context-aware similarity model applies a 3-layer Convolutional Neural Network (CNN) across a 2D temporal context window.

    Let ii denote the index of a subtitle sentence (or movie shot) and jj denote the index of a book sentence. A 3D tensor S(i,j,m)∈RI×J×M\mathbf{S}(i, j, m) \in \mathbb{R}^{I \times J \times M} is formed across a spatial-temporal neighborhood of size II subtitle sentences around ii and JJ book sentences around jj, over M=9M = 9 similarity channels:

    1. Neural sentence embedding similarity (Skip-Thought inner product).
    2. Video-text visual-semantic embedding similarity between the movie shot and the book sentence.
    3. BLEU-1 score between the subtitle and book sentence.
    4. BLEU-2 score.
    5. BLEU-3 score.
    6. BLEU-4 score.
    7. BLEU-5 score.
    8. Term frequency-inverse document frequency (TF-IDF) similarity.
    9. Uniform linear timeline prior.

    A 3-layer CNN f(S(i,j,⋅))f(\mathbf{S}(i, j, \cdot)) processes this local 3D volume using ReLU activation functions, dropout regularization, and a terminal sigmoid layer to predict a fused alignment probability score(i,j)∈[0,1]\text{score}(i, j) \in [0, 1]. The network is trained with cross-entropy loss using the Adam optimization algorithm.

  3. Knowl 3 — Global Movie-Book Alignment via Chain Conditional Random Field (CRF)

    model/method

    Global alignment between a movie and a book is framed as inference in a linear-chain Conditional Random Field (CRF). This model enforces temporal progression and consistency while allowing non-linear narrative jumps such as flashbacks.

    Let the movie sequence consist of KK sequential shot/subtitle nodes i=1,…,Ki = 1, \dots, K. Each node takes a discrete state yi∈{1,…,L}y_i \in \{1, \dots, L\}, where LL is the total number of sentences in the book. The energy of an alignment configuration y=(y1,…,yK)\mathbf{y} = (y_1, \dots, y_K) is given by: E(y)=−log⁡p(x,y;ω)=∑i=1Kωuϕu(yi)+∑i=1K∑j∈N(i)ωpψp(yi,yj)E(\mathbf{y}) = -\log p(\mathbf{x}, \mathbf{y}; \boldsymbol{\omega}) = \sum_{i=1}^K \omega_u \phi_u(y_i) + \sum_{i=1}^K \sum_{j \in \mathcal{N}(i)} \omega_p \psi_p(y_i, y_j) where N(i)\mathcal{N}(i) denotes the temporal neighbors of node ii, and ω=(ωu,ωp)\boldsymbol{\omega} = (\omega_u, \omega_p) are trainable potential weights.

    • Unary Potential: ϕu(yi)\phi_u(y_i) is the direct score output from the Context-Aware CNN evaluating the match between movie shot/subtitle ii and book sentence yiy_i.
    • Pairwise Progression Potential: Penalizes discrepancies between the elapsed time in the movie subtitles and the relative distance traversed in the book text: ψp(yi,yj)=(ds(yi,yj)−db(yi,yj))2(ds(yi,yj)−db(yi,yj))2+σ2\psi_p(y_i, y_j) = \frac{(d_s(y_i, y_j) - d_b(y_i, y_j))^2}{(d_s(y_i, y_j) - d_b(y_i, y_j))^2 + \sigma^2} where ds(yi,yj)∈[0,1]d_s(y_i, y_j) \in [0, 1] is the normalized time duration between subtitle sentences ii and jj, db(yi,yj)=∣yi−yj∣/L∈[0,1]d_b(y_i, y_j) = |y_i - y_j| / L \in [0, 1] is the normalized distance in the book, and σ2\sigma^2 is a robustness threshold that prevents over-penalizing large narrative leaps.
    • Pairwise Continuity Potential: An additional potential ψq(yi,yj)=(db(yi,yj))2(db(yi,yj))2+σ2\psi_q(y_i, y_j) = \frac{(d_b(y_i, y_j))^2}{(d_b(y_i, y_j))^2 + \sigma^2} encourages state continuity between adjacent nodes when dialogue is absent (e.g., long silences).

    Inference and Learning: Exact maximum a posteriori (MAP) configuration is found via dynamic programming. Computational speed is improved by pruning candidate book sentence states that deviate by more than 1/31/3 of the book length from the uniform timeline. Model weights ω\boldsymbol{\omega} are learned using latent-variable structured prediction on sparse ground-truth annotations.

  4. Knowl 4 — Visual-Semantic Embeddings of Movie Clips and Descriptive Video Service (DVS)

    model/method

    To score similarities between visual movie shots and book sentences, a visual-semantic joint embedding space is trained using movie video clips and Descriptive Video Service (DVS) audio descriptions.

    Video Representation: For a movie clip qq, frame-level features are extracted using GoogLeNet and a hybrid-CNN (trained on the Places scene database). Frame features are mean-pooled across the entire clip to produce clip representation q\mathbf{q}. A linear projection maps this into the embedding space: v=WIq\mathbf{v} = \mathbf{W}_I \mathbf{q}

    Text Representation: DVS description sentences are tokenized (with character names replaced by a general someone token). Word embeddings xt\mathbf{x}^t are fed sequentially into an LSTM RNN: \mathbf{i}^t = \sigma(\mathbf{W}_{xi} \mathbf{x}^t + \mathbf{W}_{hi} \mathbf{m}^{t-1} + \mathbf{W}_{ci} \mathbf{c}^{t-1})$$$ \mathbf{f}^t = \sigma(\mathbf{W}{xf} \mathbf{x}^t + \mathbf{W}{hf} \mathbf{m}^{t-1} + \mathbf{W}_{cf} \mathbf{c}^{t-1}) $$\mathbf{a}^t = \tanh(\mathbf{W}_{xc} \mathbf{x}^t + \mathbf{W}_{hc} \mathbf{m}^{t-1}) ct=ft⊙ct−1+it⊙at\mathbf{c}^t = \mathbf{f}^t \odot \mathbf{c}^{t-1} + \mathbf{i}^t \odot \mathbf{a}^t \mathbf{o}^t = \sigma(\mathbf{W}_{xo} \mathbf{x}^t + \mathbf{W}_{ho} \mathbf{m}^{t-1} + \mathbf{W}_{co} \mathbf{c}^t)$$$ \mathbf{m}^t = \mathbf{o}^t \odot \tanh(\mathbf{c}^t)$$ The final hidden memory state m=mN\mathbf{m} = \mathbf{m}^N forms the sentence representation.

    Ranking Objective: With m\mathbf{m} and v\mathbf{v} normalized to unit ℓ2\ell_2-norm, similarity is defined as s(m,v)=m⋅vs(\mathbf{m}, \mathbf{v}) = \mathbf{m} \cdot \mathbf{v}. Parameters are optimized using mini-batch stochastic gradient descent without momentum via a bidirectional margin loss: min⁡θ∑m∑kmax⁡{0,α−s(m,v)+s(m,vk)}+∑v∑kmax⁡{0,α−s(v,m)+s(v,mk)}\min_\theta \sum_{\mathbf{m}} \sum_k \max\{0, \alpha - s(\mathbf{m}, \mathbf{v}) + s(\mathbf{m}, \mathbf{v}_k)\} + \sum_{\mathbf{v}} \sum_k \max\{0, \alpha - s(\mathbf{v}, \mathbf{m}) + s(\mathbf{v}, \mathbf{m}_k)\} where α\alpha is the margin hyperparameter and vk,mk\mathbf{v}_k, \mathbf{m}_k are contrastive (non-matching) clip and text embeddings.

  5. Knowl 5 — The MovieBook and BookCorpus Datasets

    data/table

    Two datasets provide the textual and multimodal grounding for movie-to-book alignment and sentence embedding:

    1. BookCorpus Dataset: Extracted from web books written by unpublished authors to train unsupervised sentence embeddings. Books under 20,000 words were excluded. It includes 11,038 books spanning 16 genres (e.g., Romance, Fantasy, Science Fiction, Teen) containing 74,004,228 sentences, 984,846,357 words, and 1,316,420 unique vocabulary words, with a mean of 13 words per sentence (median: 11).

    2. MovieBook Dataset: Contains 11 full-length movies aligned with their original books and time-stamped subtitle files. Annotators produced 2,070 ground-truth correspondences across 90 annotation hours:

    Movie / Book Title # Sentences # Words # Unique Words Avg Words/Sent Max Words/Sent # Paragraphs # Shots # Subtitle Sents # Dialog Align # Visual Align
    Gone Girl 12,603 148,340 3,849 15 153 3,927 2,604 2,555 76 106
    Fight Club 4,229 48,946 1,833 14 90 2,082 2,365 1,864 104 42
    No Country for Old Men 8,050 69,824 1,704 10 68 3,189 1,348 889 223 47
    Harry Potter 6,458 78,596 2,363 15 227 2,925 2,647 1,227 164 73
    Shawshank Redemption 2,562 40,140 1,360 18 115 637 1,252 1,879 44 12
    The Green Mile 9,467 133,241 3,043 17 119 2,760 2,350 1,846 208 102
    American Psycho 11,992 143,631 4,632 16 422 3,945 1,012 1,311 278 85
    One Flew Over the Cuckoo's Nest 7,103 112,978 2,949 19 192 2,236 1,671 1,553 64 25
    The Firm 15,498 135,529 3,685 11 85 5,223 2,423 1,775 82 60
    Brokeback Mountain 638 10,640 470 20 173 167 1,205 1,228 80 20
    The Road 6,638 58,793 1,580 10 74 2,345 1,108 782 126 49
    Total 85,238 980,658 9,032 15 156 29,436 19,985 16,909 1,449 621
  6. Knowl 6 — Quantitative Evaluation of Movie-Book Alignment Models

    empirical result

    Movie-to-book alignment performance was evaluated across 10 test movies (Gone Girl was used to train the CNN and CRF parameters). An alignment is deemed successfully recalled if the predicted match is within 3 paragraphs in the book and within 5 subtitle sentences in the movie. Average Precision (AP) across multiple threshold levels is also evaluated.

    The 3-layer Context-Aware CNN combined with the chain CRF achieves the highest mean recall (69.10%) and mean AP (23.17%), substantially outperforming the uniform timeline baseline (UNI: 3.88% recall, 0.40% AP) and a non-contextual linear SVM baseline (38.01% recall, 10.97% AP):

    Movie Title Metric UNI SVM CNN-3 Full CRF
    Fight Club AP 1.22 0.73 1.95 5.17
    Recall 2.36 10.38 17.92 19.81
    The Green Mile AP 0.00 14.05 28.80 27.60
    Recall 0.00 51.42 74.13 78.23
    Harry Potter AP 0.00 10.30 27.17 23.65
    Recall 0.00 44.35 76.57 78.66
    American Psycho AP 0.00 14.78 34.32 32.87
    Recall 0.27 34.25 81.92 80.27
    One Flew Over the Cuckoo's Nest AP 0.00 5.68 14.83 21.13
    Recall 1.01 25.25 49.49 54.55
    Shawshank Redemption AP 0.00 8.94 19.33 19.96
    Recall 1.79 46.43 94.64 96.79
    The Firm AP 0.05 4.46 18.34 20.74
    Recall 1.38 18.62 37.93 44.83
    Brokeback Mountain AP 2.36 24.91 31.80 30.58
    Recall 27.00 74.00 98.00 100.00
    The Road AP 0.00 13.77 19.80 19.58
    Recall 1.12 41.90 65.36 65.10
    No Country for Old Men AP 0.00 12.11 28.75 30.45
    Recall 1.12 33.46 71.69 72.79
    Mean AP 0.40 10.97 22.51 23.17
    Recall 3.88 38.01 66.77 69.10

    Ablating single feature types in a 1-layer CNN baseline (full feature baseline: 52.66% recall, 9.62% AP) demonstrates that all individual features contribute:

    • Without TF-IDF: Recall drops to 47.07% (AP: 5.94%)
    • Without BOOK (Skip-Thought sentence embeddings): Recall drops to 48.75% (AP: 8.57%)
    • Without VIS (GoogLeNet visual features): Recall drops to 50.03% (AP: 8.83%)
    • Without SCENE (hybrid-CNN scene features): Recall drops to 50.77% (AP: 9.31%)
    • Without PRIOR: Recall is 52.46% (AP: 9.64%)
    • Without BLEU: Recall is 52.95% (AP: 9.88%)

    Extending CNN depth from 1 layer to 3 layers (CNN-3) boosts recall from 52.66% to 66.77% (+14.11%) and AP from 9.62% to 22.51%, while the CRF temporal smoothing yields an additional +2.33% recall and +0.66% AP.

  7. Knowl 7 — Cross-Modal Book Retrieval Using Alignment Energy

    empirical result

    Cross-modal retrieval of a movie's corresponding book is performed by computing the full alignment energy with candidate books using the CRF model. The candidate book producing the lowest alignment energy is ranked highest, with scores scaled relative to the top choice (normalized to 100.0).

    Across all 10 test movies evaluated against the 10 candidate test books, the model correctly identifies the true source book as rank 1 (100% top-1 retrieval accuracy):

    Query Movie Rank 1 (Score) Rank 2 (Score) Rank 3 (Score) Rank 4 (Score) Rank 5 (Score) Rank 6 (Score) Rank 7 (Score)
    Fight Club Fight Club (100.0) No Country (45.4) One Flew (45.2) The Road (45.1) The Firm (43.6) American Psy (43.0) Shawshank (42.7)
    The Green Mile Green Mile (100.0) One Flew (42.5) American Psy (40.1) Shawshank (39.6) No Country (38.9) The Firm (38.0) The Road (36.7)
    Harry Potter Harry Potter (100.0) Brokeback (40.5) American Psy (39.7) The Road (39.5) One Flew (39.1) Shawshank (39.0) The Firm (38.7)
    American Psycho American Psy (100.0) The Firm (55.5) One Flew (54.9) Fight Club (53.5) Shawshank (53.1) No Country (52.6) Brokeback (51.3)
    One Flew... One Flew (100.0) The Firm (84.0) Harry Potter (80.8) The Road (79.1) Shawshank (79.0) Brokeback (77.8) No Country (76.9)
    Shawshank... Shawshank (100.0) The Firm (66.0) No Country (62.0) One Flew (61.4) The Road (60.9) Brokeback (59.1) Green Mile (58.0)
    The Firm The Firm (100.0) Shawshank (75.0) Fight Club (73.9) Brokeback (73.7) One Flew (71.5) American Psy (71.4) Harry Potter (68.5)
    Brokeback... Brokeback (100.0) One Flew (54.8) Fight Club (52.2) American Psy (51.9) Green Mile (50.9) The Firm (50.7) Shawshank (50.6)
    The Road The Road (100.0) The Firm (56.0) One Flew (55.9) No Country (54.8) Fight Club (54.1) Shawshank (53.9) Green Mile (53.4)
    No Country... No Country (100.0) The Road (49.7) Brokeback (49.5) One Flew (46.8) The Firm (46.4) Shawshank (45.8) Harry Potter (45.8)
  8. Knowl 8 — CoCoBook: Story-like Description Generation for Static Images

    model/method

    The CoCoBook pipeline generates rich descriptive narratives for static images by connecting factual visual captioning with the BookCorpus literature repository:

    1. Factual Caption Generation: A multimodal neural language model generates a baseline factual caption for an input image.
    2. Sentence Embedding Retrieval: The factual caption serves as a search query. The Skip-Thought sentence embedding model retrieves the top 10 nearest-neighbor sentences from a sample of several hundred thousand BookCorpus sentences via vector dot products.
    3. Precision Re-ranking: The candidate sentences are re-ranked based on 1-gram precision of non-stop words relative to the initial caption query.
    4. Context Window Expansion: For the top re-ranked sentence, a multi-sentence story passage is generated by returning the retrieved sentence along with its surrounding paragraph context (the 2 preceding sentences and the 2 succeeding sentences from the source book).
  9. Knowl 9 — Computational Runtimes of Movie-Book Alignment Pipeline

    data/table

    The computational runtime for each stage of the alignment pipeline per movie-book pair is distributed across feature extraction, similarity scoring, and training/inference:

    Pipeline Stage Computation Time per Movie-Book Pair Nature of Computation
    BLEU calculation 6 hours Subtitle vs book sentence matching
    VIS feature extraction 2 hours GoogLeNet frame mean-pooling
    SCENE feature extraction 1 hour hybrid-CNN frame mean-pooling
    TF-IDF computation 10 minutes Text term matching
    BOOK sentence embedding scoring 3 minutes Skip-Thought vector inner products
    Context-Aware CNN Training 3 minutes Optimization on single movie pair
    Context-Aware CNN Inference 0.2 minutes (12 seconds) Forward pass over similarity tensor
    CRF Model Training 5 hours Latent structural SVM optimization
    CRF Model Inference 5 minutes Dynamic programming with pruning

    Overall, BLEU computation and visual feature extraction (VIS and SCENE) account for approximately 80% of total offline preprocessing time. During representation training, the video-text model processes 1,440 movies per day, while the Skip-Thought sentence embedding model reads 870 books per day.

Coverage note — Deliberately omitted purely qualitative figures and exploratory movie-shot matching to 10-book and 200-book corpora as they provide qualitative examples rather than standalone quantitative models or results.

References

  1. 1.D. Bahdanau, K. Cho, and Y. Bengio. Neural machine translation by jointly learning to align and translate. ICLR, 2015. 4
  2. 2.K. Cho, B. van Merrienboer, C. Gulcehre, F. Bougares, H. Schwenk, and Y. Bengio. Learning phrase representations using rnn encoder-decoder for statistical machine translation. EMNLP, 2014. 4
  3. 3.J. Chung, C. Gulcehre, K. Cho, and Y. Bengio. Empirical evaluation of gated recurrent neural networks on sequence modeling. arXiv preprint arXiv:1412.3555, 2014. 4
  4. 4.T. Cour, C. Jordan, E. Miltsakaki, and B. Taskar. Movie/script: Alignment and parsing of video and text transcription. In ECCV, 2008. 2
  5. 5.M. Everingham, J. Sivic, and A. Zisserman. “Hello! My name is... Buffy” – Automatic Naming of Characters in TV Video. BMVC, pages 899–908, 2006. 2
  6. 6.A. Farhadi, M. Hejrati, M. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. Forsyth. Every picture tells a story: Generating sentences for images. In ECCV, 2010. 2
  7. 7.S. Fidler, A. Sharma, and R. Urtasun. A sentence is worth a thousand pixels. In CVPR, 2013. 2
  8. 8.A. Gupta and L. Davis. Beyond nouns: Exploiting prepositions and comparative adjectives for learning visual classifiers. In ECCV, 2008. 1
  9. 9.S. Hochreiter and J. Schmidhuber. Long short-term memory. Neural computation, 9(8):1735–1780, 1997. 4
  10. 10.N. Kalchbrenner and P. Blunsom. Recurrent continuous translation models. In EMNLP, pages 1700–1709, 2013. 4
  11. 11.A. Karpathy and L. Fei-Fei. Deep visual-semantic alignments for generating image descriptions. In CVPR, 2015. 1, 2
  12. 12.D. Kingma and J. Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014. 5
  13. 13.R. Kiros, R. Salakhutdinov, and R. S. Zemel. Unifying visual-semantic embeddings with multimodal neural language models. CoRR, abs/1411.2539, 2014. 1, 2, 3, 5, 9, 10
  14. 14.R. Kiros, Y. Zhu, R. Salakhutdinov, R. S. Zemel, A. Torralba, R. Urtasun, and S. Fidler. Skip-Thought Vectors. In Arxiv, 2015. 3, 4
  15. 15.C. Kong, D. Lin, M. Bansal, R. Urtasun, and S. Fidler. What are you talking about? text-to-image coreference. In CVPR, 2014. 1, 2
  16. 16.G. Kulkarni, V. Premraj, S. Dhar, S. Li, Y. Choi, A. Berg, and T. Berg. Baby talk: Understanding and generating simple image descriptions. In CVPR, 2011. 2
  17. 17.D. Lin, S. Fidler, C. Kong, and R. Urtasun. Visual Semantic Search: Retrieving Videos via Complex Textual Queries. CVPR, pages 2657–2664, 2014. 1, 2
  18. 18.T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Doll'ar, and C. L. Zitnick. Microsoft coco: Common objects in context. In ECCV, pages 740–755. 2014. 1, 19
  19. 19.X. Lin and D. Parikh. Don’t just listen, use your imagination: Leveraging visual common sense for non-visual tasks. In CVPR, 2015. 1
  20. 20.M. Malinowski and M. Fritz. A multi-world approach to question answering about real-world scenes based on uncertain input. In NIPS, 2014. 1
  21. 21.J. Mao, W. Xu, Y. Yang, J. Wang, and A. L. Yuille. Explain images with multimodal recurrent neural networks. In arXiv:1410.1090, 2014. 1, 2
  22. 22.T. Mikolov, K. Chen, G. Corrado, and J. Dean. Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781, 2013. 4
  23. 23.K. Papineni, S. Roukos, T. Ward, and W. J. Zhu. BLEU: a method for automatic evaluation of machine translation. In ACL, pages 311–318, 2002. 6
  24. 24.H. Pirsiavash, C. Vondrick, and A. Torralba. Inferring the why in images. arXiv.org, jun 2014. 2
  25. 25.V. Ramanathan, A. Joulin, P. Liang, and L. Fei-Fei. Linking People in Videos with “Their” Names Using Coreference Resolution. In ECCV, pages 95–110. 2014. 2
  26. 26.V. Ramanathan, P. Liang, and L. Fei-Fei. Video event understanding using natural language descriptions. In ICCV, 2013. 1
  27. 27.A. Rohrbach, M. Rohrbach, N. Tandon, and B. Schiele. A dataset for movie description. In CVPR, 2015. 2, 5
  28. 28.P. Sankar, C. V. Jawahar, and A. Zisserman. Subtitle-free Movie to Script Alignment. In BMVC, 2009. 2
  29. 29.A. Schwing, T. Hazan, M. Pollefeys, and R. Urtasun. Efficient Structured Prediction with Latent Variables for General Graphical Models. In ICML, 2012. 6
  30. 30.J. Sivic, M. Everingham, and A. Zisserman. “Who are you?” - Learning person specific classifiers from video. CVPR, pages 1145–1152, 2009. 2
  31. 31.I. Sutskever, O. Vinyals, and Q. V. Le. Sequence to sequence learning with neural networks. In NIPS, 2014. 4
  32. 32.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich. Going deeper with convolutions. arXiv preprint arXiv:1409.4842, 2014. 5
  33. 33.M. Tapaswi, M. Bauml, and R. Stiefelhagen. Book2Movie: Aligning Video scenes with Book chapters. In CVPR, 2015. 2
  34. 34.M. Tapaswi, M. Buml, and R. Stiefelhagen. Aligning Plot Synopses to Videos for Story-based Retrieval. IJMIR, 4:3–16, 2015. 1, 2, 6
  35. 35.S. Venugopalan, H. Xu, J. Donahue, M. Rohrbach, R. J. Mooney, and K. Saenko. Translating Videos to Natural Language Using Deep Recurrent Neural Networks. CoRR abs/1312.6229, cs.CV, 2014. 1, 2
  36. 36.O. Vinyals, A. Toshev, S. Bengio, and D. Erhan. Show and tell: A neural image caption generator. In arXiv:1411.4555, 2014. 1, 2
  37. 37.K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhutdinov, R. Zemel, and Y. Bengio. Show, attend and tell: Neural image caption generation with visual attention. In arXiv:1502.03044, 2015. 2
  38. 38.B. Zhou, A. Lapedriza, J. Xiao, A. Torralba, and A. Oliva. Learning Deep Features for Scene Recognition using Places Database. In NIPS, 2014. 5, 7

Citation

MLA
Zhu, Y., et al. “Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading Books”. 2015 IEEE International Conference on Computer Vision (ICCV), 2015, pp. 19–27, https://doi.org/10.1109/ICCV.2015.11.
APA
Zhu, Y., Kiros, R., Zemel, R., Salakhutdinov, R., Urtasun, R., Torralba, A., & Fidler, S. (2015). Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading Books. 2015 IEEE International Conference on Computer Vision (ICCV), 19–27. https://doi.org/10.1109/ICCV.2015.11
Chicago
Zhu, Y., R. Kiros, R. Zemel, et al. 2015. “Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading Books”. 2015 IEEE International Conference on Computer Vision (ICCV), 19–27. https://doi.org/10.1109/ICCV.2015.11.
Harvard
Zhu, Y. et al. (2015) “Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading Books”, 2015 IEEE International Conference on Computer Vision (ICCV). IEEE, pp. 19–27. Available at: https://doi.org/10.1109/ICCV.2015.11.
Vancouver
1. Zhu Y, Kiros R, Zemel R, Salakhutdinov R, Urtasun R, Torralba A, Fidler S (2015) Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading Books. In: 2015 IEEE International Conference on Computer Vision (ICCV). IEEE, pp 19–27

BibTeX

@inproceedings{Zhu_2015, title={Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading Books}, url={http://dx.doi.org/10.1109/ICCV.2015.11}, DOI={10.1109/iccv.2015.11}, booktitle={2015 IEEE International Conference on Computer Vision (ICCV)}, publisher={IEEE}, author={Zhu, Yukun and Kiros, Ryan and Zemel, Rich and Salakhutdinov, Ruslan and Urtasun, Raquel and Torralba, Antonio and Fidler, Sanja}, year={2015}, month=Dec, pages={19–27} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: IEEE