Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading Books
Yukun ZhuRyan KirosRichard ZemelRuslan SalakhutdinovRaquel UrtasunAntonio TorralbaSanja Fidler
Introduces a context-aware neural framework to align movie scenes with their corresponding book text, generating story-level visual descriptions that capture complex character states and narrative context beyond standard video captioning.
The article addresses the challenge of generating rich, story-like explanations for visual content in movies and images, which current caption datasets fail to provide due to their limited semantic depth. Books offer detailed descriptions of characters, scenes, emotions, and narrative evolution, but lack direct visual grounding; many books have been adapted into movies, creating an opportunity to align the two for better visual-language understanding.
The work evaluates methods to align movie shots and subtitle dialogs with book sentences, aiming to produce semantically meaningful correspondences that go beyond simple keyword matching. It develops neural sentence embeddings trained unsupervised on millions of sentences from a large book corpus, extends image-text embeddings to video using DVS descriptions, combines multiple similarity measures via a context-aware CNN, and applies a CRF to enforce timeline consistency.
Evaluation on a new dataset of 11 movie-book pairs with 2,070 annotated correspondences shows strong performance: the full model achieves mean recall of 69% at the paragraph level, doubling average precision compared to baselines, with each component (sentence embedding, visual features, context CNN, CRF) contributing measurable gains. The approach also retrieves the correct book for every movie tested and generates relevant story-like passages for movie clips and static CoCo images.
These findings matter because they enable practical applications such as interactive browsing between movies and books, richer video descriptions for accessibility or analysis, and story generation from visuals, advancing AI systems that ground high-level semantics in visual data. The results outperform prior hand-crafted alignment techniques and demonstrate the value of learned embeddings over keyword-based methods.
Next steps include scaling the alignment to larger corpora of books and movies, improving visual matching for cases where descriptions are sparse or verbose, and exploring extensions to other media or tasks like question answering. Main limitations are the non-exhaustive ground truth, subjective annotation of matches, and variability in how faithfully movies follow books; caution is warranted when generalizing beyond the 11 evaluated pairs.
- Paper: Skip-Thought Vectors, Ryan Kiros et al. (2015). This paper introduces the unsupervised skip-thought sentence embedding framework trained on large novel corpora that directly underpins the neural sentence representation methods utilized in the source.
- Paper: Deep visual-semantic alignments for generating image descriptions, Andrej Karpathy et al. (2015). This paper establishes foundational multimodal neural embeddings between visual regions and descriptive text that the source adapts and extends to movie shots and temporal storylines.
- Paper: Long-term Recurrent Convolutional Networks for Visual Recognition and Description, Jeff Donahue et al. (2015). This work establishes the recurrent convolutional architecture framework for sequence modeling across video frames and natural language descriptions that informs the source's temporal visual-text alignment.
- Paper: Show and tell: A neural image caption generator, Oriol Vinyals et al. (2015). This foundational work demonstrates end-to-end neural sequence generation for visual grounding, providing the primary baseline paradigm that the source expands to rich narrative descriptions.
- Paper: Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models, Bryan A. Plummer et al. (2015). This work pioneers fine-grained phrase-to-region visual grounding benchmarks, serving as a conceptual predecessor to aligning detailed textual sentences with visual movie shots.
- Paper: From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions, Peter Young et al. (2014). This paper introduces denotational graph representations that ground linguistic expressions in visual contexts, framing the core semantic reasoning concepts utilized when linking descriptive text to video.
- Paper: Every Picture Tells a Story: Generating Sentences from Images, Ali Farhadi et al. (2010). This paper lays early groundwork for mapping images and natural language sentences into a shared semantic space using structured meaning representations.
- Paper: BabyTalk: Understanding and Generating Simple Image Descriptions, Girish Kulkarni et al. (2013). This foundational paper uses conditional random fields to infer visual-textual descriptions from detected visual elements, an alignment technique extended temporally in the source.
- Paper: MSR-VTT: A Large Video Description Dataset for Bridging Video and Language, Jun Xu et al. (2016). This work scales up open-domain video-and-language benchmarks by providing an extensive dataset of video clips paired with rich natural language descriptions.
- Paper: VideoBERT: A Joint Model for Video and Language Representation Learning, Chen Sun et al. (2019). This paper extends multimodal video-text alignment into the self-supervised transformer era by learning joint representations from uncurated video sequences and paired speech.
- Paper: Multimodal Transformer for Unaligned Multimodal Language Sequences, Yao-Hung Hubert Tsai et al. (2019). This work advances cross-modal alignment methodology by employing directional crossmodal attention to handle unaligned multimodal sequences across time.
- Paper: UNITER: UNiversal Image-TExt Representation Learning, Yen-Chun Chen et al. (2020). This paper develops universal vision-language representations via large-scale transformer pre-training across region-text matching and optimal transport alignments.
- Paper: VisualBERT: A Simple and Performant Baseline for Vision and Language, Liunian Harold Li et al. (2019). This work presents a unified transformer architecture that streamlines cross-modal grounding between visual components and linguistic descriptions.
- Paper: Phenaki: Variable Length Video Generation From Open Domain Textual Description, Ruben Villegas et al. (2023). This work extends narrative-visual modeling into generation by synthesizing variable-length, temporally coherent videos conditioned on evolving textual story scripts.
- Paper: Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis, Chaoyou Fu et al. (2025). This paper presents a comprehensive benchmark evaluating modern multimodal large language models on long-form video analysis and temporal multimodal comprehension.
