Intelligent Grimm - Open-ended Visual Storytelling via Latent Diffusion Models
Chang LiuHaoning WuYujie ZhongXiaoyun ZhangYanfeng WangWeidi Xie
Proposes StoryGen, an autoregressive latent diffusion framework conditioned on multimodal history, alongside the large-scale StorySalon dataset to synthesize visually coherent story sequences featuring unseen characters without requiring test-time optimization.
Recent advancements in text-to-image generative models have enabled the synthesis of high-quality standalone images, but creating coherent sequences of images remains a significant hurdle. Existing systems typically generate each visual frame in isolation without narrative context, rely solely on text prompts that introduce ambiguity, or depend on small datasets limited to a few specific characters. In real-world educational and creative applications, such as children's illustrated storytelling, these limitations cause visible inconsistencies in character appearances, visual style, and narrative flow. Addressing these issues requires models capable of open-ended visual generation across arbitrary storylines and new characters without needing slow, expensive per-character fine-tuning.
The article develops and evaluates StoryGen, a learning-based autoregressive framework for open-ended visual storytelling, alongside StorySalon, a large-scale multimodal dataset designed to support open-vocabulary sequential generation. The primary objective is to demonstrate that conditioning image generation on both the current text prompt and preceding image-text context maintains visual and character consistency for unseen characters without test-time optimization.
To accomplish this, the authors built StoryGen on a pre-trained Stable Diffusion model by introducing a vision-language context module. This module incorporates parallel cross-attention layers that extract and fuse diffusion denoising features from preceding frames under caption guidance, using calibrated noise levels as temporal position encodings. The framework is trained in two stages: single-frame self-attention pre-training followed by multiframe fine-tuning. To overcome the lack of suitable training data, the authors established a data collection and processing pipeline to create StorySalon. Sourced from YouTube videos and open-source e-books, the dataset comprises nearly 160,000 animation-style frames spanning 446 character categories, with an average story length of 14 frames. Performance was assessed through quantitative metrics—including Fréchet Inception Distance and similarity indicators—alongside human evaluations assessing style, content coherence, character consistency, and user preference.
The findings show that StoryGen substantially outperforms established baseline methods across both objective metrics and subjective human reviews. Quantitatively on the StorySalon test set, StoryGen achieved a Fréchet Inception Distance score of 33.90—improving upon baseline Stable Diffusion models (73.50) and prior sequential models such as StoryDALL-E (38.34) and AR-LDM (39.55)—while reaching the highest image-to-image consistency score (0.7467). In human evaluations, StoryGen earned a 67.14% win rate over standard baselines in open-ended story generation and a 96.87% win rate in story continuation tasks, consistently achieving the highest scores for style fidelity, narrative continuity, and character preservation. Ablation analyses further confirmed that utilizing diffusion-level denoising features provides significantly stronger visual consistency than relying on standard representation encoders or autoencoders.
These results demonstrate that auto-regressive context conditioning effectively eliminates the need for per-character fine-tuning methods like low-rank adaptation, drastically lowering the computational cost and latency of sequential image generation. This capability makes real-time, automated story illustration viable for educational software, digital publishing, and creative entertainment. Decision-makers looking to deploy sequential visual AI can adopt this architecture to support user-driven, interactive narratives at scale. However, practitioners should note that standard CLIP evaluation metrics exhibit a slight evaluation bias on cartoon data, and models must balance the tension between visual conditioning and text alignment. Continued work should focus on extending this contextual framework to longer narrative horizons, interactive editing workflows, and broader domains.
- Paper: Phenaki: Variable Length Video Generation From Open Domain Textual Description, Ruben Villegas et al. (2023). Introduces autoregressive generation conditioned on evolving prompt sequences over time, providing direct foundational context for multi-frame sequential visual storytelling.
- Paper: Align Your Latents: High-Resolution Video Synthesis with Latent Diffusion Models, Andreas Blattmann et al. (2023). Establishes how to adapt latent diffusion models to sequence generation by adding temporal context layers, which underpins the architectural foundations of StoryGen.
- Paper: Video Diffusion Models, Jonathan Ho et al. (2022). Pioneers video diffusion models with conditioning on prior segments, offering essential background on autoregressive sequence conditioning.
- Paper: Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading Books, Yukun Zhu et al. (2015). Pioneers the alignment of long-form story narratives with sequential visual scenes, laying the conceptual groundwork for multi-modal visual storytelling datasets.
- Paper: Hierarchical Neural Story Generation, Angela Fan et al. (2018). Presents hierarchical story generation methodologies that provide the conceptual textual basis for structured storyline-to-image sequence pipelines.
- Paper: Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning, Rohit Girdhar et al. (2024). Advances explicit image-conditioned sequential generation to synthesize continuous, high-fidelity video sequences from initial static visual frames.
- Paper: Show-o: One Single Transformer to Unify Multimodal Understanding and Generation, Jinheng Xie et al. (2025). Unifies multimodal understanding and autoregressive-diffusion image generation in a single transformer architecture, extending unified vision-language generative frameworks.
- Paper: CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer, Zhuoyi Yang et al. (2025). Scales temporal diffusion transformers to generate longer, dynamically evolving visual narratives from detailed text descriptions.
- Paper: VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation, Max Ku et al. (2024). Provides multimodal large-language-model-based explainable evaluation metrics for measuring semantic consistency and visual quality in conditional image synthesis.
