Intelligent Grimm - Open-ended Visual Storytelling via Latent Diffusion Models

Chang LiuHaoning WuYujie ZhongXiaoyun ZhangYanfeng WangWeidi Xie

article2024CVPR89 citations

Proposes StoryGen, an autoregressive latent diffusion framework conditioned on multimodal history, alongside the large-scale StorySalon dataset to synthesize visually coherent story sequences featuring unseen characters without requiring test-time optimization.

Listen

Recent advancements in text-to-image generative models have enabled the synthesis of high-quality standalone images, but creating coherent sequences of images remains a significant hurdle. Existing systems typically generate each visual frame in isolation without narrative context, rely solely on text prompts that introduce ambiguity, or depend on small datasets limited to a few specific characters. In real-world educational and creative applications, such as children's illustrated storytelling, these limitations cause visible inconsistencies in character appearances, visual style, and narrative flow. Addressing these issues requires models capable of open-ended visual generation across arbitrary storylines and new characters without needing slow, expensive per-character fine-tuning.

The article develops and evaluates StoryGen, a learning-based autoregressive framework for open-ended visual storytelling, alongside StorySalon, a large-scale multimodal dataset designed to support open-vocabulary sequential generation. The primary objective is to demonstrate that conditioning image generation on both the current text prompt and preceding image-text context maintains visual and character consistency for unseen characters without test-time optimization.

To accomplish this, the authors built StoryGen on a pre-trained Stable Diffusion model by introducing a vision-language context module. This module incorporates parallel cross-attention layers that extract and fuse diffusion denoising features from preceding frames under caption guidance, using calibrated noise levels as temporal position encodings. The framework is trained in two stages: single-frame self-attention pre-training followed by multiframe fine-tuning. To overcome the lack of suitable training data, the authors established a data collection and processing pipeline to create StorySalon. Sourced from YouTube videos and open-source e-books, the dataset comprises nearly 160,000 animation-style frames spanning 446 character categories, with an average story length of 14 frames. Performance was assessed through quantitative metrics—including Fréchet Inception Distance and similarity indicators—alongside human evaluations assessing style, content coherence, character consistency, and user preference.

The findings show that StoryGen substantially outperforms established baseline methods across both objective metrics and subjective human reviews. Quantitatively on the StorySalon test set, StoryGen achieved a Fréchet Inception Distance score of 33.90—improving upon baseline Stable Diffusion models (73.50) and prior sequential models such as StoryDALL-E (38.34) and AR-LDM (39.55)—while reaching the highest image-to-image consistency score (0.7467). In human evaluations, StoryGen earned a 67.14% win rate over standard baselines in open-ended story generation and a 96.87% win rate in story continuation tasks, consistently achieving the highest scores for style fidelity, narrative continuity, and character preservation. Ablation analyses further confirmed that utilizing diffusion-level denoising features provides significantly stronger visual consistency than relying on standard representation encoders or autoencoders.

These results demonstrate that auto-regressive context conditioning effectively eliminates the need for per-character fine-tuning methods like low-rank adaptation, drastically lowering the computational cost and latency of sequential image generation. This capability makes real-time, automated story illustration viable for educational software, digital publishing, and creative entertainment. Decision-makers looking to deploy sequential visual AI can adopt this architecture to support user-driven, interactive narratives at scale. However, practitioners should note that standard CLIP evaluation metrics exhibit a slight evaluation bias on cartoon data, and models must balance the tension between visual conditioning and text alignment. Continued work should focus on extending this contextual framework to longer narrative horizons, interactive editing workflows, and broader domains.

Cover for Intelligent Grimm - Open-ended Visual Storytelling via Latent Diffusion Models

Abstract

Generative models have recently exhibited exceptional capabilities in text-to-image generation, but still struggle to generate image sequences coherently. In this work, we focus on a novel, yet challenging task of generating a coherent image sequence based on a given storyline, denoted as open-ended visual storytelling. We make the following three contributions: (i) to fulfill the task of visual storytelling, we propose a learning-based auto-regressive image generation model, termed as StoryGen, with a novel vision-language context module, that enables to generate the current frame by conditioning on the corresponding text prompt and preceding image-caption pairs; (ii) to address the data shortage of visual storytelling, we collect paired image-text sequences by sourcing from online videos and open-source E-books, establishing processing pipeline for constructing a large-scale dataset with diverse characters, storylines, and artistic styles, named StorySalon; (iii) Quantitative experiments and human evaluations have validated the superiority of our StoryGen, where we show it can generalize to unseen characters without any optimization, and generate image sequences with coherent content and consistent character. Code, dataset, and models are available at https://haoningwu3639.github.io/StoryGen_Webpage/.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 3. Method
  • 3.1. Problem Formulation
  • 3.2. Architecture
  • 3.3. Model Training
  • 4. StorySalon Dataset
  • 5. Experiments
  • 5.1. Experimental Settings
  • 5.2. Quantitative Evaluation Results
  • 5.3. Human Evaluation Results
  • 5.4. Qualitative Results
  • 5.5. Ablation Studies
  • 6. Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Problem Formulation for Open-Ended Visual Storytelling

    model/method

    Open-ended visual storytelling is formulated as an auto-regressive text-to-image sequence generation task. Given an input storyline composed of LL sequential natural language descriptions {T1,T2,…,TL}\{\mathcal{T}_1, \mathcal{T}_2, \dots, \mathcal{T}_L\}, the objective is to synthesize a visually coherent sequence of images {I^1,I^2,…,I^L}\{\hat{\mathcal{I}}_1, \hat{\mathcal{I}}_2, \dots, \hat{\mathcal{I}}_L\}.

    The generation of the kk-th frame I^k\hat{\mathcal{I}}_k (for k>1k > 1) is conditioned on the current prompt Tk\mathcal{T}_k alongside all preceding image-caption pairs (I^<k,T<k)={(I^1,T1),…,(I^k−1,Tk−1)}(\hat{\mathcal{I}}_{<k}, \mathcal{T}_{<k}) = \{(\hat{\mathcal{I}}_1, \mathcal{T}_1), \dots, (\hat{\mathcal{I}}_{k-1}, \mathcal{T}_{k-1})\}:

    {I^1,I^2,…,I^L}=ΦStoryGen({T1,T2,…,TL};Θ)\{\hat{\mathcal{I}}_1, \hat{\mathcal{I}}_2, \dots, \hat{\mathcal{I}}_L\} = \Phi_{\text{StoryGen}}(\{\mathcal{T}_1, \mathcal{T}_2, \dots, \mathcal{T}_L\}; \Theta)

    I^k:=ΦStoryGen(I^k∣Tk,(I^<k,T<k))\hat{\mathcal{I}}_k := \Phi_{\text{StoryGen}}(\hat{\mathcal{I}}_k \mid \mathcal{T}_k, (\hat{\mathcal{I}}_{<k}, \mathcal{T}_{<k}))

    where ΦStoryGen\Phi_{\text{StoryGen}} denotes the generative storytelling model parameterized by Θ\Theta. In contrast to closed-set or optimization-based methods that require test-time fine-tuning (e.g., LoRA) for specific characters, this formulation requires direct feed-forward generalization to unseen characters and storylines.

  2. Knowl 2 — StoryGen Parallel Vision-Language Contextual Fusion Module

    model/method

    To combine narrative instructions with preceding visual-textual history, the transformer decoder blocks of the latent diffusion UNet are augmented with an image cross-attention layer placed in parallel with the standard text cross-attention layer.

    Let x∈RN×d\mathbf{x} \in \mathbb{R}^{N \times d} denote the intermediate noisy latent representation at a specific UNet block level, CT=ϕCLIP(Tk)\mathcal{C}^{\text{T}} = \phi_{\text{CLIP}}(\mathcal{T}_k) denote the CLIP text embedding of current prompt Tk\mathcal{T}_k, and CV\mathcal{C}^{\text{V}} denote the visual context features extracted from preceding frames. Linear projection matrices {WIQ,WIK,WIV}\{\mathbf{W}_I^Q, \mathbf{W}_I^K, \mathbf{W}_I^V\} and {WTQ,WTK,WTV}\{\mathbf{W}_T^Q, \mathbf{W}_T^K, \mathbf{W}_T^V\} project the representations into query, key, and value matrices:

    QI=xWIQ,KI=CVWIK,VI=CVWIV\mathbf{Q}_I = \mathbf{x}\mathbf{W}_I^Q, \quad \mathbf{K}_I = \mathcal{C}^{\text{V}}\mathbf{W}_I^K, \quad \mathbf{V}_I = \mathcal{C}^{\text{V}}\mathbf{W}_I^V

    QT=xWTQ,KT=CTWTK,VT=CTWTV\mathbf{Q}_T = \mathbf{x}\mathbf{W}_T^Q, \quad \mathbf{K}_T = \mathcal{C}^{\text{T}}\mathbf{W}_T^K, \quad \mathbf{V}_T = \mathcal{C}^{\text{T}}\mathbf{W}_T^V

    The fused block output O\mathbf{O} is computed as the direct sum of the visual cross-attention and text cross-attention operations:

    O=Softmax(QI(KI)⊤d)VI+Softmax(QT(KT)⊤d)VT\mathbf{O} = \text{Softmax}\left(\frac{\mathbf{Q}_I(\mathbf{K}_I)^\top}{\sqrt{d}}\right)\mathbf{V}_I + \text{Softmax}\left(\frac{\mathbf{Q}_T(\mathbf{K}_T)^\top}{\sqrt{d}}\right)\mathbf{V}_T

    where dd is the feature projection dimension.

  3. Knowl 3 — Multi-Frame Context Feature Extraction and Temporal Noise Scaling

    model/method

    In StoryGen, conditioning visual features CV\mathcal{C}^{\text{V}} are extracted directly from the intermediate representations of the pre-trained Stable Diffusion UNet rather than external feature extractors (e.g., CLIP or BLIP).

    For a target frame I^k\hat{\mathcal{I}}_k at diffusion timestep tt, each preceding frame I^j\hat{\mathcal{I}}_j (j<kj < k) guided by its caption Tj\mathcal{T}_j has Gaussian noise added at a reduced timestep t′=t/10t' = t / 10 and undergoes a single diffusion denoising step via the UNet ϕSDM\phi_{\text{SDM}}. The visual features after every self-attention layer across UNet blocks are extracted and concatenated across frames:

    CV=[ϕSDM(I^1,ϕCLIP(T1)),…,ϕSDM(I^k−1,ϕCLIP(Tk−1))]\mathcal{C}^{\text{V}} = [\phi_{\text{SDM}}(\hat{\mathcal{I}}_{1}, \phi_{\text{CLIP}}(\mathcal{T}_{1})), \dots, \phi_{\text{SDM}}(\hat{\mathcal{I}}_{k-1}, \phi_{\text{CLIP}}(\mathcal{T}_{k-1}))]

    To encode temporal distance, reference frames with larger temporal distance to frame kk receive progressively larger noise levels t′t'. This noise level variation functions as an implicit temporal positional encoding, down-weighting the influence of temporally distant frames on the current generation.

  4. Knowl 4 — Dual-Condition Classifier-Free Guidance Formulation

    equation

    StoryGen extends classifier-free guidance to two conditioning modalities: visual context CV\mathcal{C}^{\text{V}} and text condition CT\mathcal{C}^{\text{T}}. Using visual guidance scale wvw_v and text guidance scale wtw_t, the modified noise prediction ϵˉθ\bar{\boldsymbol{\epsilon}}_\theta at timestep tt is computed from the UNet noise estimator ϵθ\boldsymbol{\epsilon}_\theta as:

    ϵˉθ(xt,t,CV,CT)=ϵθ(xt,t,∅,∅)+wv(ϵθ(xt,t,CV,∅)−ϵθ(xt,t,∅,∅))+wt(ϵθ(xt,t,CV,CT)−ϵθ(xt,t,CV,∅))\bar{\boldsymbol{\epsilon}}_\theta(\mathbf{x}_t, t, \mathcal{C}^{\text{V}}, \mathcal{C}^{\text{T}}) = \boldsymbol{\epsilon}_\theta(\mathbf{x}_t, t, \varnothing, \varnothing) + w_v\left(\boldsymbol{\epsilon}_\theta(\mathbf{x}_t, t, \mathcal{C}^{\text{V}}, \varnothing) - \boldsymbol{\epsilon}_\theta(\mathbf{x}_t, t, \varnothing, \varnothing)\right) + w_t\left(\boldsymbol{\epsilon}_\theta(\mathbf{x}_t, t, \mathcal{C}^{\text{V}}, \mathcal{C}^{\text{T}}) - \boldsymbol{\epsilon}_\theta(\mathbf{x}_t, t, \mathcal{C}^{\text{V}}, \varnothing)\right)

    where xt\mathbf{x}_t is the noisy latent at diffusion step tt, ∅\varnothing denotes the null/unconditioned input token, and the default inference hyperparameters are set to wv=7.0w_v = 7.0 and wt=3.5w_t = 3.5.

  5. Knowl 5 — Two-Stage Training Strategy for StoryGen

    model/method

    StoryGen is trained on Stable Diffusion v1.5 using a two-stage procedure with learning rate 1×10−51 \times 10^{-5} and batch size 256 across 8 NVIDIA RTX 3090 GPUs:

    1. Single-Frame Pre-training: The standard Stable Diffusion model is trained without the image cross-attention layers. Only the self-attention layers within the UNet are optimized for 3,000 iterations to adapt single-frame generation to the storybook domain while preserving base generative capability.
    2. Multi-Frame Fine-Tuning: The newly added image cross-attention layers in the vision-language context module are trained while all pre-existing SDM parameters remain frozen. Training runs for 5,000 iterations conditioning on a single preceding image-caption pair, followed by 5,000 iterations conditioning on multiple preceding pairs.

    To enable unconditional and single-conditional evaluations during classifier-free guidance, conditioning inputs are randomly dropped during training: the current text CT\mathcal{C}^{\text{T}} is dropped with probability 5%, and the context pair CV\mathcal{C}^{\text{V}} is dropped with probability 15%.

  6. Knowl 6 — StorySalon Data Processing Pipeline

    model/method

    To build the StorySalon dataset from uncurated online storytelling videos (e.g., YouTube) and open-source CC BY 4.0 E-books, a three-step processing pipeline is applied:

    1. Visual Frame Extraction: Keyframes and subtitle timestamps are extracted from videos. DINO ViT features are computed to eliminate near-duplicate frames based on cosine similarity. YOLOv7 is applied to detect, segment, and remove human storyteller appearances and headshots. E-book text is acquired via Whisper from accompanying audio, or via OCR when audio is unavailable.
    2. Vision-Language Alignment: Subtitles and narration are temporally aligned to visual keyframes using Dynamic Time Warping (DTW). Because high-level story narrations have a semantic gap with literal visual descriptions, TextBind is used to generate dense descriptive captions for each image conditioned on both the image and the narrative text.
    3. Visual Frame Post-Processing: Storybook page borders, extraneous metadata, and embedded text regions are detected using an OCR detector and filled in using a Stable Diffusion inpainting model to prevent textual artifacts from corrupting generative training.
  7. Knowl 7 — Comparison of Story Generation Datasets

    data/table

    The StorySalon dataset provides animation-style story sequences with significantly greater length, character diversity, and total image count compared to existing story visualization datasets.

    Dataset Style #Frames Avg. Length #Categories
    PororoSV Animation 73,665 5 9
    FlintstonesSV Animation 122,560 5 7
    DiDeMoSV Real 52,905 3 -
    VIST Real 145,950 5 -
    StorySalon Animation 159,778 14 446

    Character categories were quantified by prompting MiniGPT-4 to label main character types (e.g., Dog, Cat) across all images, filtering categories appearing fewer than 3 times.

  8. Knowl 8 — Quantitative Evaluation on StorySalon Test Set

    data/table

    Automated evaluation on the StorySalon test set (5% split, ~7,000 pairs) assesses image quality via Fréchet Inception Distance (FID), image-image character/style consistency via CLIP image similarity (CLIP-I), and text-image alignment via CLIP text-image similarity (CLIP-T). Best images are selected from 10 candidates using PickScore.

    Model FID ↓\downarrow CLIP-I ↑\uparrow CLIP-T ↑\uparrow
    Ground Truth (GT) - 1.0000 0.2668
    SDM 73.50 0.6155 0.3218
    Prompt-SDM 67.35 0.6272 0.3225
    Finetuned-SDM 42.01 0.6970 0.3005
    StoryDALL E 38.34 0.6823 0.2366
    AR-LDM 39.55 0.6864 0.2614
    StoryGen 33.90 0.7467 0.2875

    Prompt-SDM uses the extra prefix prompt 'A cartoon style image', and Finetuned-SDM fine-tunes all SDM parameters on StorySalon. StoryGen achieves the best FID (33.90) and highest frame-to-frame image consistency (CLIP-I of 0.7467).

  9. Knowl 9 — Human Evaluation for Story Generation and Continuation

    data/table

    Human evaluation scores (1 to 5 scale) and user preference rates (%) across two tasks: Open-Ended Story Generation (visualizing text-only storylines generated via GPT-4) and Open-Ended Story Continuation (continuing a story from a given first frame or internet character reference).

    Model Text-Image Align. ↑\uparrow Style Consist. ↑\uparrow Content Consist. ↑\uparrow Character Consist. ↑\uparrow Quality ↑\uparrow Preference ↑\uparrow
    Story Generation
    Ground Truth 4.04 4.66 4.41 4.54 4.29 -
    SDM 3.61 2.88 2.90 2.51 3.74 14.05%
    Prompt-SDM 3.39 2.56 2.68 2.10 3.44 8.57%
    StoryGen-S 3.50 2.73 2.81 2.21 3.19 10.24%
    StoryGen 3.78 4.79 4.26 4.64 3.76 67.14%
    Story Continuation
    StoryDALL E 1.18 1.55 1.20 1.14 1.19 0.63%
    AR-LDM 2.47 2.82 2.40 1.87 2.54 2.50%
    StoryGen 4.23 4.70 4.35 4.38 4.18 96.87%

    StoryGen-S denotes StoryGen evaluated without preceding context conditions. StoryGen achieves majority user preference in both generation (67.14%) and continuation (96.87%).

  10. Knowl 10 — Ablation Study on Context Conditioning Representations

    data/table

    Ablation experiments evaluate different context feature conditioning mechanisms on the StorySalon test set using FID, CLIP-I, and CLIP-T (with PickScore candidate filtering):

    Model Variant FID ↓\downarrow CLIP-I ↑\uparrow CLIP-T ↑\uparrow
    StoryGen-Single 38.81 0.6869 0.3140
    StoryGen-VAE 36.98 0.6846 0.3061
    StoryGen-CLIP 36.66 0.6934 0.3140
    StoryGen-BLIP 34.78 0.7026 0.2838
    StoryGen-LT 36.41 0.7141 0.3025
    StoryGen (Full) 33.90 0.7467 0.2875
    • StoryGen-Single: Model without the context module (only self-attention fine-tuned).
    • StoryGen-VAE: Context conditioned using SDM VAE latent representations without text guidance.
    • StoryGen-CLIP: Context conditioned via CLIP image embeddings.
    • StoryGen-BLIP: Context conditioned via BLIP image encoder features.
    • StoryGen-LT: Context extracted using large-scale diffusion timesteps (t′=tt' = t).
    • StoryGen (Full): Context extracted using one-step UNet diffusion features with reduced timestep t′=t/10t' = t / 10, achieving the lowest FID (33.90) and highest image consistency (0.7467).

Coverage note — None was omitted; all key architectural components, dataset pipelines, training protocols, and quantitative/human evaluation benchmarks are fully represented.

References

  1. 1.Omri Avrahami, Dani Lischinski, and Ohad Fried. Blended diffusion for text-driven editing of natural images. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2022.
  2. 2.Tim Brooks, Aleksander Holynski, and Alexei A. Efros. Instructpix2pix: Learning to follow image editing instructions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2023.
  3. 3.Hong Chen, Rujun Han, Te-Lin Wu, Hideki Nakayama, and Nanyun Peng. Character-centric story visualization via visual planning and token alignment. In Proceedings of the Conference on Empirical Methods in Natural Language Processinng, 2022.
  4. 4.Zheng Chen, Yulun Zhang*, Ding Liu, Bin Xia, Jinjin Gu, Linghe Kong*, and Xin Yuan. Hierarchical integration diffusion model for realistic image deblurring. In Advances in Neural Information Processing Systems, 2023.
  5. 5.K. Dickinson David, A. Griffith Julie, Golinkoff Roberta, Michnick, and Hirsh-Pasek Kathy. How reading books fosters language development around the world. Child Development Research, 2012.
  6. 6.Patrick Esser, Johnathan Chiu, Parmida Atighehchian, Jonathan Granskog, and Anastasis Germanidis. Structure and content-guided video synthesis with diffusion models. In Proceedings of the International Conference on Computer Vision, 2023.
  7. 7.Yuan Gong, Youxin Pang, Xiaodong Cun, Menghan Xia, Yingqing He, Haoxin Chen, Longyue Wang, Yong Zhang, Xintao Wang, Ying Shan, and Yujiu Yang. Talecrafter: Interactive story visualization with multiple characters. SIGGRAPH Asia, 2023.
  8. 8.Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial networks. Communications of the ACM, 2020.
  9. 9.Tanmay Gupta, Dustin Schwenk, Ali Farhadi, Derek Hoiem, and Aniruddha Kembhavi. Imagine this! scripts to compositions to videos. In Proceedings of the European Conference on Computer Vision, 2018.
  10. 10.Amir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or. Prompt-to-prompt image editing with cross attention control. In Proceedings of the International Conference on Learning Representations, 2023.
  11. 11.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In Advances in Neural Information Processing Systems, 2017.
  12. 12.Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. In NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications, 2021.
  13. 13.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems, 2020.
  14. 14.Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P Kingma, Ben Poole, Mohammad Norouzi, David J Fleet, et al. Imagen video: High definition video generation with diffusion models. arXiv preprint arXiv:2210.02303, 2022.
  15. 15.Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J Fleet. Video diffusion models. In Advances in Neural Information Processing Systems, 2022.
  16. 16.Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. In Proceedings of the International Conference on Learning Representations, 2022.
  17. 17.Ting-Hao Huang, Francis Ferraro, Nasrin Mostafazadeh, Ishan Misra, Aishwarya Agrawal, Jacob Devlin, Ross Girshick, Xiaodong He, Pushmeet Kohli, Dhruv Batra, et al. Visual storytelling. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2016.
  18. 18.Bahjat Kawar, Shiran Zada, Oran Lang, Omer Tov, Huiwen Chang, Tali Dekel, Inbar Mosseri, and Michal Irani. Imagic: Text-based real image editing with diffusion models. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2023.
  19. 19.Diederik P Kingma and Max Welling. Auto-encoding variational bayes. In Proceedings of the International Conference on Learning Representations, 2014.
  20. 20.Yuval Kirstain, Adam Polyak, Uriel Singer, Shahbuland Matiana, Joe Penna, and Omer Levy. Pick-a-pic: An open dataset of user preferences for text-to-image generation. In Advances in Neural Information Processing Systems, 2023.
  21. 21.Bowen Li. Word-level fine-grained story visualization. In Proceedings of the European Conference on Computer Vision, 2022.
  22. 22.Dongxu Li, Junnan Li, and Steven CH Hoi. Blip-diffusion: Pre-trained subject representation for controllable text-to-image generation and editing. In Advances in Neural Information Processing Systems, 2023.
  23. 23.Huayang Li, Siheng Li, Deng Cai, Longyue Wang, Lemao Liu, Taro Watanabe, Yujiu Yang, and Shuming Shi. Textbind: Multi-turn interleaved multimodal instruction-following. arXiv preprint arXiv:2309.08637, 2023.
  24. 24.Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In Proceedings of the International Conference on Machine Learning, 2022.
  25. 25.Yitong Li, Zhe Gan, Yelong Shen, Jingjing Liu, Yu Cheng, Yuexin Wu, Lawrence Carin, David Carlson, and Jianfeng Gao. Storygan: A sequential conditional gan for story visualization. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2019.
  26. 26.Ziyi Li, Qinye Zhou, Xiaoyun Zhang, Ya Zhang, Yanfeng Wang, and Weidi Xie. Guiding text-to-image diffusion model towards grounded generation. In Proceedings of the International Conference on Computer Vision, 2023.
  27. 27.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In Proceedings of the European Conference on Computer Vision, 2014.
  28. 28.Andreas Lugmayr, Martin Danelljan, Andres Romero, Fisher Yu, Radu Timofte, and Luc Van Gool. Repaint: Inpainting using denoising diffusion probabilistic models. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2022.
  29. 29.Adyasha Maharana and Mohit Bansal. Integrating visuospatial, linguistic, and commonsense structure into story visualization. In Proceedings of the Conference on Empirical Methods in Natural Language Processinng, 2021.
  30. 30.Adyasha Maharana, Darryl Hannan, and Mohit Bansal. Improving generation and evaluation of visual stories via semantic consistency. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2021.
  31. 31.Adyasha Maharana, Darryl Hannan, and Mohit Bansal. Storydall-e: Adapting pretrained text-to-image transformers for story continuation. In Proceedings of the European Conference on Computer Vision, 2022.
  32. 32.Caron Mathilde, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. In Proceedings of the International Conference on Computer Vision, 2021.
  33. 33.Chenlin Meng, Yutong He, Yang Song, Jiaming Song, Jiajun Wu, Jun-Yan Zhu, and Stefano Ermon. Sdedit: Guided image synthesis and editing with stochastic differential equations. In Proceedings of the International Conference on Learning Representations, 2021.
  34. 34.Meinard Müller. Dynamic time warping. Information retrieval for music and motion, 2007.
  35. 35.Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. In Proceedings of the International Conference on Machine Learning, 2022.
  36. 36.Xichen Pan, Pengda Qin, Yuhong Li, Hui Xue, and Wenhu Chen. Synthesizing coherent story with auto-regressive latent diffusion models. In Winter Conference on Applications of Computer Vision, 2024.
  37. 37.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning, 2021.
  38. 38.Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. Robust speech recognition via large-scale weak supervision. In Proceedings of the International Conference on Machine Learning, 2023.
  39. 39.Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In Proceedings of the International Conference on Machine Learning, 2021.
  40. 40.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 2022.
  41. 41.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2022.
  42. 42.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al. Photorealistic text-to-image diffusion models with deep language understanding. In Advances in Neural Information Processing Systems, 2022.
  43. 43.Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, et al. Make-a-video: Text-to-video generation without text-video data. In Proceedings of the International Conference on Learning Representations, 2023.
  44. 44.Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. In Proceedings of the International Conference on Learning Representations, 2020.
  45. 45.Gabrielle A. Strouse, Angela Nyhout, and Patricia A. Ganea. The role of book features in young children’s transfer of information from picture books to real-world contexts. Frontiers in Psychology, 2018.
  46. 46.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems, 2017.
  47. 47.Chien-Yao Wang, Alexey Bochkovskiy, and Hong-Yuan Mark Liao. Yolov7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2023.
  48. 48.Shaoan Xie, Zhifei Zhang, Zhe Lin, Tobias Hinz, and Kun Zhang. Smartbrush: Text and shape guided object inpainting with diffusion model. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2023.
  49. 49.Li Xin, Chu Wenqing, Wu Ye, Yuan Weihang, Liu Fanglong, Zhang Qi, Li Fu, Feng Haocheng, Ding Errui, and Wang Jingdong. Videogen: A reference-guided latent diffusion approach for high definition text-to-video generation. arXiv preprint arXiv:2309.00398, 2023.
  50. 50.Tao Xu, Pengchuan Zhang, Qiuyuan Huang, Han Zhang, Zhe Gan, Xiaolei Huang, and Xiaodong He. Attngan: Fine-grained text to image generation with attentional generative adversarial networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018.
  51. 51.Hu Ye, Jun Zhang, Sibo Liu, Xiao Han, and Wei Yang. Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models. arXiv preprint arxiv:2308.06721, 2023.
  52. 52.Shengming Yin, Chenfei Wu, Huan Yang, Jianfeng Wang, Xiaodong Wang, Minheng Ni, Zhengyuan Yang, Linjie Li, Shuguang Liu, Fan Yang, et al. Nuwa-xl: Diffusion over diffusion for extremely long video generation. In Association for Computational Linguistics, 2023.
  53. 53.Han Zhang, Tao Xu, Hongsheng Li, Shaoting Zhang, Xiaogang Wang, Xiaolei Huang, and Dimitris N Metaxas. Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks. In Proceedings of the International Conference on Computer Vision, 2017.
  54. 54.Han Zhang, Tao Xu, Hongsheng Li, Shaoting Zhang, Xiaogang Wang, Xiaolei Huang, and Dimitris N Metaxas. Stackgan++: Realistic image synthesis with stacked generative adversarial networks. IEEE transactions on pattern analysis and machine intelligence, 2018.
  55. 55.Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. In Proceedings of the International Conference on Computer Vision, 2023.
  56. 56.Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592, 2023.

Citation

MLA
Liu, C., et al. “Intelligent Grimm -- Open-ended Visual Storytelling via Latent Diffusion Models”. arXiv, 2023, http://arxiv.org/abs/2306.00973v3.
APA
Liu, C., Wu, H., Zhong, Y., Zhang, X., Wang, Y., & Xie, W. (2023). Intelligent Grimm -- Open-ended Visual Storytelling via Latent Diffusion Models. arXiv. http://arxiv.org/abs/2306.00973v3
Chicago
Liu, C., H. Wu, Y. Zhong, X. Zhang, Y. Wang, and W. Xie. 2023. “Intelligent Grimm -- Open-ended Visual Storytelling via Latent Diffusion Models”. arXiv. http://arxiv.org/abs/2306.00973v3.
Harvard
Liu, C. et al. (2023) “Intelligent Grimm -- Open-ended Visual Storytelling via Latent Diffusion Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2306.00973v3.
Vancouver
1. Liu C, Wu H, Zhong Y, Zhang X, Wang Y, Xie W (2023) Intelligent Grimm -- Open-ended Visual Storytelling via Latent Diffusion Models. arXiv

BibTeX

@article{liu2023intelligent,
  title = {Intelligent Grimm -- Open-ended Visual Storytelling via Latent Diffusion Models},
  author = {Liu, Chang and Wu, Haoning and Zhong, Yujie and Zhang, Xiaoyun and Wang, Yanfeng and Xie, Weidi},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2306.00973v3},
  eprint = {2306.00973}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE