InfoAlign: A Human-AI Co-Creation System for Storytelling with Infographics
Jielin FengXinwu YeQianhui LiVerena PrantlJun-Hsiang YaoYuheng ZhaoYun WangSiming Chen
Introduces InfoAlign, a human-AI co-creation system that transforms unstructured text into coherent storytelling infographics through a structured workflow of story construction, visual encoding, and spatial composition while preserving user intent.
Data-driven visual communication frequently relies on infographics to translate complex narratives into engaging and digestible formats. However, existing design tools struggle to maintain story consistency across the multi-step authoring process and often fail to preserve user intent. Highly automated tools typically produce rigid, disconnected outputs, while traditional software requires extensive manual effort to align narrative structure, visual assets, and spatial layout.
The article develops and evaluates a narrative-centric workflow and an interactive system, named InfoAlign, designed to turn long or unstructured text into coherent visual storytelling infographics while supporting human-AI co-creation.
To establish design requirements, the researchers conducted formative interviews with eight professional creators and analyzed a curated corpus of 70 real-world storytelling infographics to identify layout and narrative patterns. Based on these insights, the authors designed a three-phase workflow combining large language models for story structuring and asset suggestion with rule-based optimization algorithms for spatial composition. The resulting system guides users through story construction, visual encoding, and layout arrangement while maintaining an editable canvas for interactive refinement. The system was then evaluated through a task-based user study involving 12 participants creating full infographics from various textual datasets.
The evaluation produced several key findings regarding narrative quality and user interaction. First, the automated narrative extraction demonstrated high structural quality and reliability, with 95% of generated story pieces rated as coherent with the story goal and 95.73% of extracted details verified as strictly factually accurate against source documents. Second, participants favored intervening selectively: modification rates were highest for expressive elements like text highlights (41.0%) and icons (39.3%), whereas factual text (17.9%) and chart structures (4.0%) were largely retained as recommended. Third, the hybrid authoring model proved efficient, enabling users to complete end-to-end professional infographics in an average of 20.3 minutes. Finally, system effectiveness scored consistently high on 7-point scales across usability (6.27), creativity support (6.31), perceived co-creation (6.10), and visual aesthetics (6.17).
These results demonstrate that combining automated narrative scaffolding with flexible, step-by-step human intervention provides a practical balance between productivity and authorial control. Rather than relying on end-to-end generation that frequently misaligns with spatial and narrative constraints, employing rule-based spatial layout alongside AI-driven content extraction reduces production overhead while preserving user intent and communication goals.
Organizations producing data-driven communications should explore narrative-centric, hybrid-AI workflows to lower the technical barrier and time required for content generation. For system developers, prioritizing transparency by providing explanations for AI recommendations—such as styling or layout logic—and integrating fine-grained graphic editing capabilities will further enhance user trust and creative flexibility. Future initiatives should focus on extending this narrative approach to broader formats such as slide presentations, video summaries, and multimodal data sources.
While the findings demonstrate strong effectiveness, confidence in the results should be contextualized by the small study sample size of 12 participants and the use of qualitative self-reporting during guided tasks. Additionally, the system is optimized primarily for textual inputs, meaning complex quantitative tables currently require text-based pre-processing.
- Paper: Automating the design of graphical presentations of relational information, Jock D. Mackinlay (1986). This seminal work establishes the foundational principles of graphical languages and automated graphic design criteria that underpin narrative visual encoding and layout generation.
- Paper: Hierarchical Neural Story Generation, Angela Fan et al. (2018). This paper introduces hierarchical narrative construction from high-level premises to detailed text, providing the structural conceptual foundation for decomposing unstructured text into structured story units.
- Paper: Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs, Ling Yang et al. (2024). This work demonstrates how to use multimodal LLMs for planning, prompt decomposition, and regional spatial composition, directly inspiring multi-stage semantic layout workflows.
- Paper: Adding Conditional Control to Text-to-Image Diffusion Models, Lvmin Zhang et al. (2023). This foundational work provides key spatial conditioning mechanisms that enable generative visual models to adhere strictly to layout blueprints and semantic geometries.
- Paper: TextDiffuser: Diffusion Models as Text Painters, Jingye Chen et al. (2023). This research integrates layout coordinate planning with diffusion models to render coherent visual text and graphical compositions essential for infographic generation.
- Paper: Intelligent Grimm - Open-ended Visual Storytelling via Latent Diffusion Models, Chang Liu et al. (2024). This study addresses the challenges of sequential visual storytelling and character/narrative coherence across multi-frame visual outputs.
- Paper: PaperBanana: Automating Academic Illustration for AI Scientists, Dawei Zhu et al. (2026). This work extends narrative visual layout and co-creation principles to the specialized, agentic generation of publication-ready academic methodology diagrams and illustrations.
- Paper: Idea2Story: An Automated Pipeline for Transforming Research Concepts into Complete Scientific Narratives, Tengyue Xu et al. (2026). This system builds on structured narrative generation by transforming complex scientific concepts into structured multi-agent research narratives.
- Paper: Texterial: A Text-as-Material Interaction Paradigm for LLM-Mediated Writing, Jocelyn Shen et al. (2026). This research explores interactive, direct-manipulation spatial paradigms for authoring and refining text, offering complementary interaction techniques for narrative co-creation tools.
- Paper: DuoDrama: Supporting Screenplay Refinement Through LLM-Assisted Human Reflection, Yuying Tang et al. (2026). This work applies interactive human-AI reflection workflows to story structure and character dynamics for creative screenplay refinement.
