Hierarchical Neural Story Generation
Angela FanMike LewisYann Dauphin
Presents a hierarchical story generation framework and a 300,000-example prompt-story dataset that substantially improves narrative coherence by first generating a high-level premise before expanding it into full text using gated multi-scale self-attention.
Automated story generation represents a major frontier in artificial intelligence, requiring systems to maintain thematic consistency across long passages and plan high-level narrative plots. Traditional sequential text generation models struggle with these demands, often drifting off-topic or degenerating into generic phrasing because they generate text strictly word-by-word without broader context. Addressing this challenge is vital for advancing creative AI tools, automated content creation, and long-form narrative modeling.
The article demonstrates and evaluates a hierarchical neural architecture designed to produce fluent, topically relevant stories by decomposing generation into high-level premise creation followed by full-passage drafting. To support this, the authors introduce a new large-scale dataset, a gated multi-scale self-attention mechanism to capture long-range context efficiently, and a model fusion technique that enforces strict relevance between the generated story and its guiding premise.
The authors constructed a dataset of 303,358 human-written stories paired with writing prompts scraped from Reddit's WritingPrompts community, representing over 200 million words. Using this resource, they implemented a two-step framework: first generating a prompt using a convolutional language model, and then generating the corresponding story via a convolutional sequence-to-sequence model. The system incorporates deep multi-scale self-attention to capture past context across various time scales and leverages model fusion—training a second sequence model over a fixed, pretrained model—to ensure the generated narrative adheres to the prompt rather than defaulting to generic text. Evaluation relied on perplexity metrics, prompt-ranking accuracy across 1,000 test cases, and blind human evaluation studies on Amazon Mechanical Turk comparing story preference and prompt alignment.
The evaluation yielded several key findings. First, human evaluators preferred stories produced by the hierarchical model over a non-hierarchical baseline by more than two to one (67.3% preference). Second, combining the gated multi-scale self-attention mechanism with model fusion reduced test perplexity significantly, dropping it by roughly 9 points compared to standard convolutional baselines (from 45.54 down to 36.56). Third, human judges matched stories to their source prompts with substantially higher accuracy when using the fusion model, improving pairing performance by 7% over standard ensembling. Finally, while retrieval baselines quickly lose topical relevance as more stories are created, the generative fusion model sustains consistent relevance across an unlimited number of generated samples while maintaining lower word-overlap copying rates.
These findings indicate that introducing an explicit planning hierarchy alongside residual fusion architectures substantially mitigates the tendency of sequence models to lose coherence over multi-paragraph texts. By enabling the primary network to focus on narrative premises while the fused secondary layer generates rare, prompt-specific words, this strategy reduces the risk of repetitive, generic outputs. Consequently, the approach significantly improves text generation performance without requiring costly retrieval search over massive static databases.
Organizations developing creative AI tools or long-form document generators should adopt hierarchical planning architectures and multi-scale attention rather than single-pass language models. When implementing these systems, teams should utilize restricted top-k random sampling rather than beam search to avoid repetitive phrasing. Further work should focus on refining premise-generation models to produce more diverse, imaginative concepts, as well as addressing formatting artifacts and local repetition.
Confidence in these findings is high given the large dataset and strong alignment between automated metrics and human evaluations. However, practitioners should note certain limitations: prompt generation remains prone to producing common, generic themes, and sampling occasionally introduces minor grammatical artifacts, dialogue formatting errors, or local word repetition.
No sufficiently relevant recommendations were found.
- Paper: Re3: Generating Longer Stories With Recursive Reprompting and Revision, Kevin Yang et al. (2022). Re3 carries the source’s premise-to-story planning idea into longer narratives, adding recursive drafting, revision, and consistency checks to sustain plots across thousands of words.
