Re3: Generating Longer Stories With Recursive Reprompting and Revision

Kevin YangYuandong TianNanyun PengDan Klein

article2022EMNLP292 citations

Proposes a fully automatic framework that enables general-purpose language models to generate coherent, multi-thousand-word stories by combining structured plan generation, dynamic contextual prompt composition, reranking, and targeted factual revision.

Listen

Generating coherent long-form narratives of several thousand words remains a major challenge for artificial intelligence. Standard language models typically lose narrative focus, contradict earlier details, or wander away from the initial premise when generating long texts. The article addresses this challenge by evaluating whether decomposing the writing process into structured planning, drafting, and revising stages can enable general-purpose language models to produce high-quality, multi-thousand-word stories automatically.

The main objective of the article is to demonstrate and evaluate the Recursive Reprompting and Revision framework, an automated system designed to generate plot-coherent stories of over two thousand words from brief initial premises. The framework emulates the human writing process without requiring human intervention or task-specific training for the generation components.

The authors designed a four-part modular approach: a planning module that generates story settings, character descriptions, and numbered outlines; a drafting module that recursively composes prompts combining high-level plans with recent narrative summaries; a rewrite module that reranks alternative continuations based on learned coherence and relevance models; and an edit module that tracks character attributes in a structured dictionary to detect and correct factual contradictions. To test the approach, the authors generated 2,000 to 2,500-word stories across 100 diverse premises and conducted pairwise human evaluations comparing the framework against two standard rolling-window language model baselines.

The evaluation revealed several key findings. First, human evaluators judged the framework's stories to have a coherent overarching plot significantly more often than the baselines, achieving up to a 14% absolute increase in perceived coherence. Second, faithfulness to the starting premise increased by up to 20%. Third, evaluators judged between 80.0% and 83.3% of the generated stories to be human-written. Finally, component ablation analyses demonstrated that the planning and rewrite modules were critical to maintaining coherence and relevance, whereas the factual edit module provided negligible measurable improvement to the overall story quality.

These findings imply that structured prompt management and discriminative reranking can effectively extend the capabilities of out-of-the-box language models to complex, long-horizon text generation. The framework demonstrates an ability to self-correct and return to high-level outlines even after minor narrative deviations. However, resolving fine-grained factual continuity remains a bottleneck, as current error detection and rewriting subroutines suffer from compounding inaccuracies.

To advance automated long-form generation, the article recommends developing hierarchical outline schemes for even longer texts, such as novellas, along with adaptive mechanisms to control narrative pacing. Crucially, researchers must prioritize creating automated evaluation metrics for long-range plot coherence and factual consistency to overcome the prohibitive cost and noise of relying exclusively on human annotators.

Readers should interpret these conclusions in light of several limitations. Evaluator agreement on subjective story quality metrics was low, and the study relied on a constrained sample size due to evaluation costs. Additionally, the inconsistency detection mechanism was limited to character attributes and exhibited modest accuracy, achieving a classification area under the curve score of only 0.684 in controlled testing.

arXiv: 2210.06774
Cover for Re3: Generating Longer Stories With Recursive Reprompting and Revision

Abstract

We consider the problem of automatically generating longer stories of over two thousand words. Compared to prior work on shorter stories, long-range plot coherence and relevance are more central challenges here. We propose the Recursive Reprompting and Revision framework (Re³) to address these challenges by (a) prompting a general-purpose language model to construct a structured overarching plan, and (b) generating story passages by repeatedly injecting contextual information from both the plan and current story state into a language model prompt. We then revise by (c) reranking different continuations for plot coherence and premise relevance, and finally (d) editing the best continuation for factual consistency. Compared to similar-length stories generated directly from the same base model, human evaluators judged substantially more of Re³'s stories as having a coherent overarching plot (by 14% absolute increase), and relevant to the given initial premise (by 20%).

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Recursive Reprompting and Revision
  • 3.1 Plan Module
  • 3.2 Draft Module
  • 3.3 Rewrite Module
  • 3.4 Edit Module
  • 4 Evaluation
  • 5 Analysis
  • 5.1 Ablation Study
  • 5.2 Further Analysis of Edit Module
  • 6 Discussion
  • Limitations
  • Acknowledgements
  • Ethics Statement
  • References
  • A Character Name Generation
  • B Details on Additional Reranking Heuristics
  • C Details on Editing System Information Extraction
  • D Data on API Usage
  • E Dataset Usage
  • F Length vs. Story Quality Analysis
  • G Full Metrics for Miscellaneous Writing Problems
  • H Mechanical Turk Evaluation Details
  • I Annotator Agreement
  • J Example Stories
  • J.1 Examples for Premise 1
  • J.2 Examples for Premise 2
  • J.3 Examples for Premise 3
  • J.4 Examples for Premise 4
  • K Example Story Generation Steps
  • Initial Setup and Outline For Story Generation Step Example
  • Initial Inferred Attributes
  • Prompt for First Story Generation Step
  • Best Continuation for First Story Generation Step
  • Prompt for Later Story Generation Step
  • Best Continuation for Later Story Generation Step
  • Prompt for Second Later Story Generation Step
  • L Example Data for Editing System Evaluation
  • Initial Setups
  • Altered Setups
  • Initial Story
  • Altered Story
  • M Longer Story Example
  • Longer Story Premise
  • Initial Setup and Outline for RE 3, Longer Story Example
  • RE 3 Longer Story Example

Knowls

  1. Knowl 1 — Recursive Reprompting and Revision framework

    model/method

    Re3 is a fully automatic framework for generating multi-thousand-word stories from an initial premise. It decomposes long-form writing into four stages: Plan constructs a structured story plan; Draft generates passages sequentially by repeatedly rebuilding a prompt from the plan and the story so far; Rewrite reranks alternative continuations for coherence and premise relevance; and Edit makes local corrections for long-range factual consistency.

    The text-generation components are zero-shot: they use general-purpose GPT-3 models without training on story datasets. The only component trained on prior story data is the Rewrite module’s pair of rerankers. In the main implementation, Re3 generates approximately 2,000–2,500-word stories, while a hierarchical extension demonstrates generation of a roughly 7,500-word story without increasing the generator’s context window beyond 1,024 tokens.

  2. Knowl 2 — Structured plan generation from a premise

    model/method

    The Plan module augments an initial story premise with three components: a setting, character descriptions, and a numbered plot outline. GPT3-Instruct-175B generates a one-sentence setting using a prompt beginning with “The story is set in.” The same model generates up to three character names and descriptions conditioned on the premise and setting. For character names, the system samples candidates and rejects malformed outputs using heuristics for punctuation, repeated names, story-role words, and attribute words; it prefers two-word names to obtain full names.

    GPT3-Instruct-175B then produces a numbered outline of the story. The system parses the output into outline points and resamples until the outline is well formed. These generated plan elements are later reused as structured context during passage generation, rather than being used only once at the beginning.

  3. Knowl 3 — Recursive context composition for passage drafting

    model/method

    For every outline point, the Draft module generates several fixed-length story continuations before advancing to the next point. Each generation prompt combines information at multiple levels of detail:

    • Relevant context: premise, setting, and character descriptions selected for relevance to the recent passage. New named entities are detected with Flair-based named-entity recognition, and GPT3-Instruct-175B writes descriptions for them. A Dense Passage Retrieval model selects context that is relevant to the recent story.
    • Previous sections’ outlines: high-level summaries of completed outline sections.
    • Recent story summary: a summary of several penultimate passages generated by GPT3-Instruct-13B.
    • Current outline point: the plot point that the next passage should address.
    • Autoregressive context: the immediately preceding passage copied verbatim.

    The resulting prompt is sent to GPT3-175B, which generates the next passage. This coarse-to-fine prompt composition preserves detailed information near the current generation point while retaining high-level information about earlier events. In the main experiments, each passage contains 256 tokens.

  4. Knowl 4 — Coherence-and-relevance reranking for rewriting

    model/method

    The Rewrite module approximates a full passage rewrite by generating alternative Draft continuations and selecting the highest-scoring one rather than editing the first continuation directly. It scores each candidate for both coherence with the preceding story and relevance to the current outline point.

    The coherence reranker is a Longformer-Base classifier fine-tuned on WritingPrompts passages of at most 1,000 tokens. The final 200 tokens are treated as the true continuation, while negative examples replace them with a continuation sampled from the same or another story. The relevance reranker uses the same architecture and is trained on 2,000 examples consisting of a 200-token passage paired with a GPT3-Instruct-13B summary; mismatched summaries provide negative examples.

    Rule-based filters remove empty outputs, repeated five-word sequences, near-duplicate sentences, prompt repetitions, analysis-like headers, undesirable strings such as “Comment” or “copyright,” and passages that switch away from third-person narration. These filters are applied before selecting the continuation with the best combined coherence and relevance score.

  5. Knowl 5 — Structured attribute editing for factual continuity

    model/method

    The Edit module targets long-range contradictions in character attributes such as gender, age, occupation, and relationships. For each character, it maintains an attribute dictionary containing attribute–value pairs extracted from the generated story. When a new passage is produced, the system does not compare every statement with the entire preceding story; it compares newly extracted attributes only with the corresponding dictionary entries.

    GPT3-Instruct-175B first lists natural-language facts about each character in the new passage. GPT3-Instruct-13B extracts attribute keys from those facts and then regenerates the corresponding values. The system uses three outputs and retains repeated or mutually entailed results to reduce hallucinated facts. A UnifiedQA model filters unreliable attribute keys, and a BART-Large MNLI entailment model checks whether new values contradict old values. New characters receive new dictionaries, and reciprocal relationships between characters can be added automatically.

    When a contradiction is detected, the original natural-language fact becomes an editing instruction. The selected continuation and that instruction are passed to the GPT-3 Edit API, which makes a local correction. The procedure is designed for isolated character-attribute inconsistencies rather than broad changes to plot, setting, or temporal state.

  6. Knowl 6 — Long-form story-generation evaluation design

    experimental setup

    The evaluation task is to generate an English story from a brief premise. GPT3-Instruct-175B is prompted at high temperature to produce 100 diverse premises. For the fixed-length implementation, outlines are resampled until they contain exactly three points, and Re3 generates four 256-token continuations per point, yielding 3,072 generated tokens before story ending. GPT3-175B’s Insert API completes the ending with the suffix “The End.”

    The primary baselines are GPT3-175B ROLLING, which generates 256 tokens at a time from the premise and previous story using a rolling 1,024-token context window, and ROLLING-FT, which is identical except that the generator is fine-tuned on several hundred WritingPrompts passages of at least 3,000 tokens.

    Amazon Mechanical Turk workers compare a Re3 story with a baseline story from the same premise. They judge interest, overarching plot coherence, premise relevance, and whether each story appears human-written. They also mark narration changes, factual inconsistencies or odd details, confusing passages, repetition, and disfluency; these binary problem indicators are summed as “Misc. Problems.” Each story pair is judged by three workers.

  7. Knowl 7 — Re3 improves coherence and premise relevance over rolling baselines

    data/table

    Human pairwise evaluations show that Re3 improves long-range plot coherence and premise relevance over both GPT3 rolling-window baselines. The percentages are the fractions of stories judged interesting, coherent, relevant, or humanlike; lower values are better for Misc. Problems. The two Re3 rows belong to separate pairwise experiments, so they should not be combined.

    Could not parse LaTeX table

    Relative to ROLLING, Re3 raises the coherent-story fraction by 14.3 percentage points and the relevant-story fraction by 20.0 points. Relative to ROLLING-FT, the increases are 11.3 and 16.0 points. The paper reports statistically significant improvements in coherence and relevance, along with fewer aggregate writing problems. Annotators judged 83.3% and 80.0% of Re3 stories humanlike in the two comparisons.

  8. Knowl 8 — Planning and reranking are the critical Re3 components

    data/table

    Ablation experiments remove the Plan, Rewrite, or Edit module while retaining the other applicable components. The values are human-evaluation percentages for interest, coherence, relevance, and humanlikeness, followed by the average number of miscellaneous writing problems; lower is better for the last metric.

    Could not parse LaTeX table

    Removing the Plan module substantially reduces coherence and relevance, showing that the structured plan and recursive prompt composition are important. Removing Rewrite also substantially reduces coherence and relevance, showing that selecting among alternative continuations is important. Removing Edit has little effect on the main metrics, and the full system does not consistently reduce continuity errors through its current editing implementation.

  9. Knowl 9 — Structured detection outperforms sentence-level entailment for character contradictions

    empirical result

    The Edit module’s contradiction detector was evaluated in a controlled dataset of 50 setups. Each original setup and its altered version contained a manually introduced contradiction in one character description; stories were generated from both setups until the contradicted attribute appeared. The resulting 200 setup–story pairs included consistent pairs (s,t)(s,t) and (s′,t′)(s',t') and contradictory pairs (s,t′)(s,t') and (s′,t)(s',t).

    The structured detector was compared with two baselines using the same BART-Large MNLI entailment model. ENTAILMENT compares every setup sentence with every story sentence, while ENTAILMENT-DPR compares each story sentence only with the setup sentence judged most relevant by DPR. ROC-AUC scores for predicting contradiction probability were:

    Could not parse LaTeX table

    The structured detector therefore outperforms both entailment baselines in this simplified character-attribute setting, but its absolute performance remains low. The paper reports that full stories also contain contradictions outside character attributes and that the correction API can introduce unwanted edits or fail on larger changes.

  10. Knowl 10 — Evaluation and Edit-module limitations constrain the conclusions

    limitation

    The paper’s long-form evaluation is limited by the cost and noisiness of human annotation. The experiments use relatively small samples, and workers quickly read or skim stories of several thousand words; agreement is consequently low. Many prompt designs, reranking rules, and contradiction thresholds were selected manually rather than through systematic validation.

    The current Edit module does not improve the main human metrics and handles only character-attribute contradictions. It can miss non-character inconsistencies involving settings, scenes, themes, pacing, foreshadowing, or facts that change over time, while false positives can arise when an attribute legitimately changes. The contradiction detector’s controlled ROC-AUC is still low, and the GPT-3 editing API may make unnecessary changes.

    The system is also strongly dependent on a high-quality general-purpose language model and may perform worse in languages without models comparable to GPT-3. Adapting the prompts and structured attribute representation to domains other than story generation would require additional manual redesign.

Coverage note — Detailed generated-story examples, API-usage costs, individual writing-problem breakdowns, annotator-agreement tables, and the exploratory length comparison were omitted because they provide supporting implementation or diagnostic detail rather than separate load-bearing contributions.

References

  1. 1.Alan Akbik, Duncan Blythe, and Roland Vollgraf. 2018. Contextual string embeddings for sequence labeling. In COLING 2018, 27th International Conference on Computational Linguistics, pages 1638–1649.
  2. 2.Nader Akoury, Shufan Wang, Josh Whiting, Stephen Hood, Nanyun Peng, and Mohit Iyyer. 2020. Storium: A dataset and evaluation platform for machine-in-the-loop story generation. In the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP).
  3. 3.Regina Barzilay and Mirella Lapata. 2008. Modeling local coherence: An entity-based approach. Computational Linguistics, 34(1):1–34.
  4. 4.Iz Beltagy, Matthew E Peters, and Arman Cohan. 2020. Longformer: The long-document transformer. arXiv preprint arXiv:2004.05150.
  5. 5.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  6. 6.Louis Castricato, Stella Biderman, David Thue, and Rogelio Cardona-Rivera. 2021. Towards a model-theoretic view of narratives. In Proceedings of the Third Workshop on Narrative Understanding, pages 95–104.
  7. 7.Eugene Charniak. 1972. Toward a model of children’s story comprehension. Ph.D. thesis, Massachusetts Institute of Technology.
  8. 8.John Joon Young Chung, Wooseok Kim, Kang Min Yoo, Hwaran Lee, Eytan Adar, and Minsuk Chang. 2022. Talebrush: Sketching stories with generative pretrained language models. In CHI Conference on Human Factors in Computing Systems, pages 1–19.
  9. 9.Andy Coenen, Luke Davis, Daphne Ippolito, Emily Reif, and Ann Yuan. 2021. Wordcraft: a human-ai collaborative editor for story writing. arXiv preprint arXiv:2107.07430.
  10. 10.Sumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung, Eric Frank, Piero Molino, Jason Yosinski, and Rosanne Liu. 2019. Plug and play language models: A simple approach to controlled text generation. arXiv preprint arXiv:1912.02164.
  11. 11.Oren Etzioni, Michele Banko, Stephen Soderland, and Daniel S Weld. 2008. Open information extraction from the web. Communications of the ACM, 51(12):68–74.
  12. 12.Angela Fan, Mike Lewis, and Yann Dauphin. 2018. Hierarchical neural story generation. arXiv preprint arXiv:1805.04833.
  13. 13.Angela Fan, Mike Lewis, and Yann Dauphin. 2019. Strategies for structuring story generation. arXiv preprint arXiv:1902.01109.
  14. 14.Seraphina Goldfarb-Tarrant, Tuhin Chakrabarty, Ralph Weischedel, and Nanyun Peng. 2020. Content planning for neural story generation with aristotelian rescoring. arXiv preprint arXiv:2009.09870.
  15. 15.Seraphina Goldfarb-Tarrant, Haining Feng, and Nanyun Peng. 2019. Plan, write, and revise: an interactive system for open-domain story generation. arXiv preprint arXiv:1904.02357.
  16. 16.Albert Gu, Karan Goel, and Christopher Ré. 2021. Efficiently modeling long sequences with structured state spaces. arXiv preprint arXiv:2111.00396.
  17. 17.Jian Guan, Fei Huang, Zhihao Zhao, Xiaoyan Zhu, and Minlie Huang. 2020. A knowledge-enhanced pre-training model for commonsense story generation. Transactions of the Association for Computational Linguistics, 8:93–108.
  18. 18.Rujun Han, Hong Chen, Yufei Tian, and Nanyun Peng. 2022. Go back in time: Generating flashbacks in stories with event temporal prompts. In 2022 Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL).
  19. 19.James A Hanley and Barbara J McNeil. 1982. The meaning and use of the area under a receiver operating characteristic (roc) curve. Radiology, 143(1):29–36.
  20. 20.Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2019. The curious case of neural text degeneration. arXiv preprint arXiv:1904.09751.
  21. 21.Daphne Ippolito, David Grangier, Chris Callison-Burch, and Douglas Eck. 2019. Unsupervised hierarchical story infilling. In Proceedings of the First Workshop on Narrative Understanding, pages 37–43.
  22. 22.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. arXiv preprint arXiv:2004.04906.
  23. 23.Daniel Khashabi, Sewon Min, Tushar Khot, Ashish Sabharwal, Oyvind Tafjord, Peter Clark, and Hannaneh Hajishirzi. 2020. Unifiedqa: Crossing format boundaries with a single qa system. arXiv preprint arXiv:2005.00700.
  24. 24.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. arXiv preprint arXiv:2205.11916.
  25. 25.Ben Krause, Akhilesh Deepak Gotmare, Bryan McCann, Nitish Shirish Keskar, Shafiq Joty, Richard Socher, and Nazneen Fatema Rajani. 2020. Gedi: Generative discriminator guided sequence generation. arXiv preprint arXiv:2009.06367.
  26. 26.Chia-Hsuan Lee, Hao Cheng, and Mari Ostendorf. 2021. Dialogue state tracking with a language model using schema-driven prompting. arXiv preprint arXiv:2109.07506.
  27. 27.Mina Lee, Percy Liang, and Qian Yang. 2022. Coauthor: Designing a human-ai collaborative writing dataset for exploring language model capabilities. arXiv preprint arXiv:2201.06796.
  28. 28.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2019. BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. CoRR, abs/1910.13461.
  29. 29.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tomáš Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing Systems, 33:9459–9474.
  30. 30.Boyang Li, Stephen Lee-Urban, George Johnston, and Mark Riedl. 2013. Story generation with crowdsourced plot graphs. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 27, pages 598–604.
  31. 31.Zhiyu Lin and Mark O Riedl. 2021. Plug-and-blend: a framework for plug-and-play controllable story generation with sketches. In Proceedings of the AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment, volume 17, pages 58–65.
  32. 32.Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2021. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. arXiv preprint arXiv:2107.13586.
  33. 33.Li Lucy and David Bamman. 2021. Gender and representation bias in gpt-3 generated stories. In Proceedings of the Third Workshop on Narrative Understanding, pages 48–55.
  34. 34.Yishu Miao and Phil Blunsom. 2016. Language as a latent variable: Discrete generative models for sentence compression. arXiv preprint arXiv:1609.07317.
  35. 35.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. arXiv preprint arXiv:2203.02155.
  36. 36.Nanyun Peng, Marjan Ghazvininejad, Jonathan May, and Kevin Knight. 2018. Towards controllable story generation. In NAACL Story-NLP Workshop.
  37. 37.Lianhui Qin, Antoine Bosselut, Ari Holtzman, Chandra Bhagavatula, Elizabeth Clark, and Yejin Choi. 2019. Counterfactual story reasoning and generation. arXiv preprint arXiv:1909.04076.
  38. 38.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners.
  39. 39.Hannah Rashkin, Asli Celikyilmaz, Yejin Choi, and Jianfeng Gao. 2020. Plotmachines: Outline-conditioned generation with dynamic plot state tracking. arXiv preprint arXiv:2004.14967.
  40. 40.Victor Sanh, Albert Webson, Colin Raffel, Stephen H Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, et al. 2021. Multitask prompted training enables zero-shot task generalization. arXiv preprint arXiv:2110.08207.
  41. 41.Abigail See, Aneesh Pappu, Rohun Saxena, Akhila Yerukola, and Christopher D Manning. 2019. Do massively pretrained language models make better storytellers? arXiv preprint arXiv:1909.10705.
  42. 42.Emma Strubell, Ananya Ganesh, and Andrew McCallum. 2019. Energy and policy considerations for deep learning in nlp. arXiv preprint arXiv:1906.02243.
  43. 43.Yufei Tian and Nanyun Peng. 2022. Zero-shot sonnet generation with discourse-level planning and aesthetics features. In 2022 Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL).
  44. 44.Scott R Turner. 1994. The creative process: A computer model of storytelling and creativity.
  45. 45.Rose E Wang, Esin Durmus, Noah Goodman, and Tatsunori Hashimoto. 2022. Language modeling via stochastic processes. arXiv preprint arXiv:2203.11370.
  46. 46.Su Wang, Greg Durrett, and Katrin Erk. 2020. Narrative interpolation for generating and understanding stories. arXiv preprint arXiv:2008.07466.
  47. 47.Tianming Wang and Xiaojun Wan. 2019. T-cvae: Transformer-based conditioned variational autoencoder for story completion. In IJCAI, pages 5233–5239.
  48. 48.Adina Williams, Nikita Nangia, and Samuel Bowman. 2018. A broad-coverage challenge corpus for sentence understanding through inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1112–1122. Association for Computational Linguistics.
  49. 49.Yuhuai Wu, Albert Q Jiang, Wenda Li, Markus N Rabe, Charles Staats, Mateja Jamnik, and Christian Szegedy. 2022. Autoformalization with large language models. arXiv preprint arXiv:2205.12615.
  50. 50.Peng Xu, Mostofa Patwary, Mohammad Shoeybi, Raul Puri, Pascale Fung, Anima Anandkumar, and Bryan Catanzaro. 2020. Megatron-cntrl: Controllable story generation with external knowledge using large-scale language models. arXiv preprint arXiv:2010.00840.
  51. 51.Kevin Yang and Dan Klein. 2021. Fudge: Controlled text generation with future discriminators. arXiv preprint arXiv:2104.05218.
  52. 52.Lili Yao, Nanyun Peng, Ralph Weischedel, Kevin Knight, Dongyan Zhao, and Rui Yan. 2019. Plan-and-write: Towards better automatic storytelling. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 7378–7385.
  53. 53.Meng-Hsuan Yu, Juntao Li, Danyang Liu, Dongyan Zhao, Rui Yan, Bo Tang, and Haisong Zhang. 2020. Draft and edit: Automatic storytelling through multi-pass hierarchical conditional variational autoencoder. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 1741–1748.
  54. 54.Ruiqi Zhong, Kristy Lee, Zheng Zhang, and Dan Klein. 2021. Adapting language models for zero-shot learning by meta-tuning on dataset and prompt collections. arXiv preprint arXiv:2104.04670.

Citation

MLA
Yang, K., et al. “Re3: Generating Longer Stories With Recursive Reprompting and Revision”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 4393–479, https://doi.org/10.18653/v1/2022.emnlp-main.296.
APA
Yang, K., Tian, Y., Peng, N., & Klein, D. (2022). Re3: Generating Longer Stories With Recursive Reprompting and Revision. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 4393–4479. https://doi.org/10.18653/v1/2022.emnlp-main.296
Chicago
Yang, K., Y. Tian, N. Peng, and D. Klein. 2022. “Re3: Generating Longer Stories With Recursive Reprompting and Revision”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 4393–4479. https://doi.org/10.18653/v1/2022.emnlp-main.296.
Harvard
Yang, K. et al. (2022) “Re3: Generating Longer Stories With Recursive Reprompting and Revision”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 4393–4479. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.296.
Vancouver
1. Yang K, Tian Y, Peng N, Klein D (2022) Re3: Generating Longer Stories With Recursive Reprompting and Revision. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 4393–4479

BibTeX

@inproceedings{yang-etal-2022-re3,
    title = "Re3: Generating Longer Stories With Recursive Reprompting and Revision",
    author = "Yang, Kevin  and
      Tian, Yuandong  and
      Peng, Nanyun  and
      Klein, Dan",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.296/",
    doi = "10.18653/v1/2022.emnlp-main.296",
    pages = "4393--4479"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/