Hierarchical Neural Story Generation

Angela FanMike LewisYann Dauphin

article2018ACL2,105 citations

Presents a hierarchical story generation framework and a 300,000-example prompt-story dataset that substantially improves narrative coherence by first generating a high-level premise before expanding it into full text using gated multi-scale self-attention.

Listen

Automated story generation represents a major frontier in artificial intelligence, requiring systems to maintain thematic consistency across long passages and plan high-level narrative plots. Traditional sequential text generation models struggle with these demands, often drifting off-topic or degenerating into generic phrasing because they generate text strictly word-by-word without broader context. Addressing this challenge is vital for advancing creative AI tools, automated content creation, and long-form narrative modeling.

The article demonstrates and evaluates a hierarchical neural architecture designed to produce fluent, topically relevant stories by decomposing generation into high-level premise creation followed by full-passage drafting. To support this, the authors introduce a new large-scale dataset, a gated multi-scale self-attention mechanism to capture long-range context efficiently, and a model fusion technique that enforces strict relevance between the generated story and its guiding premise.

The authors constructed a dataset of 303,358 human-written stories paired with writing prompts scraped from Reddit's WritingPrompts community, representing over 200 million words. Using this resource, they implemented a two-step framework: first generating a prompt using a convolutional language model, and then generating the corresponding story via a convolutional sequence-to-sequence model. The system incorporates deep multi-scale self-attention to capture past context across various time scales and leverages model fusion—training a second sequence model over a fixed, pretrained model—to ensure the generated narrative adheres to the prompt rather than defaulting to generic text. Evaluation relied on perplexity metrics, prompt-ranking accuracy across 1,000 test cases, and blind human evaluation studies on Amazon Mechanical Turk comparing story preference and prompt alignment.

The evaluation yielded several key findings. First, human evaluators preferred stories produced by the hierarchical model over a non-hierarchical baseline by more than two to one (67.3% preference). Second, combining the gated multi-scale self-attention mechanism with model fusion reduced test perplexity significantly, dropping it by roughly 9 points compared to standard convolutional baselines (from 45.54 down to 36.56). Third, human judges matched stories to their source prompts with substantially higher accuracy when using the fusion model, improving pairing performance by 7% over standard ensembling. Finally, while retrieval baselines quickly lose topical relevance as more stories are created, the generative fusion model sustains consistent relevance across an unlimited number of generated samples while maintaining lower word-overlap copying rates.

These findings indicate that introducing an explicit planning hierarchy alongside residual fusion architectures substantially mitigates the tendency of sequence models to lose coherence over multi-paragraph texts. By enabling the primary network to focus on narrative premises while the fused secondary layer generates rare, prompt-specific words, this strategy reduces the risk of repetitive, generic outputs. Consequently, the approach significantly improves text generation performance without requiring costly retrieval search over massive static databases.

Organizations developing creative AI tools or long-form document generators should adopt hierarchical planning architectures and multi-scale attention rather than single-pass language models. When implementing these systems, teams should utilize restricted top-k random sampling rather than beam search to avoid repetitive phrasing. Further work should focus on refining premise-generation models to produce more diverse, imaginative concepts, as well as addressing formatting artifacts and local repetition.

Confidence in these findings is high given the large dataset and strong alignment between automated metrics and human evaluations. However, practitioners should note certain limitations: prompt generation remains prone to producing common, generic themes, and sampling occasionally introduces minor grammatical artifacts, dialogue formatting errors, or local word repetition.

No sufficiently relevant recommendations were found.

Cover for Hierarchical Neural Story Generation

Abstract

We explore story generation: creative systems that can build coherent and fluent passages of text about a topic. We collect a large dataset of 300K human-written stories paired with writing prompts from an online forum. Our dataset enables hierarchical story generation, where the model first generates a premise, and then transforms it into a passage of text. We gain further improvements with a novel form of model fusion that improves the relevance of the story to the prompt, and adding a new gated multi-scale self-attention mechanism to model long-range context. Experiments show large improvements over strong baselines on both automated and human evaluations. Human judges prefer stories generated by our approach to those from a strong non-hierarchical model by a factor of two to one.

Table of Contents

  • 1 Introduction
  • 2 Writing Prompts Dataset
  • 3 Approach
  • 3.1 Hierarchical Story Generation
  • 3.2 Efficient Learning with Convolutional Sequence-to-Sequence Model
  • 3.3 Modeling Unbounded Context with Gated Multi-Scale Self-attention
  • 3.4 Improving Relevance to Input Prompt with Model Fusion
  • 4 Related Work
  • 4.1 Story Generation
  • 4.2 Hierarchical Text Generation
  • 4.3 Fusion Models
  • 5 Experimental Setup
  • 5.1 Baselines
  • 5.2 Fusion Training
  • 5.3 Training
  • 5.4 Generation
  • 5.5 Evaluation
  • 6 Results
  • 7 Discussion
  • 7.1 Generation Quality
  • 7.2 Use of Attention
  • 8 Conclusion
  • References
  • 9 Appendix of Model Architectures
  • 9.1 GCNN Language Model + Self-Attention
  • 9.2 Conv seq2seq + self-attention
  • 9.3 Ensemble: Conv seq2seq + self-attention
  • 9.4 Fusion: Conv seq2seq + self-attention

Knowls

  1. Knowl 1 — Hierarchical Story Generation Framework

    model/method

    Hierarchical story generation decomposes open-ended creative story writing into a two-level generation pipeline:

    1. Prompt/Premise Generation: A convolutional language model (such as a Gated Convolutional Neural Network) first generates a high-level sentence or premise p=(p1,…,pM)\mathbf{p} = (p_1, \dots, p_M) that establishes the plot structure, topic, and character setting of the story.
    2. Story Generation: A sequence-to-sequence (seq2seq) model conditions on the generated premise p\mathbf{p} to produce the full passage of story text s=(s1,…,sN)\mathbf{s} = (s_1, \dots, s_N).

    Conditioning on an explicit textual premise provides high-level plot grounding, prevents language models from drifting off-topic across long texts, and enables macro-level narrative planning beyond local next-word prediction.

  2. Knowl 2 — Model Fusion for Sequence-to-Sequence Story Generation

    model/method

    To prevent sequence-to-sequence models from ignoring the input prompt and degenerating into unconditional language models, model fusion integrates a fixed, pre-trained sequence-to-sequence model with a secondary sequence-to-sequence model trained concurrently.

    At decoder timestep tt, let htPretrainedh_t^{\text{Pretrained}} denote the hidden state of the pre-trained seq2seq model and htTrainingh_t^{\text{Training}} denote the hidden state of the model being trained. Dynamic gating vectors gtg_t are computed via a linear projection of the concatenated hidden states followed by a sigmoid activation:

    gt=σ(W[htTraining;htPretrained]+b)g_t = \sigma\left(W [h_t^{\text{Training}}; h_t^{\text{Pretrained}}] + b\right)

    ht=gt⊙[htTraining;htPretrained]h_t = g_t \odot [h_t^{\text{Training}}; h_t^{\text{Pretrained}}]

    where WW and bb are trainable weights and biases, and ⊙\odot is element-wise multiplication. The gated representation hth_t is passed through fully connected layers with Gated Linear Unit (GLU) activations and layer normalization before final vocabulary projection.

    This mechanism acts as a residual/boosting learner: the pre-trained seq2seq network handles standard grammar and frequent words, forcing the secondary network to focus on prompt dependencies and domain-specific vocabulary.

  3. Knowl 3 — Gated Multi-Scale Self-Attention for Convolutional Decoders

    model/method

    To overcome the bounded context window of standard convolutional decoders without recurrent bottlenecking, decoders are augmented with multi-head gated self-attention operating over past generation timesteps at varying time scales.

    Key components of the self-attention mechanism include:

    • Gated Transformations: Queries qq, keys kk, and values vv are computed using deep neural sub-networks with Gated Linear Unit (GLU) activations rather than standard linear projections, providing fine-grained selection capacity.
    • Multi-Scale Downsampling: Different attention heads operate at distinct temporal scales by downsampling inputs by varying stride factors (e.g., head 1 receives every timestep, head 2 every second timestep, head 3 every third timestep), enforcing distinct temporal resolution per head and sharpening attention distributions.
    • Null Attention Option and Past Masking: The model can optionally attend to a 0\mathbf{0} vector if past context is uninformative, allowing the architecture to revert to local convolution. Attention is masked strictly to past timesteps 0,…,t−10, \dots, t-1.

    The output representation of a single attention head for hidden states h0:tLh_{0:t}^L at decoder layer LL is:

    h0:tL+1=Linear(v(h0:t−1L)⊙softmax(q(h0:tL)k(h0:tL)⊤))h_{0:t}^{L+1} = \text{Linear}\left( v(h_{0:t-1}^L) \odot \text{softmax}\left( q(h_{0:t}^L) k(h_{0:t}^L)^\top \right) \right)

    Outputs across multiple heads are concatenated and linearly projected.

  4. Knowl 4 — WritingPrompts Story Generation Benchmark Dataset

    data/table

    The WRITINGPROMPTS dataset is a large corpus scraped from Reddit's r/WritingPrompts community over three years, pairing diverse user-submitted writing prompts with creative story responses.

    Stories were filtered to remove moderator notices, automated bot posts, deleted submissions, and stories shorter than 30 words. Prompts and stories are tokenized with NLTK, limiting the vocabulary to tokens appearing more than 10 times (yielding 19,025 prompt tokens and 104,960 story tokens). Stories are truncated to a maximum length of 1,000 words.

    Statistic Value
    # Train Stories 272,600
    # Test Stories 15,138
    # Validation Stories 15,620
    # Prompt Words 7.7M
    # Story Words 200M
    Average Length of Prompts 28.4 words
    Average Length of Stories 734.5 words
  5. Knowl 5 — Top-k Random Sampling Generation Strategy

    algorithm

    To prevent repetition, generic phrase loops, and truncation associated with beam search, stories are generated using top-kk random sampling combined with softmax temperature scaling.

    Input: Decoder model parameterized with vocabulary VV, context prompt p\mathbf{p}, top candidates cutoff k=10k = 10, temperature τ>0\tau > 0, max generation length TT, end-of-story token <eos>\text{<eos>}
    Output: Generated story tokens s1:Ls_{1:L}
    s0←<sos>s_0 \leftarrow \text{<sos>}
    t←1t \leftarrow 1
    while t≤Tt \le T do
        zt←DecoderLogits(s0:t−1,p)z_t \leftarrow \text{DecoderLogits}(s_{0:t-1}, \mathbf{p})
        pt(w)←exp⁡(zt,w/τ)∑v∈Vexp⁡(zt,v/τ)∀w∈Vp_t(w) \leftarrow \frac{\exp(z_{t,w} / \tau)}{\sum_{v \in V} \exp(z_{t,v} / \tau)} \quad \forall w \in V
        Vk←{w∈V:pt(w) is among the top k largest probabilities}V_k \leftarrow \{w \in V : p_t(w) \text{ is among the top } k \text{ largest probabilities}\}
        Pk(w)←pt(w)∑u∈Vkpt(u)∀w∈VkP_k(w) \leftarrow \frac{p_t(w)}{\sum_{u \in V_k} p_t(u)} \quad \forall w \in V_k
        st∼Categorical(Vk,Pk)s_t \sim \text{Categorical}(V_k, P_k)
        if st=<eos>s_t = \text{<eos>} then
            break
        end if
        t←t+1t \leftarrow t + 1
    end while
    return s1:ts_{1:t}
  6. Knowl 6 — Language Modeling Perplexity on WritingPrompts

    data/table

    Language modeling perplexity results evaluated on the WRITINGPROMPTS dataset show that sequence-to-sequence convolutional architectures with self-attention and model fusion substantially outperform recurrent baselines and standard convolutional seq2seq models.

    Model # Parameters (mil) Valid Perplexity Test Perplexity
    GCNN LM 123.4 54.50 54.79
    GCNN + self-attention LM 126.4 51.84 51.18
    LSTM seq2seq 110.3 46.83 46.79
    Conv seq2seq 113.0 45.27 45.54
    Conv seq2seq + self-attention 134.7 37.37 37.94
    Ensemble: Conv seq2seq + self-attention 270.3 36.63 36.93
    Fusion: Conv seq2seq + self-attention 255.4 36.08 36.56

    The fusion model achieves the lowest validation and test perplexity (36.08 and 36.56), outperforming both an ensemble of two self-attentive convolutional seq2seq models (36.93 test perplexity) with fewer total parameters, and improving upon the basic Conv seq2seq model by approximately 9 perplexity points.

  7. Knowl 7 — Component Ablation of Gated Multi-Scale Self-Attention

    data/table

    Ablation of individual architectural additions to the self-attention mechanism within the convolutional seq2seq decoder reveals cumulative improvements in validation and test perplexity on WRITINGPROMPTS:

    Model Configuration Valid Perplexity Test Perplexity
    Conv seq2seq 45.27 45.54
    + self-attention 42.01 42.32
    + multihead 40.12 40.39
    + multiscale 38.76 38.91
    + gating 37.37 37.94

    Adding baseline self-attention reduces test perplexity by 3.22 points. Implementing multi-head attention provides an additional 1.93 point drop, multi-scale downsampling adds another 1.48 point drop, and GLU gating across queries, keys, and values provides an additional 0.97 point drop, totaling a 7.6 point reduction over standard convolutional seq2seq.

  8. Knowl 8 — Human and Automatic Evaluation of Story Adherence and Hierarchical Generation

    empirical result

    Empirical evaluations on WRITINGPROMPTS demonstrate strong gains in topical adherence and human narrative preference when using model fusion and hierarchical generation:

    • Human Triple-Pairing Task: Human evaluators matched 105 generated 150-word stories back to their original test prompts (presented in shuffled groups of 3 with 15 judges per question). Human accuracy was 72.3% for the Fusion model (approaching human ground truth stories at 88.4%), compared to 66.1% for an Ensemble, 64.8% for Conv Seq2Seq + Self-Attention, 51.4% for Conv Seq2Seq, 71.9% for 1-NN retrieval, and 17.0% for unconditional GCNN + Self-Attention LM.
    • Prompt Ranking Accuracy: When scoring the likelihood of 1,000 test stories decoded under 10 candidate prompts (1 true, 9 random distractors), the Fusion model identified the true prompt with 16.3% accuracy, outperforming the Ensemble (13.8%), Conv Seq2Seq + Self-Att (12.6%), and baseline Conv Seq2Seq (10.7%).
    • Hierarchical Generation Blind Preference: In a blind A/B evaluation of 400 story pairs evaluated by 5 judges each, human raters preferred stories generated by the hierarchical system (GCNN LM prompt generation followed by fusion seq2seq story generation) over an unconditional GCNN LM by a margin of 67.32% to 32.68% (over 2 to 1).
  9. Knowl 9 — Text Novelty and Story Diversity vs. Retrieval Baselines

    empirical result

    Evaluation of generative story models versus retrieval baselines (TF-IDF FASTTEXT vectors queried with FAISS kk-nearest neighbors) indicates:

    • Novelty vs. Copying: On 500 generated 150-word test stories, the average Longest Common Subsequence (LCS) between the generated text and the training set was 8.9 words for the Fusion model, 10.2 words for the standard Conv seq2seq model, and 150 words for the kk-NN baseline (which copies entire training stories verbatim).
    • Prompt Diversity: When evaluating the connection between a prompt and multiple generated stories (kk-th best story ranking), kk-NN relevance degrades sharply after the single nearest neighbor, whereas the generative fusion model sustains prompt topicality across repeated story generations.
  10. Knowl 10 — Limitations of Hierarchical Story Generation and Sampling

    limitation

    The hierarchical neural story generation framework exhibits several specific operational limitations:

    • Subword and Formatting Artifacts: Top-kk sampling occasionally samples isolated subword tokens without subsequent completions (e.g., generating ca without n't) or fails to output delimiter/newline tokens between consecutive dialogue turns.
    • Repetition Loops: The self-attention mechanism frequently attends to recently generated words rather than distant context, leading to repetitive phrases and local looping.
    • Generic Prompt Generation: The prompt-generating language model struggles with rare vocabulary and named entities characteristic of creative fantasy domains (e.g., Harry Potter or Game of Thrones), falling back onto generic sentence openings (e.g., frequently opening with The man...).
    • Static Encoder Attention: Unlike machine translation where encoder attention shifts continuously across decoding steps, encoder-decoder attention maps in story fusion remain largely stationary throughout the entire generation process, focusing continuously on a small set of salient prompt keywords.

Coverage note — None: all main contributions (the WRITINGPROMPTS dataset, the hierarchical generation framework, gated multi-scale self-attention, model fusion, top-k decoding, quantitative perplexity and ablation tables, human and automatic adherence benchmarks, and qualitative error analyses) are fully represented.

References

  1. 1.Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016. Layer normalization. arXiv preprint arXiv:1607.06450.
  2. 2.Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015. Neural machine translation by jointly learning to align and translate. International Conference on Learning Representation (ICLR).
  3. 3.Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. 2016. Enriching word vectors with subword information. arXiv preprint arXiv:1607.04606.
  4. 4.Jan Chorowski and Navdeep Jaitly. 2016. Towards better decoding and language model integration in sequence to sequence models. arXiv preprint arXiv:1612.02695.
  5. 5.Yann N. Dauphin, Angela Fan, Michael Auli, and David Grangier. 2017. Language modeling with gated convolutional networks.
  6. 6.Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann Dauphin. 2017. Convolutional sequence to sequence learning.
  7. 7.Caglar Gulcehre, Orhan Firat, Kelvin Xu, Kyunghyun Cho, Loic Barrault, Huei-Chi Lin, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2015. On using monolingual corpora in neural machine translation. arXiv preprint arXiv:1503.03535.
  8. 8.Brent Harrison, Christopher Purdy, and Mark O Riedl. 2017. Toward automated story generation with markov chain monte carlo methods and deep neural networks.
  9. 9.Parag Jain, Priyanka Agrawal, Abhijit Mishra, Mohak Sukhwani, Anirban Laha, and Karthik Sankaranarayanan. 2017. Story generation from sequence of independent short descriptions. arXiv preprint arXiv:1707.05501.
  10. 10.Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2017. Billion-scale similarity search with gpus. arXiv preprint arXiv:1702.08734.
  11. 11.Rafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu. 2016. Exploring the limits of language modeling. arXiv preprint arXiv:1602.02410.
  12. 12.Ryan Kiros, Yukun Zhu, Ruslan Salakhutdinov, Richard S Zemel, Antonio Torralba, Raquel Urtasun, and Sanja Fidler. 2015. Skip-thought vectors. arXiv preprint arXiv:1506.06726.
  13. 13.Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan. 2015a. A diversity-promoting objective function for neural conversation models. arXiv preprint arXiv:1510.03055.
  14. 14.Jiwei Li, Minh-Thang Luong, and Dan Jurafsky. 2015b. A hierarchical neural autoencoder for paragraphs and documents. arXiv preprint arXiv:1506.01057.
  15. 15.Peter J. Liu, Mohammad Saleh, Etienne Pot, Ben Goodrich, Ryan Sepassi, Lukasz Kaiser, and Noam Shazeer. 2018. Generating wikipedia by summarizing long sequences. arXiv preprint arXiv:1801.10198.
  16. 16.Lara J Martin, Prithviraj Ammanabrolu, William Hancock, Shruti Singh, Brent Harrison, and Mark O Riedl. 2017. Event representations for automated story generation with deep neural nets. arXiv preprint arXiv:1706.01331.
  17. 17.Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio. 2013. On the difficulty of training recurrent neural networks. In ICML.
  18. 18.Prajit Ramachandran, Peter J Liu, and Quoc V Le. 2016. Unsupervised pretraining for sequence to sequence learning. arXiv preprint arXiv:1611.02683.
  19. 19.Melissa Roemmele. 2016. Writing stories with help from recurrent neural networks. In AAAI.
  20. 20.Alexander M Rush, Sumit Chopra, and Jason Weston. 2015. A neural attention model for abstractive sentence summarization. arXiv preprint arXiv:1509.00685.
  21. 21.Louis Shao, Stephan Gouws, Denny Britz, Anna Goldie, Brian Strope, and Ray Kurzweil. 2017. Generating long and diverse responses with neural conversation models. arXiv preprint arXiv:1701.03185.
  22. 22.Anuroop Sriram, Heewoo Jun, Sanjeev Satheesh, and Adam Coates. 2017. Cold fusion: Training seq2seq models together with language models. arXiv preprint arXiv:1708.06426.
  23. 23.Sainbayar Sukhbaatar, Jason Weston, Rob Fergus, et al. 2015. End-to-end memory networks. In Advances in neural information processing systems, pages 2440–2448.
  24. 24.Ilya Sutskever, James Martens, George E. Dahl, and Geoffrey E. Hinton. 2013. On the importance of initialization and momentum in deep learning. In ICML.
  25. 25.Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014. Sequence to sequence learning with neural networks. In Neural Information Processing Systems (NIPS).
  26. 26.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems, pages 6000–6010.
  27. 27.Ashwin K Vijayakumar, Michael Cogswell, Ramprasath R Selvaraju, Qing Sun, Stefan Lee, David Crandall, and Dhruv Batra. 2016. Diverse beam search: Decoding diverse solutions from neural sequence models. arXiv preprint arXiv:1610.02424.
  28. 28.Sam Wiseman, Stuart M Shieber, and Alexander M Rush. 2017. Challenges in data-to-document generation. arXiv preprint arXiv:1707.08052.
  29. 29.Denis Yarats and Mike Lewis. 2017. Hierarchical text generation and planning for strategic dialogue. arXiv preprint arXiv:1712.05846.
  30. 30.Xingxing Zhang and Mirella Lapata. 2014. Chinese poetry generation with recurrent neural networks. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 670–680.

Citation

MLA
Fan, A., et al. “Hierarchical Neural Story Generation”. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2018, pp. 889–98, https://doi.org/10.18653/v1/P18-1082.
APA
Fan, A., Lewis, M., & Dauphin, Y. (2018). Hierarchical Neural Story Generation. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 889–898. https://doi.org/10.18653/v1/P18-1082
Chicago
Fan, A., M. Lewis, and Y. Dauphin. 2018. “Hierarchical Neural Story Generation”. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 889–98. https://doi.org/10.18653/v1/P18-1082.
Harvard
Fan, A., Lewis, M. and Dauphin, Y. (2018) “Hierarchical Neural Story Generation”, Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 889–898. Available at: https://doi.org/10.18653/v1/P18-1082.
Vancouver
1. Fan A, Lewis M, Dauphin Y (2018) Hierarchical Neural Story Generation. In: Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 889–898

BibTeX

@inproceedings{fan-etal-2018-hierarchical,
    title = "Hierarchical Neural Story Generation",
    author = "Fan, Angela  and
      Lewis, Mike  and
      Dauphin, Yann",
    editor = "Gurevych, Iryna  and
      Miyao, Yusuke",
    booktitle = "Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2018",
    address = "Melbourne, Australia",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/P18-1082/",
    doi = "10.18653/v1/P18-1082",
    pages = "889--898"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/