Prompt Waywardness: The Curious Case of Discretized Interpretation of Continuous Prompts

Daniel KhashabiXinxi LyuSewon MinLianhui QinKyle RichardsonSean WelleckHannaneh HajishirziTushar KhotAshish SabharwalSameer Singh

article2022NAACL93 citations

Reveals that continuous prompts optimized for language models can be projected onto completely arbitrary or contradictory discrete text without losing task performance, exposing critical flaws in standard prompt interpretability methods.

Listen

Prompt tuning optimizes continuous numerical vectors rather than updating an entire language model's parameters, offering a lightweight approach to controlling AI systems. However, these continuous prompts are not inherently human-readable. Practitioners often attempt to interpret them by mapping continuous vectors to their nearest discrete words. As organizations increasingly rely on prompt-based AI for critical tasks, understanding whether these discrete textual interpretations faithfully reflect how models actually behave is essential for safety, governance, and model reliability.

The main objective of the article is to evaluate whether continuous prompts can be faithfully interpreted into human-readable text by mapping them to their nearest discrete words. The authors formulate and test the Prompt Waywardness hypothesis, which posits a fundamental disconnect between what a continuous prompt instructs a model to do and what its nearest-neighbor discrete text projection appears to say.

To evaluate this behavior, the authors conducted empirical experiments using the GPT-2 language model family across five distinct classification tasks. They modified standard continuous prompt tuning to jointly optimize downstream task accuracy while pulling the continuous prompt toward an arbitrarily chosen target text. The evaluation tested 62 distinct target texts—including instructions from unrelated tasks and random sentences—measuring both task accuracy retention and textual overlap. They also evaluated scaling across model sizes from 124 million to 1.5 billion parameters and tested various prompt lengths.

The findings reveal that continuous prompts can be optimized to solve a given task while simultaneously projecting to virtually any arbitrary target text, retaining performance within 2% of unconstrained prompts while achieving over 94% text overlap with the arbitrary target. Furthermore, continuous prompts forced to project to the true definitions of tasks showed no meaningful performance advantage over those projecting to completely irrelevant text, with both exhibiting a similar 1.5% average accuracy drop compared to unconstrained baselines. The disconnect between continuous prompts and their textual interpretations becomes more pronounced as model size increases and prompt length grows, due to the high expressive power in early network layers and the mathematical reality that infinitely many continuous vectors map to the same discrete token.

These results demonstrate that projecting continuous prompts to nearest-neighbor words does not yield a faithful explanation of model behavior. This creates severe security and governance risks, as malicious or biased model behaviors can be concealed beneath benign, harmless-looking textual descriptions, creating a false sense of security for auditors. Additionally, the findings indicate that attempting to discover human-readable discrete prompts purely through unconstrained continuous optimization will often yield degenerate, unfaithful solutions.

Organizations and developers should avoid relying on nearest-neighbor word projections to interpret or audit continuous prompts. AI teams developing interpretable prompting methods must incorporate domain-specific constraints rather than depending solely on continuous gradients. Further research should focus on developing novel architectural mechanisms and robust explanation frameworks that bridge the continuous-discrete representation gap.

The conclusions are supported by consistent results across multiple benchmark tasks, prompt lengths, and model sizes within the auto-regressive GPT-2 architecture. Readers should note that these findings specifically reflect standard nearest-neighbor projection techniques within existing transformer frameworks; evaluating alternative projection methods or newer model architectures remains an important area for future confirmation.

arXiv: 2112.08348
Cover for Prompt Waywardness: The Curious Case of Discretized Interpretation of Continuous Prompts

Abstract

Fine-tuning continuous prompts for target tasks has recently emerged as a compact alternative to full model fine-tuning. Motivated by these promising results, we investigate the feasibility of extracting a discrete (textual) interpretation of continuous prompts that is faithful to the problem they solve. In practice, we observe a “wayward” behavior between the task solved by continuous prompts and the nearest neighbor discrete projections of these prompts: One can find continuous prompts that solve a task while being projected to an arbitrary text (e.g., definition of a different or even a contradictory task) and simultaneously being within a very small (2%) margin of the best continuous prompt of the same size for the task. We provide intuitions behind this odd and surprising behavior, as well as extensive empirical analyses quantifying the effect of design choices. For instance, larger models exhibit higher waywardness, i.e, we can find prompts that more closely map to any arbitrary text with a smaller drop of accuracy. These findings have important implications relating to the difficulty of faithfully interpreting continuous prompts and their generalization across models and tasks, providing guidance for future progress in prompting language models.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Prompt Waywardness
  • 3.1 Preliminaries: Setup and Terminology
  • 3.2 The Waywardness Hypothesis
  • 3.3 Finding Wayward Prompts
  • 4 Empirical Support of Waywardness
  • 4.1 Setup
  • 4.2 Main Results
  • 4.3 Further Analysis
  • 5 Explaining Waywardness
  • 6 Implications of Prompt Waywardness
  • 7 Conclusion
  • Acknowledgment
  • References
  • Supplementary Material
  • A Additional Experimental Details
  • B The mapping between continuous and discrete space is not one-to-one
  • B.1 Proofs

Knowls

  1. Knowl 1 — Prompt Waywardness hypothesis

    theoretical result

    The Prompt Waywardness hypothesis proposes that a continuous prompt can perform a downstream task well while projecting to an arbitrary, even irrelevant or contradictory, text. Let MM be a frozen language model, DD a dataset for a target task, and LL the prompt length. A discrete prompt pd∈{0,1}L×Vp_d\in\{0,1\}^{L\times V} is a sequence of one-hot vectors over a vocabulary of size VV; a continuous prompt pc∈RL×dp_c\in\mathbb{R}^{L\times d} is a sequence of dd-dimensional vectors. Given an embedding matrix E∈RV×dE\in\mathbb{R}^{V\times d}, the continuous projection is c-proj(pd)=pdEc\text{-proj}(p_d)=p_dE. The discrete projection d-proj(pc)d\text{-proj}(p_c) replaces each continuous vector with the vocabulary item whose embedding has the greatest dot product with it.

    Let ℓ(p;D)\ell(p;D) denote the expected task loss of MM when prompt pp is prepended to examples from DD, and let pc∗p_c^* minimize this loss on the training data. The hypothesis is that, for any target pdp_d, a continuous prompt p~c\tilde p_c can be found such that d-proj(p~c)=pdd\text{-proj}(\tilde p_c)=p_d while its test loss differs from that of the optimal continuous prompt by less than a small gap Δ\Delta: ∣ℓ(p~c;Dtest)−ℓ(pc∗;Dtest)∣<Δ|\ell(\tilde p_c;D_{\mathrm{test}})-\ell(p_c^*;D_{\mathrm{test}})|<\Delta. The paper presents this as a hypothesis supported empirically, not as a general theorem; the gap may depend on prompt length, model, and dataset.

  2. Knowl 2 — Wayward prompts retain task accuracy across five classification datasets

    empirical result

    The central experiment compared unconstrained continuous prompts with continuous prompts trained to project to 62 irrelevant texts: 32 task instructions from Natural Instructions and 30 sentences sampled from The PILE. The downstream tasks were SST-2, SST-5, AGNews, Subj, and TREC. The model was GPT-2 Large (774M parameters); each constrained prompt used the same length as its unconstrained comparison. Scores are means over three random seeds and the target prompts in each source category. Prompt F1 measures word-level token overlap between the projected text and target text, ignoring punctuation and articles and applying lemmatization. The table reports relative accuracy drop, accuracy before and after the constraint, and prompt F1.

    Could not parse LaTeX table

    With the distance weight set to γ=0.01\gamma=0.01, four datasets averaged at least 94% prompt F1 with less than a 2% reported relative accuracy drop. TREC was the main exception: its average prompt F1 was 86.1% and its reported drop was 2.3%. Thus, continuous prompts could closely reproduce the target text under nearest-neighbor projection while largely preserving performance on a different task.

  3. Knowl 3 — Joint objective for finding wayward prompts

    model/method

    The paper searches for a wayward prompt by jointly minimizing task loss and distance from a chosen discrete target. A frozen model MM receives a continuous prompt pc∈RL×dp_c\in\mathbb{R}^{L\times d} together with task input; the prompt is the only learned parameter. For a discrete target pd∈{0,1}L×Vp_d\in\{0,1\}^{L\times V} and embedding matrix E∈RV×dE\in\mathbb{R}^{V\times d}, training uses the continuous-space distance to its embedding, normalized by prompt length:

    ℓ′(pc;D,γ)=ℓ(pc;D)+γ∥pc−pdE∥F2L.\ell'(p_c;D,\gamma)=\ell(p_c;D)+\gamma\frac{\|p_c-p_dE\|_F^2}{L}.

    Here DD is the task training dataset, ℓ\ell is the task loss, LL is prompt length, ∥⋅∥F\|\cdot\|_F is the Frobenius norm, and γ≥0\gamma\geq0 weights the projection constraint. Setting γ=0\gamma=0 gives the unconstrained prompt-tuning objective; increasing γ\gamma encourages proximity to the target embedding but can reduce task accuracy. The paper used γ=0.01\gamma=0.01 as a typical trade-off and initialized both constrained and unconstrained searches from pdEp_dE. Evaluation compared the nearest-neighbor discrete projection of the learned prompt with pdp_d using token-overlap F1.

  4. Knowl 4 — Continuous-to-discrete projection is infinitely many-to-one

    theoretical result

    For a lexicon embedded in Rd\mathbb{R}^d, the paper establishes that a nearest-neighbor projection under any metric maps infinitely many continuous vectors to each one-hot vocabulary item; these preimages include continuous regions. Applied positionwise, this means many continuous prompts can share the same discrete prompt projection. The result is stated for nearest-neighbor projection with arbitrary tie-breaking.

    The paper also states a broader result: if a projection operator from Rd\mathbb{R}^d to the one-hot vectors over a vocabulary of size VV is drawn uniformly from the space of such operators, then with probability one every one-hot vector has infinitely many preimages. These results concern the geometry of the projection, not whether any particular point in a preimage performs a desired task.

  5. Knowl 5 — Larger language models show greater waywardness

    empirical result

    On SST-2, the paper compared GPT-2 Small (124M parameters), Medium (355M), Large (774M), and XL (1.5B). For each size, it evaluated 30 PILE target prompts with three random seeds and selected among γ∈{0.01,0.005,0.003}\gamma\in\{0.01,0.005,0.003\} to obtain prompt F1 above 0.98. The constrained and unconstrained prompts were compared at the same size. The relative accuracy drop was 1.2% for Small, 0.5–0.7% for Medium and Large, and 0.2% for XL. The ability to match the target projection at high F1 therefore persisted across model sizes, while the accuracy cost generally became smaller for larger models.

  6. Knowl 6 — Prompt length improves projection fidelity and reduces the accuracy gap

    empirical result

    In the AGNews length analysis, prompt lengths were varied across L∈{4,7,14,28,56}L\in\{4,7,14,28,56\} with γ=0.01\gamma=0.01. The plotted results average over 32 target prompts and three random seeds. At length 4, the learned continuous prompts had difficulty matching the target text, with prompt F1 below 60%. At lengths around 14 or longer, prompt F1 was near 1.0 while task-accuracy loss remained small. Accuracy for both constrained and unconstrained prompts increased with prompt length, and their gap tended to narrow; the relative accuracy drop was marginal once prompts were at least 7 tokens long.

  7. Knowl 7 — Projecting to a true task definition gives no accuracy advantage

    empirical result

    The authors manually wrote one task-defining target text for each of the five classification datasets, then trained a constrained prompt with γ=0.01\gamma=0.01 and compared it with an unconstrained prompt of the same length. The table gives the reported relative accuracy drop, unconstrained-to-constrained accuracy, and prompt F1.

    Could not parse LaTeX table

    The mean reported gap for true task definitions was 1.6%, compared with 1.4% for the irrelevant target texts in the main experiment. Projecting a continuous prompt to text that genuinely describes the task therefore did not improve task-solving performance relative to projecting it to unrelated text.

  8. Knowl 8 — Projection fidelity and task accuracy trade off as distance weight changes

    empirical result

    For SST-2 and AGNews, the authors varied the distance weight γ\gamma from 0 to 0.03, averaging over 32 Natural Instructions target prompts and three random seeds. As γ\gamma increased, token-overlap F1 between the target text and the learned prompt’s discrete projection increased, while task accuracy generally decreased. The trade-off was often modest: the paper reports that projection F1 could be close to 1.00 with less than a 1% relative accuracy drop from the unconstrained prompt. This analysis shows that the degree of projection matching can be adjusted through the objective’s distance weight.

  9. Knowl 9 — Deep-model expressivity helps explain waywardness

    model/method

    The paper’s explanation combines the many-to-one geometry of nearest-neighbor projection with the expressive capacity of deep models. A fixed discrete projection corresponds to a region containing many continuous prompt vectors, rather than a unique vector. Continuous prompts enter the model before its first layer, where their effects can be transformed by the network’s successive layers. The authors therefore argue that even prompts within one projection region can have diverse task behavior, and that deeper models’ greater input expressivity helps explain why waywardness is stronger for larger models. This is offered as an explanatory intuition, alongside the experiments, rather than as a separately established causal result.

  10. Knowl 10 — A benign discrete projection does not certify prompt behavior

    limitation

    Under the architectures and projection methods studied, nearest-neighbor text interpretations are not reliable certificates of what a continuous prompt makes a model do. A continuous prompt could project to a benign task description while producing harmful or otherwise different behavior, and a limited evaluation set might fail to reveal that behavior—for example, if it affects a group absent from the evaluation cases. The same disconnect can make differentiable optimization for readable prompts degenerate: an optimization may achieve task utility and readable projected text even when that text is irrelevant or contradictory to the behavior. The paper frames these as risks and implications of its findings, not as demonstrated attacks in its experiments.

Coverage note — The discussion of task-specific initialization and gradient-based reverse engineering is omitted as a separate knowl because the paper presents these as implications rather than independently tested results; proof derivations are also excluded.

References

  1. 1.Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. arXiv preprint arXiv:2005.14165.
  2. 2.Sumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung, Eric Frank, Piero Molino, Jason Yosinski, and Rosanne Liu. 2019. Plug and play language models: A simple approach to controlled text generation. In Proceedings of ICLR.
  3. 3.Tianyu Gao, Adam Fisch, and Danqi Chen. 2021. Making pre-trained language models better few-shot learners. In Proceedings of ACL.
  4. 4.Chuan Guo, Alexandre Sablayrolles, Hervé Jégou, and Douwe Kiela. 2021. Gradient-based adversarial attacks against text transformers. In Proceedings of EMNLP.
  5. 5.Tatsunori B Hashimoto, David Alvarez-Melis, and Tommi S Jaakkola. 2016. Word embeddings as metric recovery in semantic spaces. TACL, 4:273–286.
  6. 6.Zhengbao Jiang, Frank F Xu, Jun Araki, and Graham Neubig. 2020. How can we know what language models know? TACL, 8:423–438.
  7. 7.Tushar Khot, Daniel Khashabi, Kyle Richardson, Peter Clark, and Ashish Sabharwal. 2021. Text Modular Networks: Learning to decompose tasks in the language of existing models. Proceedings of NAACL, page 1264–1279.
  8. 8.Sachin Kumar, Eric Malmi, Aliaksei Severyn, and Yulia Tsvetkov. 2021. Controlled text generation as continuous optimization with multiple constraints. Proceedings of NeurIPS, 34.
  9. 9.Teven Le Scao and Alexander M Rush. 2021. How many data points is a prompt worth? In Proceedings of NAACL, pages 2627–2636.
  10. 10.Brian Lester, Rami Al-Rfou, and Noah Constant. 2021. The power of scale for parameter-efficient prompt tuning. In Proceedings of EMNLP.
  11. 11.Xiang Lisa Li and Percy Liang. 2021. Prefix-tuning: Optimizing continuous prompts for generation. In Proceedings of ACL.
  12. 12.Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013. Distributed representations of words and phrases and their compositionality. In Proceedings of NeurIPS, pages 3111–3119.
  13. 13.Sewon Min, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022. Noisy channel language model prompting for few-shot text classification. In Proceedings of ACL.
  14. 14.Swaroop Mishra, Daniel Khashabi, Chitta Baral, Yejin Choi, and Hannaneh Hajishirzi. 2022a. Reframing instructional prompts to GPTk’s language. In Proceedings of ACL - Findings.
  15. 15.Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2022b. Cross-task generalization via natural language crowdsourcing instructions. In Proceedings of ACL.
  16. 16.Bo Pang and Lillian Lee. 2004. A sentimental education: Sentiment analysis using subjectivity summarization based on minimum cuts. In Proceedings of ACL.
  17. 17.Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019. Language models as knowledge bases? In Proceedings of EMNLP.
  18. 18.Guanghui Qin and Jason Eisner. 2021. Learning how to ask: Querying lms with mixtures of soft prompts. In Proceedings of NAACL.
  19. 19.Lianhui Qin, Vered Shwartz, Peter West, Chandra Bhagavatula, Jena D Hwang, Ronan Le Bras, Antoine Bosselut, and Yejin Choi. 2020. Backpropagation-based decoding for unsupervised counterfactual and abductive reasoning. In Proceedings of EMNLP.
  20. 20.Lianhui Qin, Sean Welleck, Daniel Khashabi, and Yejin Choi. 2022. COLD decoding: Energy-based constrained text generation with Langevin dynamics. arXiv preprint arXiv:2202.11705.
  21. 21.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  22. 22.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. JMLR, 21:1–67.
  23. 23.Maithra Raghu, Ben Poole, Jon Kleinberg, Surya Ganguli, and Jascha Sohl-Dickstein. 2017. On the expressive power of deep neural networks. In Proceedings of ICML, pages 2847–2854.
  24. 24.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. SQuAD: 100,000+ questions for machine comprehension of text. In Proceedings of EMNLP, pages 2383–2392.
  25. 25.Laria Reynolds and Kyle McDonell. 2021. Prompt programming for large language models: Beyond the few-shot paradigm. In Proceedings of CHI.
  26. 26.Timo Schick and Hinrich Schütze. 2021. Exploiting cloze-questions for few-shot text classification and natural language inference. In Proceedings of EACL, pages 255–269.
  27. 27.Lei Sha. 2020. Gradient-guided unsupervised lexically constrained text generation. In Proceedings of EMNLP, pages 8692–8703.
  28. 28.Taylor Shin, Yasaman Razeghi, Robert L Logan IV, Eric Wallace, and Sameer Singh. 2020. Eliciting knowledge from language models using automatically generated prompts. In Proceedings of EMNLP.
  29. 29.Dylan Slack, Sophie Hilgard, Emily Jia, Sameer Singh, and Himabindu Lakkaraju. 2020. Fooling lime and shap: Adversarial attacks on post hoc explanation methods. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, pages 180–186.
  30. 30.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of EMNLP.
  31. 31.Matus Telgarsky. 2016. Benefits of depth in neural networks. In Proceedings of COLT, pages 1517–1539.
  32. 32.Ellen M Voorhees and Dawn M Tice. 2000. Building a question answering test collection. In Proceedings of SIGIR.
  33. 33.Eric Wallace, Tony Zhao, Shi Feng, and Sameer Singh. 2021. Concealed data poisoning attacks on NLP models. In Proceedings of NAACL.
  34. 34.Ben Wang and Aran Komatsuzaki. 2021. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https://github.com/kingoflolz/mesh-transformer-jax.
  35. 35.Lifan Yuan, Yichi Zhang, Yangyi Chen, and Wei Wei. 2021. Bridge the gap between CV and NLP! a gradient-based textual adversarial attack framework. arXiv preprint arXiv:2110.15317.
  36. 36.Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015. Character-level convolutional networks for text classification. In Proceedings of NeurIPS.
  37. 37.Zexuan Zhong, Dan Friedman, and Danqi Chen. 2021. Factual probing is [mask]: Learning vs. learning to recall. In Proceedings of NAACL, pages 5017–5033.
  38. 38.Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. 2021. Learning to prompt for vision-language models. arXiv preprint arXiv:2109.01134.

Citation

MLA
Khashabi, D., et al. “Prompt Waywardness: The Curious Case of Discretized Interpretation of Continuous Prompts”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2022, pp. 3631–43, https://doi.org/10.18653/v1/2022.naacl-main.266.
APA
Khashabi, D., Lyu, X., Min, S., Qin, L., Richardson, K., Welleck, S., Hajishirzi, H., Khot, T., Sabharwal, A., Singh, S., & Choi, Y. (2022). Prompt Waywardness: The Curious Case of Discretized Interpretation of Continuous Prompts. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 3631–3643. https://doi.org/10.18653/v1/2022.naacl-main.266
Chicago
Khashabi, D., X. Lyu, S. Min, et al. 2022. “Prompt Waywardness: The Curious Case of Discretized Interpretation of Continuous Prompts”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 3631–43. https://doi.org/10.18653/v1/2022.naacl-main.266.
Harvard
Khashabi, D. et al. (2022) “Prompt Waywardness: The Curious Case of Discretized Interpretation of Continuous Prompts”, Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp. 3631–3643. Available at: https://doi.org/10.18653/v1/2022.naacl-main.266.
Vancouver
1. Khashabi D, Lyu X, Min S, et al (2022) Prompt Waywardness: The Curious Case of Discretized Interpretation of Continuous Prompts. In: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp 3631–3643

BibTeX

@inproceedings{khashabi-etal-2022-prompt,
    title = "Prompt Waywardness: The Curious Case of Discretized Interpretation of Continuous Prompts",
    author = "Khashabi, Daniel  and
      Lyu, Xinxi  and
      Min, Sewon  and
      Qin, Lianhui  and
      Richardson, Kyle  and
      Welleck, Sean  and
      Hajishirzi, Hannaneh  and
      Khot, Tushar  and
      Sabharwal, Ashish  and
      Singh, Sameer  and
      Choi, Yejin",
    editor = "Carpuat, Marine  and
      de Marneffe, Marie-Catherine  and
      Meza Ruiz, Ivan Vladimir",
    booktitle = "Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
    month = jul,
    year = "2022",
    address = "Seattle, United States",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.naacl-main.266/",
    doi = "10.18653/v1/2022.naacl-main.266",
    pages = "3631--3643"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/