Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data

Emily M. BenderAlexander Koller

article2020ACL1,415 citationsBest Theme Paper

Establishes foundational theoretical limits for large language models by demonstrating that systems trained exclusively on raw text form cannot, in principle, acquire semantic meaning or achieve genuine natural language understanding without grounding in external intent.

Listen

Recent advancements in large neural language models have spurred widespread enthusiasm and claims that artificial intelligence systems can now "understand" human language, comprehend text, and recall factual knowledge. However, this hype obscures a fundamental theoretical limitation in how these models operate. The article addresses this critical issue by evaluating whether computational systems trained solely on text prediction can learn the true meaning of natural language expressions.

The main objective of the article is to demonstrate that systems trained exclusively on linguistic form cannot, in principle, acquire meaning. To establish this, the article employs theoretical analysis, conceptual thought experiments, and a review of research in linguistics, developmental psychology, and diagnostic machine learning studies.

The article outlines several foundational findings. First, linguistic meaning is defined as the relation connecting an observable language form to non-linguistic communicative intent and the external world. Second, statistical models trained only on string sequences—exemplified by an isolated octopus learning to mimic human telegraph patterns—can replicate surface lexical correlations but cannot ground words to real-world objects or intentions. Third, empirical diagnostics demonstrate that models that appear to understand complex tasks often rely on statistical artifacts and syntactic heuristics rather than genuine reasoning, causing their performance to collapse when presented with adversarial tests. Fourth, evidence from child language acquisition confirms that humans require interactive joint attention and real-world grounding, rather than passive text exposure, to learn meaning.

These findings have significant implications for technology leadership and AI strategy. Overstating language model capabilities creates severe operational, safety, and reputational risks, particularly if organizations deploy text-only models into environments that demand reliable real-world reasoning. While these models capture useful statistical properties of language structure, they are fundamentally incomplete for human-analogous language understanding.

To build more robust systems, the article recommends pairing linguistic forms with grounding data, such as perceptual inputs or interactive task feedback. Organizations and researchers should also design benchmark tasks that minimize dataset-specific cues and implement rigorous diagnostic evaluations to verify whether systems perform well for valid reasons. While the article acknowledges that language models serve as effective components in broader pipelines and can manipulate surface form exceptionally well, leaders must maintain healthy skepticism regarding claims of genuine machine comprehension.

Bender et al (2020).pdf
Cover for Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data

Abstract

The success of the large neural language models on many NLP tasks is exciting. However, we find that these successes sometimes lead to hype in which these models are being described as "understanding" language or capturing "meaning". In this position paper, we argue that a system trained only on form has a priori no way to learn meaning. In keeping with the ACL 2020 theme of "Taking Stock of Where We've Been and Where We're Going", we argue that a clear understanding of the distinction between form and meaning will help guide the field towards better science around natural language understanding.

Table of Contents

  • 1 Introduction
  • 2 Large LMs: Hype and analysis
  • 3 What is meaning?
  • 3.1 Meaning and communicative intent
  • 3.2 Meaning and intelligence
  • 4 The octopus test
  • 5 More constrained thought experiments
  • 6 Human language acquisition
  • 7 Distributional semantics
  • 8 On climbing the right hills
  • 8.1 Top-down and bottom-up theory-building
  • 8.2 Hillclimbing diagnostics
  • 9 Some possible counterarguments
  • 10 Conclusion
  • References
  • A GPT-2 on fighting bears with sticks
  • B GPT-2 and arithmetic

Knowls

  1. Knowl 1 — Formal Distinction between Form, Conventional Meaning, and Communicative Intent

    definition

    In linguistic analysis and natural language processing, linguistic communication is decomposed into three distinct components:

    1. Form: Any observable physical or digital realization of language, such as printed text, digital byte sequences, acoustic speech waveforms, or manual/facial articulator movements.
    2. Communicative Intent: An agent's non-linguistic purpose or intended message i∈Ii \in I, which is grounded in the agent's physical environment, social context, abstract domains (e.g., databases, file systems), or mental states. Communicative intents are external to language.
    3. Conventional (Standing) Meaning: An abstract object s∈Ss \in S assigned by a linguistic system that represents the constant communicative potential of an expression across all possible contexts of use. Conventional meanings possess interpretations (such as truth conditions evaluated against a model of the world).

    Let EE denote the set of natural language expressions, II the set of communicative intents, and SS the set of conventional meanings. The linguistic system defines a conventional meaning relation:

    C⊆E×SC \subseteq E \times S

    which pairs expressions e∈Ee \in E with their standing meanings s∈Ss \in S. The overall meaning relation is:

    M⊆E×IM \subseteq E \times I

    which pairs expressions e∈Ee \in E with the communicative intents i∈Ii \in I they can evoke. Human language understanding is defined as the process of recovering ii given an observed form ee, mediated by the shared conventional system CC and context.

  2. Knowl 2 — Theoretical Impossibility of Learning Linguistic Meaning from Pure Form

    theoretical result

    A language model or system trained exclusively on linguistic form e∈Ee \in E via string prediction (character, word, or sentence prediction) cannot in principle learn the meaning relations M⊆E×IM \subseteq E \times I or C⊆E×SC \subseteq E \times S, where II denotes communicative intents and SS denotes conventional meanings.

    Because training data consisting solely of text contains only distributions of forms and lacks access to extra-linguistic referents, world models, perceptual stimuli, or communicative intents, there is no signal from which the mapping between form and non-linguistic intent can be deduced. Consequently, pure language modeling cannot solve the symbol grounding problem. Increasing the volume of training text or model parameters does not overcome this limitation, as scaling form data provides only a richer reflection of form distributions rather than access to the underlying grounded relations.

  3. Knowl 3 — The Octopus Test for Language Model Understanding

    theoretical result

    The Octopus Test is a thought experiment establishing that observing statistical regularities of linguistic forms alone is insufficient for language understanding:

    • Setup: Two interlocutors (AA and BB) on isolated islands communicate via an underwater telegraph wire. An underwater entity (OO, the octopus) intercepts the signal. OO has access only to the transmitted strings of text (linguistic forms) and no access to the terrestrial physical world or the speakers' environments.
    • Form Modeling: By observing message exchanges over time, OO builds an accurate statistical model of how BB responds to AA's messages and learns context-dependent lexical distributions.
    • Phatic/Chatbot Behavior: If OO intercepts the wire and impersonates BB, OO can successfully maintain superficial, phatic social dialogue where responses need only be internally coherent and do not require grounding in physical reality.
    • Failure on Grounded Tasks: When AA introduces a novel situation requiring action in the physical world (e.g., instructions for building a novel device like a coconut catapult or advice on using physical objects during a bear attack), OO fails. Because OO has never observed the physical referents of terms or their real-world causal dynamics, OO cannot determine what actions or objects the forms refer to, causing communication to break down.
  4. Knowl 4 — Active Listener Bias in Attributing Machine Understanding

    theoretical result

    Human communication fundamentally depends on the active interpretive role of the listener. When humans encounter linguistic forms e∈Ee \in E, they automatically assume the producer possesses communicative intent i∈Ii \in I and reconstruct ii using shared conventional meaning C⊆E×SC \subseteq E \times S alongside hypotheses about the speaker's mental state.

    When humans interact with artificial language models trained purely on form, this cognitive bias causes human interlocutors to project communicative intent and understanding onto the system's outputs (analogous to the ELIZA effect). The perceived coherence and meaningfulness of the generated text originate in the human listener's interpretive work rather than in any intrinsic comprehension or intent within the language model.

  5. Knowl 5 — Execution and Multimodal Thought Experiments on Semantic Grounding

    theoretical result

    Two constrained thought experiments demonstrate the necessity of external grounding signals for acquiring semantic relations:

    1. Execution Semantics in Code (Java): Consider a language model trained on all syntactically valid Java source code (e∈Ee \in E). If the training input consists solely of raw code without compilers, virtual machine bytecodes, execution traces, or input-output pairs (x,y)(x, y), the model cannot learn the semantic meaning relation J⊆E×IJ \subseteq E \times I (where ii is the executable function mapping inputs xx to outputs yy). The system cannot execute a novel program and predict its computational output from form alone.
    2. Multimodal Grounding (English and Vision): Consider a language model trained on a large corpus of text alongside an uncurated, unaligned collection of images. If no training signal links textual expressions to visual entities, the system cannot answer questions grounded in the images (e.g., counting objects or identifying referents depicted in a photo), because the text-only training distribution provides no mapping between words and visual percepts.
  6. Knowl 6 — Grounding and Intersubjectivity Constraints from Human Language Acquisition

    theoretical result

    Empirical psycholinguistic findings establish that human language acquisition does not occur via passive exposure to linguistic form alone (e.g., children exposed exclusively to television or radio in a foreign language fail to acquire it).

    Human language learning requires embodied, situated interaction mediated by:

    1. Joint Attention: Interactions where both the learner and caregiver simultaneously attend to the same physical object or event and are mutually aware of their shared focus. Caregiver labelling during joint attention strongly predicts vocabulary size and comprehension.
    2. Intersubjectivity: The capacity to infer another agent's attentional focus and communicative intent.

    Because biological language acquisition requires grounded physical and social interaction, artificial systems exposed only to ungrounded text streams cannot replicate human-like language acquisition.

  7. Knowl 7 — Distributional Semantics and Grounded Use

    theoretical result

    In distributional semantics, the slogan 'meaning is use' refers to the pragmatic deployment of language by situated agents in real-world contexts to achieve communicative intents i∈Ii \in I, rather than textual word co-occurrence distributions within corpora.

    While distributional vector spaces capture syntactic regularities and lexical similarity relations (which act as formal reflections of meaning), they do not connect words to external referents. Grounding distributional representations requires supplementing linguistic text with perceptual data (e.g., visual, auditory, or olfactory features) or interactive dialogue data with grounded feedback signals (e.g., task success, gaze tracking, or physiological cues).

  8. Knowl 8 — Methodological Diagnostics for Evaluating Natural Language Understanding

    model/method

    To prevent equating surface heuristic performance with true natural language understanding (NLU), research evaluating language models should follow five diagnostic practices:

    1. Top-Down Perspective: Evaluate whether incremental performance gains on specific tasks contribute to the end-goal of grounded understanding rather than merely climbing local optima on form-based datasets.
    2. Dataset Artifact Scrutiny: Audit benchmarks for statistical shortcuts and lexical artifacts (e.g., hypothesis negation or bag-of-words heuristics in natural language inference) where models exceed human baseline performance without semantic reasoning.
    3. Discrete Reasoning Benchmarks: Develop tasks that enforce semantic integration across contexts via operations such as arithmetic or discrete logic (e.g., DROP) rather than simple pattern matching.
    4. Cross-Task Evaluation: Evaluate representations across diverse, unified benchmark suites (e.g., SuperGLUE) to verify that learned semantic capabilities generalize across distinct tasks.
    5. Probing and Adversarial Testing: Use adversarial perturbations and structural probing tasks to identify the precise internal representations learned by models and verify whether success is achieved for the correct semantic reasons.

Coverage note — Specific anecdotal GPT-2 prompt completions regarding bear attacks and arithmetic from Appendices A and B were omitted as standalone knowls because they serve as illustrative qualitative examples rather than generalizable empirical contributions.

References

  1. 1.Yossi Adi, Einat Kermany, Yonatan Belinkov, Ofer Lavi, and Yoav Goldberg. 2017. Fine-grained analysis of sentence embeddings using auxiliary prediction tasks. In Proceedings of ICLR.
  2. 2.Dare A. Baldwin. 1995. Understanding the link between joint attention and language. In Chris Moore and Philip J. Dunham, editors, Joint Attention: Its Origins and Role in Development, pages 131–158. Psychology Press.
  3. 3.Andrei Barbu, David Mayo, Julian Alverio, William Luo, Christopher Wang, Dan Gutfreund, Josh Tenenbaum, and Boris Katz. 2019. ObjectNet: A large-scale bias-controlled dataset for pushing the limits of object recognition models. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d’Alche Buc, E. Fox, and R. Garnett, editors, Advances in Neural Information Processing Systems 32, pages 9453–9463. Curran Associates, Inc.
  4. 4.Marco Baroni, Raffaella Bernardi, Roberto Zamparelli, et al. 2014. Frege in space: A program for compositional distributional semantics. Linguistic Issues in Language Technology, 9(6):5–110.
  5. 5.Yonatan Bisk, Ari Holtzman, Jesse Thomason, Jacob Andreas, Yoshua Bengio, Joyce Chai, Mirella Lapata, Angeliki Lazaridou, Jonathan May, Aleksandr Nisnevich, Nicolas Pinto, and Joseph Turian. 2020. Experience grounds language. ArXiv preprint.
  6. 6.Ned Block. 1981. Psychologism and behaviorism. The Philosophical Review, 90(1):5–43.
  7. 7.Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015. A large annotated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 632–642, Lisbon, Portugal. Association for Computational Linguistics.
  8. 8.Rechele Brooks and Andrew N. Meltzoff. 2005. The development of gaze following and its relation to language. Developmental Science, 8(6):535–543.
  9. 9.Herbert H. Clark. 1996. Using Language. Cambridge University Press, Cambridge.
  10. 10.Stephen Clark. 2015. Vector space models of lexical meaning. In Shalom Lappin and Chris Fox, editors, Handbook of Contemporary Semantic Theory, second edition, pages 493–522. Wiley-Blackwell.
  11. 11.Ido Dagan, Oren Glickman, and Bernardo Magnini. 2006. The PASCAL Recognising Textual Entailment Challenge. In Machine Learning Challenges. Evaluating Predictive Uncertainty, Visual Object Classification, and Recognising Tectual Entailment, pages 177–190, Berlin, Heidelberg. Springer.
  12. 12.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  13. 13.Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019. DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 2368–2378, Minneapolis, Minnesota. Association for Computational Linguistics.
  14. 14.Guy Emerson. 2020. What are the goals of distributional semantics? In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Seattle, Washington. Association for Computational Linguistics.
  15. 15.Katrin Erk. 2016. What do you know about an alligator when you know the company it keeps? Semantics & Pragmatics, 9(17):1–63.
  16. 16.Allyson Ettinger, Ahmed Elgohary, Colin Phillips, and Philip Resnik. 2018. Assessing composition in sentence vector representations. In Proceedings of the 27th International Conference on Computational Linguistics, pages 1790–1801, Santa Fe, New Mexico, USA. Association for Computational Linguistics.
  17. 17.Yoav Goldberg. 2019. Assessing BERT’s syntactic abilities. ArXiv preprint.
  18. 18.H. Paul Grice. 1968. Utterer’s meaning, sentencemeaning, and word-meaning. Foundations of Language, 4(3):225–242.
  19. 19.Ivan Habernal, Henning Wachsmuth, Iryna Gurevych, and Benno Stein. 2018. The argument reasoning comprehension task: Identification and reconstruction of implicit warrants. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1930–1940, New Orleans, Louisiana. Association for Computational Linguistics.
  20. 20.William L. Hamilton, Jure Leskovec, and Dan Jurafsky. 2016. Diachronic word embeddings reveal statistical laws of semantic change. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1489–1501, Berlin, Germany. Association for Computational Linguistics.
  21. 21.Stevan Harnad. 1990. The symbol grounding problem. Physica D, 42:335–346.
  22. 22.Junxian He, Graham Neubig, and Taylor Berg-Kirkpatrick. 2018. Unsupervised learning of syntactic structure with invertible neural projections. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 1292–1302, Brussels, Belgium. Association for Computational Linguistics.
  23. 23.Benjamin Heinzerling. 2019. NLP’s Clever Hans moment has arrived. Blog post, accessed 12/4/2019.
  24. 24.Aurelie Herbelot. 2013. What is in a text, what isn’t, and what this has to do with lexical semantics. In Proceedings of the 10th International Conference on Computational Semantics (IWCS 2013) – Short Papers, pages 321–327, Potsdam, Germany. Association for Computational Linguistics.
  25. 25.Aurelie Herbelot, Eva von Redecker, and Johanna Muller. 2012. Distributional techniques for philosophical enquiry. In Proceedings of the 6th Workshop on Language Technology for Cultural Heritage, Social Sciences, and Humanities, pages 45–54, Avignon, France. Association for Computational Linguistics.
  26. 26.John Hewitt and Christopher D. Manning. 2019. A structural probe for finding syntax in word representations. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4129–4138, Minneapolis, Minnesota. Association for Computational Linguistics.
  27. 27.MD. Zakir Hossain, Ferdous Sohel, Mohd Fairuz Shiratuddin, and Hamid Laga. 2019. A comprehensive survey of deep learning for image captioning. ACM Comput. Surv., 51(6):118:1–118:36.
  28. 28.Ganesh Jawahar, Benoıt Sagot, and Djame Seddah. 2019. What does BERT learn about the structure of language? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3651–3657, Florence, Italy. Association for Computational Linguistics.
  29. 29.Nitish Shirish Keskar, Bryan McCann, Lav R. Varshney, Caiming Xiong, and Richard Socher. 2019. CTRL: A conditional transformer language model for controllable generation. ArXiv preprint.
  30. 30.Douwe Kiela, Luana Bulat, and Stephen Clark. 2015. Grounding semantics in olfactory perception. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pages 231–236, Beijing, China. Association for Computational Linguistics.
  31. 31.Douwe Kiela and Stephen Clark. 2015. Multi- and cross-modal semantics beyond vision: Grounding in auditory perception. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 2461–2470, Lisbon, Portugal. Association for Computational Linguistics.
  32. 32.Alexander Koller, Konstantina Garoufi, Maria Staudte, and Matthew Crocker. 2012. Enhancing referential success by tracking hearer gaze. In Proceedings of the 13th Annual Meeting of the Special Interest Group on Discourse and Dialogue, pages 30–39, Seoul, South Korea. Association for Computational Linguistics.
  33. 33.Patricia K. Kuhl. 2007. Is speech learning ‘gated’ by the social brain? Developmental Science, 10(1):110–120.
  34. 34.Guillaume Lample, Myle Ott, Alexis Conneau, Ludovic Denoyer, and Marc’Aurelio Ranzato. 2018. Phrase-based & neural unsupervised machine translation. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 5039–5049, Brussels, Belgium. Association for Computational Linguistics.
  35. 35.Chu-Cheng Lin, Waleed Ammar, Chris Dyer, and Lori Levin. 2015. Unsupervised POS induction with word embeddings. In Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1311–1316, Denver, Colorado. Association for Computational Linguistics.
  36. 36.Sally McConnell-Ginet. 1984. The origins of sexist language in discourse. Annals of the New York Academy of Sciences, 433(1):123–135.
  37. 37.Tom McCoy, Ellie Pavlick, and Tal Linzen. 2019. Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3428–3448, Florence, Italy. Association for Computational Linguistics.
  38. 38.Daniel McDuff and Ashish Kapoor. 2019. Visceral machines: Reinforcement learning with intrinsic physiological rewards. In International Conference on Learning Representations.
  39. 39.Tomas Mikolov, Wen-tau Yih, and Geoffrey Zweig. 2013. Linguistic regularities in continuous space word representations. In Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 746–751, Atlanta, Georgia. Association for Computational Linguistics.
  40. 40.Timothy Niven and Hung-Yu Kao. 2019. Probing neural network comprehension of natural language arguments. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4658–4664, Florence, Italy. Association for Computational Linguistics.
  41. 41.Stephan Oepen, Marco Kuhlmann, Yusuke Miyao, Daniel Zeman, Silvie Cinkova, Dan Flickinger, Jan Hajic, and Zdenka Uresova. 2015. SemEval 2015 Task 18: Broad-coverage semantic dependency parsing. In Proceedings of the 9th International Workshop on Semantic Evaluation (SemEval 2015).
  42. 42.Yasuhito Ohsugi, Itsumi Saito, Kyosuke Nishida, Hisako Asano, and Junji Tomita. 2019. A simple but effective method to incorporate multi-turn context with BERT for conversational machine comprehension. In Proceedings of the First Workshop on NLP for Conversational AI, pages 11–17, Florence, Italy. Association for Computational Linguistics.
  43. 43.Simon Ostermann, Michael Roth, and Manfred Pinkal. 2019. MCScript2.0: A machine comprehension corpus focused on script events and participants. In *Proceedings of the Eighth Joint Conference on Lexical and Computational Semantics (SEM 2019), pages 103–117, Minneapolis, Minnesota. Association for Computational Linguistics.
  44. 44.Fabio Petroni, Tim Rocktaschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019. Language models as knowledge bases? In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2463–2473, Hong Kong, China. Association for Computational Linguistics.
  45. 45.W. V. O. Quine. 1960. Word and Object. MIT Press.
  46. 46.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. Open AI Blog, accessed 12/4/2019.
  47. 47.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. SQuAD: 100,000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2383–2392, Austin, Texas. Association for Computational Linguistics.
  48. 48.Michael J. Reddy. 1979. The conduit metaphor: A case of frame conflict in our language about language. In A. Ortony, editor, Metaphor and Thought, pages 284–310. Cambridge University Press.
  49. 49.Herbert Rubenstein and John B. Goodenough. 1965. Contextual correlates of synonymy. Communications of the ACM, 8(10):627–633.
  50. 50.John Searle. 1980. Minds, brains, and programs. Behavioral and Brain Sciences, 3(3):417–457.
  51. 51.Catherine E Snow, Anjo Arlman-Rupp, Yvonne Hassing, Jan Jobse, Jan Joosten, and Jan Vorster. 1976. Mothers’ speech in three social classes. Journal of Psycholinguistic Research, 5(1):1–20.
  52. 52.Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R Thomas McCoy, Najoung Kim, Benjamin Van Durme, Sam Bowman, Dipanjan Das, and Ellie Pavlick. 2019. What do you learn from context? Probing for sentence structure in contextualized word representations. In International Conference on Learning Representations.
  53. 53.Michael Tomasello and Michael Jeffrey Farrar. 1986. Joint attention and early language. Child Development, 57(6):1454–1463.
  54. 54.Alan Turing. 1950. Computing machinery and intelligence. Mind, 59(236):433–460.
  55. 55.Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2019. SuperGLUE: A stickier benchmark for general-purpose language understanding systems. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d’Alche Buc, E. Fox, and R. Garnett, editors, Advances in Neural Information Processing Systems 32, pages 3266–3280. Curran Associates, Inc.
  56. 56.Alex Warstadt, Yu Cao, Ioana Grosu, Wei Peng, Hagen Blix, Yining Nie, Anna Alsop, Shikha Bordia, Haokun Liu, Alicia Parrish, Sheng-Fu Wang, Jason Phang, Anhad Mohananey, Phu Mon Htut, Paloma Jeretic, and Samuel R. Bowman. 2019. Investigating BERT’s knowledge of language: Five analysis methods with NPIs. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2877–2887, Hong Kong, China. Association for Computational Linguistics.
  57. 57.Joseph Weizenbaum. 1966. ELIZA—A computer program for the study of natural language communication between men and machines. Communications of the ACM, 9:36–45.
  58. 58.Jason Weston, Antoine Bordes, Sumit Chopra, Alexander M. Rush, Bart van Merrienboer, Armand Joulin, and Tomas Mikolov. 2016. Towards AI-complete question answering: A set of prerequisite toy tasks. In Proceedings of ICLR.
  59. 59.Adina Williams, Nikita Nangia, and Samuel Bowman. 2018. A broad-coverage challenge corpus for sentence understanding through inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1112–1122, New Orleans, Louisiana. Association for Computational Linguistics.
  60. 60.Ludwig Wittgenstein. 1953. Philosophical Investigations. MacMillan, New York.
  61. 61.Thomas Wolf. 2018. Learning meaning in natural language processing — The semantics mega-thread. Blog post, accessed 4/15/2020.
  62. 62.Dani Yogatama, Cyprien de Masson d’Autume, Jerome Connor, Tomas Kocisky, Mike Chrzanowski, Lingpeng Kong, Angeliki Lazaridou, Wang Ling, Lei Yu, Chris Dyer, and Phil Blunsom. 2019. Learning and evaluating general linguistic intelligence. ArXiv preprint.

Citation

MLA
Bender, E. M., and A. Koller. “Climbing Towards NLU: On Meaning, Form, and Understanding in the Age of Data”. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020, pp. 5185–98, https://doi.org/10.18653/v1/2020.acl-main.463.
APA
Bender, E. M., & Koller, A. (2020). Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 5185–5198. https://doi.org/10.18653/v1/2020.acl-main.463
Chicago
Bender, E. M., and A. Koller. 2020. “Climbing Towards NLU: On Meaning, Form, and Understanding in the Age of Data”. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 5185–98. https://doi.org/10.18653/v1/2020.acl-main.463.
Harvard
Bender, E.M. and Koller, A. (2020) “Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data”, Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, pp. 5185–5198. Available at: https://doi.org/10.18653/v1/2020.acl-main.463.
Vancouver
1. Bender EM, Koller A (2020) Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, pp 5185–5198

BibTeX

@inproceedings{bender-koller-2020-climbing,
    title = "Climbing towards {NLU}: {On} Meaning, Form, and Understanding in the Age of Data",
    author = "Bender, Emily M.  and
      Koller, Alexander",
    editor = "Jurafsky, Dan  and
      Chai, Joyce  and
      Schluter, Natalie  and
      Tetreault, Joel",
    booktitle = "Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics",
    month = jul,
    year = "2020",
    address = "Online",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2020.acl-main.463/",
    doi = "10.18653/v1/2020.acl-main.463",
    pages = "5185--5198"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/