The Symbol Grounding Problem

Stevan Harnad

article1990Physica D: Nonlinear Phenomena3,856 citations

Formulates the fundamental problem of how computational symbols acquire real-world meaning and proposes a framework that grounds symbolic reasoning bottom-up in sensory and categorical perception.

Listen

The Symbol Grounding Problem highlights a core limitation in computational language processing: symbolic systems and text-based models lack direct ties to real-world experience, leaving them unable to capture the full meaning of many words. This issue persists despite advances in deep learning, as models trained only on text derive patterns from word co-occurrence but miss perceptual and experiential content that humans use to understand language.

This paper reviews recent work on symbol grounding, argues that the problem has been misinterpreted as solvable through vision alone or larger text models, and reframes it around the distinction between concrete words (such as chair or red, which require grounding in perception) and abstract words (such as democracy, which can be learned from linguistic context). The authors conducted a small-scale experiment using logistic regression classifiers trained on image features from the CLIP model to test whether abstract category meanings could be built from concrete word representations. They also drew on child development research to outline how meaning might progress from concrete to abstract through spoken interaction and emotional cues.

The experiment achieved 88 percent accuracy in classifying nine test words into five abstract categories when using coefficients from concrete word classifiers as input features. The analysis shows that distributional models capture abstractness effectively but treat all words as ungrounded, while concrete words benefit from direct perceptual grounding. It further indicates that emotion functions as a parallel modality that scaffolds early learning and remains intertwined with abstract concepts, rather than serving as an optional add-on.

These results matter because current language-and-vision models still rely on symbolic object labels and text-heavy training, limiting their ability to handle the full range of concrete and intermediate terms that appear in real dialogue. Without addressing multiple modalities and the concrete-to-abstract progression, systems will continue to produce fluent but ungrounded output that fails on tasks requiring genuine world knowledge.

The authors recommend scaling the classifier approach into larger experiments that combine grounded coefficients with models such as BERT, using concreteness ratings to route words to the appropriate learning pathway. They also call for exploration of additional modalities beyond vision, including haptics and interoception, ideally within embodied, interactive settings. The work rests on a toy dataset with only five categories and nine test items, assumes word independence during training, and leaves open how positive examples for abstract categories should be selected in realistic data. The core reframing and experimental demonstration provide a credible direction, though broader validation is needed before deployment decisions.

Cover for The Symbol Grounding Problem

Abstract

How can the semantic interpretation of a formal symbol system be made intrinsic to the system, rather than just parasitic on the meanings in our heads? How can the meanings of the meaningless symbol tokens, manipulated solely on the basis of their (arbitrary) shapes, be grounded in anything but other meaningless symbols? The problem is analogous to trying to learn Chinese from a Chinese/Chinese dictionary alone. A candidate solution is sketched: Symbolic representations must be grounded bottom-up in nonsymbolic representations of two kinds: (1) "iconic representations," which are analogs of the proximal sensory projections of distal objects and events, and (2) "categorical representations," which are learned and innate feature-detectors that pick out the invariant features of object and event categories from their sensory projections. Elementary symbols are the names of these object and event categories, assigned on the basis of their (nonsymbolic) categorical representations. Higher-order (3) "symbolic representations," grounded in these elementary symbols, consist of symbol strings describing category membership relations (e.g., "An X is a Y that is Z").

Table of Contents

  • 1 Introduction
  • 2 The challenge of symbol grounding
  • 3 Reframing the problem: concreteness & abstractness
  • 3.1 A toy experiment: grounding into concrete words meanings
  • Train:
  • Test:
  • 4 Learning meaning from concrete to abstract
  • 4.1 The setting of spoken interaction
  • 4.2 Concrete-affect; abstract-emotion
  • 5 Open questions
  • 5.1 Modality questions
  • 5.2 Modeling questions
  • 5.3 Philosophical Questions
  • 6 Conclusion
  • References

Knowls

  1. Knowl 1 — Re-framing the Symbol Grounding Problem along the Concreteness-Abstractness Continuum

    theoretical result

    The Symbol Grounding Problem—the challenge that purely computational symbols and text-only distributional models lack access to real-world experience—is re-framed as a continuous spectrum between concreteness and abstractness rather than a binary dichotomy.

    • Concrete words (e.g., ball, red, chair) denote physical entities, shapes, or sensory qualities. Their meaning fundamentally requires perceptual symbol grounding into sensorimotor experience.
    • Abstract words (e.g., democracy, utopia) denote ideas that cannot be grounded directly into physical perception. Their meaning is defined through relations to other words and can be captured distributionally from lexical context.
    • Intermediate words (e.g., color, farm, animal) lie between these extremes. Their semantics can be acquired hierarchically by grounding them into the computational intensions (representations) of more concrete constituent concepts rather than directly into raw sensorimotor perception or pure lexical co-occurrence.
  2. Knowl 2 — Hierarchical Words-as-Classifiers Grounding Mechanism

    model/method

    A two-tier extension of the Words-as-Classifiers (WAC) framework enables more abstract words to be grounded into concrete word representations:

    1. Concrete Word Classifiers: For each concrete word wcw_c, a binary logistic regression classifier is trained on visual feature vectors extracted from images. The resulting vector of learned classifier coefficients θwc\mathbf{\theta}_{w_c} serves as a computational intension representing the perceptual properties associated with wcw_c.
    2. Abstract Category Classifiers: For a higher-level abstract category word waw_a (e.g., color, furniture), a binary classifier is trained using the coefficient vectors θwc\mathbf{\theta}_{w_c} of its constituent concrete words as positive training instances, and coefficient vectors from other categories as negative instances.
    3. Category Assignment for Unseen Concrete Words: Given a new concrete word wtestw_{\text{test}}, its visual classifier is trained to obtain coefficient vector θwtest\mathbf{\theta}_{w_{\text{test}}}, which is then provided as input to all category classifiers to predict the corresponding abstract category via maximum probability.
  3. Knowl 3 — Experimental Setup for Hierarchical WAC Classification of Unseen Concrete Words

    experimental setup

    The hierarchical Words-as-Classifiers (WAC) grounding mechanism was evaluated across five abstract categories using separate training and test vocabularies of concrete words:

    • Training Categories and Concrete Words:

      • color: red, blue, green, yellow, brown
      • animal: dog, cow, cat, mouse, bird
      • furniture: couch, chair, desk, bed
      • vehicle: car, van, truck, pickup, tractor
      • appliance: stove, oven, microwave
    • Test Words (9 unseen concrete words):

      • color: orange, purple
      • animal: horse, sheep
      • furniture: table, sofa
      • vehicle: taxi, jeep
      • appliance: mixer
    • Image Representation and Base Classifier Training: For each concrete word, the top 100 images from Google Image Search were passed through the CLIP model to obtain 512-dimensional feature vectors. A binary logistic regression classifier (C=0.25C = 0.25, max-iter=1000\text{max-iter} = 1000) was trained for each concrete word using positive images and randomly sampled negative images from other words (ratio 1:3 positive to negative).

    • Abstract Category Training & Testing: Abstract category classifiers were trained using the 512-dimensional coefficient vectors of the concrete training classifiers (with 3 negative category exemplars per positive). Evaluation was conducted by passing the 512-dimensional coefficient vectors of the 9 test words into the 5 category classifiers and assigning the category with the highest predicted probability.

  4. Knowl 4 — Hierarchical WAC Classification Accuracy on Unseen Concrete Concepts

    empirical result

    Evaluating the 9 unseen test words across the 5 trained abstract category classifiers achieved an overall classification accuracy of 88.89%88.89\% (8 out of 9 correct).

    • Correctly classified instances:
      • orange, purple \to color
      • horse, sheep \to animal
      • table, sofa \to furniture
      • taxi, jeep \to vehicle
    • Misclassified instance:
      • mixer \to misclassified as furniture (true category was appliance)

    This confirms that the learned parameter coefficients of visual classifiers carry consistent structural information that higher-level category models can leverage to organize concrete concepts without direct image supervision.

  5. Knowl 5 — Developmental Progression of Grounding through Spoken Interaction

    model/method

    A developmentally inspired computational model of language grounding structures learning across an interactive, spoken progression:

    • Communicative Grounding as Prerequisite: Symbol grounding is facilitated as a side effect of communicative grounding—the active, mutual coordination of meaning between agents in participatory spoken interaction.
    • Concrete-to-Abstract Progression: Language acquisition naturally shifts over developmental time from perceptual grounding to distributional learning. Early vocabulary is dominated by sensorimotor-grounded concrete words (abstract words make up only 10%\approx 10\% of vocabulary at age 4), providing the initial foundation upon which abstract concepts (25%\approx 25\% at age 5, >40%>40\% at age 12) are later learned distributionally from language use.
    • Participatory vs. Passive Observation: Passive observation of text or pre-labeled visual scenes is insufficient for human-like grounding; interactive dialogue provides the necessary social feedback and grounding mechanisms for concept acquisition.
  6. Knowl 6 — Role of Affect and Emotion in Multi-Level Symbol Grounding

    theoretical result

    Emotion and affect play distinct, dual roles in language grounding that parallel the concrete-to-abstract continuum:

    1. Pre-linguistic Communicative Scaffolding: Affect and vocal emotion provide immediate social signals that allow pre-verbal learners to infer speaker intent before understanding lexical content, facilitating early concrete word grounding.
    2. Abstract Concept Grounding via Interoception: Abstract concepts (particularly mental, social, and emotional states) are intrinsically grounded in interoception (internal bodily, physiological, and affective sensations) and motor systems. While sentiment can be statistically inferred from text distributions, computational models cannot capture holistic semantics without explicitly representing internal affective states alongside external sensorimotor perception.
  7. Knowl 7 — Limitations of the Hierarchical WAC Categorization Model

    limitation

    The proposed hierarchical Words-as-Classifiers (WAC) model has two main limitations:

    • Assumption of Word Independence: Concrete classifiers are trained independently, preventing the model from capturing cross-category semantic knowledge that distributional models easily capture (e.g., the model has no knowledge that appliances can have color).
    • Exemplar Selection for Non-Taxonomic Abstract Words: The framework requires an explicit set of concrete exemplar words to train a category classifier. While this works for taxonomic hypernyms (e.g., animal, furniture), it does not easily generalize to complex abstract concepts (e.g., democracy, utopia) whose definitions cannot be framed as groupings of concrete object classifiers.

Coverage note — None was omitted; all key contributions—the theoretical re-framing of symbol grounding along the concrete-abstract continuum, the hierarchical WAC model and its experimental evaluation, the developmental and emotional grounding proposals, and the stated limitations—have been captured as knowls.

References

  1. 1.L Alan Sroufe, Byron Egeland, Elizabeth A Carlson, and W Andrew Collins. 2009. The Development of the Person: The Minnesota Study of Risk and Adaptation from Birth to Adulthood. Guilford Press.
  2. 2.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karen Simonyan. 2022. Flamingo: a visual language model for Few-Shot learning.
  3. 3.Lawrence W Barsalou. 2008. Grounded cognition. Annu. Rev. Psychol., (59):617–645.
  4. 4.Emily M Bender and Alexander Koller. 2020. Climbing towards NLU: On meaning, form, and understanding in the age of data. In Association for Computational Linguistics, pages 5185–5198.
  5. 5.Yonatan Bisk, Ari Holtzman, Jesse Thomason, Jacob Andreas, Yoshua Bengio, Joyce Chai, Mirella Lapata, Angeliki Lazaridou, Jonathan May, Aleksandr Nisnevich, Nicolas Pinto, and Joseph Turian. 2020. Experience grounds language. arXiv.
  6. 6.Yuri Bizzoni and Simon Dobnik. Sky + fire = sunset exploring parallels between visually grounded metaphors and image classifiers.
  7. 7.Anna M Borghi, Laura Barca, Ferdinand Binkofski, Cristiano Castelfranchi, Giovanni Pezzulo, and Luca Tummolini. 2019. Words as social tools: Language, sociality and inner grounding in abstract concepts. Phys. Life Rev., 29:120–153.
  8. 8.Marc Brysbaert, Amy Beth Warriner, and Victor Kuperman. 2014. Concreteness ratings for 40 thousand generally known english word lemmas. Behav. Res. Methods, 46(3):904–911.
  9. 9.Angelo Cangelosi and Matthew Schlesinger. 2015. Developmental robotics: From babies to robots. MIT press.
  10. 10.Eve V Clark. 2013. First language acquisition. Cambridge University Press.
  11. 11.Herbert H Clark. 1996. Using Language. Cambridge University Press.
  12. 12.Pasquale A Della Rosa, Eleonora Catricalà, Gabriella Vigliocco, and Stefano F Cappa. 2010. Beyond the abstract—concrete dichotomy: Mode of acquisition, concreteness, imageability, familiarity, age of acquisition, context availability, and abstractness norms for a set of 417 italian words. Behav. Res. Methods, 42(4):1042–1048.
  13. 13.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. BERT: Pre-training of deep bidirectional transformers for language understanding.
  14. 14.Felix R Dreyer and Friedemann Pulvermüller. 2018. Abstract semantics in the motor system? – an event-related fMRI study on passive reading of semantic word categories carrying abstract emotional and mental meaning. Cortex, 100:52–70.
  15. 15.Leonardo Fernandino, Jia-Qing Tong, Lisa L Conant, Colin J Humphries, and Jeffrey R Binder. 2022. Decoding the information structure underlying the neural representation of concepts. Proc. Natl. Acad. Sci. U. S. A., 119(6).
  16. 16.Stevan Harnad. 1990. The symbol grounding problem. Physica D, 42(1-3):335–346.
  17. 17.Stevan Harnad. 2017. To cognize is to categorize: Cognition is categorization. In Handbook of Categorization in Cognitive Science, pages 21–54.
  18. 18.Lisa Anne Hendricks, John Mellor, Rosalia Schneider, Jean-Baptiste Alayrac, and Aida Nematzadeh. 2021. Decoupling the role of data, attention, and losses in multimodal transformers.
  19. 19.Aurélie Herbelot. 2013. What is in a text, what isn’t, and what this has to do with lexical semantics. In Proceedings of the 10th International Conference on Computational Semantics (IWCS 2013) – Short Papers, pages 321–327, Potsdam, Germany. Association for Computational Linguistics.
  20. 20.Felix Hill, Olivier Tieleman, Tamara von Glehn, Nathaniel Wong, Hamza Merzic, and Stephen Clark. 2020. Grounded language learning fast and slow.
  21. 21.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V Le, Yunhsuan Sung, Zhen Li, and Tom Duerig. 2021. Scaling up visual and Vision-Language representation learning with noisy text supervision.
  22. 22.Mark Johnson. 2008. The meaning of the body: Aesthetics of human understanding. University of Chicago Press.
  23. 23.Casey Kennington. 2021. Enriching language models with visually-grounded word vectors and the Lancaster sensorimotor norms. In Proceedings of the 25th Conference on Computational Natural Language Learning, pages 148–157, Online. Association for Computational Linguistics.
  24. 24.Douwe Kiela, Luana Bulat, and Stephen Clark. 2015. Grounding semantics in olfactory perception. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pages 231–236, Beijing, China. Association for Computational Linguistics.
  25. 25.Victor Kuperman, Hans Stadthagen-Gonzalez, and Marc Brysbaert. 2012. Age-of-acquisition ratings for 30,000 english words. Behav. Res. Methods, 44(4):978–990.
  26. 26.George Lakoff and Mark Johnson. 2008. Metaphors We Live By. University of Chicago Press.
  27. 27.Richard D Lane and Lynn Nadel. 2002. Cognitive Neuroscience of Emotion. Oxford University Press.
  28. 28.Staffan Larsson. 2018. Grounding as a Side-Effect of grounding. Top. Cogn. Sci.
  29. 29.John L Locke. 1995. The Child’s Path to Spoken Language. Harvard University Press.
  30. 30.Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019. ViLBERT: Pretraining Task-Agnostic visiolinguistic representations for Vision-and-Language tasks.
  31. 31.Dermot Lynott, Louise Connell, Marc Brysbaert, James Brand, and James Carney. 2019. The lancaster sensorimotor norms: multidimensional measures of perceptual and action strength for 40,000 english words. Behav. Res. Methods, pages 1–21.
  32. 32.Gary Marcus, Ernest Davis, and Scott Aaronson. 2022. A very preliminary analysis of DALL-E 2.
  33. 33.Claudia Mazzuca, Luisa Lugli, Mariagrazia Benassi, Roberto Nicoletti, and Anna M Borghi. 2018. Abstract, emotional and concrete concepts and the activation of mouth-hand effectors. PeerJ, 6:e5987.
  34. 34.David McNeill and Casey Kennington. 2020. Learning word groundings from humans facilitated by robot emotional displays. In Proceedings of the 21st Annual SIGdial Meeting on Discourse and Dialogue, Virtual. Association for Computational Linguistics.
  35. 35.Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2015. Efficient estimation of word representations in vector space. In International Conference on Learning Representations (ICLR).
  36. 36.Daniele Moro, Gerardo Caracas, David McNeill, and Casey Kennington. 2020. Semantics with feeling: Emotions for abstract embedding, affect for concrete grounding. In Proceedings of the 24th Workshop on the Semantics and Pragmatics of Dialogue - Full Papers, Virtual.
  37. 37.Daniele Moro and Casey Kennington. 2018. Multimodal visual and simulated muscle activations for grounded semantics of hand-related descriptions. In Proceedings of the 22nd Workshop onthe Semantics and Pragmatics of Dialogue.
  38. 38.Letitia Parcalabescu, Michele Cafagna, Lilitta Muradjan, Anette Frank, Iacer Calixto, and Albert Gatt. 2021. VALSE: A Task-Independent benchmark for vision and language models centered on linguistic phenomena.
  39. 39.Letitia Parcalabescu, Albert Gatt, Anette Frank, and Iacer Calixto. 2020. Seeing past words: Testing the cross-modal capabilities of pretrained V&L models on counting tasks.
  40. 40.Marta Ponari, Courtenay Frazier Norbury, and Gabriella Vigliocco. 2018. Acquisition of abstract concepts is influenced by emotional valence. Dev. Sci., 21(2).
  41. 41.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. 2022. Hierarchical Text-Conditional image generation with CLIP latents.
  42. 42.Anna Rogers, Olga Kovaleva, and Anna Rumshisky. 2020. A primer in BERTology: What we know about how BERT works. arXiv.
  43. 43.Jacqueline Sachs, Barbara Bard, and Marie L Johnson. 1981. Language learning with restricted input: Case studies of two hearing children of deaf parents. Appl. Psycholinguist., 2(01):33–54.
  44. 44.David Schlangen, Sina Zarriess, and Casey Kennington. 2016. Resolving references to objects in photographs using the Words-As-Classifiers model. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, pages 1213–1223.
  45. 45.John R Searle. 1980. Minds, brains, and programs. Behav. Brain Sci., 3(03):417.
  46. 46.Linda Smith and Michael Gasser. 2005. The development of embodied cognition: Six lessons from babies. Artif. Life, (11):13–29.
  47. 47.Mariarosaria Taddeo and Luciano Floridi. 2005. Solving the symbol grounding problem: a critical review of fifteen years of research. J. Exp. Theor. Artif. Intell., 17(4):419–445.
  48. 48.Jesse Thomason, Jivko Sinapov, Raymond J Mooney, and Peter Stone. 2018. Guiding exploratory behaviors for multi-modal grounding of linguistic descriptions. In 32nd AAAI Conference on Artificial Intelligence, AAAI 2018, pages 5520–5527. AAAI.
  49. 49.Caterina Villani, Luisa Lugli, Marco Tullio Liuzza, Roberto Nicoletti, and Anna M Borghi. 2021. Sensorimotor and interoceptive dimensions in concrete and abstract concepts. J. Mem. Lang., 116:104173.
  50. 50.Philippe Vincent-Lamarre, Alexandre Blondin Massé, Marcos Lopes, Mélanie Lord, Odile Marcotte, and Stevan Harnad. 2016. The latent structure of dictionaries. Top. Cogn. Sci., 8(3):625–659.
  51. 51.L Wittgenstein. 2010. Philosophische untersuchungen. In Sprachwissenschaft, pages 105–111. De Gruyter.
  52. 52.Benfeng Xu, Licheng Zhang, Zhendong Mao, Quan Wang, Hongtao Xie, and Yongdong Zhang. 2020. Curriculum learning for natural language understanding. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6095–6104, Online. Association for Computational Linguistics.

Citation

MLA
Harnad, S. “The Symbol Grounding Problem”. Physica D: Nonlinear Phenomena, vol. 42, nos. 1-3, 1990, pp. 335–46, https://doi.org/10.1016/0167-2789(90)90087-6.
APA
Harnad, S. (1990). The symbol grounding problem. Physica D: Nonlinear Phenomena, 42(1-3), 335–346. https://doi.org/10.1016/0167-2789(90)90087-6
Chicago
Harnad, S. 1990. “The Symbol Grounding Problem”. Physica D: Nonlinear Phenomena 42 (1-3): 335–46. https://doi.org/10.1016/0167-2789(90)90087-6.
Harvard
Harnad, S. (1990) “The symbol grounding problem”, Physica D: Nonlinear Phenomena, 42(1-3), pp. 335–346. Available at: https://doi.org/10.1016/0167-2789(90)90087-6.
Vancouver
1. Harnad S (1990) The symbol grounding problem. Physica D: Nonlinear Phenomena 42:335–346

BibTeX

@article{Harnad_1990, title={The symbol grounding problem}, volume={42}, ISSN={0167-2789}, url={http://dx.doi.org/10.1016/0167-2789(90)90087-6}, DOI={10.1016/0167-2789(90)90087-6}, number={1-3}, journal={Physica D: Nonlinear Phenomena}, publisher={Elsevier BV}, author={Harnad, Stevan}, year={1990}, month=June, pages={335–346} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: Authors