Analyzing Encoded Concepts in Transformer Language Models

Hassan SajjadNadir DurraniFahim DalviFiroj AlamAbdul Rafae KhanJia Xu

article2022NAACL66 citations

Proposes ConceptX, an unsupervised framework that clusters latent contextual representations and aligns them with human-defined linguistic categories to explain how transformer language models organize knowledge across layers without relying on probing classifiers.

Listen

Deep neural network language models have become foundational across modern natural language processing applications, yet their "black-box" nature creates significant risks regarding reliability, fairness, and governance. Understanding how these models internally organize linguistic and conceptual information is critical for ensuring safe deployment and control. The article addresses this challenge by introducing ConceptX, a framework designed to analyze and interpret the latent, context-aware representations learned across various layers of pre-trained transformer models.

The main objective of the article is to evaluate how internal model representations correspond to human-defined linguistic concepts and to establish where and how this knowledge is structured across network layers. To achieve this, the authors used an unsupervised agglomerative hierarchical clustering method to group contextualized word representations into 1,000 clusters per layer across seven prominent 12-layer transformer architectures (including BERT, RoBERTa, XLNet, ALBERT, and multilingual variants). These learned clusters—termed encoded concepts—were then evaluated against a broad suite of human-defined linguistic categories, spanning lexical units, morphology, syntax, semantics, and psycholinguistic ontologies, using a strict 90% alignment threshold on a standardized news dataset.

The evaluation revealed several key findings regarding how language models internally process text. First, between 43.6% and 72.4% of all learned clusters align directly with standard human-defined concepts, with multilingual models (such as XLM-RoBERTa at 72.4%) exhibiting significantly higher concept alignment than monolingual models (such as BERT-cased at 47.2%). Second, the internal architecture organizes information hierarchically: lower layers are dominated by shallow lexical patterns (like subword ngrams and affixes) and static semantic ontologies, whereas middle and higher layers (layers 8–10) primarily capture core linguistic properties such as morphology, parts-of-speech, and syntax. Third, morphological properties align much more strongly across all models (up to 26% fine-grained and 53% coarse alignment) than semantic (up to 16%) or complex syntactic categories (up to 14%). Finally, 28% to 56% of internal clusters do not match single traditional categories because they represent multi-faceted or compositional concepts—such as grouping words by both specific verb tense and semantic meaning simultaneously—which can be explained when combining multiple linguistic labels.

These findings demonstrate that language models naturally learn structured, hierarchical abstractions of language without explicit supervision, though their internal logic often blurs the boundaries of traditional human grammar rules. For practitioners and decision-makers, this means model performance does not always correlate directly with adherence to formal linguistic concepts; for instance, higher GLUE benchmark performance in some architectures does not necessarily mean higher alignment with human categories. Furthermore, training complexity (such as handling multiple languages or lacking capitalization cues) forces models to encode richer internal linguistic structures.

To build upon these insights, organizations and researchers deploying large transformer models should utilize unsupervised concept clustering as an interpretability check to audit model internals. Where traditional single-label taxonomies fail to explain model behaviors, teams should adopt compositional concept evaluations or human-in-the-loop assessments to capture multi-faceted internal representations. Future work should conduct controlled experiments to isolate how specific architectural parameters, pre-training objectives, and vocabulary tokenization schemes influence concept formation, as well as extend evaluations to larger, modern generative architectures.

arXiv: 2206.13289
Cover for Analyzing Encoded Concepts in Transformer Language Models

Abstract

We propose a novel framework ConceptX, to analyze how latent concepts are encoded in representations learned within pre-trained language models. It uses clustering to discover the encoded concepts and explains them by aligning with a large set of human-defined concepts. Our analysis on seven transformer language models reveal interesting insights: i) the latent space within the learned representations overlap with different linguistic concepts to a varying degree, ii) the lower layers in the model are dominated by lexical concepts (e.g., affixation), whereas the core-linguistic concepts (e.g., morphological or syntactic relations) are better represented in the middle and higher layers, iii) some encoded concepts are multi-faceted and cannot be adequately explained using the existing human-defined concepts.¹

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methodology
  • 3.1 Clustering
  • 3.2 Alignment
  • 4 Experimental Setup
  • 4.1 Dataset
  • 4.2 Pre-trained Models
  • 4.3 Clustering and Alignment
  • 4.4 Human-defined concepts
  • 5 Analysis
  • 5.1 Overall Alignment
  • 5.2 Layer-wise Alignment
  • 5.3 Unaligned Concepts
  • 5.4 Generalization of Results
  • 6 Conclusion
  • References
  • Appendix
  • A Human-defined concept labels
  • A.1 Lexical Concepts
  • A.2 Morphology and Semantics
  • A.3 Syntactic
  • A.4 Linguistic Ontologies
  • B BERT-based Sequence Tagger
  • C Clustering details
  • C.1 Selection of the number of Clusters
  • D Coarse vs. Fine-grained Categories
  • D.1 Coarse vs. Fine-grained Categories
  • D.2 Coarse POS and SEM labels
  • D.3 Results
  • E Compositional Coverage
  • F Robustness of Methodology across Datasets and Settings
  • G Layer-wise results

Knowls

  1. Knowl 1 — ConceptX Interpretability Framework

    model/method

    ConceptX is an unsupervised interpretability framework designed to analyze how latent linguistic and non-linguistic concepts are structured within the representations of deep pre-trained transformer language models.

    Given a neural network MM with LL encoder layers and HH hidden units per layer, an input sequence of words w1,w2,…,wMw_1, w_2, \dots, w_M is fed forward to produce hidden activation vectors y⃗ l(wi)∈RH\vec{y}^{\,l}(w_i) \in \mathbb{R}^H for each token wiw_i at layer ll. ConceptX operates in three stages:

    1. Representation Extraction: Contextualized representations y⃗ l(wi)\vec{y}^{\,l}(w_i) are extracted for a large vocabulary sampled across multiple sentence contexts.
    2. Unsupervised Clustering: At each layer ll, the representations are clustered into KK discrete clusters. Each cluster of context-aware latent representations constitutes an encoded concept cc.
    3. Concept Alignment: Each encoded concept cc is compared against a broad suite of pre-annotated human-defined concepts zz via a quantitative alignment metric to provide human-interpretable explanations of the model's latent spaces across layers.
  2. Knowl 2 — Theta-Alignment Metric for Encoded and Human-Defined Concepts

    equation

    Let zz denote a human-defined concept category (e.g., a specific part-of-speech tag or semantic role), and let its inverse mapping z−1(z)={w1,w2,…,wJ}z^{-1}(z) = \{w_1, w_2, \dots, w_J\} be the set of unique word instances annotated with zz, where J=∣z−1(z)∣J = |z^{-1}(z)| is the total count of words bearing concept zz. Let cc be an encoded concept representing a cluster of contextualized word embeddings, with its instance set denoted by c−1(c)={w1,w2,…,wI}c^{-1}(c) = \{w_1, w_2, \dots, w_I\}, where I=∣c−1(c)∣I = |c^{-1}(c)|.

    An encoded concept cc is defined to be θ\theta-aligned (denoted Λθ(z,c)=1\Lambda_\theta(z, c) = 1) with the human-defined concept zz if the fraction of instances of zz that belong to cluster cc meets or exceeds the threshold θ∈[0,1]\theta \in [0, 1]:

    Λθ(z,c)={1if 1J∑w′∈z−1(z)∑w∈c−1(c)δ(w,w′)≥θ0otherwise\Lambda_\theta(z, c) = \begin{cases} 1 & \text{if } \frac{1}{J} \sum_{w' \in z^{-1}(z)} \sum_{w \in c^{-1}(c)} \delta(w, w') \ge \theta \\ 0 & \text{otherwise} \end{cases}

    where δ(w,w′)\delta(w, w') is the Kronecker delta function:

    δ(w,w′)={1if w=w′0otherwise\delta(w, w') = \begin{cases} 1 & \text{if } w = w' \\ 0 & \text{otherwise} \end{cases}

    To compute the network-wide alignment score for a given concept, Λθ(z,c)\Lambda_\theta(z, c) is computed for each encoder layer ll and averaged across all layers. In standard evaluations, the threshold is set to θ=0.90\theta = 0.90 (requiring at least a 90% overlap).

  3. Knowl 3 — Agglomerative Hierarchical Clustering of Latent Representations

    algorithm

    ConceptX groups contextualized word representations at each layer using agglomerative hierarchical clustering with Ward's minimum variance criterion, which minimizes intra-cluster variance at each step based on squared Euclidean distance.

    Input: Set of contextualized representations {y⃗ l(wi)}i=1N\{\vec{y}^{\,l}(w_i)\}_{i=1}^N for input words
    Parameter: Target number of clusters KK
    for each word wiw_i do
        assign wiw_i to singleton cluster cic_i
    end for
    while current number of clusters >K> K do
        for each pair of distinct clusters (ci,ci′)(c_i, c_{i'}) do
            compute merge cost di,i′=ΔVar(ci∪ci′)d_{i, i'} = \Delta \text{Var}(c_i \cup c_{i'})
        end for
        select pair (cj,cj′)(c_j, c_{j'}) that minimizes dj,j′d_{j, j'}
        merge clusters cjc_j and cj′c_{j'} into a single cluster
    end while
    Output: Set of KK clusters (encoded concepts)

    The algorithm terminates when exactly KK clusters remain. Distance between vector representations is evaluated using squared Euclidean distance. For transformer representations of dimension H=768H=768, K=1000K=1000 provides an empirically stable balance between under-clustering and over-clustering.

  4. Knowl 4 — Layer-Wise Hierarchy and Evolution of Encoded Concepts

    empirical result

    Analyzing concept alignment layer-by-layer across 12-layer transformer language models (including BERT-cased, BERT-uncased, RoBERTa, XLNet, mBERT, and XLM-R) reveals a consistent hierarchical progression:

    1. Lower layers (layers 0–4): Dominated by shallow lexical concepts (e.g., n-grams and subword suffixes generated as artifacts of BPE/WordPiece segmentation) and static, context-independent ontologies (WordNet supersenses and LIWC psycholinguistic categories).
    2. Middle to higher layers (layers 7–10): Core linguistic properties requiring contextual integration (Part-of-Speech tags, chunking phrases, CCG super-tags, and semantic tags) reach their peak alignment, indicating that representations evolve from lexical/semantic clusters into linguistic manifolds.
    3. Final layers (layers 11–12): Alignment with general linguistic concepts drops significantly as representations specialize toward the pre-training objective. In BERT, core-linguistic information is retained deeper into the final layers than in XLNet or mBERT.
    4. Cross-layer parameter sharing exception (ALBERT): Due to weight tying across layers, ALBERT exhibits minimal layer-wise variation; its concept alignment remains flat and preserves all human-defined concepts up to the final layer.
  5. Knowl 5 — Concept Alignment Coverage Across Monolingual and Multilingual Models

    data/table

    The overall alignment coverage measures the percentage of discovered latent clusters (encoded concepts) across all layers that align (theta≥0.90\\theta \ge 0.90) with at least one human-defined concept category under fine-grained labeling.

    BERT-c BERT-uc mBERT XLM-R RoBERTa ALBERT XLNet
    Overall alignment 47.2% 50.4% 66.0% 72.4% 50.1% 51.6% 43.6%

    Multilingual models (mBERT: 66.0%, XLM-R: 72.4%) demonstrate substantially higher alignment with human-defined concepts than monolingual models (43.6%–51.6%). This is driven by their aggressive subword segmentation (which produces high n-gram/suffix cluster matches in lower layers) and their need to generalize across diverse linguistic structures, which forces stronger encoding of core morphological (POS) properties. Additionally, BERT-uncased achieves higher core-linguistic alignment than BERT-cased because the lack of casing cues requires the model to rely more heavily on contextual linguistic structure.

  6. Knowl 6 — Category-Wise Distribution of Encoded Concepts

    empirical result

    Comparing alignment across different categories of human-defined concepts demonstrates how transformers prioritize linguistic knowledge:

    • Morphology vs. Semantics: Part-of-Speech (POS) tags show the highest alignment among core-linguistic concepts across all models (13% in BERT-c to 26% in mBERT). Semantic tags (SEM) achieve lower match rates (7%–16%), showing a relative structural preference for morphological over semantic ontologies.
    • Syntactic Structures: Syntactic chunking (4%–7%) and CCG super-tags (6%–14%) show consistently low alignment, suggesting that transformer representations do not explicitly mirror classical syntactic hierarchies.
    • Linguistic Ontologies: Static semantic ontologies (WordNet) represent 11%–21% of clusters, making them the second most aligned category after POS. LIWC psycholinguistic categories match between 5% and 16% of clusters, particularly in lower layers.
    • Subwords vs. Affixes: N-gram matching (10%–48%) exceeds affix matching (1%–25%), indicating that models encode statistical segmentation artifacts more strongly than pure morphological affixes.
  7. Knowl 7 — Coarse-Grained Labeling and Concept Compositionality

    data/table

    Many latent clusters unaligned with single fine-grained tags are compositional concepts that combine multiple related linguistic categories (e.g., combining semantic geopolitical entities SEM:GPE with adjectives POS:JJ, or clustering multiple verb inflections VB, VBD, VBG, VBN into a general verb concept).

    When fine-grained POS (48 tags) and SEM (73 tags) are collapsed into coarse categories (27 POS and 15 SEM tags), concept alignment rates approximately double, and total coverage increases substantially across all architectures:

    Model POS (Fine) POS (Coarse) SEM (Fine) SEM (Coarse)
    BERT-cased 13% 42% 7% 15%
    BERT-uncased 16% 43% 9% 18%
    mBERT 26% 53% 16% 26%
    XLM-RoBERTa 24% 47% 11% 21%
    RoBERTa 18% 43% 10% 20%
    ALBERT 17% 42% 9% 17%
    XLNet 17% 39% 10% 18%

    Overall alignment coverage with coarse POS and SEM labels reaches 61.5% (BERT-c), 63.6% (BERT-uc), 77.7% (mBERT), 81.0% (XLM-R), 62.9% (RoBERTa), 64.0% (ALBERT), and 55.3% (XLNet). Furthermore, evaluating compositionality by allowing clusters to be explained by up to NN morphological concepts reveals that allowing N≤3N \le 3 concepts accounts for almost all clustered representations.

  8. Knowl 8 — Suite of Human-Defined Concepts for ConceptX Alignment

    experimental setup

    ConceptX benchmarks latent clusters against four groups of human-defined concepts across language abstraction levels:

    1. Lexical Concepts: Word n-grams, morphological affixes (prefixes/suffixes), casing patterns (title case, uppercase), and sentence positional tokens (first/last word).
    2. Morphology and Semantics: 48 Penn Treebank POS tags (Marcus et al., 1993) and 73 Parallel Meaning Bank SEM tags grouped into 13 meta-tags (Abzianidze et al., 2017).
    3. Syntactic Concepts: 22 CoNLL-2000 chunking tags in IOB format (Tjong Kim Sang and Buchholz, 2000) and 1,272 CCGbank Combinatory Categorial Grammar super-tags (Hockenmaier, 2006).
    4. Linguistic Ontologies: WordNet supersenses (26 noun, 2 adjective, 1 adverb lexicographic senses) and LIWC psycholinguistic category dictionaries (Pennebaker et al., 2001).

    Gold-standard splits for POS, SEM, Chunking, and CCG tagging are used to train BERT-based sequence tagger classifiers (achieving test F1 scores of 96.69%, 96.22%, 96.91%, and 94.90% respectively) to auto-annotate evaluation text.

  9. Knowl 9 — ConceptX Experimental Configuration and Clustering Stability

    experimental setup

    The evaluation dataset is constructed by sampling 250k sentences (~5M tokens) from the WMT News 2018 corpus. To avoid high-frequency bias and ensure memory tractability without applying lossy dimensionality reduction (e.g., PCA), words with frequency <10< 10 are removed, and a maximum of 10 contexts per word type are retained, resulting in 25,000 word types (250k total token representations per layer).

    Cluster hyperparameter selection was analyzed using Elbow and Silhouette metrics:

    • The Elbow curve showed continuous distortion reduction without a sharp bend, causing over-clustering (many clusters with <5<5 words yielding spuriously high alignment matches).
    • Silhouette optimization produced severe under-clustering (K=10K=10).

    Setting K=1000K=1000 empirically balances cluster granularity. Stability checks with K=600K=600 and K∈[200,1600]K \in [200, 1600] as well as resampling on a low-frequency vocabulary split (2–10 occurrences) confirmed that the qualitative and layer-wise alignment findings remain consistent.

Coverage note — None was omitted; all key methodology, definitions, equations, algorithms, empirical findings, and ablation/robustness data from the main paper and appendices were included.

References

  1. 1.Lasha Abzianidze, Johannes Bjerva, Kilian Evang, Hessel Haagsma, Rik van Noord, Pierre Ludmann, Duc-Duy Nguyen, and Johan Bos. 2017. The parallel meaning bank: Towards a multilingual corpus of translations annotated with compositional meaning representations. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics, EACL '17, pages 242–247, Valencia, Spain.
  2. 2.Yossi Adi, Einat Kermany, Yonatan Belinkov, Ofer Lavi, and Yoav Goldberg. 2016. Fine-grained Analysis of Sentence Embeddings Using Auxiliary Prediction Tasks. arXiv preprint arXiv:1608.04207.
  3. 3.Anthony Bau, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James Glass. 2019. Identifying and controlling important neurons in neural machine translation. In International Conference on Learning Representations.
  4. 4.Yonatan Belinkov. 2021. Probing classifiers: Promises, shortcomings, and alternatives. CoRR, abs/2102.12452.
  5. 5.Yonatan Belinkov, Nadir Durrani, Fahim Dalvi, Hassan Sajjad, and James Glass. 2017a. What do Neural Machine Translation Models Learn about Morphology? In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (ACL), Vancouver. Association for Computational Linguistics.
  6. 6.Yonatan Belinkov, Nadir Durrani, Fahim Dalvi, Hassan Sajjad, and James Glass. 2020. On the linguistic representational power of neural machine translation models. Computational Linguistics, 45(1):1–57.
  7. 7.Yonatan Belinkov, Lluís Màrquez, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James Glass. 2017b. Evaluating Layers of Representation in Neural Machine Translation on Part-of-Speech and Semantic Tagging Tasks. In Proceedings of the 8th International Joint Conference on Natural Language Processing (IJCNLP).
  8. 8.Terra Blevins, Omer Levy, and Luke Zettlemoyer. 2018. Deep RNNs encode soft hierarchical syntax. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 14–19, Melbourne, Australia. Association for Computational Linguistics.
  9. 9.Zhi Chen, Yijie Bei, and Cynthia Rudin. 2020. Concept whitening for interpretable image recognition. Nature Machine Intelligence, 2(12):772–782.
  10. 10.Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019. What does BERT look at? an analysis of BERT's attention. In Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 276–286, Florence, Italy. Association for Computational Linguistics.
  11. 11.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8440–8451. Association for Computational Linguistics.
  12. 12.Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni. 2018. What you can cram into a single vector: Probing sentence embeddings for linguistic properties. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (ACL).
  13. 13.Fahim Dalvi, Nadir Durrani, Hassan Sajjad, Yonatan Belinkov, D. Anthony Bau, and James Glass. 2019a. What is one grain of sand in the desert? analyzing individual neurons in deep nlp models. In Proceedings of the Thirty-Third AAAI Conference on Artificial Intelligence (AAAI, Oral presentation).
  14. 14.Fahim Dalvi, Abdul Rafae Khan, Firoj Alam, Nadir Durrani, Jia Xu, and Hassan Sajjad. 2022. Discovering latent concepts learned in BERT. In International Conference on Learning Representations.
  15. 15.Fahim Dalvi, Avery Nortonsmith, D. Anthony Bau, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, and James Glass. 2019b. Neurox: A toolkit for analyzing individual neurons in neural networks. In AAAI Conference on Artificial Intelligence (AAAI).
  16. 16.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), Minneapolis, Minnesota. Association for Computational Linguistics.
  17. 17.Nadir Durrani, Fahim Dalvi, Hassan Sajjad, Yonatan Belinkov, and Preslav Nakov. 2019. One size does not fit all: Comparing NMT representations of different granularities. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 1504–1516, Minneapolis, Minnesota. Association for Computational Linguistics.
  18. 18.Nadir Durrani, Hassan Sajjad, and Fahim Dalvi. 2021. How transfer learning impacts linguistic knowledge in deep NLP models? In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 4947–4957, Online. Association for Computational Linguistics.
  19. 19.Nadir Durrani, Hassan Sajjad, Fahim Dalvi, and Yonatan Belinkov. 2020. Analyzing individual neurons in pre-trained language models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4865–4880, Online. Association for Computational Linguistics.
  20. 20.Amirata Ghorbani, James Wexler, James Y Zou, and Been Kim. 2019. Towards automatic concept-based explanations. Advances in Neural Information Processing Systems, 32:9277–9286.
  21. 21.K Chidananda Gowda and G Krishna. 1978. Agglomerative clustering using the concept of mutual nearest neighbourhood. Pattern recognition, 10(2):105–112.
  22. 22.Kristina Gulordava, Piotr Bojanowski, Edouard Grave, Tal Linzen, and Marco Baroni. 2018. Colorless green recurrent networks dream hierarchically. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1195–1205, New Orleans, Louisiana. Association for Computational Linguistics.
  23. 23.Julia Hockenmaier. 2006. Creating a CCGbank and a wide-coverage CCG lexicon for German. In Proceedings of the 21st International Conference on Computational Linguistics and 44th Annual Meeting of the Association for Computational Linguistics, ACL '06, pages 505–512, Sydney, Australia.
  24. 24.Dieuwke Hupkes, Sara Veldhoen, and Willem Zuidema. 2018. Visualisation and 'diagnostic classifiers' reveal how recurrent and recursive neural networks process hierarchical structure.
  25. 25.Akos Kádár, Grzegorz Chrupała, and Afra Alishahi. 2017. Representation of linguistic form and function in recurrent neural networks. Computational Linguistics, 43(4):761–780.
  26. 26.Andrej Karpathy, Justin Johnson, and Li Fei-Fei. 2015. Visualizing and understanding recurrent networks. arXiv preprint arXiv:1506.02078.
  27. 27.Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, et al. 2018. Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav). In International conference on machine learning, pages 2668–2677. PMLR.
  28. 28.Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019. Albert: A lite bert for self-supervised learning of language representations. ArXiv:1909.11942.
  29. 29.Jiwei Li, Xinlei Chen, Eduard Hovy, and Dan Jurafsky. 2016. Visualizing and understanding neural models in NLP. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 681–691, San Diego, California. Association for Computational Linguistics.
  30. 30.Tal Linzen, Emmanuel Dupoux, and Yoav Goldberg. 2016. Assessing the ability of LSTMs to learn syntax-sensitive dependencies. Transactions of the Association for Computational Linguistics, 4:521– 535.
  31. 31.Nelson F. Liu, Matt Gardner, Yonatan Belinkov, Matthew E. Peters, and Noah A. Smith. 2019a. Linguistic knowledge and transferability of contextual representations. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 1073–1094, Minneapolis, Minnesota. Association for Computational Linguistics.
  32. 32.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019b. RoBERTa: A robustly optimized BERT pretraining approach. ArXiv:1907.11692.
  33. 33.Jonathan Mamou, Hang Le, Miguel Del Rio, Cory Stephenson, Hanlin Tang, Yoon Kim, and Sueyeon Chung. 2020. Emergence of separable manifolds in deep language representations. In International Conference on Machine Learning, pages 6713–6723. PMLR.
  34. 34.Mitchell P. Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz. 1993. Building a large annotated corpus of English: The Penn Treebank. Computational Linguistics, 19(2):313–330.
  35. 35.Rebecca Marvin and Tal Linzen. 2018. Targeted syntactic evaluation of language models. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 1192–1202, Brussels, Belgium. Association for Computational Linguistics.
  36. 36.Julian Michael, Jan A. Botha, and Ian Tenney. 2020. Asking without telling: Exploring latent ontologies in contextual representations. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6792–6812, Online. Association for Computational Linguistics.
  37. 37.George A Miller. 1995. Wordnet: a lexical database for english. Communications of the ACM, 38(11):39–41.
  38. 38.Jesse Mu and Jacob Andreas. 2020. Compositional explanations of neurons. CoRR, abs/2006.14032.
  39. 39.James W Pennebaker, Martha E Francis, and Roger J Booth. 2001. Linguistic inquiry and word count: Liwc 2001. Mahway: Lawrence Erlbaum Associates, 71(2001):2001.
  40. 40.Peng Qian, Xipeng Qiu, and Xuanjing Huang. 2016. Investigating Language Universal and Specific Properties in Word Embeddings. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1478–1488, Berlin, Germany. Association for Computational Linguistics.
  41. 41.Emily Reif, Ann Yuan, Martin Wattenberg, Fernanda B Viegas, Andy Coenen, Adam Pearce, and Been Kim. 2019. Visualizing and measuring the geometry of bert. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc.
  42. 42.Anna Rogers, Olga Kovaleva, and Anna Rumshisky. 2020. A primer in BERTology: What we know about how BERT works. Transactions of the Association for Computational Linguistics, 8:842–866.
  43. 43.Peter Rousseeuw. 1987. Silhouettes: a graphical aid to the interpretation and validation of cluster analysis. J. Comput. Appl. Math., 20(1):53–65.
  44. 44.Hassan Sajjad, Narine Kokhlikyan, Fahim Dalvi, and Nadir Durrani. 2021. Fine-grained interpretation and causation analysis in deep NLP models. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: Tutorials, pages 5–10, Online. Association for Computational Linguistics.
  45. 45.Mike Schuster and Kaisuke Nakajima. 2012. Japanese and korean voice search. In 2012 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 5149–5152. IEEE.
  46. 46.Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016. Neural machine translation of rare words with subword units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1715–1725, Berlin, Germany. Association for Computational Linguistics.
  47. 47.Wolfgang G Stock. 2010. Concepts and semantic relations in information science. Journal of the American Society for Information Science and Technology, 61(10):1951–1969.
  48. 48.Xavier Suau, Luca Zappella, and Nicholas Apostoloff. 2020. Finding experts in transformer models. CoRR, abs/2005.07647.
  49. 49.Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019. BERT rediscovers the classical NLP pipeline. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4593–4601, Florence, Italy. Association for Computational Linguistics.
  50. 50.Robert L. Thorndike. 1953. Who belongs in the family. Psychometrika, pages 267–276.
  51. 51.Erik F. Tjong Kim Sang and Sabine Buchholz. 2000. Introduction to the CoNLL-2000 shared task chunking. In Fourth Conference on Computational Natural Language Learning and the Second Learning Language in Logic Workshop.
  52. 52.Jesse Vig. 2019. A multiscale visualization of attention in the transformer model. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pages 37–42, Florence, Italy. Association for Computational Linguistics.
  53. 53.Ekaterina Vylomova, Trevor Cohn, Xuanli He, and Gholamreza Haffari. 2016. Word Representation Models for Morphologically Rich Languages in Neural Machine Translation. arXiv preprint arXiv:1606.04217.
  54. 54.John Wu, Hassan Belinkov, Yonatan Sajjad, Nadir Durrani, Fahim Dalvi, and James Glass. 2020. Similarity Analysis of Contextual Word Representation Models. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), Seattle. Association for Computational Linguistics.
  55. 55.Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. 2019. Xlnet: Generalized autoregressive pretraining for language understanding. Advances in neural information processing systems, 32.

Citation

MLA
Sajjad, H., et al. “Analyzing Encoded Concepts in Transformer Language Models”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2022, pp. 3082–101, https://doi.org/10.18653/v1/2022.naacl-main.225.
APA
Sajjad, H., Durrani, N., Dalvi, F., Alam, F., Khan, A., & Xu, J. (2022). Analyzing Encoded Concepts in Transformer Language Models. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 3082–3101. https://doi.org/10.18653/v1/2022.naacl-main.225
Chicago
Sajjad, H., N. Durrani, F. Dalvi, F. Alam, A. Khan, and J. Xu. 2022. “Analyzing Encoded Concepts in Transformer Language Models”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 3082–3101. https://doi.org/10.18653/v1/2022.naacl-main.225.
Harvard
Sajjad, H. et al. (2022) “Analyzing Encoded Concepts in Transformer Language Models”, Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp. 3082–3101. Available at: https://doi.org/10.18653/v1/2022.naacl-main.225.
Vancouver
1. Sajjad H, Durrani N, Dalvi F, Alam F, Khan A, Xu J (2022) Analyzing Encoded Concepts in Transformer Language Models. In: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp 3082–3101

BibTeX

@inproceedings{sajjad-etal-2022-analyzing,
    title = "Analyzing Encoded Concepts in Transformer Language Models",
    author = "Sajjad, Hassan  and
      Durrani, Nadir  and
      Dalvi, Fahim  and
      Alam, Firoj  and
      Khan, Abdul  and
      Xu, Jia",
    editor = "Carpuat, Marine  and
      de Marneffe, Marie-Catherine  and
      Meza Ruiz, Ivan Vladimir",
    booktitle = "Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
    month = jul,
    year = "2022",
    address = "Seattle, United States",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.naacl-main.225/",
    doi = "10.18653/v1/2022.naacl-main.225",
    pages = "3082--3101"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/