Analyzing Encoded Concepts in Transformer Language Models
Hassan SajjadNadir DurraniFahim DalviFiroj AlamAbdul Rafae KhanJia Xu
Proposes ConceptX, an unsupervised framework that clusters latent contextual representations and aligns them with human-defined linguistic categories to explain how transformer language models organize knowledge across layers without relying on probing classifiers.
Deep neural network language models have become foundational across modern natural language processing applications, yet their "black-box" nature creates significant risks regarding reliability, fairness, and governance. Understanding how these models internally organize linguistic and conceptual information is critical for ensuring safe deployment and control. The article addresses this challenge by introducing ConceptX, a framework designed to analyze and interpret the latent, context-aware representations learned across various layers of pre-trained transformer models.
The main objective of the article is to evaluate how internal model representations correspond to human-defined linguistic concepts and to establish where and how this knowledge is structured across network layers. To achieve this, the authors used an unsupervised agglomerative hierarchical clustering method to group contextualized word representations into 1,000 clusters per layer across seven prominent 12-layer transformer architectures (including BERT, RoBERTa, XLNet, ALBERT, and multilingual variants). These learned clusters—termed encoded concepts—were then evaluated against a broad suite of human-defined linguistic categories, spanning lexical units, morphology, syntax, semantics, and psycholinguistic ontologies, using a strict 90% alignment threshold on a standardized news dataset.
The evaluation revealed several key findings regarding how language models internally process text. First, between 43.6% and 72.4% of all learned clusters align directly with standard human-defined concepts, with multilingual models (such as XLM-RoBERTa at 72.4%) exhibiting significantly higher concept alignment than monolingual models (such as BERT-cased at 47.2%). Second, the internal architecture organizes information hierarchically: lower layers are dominated by shallow lexical patterns (like subword ngrams and affixes) and static semantic ontologies, whereas middle and higher layers (layers 8–10) primarily capture core linguistic properties such as morphology, parts-of-speech, and syntax. Third, morphological properties align much more strongly across all models (up to 26% fine-grained and 53% coarse alignment) than semantic (up to 16%) or complex syntactic categories (up to 14%). Finally, 28% to 56% of internal clusters do not match single traditional categories because they represent multi-faceted or compositional concepts—such as grouping words by both specific verb tense and semantic meaning simultaneously—which can be explained when combining multiple linguistic labels.
These findings demonstrate that language models naturally learn structured, hierarchical abstractions of language without explicit supervision, though their internal logic often blurs the boundaries of traditional human grammar rules. For practitioners and decision-makers, this means model performance does not always correlate directly with adherence to formal linguistic concepts; for instance, higher GLUE benchmark performance in some architectures does not necessarily mean higher alignment with human categories. Furthermore, training complexity (such as handling multiple languages or lacking capitalization cues) forces models to encode richer internal linguistic structures.
To build upon these insights, organizations and researchers deploying large transformer models should utilize unsupervised concept clustering as an interpretability check to audit model internals. Where traditional single-label taxonomies fail to explain model behaviors, teams should adopt compositional concept evaluations or human-in-the-loop assessments to capture multi-faceted internal representations. Future work should conduct controlled experiments to isolate how specific architectural parameters, pre-training objectives, and vocabulary tokenization schemes influence concept formation, as well as extend evaluations to larger, modern generative architectures.
- Paper: BERT Rediscovers the Classical NLP Pipeline, Ian Tenney et al. (2019). This paper establishes the foundational edge-probing framework showing that transformer layers hierarchically encode classical linguistic stages, directly informing ConceptX's layer-wise analysis.
- Paper: What Does BERT Learn about the Structure of Language?, Ganesh Jawahar et al. (2019). This study demonstrates how BERT organizes surface, syntactic, and semantic knowledge across its layers using diagnostic probing and clustering, providing the structural basis for analyzing encoded concepts.
- Paper: Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV), Been Kim et al. (2018). This work introduces testing with Concept Activation Vectors (TCAV) to quantify whether neural networks learn human-interpretable concepts, laying the groundwork for concept-level interpretability.
- Paper: Network Dissection: Quantifying Interpretability of Deep Visual Representations, David Bau et al. (2017). This foundational paper presents Network Dissection to quantify unit-level concept alignment against human-defined ontologies, inspiring unsupervised concept extraction and evaluation methodologies.
- Paper: How Contextual are Contextualized Word Representations? Comparing the Geometry of BERT, ELMo, and GPT-2 Embeddings, Kawin Ethayarajh (2019). This research analyzes the geometric properties and layer-wise contextuality of contextualized representations in models like BERT and GPT-2, providing essential geometric context for clustering hidden states.
- Paper: A Primer in BERTology: What We Know About How BERT Works, Anna Rogers et al. (2020). This comprehensive survey synthesizes empirical findings on how transformer language models encode linguistic knowledge across layers, serving as essential background for interpretability research.
- Paper: A Structural Probe for Finding Syntax in Word Representations, John Hewitt et al. (2019). This paper develops structural probes to identify syntactic trees in contextualized word vectors, illustrating how linguistic properties are embedded in transformer representation spaces.
- Paper: Concept Bottleneck Models, Pang Wei Koh et al. (2020). This work introduces Concept Bottleneck Models to structure neural predictions around human-interpretable intermediate concepts, highlighting the importance of evaluating concept-level representations.
- Paper: The Linear Representation Hypothesis and the Geometry of Large Language Models, Kiho Park et al. (2024). This paper advances the study of internal concept representations by formalizing the Linear Representation Hypothesis and developing causal inner products to geometrically isolate linguistic and semantic concepts in large language models.
- Paper: Representation Engineering: A Top-Down Approach to AI Transparency, Andy Zou et al. (2023). This work moves beyond post-hoc concept extraction by introducing representation engineering to actively read and steer higher-level cognitive and behavioral concepts across network layers.
- Paper: Sparse Autoencoders Find Highly Interpretable Features in Language Models, Hoagy Cunningham et al. (2023). This study extends unsupervised concept discovery by training sparse autoencoders to disentangle polysemantic features and superposed concept directions in language model activations.
- Paper: RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations, Jing Huang et al. (2024). This paper presents the RAVEL benchmark to rigorously evaluate how effectively interpretability techniques can isolate and disentangle specific attribute concepts in language models.
- Paper: A Holistic Approach to Unifying Automatic Concept Extraction and Concept Importance Estimation, Thomas Fel et al. (2023). This article unifies automated concept extraction and importance estimation within a dictionary-learning framework, providing formalized metrics that build upon empirical concept clustering.
- Paper: On the Origins of Linear Representations in Large Language Models, Yibo Jiang et al. (2024). This research provides mathematical foundations for why language models naturally form linear, orthogonal concept geometries during pre-training.
- Paper: From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs after Instruction Tuning, Xuansheng Wu et al. (2024). This paper applies concept extraction and representation analysis to investigate how instruction tuning alters internal knowledge organization and attention dynamics across transformer layers.
