keyword
contextualized representations
Contextualized representations are dynamic numerical vector embeddings in natural language processing that encode the meaning of a word, subword, or token based on the specific surrounding context in which it appears. Unlike static embeddings that assign a single, fixed vector to a word regardless of how it is used, contextualized representations generate distinct vector values for the same word depending on its syntactic role, semantic nuances, and neighboring text. Typically generated by deep neural architectures such as transformers, these representations are refined across successive model layers as attention mechanisms integrate information from across the entire sequence. This adaptability allows models to resolve lexical ambiguity, handle polysemy, and produce rich linguistic representations that can be transferred and fine-tuned across various language understanding and generation tasks.
4 items

On the Emergence of Position Bias in Transformers
Xinyi Wu, Yifei Wang, Stefanie Jegelka, Ali Jadbabaie
Why you should read this
Develops a graph-theoretic framework that mathematically explains how multi-layer causal masking and relative positional encodings interact across depth to produce systematic position biases such as attention sinks and the lost-in-the-middle effect in transformers.
Recent studies have revealed various manifestations of position bias in transformer architectures, from the “lost-in-the-middle” phenomenon to attention sinks, yet a comprehensive theoretical understanding of how attention masks and positional encodings shape these biases remains elusive. This paper presents a graph-theoretic framework for analyzing position bias in multi-layer attention. Modeling attention masks as directed graphs, we quantify how tokens interact with contextual information based on their sequential positions. We uncover two key insights: First, causal masking inherently biases attention toward earlier positions, as tokens in deeper layers attend to increasingly more contextualized representations of earlier tokens. Second, we characterize the competing effects of the causal mask and relative positional encodings, such as the decay mask and rotary positional encoding (RoPE): while both mechanisms introduce distance-based decay within individual attention maps, their aggregate effect across multiple attention layers—coupled with the causal mask—leads to a trade-off between the long-term decay effects and the cumulative importance of early sequence positions. Through controlled numerical experiments, we not only validate our theoretical findings but also reproduce position biases observed in real-world LLMs. Our framework offers a principled foundation for understanding positional biases in transformers, shedding light on the complex interplay of attention mechanism components and guiding more informed architectural design.
Added
2026-10-05

Finding Skill Neurons in Pre-trained Transformer-based Language Models
Xiaozhi Wang, Kaiyue Wen, Zhengyan Zhang, Lei Hou, Zhiyuan Liu, Juanzi Li
Why you should read this
Reveals that pre-trained Transformers contain task-predictive "skill neurons" formed during pre-training, providing a concrete method to identify them via prompt tuning and apply them to model pruning and transferability estimation.
Transformer-based pre-trained language models have demonstrated superior performance on various natural language processing tasks. However, it remains unclear how the skills required to handle these tasks distribute among model parameters. In this paper, we find that after prompt tuning for specific tasks, the activations of some neurons within pre-trained Transformers¹ are highly predictive of the task labels. We dub these neurons skill neurons and confirm they encode task-specific skills by finding that: (1) Skill neurons are crucial for handling tasks. Performances of pre-trained Transformers on a task significantly drop when corresponding skill neurons are perturbed. (2) Skill neurons are task-specific. Similar tasks tend to have similar distributions of skill neurons. Furthermore, we demonstrate the skill neurons are most likely generated in pre-training rather than fine-tuning by showing that the skill neurons found with prompt tuning are also crucial for other fine-tuning methods freezing neuron weights, such as the adapter-based tuning and BitFit. We also explore the applications of skill neurons, including accelerating Transformers with network pruning and building better transferability indicators. These findings may promote further research on understanding Transformers. The source code can be obtained from https://github.com/THU-KEG/Skill-Neuron.
Added
2026-10-02

A Closer Look at How Fine-tuning Changes BERT
Yichu Zhou, Vivek Srikumar
Why you should read this
Reveals how fine-tuning improves BERT representations by widening separation distances between distinct label clusters while preserving the model's underlying geometric structure across downstream tasks.
Given the prevalence of pre-trained contextualized representations in today's NLP, there have been many efforts to understand what information they contain, and why they seem to be universally successful. The most common approach to use these representations involves fine-tuning them for an end task. Yet, how fine-tuning changes the underlying embedding space is less studied. In this work, we study the English BERT family and use two probing techniques to analyze how fine-tuning changes the space. We hypothesize that fine-tuning affects classification performance by increasing the distances between examples associated with different labels. We confirm this hypothesis with carefully designed experiments on five different NLP tasks. Via these experiments, we also discover an exception to the prevailing wisdom that "fine-tuning always improves performance". Finally, by comparing the representations before and after fine-tuning, we discover that fine-tuning does not introduce arbitrary changes to representations; instead, it adjusts the representations to downstream tasks while largely preserving the original spatial structure of the data points.
Added
2026-10-02

How Can We Know What Language Models Know?
Zhengbao Jiang, Frank F. Xu, Jun Araki, Graham Neubig
Why you should read this
Demonstrates that manual prompts underestimate the factual knowledge stored in language models and introduces automated mining, paraphrasing, and ensembling methods to substantially improve relation extraction accuracy on the LAMA benchmark.
Recent work has presented intriguing results examining the knowledge contained in language models (LM) by having the LM fill in the blanks of prompts such as "Obama is a _ by profession". These prompts are usually manually created, and quite possibly sub-optimal; another prompt such as "Obama worked as a _" may result in more accurately predicting the correct profession. Because of this, given an inappropriate prompt, we might fail to retrieve facts that the LM does know, and thus any given prompt only provides a lower bound estimate of the knowledge contained in an LM. In this paper, we attempt to more accurately estimate the knowledge contained in LMs by automatically discovering better prompts to use in this querying process. Specifically, we propose mining-based and paraphrasing-based methods to automatically generate high-quality and diverse prompts, as well as ensemble methods to combine answers from different prompts. Extensive experiments on the LAMA benchmark for extracting relational knowledge from LMs demonstrate that our methods can improve accuracy from 31.1% to 39.6%, providing a tighter lower bound on what LMs know. We have released the code and the resulting LM Prompt And Query Archive (LPAQA) at this https URL.
Added
2026-09-24
