Layer by Layer: Uncovering Hidden Representations in Language Models
Oscar SkeanMd Rifat ArefinDan ZhaoNiket PatelJalal NaghiyevYann LeCunRavid Shwartz-Ziv
Demonstrates that intermediate layers in language models consistently produce richer representations than the final layer, introducing a geometric and information-theoretic framework that explains why mid-depth embeddings achieve superior performance across diverse downstream tasks.
Artificial intelligence applications typically rely on the final layer of large language models to extract text embeddings and semantic features. This standard practice assumes that deeper layers capture the most refined abstractions while earlier layers hold only low-level cues. However, final layers often become overly specialized to the specific pretraining objective, such as next-token prediction, which can degrade their utility for general downstream applications. Determining which layers actually generate the highest quality representations is essential for maximizing model performance, improving computational efficiency, and deploying more robust systems.
The article evaluates how internal representations evolve across model depth and demonstrates that intermediate hidden layers systematically provide stronger, more useful representations than the final layers across a variety of architectures and tasks. It also establishes a unified theoretical framework based on matrix-based entropy to quantify and explain how neural networks balance information compression, geometric structure, and invariance to input perturbations.
To conduct this evaluation, the analysis tested embeddings across every layer of diverse models—including autoregressive transformers like Pythia and Llama 3, state-space models like Mamba, and encoder-based models like BERT—on 32 tasks from the Massive Text Embedding Benchmark spanning classification, clustering, and reranking. The investigation analyzed parameter scales ranging from 14 million to several billion and evaluated representations across training checkpoints, fine-tuning regimes, extreme input conditions, and computer vision models.
The findings show that intermediate layers consistently outperform final layers on downstream embedding benchmarks, achieving absolute performance gains ranging from 2% to 16%. In addition, autoregressive models develop a pronounced mid-depth compression bottleneck driven by residual connections, whereas bidirectional models maintain stable entropy across layers. In computer vision, autoregressive patch-prediction models display the exact same mid-layer bottleneck and intermediate performance peak, proving that the training objective rather than the data modality drives this behavior. Furthermore, unsupervised quality metrics strongly correlate with downstream task performance, enabling automated, zero-label selection of optimal intermediate layers that boosts benchmark scores by 3% over standard final layers.
These findings indicate that conventional AI deployment pipelines routinely discard the most effective representations by defaulting to final-layer outputs. Harnessing mid-depth representations can improve task accuracy without retraining models and presents opportunities to prune subsequent layers to reduce inference costs and latency. In addition, chain-of-thought fine-tuning was shown to expand intermediate entropy, providing a mechanistic explanation for how reasoning models retain necessary contextual breadth during multi-step problem solving.
Organizations developing or deploying language and vision models should audit intermediate layers rather than defaulting to final outputs, using unsupervised metrics to select the most effective layer for each application. Engineering teams should explore early-exit architectures and layer pruning to decrease compute costs where mid-depth representations suffice. Decision-makers should also invest in further exploration of explicitly regularized training regimes that optimize intermediate compression directly.
While the empirical findings are consistent across 32 text benchmarks and standard vision evaluations, the analysis primarily focuses on linear kernel representations and feature extraction tasks. Stakeholders should exercise caution because intermediate representations may also expose latent biases encoded earlier in the network. Overall confidence in the empirical observations remains high across evaluated model families, but practitioners should validate layer selection on domain-specific workloads before full-scale deployment.
- Paper: BERT Rediscovers the Classical NLP Pipeline, Ian Tenney et al. (2019). This paper establishes the layer-by-layer probing methodology showing how linguistic representations evolve across network depth, providing foundational context for analyzing intermediate versus final layers.
- Paper: What Does BERT Learn about the Structure of Language?, Ganesh Jawahar et al. (2019). It provides essential empirical evidence on how surface, syntactic, and semantic information are hierarchically distributed across hidden layers in transformer language models.
- Paper: How Contextual are Contextualized Word Representations? Comparing the Geometry of BERT, ELMo, and GPT-2 Embeddings, Kawin Ethayarajh (2019). It examines the geometry and layer-wise contextuality of transformer representations, establishing the geometric properties and anisotropy evaluated in intermediate representations.
- Paper: Deep learning and the information bottleneck principle, Naftali Tishby et al. (2015). It provides the foundational theoretical framework using the information bottleneck principle to analyze how successive neural network layers balance information compression and preservation.
- Paper: Opening the Black Box of Deep Neural Networks via Information, Ravid Shwartz-Ziv et al. (2017). It formalizes mutual information tracking across deep neural network layers, establishing the analytical basis for evaluating representation quality across network depth.
- Paper: A Structural Probe for Finding Syntax in Word Representations, John Hewitt et al. (2019). It introduces geometric structural probing to show that intermediate representations often capture structural linguistic relationships better than final layers.
- Paper: Understanding Contrastive Representation Learning through Alignment and Uniformity on the Hypersphere, Tongzhou Wang et al. (2020). It defines crucial representation quality metrics—specifically alignment and uniformity on hyperspheres—that underpin geometric evaluations of neural embeddings.
- Paper: Deep contextualized word representations, Matthew E. Peters et al. (2018). It introduces the paradigm of leveraging multi-layer representations from deep contextual models to improve downstream task performance.
- Paper: TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI Agents, Bofei Zhang et al. (2026). It builds upon the mechanistic analysis of layer-specific LLM representations by examining how different network layers localize and route cross-lingual concepts in foundation models.
- Paper: Learn from your own latents and not from tokens: A sample-complexity theory, Daniel J. Korchinski et al. (2026). It extends the study of internal representations by demonstrating the theoretical sample-complexity benefits of training models on self-supervised intermediate latents rather than surface tokens.
