Layer by Layer: Uncovering Hidden Representations in Language Models

Oscar SkeanMd Rifat ArefinDan ZhaoNiket PatelJalal NaghiyevYann LeCunRavid Shwartz-Ziv

article2025ICML192 citations

Demonstrates that intermediate layers in language models consistently produce richer representations than the final layer, introducing a geometric and information-theoretic framework that explains why mid-depth embeddings achieve superior performance across diverse downstream tasks.

Listen

Artificial intelligence applications typically rely on the final layer of large language models to extract text embeddings and semantic features. This standard practice assumes that deeper layers capture the most refined abstractions while earlier layers hold only low-level cues. However, final layers often become overly specialized to the specific pretraining objective, such as next-token prediction, which can degrade their utility for general downstream applications. Determining which layers actually generate the highest quality representations is essential for maximizing model performance, improving computational efficiency, and deploying more robust systems.

The article evaluates how internal representations evolve across model depth and demonstrates that intermediate hidden layers systematically provide stronger, more useful representations than the final layers across a variety of architectures and tasks. It also establishes a unified theoretical framework based on matrix-based entropy to quantify and explain how neural networks balance information compression, geometric structure, and invariance to input perturbations.

To conduct this evaluation, the analysis tested embeddings across every layer of diverse models—including autoregressive transformers like Pythia and Llama 3, state-space models like Mamba, and encoder-based models like BERT—on 32 tasks from the Massive Text Embedding Benchmark spanning classification, clustering, and reranking. The investigation analyzed parameter scales ranging from 14 million to several billion and evaluated representations across training checkpoints, fine-tuning regimes, extreme input conditions, and computer vision models.

The findings show that intermediate layers consistently outperform final layers on downstream embedding benchmarks, achieving absolute performance gains ranging from 2% to 16%. In addition, autoregressive models develop a pronounced mid-depth compression bottleneck driven by residual connections, whereas bidirectional models maintain stable entropy across layers. In computer vision, autoregressive patch-prediction models display the exact same mid-layer bottleneck and intermediate performance peak, proving that the training objective rather than the data modality drives this behavior. Furthermore, unsupervised quality metrics strongly correlate with downstream task performance, enabling automated, zero-label selection of optimal intermediate layers that boosts benchmark scores by 3% over standard final layers.

These findings indicate that conventional AI deployment pipelines routinely discard the most effective representations by defaulting to final-layer outputs. Harnessing mid-depth representations can improve task accuracy without retraining models and presents opportunities to prune subsequent layers to reduce inference costs and latency. In addition, chain-of-thought fine-tuning was shown to expand intermediate entropy, providing a mechanistic explanation for how reasoning models retain necessary contextual breadth during multi-step problem solving.

Organizations developing or deploying language and vision models should audit intermediate layers rather than defaulting to final outputs, using unsupervised metrics to select the most effective layer for each application. Engineering teams should explore early-exit architectures and layer pruning to decrease compute costs where mid-depth representations suffice. Decision-makers should also invest in further exploration of explicitly regularized training regimes that optimize intermediate compression directly.

While the empirical findings are consistent across 32 text benchmarks and standard vision evaluations, the analysis primarily focuses on linear kernel representations and feature extraction tasks. Stakeholders should exercise caution because intermediate representations may also expose latent biases encoded earlier in the network. Overall confidence in the empirical observations remains high across evaluated model families, but practitioners should validate layer selection on domain-specific workloads before full-scale deployment.

arXiv: 2502.02013
Cover for Layer by Layer: Uncovering Hidden Representations in Language Models

Abstract

From extracting features to generating text, the outputs of large language models (LLMs) typically rely on the final layers, following the conventional wisdom that earlier layers capture only low-level cues. However, our analysis shows that intermediate layers can encode even richer representations, often improving performance on a range of downstream tasks. To explain and quantify these hidden-layer properties, we propose a unified framework of representation quality metrics based on information theory, geometry, and invariance to input perturbations. Our framework highlights how each layer balances information compression and signal preservation, revealing why mid-depth embeddings can exceed the last layer's performance. Through extensive experiments on 32 text-embedding tasks across various architectures (transformers, state-space models) and domains (language, vision), we demonstrate that intermediate layers consistently provide stronger features, challenging the standard view on final-layer embeddings and opening new directions on using mid-layer representations for more robust and accurate representations.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 A Unified Framework for Neural Representations
  • 3.1 Notation and Motivation
  • 3.2 Matrix-Based Entropy: A Common Theoretical Thread
  • 3.3 Representation Evaluation Metrics
  • 3.3.1 Information-Theoretic Metrics
  • 3.3.2 Geometric Metrics
  • 3.3.3 Augmentation Invariance Metrics
  • 3.4 Core Theoretical Results
  • 4 Empirical Results
  • 4.1 Downstream Task Performance
  • 4.1.1 Experimental Setup
  • 4.1.2 Intermediate Layers Often Outperform Final Layers
  • 4.1.3 Layer-Wise Metrics Correlate with Downstream Performance
  • 4.2 Architectural and Scale Differences
  • 4.3 Impact of Training Progression
  • 4.4 Impact of Chain-of-Thought Finetuning
  • 5 Extreme Input Conditions
  • 6 Comparison to Vision Transformers
  • 7 Discussion and Conclusion
  • References
  • A Architectural Details
  • A.1 Transformer
  • A.2 State Space Models
  • B Discussion on Prompt Entropy
  • C Dataset Details
  • C.1 Wikitext Dataset
  • C.2 MTEB
  • D Prompt Augmentations
  • E Using Evaluation Metrics as a Performance Proxy
  • F Extreme Prompts
  • F.1 Increasing Repetition
  • F.2 Increasing Randomness
  • G Theorems
  • H Additional Plots & Visualizations

Knowls

  1. Knowl 1 — Matrix-Based Entropy as a Unified Metric for Hidden Representations

    definition

    Let Z∈RN×DZ \in \mathbb{R}^{N \times D} represent a hidden representation matrix corresponding to NN data samples (or token vectors) in a DD-dimensional feature space. The Gram matrix is defined as K=ZZ⊤∈RN×NK = Z Z^\top \in \mathbb{R}^{N \times N}, with non-negative eigenvalues {λi(K)}i=1r\{\lambda_i(K)\}_{i=1}^r, where r=rank(K)≤min⁡(N,D)r = \text{rank}(K) \le \min(N, D). The matrix-based Rényi entropy of order α>0\alpha > 0 for ZZ is defined as:

    Sα(Z)=11−αlog⁡∑i=1r(λi(K)tr(K))αS_\alpha(Z) = \frac{1}{1 - \alpha} \log \sum_{i=1}^r \left( \frac{\lambda_i(K)}{\text{tr}(K)} \right)^\alpha

    Normalizing by the trace tr(K)=∑i=1rλi(K)\text{tr}(K) = \sum_{i=1}^r \lambda_i(K) guarantees that the normalized eigenvalues form a valid probability distribution over principal components. When α→1\alpha \to 1, S1(Z)S_1(Z) converges to the matrix-based Shannon (von Neumann) entropy, and when α=2\alpha = 2, it yields collision entropy.

    This framework defines two distinct representational metrics:

    1. Prompt Entropy: Z∈RL×DZ \in \mathbb{R}^{L \times D} contains the LL token embeddings of an individual sequence; lower prompt entropy indicates intra-prompt compression and feature redundancy, whereas higher prompt entropy indicates diverse token-level features.
    2. Dataset Entropy: Z∈RN×DZ \in \mathbb{R}^{N \times D} consists of mean-pooled token vectors representing NN distinct prompts; higher dataset entropy indicates global separation and distinctiveness between different inputs.
  2. Knowl 2 — Downstream Performance Superiority of Intermediate Layer Representations

    empirical result

    Across multiple neural language model architectures—including decoder-only transformers (such as Pythia-410M and Llama 3), state-space models (such as Mamba-370M), and encoder-only transformers (BERT-base)—intermediate hidden layer embeddings consistently outperform the final layer representations when evaluated on the 32 downstream text-embedding tasks of the Massive Text Embedding Benchmark (MTEB).

    Key performance characteristics include:

    • Downstream embedding performance (spanning classification, clustering, reranking, and semantic textual similarity) peaks near the mid-depth layers of the network rather than at the final layer.
    • Across tasks, selecting optimal mid-depth representations yields an absolute accuracy improvement over the final layer ranging from 2%2\% to 16%16\%.
    • Final layer representations become overly specialized to the next-token prediction or masked language pretraining objectives, whereas intermediate layers maintain an optimal trade-off between semantic abstraction and the preservation of non-local contextual information.
  3. Knowl 3 — Lower Bound on Matrix-Based Shannon Entropy via Effective Rank

    theoretical result

    For an embedding matrix Z∈RN×DZ \in \mathbb{R}^{N \times D} with singular values σ1≥σ2≥⋯≥σmin⁡(N,D)≥0\sigma_1 \ge \sigma_2 \ge \dots \ge \sigma_{\min(N,D)} \ge 0 and covariance eigenvalues λi(Z⊤Z)=σi2\lambda_i(Z^\top Z) = \sigma_i^2, the effective rank EffRank(Z)=exp⁡(−∑ipilog⁡pi)\text{EffRank}(Z) = \exp\left(-\sum_{i} p_i \log p_i\right) (where pi=σi/∑jσjp_i = \sigma_i / \sum_j \sigma_j) is upper-bounded by the exponential of the matrix-based Shannon entropy S1(Z)S_1(Z):

    EffRank(Z)≤exp⁡(S1(Z))\text{EffRank}(Z) \le \exp(S_1(Z))

    Because σi≤1\sigma_i \le 1 under trace normalization, σi≥σi2=λi\sigma_i \ge \sigma_i^2 = \lambda_i, causing the eigenvalue distribution λ\lambda to be majorized by the singular value distribution σ\sigma (i.e., λ≺σ\lambda \prec \sigma). By the Schur-concavity of matrix-based entropy S1S_1, S1(σ)≤S1(λ)S_1(\sigma) \le S_1(\lambda), establishing that a high effective dimensionality strictly requires high matrix-based representation entropy.

  4. Knowl 4 — Dataset Entropy as an Upper Bound on InfoNCE Mutual Information

    theoretical result

    Let X∼DX \sim \mathcal{D} be a discrete input variable and Z=Z(X)Z = Z(X) be its learned representation via a Markov chain X→ZX \to Z. For a batch of NN samples, the self-supervised InfoNCE loss satisfies the bound:

    log⁡(N)−InfoNCE≤I(X;Z)≤H(Z)\log(N) - \text{InfoNCE} \le I(X; Z) \le H(Z)

    where I(X;Z)I(X; Z) is the mutual information between the inputs and representations, and H(Z)H(Z) is the representation entropy evaluated at the dataset level. Consequently, reducing the InfoNCE loss (maximizing invariance to augmentations while contrasting negative pairs) lower-bounds the mutual information I(X;Z)I(X; Z) and requires representations to exhibit sufficiently high dataset-level entropy H(Z)H(Z) to prevent representational collapse.

  5. Knowl 5 — Scaling Limits of Dataset Entropy Under Prompt Diversity Extremes

    theoretical result

    Let Z(⋅)Z(\cdot) be an orthogonally equivariant model mapping sequences of length L=DL = D sampled from an isotropic Gaussian distribution to unit-norm row representations, and let Zˉ∈RN×D\bar{Z} \in \mathbb{R}^{N \times D} denote the dataset matrix formed by sequence-averaged embeddings across NN independent inputs. The collision dataset entropy S2(ZˉZˉ⊤)=−log⁡∥ZˉZˉ⊤∥F2S_2(\bar{Z}\bar{Z}^\top) = -\log \|\bar{Z}\bar{Z}^\top\|_F^2 exhibits distinct scaling behaviors depending on token-level prompt entropy:

    1. Maximum Prompt Entropy: When token representations within every prompt are mutually orthonormal (maximal prompt entropy), then with probability at least 1−N22πexp⁡(−Dϵ22N2)1 - N^2 \sqrt{2\pi} \exp\left(-\frac{D\epsilon^2}{2N^2}\right):

    ∣e−S2(ZˉZˉ⊤)−NL2∣≤ϵ\left| e^{-S_2(\bar{Z}\bar{Z}^\top)} - \frac{N}{L^2} \right| \le \epsilon

    meaning S2(ZˉZˉ⊤)S_2(\bar{Z}\bar{Z}^\top) grows on the order of log⁡(L2N)\log\left(\frac{L^2}{N}\right).

    1. Minimum Prompt Entropy: When token representations within every prompt collapse to rank-one identical vectors (minimal prompt entropy), then with probability at least 1−N22πexp⁡(−D3ϵ22N8)1 - N^2 \sqrt{2\pi} \exp\left(-\frac{D^3\epsilon^2}{2N^8}\right):

    ∣e−S2(ZˉZˉ⊤)−N3L2∣≤ϵ\left| e^{-S_2(\bar{Z}\bar{Z}^\top)} - \frac{N^3}{L^2} \right| \le \epsilon

    meaning S2(ZˉZˉ⊤)S_2(\bar{Z}\bar{Z}^\top) grows more slowly on the order of log⁡(L2N3)\log\left(\frac{L^2}{N^3}\right).

  6. Knowl 6 — Architectural Disparities in Hidden Layer Information Compression

    empirical result

    Different language model architectures display distinct layer-wise compression profiles when evaluated using prompt entropy, trajectory curvature Cˉ\bar{C}, and augmentation invariance metrics (LiDAR, DiME, InfoNCE):

    • Autoregressive Transformers (e.g., Pythia): Exhibit a pronounced mid-layer "compression valley" where prompt entropy sharply decreases at 40%–60%40\%\text{--}60\% network depth before rising at the final layers. Concurrently, token trajectory curvature peaks and augmentation invariance (LiDAR/DiME) increases in these middle layers, reflecting aggressive pruning of local details to extract high-level semantic context.
    • Bidirectional Encoders (e.g., BERT-base): Maintain high and flat prompt entropy throughout all layers with minimal layer-wise variation, reflecting the absence of a causal mask and eliminating the need to discard non-local features early.
    • State Space Models (e.g., Mamba-370M): Produce smooth, intermediate, and relatively uniform metric trajectories across depth, compressing information moderately without forming the deep mid-layer entropy bottleneck characteristic of autoregressive transformers.
  7. Knowl 7 — Autoregressive Pretraining Objective as the Root Driver of Mid-Layer Compression Across Modalities

    empirical result

    The mid-depth representation bottleneck is governed by the autoregressive training objective rather than the underlying data modality (language vs. vision). When comparing 24-layer ViT-L architectures on ImageNet-100 and ImageNet-1k using attention probing across layers:

    • Autoregressive Vision Models (AIM): Predicting image patches autoregressively (GPT-style next-patch prediction) produces the exact same intermediate prompt entropy dip and a mid-layer accuracy peak observed in autoregressive language models, outperforming the final layer by up to 4%4\% top-1 probing accuracy.
    • Non-Autoregressive Vision Models (ViT, CLIP, DINOv2, MAE): Supervised, contrastive, and masked autoencoding vision models display monotonically increasing downstream probing accuracy toward their final layers and steadily increasing dataset entropy across depth, with no mid-depth accuracy peaks.
    • Masked Token Pretraining (BEiT): Exhibits an intermediate entropy drop without the corresponding mid-layer probing performance peak seen in purely autoregressive models.
  8. Knowl 8 — Residual Connections Drive Transformer Mid-Layer Information Compression

    empirical result

    Decomposing transformer blocks into individual sub-components (pre-attention normalization, multi-head self-attention, pre-MLP representations, MLP outputs, and residual stream additions) reveals that residual connections are the primary mechanical driver of the mid-network entropy valley in autoregressive models:

    • Sub-layers evaluated prior to residual addition (such as raw self-attention outputs and pre-residual MLP projections) maintain high effective rank and high entropy across intermediate layers.
    • Passing representations through the residual summation (z(l)=z(l−1)+SubLayer(z(l−1))z^{(l)} = z^{(l-1)} + \text{SubLayer}(z^{(l-1)})) causes an abrupt drop in effective rank and entropy specifically in the middle layers (layers 6 to 18 in Pythia-410M).
    • Residual paths serve as an implicit regularizer, filtering out high-frequency noise and merging new contextual updates into a compressed low-rank subspace.
  9. Knowl 9 — Training Dynamics of Representation Metrics and the Early Detokenization Hypothesis

    empirical result

    Tracking representation quality metrics across training steps (from checkpoint step 1 to step 143k in Pythia-410M) reveals two distinct layer-wise regimes:

    1. Early Layer Rapid Stabilization: The earliest layers (0%–20%0\%\text{--}20\% depth) stabilize their representation metrics (entropy, LiDAR, and curvature) within the initial phase of pretraining, remaining virtually invariant thereafter. This supports the detokenization hypothesis, indicating that early layers rapidly learn a fixed mapping from discrete token inputs into a continuous embedding space.
    2. Continuous Intermediate Layer Evolution: The mid-depth layers (40%–70%40\%\text{--}70\% depth) undergo continuous structural changes throughout the entirety of pretraining. Prompt entropy steadily decreases, token trajectory curvature smooths, and augmentation invariance tightens over 143k steps, demonstrating that high-level feature distillation and abstraction develop progressively in intermediate layers.
  10. Knowl 10 — Chain-of-Thought Finetuning Sustains Latent Context Diversity Across Sequence Positions

    empirical result

    Comparing base language models (Qwen 2.5) against reasoning-specialized models trained with math pretraining and Chain-of-Thought finetuning (Qwen 2.5-Math) reveals that CoT finetuning alters token-level representation entropy across sequence positions:

    • Standard autoregressive base models display heavy prompt entropy compression across token positions, reducing the representational capacity allocated to later tokens.
    • Chain-of-thought finetuned models maintain systematically higher token-level prompt entropy with reduced across-sample variance across long context lengths.
    • This higher latent entropy prevents representational collapse, enabling the model to retain rich multi-step reasoning context throughout deep transformer layers.
  11. Knowl 11 — Unsupervised Layer Selection for Downstream Tasks via Representation Quality Metrics

    data/table

    Unsupervised representation metrics (DiME, InfoNCE, and Dataset Entropy) computed in a single forward pass without task labels can identify high-performing intermediate layers that surpass standard final-layer representations on downstream tasks.

    Model Supervised (Best) Naive (Last Layer) min-DiME min-InfoNCE min-Dataset Entropy
    Pythia-410M 52.0% 45.5% 48.5% 46.2% 48.1%
    LLM2Vec-8B 66.3% 63.9% 60.0% 64.3% 50.4%

    The table demonstrates average accuracy across 32 MTEB tasks:

    • On Pythia-410M, selecting the layer that minimizes DiME achieves 48.5%48.5\%, outperforming the naive final layer (45.5%45.5\%) by 3.0%3.0\% without supervision.
    • On LLM2Vec-8B, selecting the layer minimizing InfoNCE achieves 64.3%64.3\%, exceeding the final layer baseline (63.9%63.9\%).
    • In comparison, an oracle supervised selection choosing the optimal layer per task achieves 52.0%52.0\% on Pythia-410M and 66.3%66.3\% on LLM2Vec-8B.
  12. Knowl 12 — Layer-Specific Sensitivity to Token Repetition and Random Token Injections

    empirical result

    Evaluating Pythia-410M under extreme input perturbations demonstrates differential sensitivity across network depth:

    • Token Repetition: Replacing tokens with repeated intra-prompt tokens with probability p∈[0,1]p \in [0, 1] causes a sharp, selective collapse in normalized prompt entropy in mid-depth layers (dropping from ∼0.7\sim 0.7 at p=0p=0 to ∼0.2\sim 0.2 at p=1.0p=1.0), demonstrating that intermediate layers actively compress repetitive and redundant patterns.
    • Random Token Noise: Replacing tokens with uniformly sampled random vocabulary tokens with probability p∈[0,1]p \in [0, 1] sharply inflates prompt entropy in early layers (layers 0–5 rising towards 1.0), whereas intermediate and deep layers filter out the noise and maintain stable entropy trajectories.

Coverage note — All primary contributions—including the mathematical framework of matrix-based entropy, theoretical bounding theorems, empirical layer-wise evaluations on MTEB, architectural comparisons (transformers, SSMs, vision models), training dynamics, sub-layer residual analysis, CoT finetuning effects, extreme input tests, and unsupervised layer selection—are covered. Specific minor implementation details (e.g., exact QWERTY keyboard augmentation neighbor lists and the exhaustive list of individual MTEB task dataset names) were omitted to maintain granularity.

References

  1. 1.Agrawal, K. K., Mondal, A. K., Ghosh, A., and Richards, B. α\alpha-ReQ: Assessing representation quality in self-supervised learning by measuring eigenspectrum decay. NeurIPs, 2022.
  2. 2.Alain, G. and Bengio, Y. Understanding intermediate layers using linear classifier probes. ICLR, 2017.
  3. 3.Arefin, M. R., Subbaraj, G., Gontier, N., LeCun, Y., Rish, I., Shwartz-Ziv, R., and Pal, C. Seq-VCR: Preventing collapse in intermediate transformer representations for enhanced reasoning. ICLR, 2025.
  4. 4.Bach, F. Information theory with kernel methods. IEEE Transactions on Information Theory, 2022.
  5. 5.Bao, H., Dong, L., Piao, S., and Wei, F. BeIT: Bert pre-training of image transformers. ICLR, 2022.
  6. 6.Barbero, F., Arroyo, A., Gu, X., Perivolaropoulos, C., Bronstein, M., Veličković, P., and Pascanu, R. Why do LLMs attend to the first token? arXiv, 2025.
  7. 7.Behnam Ghader, P., Adlakha, V., Mosbach, M., Bahdanau, D., Chapados, N., and Reddy, S. LLM2Vec: Large language models are secretly powerful text encoders. COLM, 2024.
  8. 8.Biderman, S., Schoelkopf, H., Anthony, Q. G., Bradley, H., O’Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., et al. Pythia: A suite for analyzing large language models across training and scaling. ICML, 2023.
  9. 9.Boes, P., Eisert, J., Gallego, R., Müller, M. P., and Wilming, H. Von neumann entropy from unitarity. Physical review letters, 2019.
  10. 10.Bordes, F., Balestriero, R., Garrido, Q., Bardes, A., and Vincent, P. Guillotine regularization: Why removing layers is needed to improve generalization in self-supervised learning. TMLR, 2023.
  11. 11.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. Language models are few-shot learners. NeurIPs, 2020.
  12. 12.Brunner, G., Liu, Y., Pascual, D., Richter, O., Ciaramita, M., and Wattenhofer, R. On identifiability in transformers. ICLR, 2020.
  13. 13.Burns, C., Ye, H., Klein, D., and Steinhardt, J. Discovering latent knowledge in language models without supervision. ICLR, 2023.
  14. 14.Chen, M., Radford, A., Child, R., Wu, J., Jun, H., Luan, D., and Sutskever, I. Generative pretraining from pixels. ICML, 2020.
  15. 15.Cheng, E., Doimo, D., Kervadec, C., Macocco, I., Yu, J., Laio, A., and Baroni, M. Emergence of a high-dimensional abstraction phase in language transformers. ICLR, 2025.
  16. 16.Csordás, R., Manning, C. D., and Potts, C. Do language models use their depth efficiently? arXiv, 2025.
  17. 17.DeepSeek-AI. Deepseek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning. arXiv, 2025.
  18. 18.Deletang, G., Ruoss, A., Duquenne, P.-A., Catt, E., Genewein, T., Mattern, C., Grau-Moya, J., Wenliang, L. K., Aitchison, M., Orseau, L., et al. Language modeling is compression. ICLR, 2024.
  19. 19.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. BERT: Pre-training of deep bidirectional transformers for language understanding. NAACL, 2019.
  20. 20.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al. An image is worth 16x16 words: Transformers for image recognition at scale. ICLR, 2021.
  21. 21.Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al. The Llama 3 herd of models. arXiv, 2024.
  22. 22.El-Nouby, A., Klein, M., Zhai, S., Bautista, M. A., Toshev, A., Shankar, V., Susskind, J. M., and Joulin, A. Scalable pre-training of large autoregressive image models. ICML, 2024.
  23. 23.Fan, S., Jiang, X., Li, X., Meng, X., Han, P., Shang, S., Sun, A., Wang, Y., and Wang, Z. Not all layers of LLMs are necessary during inference. arXiv, 2024.
  24. 24.Fini, E., Shukor, M., Li, X., Dufter, P., Klein, M., Haldimann, D., Aitharaju, S., da Costa, V. G. T., Béthune, L., Gan, Z., et al. Multimodal autoregressive pre-training of large vision encoders. CVPR, 2025.
  25. 25.Garrido, Q., Balestriero, R., Najman, L., and Lecun, Y. RankMe: Assessing the downstream performance of pretrained self-supervised representations by their rank. ICML, 2023.
  26. 26.Giraldo, L. G. S., Rao, M., and Principe, J. C. Measures of entropy from data using infinitely divisible kernels. IEEE Transactions on Information Theory, 2014.
  27. 27.Gu, A. and Dao, T. Mamba: Linear-time sequence modeling with selective state spaces. COLM, 2024.
  28. 28.Gu, X., Pang, T., Du, C., Liu, Q., Zhang, F., Du, C., Wang, Y., and Lin, M. When attention sink emerges in language models: An empirical view. ICLR, 2025.
  29. 29.Gurnee, W. and Tegmark, M. Language models represent space and time. arXiv, 2023.
  30. 30.Hao, S., Sukhbaatar, S., Su, D., Li, X., Hu, Z., Weston, J., and Tian, Y. Training large language models to reason in a continuous latent space. arXiv, 2024.
  31. 31.He, K., Chen, X., Xie, S., Li, Y., Dollár, P., and Girshick, R. Masked autoencoders are scalable vision learners. CVPR, 2022.
  32. 32.Hosseini, E. and Fedorenko, E. Large language models implicitly learn to straighten neural sentence trajectories to construct a predictive representation of natural language. NeurIPs, 2023.
  33. 33.Jin, M., Yu, Q., Huang, J., Zeng, Q., Wang, Z., Hua, W., Zhao, H., Mei, K., Meng, Y., Ding, K., et al. Exploring concept depth: How large language models acquire knowledge at different layers? arXiv, 2024.
  34. 34.Lad, V., Gurnee, W., and Tegmark, M. The remarkable robustness of LLMs: Stages of inference? arXiv, 2024.
  35. 35.Li, Y., Choi, D., Chung, J., Kushman, N., Schrittwieser, J., Leblond, R., Eccles, T., Keeling, J., Gimeno, F., Dal Lago, A., et al. Competition-level code generation with alphacode. Science, 2022.
  36. 36.Liu, N. F., Gardner, M., Belinkov, Y., Peters, M. E., and Smith, N. A. Linguistic knowledge and transferability of contextual representations. NAACL, 2019.
  37. 37.Ma, E. NLP Augmentation, 2019. URL https://github.com/makcedward/nlpaug.
  38. 38.Mallen, A. T. and Belrose, N. Eliciting latent knowledge from quirky language models. ICLR 2024 Workshop on Mathematical and Empirical Understanding of Foundation Models, 2024.
  39. 39.Mamou, J., Le, H., Del Rio, M. A., Stephenson, C., Tang, H., Kim, Y., and Chung, S. Emergence of separable manifolds in deep language representations. ICML, 2020.
  40. 40.Marion, P., Wu, Y.-H., Sander, M. E., and Biau, G. Implicit regularization of deep residual networks towards neural odes. ICLR, 2024.
  41. 41.Merity, S., Xiong, C., Bradbury, J., and Socher, R. Pointer sentinel mixture models. ICLR, 2017.
  42. 42.Muennighoff, N., Tazi, N., Magne, L., and Reimers, N. MTEB: Massive text embedding benchmark. EACL, 2022.
  43. 43.Oord, A. v. d., Li, Y., and Vinyals, O. Representation learning with contrastive predictive coding. ICLR, 2018.
  44. 44.Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., et al. DINOv2: Learning robust visual features without supervision. TMLR, 2024.
  45. 45.Park, K., Choe, Y. J., Jiang, Y., and Veitch, V. The geometry of categorical and hierarchical concepts in large language models. ICML 2024 Workshop on Mechanistic Interpretability, 2024a.
  46. 46.Park, K., Choe, Y. J., and Veitch, V. The linear representation hypothesis and the geometry of large language models. ICML, 2024b.
  47. 47.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. Learning transferable visual models from natural language supervision. ICML, 2021.
  48. 48.Raghu, M., Gilmer, J., Yosinski, J., and Sohl-Dickstein, J. SVCCA: Singular vector canonical correlation analysis for deep learning dynamics and interpretability. NeurIPs, 2017.
  49. 49.Razzhigaev, A., Mikhalchuk, M., Goncharova, E., Oseledets, I., Dimitrov, D., and Kuznetsov, A. The shape of learning: Anisotropy and intrinsic dimensions in transformer-based models. EACL, 2024.
  50. 50.Rényi, A. On measures of entropy and information. Proceedings of the fourth Berkeley symposium on mathematical statistics and probability, 1961.
  51. 51.Roy, O. and Vetterli, M. The effective rank: A measure of effective dimensionality. European signal processing conference, 2007.
  52. 52.Saponati, M., Sager, P., Aceituno, P. V., Stadelmann, T., and Grewe, B. The underlying structures of self-attention: symmetry, directionality, and emergent dynamics in transformer training. arXiv, 2025.
  53. 53.Scholkopf, B. and Smola, A. J. Learning with kernels: support vector machines, regularization, optimization, and beyond. MIT press, 2018.
  54. 54.Shwartz-Ziv, R. Information flow in deep neural networks. PhD thesis, Hebrew University, 2022.
  55. 55.Shwartz-Ziv, R. and Tishby, N. Opening the black box of deep neural networks via information. Entropy, 2019.
  56. 56.Shwartz-Ziv, R., Balestriero, R., Kawaguchi, K., Rudner, T. G., and LeCun, Y. An information theory perspective on variance-invariance-covariance regularization. NeurIPs, 2023.
  57. 57.Skean, O., Osorio, J. K. H., Brockmeier, A. J., and Giraldo, L. G. S. DiME: Maximizing mutual information by a difference of matrix-based entropies. arXiv, 2023.
  58. 58.Skean, O., Dhakal, A., Jacobs, N., and Giraldo, L. G. S. FroSSL: Frobenius norm minimization for self-supervised learning. ECCV, 2024.
  59. 59.Sorscher, B., Ganguli, S., and Sompolinsky, H. Neural representational geometry underlies few-shot concept learning. Proceedings of the National Academy of Sciences, 2022.
  60. 60.Tenney, I., Das, D., and Pavlick, E. BERT rediscovers the classical nlp pipeline. NAACL, 2019.
  61. 61.Thilak, V., Huang, C., Saremi, O., Dinh, L., Goh, H., Nakkiran, P., Susskind, J. M., and Littwin, E. LiDAR: Sensing linear probing performance in joint embedding ssl architectures. ICLR, 2024.
  62. 62.Tian, Y., Krishnan, D., and Isola, P. Contrastive multiview coding. ECCV, 2020.
  63. 63.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv, 2023.
  64. 64.Valeriani, L., Doimo, D., Cuturello, F., Laio, A., Ansuini, A., and Cazzaniga, A. The geometry of hidden representations of large transformer models. NeurIPs, 2023.
  65. 65.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I. Attention is all you need. NeurIPs, 2017.
  66. 66.Voita, E., Sennrich, R., and Titov, I. The bottom-up evolution of representations in the transformer: A study with machine translation and language modeling objectives. EMNLP-IJCNLP, 2019.
  67. 67.Wei, L., Tan, Z., Li, C., Wang, J., and Huang, W. DiffeRank: A novel rank-based metric for evaluating large language models. NeurIPs, 2024.
  68. 68.Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M. Efficient streaming language models with attention sinks. ICLR, 2024.
  69. 69.Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D., Huang, F., Wei, H., Lin, H., Yang, J., Tu, J., Zhang, J., Yang, J., Yang, J., Zhou, J., Lin, J., Dang, K., Lu, K., Bao, K., Yang, K., Yu, L., Li, M., Xue, M., Zhang, P., Zhu, Q., Men, R., Lin, R., Li, T., Xia, T., Ren, X., Ren, X., Fan, Y., Su, Y., Zhang, Y., Wan, Y., Liu, Y., Cui, Z., Zhang, Z., and Qiu, Z. Qwen2.5 technical report. arXiv, 2024.
  70. 70.Zhao, Z., Ziser, Y., and Cohen, S. B. Layer by layer: Uncovering where multi-task learning happens in instruction-tuned large language models. EMNLP-IJCNLP, 2024.
  71. 71.Zhouyin, Z. and Liu, D. Understanding neural networks with logarithm determinant entropy estimator. arXiv, 2021.

Citation

MLA
Skean, O., et al. “Layer by Layer: Uncovering Hidden Representations in Language Models”. arXiv, 2025, http://arxiv.org/abs/2502.02013v2.
APA
Skean, O., Arefin, M. R., Zhao, D., Patel, N., Naghiyev, J., LeCun, Y., & Shwartz-Ziv, R. (2025). Layer by Layer: Uncovering Hidden Representations in Language Models. arXiv. http://arxiv.org/abs/2502.02013v2
Chicago
Skean, O., M. R. Arefin, D. Zhao, et al. 2025. “Layer by Layer: Uncovering Hidden Representations in Language Models”. arXiv. http://arxiv.org/abs/2502.02013v2.
Harvard
Skean, O. et al. (2025) “Layer by Layer: Uncovering Hidden Representations in Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2502.02013v2.
Vancouver
1. Skean O, Arefin MR, Zhao D, Patel N, Naghiyev J, LeCun Y, Shwartz-Ziv R (2025) Layer by Layer: Uncovering Hidden Representations in Language Models. arXiv

BibTeX

@article{skean2025layer,
  title = {Layer by Layer: Uncovering Hidden Representations in Language Models},
  author = {Skean, Oscar and Arefin, Md Rifat and Zhao, Dan and Patel, Niket and Naghiyev, Jalal and LeCun, Yann and Shwartz-Ziv, Ravid},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2502.02013v2},
  eprint = {2502.02013}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/