The Dual Form of Neural Networks Revisited: Connecting Test Time Predictions to Training Patterns via Spotlights of Attention
Kazuki IrieRóbert CsordásJürgen Schmidhuber
Reformulates linear layers in gradient-trained neural networks as key-value attention over the entire training history, providing a direct method to trace and interpret how specific training samples drive test-time predictions.
Modern deep neural networks deliver powerful predictions across many domains, but their internal decision-making processes remain largely opaque black boxes. Because models compress vast amounts of training data into static weight matrices, it is difficult to identify which specific training examples drive individual predictions at test time. This lack of transparency presents a growing challenge for accountability, auditability, and trust as artificial intelligence systems are deployed into high-stakes environments.
The article demonstrates how every linear layer in a neural network trained by gradient descent can be exactly reformulated as a key-value memory system using dot-product attention over its entire training history. By adopting this dual formulation, the authors evaluate whether examining attention scores across stored training examples provides a direct, interpretable link between training data and test-time outputs.
To test this approach empirically, the authors recorded layer inputs during standard gradient descent training and evaluated attention weights at test time across multiple setups. They studied small-scale feedforward networks with two hidden layers on MNIST and Fashion-MNIST image classification across single-task, joint multi-task, and sequential continual learning environments. They also evaluated one-layer recurrent language models on small text corpora, including WikiText-2 and Aesop's Fables.
The investigation produced four primary findings. First, while individual top-ranked training examples do not always match the target class, the aggregate sum of attention weights across classes strongly correlates with correct model outputs, with accuracy reaching 84.7% in the final layer for correct predictions compared to roughly 20.6% for incorrect predictions. Second, in joint multi-task training, representations in deeper layers exhibit cross-task attention, drawing on relevant structural features across datasets. Third, the dual view explains catastrophic forgetting in continual learning as a retrieval interference failure; stored memory of earlier tasks remains intact within the network history but is overwhelmed by later training patterns, causing test accuracy on the initial task to plummet from 97% to 45%. Fourth, in language modeling, attention retrieves training passages sharing context, grammatical structure, and semantic concepts rather than superficial token matches.
These findings provide fundamental insights into how neural networks retain and retrieve information. They show that neural networks do not truly lose previous data in their historical formulation, but rather suffer from retrieval bottlenecks. This understanding opens concrete pathways to analyze safety risks, diagnose misclassifications, and explain algorithmic bias by pinpointing influential training examples.
Organizations developing or auditing machine learning models should consider capturing layer activations during training for pilot interpretability and diagnostic workflows where storage allows. In continual learning scenarios, exploring selective retrieval mechanisms or attention masks represents a promising option to mitigate interference without architectural redesigns. Future research should prioritize developing memory-efficient approximations, such as re-computation pipelines during testing, to scale this diagnostic analysis to industrial foundation models.
The findings are derived from exact mathematical formulations, providing high confidence in the underlying duality. However, practical application is constrained by high time and space complexity, as storage requirements scale linearly with the total number of training iterations. In addition, the method requires logging activations during training and cannot be applied retroactively to already-trained models.
- Paper: Transformer Feed-Forward Layers Are Key-Value Memories, Mor Geva et al. (2020). This paper establishes that feed-forward layers in deep models operate as key-value memory retrieval mechanisms, providing the foundational conceptual framing for viewing network layers as associative memories.
- Paper: Understanding Black-box Predictions via Influence Functions, Pang Wei Koh et al. (2017). This work introduces the classic influence function paradigm for attributing test-time predictions to individual training examples, providing essential context for alternative data-attribution formulations.
- Paper: End-To-End Memory Networks, Sainbayar Sukhbaatar et al. (2015). This paper introduces end-to-end memory networks with attention-based addressing, foundational to understanding how neural networks store and retrieve historical data through soft attention.
- Paper: Attention is not Explanation, Sarthak Jain et al. (2019). This work critically analyzes the fidelity of attention mechanisms as interpretability tools, which motivates the source's empirical evaluation of attention-based dual formulations.
- Paper: Understanding intermediate layers using linear classifier probes, Guillaume Alain et al. (2016). This text details the use of linear probes on intermediate representations to track representation evolution, establishing key diagnostic methods used to evaluate internal network states.
- Paper: Overcoming catastrophic forgetting with hard attention to the task, Joan Serrà et al. (2018). This study analyzes catastrophic forgetting in continual learning through attention masking mechanisms, offering crucial background for the source's dual-retrieval interpretation of forgetting.
- Paper: Test-time regression: a unifying framework for designing sequence models with associative memory, Ke Alexander Wang et al. (2025). This work formalizes test-time retrieval and associative recall into a unified test-time regression framework, expanding the view of sequence models and attention as key-value memory systems.
- Paper: Test-Time Training with KV Binding Is Secretly Linear Attention, Junchen Liu et al. (2026). This research proves that test-time training with key-value binding mathematically reduces to linear attention operators, extending the duality between gradient-based training dynamics and attention-based retrieval.
- Paper: Simple linear attention language models balance the recall-throughput tradeoff, Simran Arora et al. (2024). This paper explores the fundamental tradeoffs between inference memory state and associative recall capacity in linear attention models, continuing the investigation of memory bottlenecks in attention-driven systems.
- Paper: Remember the Past: Distilling Datasets into Addressable Memories for Neural Networks, Zhiwei Deng et al. (2022). This paper leverages addressable memory architectures to distill datasets and overcome catastrophic forgetting, directly applying memory-retrieval concepts to lifelong learning problems.
- Paper: Probing Representation Forgetting in Supervised and Unsupervised Continual Learning, MohammadReza Davari et al. (2022). This empirical study proves that intermediate representations survive catastrophic forgetting during continual learning, reinforcing and extending the source's finding that forgetting is primarily a retrieval failure.
