Built independently by an author, for readers. Read the story and support ChapterPal

keyword

test-time training

Test-time training is a machine learning paradigm in which a pre-trained model updates its parameters or internal representation during inference using the test inputs themselves. Instead of keeping model weights completely fixed after the initial training phase, test-time training optimizes the model on the fly using self-supervised learning objectives, consistency criteria, or available unlabelled test samples. This enables the system to adapt dynamically to distribution shifts, changing environments, novel tasks, or extended context sequences without requiring ground-truth labels or access to the original training dataset.

12 items

Robust Test-Time Adaptation in Dynamic Scenarios

Robust Test-Time Adaptation in Dynamic Scenarios

Longhui Yuan, Binhui Xie, Shuang Li

Why you should read this

Proposes a practical test-time adaptation framework that combines an exponential moving average normalization scheme, a timeliness- and uncertainty-aware memory bank, and a teacher-student model to prevent model collapse on continually shifting, temporally correlated test streams.

Test-time adaptation (TTA) intends to adapt the pre-trained model to test distributions with only unlabeled test data streams. Most of the previous TTA methods have achieved great success on simple test data streams such as independently sampled data from single or multiple distributions. However, these attempts may fail in dynamic scenarios of real-world applications like autonomous driving, where the environments gradually change and the test data is sampled correlatively over time. In this work, we explore such practical test data streams to deploy the model on the fly, namely practical test-time adaptation (PTTA). To do so, we elaborate a Robust Test-Time Adaptation (RoTTA) method against the complex data stream in PTTA. More specifically, we present a robust batch normalization scheme to estimate the normalization statistics. Meanwhile, a memory bank is utilized to sample category-balanced data with consideration of timeliness and uncertainty. Further, to stabilize the training procedure, we develop a time-aware reweighting strategy with a teacher-student model. Extensive experiments prove that RoTTA enables continual test-time adaptation on the correlatively sampled data streams. Our method is easy to implement, making it a good choice for rapid deployment. The code is publicly available at https://github.com/BIT-DA/RoTTA

Added

2026-10-05

Parameter-free Online Test-time Adaptation

Parameter-free Online Test-time Adaptation

Malik Boudiaf, Romain Müller, Ismail Ben Ayed, Luca Bertinetto

Why you should read this

Proposes Laplacian Adjusted Maximum-likelihood Estimation (LAME), a parameter-free online test-time adaptation method that corrects output predictions instead of updating network weights, preventing catastrophic failure under unpredictable domain shifts while cutting memory usage and inference latency in half.

Training state-of-the-art vision models has become pro-hibitively expensive for researchers and practitioners. For the sake of accessibility and resource reuse, it is important to focus on adapting these models to a variety of down-stream scenarios. An interesting and practical paradigm is online test-time adaptation, according to which training data is inaccessible, no labelled data from the test distribution is available, and adaptation can only happen at test time and on a handful of samples. In this paper, we investigate how test-time adaptation methods fare for a number of pre-trained models on a variety of real-world scenarios, significantly extending the way they have been originally evaluated. We show that they perform well only in narrowly-defined experimental setups and sometimes fail catastrophically when their hyperparameters are not selected for the same scenario in which they are being tested. Motivated by the inherent uncertainty around the conditions that will ultimately be encountered at test time, we propose a particularly “conservative” approach, which addresses the problem with a Laplacian Adjusted Maximum-likelihood Estimation (LAME) objective. By adapting the model’s output (not its parameters), and solving our objective with an efficient concave-convex procedure, our approach exhibits a much higher average accuracy across scenarios than existing methods, while being notably faster and have a much lower memory footprint. The code is available at https://github.com/fiveai/LAME.

Added

2026-10-05

Test-Time Training Can Close the Natural Distribution Shift Performance Gap in Deep Learning Based Compressed Sensing

Test-Time Training Can Close the Natural Distribution Shift Performance Gap in Deep Learning Based Compressed Sensing

Mohammad Zalbagi Darestani, Jiayu Liu, Reinhard Heckel

OrganizationsRice UniversityTechnical University of Munich

Why you should read this

Demonstrates that combining self-supervised pre-training with per-sample test-time training closes up to 99% of the distribution shift performance gap in deep learning-based accelerated MRI reconstruction across diverse anatomies, datasets, and acquisition parameters.

Deep learning based image reconstruction methods outperform traditional methods. However, neural networks suffer from a performance drop when applied to images from a different distribution than the training images. For example, a model trained for reconstructing knees in accelerated magnetic resonance imaging (MRI) does not reconstruct brains well, even though the same network trained on brains reconstructs brains perfectly well. Thus there is a distribution shift performance gap for a given neural network, defined as the difference in performance when training on a distribution P and training on another distribution Q, and evaluating both models on Q. In this work, we propose a domain adaptation method for deep learning based compressive sensing that relies on self-supervision during training paired with test-time training at inference. We show that for four natural distribution shifts, this method essentially closes the distribution shift performance gap for state-of-the-art architectures for accelerated MRI.

Added

2026-10-03

Learning to (Learn at Test Time): RNNs with Expressive Hidden States

Learning to (Learn at Test Time): RNNs with Expressive Hidden States

Yu Sun, Xinhao Li, Karan Dalal, Jiarui Xu, Arjun Vikram, Genghan Zhang, Yann Dubois, Xinlei Chen, Xiaolong Wang, Sanmi Koyejo, Tatsunori Hashimoto, Carlos Guestrin

OrganizationsMetaStanford UniversityUniversity of California BerkeleyUniversity of California, San Diego

Why you should read this

Introduces Test-Time Training layers that treat RNN hidden states as internal machine learning models updated via self-supervised gradient steps, achieving linear-time sequence modeling that continues to improve across long contexts where existing architectures plateau.

Self-attention performs well in long context but has quadratic complexity. Existing RNN layers have linear complexity, but their performance in long context is limited by the expressive power of their hidden states. We present a practical framework for instantiating sequence modeling layers with linear complexity and expressive hidden states. The key idea is to make the hidden state a machine learning model itself, and the update rule a step of self-supervised learning. Since the hidden state is updated by training even on test sequences, our layers are called Test-Time Training (TTT) layers. We consider two instantiations: TTT-Linear and TTT-MLP, whose hidden state is a linear model and a two-layer MLP respectively. We evaluate our instantiations at the scale of 125M to 1.3B parameters, comparing with a strong Transformer and Mamba, a modern RNN. Similar to Transformer, TTT-Linear and TTT-MLP can keep reducing perplexity by conditioning on more tokens, while Mamba cannot after 16k context. TTT-MLP still faces challenges in memory I/O, but shows larger potential in long context, pointing to a promising direction for future research.

Added

2026-09-26

Continual Test-Time Domain Adaptation

Continual Test-Time Domain Adaptation

Qin Wang, Olga Fink, Luc Van Gool, Dengxin Dai

OrganizationsÉcole Polytechnique Fédérale de LausanneETH ZurichKU LeuvenMax Planck Institute for Informatics

Why you should read this

Introduces CoTTA, a continual test-time adaptation framework that prevents error accumulation and catastrophic forgetting in dynamically shifting target domains through averaged predictions and stochastic weight restoration.

Test-time domain adaptation aims to adapt a source pre-trained model to a target domain without using any source data. Existing works mainly consider the case where the target domain is static. However, real-world machine perception systems are running in non-stationary and continually changing environments where the target domain distribution can change over time. Existing methods, which are mostly based on self-training and entropy regularization, can suffer from these non-stationary environments. Due to the distribution shift over time in the target domain, pseudo-labels become unreliable. The noisy pseudo-labels can further lead to error accumulation and catastrophic forgetting. To tackle these issues, we propose a continual test-time adaptation approach~(CoTTA) which comprises two parts. Firstly, we propose to reduce the error accumulation by using weight-averaged and augmentation-averaged predictions which are often more accurate. On the other hand, to avoid catastrophic forgetting, we propose to stochastically restore a small part of the neurons to the source pre-trained weights during each iteration to help preserve source knowledge in the long-term. The proposed method enables the long-term adaptation for all parameters in the network. CoTTA is easy to implement and can be readily incorporated in off-the-shelf pre-trained models. We demonstrate the effectiveness of our approach on four classification tasks and a segmentation task for continual test-time adaptation, on which we outperform existing methods. Our code is available at \url{this https URL}.

Added

2026-09-26

The Surprising Effectiveness of Test-Time Training for Few-Shot Learning

The Surprising Effectiveness of Test-Time Training for Few-Shot Learning

Ekin Akyrek, Mehul Damani, Adam Zweiger, Linlu Qiu, Han Guo, Jyothish Pari, Yoon Kim, Jacob Andreas

OrganizationsMassachusetts Institute of Technology

Why you should read this

Demonstrates that updating language model parameters at inference time using losses derived from in-context examples dramatically improves reasoning performance on out-of-distribution tasks, achieving human-level accuracy on the Abstraction and Reasoning Corpus.

Language models (LMs) have shown impressive performance on tasks within their training distribution, but often struggle with structurally novel tasks even when given a small number of in-context task examples. We investigate the effectiveness of test-time training (TTT)—temporarily updating model parameters during inference using a loss derived from in-context examples—as a mechanism for improving LMs’ reasoning and few-shot learning capabilities. On the Abstraction and Reasoning Corpus (ARC), performing TTT with in-context examples yields up to 6× higher accuracy compared to fine-tuned baselines—reaching 53.0% on the public validation set with an 8B-parameter LM and 61.9% when ensembled with program-synthesis methods, matching average human performance. On BIG-Bench Hard (BBH), TTT on in-context examples surpasses standard few-shot prompting in the 10-shot setting by 7.3 percentage points (50.5% to 57.8%). Our findings highlight the limitations of in-context learning for novel tasks and demonstrate the potential of test-time training to enhance language model adaptability.

Added

2026-09-26

Robust Mean Teacher for Continual and Gradual Test-Time Adaptation

Robust Mean Teacher for Continual and Gradual Test-Time Adaptation

Mario Döbler, Robert A. Marsden, Bin Yang

OrganizationsUniversity of Stuttgart

Why you should read this

Proposes a continual test-time adaptation framework that uses symmetric cross-entropy and multi-view contrastive learning within a mean-teacher architecture to prevent error accumulation across shifting data distributions.

Since experiencing domain shifts during test-time is inevitable in practice, test-time adaption (TTA) continues to adapt the model after deployment. Recently, the area of continual and gradual test-time adaptation (TTA) emerged. In contrast to standard TTA, continual TTA considers not only a single domain shift, but a sequence of shifts. Gradual TTA further exploits the property that some shifts evolve gradually over time. Since in both settings long test sequences are present, error accumulation needs to be addressed for methods relying on self-training. In this work, we propose and show that in the setting of TTA, the symmetric cross-entropy is better suited as a consistency loss for mean teachers compared to the commonly used cross-entropy. This is justified by our analysis with respect to the (symmetric) cross-entropy's gradient properties. To pull the test feature space closer to the source domain, where the pre-trained model is well posed, contrastive learning is leveraged. Since applications differ in their requirements, we address several settings, including having source data available and the more challenging source-free setting. We demonstrate the effectiveness of our proposed method “robust mean teacher“ (RMT) on the continual and gradual corruption benchmarks CIFAR10C, CIFAR100C, and Imagenet-C. We further consider ImageNet-R and propose a new continual DomainNet-126 benchmark. State-of-the-art results are achieved on all benchmarks.

Added

2026-09-26

Fast Weight Attention for Continual Learning

Fast Weight Attention for Continual Learning

Yifan Zhang, Steve Ta, Jasper Zhang, Jichen Feng, Shuzhen Li, Yongxin Zhang, Yifeng Liu, Huizhuo Yuan, Mengdi Wang, Quanquan Gu, Andrew Chi-Chih Yao

OrganizationsByteDanceHyperbolic LabsPrinceton UniversityTsinghua UniversityUniversity of California, Los Angeles

Why you should read this

Proposes a family of fast-weight attention mechanisms that reformulate recurrent sequence transitions as normalized online learning rules with parallelizable implementations, improving length extrapolation without the memory overhead of standard key-value caching.

Recurrent fast-weight memories and selective state-space models compress an expanding context into a fixed-size recurrent state, making the state transition an online learning rule. We study this rule under read-after-write autoregressive semantics. For the prefix-prediction objective considered here, the local fast-memory example revealed at step tt is the prefix-aligned pair (xt,yt)=(ϕ(kt−1),vt)(\mathbf{x}_t,\mathbf{y}_t)=(\phi(\mathbf{k}_{t-1}),\mathbf{v}_t). The common same-step association (ϕ(kt),vt)(\phi(\mathbf{k}_t),\mathbf{v}_t) remains causal, but optimizes a different internal objective. We derive normalized first-order updates for squared-error regression and negative inner-product objectives. The regression family comprises Falcon-1 (a scalar NLMS update), Falcon-2 (its per-column extension), and Falcon-3 (a sliding-window mini-batch update); Falcon-1A/Falcon-2A/Falcon-3A are the corresponding inner-product variants. We provide recurrent, masked-parallel, and chunk-parallel forms, together with numerically stable positive-decay renormalization. Representative variants remain competitive in language modeling and improve length extrapolation on variable-digit addition. This framework separates temporal alignment, plasticity, forgetting, and bounded rehearsal in recurrent sequence models.

Added

2026-09-02

License

Published with permission

Test-Time Training with KV Binding Is Secretly Linear Attention

Test-Time Training with KV Binding Is Secretly Linear Attention

Junchen Liu, Sven Elflein, Or Litany, Zan Gojcic, Ruilong Li

OrganizationsNVIDIATechnion – Israel Institute of TechnologyUniversity of TorontoVector Institute

Why you should read this

Shows that test-time training with key-value binding functions as learned linear attention rather than online memorization, resolving empirical anomalies while enabling fully parallel, computationally efficient implementations without loss of performance.

Test-time training (TTT) with KV binding as sequence modeling layer is commonly interpreted as a form of online meta-learning that memorizes a key-value mapping at test time. However, our analysis reveals multiple phenomena that contradict this memorization-based interpretation. Motivated by these findings, we revisit the formulation of TTT and show that a broad class of TTT architectures can be expressed as a form of learned linear attention operator. Beyond explaining previously puzzling model behaviors, this perspective yields multiple practical benefits: it enables principled architectural simplifications, admits fully parallel formulations that preserve performance while improving efficiency, and provides a systematic reduction of diverse TTT variants to a standard linear attention form. Overall, our results reframe TTT not as test-time memorization, but as learned linear attention with enhanced representational capacity. Project page: this https URL.

Added

2026-08-30

Creative Commons License
Dynamic Compression in Recurrent Networks

Dynamic Compression in Recurrent Networks

Jyothish Pari, Ryan Bahlous-Boldi, Pulkit Agrawal

OrganizationsImprobable AI LabMassachusetts Institute of Technology

Why you should read this

Proposes dynamic compression, an approach allowing recurrent neural networks to selectively re-scan raw input history and revise their fixed-size hidden state on demand, establishing a computation-memory tradeoff that drastically reduces state capacity requirements for long-context task reuse.

Recurrent models process long contexts efficiently by compressing their history into a fixed-size state, but modern architectures typically do so in a single causal pass over the sequence. Each input must therefore be compressed before the model knows how it will later be used, forcing a limited state to compromise across possible future demands. We introduce dynamic compression, which allows a recurrent model to selectively revisit past tokens and revise its fixed-size state through additional recurrent updates. The model need not preserve every part of the history at uniformly high fidelity in its recurrent state, because lower-fidelity information can be revisited from the retained raw sequence when it becomes relevant. We study this in a controlled setting where the model first learns multiple functions in-context and, later in the same sequence, encounters a series of few-shot tasks that each require it to identify and reuse one of those functions. A single-pass model must preserve every function at sufficient fidelity for any future task, whereas selective re-scanning allows the model to revisit and refine only the function currently needed. We find that dynamic compression substantially reduces the recurrent state required for accurate reuse and scales more favorably as the number of stored functions grows. These results demonstrate a computation--memory tradeoff in which recurrent models can spend more computation revisiting their history to make more effective use of a fixed-size state.

Added

2026-08-25

Creative Commons License