Built independently by an author, for readers. Read the story and support ChapterPal

keyword

continual learning

Continual learning is a machine learning paradigm in which an artificial intelligence model sequentially acquires new knowledge, tasks, or capabilities over time without forgetting what it previously learned. Also referred to as lifelong learning or incremental learning, this approach aims to solve the stability-plasticity dilemma by balancing the plasticity needed to integrate new information with the stability required to prevent catastrophic forgetting of past skills. Unlike conventional machine learning methods that assume fixed datasets and require complete retraining when new data arrives, continual learning allows systems to adapt dynamically to continuous data streams and shifting distributions, reducing computational retraining costs and enabling more efficient long-term adaptation.

70 items

Fine-tuned Language Models are Continual Learners

Fine-tuned Language Models are Continual Learners

Thomas Scialom, Tuhin Chakrabarty, Smaranda Muresan

OrganizationsColumbia UniversityMeta

Why you should read this

Demonstrates that instruction-tuned language models can sequentially acquire new generation tasks with minimal rehearsal while preventing catastrophic forgetting, identifying self-supervised pre-training as the primary driver of this capability.

Recent work on large language models relies on the intuition that most natural language processing tasks can be described via natural language instructions and that models trained on these instructions show strong zero-shot performance on several standard datasets. However, these models even though impressive still perform poorly on a wide range of tasks outside of their respective training and evaluation sets. To address this limitation, we argue that a model should be able to keep extending its knowledge and abilities, without forgetting previous skills. In spite of the limited success of Continual Learning we show that Fine-tuned Language Models can be continual learners. We empirically investigate the reason for this success and conclude that Continual Learning emerges from self-supervision pre-training. Our resulting model Continual-T0 (CT0) is able to learn 8 new diverse language generation tasks, while still maintaining good performance on previous tasks, spanning in total 70 datasets. Finally, we show that CT0 is able to combine instructions in ways it was never trained for, demonstrating some level of instruction compositionality.^1

Added

2026-10-05

Dataset Distillation using Neural Feature Regression

Dataset Distillation using Neural Feature Regression

Yongchao Zhou, Ehsan Nezhadarya, Jimmy Ba

Why you should read this

Proposes an efficient dataset distillation algorithm that trains the final linear layer to convergence via kernel ridge regression across a pool of feature extractors, cutting training time by two orders of magnitude while scaling effectively to ImageNet-1K.

Dataset distillation aims to learn a small synthetic dataset that preserves most of the information from the original dataset. Dataset distillation can be formulated as a bi-level meta-learning problem where the outer loop optimizes the meta-dataset and the inner loop trains a model on the distilled data. Meta-gradient computation is one of the key challenges in this formulation, as differentiating through the inner loop learning procedure introduces significant computation and memory costs. In this paper, we address these challenges using neural Feature Regression with Pooling (FRePo), achieving the state-of-the-art performance with an order of magnitude less memory requirement and two orders of magnitude faster training than previous methods. The proposed algorithm is analogous to truncated backpropagation through time with a pool of models to alleviate various types of overfitting in dataset distillation. FRePo significantly outperforms the previous methods on CIFAR100, Tiny ImageNet, and ImageNet-1K. Furthermore, we show that high-quality distilled data can greatly improve various downstream applications, such as continual learning and membership inference defense. Please check out our webpage at https://sites.google.com/view/frepo.

Added

2026-10-05

Generalized Neural Collapse for a Large Number of Classes

Generalized Neural Collapse for a Large Number of Classes

Jiachen Jiang, Jinxin Zhou, Peng Wang, Qing Qu, Dustin G. Mixon, Chong You, Zhihui Zhu

OrganizationsGoogleThe Ohio State UniversityUniversity of Michigan

Why you should read this

Extends neural collapse theory to regimes where the number of classes exceeds the feature dimension by introducing softmax codes to characterize the geometric structure of learned representations, providing both theoretical guarantees and empirical validation for high-cardinality classification tasks.

Neural collapse provides an elegant mathematical characterization of learned last-layer representations, also known as features, and classifier weights within deep classification models. The result not only provides insights into deep models but also catalyzes the development of new techniques for improving them. However, most of the existing empirical and theoretical studies into neural collapse center around scenarios where the number of classes is small relative to the dimensionality of the feature space. This paper introduces a generalization of neural collapse to encompass scenarios where the number of classes surpasses the dimension of feature space, which broadly occurs for language models, information retrieval systems, and face recognition applications. A key technical contribution is the introduction of the concept of softmax code, defined as a collection of points that maximizes the minimum one-vs-rest margin, to describe the arrangement of class-mean features. We provide empirical study to verify the prevalence of generalized neural collapse in practical deep neural networks. Moreover, we provide theoretical study to show that the generalized neural collapse provably occurs under an unconstrained feature model with spherical constraint, subject to specific technical conditions on feature dimension and the number of classes.

Added

2026-10-05

Temporal-Difference Variational Continual Learning

Temporal-Difference Variational Continual Learning

Luckeciano Carvalho Melo, Alessandro Abate, Yarin Gal

OrganizationsUniversity of Oxford

Why you should read this

Introduces a temporal-difference-inspired variational continual learning objective that regularizes model updates using multiple past posterior estimates to prevent compounding approximation errors and reduce catastrophic forgetting.

Machine Learning models in real-world applications must continuously learn new tasks to adapt to shifts in the data-generating distribution. Yet, for Continual Learning (CL), models often struggle to balance learning new tasks (plasticity) with retaining previous knowledge (memory stability). Consequently, they are susceptible to Catastrophic Forgetting, which degrades performance and undermines the reliability of deployed systems. In the Bayesian CL literature, variational methods tackle this challenge by employing a learning objective that recursively updates the posterior distribution while constraining it to stay close to its previous estimate. Nonetheless, we argue that these methods may be ineffective due to compounding approximation errors over successive recursions. To mitigate this, we propose new learning objectives that integrate the regularization effects of multiple previous posterior estimations, preventing individual errors from dominating future posterior updates and compounding over time. We reveal insightful connections between these objectives and Temporal-Difference methods, a popular learning mechanism in Reinforcement Learning and Neuroscience. Experiments on challenging CL benchmarks show that our approach effectively mitigates Catastrophic Forgetting, outperforming strong Variational CL methods.

Added

2026-10-05

Meta-Learning Online Adaptation of Language Models

Meta-Learning Online Adaptation of Language Models

Nathan Hu, Eric Mitchell, Christopher D. Manning, Chelsea Finn

OrganizationsStanford University

Why you should read this

Introduces Context-aware Meta-learned Loss Scaling (CaMeLS), a meta-learning method that trains a lightweight model to dynamically upweight informative tokens during online document streams, substantially boosting factual knowledge uptake in language models over standard fine-tuning.

Large language models encode impressively broad world knowledge in their parameters. However, the knowledge in static language models falls out of date, limiting the model’s effective “shelf life.” While online fine-tuning can reduce this degradation, we find that naively fine-tuning on a stream of documents leads to a low level of information uptake. We hypothesize that online fine-tuning does not sufficiently attend to important information. That is, the gradient signal from important tokens representing factual information is drowned out by the gradient from inherently noisy tokens, suggesting that a dynamic, context-aware learning rate may be beneficial. We therefore propose learning which tokens to upweight. We meta-train a small, autoregressive model to reweight the language modeling loss for each token during online fine-tuning, with the objective of maximizing the out-of-date base question-answering model’s ability to answer questions about a document after a single weighted gradient step. We call this approach Context-aware Meta-learned Loss Scaling (CaMeLS). Across three different distributions of documents, our experiments find that CaMeLS provides substantially improved information uptake on streams of thousands of documents compared with standard fine-tuning and baseline heuristics for reweighting token losses.

Added

2026-10-04

Towards Teachable Reasoning Systems: Using a Dynamic Memory of User Feedback for Continual System Improvement

Towards Teachable Reasoning Systems: Using a Dynamic Memory of User Feedback for Continual System Improvement

Bhavana Dalvi Mishra, Oyvind Tafjord, Peter Clark

OrganizationsAllen Institute for AI

Why you should read this

Presents TeachMe, a teachable question-answering architecture that leverages dynamic memory to store user-corrected beliefs from faithful reasoning explanations, enabling frozen language models to continually improve performance on unseen questions without retraining.

Our goal is a teachable reasoning system for question-answering (QA), where a user can interact with faithful answer explanations, and correct its errors so that the system improves over time. Our approach is to augment a QA model with a dynamic memory of user feedback, containing user-supplied corrections to erroneous model beliefs that users identify during interaction. Retrievals from memory are used as additional context for QA, to help avoid previous mistakes in similar new situations - a novel application of memory-based continual learning. With simulated feedback, we find that our system (called TeachMe¹) continually improves with time, and without model retraining, requiring feedback on only 25% of training examples to reach within 1% of the upper-bound (feedback on all examples). Similarly, in experiments with real users, we observe a similar trend, with performance improving by over 15% on a hidden test set after teaching. This suggests new opportunities for using frozen language models in an interactive setting where users can inspect, debug, and correct the model's beliefs, leading to improved system's performance over time.

Added

2026-10-03

Model Tailor: Mitigating Catastrophic Forgetting in Multi-modal Large Language Models

Model Tailor: Mitigating Catastrophic Forgetting in Multi-modal Large Language Models

Didi Zhu, Zhongyi Sun, Zexi Li, Tao Shen, Ke Yan, Shouhong Ding, Chao Wu, Kun Kuang

OrganizationsTencentZhejiang University

Why you should read this

Proposes Model Tailor, a parameter-efficient post-training method that updates fewer than ten percent of model parameters through sparse masking and Hessian-based compensation, preventing catastrophic forgetting in multi-modal large language models while maintaining performance on both original and target tasks.

Catastrophic forgetting emerges as a critical challenge when fine-tuning multi-modal large language models (MLLMs), where improving performance on target tasks often leads to a significant performance drop on the original tasks. This paper presents a comprehensive analysis of catastrophic forgetting in MLLMs and introduces a post-training adjustment method called Model Tailor. Our method primarily preserves the pre-trained parameters while replacing a small number (≤ 10%) of fine-tuned parameters, maintaining ~ 99% effectiveness on original tasks versus pre-training, and achieving ~ 97% on new tasks compared to standard fine-tuning. Specifically, we derive a sparse mask to identify the “model patch”, based on a fusion strategy that integrates salience and sensitivity analysis. Subsequently, a compensation mechanism is introduced to “decorate the patch”, enhancing the model’s performance on both target and original tasks. Additionally, our method is adaptable to multi-task scenarios. Through extensive experiments on Instruct-BLIP and LLaVA-1.5 in both image captioning and visual question answering tasks, our approach demonstrates significant task adaptability while preserving inherent pre-trained capabilities.

Added

2026-10-03

Finetuning with Sampling: SFT Learns Better Than You Think

Finetuning with Sampling: SFT Learns Better Than You Think

Aayush Karan, Sitan Chen, Yilun Du

OrganizationsHarvard University

Why you should read this

Proposes an MCMC sampling algorithm that transforms off-policy expert trajectories toward a model's on-policy distribution, enabling supervised finetuning to match reinforcement learning in task generalization while mitigating catastrophic forgetting.

Introducing new capabilities to frontier models has long been the goal of posttraining, which predominantly employs supervised finetuning (SFT) and reinforcement learning (RL) to this end. Conventional wisdom dictates that RL enables strong generalization on new tasks without losing existing capabilities, while SFT is prone to weak generalization and catastrophic forgetting. At the same time, SFT can learn from off-policy expert data, whereas RL must rely on a model's ability to find successful trajectories with repeated sampling. In our work, we seek to leverage the strength of on-policy learning while utilizing the privileged information contained in off-policy data. However, rather than modifying the learning objective to accommodate this data, we instead tailor the data distribution to better suit the learner. We introduce a Markov chain Monte Carlo (MCMC) sampling algorithm that progressively transforms off-policy traces to be more on-policy given a reference model for finetuning. Across tasks like scientific skill acquisition, mathematical reasoning, and open-ended expertise, our sampling algorithm enables SFT to rival prevailing posttraining techniques, often generalizing better and forgetting less than strong on-policy baselines. In addition, the resulting finetuned models exhibit strong distributional performance and are capable of learning beyond sharpening the base model distribution. At a higher level, our approach presents sampling as a model-native operator that shapes data for learnability, offering broader utility as a general-purpose primitive throughout the posttraining stack.

Added

2026-10-03

Creative Commons License
Overcoming Catastrophic Forgetting During Domain Adaptation of Seq2seq Language Generation

Overcoming Catastrophic Forgetting During Domain Adaptation of Seq2seq Language Generation

Dingcheng Li, Zheng Chen, Eunah Cho, Jie Hao, Xiaohu Liu, Fan Xing, Chenlei Guo, Yang Liu

OrganizationsAmazon

Why you should read this

Proposes a framework combining adaptive parameter regularization with embedding-space domain drift estimation to prevent catastrophic forgetting in sequential sequence-to-sequence language generation without storing past task data.

Seq2seq language generation models that are trained offline with multiple domains in a sequential fashion often suffer from catastrophic forgetting. Lifelong learning has been proposed to handle this problem. However, existing work such as experience replay or elastic weighted consolidation requires incremental memory space. In this work, we propose an innovative framework, RMR_DSE that leverages a recall optimization mechanism to selectively memorize important parameters of previous tasks via regularization, and uses a domain drift estimation algorithm to compensate for the drift between different domains in the embedding space. These designs enable the model to be trained on the current task while keeping the memory of previous tasks, and avoid much additional data storage. Furthermore, RMR_DSE can be combined with existing lifelong learning approaches. Our experiments on two seq2seq language generation tasks, paraphrase and dialog response generation, show that RMR_DSE outperforms state-of-the-art models by a considerable margin and greatly reduces forgetting.

Added

2026-10-03

Wide Neural Networks Forget Less Catastrophically

Wide Neural Networks Forget Less Catastrophically

Seyed-Iman Mirzadeh, Arslan Chaudhry, Dong Yin, Huiyi Hu, Razvan Pascanu, Dilan Görür, Mehrdad Farajtabar

OrganizationsGoogleWashington State University

Why you should read this

Demonstrates that increasing network width significantly reduces catastrophic forgetting in continual learning—even matching the benefits of replay buffers—and explains this phenomenon through gradient orthogonality, activation sparsity, and the lazy training regime.

A primary focus area in continual learning research is alleviating the “catastrophic forgetting” problem in neural networks by designing new algorithms that are more robust to the distribution shifts. While the recent progress in continual learning literature is encouraging, our understanding of what properties of neural networks contribute to catastrophic forgetting is still limited. To address this, instead of focusing on continual learning algorithms, in this work, we focus on the model itself and study the impact of “width” of the neural network architecture on catastrophic forgetting, and show that width has a surprisingly significant effect on forgetting. To explain this effect, we study the learning dynamics of the network from various perspectives such as gradient orthogonality, sparsity, and lazy training regime. We provide potential explanations that are consistent with the empirical results across different architectures and continual learning benchmarks.

Added

2026-10-02

LLaMA Pro: Progressive LLaMA with Block Expansion

LLaMA Pro: Progressive LLaMA with Block Expansion

Chengyue Wu, Yukang Gan, Yixiao Ge, Zeyu Lu, Jiahao Wang, Ye Feng, Ying Shan, Ping Luo

OrganizationsBeijing Language and Culture UniversityShanghai Jiao Tong UniversityTencentUniversity of Hong Kong

Why you should read this

Proposes a block-expansion post-pretraining method that adds and tunes zero-initialized Transformer blocks on domain-specific data while freezing the base model, enabling large language models to master specialized coding and math skills without suffering from catastrophic forgetting of their general capabilities.

Humans generally acquire new skills without compromising the old; however, the opposite holds for Large Language Models (LLMs), e.g., from LLaMA to CodeLLaMA. To this end, we propose a new post-pretraining method for LLMs with an expansion of Transformer blocks. We tune the expanded blocks using only new corpus, efficiently and effectively improving the model’s knowledge while mitigating forgetting. In this paper, we experiment on the corpus of code and math, yielding LLaMA Pro-8.3B, a versatile foundation model initialized from LLaMA2-7B, excelling in general tasks, programming, and mathematics. LLaMA Pro and its instruction-following counterpart (LLaMA Pro - Instruct) achieve advanced performance among various benchmarks, demonstrating superiority over existing open models in the LLaMA family and the immense potential of reasoning and addressing diverse tasks as an intelligent agent. Our findings provide valuable insights into integrating natural and programming languages, laying a solid foundation for developing advanced language agents that operate effectively in various environments.

Added

2026-10-01

A Survey of Zero-shot Generalisation in Deep Reinforcement Learning

A Survey of Zero-shot Generalisation in Deep Reinforcement Learning

Robert Kirk, Amy Zhang, Edward Grefenstette, Tim Rocktäschel

OrganizationsMetaUniversity College LondonUniversity of California Berkeley

Why you should read this

Presents a unifying mathematical formalism and taxonomy for zero-shot generalization in deep reinforcement learning, critically evaluating existing benchmarks and methods to guide the development of policies that successfully transfer to unseen environments.

The study of zero-shot generalisation (ZSG) in deep Reinforcement Learning (RL) aims to produce RL algorithms whose policies generalise well to novel unseen situations at deployment time, avoiding overfitting to their training environments. Tackling this is vital if we are to deploy reinforcement learning algorithms in real world scenarios, where the environment will be diverse, dynamic and unpredictable. This survey is an overview of this nascent field. We rely on a unifying formalism and terminology for discussing different ZSG problems, building upon previous works. We go on to categorise existing benchmarks for ZSG, as well as current methods for tackling these problems. Finally, we provide a critical discussion of the current state of the field, including recommendations for future work. Among other conclusions, we argue that taking a purely procedural content generation approach to benchmark design is not conducive to progress in ZSG, we suggest fast online adaptation and tackling RL-specific problems as some areas for future work on methods for ZSG, and we recommend building benchmarks in underexplored problem settings such as offline RL ZSG and reward-function variation.

Added

2026-09-28

Dense Network Expansion for Class Incremental Learning

Dense Network Expansion for Class Incremental Learning

Zhiyuan Hu, Yunsheng Li, Jiancheng Lyu, Dashan Gao, Nuno Vasconcelos

OrganizationsMicrosoftQualcommUniversity of California, San Diego

Why you should read this

Proposes a dense network expansion method for class incremental learning that transfers knowledge across frozen task experts using a cross-task attention block, achieving superior accuracy while significantly curbing parameter growth.

The problem of class incremental learning (CIL) is considered. State-of-the-art approaches use a dynamic architecture based on network expansion (NE), in which a task expert is added per task. While effective from a computational standpoint, these methods lead to models that grow quickly with the number of tasks. A new NE method, dense network expansion (DNE), is proposed to achieve a better trade-off between accuracy and model complexity. This is accomplished by the introduction of dense connections between the intermediate layers of the task expert networks, that enable the transfer of knowledge from old to new tasks via feature sharing and reusing. This sharing is implemented with a cross-task attention mechanism, based on a new task attention block (TAB), that fuses information across tasks. Unlike traditional attention mechanisms, TAB operates at the level of the feature mixing and is decoupled with spatial attentions. This is shown more effective than a joint spatial-and-task attention for CIL. The proposed DNE approach can strictly maintain the feature space of old classes while growing the network and feature scale at a much slower rate than previous methods. In result, it outperforms the previous SOTA methods by a margin of 4% in terms of accuracy, with similar or even smaller model scale.

Added

2026-09-26

Remember the Past: Distilling Datasets into Addressable Memories for Neural Networks

Remember the Past: Distilling Datasets into Addressable Memories for Neural Networks

Zhiwei Deng, Olga Russakovsky

OrganizationsPrinceton University

Why you should read this

Proposes a dataset distillation framework that stores shared memory bases combined via learned addressing functions, breaking the linear scaling bottleneck with class count and substantially outperforming prior distillation and continual learning baselines.

We propose an algorithm that compresses the critical information of a large dataset into compact addressable memories. These memories can then be recalled to quickly re-train a neural network and recover the performance (instead of storing and re-training on the full original dataset). Building upon the dataset distillation framework, we make a key observation that a shared common representation allows for more efficient and effective distillation. Concretely, we learn a set of bases (aka “memories”) which are shared between classes and combined through learned flexible addressing functions to generate a diverse set of training examples. This leads to several benefits: 1) the size of compressed data does not necessarily grow linearly with the number of classes; 2) an overall higher compression rate with more effective distillation is achieved; and 3) more generalized queries are allowed beyond recalling the original classes. We demonstrate state-of-the-art results on the dataset distillation task across six benchmarks, including up to 16.5% and 9.7% in retained accuracy improvement when distilling CIFAR10 and CIFAR100 respectively. We then leverage our framework to perform continual learning, achieving state-of-the-art results on four benchmarks, with 23.2% accuracy improvement on MANY. The code is released on our project webpage¹.

Added

2026-09-26

Self-Sustaining Representation Expansion for Non-Exemplar Class-Incremental Learning

Self-Sustaining Representation Expansion for Non-Exemplar Class-Incremental Learning

Kai Zhu, Wei Zhai, Yang Cao, Jiebo Luo, Zhengjun Zha

OrganizationsInstitute of Artificial Intelligence, Hefei Comprehensive National Science CenterUniversity of RochesterUniversity of Science and Technology of China

Why you should read this

Develops a self-sustaining representation expansion framework for non-exemplar class-incremental learning that prevents catastrophic forgetting and parameter explosion by dynamically reorganizing network structures and selectively distilling knowledge using new class prototypes.

Non-exemplar class-incremental learning is to recognize both the old and new classes when old class samples cannot be saved. It is a challenging task since representation optimization and feature retention can only be achieved under supervision from new classes. To address this problem, we propose a novel self-sustaining representation expansion scheme. Our scheme consists of a structure reorganization strategy that fuses main-branch expansion and side-branch updating to maintain the old features, and a main-branch distillation scheme to transfer the invariant knowledge. Furthermore, a prototype selection mechanism is proposed to enhance the discrimination between the old and new classes by selectively incorporating new samples into the distillation process. Extensive experiments on three benchmarks demonstrate significant incremental performance, outperforming the state-of-the-art methods by a margin of 3%, 3% and 6%, respectively.

Added

2026-09-26

Robust and Resource-Efficient Data-Free Knowledge Distillation by Generative Pseudo Replay

Robust and Resource-Efficient Data-Free Knowledge Distillation by Generative Pseudo Replay

Kuluhan Binici, Shivam Aggarwal, Nam Trung Pham, Karianto Leman, Tulika Mitra

OrganizationsInstitute for Infocomm ResearchNational University of Singapore

Why you should read this

Presents a data-free knowledge distillation framework that uses a variational autoencoder to generate replay samples, preventing student accuracy degradation over training epochs without storing past synthetic data in memory.

Data-Free Knowledge Distillation (KD) allows knowledge transfer from a trained neural network (teacher) to a more compact one (student) in the absence of original training data. Existing works use a validation set to monitor the accuracy of the student over real data and report the highest performance throughout the entire process. However, validation data may not be available at distillation time either, making it infeasible to record the student snapshot that achieved the peak accuracy. Therefore, a practical data-free KD method should be robust and ideally provide monotonically increasing student accuracy during distillation. This is challenging because the student experiences knowledge degradation due to the distribution shift of the synthetic data. A straightforward approach to overcome this issue is to store and rehearse the generated samples periodically, which increases the memory footprint and creates privacy concerns. We propose to model the distribution of the previously observed synthetic samples with a generative network. In particular, we design a Variational Autoencoder (VAE) with a training objective that is customized to learn the synthetic data representations optimally. The student is rehearsed by the generative pseudo replay technique, with samples produced by the VAE. Hence knowledge degradation can be prevented without storing any samples. Experiments on image classification benchmarks show that our method optimizes the expected value of the distilled model accuracy while eliminating the large memory overhead incurred by the sample-storing methods.

Added

2026-09-26

Continual Sequence Generation with Adaptive Compositional Modules

Continual Sequence Generation with Adaptive Compositional Modules

Yanzhe Zhang, Xuezhi Wang, Diyi Yang

OrganizationsGeorgia Institute of TechnologyGoogle

Why you should read this

Proposes an adaptive modular framework for continual sequence generation that dynamically decides whether to reuse existing transformer adapters or insert new ones based on task similarity, preventing catastrophic forgetting while maximizing parameter efficiency.

Continual learning is essential for real-world deployment when there is a need to quickly adapt the model to new tasks without forgetting knowledge of old tasks. Existing work on continual sequence generation either always reuses existing parameters to learn new tasks, which is vulnerable to catastrophic forgetting on dissimilar tasks, or blindly adds new parameters for every new task, which could prevent knowledge sharing between similar tasks. To get the best of both worlds, in this work, we propose continual sequence generation with adaptive compositional modules to adaptively add modules in transformer architectures and compose both old and new modules for new tasks. We also incorporate pseudo experience replay to facilitate knowledge transfer in those shared modules. Experiment results on various sequences of generation tasks show that our framework can adaptively add modules or reuse modules based on task similarity, outperforming state-of-the-art baselines in terms of both performance and parameter efficiency. We make our code public at https://github.com/GT-SALT/Adaptive-Compositional-Modules.

Added

2026-09-26

Continual Test-Time Domain Adaptation

Continual Test-Time Domain Adaptation

Qin Wang, Olga Fink, Luc Van Gool, Dengxin Dai

OrganizationsÉcole Polytechnique Fédérale de LausanneETH ZurichKU LeuvenMax Planck Institute for Informatics

Why you should read this

Introduces CoTTA, a continual test-time adaptation framework that prevents error accumulation and catastrophic forgetting in dynamically shifting target domains through averaged predictions and stochastic weight restoration.

Test-time domain adaptation aims to adapt a source pre-trained model to a target domain without using any source data. Existing works mainly consider the case where the target domain is static. However, real-world machine perception systems are running in non-stationary and continually changing environments where the target domain distribution can change over time. Existing methods, which are mostly based on self-training and entropy regularization, can suffer from these non-stationary environments. Due to the distribution shift over time in the target domain, pseudo-labels become unreliable. The noisy pseudo-labels can further lead to error accumulation and catastrophic forgetting. To tackle these issues, we propose a continual test-time adaptation approach~(CoTTA) which comprises two parts. Firstly, we propose to reduce the error accumulation by using weight-averaged and augmentation-averaged predictions which are often more accurate. On the other hand, to avoid catastrophic forgetting, we propose to stochastically restore a small part of the neurons to the source pre-trained weights during each iteration to help preserve source knowledge in the long-term. The proposed method enables the long-term adaptation for all parameters in the network. CoTTA is easy to implement and can be readily incorporated in off-the-shelf pre-trained models. We demonstrate the effectiveness of our approach on four classification tasks and a segmentation task for continual test-time adaptation, on which we outperform existing methods. Our code is available at \url{this https URL}.

Added

2026-09-26

Towards Modular LLMs by Building and Reusing a Library of LoRAs

Towards Modular LLMs by Building and Reusing a Library of LoRAs

Oleksiy Ostapenko, Zhan Su, Edoardo M. Ponti, Laurent Charlin, Nicolas Le Roux, Lucas Caccia, Alessandro Sordoni

OrganizationsCIFARHEC MontréalMicrosoftMila – Québec Artificial Intelligence InstituteUniversité de MontréalUniversity of CopenhagenUniversity of Edinburgh

Why you should read this

Proposes a modular framework that clusters LoRA adapters by parameter similarity and dynamically routes hidden states to the most relevant adapters at inference time, achieving superior zero-shot and supervised generalization without requiring joint retraining.

Given the increasing number of parameter-efficient adapters of large language models (LLMs), how can we reuse them to improve LLM performance on new tasks? We study how to best build a library of adapters given multi-task data and devise techniques for both zero-shot and supervised task generalization through routing in such library. We benchmark existing approaches to build this library and introduce model-based clustering, MBC, a method that groups tasks based on the similarity of their adapter parameters, indirectly optimizing for transfer across tasks. In order to reuse the library, we present a novel zero-shot routing mechanism, Arrow, which enables dynamic selection of the most relevant adapters for new inputs without the need for retraining. We experiment with several LLMs, such as Phi-2 and Mistral, on a wide array of held-out tasks, verifying that MBC-based adapters and Arrow routing lead to superior generalization to new tasks. Thus, we make steps towards creating modular, adaptable LLMs that can match or outperform traditional joint training.

Added

2026-09-26