keyword
representation similarity
Representation similarity is a measure of the degree to which different neural networks, model layers, or biological neural systems encode and structure information in comparable ways within their internal feature spaces. Rather than requiring direct one-to-one matching between individual artificial or biological neurons, representation similarity evaluates the geometric, relational, or statistical alignment among latent activations elicited by a common set of inputs. Standard analytical frameworks, including centered kernel alignment, canonical correlation analysis, and representational similarity analysis, determine whether relative distances and relational structures among data points are preserved across different models. This property is widely applied in machine learning and cognitive neuroscience to assess how closely different architectures, training runs, or learning stages converge on shared representations, to monitor feature consistency during sequential learning, and to examine the universality of learned computational circuits.
2 items

A Toy Model of Universality: Reverse Engineering how Networks Learn Group Operations
Bilal Chughtai, Lawrence Chan, Neel Nanda
Why you should read this
Reverse-engineers how neural networks learn finite group composition using mathematical representation theory, revealing that while networks share an underlying algorithmic family, the specific circuits they develop remain arbitrary.
Universality is a key hypothesis in mechanistic interpretability – that different models learn similar features and circuits when trained on similar tasks. In this work, we study the universality hypothesis by examining how small neural networks learn to implement group composition. We present a novel algorithm by which neural networks may implement composition for any finite group via mathematical representation theory. We then show that networks consistently learn this algorithm by reverse engineering model logits and weights, and confirm our understanding using ablations. By studying networks of differing architectures trained on various groups, we find mixed evidence for universality: using our algorithm, we can completely characterize the family of circuits and features that networks learn on this task, but for a given network the precise circuits learned – as well as the order they develop – are arbitrary.
Added
2026-10-01

Endpoints Weight Fusion for Class Incremental Semantic Segmentation
Jia-Wen Xiao, Chang-Bin Zhang, Jiekang Feng, Xialei Liu, Joost van de Weijer, Ming-Ming Cheng
Why you should read this
Proposes a dynamic parameter-averaging strategy that fuses starting and ending network weights from each incremental step to mitigate catastrophic forgetting in semantic segmentation without increasing model size or requiring extra training.
Class incremental semantic segmentation (CISS) focuses on alleviating catastrophic forgetting to improve discrimination. Previous work mainly exploits regularization (e.g., knowledge distillation) to maintain previous knowledge in the current model. However, distillation alone often yields limited gain to the model since only the representations of old and new models are restricted to be consistent. In this paper, we propose a simple yet effective method to obtain a model with a strong memory of old knowledge, named Endpoints Weight Fusion (EWF). In our method, the model containing old knowledge is fused with the model retaining new knowledge in a dynamic fusion manner, strengthening the memory of old classes in ever-changing distributions. In addition, we analyze the relationship between our fusion strategy and a popular moving average technique EMA, which reveals why our method is more suitable for class-incremental learning. To facilitate parameter fusion with closer distance in the parameter space, we use distillation to enhance the optimization process. Furthermore, we conduct experiments on two widely used datasets, achieving state-of-the-art performance.
Added
2026-09-26
