keyword
similarity matrices
A similarity matrix is a square array of numbers that quantifies the pairwise resemblance or closeness between a set of objects. In such a matrix, each row and column corresponds to an individual item in a dataset, and the numerical entry at the intersection of a row and a column indicates how similar the corresponding pair of items is, typically computed using distance metrics, correlation coefficients, or kernel functions. These matrices are usually symmetric, with larger values denoting greater similarity and the main diagonal reflecting each item's maximal self-similarity. In data analysis and machine learning, similarity matrices provide a standard structured representation of relational data and are widely utilized in algorithms for spectral clustering, dimensionality reduction, graph-based learning, and comparing internal data representations across different models.
2 items

Co-regularized Multi-view Spectral Clustering
Abhishek Kumar, Piyush Rai, Hal Daumé
Why you should read this
Proposes a multi-view spectral clustering framework that co-regularizes graph Laplacians across different data representations to find consistent cluster assignments across diverse views.
In many clustering problems, we have access to multiple views of the data each of which could be individually used for clustering. Exploiting information from multiple views, one can hope to find a clustering that is more accurate than the ones obtained using the individual views. Often these different views admit same underlying clustering of the data, so we can approach this problem by looking for clusterings that are consistent across the views, i.e., corresponding data points in each view should have same cluster membership. We propose a spectral clustering framework that achieves this goal by co-regularizing the clustering hypotheses, and propose two co-regularization schemes to accomplish this. Experimental comparisons with a number of baselines on two synthetic and three real-world datasets establish the efficacy of our proposed approaches.
Added
2026-09-25

Similarity of Neural Network Representations Revisited
Simon Kornblith, Mohammad Norouzi, Honglak Lee, Geoffrey E. Hinton
Why you should read this
Establishes centered kernel alignment (CKA) as a reliable method for comparing hidden representations across differently initialized neural networks while exposing the mathematical limitations of canonical correlation analysis.
Recent work has sought to understand the behavior of neural networks by comparing representations between layers and between different trained models. We examine methods for comparing neural network representations based on canonical correlation analysis (CCA). We show that CCA belongs to a family of statistics for measuring multivariate similarity, but that neither CCA nor any other statistic that is invariant to invertible linear transformation can measure meaningful similarities between representations of higher dimension than the number of data points. We introduce a similarity index that measures the relationship between representational similarity matrices and does not suffer from this limitation. This similarity index is equivalent to centered kernel alignment (CKA) and is also closely connected to CCA. Unlike CCA, CKA can reliably identify correspondences between representations in networks trained from different initializations.
Added
2026-09-14
