keyword
Rand index
The Rand index is an external evaluation metric in statistics and machine learning used to measure the similarity between two data clusterings or partitions of the same dataset. It evaluates agreement by considering all possible pairs of data points and calculating the proportion of pairs that are either consistently placed in the same cluster or consistently placed in different clusters across both partitions. The resulting score ranges from 0 to 1, where a value of 1 indicates that the two clusterings are identical and a value of 0 indicates complete disagreement. Because the standard metric does not account for agreement occurring purely by chance, an adjusted version known as the adjusted Rand index is frequently utilized to establish a baseline expected value of zero for random clusterings.
3 items

Orthogonal nonnegative matrix t-factorizations for clustering
C. Ding, Tao Li, Wei Peng, Haesun Park
Why you should read this
Establishes a rigorous mathematical foundation and convergent update algorithms for orthogonal three-factor nonnegative matrix factorization, enabling simultaneous, interpretable co-clustering of rows and columns in complex data matrices.
Currently, most research on nonnegative matrix factorization (NMF) focus on 2-factor X = FG^T factorization. We provide a systematic analysis of 3-factor X = FSG^T NMF. While unconstrained 3-factor NMF is equivalent to unconstrained 2-factor NMF, constrained 3-factor NMF brings new features to constrained 2-factor NMF. We study the orthogonality constraint because it leads to rigorous clustering interpretation. We provide new rules for updating F,S,G and prove the convergence of these algorithms. Experiments on 5 datasets and a real world case study are performed to show the capability of bi-orthogonal 3-factor NMF on simultaneously clustering rows and columns of the input data matrix. We provide a new approach of evaluating the quality of clustering on words using class aggregate distribution and multi-peak distribution. We also provide an overview of various NMF extensions and examine their relationships.
Added
2026-09-25

V-Measure: A Conditional Entropy-Based External Cluster Evaluation Measure
Andrew Rosenberg, Julia Hirschberg
Why you should read this
Introduces V-measure, an entropy-based external evaluation metric that transparently balances clustering homogeneity and completeness without depending on dataset size or cluster counts.
We present V-measure, an external entropy-based cluster evaluation measure. V-measure provides an elegant solution to many problems that affect previously defined cluster evaluation measures including 1) dependence on clustering algorithm or data set, 2) the “problem of matching”, where the clustering of only a portion of data points are evaluated and 3) accurate evaluation and combination of two desirable aspects of clustering, homogeneity and completeness. We compare V-measure to a number of popular cluster evaluation measures and demonstrate that it satisfies several desirable properties of clustering solutions, using simulated clustering results. Finally, we use V-measure to evaluate two clustering tasks: document clustering and pitch accent type clustering.
Added
2026-09-18

Information Theoretic Measures for Clusterings Comparison: Variants, Properties, Normalization and Correction for Chance
X. Nguyen, Julien Epps, James Bailey
Why you should read this
Establishes which information-theoretic clustering comparison measures satisfy metric, normalization, and chance-correction properties, motivating normalized information distance as a principled default.
Information theoretic measures form a fundamental class of measures for comparing clusterings, and have recently received increasing interest. Nevertheless, a number of questions concerning their properties and inter-relationships remain unresolved. In this paper, we perform an organized study of information theoretic measures for clustering comparison, including several existing popular measures in the literature, as well as some newly proposed ones. We discuss and prove their important properties, such as the metric property and the normalization property. We then highlight to the clustering community the importance of correcting information theoretic measures for chance, especially when the data size is small compared to the number of clusters present therein. Of the available information theoretic based measures, we advocate the normalized information distance (NID) as a general measure of choice, for it possesses concurrently several important properties, such as being both a metric and a normalized measure, admitting an exact analytical adjusted-for-chance form, and using the nominal [0, 1] range better than other normalized variants.
Added
2026-09-14
