Built independently by an author, for readers. Read the story and support ChapterPal

keyword

Rand index

The Rand index is an external evaluation metric in statistics and machine learning used to measure the similarity between two data clusterings or partitions of the same dataset. It evaluates agreement by considering all possible pairs of data points and calculating the proportion of pairs that are either consistently placed in the same cluster or consistently placed in different clusters across both partitions. The resulting score ranges from 0 to 1, where a value of 1 indicates that the two clusterings are identical and a value of 0 indicates complete disagreement. Because the standard metric does not account for agreement occurring purely by chance, an adjusted version known as the adjusted Rand index is frequently utilized to establish a baseline expected value of zero for random clusterings.

3 items

Information Theoretic Measures for Clusterings Comparison: Variants, Properties, Normalization and Correction for Chance

Information Theoretic Measures for Clusterings Comparison: Variants, Properties, Normalization and Correction for Chance

X. Nguyen, Julien Epps, James Bailey

OrganizationsCSIRO’s Data61University of MelbourneUniversity of New South Wales

Why you should read this

Establishes which information-theoretic clustering comparison measures satisfy metric, normalization, and chance-correction properties, motivating normalized information distance as a principled default.

Information theoretic measures form a fundamental class of measures for comparing clusterings, and have recently received increasing interest. Nevertheless, a number of questions concerning their properties and inter-relationships remain unresolved. In this paper, we perform an organized study of information theoretic measures for clustering comparison, including several existing popular measures in the literature, as well as some newly proposed ones. We discuss and prove their important properties, such as the metric property and the normalization property. We then highlight to the clustering community the importance of correcting information theoretic measures for chance, especially when the data size is small compared to the number of clusters present therein. Of the available information theoretic based measures, we advocate the normalized information distance (NID) as a general measure of choice, for it possesses concurrently several important properties, such as being both a metric and a normalized measure, admitting an exact analytical adjusted-for-chance form, and using the nominal [0, 1] range better than other normalized variants.

Added

2026-09-14