Built independently by an author, for readers. Read the story and support ChapterPal

keyword

consensus clustering

Consensus clustering is an unsupervised machine learning approach that combines multiple distinct clustering solutions into a single consolidated partition of a dataset. In this framework, diverse clusterings are generated by applying different algorithms, varying parameter settings or random initializations, utilizing alternative feature subsets, or repeatedly perturbing the data through resampling techniques. By evaluating the agreement across these varied partitions—often by calculating how frequently pairs of data points are grouped together—a consensus function determines the final unified clustering. This aggregation helps overcome the sensitivity and bias of individual algorithms, improving overall clustering robustness, validating cluster stability, and assisting in the discovery of the most reliable number of underlying classes.

5 items

Information Theoretic Measures for Clusterings Comparison: Variants, Properties, Normalization and Correction for Chance

Information Theoretic Measures for Clusterings Comparison: Variants, Properties, Normalization and Correction for Chance

X. Nguyen, Julien Epps, James Bailey

OrganizationsCSIRO’s Data61University of MelbourneUniversity of New South Wales

Why you should read this

Establishes which information-theoretic clustering comparison measures satisfy metric, normalization, and chance-correction properties, motivating normalized information distance as a principled default.

Information theoretic measures form a fundamental class of measures for comparing clusterings, and have recently received increasing interest. Nevertheless, a number of questions concerning their properties and inter-relationships remain unresolved. In this paper, we perform an organized study of information theoretic measures for clustering comparison, including several existing popular measures in the literature, as well as some newly proposed ones. We discuss and prove their important properties, such as the metric property and the normalization property. We then highlight to the clustering community the importance of correcting information theoretic measures for chance, especially when the data size is small compared to the number of clusters present therein. Of the available information theoretic based measures, we advocate the normalized information distance (NID) as a general measure of choice, for it possesses concurrently several important properties, such as being both a metric and a normalized measure, admitting an exact analytical adjusted-for-chance form, and using the nominal [0, 1] range better than other normalized variants.

Added

2026-09-14

Cluster Ensembles – A Knowledge Reuse Framework for Combining Multiple Partitions

Cluster Ensembles – A Knowledge Reuse Framework for Combining Multiple Partitions

Alexander Strehl, Joydeep Ghosh

OrganizationsUniversity of Texas at Austin

Why you should read this

Proposes a framework for cluster ensembles that combines multiple data partitions without accessing raw features, introducing three graph- and similarity-based consensus algorithms optimized through shared mutual information.

This paper introduces the problem of combining multiple partitionings of a set of objects into a single consolidated clustering without accessing the features or algorithms that determined these partitionings. We first identify several application scenarios for the resultant ‘knowledge reuse’ framework that we call cluster ensembles. The cluster ensemble problem is then formalized as a combinatorial optimization problem in terms of shared mutual information. In addition to a direct maximization approach, we propose three effective and efficient techniques for obtaining high-quality combiners (consensus functions). The first combiner induces a similarity measure from the partitionings and then reclusters the objects. The second combiner is based on hypergraph partitioning. The third one collapses groups of clusters into meta-clusters which then compete for each object to determine the combined clustering. Due to the low computational costs of our techniques, it is quite feasible to use a supra-consensus function that evaluates all three approaches against the objective function and picks the best solution for a given situation. We evaluate the effectiveness of cluster ensembles in three qualitatively different application scenarios: (i) where the original clusters were formed based on non-identical sets of features, (ii) where the original clustering algorithms worked on non-identical sets of objects, and (iii) where a common data-set is used and the main purpose of combining multiple clusterings is to improve the quality and robustness of the solution. Promising results are obtained in all three situations for synthetic as well as real data-sets.

Added

2026-09-10

Creative Commons License