keyword
consensus clustering
Consensus clustering is an unsupervised machine learning approach that combines multiple distinct clustering solutions into a single consolidated partition of a dataset. In this framework, diverse clusterings are generated by applying different algorithms, varying parameter settings or random initializations, utilizing alternative feature subsets, or repeatedly perturbing the data through resampling techniques. By evaluating the agreement across these varied partitions—often by calculating how frequently pairs of data points are grouped together—a consensus function determines the final unified clustering. This aggregation helps overcome the sensitivity and bias of individual algorithms, improving overall clustering robustness, validating cluster stability, and assisting in the discovery of the most reliable number of underlying classes.
5 items

Co-regularized Multi-view Spectral Clustering
Abhishek Kumar, Piyush Rai, Hal Daumé
Why you should read this
Proposes a multi-view spectral clustering framework that co-regularizes graph Laplacians across different data representations to find consistent cluster assignments across diverse views.
In many clustering problems, we have access to multiple views of the data each of which could be individually used for clustering. Exploiting information from multiple views, one can hope to find a clustering that is more accurate than the ones obtained using the individual views. Often these different views admit same underlying clustering of the data, so we can approach this problem by looking for clusterings that are consistent across the views, i.e., corresponding data points in each view should have same cluster membership. We propose a spectral clustering framework that achieves this goal by co-regularizing the clustering hypotheses, and propose two co-regularization schemes to accomplish this. Experimental comparisons with a number of baselines on two synthetic and three real-world datasets establish the efficacy of our proposed approaches.
Added
2026-09-25

Community detection in networks: A user guide
Santo Fortunato, Darko Hric
Why you should read this
Clarifies the foundations of network community detection by evaluating the strengths and limitations of popular clustering algorithms, dispelling widespread methodological misconceptions, and providing practical criteria for performance validation.
Community detection in networks is one of the most popular topics of modern network science. Communities, or clusters, are usually groups of vertices having higher probability of being connected to each other than to members of other groups, though other patterns are possible. Identifying communities is an ill-defined problem. There are no universal protocols on the fundamental ingredients, like the definition of community itself, nor on other crucial issues, like the validation of algorithms and the comparison of their performances. This has generated a number of confusions and misconceptions, which undermine the progress in the field. We offer a guided tour through the main aspects of the problem. We also point out strengths and weaknesses of popular methods, and give directions to their use.
Added
2026-09-16

Consensus Clustering: A Resampling-Based Method for Class Discovery and Visualization of Gene Expression Microarray Data
Stefano Monti, Pablo Tamayo, Jill Mesirov, Todd Golub
In this paper we present a new methodology of class discovery and clustering validation tailored to the task of analyzing gene expression data. The method can be best thought of as an analysis approach, to guide and assist in the use of any of a wide range of available clustering algorithms. We call the new methodology consensus clustering, and in conjunction with resampling techniques, it provides for a method to represent the consensus across multiple runs of a clustering algorithm and to assess the stability of the discovered clusters. The method can also be used to represent the consensus over multiple runs of a clustering algorithm with random restart (such as K-means, model-based Bayesian clustering, SOM, etc.), so as to account for its sensitivity to the initial conditions. Finally, it provides for a visualization tool to inspect cluster number, membership, and boundaries. We present the results of our experiments on both simulated data and real gene expression data aimed at evaluating the effectiveness of the methodology in discovering biologically meaningful clusters.
Added
2026-09-16

Information Theoretic Measures for Clusterings Comparison: Variants, Properties, Normalization and Correction for Chance
X. Nguyen, Julien Epps, James Bailey
Why you should read this
Establishes which information-theoretic clustering comparison measures satisfy metric, normalization, and chance-correction properties, motivating normalized information distance as a principled default.
Information theoretic measures form a fundamental class of measures for comparing clusterings, and have recently received increasing interest. Nevertheless, a number of questions concerning their properties and inter-relationships remain unresolved. In this paper, we perform an organized study of information theoretic measures for clustering comparison, including several existing popular measures in the literature, as well as some newly proposed ones. We discuss and prove their important properties, such as the metric property and the normalization property. We then highlight to the clustering community the importance of correcting information theoretic measures for chance, especially when the data size is small compared to the number of clusters present therein. Of the available information theoretic based measures, we advocate the normalized information distance (NID) as a general measure of choice, for it possesses concurrently several important properties, such as being both a metric and a normalized measure, admitting an exact analytical adjusted-for-chance form, and using the nominal [0, 1] range better than other normalized variants.
Added
2026-09-14

Cluster Ensembles – A Knowledge Reuse Framework for Combining Multiple Partitions
Alexander Strehl, Joydeep Ghosh
Why you should read this
Proposes a framework for cluster ensembles that combines multiple data partitions without accessing raw features, introducing three graph- and similarity-based consensus algorithms optimized through shared mutual information.
This paper introduces the problem of combining multiple partitionings of a set of objects into a single consolidated clustering without accessing the features or algorithms that determined these partitionings. We first identify several application scenarios for the resultant ‘knowledge reuse’ framework that we call cluster ensembles. The cluster ensemble problem is then formalized as a combinatorial optimization problem in terms of shared mutual information. In addition to a direct maximization approach, we propose three effective and efficient techniques for obtaining high-quality combiners (consensus functions). The first combiner induces a similarity measure from the partitionings and then reclusters the objects. The second combiner is based on hypergraph partitioning. The third one collapses groups of clusters into meta-clusters which then compete for each object to determine the combined clustering. Due to the low computational costs of our techniques, it is quite feasible to use a supra-consensus function that evaluates all three approaches against the objective function and picks the best solution for a given situation. We evaluate the effectiveness of cluster ensembles in three qualitatively different application scenarios: (i) where the original clusters were formed based on non-identical sets of features, (ii) where the original clustering algorithms worked on non-identical sets of objects, and (iii) where a common data-set is used and the main purpose of combining multiple clusterings is to improve the quality and robustness of the solution. Promising results are obtained in all three situations for synthetic as well as real data-sets.
Added
2026-09-10

