Consensus Clustering: A Resampling-Based Method for Class Discovery and Visualization of Gene Expression Microarray Data
Stefano MontiPablo TamayoJill MesirovTodd Golub
- Paper: Cluster Ensembles – A Knowledge Reuse Framework for Combining Multiple Partitions, Alexander Strehl et al. (2002). Establishes the foundational framework for cluster ensembles and combining multiple cluster partitions into a unified consensus representation.
- Paper: X-means: Extending K-means with Efficient Estimation of the Number of Clusters, Dan Pelleg et al. (2000). Presents Bayesian Information Criterion-based model selection to determine the optimal number of clusters, a core problem consensus clustering addresses through resampling stability.
- Paper: Gene Selection for Cancer Classification using Support Vector Machines, ISABELLE GUYON et al. (2002). Introduces key benchmark cancer microarray datasets and class discovery challenges in high-dimensional gene expression data that motivate the consensus clustering methodology.
- Paper: On Spectral Clustering: Analysis and an algorithm, Andrew Y. Ng et al. (2001). Provides fundamental graph-theoretic and spectral clustering foundations that inform partition stability and similarity matrix analysis.
- Paper: Information Theoretic Measures for Clusterings Comparison: Variants, Properties, Normalization and Correction for Chance, X. Nguyen et al. (2010). Develops information-theoretic evaluation metrics and chance corrections essential for rigorously quantifying the agreement between consensus partitions and ground-truth classes.
- Paper: V-Measure: A Conditional Entropy-Based External Cluster Evaluation Measure, Andrew Rosenberg et al. (2007). Introduces the V-Measure to evaluate clustering validity via entropy-based homogeneity and completeness, extending external cluster evaluation beyond consensus stability metrics.
- Paper: Toward integrating feature selection algorithms for classification and clustering, Huan Liu et al. (2005). Integrates feature selection algorithms with clustering to address the high-dimensional noise and redundancy inherent in gene expression and related datasets.
- Paper: Ward’s Hierarchical Agglomerative Clustering Method: Which Algorithms Implement Ward’s Criterion?, Fionn Murtagh et al. (2011). Clarifies the algorithmic implementations and variance criteria of Ward's agglomerative clustering, a method widely evaluated within consensus clustering pipelines.
