keyword
cluster validation
Cluster validation is the process of quantitatively and qualitatively assessing the quality, correctness, and stability of groupings produced by a clustering algorithm. In unsupervised learning and exploratory data analysis, clustering algorithms can impose artificial partitions even on random data, making validation essential to verify whether genuine patterns exist. The procedure typically involves determining if a dataset has an inherent clustering tendency, selecting the optimal number of clusters, and evaluating the reproducibility and stability of the discovered groups across different algorithmic runs or data perturbations. Validation techniques are generally categorized into internal measures, which evaluate cluster compactness and separation using the intrinsic properties of the data; external measures, which benchmark the clusters against known ground-truth labels or prior domain knowledge; and relative measures, which compare alternative clustering schemes or parameter settings to identify the most appropriate model.
2 items

Cluster analysis for gene expression data: a survey
Daxin Jiang, Chun Tang, Aidong Zhang
Why you should read this
Categorizes clustering methods for microarray gene expression data into gene-based, sample-based, and subspace approaches while reviewing specific algorithms, proximity measures, and validation techniques to guide functional genomics research.
DNA microarray technology has now made it possible to simultaneously monitor the expression levels of thousands of genes during important biological processes and across collections of related samples. Elucidating the patterns hidden in gene expression data offers a tremendous opportunity for an enhanced understanding of functional genomics. However, the large number of genes and the complexity of biological networks greatly increases the challenges of comprehending and interpreting the resulting mass of data, which often consists of millions of measurements. A first step toward addressing this challenge is the use of clustering techniques, which is essential in the data mining process to reveal natural structures and identify interesting patterns in the underlying data. Cluster analysis seeks to partition a given data set into groups based on specified features so that the data points within a group are more similar to each other than the points in different groups. A very rich literature on cluster analysis has developed over the past three decades. Many conventional clustering algorithms have been adapted or directly applied to gene expression data, and also new algorithms have recently been proposed specifically aiming at gene expression data. These clustering algorithms have been proven useful for identifying biologically relevant groups of genes and samples. In this paper, we first briefly introduce the concepts of microarray technology and discuss the basic elements of clustering on gene expression data. In particular, we divide cluster analysis for gene expression data into three categories. Then, we present specific challenges pertinent to each clustering category and introduce several representative approaches. We also discuss the problem of cluster validation in three aspects and review various methods to assess the quality and reliability of clustering results. Finally, we conclude this paper and suggest the promising trends in this field.
Added
2026-09-25

Consensus Clustering: A Resampling-Based Method for Class Discovery and Visualization of Gene Expression Microarray Data
Stefano Monti, Pablo Tamayo, Jill Mesirov, Todd Golub
In this paper we present a new methodology of class discovery and clustering validation tailored to the task of analyzing gene expression data. The method can be best thought of as an analysis approach, to guide and assist in the use of any of a wide range of available clustering algorithms. We call the new methodology consensus clustering, and in conjunction with resampling techniques, it provides for a method to represent the consensus across multiple runs of a clustering algorithm and to assess the stability of the discovered clusters. The method can also be used to represent the consensus over multiple runs of a clustering algorithm with random restart (such as K-means, model-based Bayesian clustering, SOM, etc.), so as to account for its sensitivity to the initial conditions. Finally, it provides for a visualization tool to inspect cluster number, membership, and boundaries. We present the results of our experiments on both simulated data and real gene expression data aimed at evaluating the effectiveness of the methodology in discovering biologically meaningful clusters.
Added
2026-09-16
