Cluster analysis for gene expression data: a survey
Daxin JiangChun TangAidong Zhang
Categorizes clustering methods for microarray gene expression data into gene-based, sample-based, and subspace approaches while reviewing specific algorithms, proximity measures, and validation techniques to guide functional genomics research.
Modern DNA microarray technology enables the parallel measurement of expression levels for tens of thousands of genes across various biological conditions, generating massive datasets of millions of measurements. Interpreting this high-dimensional and noisy data is essential for understanding gene functions, biological pathways, and disease classifications, such as cancer subtypes. The article surveys cluster analysis methods applied to gene expression data, detailing their operational challenges, technical frameworks, and validation techniques.
The analysis systematically organizes clustering tasks into three main categories: gene-based clustering, sample-based clustering, and subspace clustering (biclustering). Gene-based approaches group coexpressed genes to identify shared biological functions and regulatory mechanisms, utilizing algorithms ranging from conventional partition and hierarchical methods to specialized graph-theoretical models (such as CLICK and CAST) and density-based techniques (such as DHC). Sample-based approaches aim to classify patient or tissue phenotypes, which requires addressing high gene dimensionality and low signal-to-noise ratios where informative genes account for less than 10 percent of the total dataset. Subspace clustering captures complex biological mechanisms where specific gene subsets are active only within particular sample subsets, using heuristic search models to address the computationally difficult combinatorial problem.
The key findings indicate that standard distance measures, such as Euclidean distance after data standardization, perform equivalently to Pearson's correlation coefficient, though rank-based correlation often loses critical quantitative information. Furthermore, traditional global clustering algorithms frequently struggle with the severe background noise, high interconnectedness, and embedded structures common in genomic data. In sample-based clustering, supervised selection of informative genes provides high diagnostic accuracy, whereas unsupervised sample classification requires iterative filtering to progressively isolate phenotype-relevant genes from overwhelming noise. Additionally, no single algorithm universally outperforms others across all datasets; algorithm effectiveness depends heavily on the underlying data distribution and experimental context.
These findings imply that applying unsuitable clustering methods or failing to isolate informative genes poses substantial risks of producing false positives or misleading biological conclusions. For decision-makers and research leaders, the choice of computational tools directly affects the cost, accuracy, and reliability of discovering diagnostic biomarkers and therapeutic targets. Robust cluster validation—evaluating internal homogeneity and separation, agreement with known biological reference standards, and statistical reliability via cross-validation and predictive strength—is essential before relying on identified patterns for clinical or pharmaceutical pipelines.
The article recommends adopting specialized algorithms tailored to specific data properties, such as using graph- or density-based methods when cluster counts are unknown and noise is high, and employing iterative feature filtering for sample classification. Moving forward, analytics platforms should transition from rigid, fully unsupervised clustering toward interactive exploration tools that integrate prior biological knowledge and provide scalable visualizations adaptable to varying levels of cluster detail. Readers should note that current methodologies carry uncertainties arising from experimental noise, unverified parametric distribution assumptions, and the lack of a universally superior validation metric, warranting careful cross-validation of results.
- Paper: Consensus Clustering: A Resampling-Based Method for Class Discovery and Visualization of Gene Expression Microarray Data, Stefano Monti et al. (2003). Introduces consensus clustering and resampling techniques specifically tailored to discover classes and assess cluster stability in gene expression microarray data.
- Paper: Automatic subspace clustering of high dimensional data for data mining applications, Rakesh Agrawal et al. (1998). Pioneers subspace clustering on high-dimensional data, a foundational methodology surveyed for identifying gene patterns across subsets of conditions.
- Paper: Co-clustering documents and words using bipartite spectral graph partitioning, Inderjit S. Dhillon (2001). Establishes bipartite spectral graph partitioning for simultaneous co-clustering of rows and columns, a core paradigm discussed in gene expression analysis.
- Paper: On Spectral Clustering: Analysis and an algorithm, Andrew Y. Ng et al. (2001). Provides the foundational mathematical framework and normalized graph Laplacian formulation for spectral clustering applied to complex geometric data distributions.
- Paper: Support Vector Clustering, Asa Ben-Hur et al. (2002). Presents support vector clustering for delineating non-spherical clusters in high-dimensional feature spaces with noise handling.
- Paper: Performance Evaluation of Some Clustering Algorithms and Validity Indices, Ujjwal Maulik et al. (2002). Evaluates foundational cluster validity indices and grouping performance metrics that underpin the survey's discussion of cluster validation.
- Paper: A Density-Based Algorithm for Discovering Clusters in Large Spatial Databases with Noise, Martin Ester et al. (1996). Introduces DBSCAN, the standard density-based clustering algorithm that the survey categorizes and contrasts with partitional and hierarchical methods.
- Paper: CURE: an efficient clustering algorithm for large databases, Sudipto Guha et al. (1998). Develops the CURE hierarchical clustering algorithm utilizing multiple representative points to handle non-spherical shapes and outliers in large datasets.
- Paper: Cluster Ensembles – A Knowledge Reuse Framework for Combining Multiple Partitions, Alexander Strehl et al. (2002). Formulates the cluster ensemble framework for combining multiple clusterings into a robust consensus partition.
- Paper: Gene Selection for Cancer Classification using Support Vector Machines, ISABELLE GUYON et al. (2002). Demonstrates feature and gene selection using support vector machine elimination on cancer microarray benchmarks, establishing key preprocessing practices.
- Paper: Information Theoretic Measures for Clusterings Comparison: Variants, Properties, Normalization and Correction for Chance, X. Nguyen et al. (2010). Establishes rigorous information-theoretic metrics with correction for chance to evaluate and compare clusterings, tested directly on genomic datasets.
- Paper: V-Measure: A Conditional Entropy-Based External Cluster Evaluation Measure, Andrew Rosenberg et al. (2007). Introduces the V-Measure external cluster evaluation metric combining homogeneity and completeness, advancing the validation methods reviewed in the survey.
- Paper: Orthogonal nonnegative matrix t-factorizations for clustering, C. Ding et al. (2006). Extends matrix factorization models by using orthogonal non-negative matrix tri-factorization for simultaneous row-column co-clustering.
- Paper: A tutorial on spectral clustering, Ulrike von Luxburg (2007). Provides a comprehensive mathematical tutorial unifying spectral clustering theory, graph Laplacians, and practical implementation.
- Paper: Clustering with Bregman Divergences, Arindam Banerjee et al. (2005). Generalizes centroid-based and mixture-model clustering across the entire family of Bregman divergences.
- Paper: Unsupervised Deep Embedding for Clustering Analysis, Junyuan Xie et al. (2015). Advances high-dimensional clustering by jointly optimizing deep neural network feature representations and cluster centroids in an unsupervised framework.
- Paper: Toward integrating feature selection algorithms for classification and clustering, Huan Liu et al. (2005). Synthesizes feature selection strategies across classification and clustering into a unified categorization framework.
- Paper: Community detection in graphs, Santo Fortunato (2009). Surveys graph community detection algorithms and validation criteria for discovering functional modules across biological and complex networks.
