Cluster analysis for gene expression data: a survey

Daxin JiangChun TangAidong Zhang

article2004TKDE1,348 citations

Categorizes clustering methods for microarray gene expression data into gene-based, sample-based, and subspace approaches while reviewing specific algorithms, proximity measures, and validation techniques to guide functional genomics research.

Listen

Modern DNA microarray technology enables the parallel measurement of expression levels for tens of thousands of genes across various biological conditions, generating massive datasets of millions of measurements. Interpreting this high-dimensional and noisy data is essential for understanding gene functions, biological pathways, and disease classifications, such as cancer subtypes. The article surveys cluster analysis methods applied to gene expression data, detailing their operational challenges, technical frameworks, and validation techniques.

The analysis systematically organizes clustering tasks into three main categories: gene-based clustering, sample-based clustering, and subspace clustering (biclustering). Gene-based approaches group coexpressed genes to identify shared biological functions and regulatory mechanisms, utilizing algorithms ranging from conventional partition and hierarchical methods to specialized graph-theoretical models (such as CLICK and CAST) and density-based techniques (such as DHC). Sample-based approaches aim to classify patient or tissue phenotypes, which requires addressing high gene dimensionality and low signal-to-noise ratios where informative genes account for less than 10 percent of the total dataset. Subspace clustering captures complex biological mechanisms where specific gene subsets are active only within particular sample subsets, using heuristic search models to address the computationally difficult combinatorial problem.

The key findings indicate that standard distance measures, such as Euclidean distance after data standardization, perform equivalently to Pearson's correlation coefficient, though rank-based correlation often loses critical quantitative information. Furthermore, traditional global clustering algorithms frequently struggle with the severe background noise, high interconnectedness, and embedded structures common in genomic data. In sample-based clustering, supervised selection of informative genes provides high diagnostic accuracy, whereas unsupervised sample classification requires iterative filtering to progressively isolate phenotype-relevant genes from overwhelming noise. Additionally, no single algorithm universally outperforms others across all datasets; algorithm effectiveness depends heavily on the underlying data distribution and experimental context.

These findings imply that applying unsuitable clustering methods or failing to isolate informative genes poses substantial risks of producing false positives or misleading biological conclusions. For decision-makers and research leaders, the choice of computational tools directly affects the cost, accuracy, and reliability of discovering diagnostic biomarkers and therapeutic targets. Robust cluster validation—evaluating internal homogeneity and separation, agreement with known biological reference standards, and statistical reliability via cross-validation and predictive strength—is essential before relying on identified patterns for clinical or pharmaceutical pipelines.

The article recommends adopting specialized algorithms tailored to specific data properties, such as using graph- or density-based methods when cluster counts are unknown and noise is high, and employing iterative feature filtering for sample classification. Moving forward, analytics platforms should transition from rigid, fully unsupervised clustering toward interactive exploration tools that integrate prior biological knowledge and provide scalable visualizations adaptable to varying levels of cluster detail. Readers should note that current methodologies carry uncertainties arising from experimental noise, unverified parametric distribution assumptions, and the lack of a universally superior validation metric, warranting careful cross-validation of results.

Jiang et al (2004).pdf
Cover for Cluster analysis for gene expression data: a survey

Abstract

DNA microarray technology has now made it possible to simultaneously monitor the expression levels of thousands of genes during important biological processes and across collections of related samples. Elucidating the patterns hidden in gene expression data offers a tremendous opportunity for an enhanced understanding of functional genomics. However, the large number of genes and the complexity of biological networks greatly increases the challenges of comprehending and interpreting the resulting mass of data, which often consists of millions of measurements. A first step toward addressing this challenge is the use of clustering techniques, which is essential in the data mining process to reveal natural structures and identify interesting patterns in the underlying data. Cluster analysis seeks to partition a given data set into groups based on specified features so that the data points within a group are more similar to each other than the points in different groups. A very rich literature on cluster analysis has developed over the past three decades. Many conventional clustering algorithms have been adapted or directly applied to gene expression data, and also new algorithms have recently been proposed specifically aiming at gene expression data. These clustering algorithms have been proven useful for identifying biologically relevant groups of genes and samples. In this paper, we first briefly introduce the concepts of microarray technology and discuss the basic elements of clustering on gene expression data. In particular, we divide cluster analysis for gene expression data into three categories. Then, we present specific challenges pertinent to each clustering category and introduce several representative approaches. We also discuss the problem of cluster validation in three aspects and review various methods to assess the quality and reliability of clustering results. Finally, we conclude this paper and suggest the promising trends in this field.

Table of Contents

  • 1 INTRODUCTION
  • 1.1 Introduction to Microarray Technology
  • 1.1.1 Measuring mRNA Levels
  • 1.1.2 Preprocessing of Gene Expression Data
  • 1.1.3 Applications of Clustering Gene Expression Data
  • 1.2 Introduction to Clustering Techniques
  • 1.2.1 Clusters and Clustering
  • 1.2.2 Categories of Gene Expression Data Clustering
  • 1.2.3 Proximity Measurement for Gene Expression Data
  • 2 CLUSTERING ALGORITHMS
  • 2.1 Gene-Based Clustering
  • 2.1.1 Challenges of Gene Clustering
  • 2.1.2 K-Means
  • 2.1.3 Self-Organizing Map
  • 2.1.4 Hierarchical Clustering
  • 2.1.5 Graph-Theoretical Approaches
  • 2.1.6 Model-Based Clustering
  • 2.1.7 A Density-Based Hierarchical Approach: DHC
  • 2.1.8 Summary
  • 2.2 Sample-Based Clustering
  • 2.2.1 Clustering Based on Supervised Informative Gene Selection
  • 2.2.2 Unsupervised Clustering and Informative Gene Selection
  • 2.2.3 Summary
  • 2.3 Subspace Clustering
  • 2.3.1 Coupled Two-Way Clustering (CTWC)
  • 2.3.2 Plaid Model
  • 2.3.3 Biclustering and -Clusters
  • 2.3.4 Summary
  • 3 CLASS VALIDATION
  • 3.1 Homogeneity and Separation
  • 3.2 Agreement with Reference Partition
  • 3.3 Reliability of Clusters
  • 4 CURRENT AND FUTURE RESEARCH DIRECTIONS
  • REFERENCES

Knowls

  1. Knowl 1 — Taxonomy of Gene Expression Clustering: Gene-Based, Sample-Based, and Subspace Clustering

    definition

    Given a gene expression matrix M={wij∣1≤i≤n,1≤j≤m}M = \{w_{ij} \mid 1 \le i \le n, 1 \le j \le m\} representing the measured expression levels of nn genes across mm samples/conditions, cluster analysis is categorized into three distinct paradigms:

    1. Gene-Based Clustering: Treats rows (genes G={g⃗1,…,g⃗n}G = \{\vec{g}_1, \dots, \vec{g}_n\}) as data objects in an mm-dimensional sample feature space. The objective is to identify coexpressed genes that share similar expression patterns across all conditions to infer cofunction, coregulation, and shared transcriptional regulatory networks.

    2. Sample-Based Clustering: Treats columns (samples S={s⃗1,…,s⃗m}S = \{\vec{s}_1, \dots, \vec{s}_m\}) as data objects in an nn-dimensional gene feature space. The goal is to discover macroscopic phenotypes or cellular subtypes (such as cancer classifications or normal vs. diseased tissues). Because the feature space is very high-dimensional (n≫mn \gg m) and fewer than 10%10\% of genes are typically informative for a specific phenotype distinction (a signal-to-noise ratio below 1:101:10), informative gene selection is mandatory.

    3. Subspace Clustering (Biclustering): Treats genes and samples symmetrically, searching for submatrices formed by a subset of genes across a subset of samples (G′×S′⊆G×SG' \times S' \subseteq G \times S). This reflects the biological reality that specific cellular processes involve only a subset of genes active in a subset of physiological conditions or disease states.

  2. Knowl 2 — Relationship Between Pearson Correlation and Standardized Euclidean Distance

    theoretical result

    Let Oi,Oj∈RpO_i, O_j \in \mathbb{R}^p be two numerical expression vectors across pp features. Standardizing each vector to have zero mean and unit variance yields the transformed vector Oi′O'_i with elements:

    Oid′=Oid−μOiσOiO'_{id} = \frac{O_{id} - \mu_{O_i}}{\sigma_{O_i}}

    where μOi=1p∑d=1pOid\mu_{O_i} = \frac{1}{p}\sum_{d=1}^p O_{id} and σOi=1p∑d=1p(Oid−μOi)2\sigma_{O_i} = \sqrt{\frac{1}{p}\sum_{d=1}^p (O_{id} - \mu_{O_i})^2}.

    Pearson's correlation coefficient is invariant to standardization:

    Pearson(Oi,Oj)=Pearson(Oi′,Oj′)\text{Pearson}(O_i, O_j) = \text{Pearson}(O'_i, O'_j)

    Furthermore, the Euclidean distance between the standardized vectors Oi′O'_i and Oj′O'_j is strictly determined by their Pearson correlation coefficient:

    Euclidean(Oi′,Oj′)=2p(1−Pearson(Oi′,Oj′))\text{Euclidean}(O'_i, O'_j) = \sqrt{2p \left(1 - \text{Pearson}(O'_i, O'_j)\right)}

    Consequently, pairwise similarity rankings are preserved: a pair (Oi1,Oj1)(O_{i_1}, O_{j_1}) with a higher Pearson correlation than (Oi2,Oj2)(O_{i_2}, O_{j_2}) has a strictly smaller Euclidean distance between its standardized vectors.

  3. Knowl 3 — Cluster Affinity Search Technique (CAST) Algorithm

    algorithm

    CAST partitions data objects using a corrupted clique graph model, which assumes the true underlying cluster structure is a collection of disjoint cliques perturbed by random edge flips with error probability α\alpha.

    Input: Similarity matrix S∈[0,1]n×nS \in [0, 1]^{n \times n}, affinity threshold t∈[0,1]t \in [0, 1]
    Output: Set of clusters C\mathcal{C}
    U←{1,2,…,n}\mathcal{U} \leftarrow \{1, 2, \dots, n\}
    C←∅\mathcal{C} \leftarrow \emptyset
    while U≠∅\mathcal{U} \ne \emptyset do
        Copen←∅C_{\text{open}} \leftarrow \emptyset
        Pick an unassigned object u∈Uu \in \mathcal{U} and set Copen←{u}C_{\text{open}} \leftarrow \{u\}
        while CopenC_{\text{open}} is not stable do
            for each element x∈U∖Copenx \in \mathcal{U} \setminus C_{\text{open}} do
                a(x)←∑y∈CopenS(x,y)a(x) \leftarrow \sum_{y \in C_{\text{open}}} S(x, y)
                if a(x)≥t⋅∣Copen∣a(x) \ge t \cdot |C_{\text{open}}| then
                    Copen←Copen∪{x}C_{\text{open}} \leftarrow C_{\text{open}} \cup \{x\}
            for each element x∈Copenx \in C_{\text{open}} do
                a(x)←∑y∈CopenS(x,y)a(x) \leftarrow \sum_{y \in C_{\text{open}}} S(x, y)
                if a(x)<t⋅∣Copen∣a(x) < t \cdot |C_{\text{open}}| then
                    Copen←Copen∖{x}C_{\text{open}} \leftarrow C_{\text{open}} \setminus \{x\}
        C←C∪{Copen}\mathcal{C} \leftarrow \mathcal{C} \cup \{C_{\text{open}}\}
        U←U∖Copen\mathcal{U} \leftarrow \mathcal{U} \setminus C_{\text{open}}
    return C\mathcal{C}

    The affinity threshold tt controls the average intra-cluster similarity requirement. CAST does not require specifying the number of clusters in advance and isolates noise as singletons or unclustered objects.

  4. Knowl 4 — Cluster Identification via Connectivity Kernels (CLICK)

    model/method

    CLICK models pairwise similarity values between standardized gene expression profiles as normally distributed random variables. A weighted proximity graph G=(V,E)G=(V,E) is constructed where vertices VV represent genes and edge weights ωij\omega_{ij} denote the posterior probability that objects ii and jj belong to the same true cluster.

    The algorithm clusters the data through three primary steps:

    1. Recursive Min-Cut Partitioning: Identifies the minimum-weight cut in GG. If a sub-graph is not deemed homogeneous by a connectivity kernel statistical test, it is recursively partitioned along the cut.
    2. Adoption: Examines singleton data objects that were separated during recursive partitioning and iteratively assigns (adopts) each into an established cluster if its average similarity to that cluster exceeds a threshold.
    3. Cluster Merging: Iteratively merges pairs of formed clusters if their inter-cluster similarity exceeds a predefined merge threshold.

    CLICK automatically identifies the number of clusters, but it can be susceptible to generating unbalanced partitions (separating single outliers) or failing to separate clusters that intersect heavily.

  5. Knowl 5 — Mean-Squared Residue Score and $\delta$-Biclustering

    model/method

    A bicluster is defined by a submatrix specified by a subset of genes G′⊆{1,…,n}G' \subseteq \{1, \dots, n\} and a subset of conditions S′⊆{1,…,m}S' \subseteq \{1, \dots, m\} within an expression matrix M=(wij)M = (w_{ij}). The coherence of the entries is measured by the mean-squared residue H(G′,S′)H(G', S'):

    H(G′,S′)=1∣G′∣∣S′∣∑i∈G′,j∈S′(wij−wiS′−wG′j+wG′S′)2H(G', S') = \frac{1}{|G'||S'|} \sum_{i \in G', j \in S'} (w_{ij} - w_{iS'} - w_{G'j} + w_{G'S'})^2

    where the row mean wiS′w_{iS'}, column mean wG′jw_{G'j}, and overall submatrix mean wG′S′w_{G'S'} are given by:

    wiS′=1∣S′∣∑j∈S′wij,wG′j=1∣G′∣∑i∈G′wij,wG′S′=1∣G′∣∣S′∣∑i∈G′,j∈S′wijw_{iS'} = \frac{1}{|S'|} \sum_{j \in S'} w_{ij}, \quad w_{G'j} = \frac{1}{|G'|} \sum_{i \in G'} w_{ij}, \quad w_{G'S'} = \frac{1}{|G'||S'|} \sum_{i \in G', j \in S'} w_{ij}

    A submatrix (G′,S′)(G', S') is a δ\delta-bicluster if H(G′,S′)≤δH(G', S') \le \delta for a chosen residue threshold δ≥0\delta \ge 0.

    Because finding a minimal set of biclusters covering the matrix is NP-hard, heuristic algorithms find biclusters by greedily adding or deleting rows and columns to minimize H(G′,S′)H(G', S'). Once a bicluster is extracted, its entries in the matrix are replaced with random noise (masked) before finding the next bicluster, which prevents redundant discovery but obscures genuine overlapping biclusters.

  6. Knowl 6 — Plaid Additive Model for Layered Subspace Clustering

    model/method

    The plaid model decomposes a gene expression matrix Y∈Rn×mY \in \mathbb{R}^{n \times m} into a superposition of a background level and KK biological "layers" (subspace clusters):

    Yij=θij0+∑k=1KθijkρikκjkY_{ij} = \theta_{ij0} + \sum_{k=1}^K \theta_{ijk} \rho_{ik} \kappa_{jk}

    where:

    • θij0\theta_{ij0} is the background expression level across the dataset.
    • θijk\theta_{ijk} is the contribution of layer kk to the expression of gene ii in sample jj.
    • ρik∈{0,1}\rho_{ik} \in \{0, 1\} indicates whether gene ii belongs to layer kk.
    • κjk∈{0,1}\kappa_{jk} \in \{0, 1\} indicates whether sample jj belongs to layer kk.

    Layers are extracted iteratively. After finding K−1K-1 layers, the KK-th layer parameters are estimated via the Expectation-Maximization (EM) algorithm by minimizing the residual sum of squares:

    Q=12∑i=1n∑j=1m(Zij−θijKρiKκjK)2Q = \frac{1}{2}\sum_{i=1}^n \sum_{j=1}^m \left( Z_{ij} - \theta_{ijK} \rho_{iK} \kappa_{jK} \right)^2

    where Zij=Yij−θij0−∑k=1K−1θijkρikκjkZ_{ij} = Y_{ij} - \theta_{ij0} - \sum_{k=1}^{K-1} \theta_{ijk} \rho_{ik} \kappa_{jk} is the residual from the prior K−1K-1 layers. The process stops when the variance within the newly identified layer falls below a threshold.

  7. Knowl 7 — Interrelated Iterative Clustering and Gene Selection for Sample Subtyping

    model/method

    Because sample-based clustering operates with very low signal-to-noise ratios (<1:10< 1:10 informative genes relative to noise genes) and sample phenotype labels are unavailable a priori, interrelated clustering algorithms alternate iteratively between sample partition discovery and gene feature pruning:

    1. Initialization: Samples and genes are pre-grouped into coarse subsets using standard partition algorithms (e.g., K-means or SOM).
    2. Interrelated Iteration: High-coherence sample subsets (representative patterns) are identified based on inter-group distance and internal consistency. Using these sample groupings as pseudo-labels, genes are scored and filtered using discriminability metrics (e.g., Information Gain, tt-test scores, or Markov blanket filtering). The reduced informative gene set is then used as the feature space to re-partition the samples (e.g., via Normalized Cut or K-means).
    3. Validation and Stopping: Convergence is assessed using the coefficient of variation (CVCV):

    CV=1K∑t=1Kσt∥μ⃗t∥CV = \frac{1}{K}\sum_{t=1}^K \frac{\sigma_t}{\|\vec{\mu}_t\|}

    where KK is the number of sample groups, μ⃗t\vec{\mu}_t is the centroid vector of group tt, and σt\sigma_t is the standard deviation within group tt. The loop terminates when stable sample classes and a corresponding set of informative genes are obtained.

  8. Knowl 8 — External Validation of Gene Expression Clusters via Pairwise Agreement Indices

    equation

    Let a clustering solution C={C1,…,Cp}\mathcal{C} = \{C_1, \dots, C_p\} and a known reference partition (ground truth) P={P1,…,Ps}\mathcal{P} = \{P_1, \dots, P_s\} over nn data objects be represented as binary co-membership matrices CC and PP, where Cij=1C_{ij}=1 (or Pij=1P_{ij}=1) if objects Oi,OjO_i, O_j share a cluster (or class), and 00 otherwise. Agreement is categorized across all (n2)\binom{n}{2} pairs into four counts:

    • n11n_{11}: pairs where Cij=1C_{ij} = 1 and Pij=1P_{ij} = 1
    • n10n_{10}: pairs where Cij=1C_{ij} = 1 and Pij=0P_{ij} = 0
    • n01n_{01}: pairs where Cij=0C_{ij} = 0 and Pij=1P_{ij} = 1
    • n00n_{00}: pairs where Cij=0C_{ij} = 0 and Pij=0P_{ij} = 0

    The agreement indices are defined as:

    Rand Index=n11+n00n11+n10+n01+n00\text{Rand Index} = \frac{n_{11} + n_{00}}{n_{11} + n_{10} + n_{01} + n_{00}}

    Jaccard Coefficient=n11n11+n10+n01\text{Jaccard Coefficient} = \frac{n_{11}}{n_{11} + n_{10} + n_{01}}

    Minkowski Measure=n10+n01n11+n01\text{Minkowski Measure} = \sqrt{\frac{n_{10} + n_{01}}{n_{11} + n_{01}}}

    In gene-based clustering, because most gene pairs belong to distinct functional clusters, n00n_{00} heavily dominates all other counts. Therefore, indices that exclude n00n_{00} (such as the Jaccard coefficient and Minkowski measure) provide more sensitive assessments of cluster quality than the Rand Index.

  9. Knowl 9 — Figure of Merit (FOM) for Predictive Cluster Validation

    equation

    The Figure of Merit (FOM) evaluates the predictive quality of a gene clustering algorithm by measuring how well gene clusters derived from m−1m-1 conditions predict expression levels in a left-out test condition e∈{1,…,m}e \in \{1, \dots, m\}.

    For a clustering of nn genes into kk clusters C1,…,CkC_1, \dots, C_k computed on all samples except ee, the predictive error on sample ee is:

    FOM(e,k)=1n∑i=1k∑x∈Ci(R(x,e)−μCi(e))2FOM(e, k) = \sqrt{\frac{1}{n} \sum_{i=1}^k \sum_{x \in C_i} \left( R(x, e) - \mu_{C_i}(e) \right)^2}

    where R(x,e)R(x, e) is the expression level of gene xx in sample ee, and μCi(e)=1∣Ci∣∑y∈CiR(y,e)\mu_{C_i}(e) = \frac{1}{|C_i|} \sum_{y \in C_i} R(y, e) is the average expression of cluster CiC_i in sample ee.

    The aggregate Figure of Merit across all mm jackknife leave-one-out iterations is:

    FOM(k)=∑e=1mFOM(e,k)FOM(k) = \sum_{e=1}^m FOM(e, k)

    Smaller values of FOM(k)FOM(k) indicate greater cluster stability and higher predictive performance.

  10. Knowl 10 — Hypergeometric Statistical Significance for Biological Functional Enrichment

    equation

    To evaluate whether a cluster of coexpressed genes is biologically meaningful, genes within the cluster are mapped to known functional categories (e.g., MIPS or Gene Ontology). Given a total genome of gg genes in which ff genes belong to a specific functional category, the probability PP of observing at least kk genes from that category in a cluster of size nn is given by the cumulative hypergeometric distribution:

    P=1−∑i=0k−1(fi)(g−fn−i)(gn)P = 1 - \sum_{i=0}^{k-1} \frac{\binom{f}{i} \binom{g - f}{n - i}}{\binom{g}{n}}

    where (ab)\binom{a}{b} is the binomial coefficient. A cluster is considered statistically significantly enriched for a biological function if its PP-value falls below a strict significance threshold (such as P<3×10−4P < 3 \times 10^{-4}), indicating that the observed functional coherence is unlikely to have arisen by chance.

Coverage note — Standard general-purpose clustering algorithms (such as standard K-means, basic SOM, and UPGMA) and specific biological benchmark dataset tables were omitted as standalone knowls to focus on the survey's domain-specific taxonomy, mathematical relationships, specialized algorithms (CAST, CLICK, Biclustering, Plaid, Interrelated Clustering), and cluster validation metrics.

References

  1. 1.R. Agrawal, J. Gehrke, D. Gunopulos, and P. Raghavan, “Automatic Subspace Clustering of High Dimensional Data for Data Mining Applications,” SIGMOD 1998, Proc. ACM SIGMOD Int’l Conf. Management of Data, pp. 94-105, 1998.
  2. 2.A.A. Alizadeh et al., “Distinct Types of Diffuse Large B-Cell Lymphoma Identified by Gene Expression Profiling,” Nature, vol. 403, pp. 503-511, Feb. 2000.
  3. 3.U. Alon, N. Barkai, D.A. Notterman, K. Gish, S. Ybarra, D. Mack, and A.J. Levine, “Broad Patterns of Gene Expression Revealed by Clustering Analysis of Tumor and Normal Colon Tissues Probed by Oligonucleotide Array,” Proc. Nat’l Academy of Science, vol. 96, no. 12, pp. 6745-6750, June 1999.
  4. 4.O. Alter, P.O. Brown, and D. Bostein, “Singular Value Decomposition for Genome-Wide Expression Data Processing and Modeling,” Proc. Nat’l Academy of Science, vol. 97, no. 18, pp. 10101-10106, Aug. 2000.
  5. 5.M. Ankerst, M.M. Breunig, H.-P. Kriegel, and J. Sander, “OPTICS: Ordering Points to Identify the Clustering Structure,” Sigmod, pp. 49-60, 1999.
  6. 6.A. Ben-Dor, N. Friedman, and Z. Yakhini, “Class Discovery in Gene Expression Data,” Proc. Fifth Ann. Int’l Conf. Computational Molecular Biology (RECOMB 2001), pp. 31-38, 2001.
  7. 7.A. Ben-Dor, R. Shamir, and Z. Yakhini, “Clustering Gene Expression Patterns,” J. Computational Biology, vol. 6, nos. 3/4, pp. 281-297, 1999.
  8. 8.M. Blat, S. Wiseman, and E. Domany, “Super-Paramagnetic Clustering of Data,” Physical Review Letters, vol. 76, pp. 3251-3255, 1996.
  9. 9.A. Brazma and J. Vilo, “Minireview: Gene Expression Data Analysis,” Federation of European Biochemical Soc., vol. 480, pp. 17-24, June 2000.
  10. 10.M.P.S. Brown, W.N. Grundy, D. Lin, N. Cristianini, C.W. Sugnet, T.S. Furey, M. Ares Jr., and D. Haussler, “Knowledge-Based Analysis of Microarray Gene Expression Data Using Support Vector Machines,” Proc. Nat’l Academy of Science, vol. 97, no. 1, pp. 262-267, Jan. 2000.
  11. 11.Y. Cheng and G.M. Church, “Biclustering of Expression Data,” Proc. Eighth Int’l Conf. Intelligent Systems for Molecular Biology (ISMB), vol. 8, pp. 93-103, 2000.
  12. 12.R.J. Cho, M.J. Campbell, E.A. Winzeler, L. Steinmetz, A. Conway, L. Wodicka, T.G. Wolfsberg, A.E. Gabrielian, D. Landsman, D.J. Lockhart, and R.W. Davis, “A Genome-Wide Transcriptional Analysis of the Mitotic Cell Cycle,” Molecular Cell, vol. 2, no. 1, pp. 65-73, July 1998.
  13. 13.S. Chu et al., “The Transcriptional Program of Sporulation in Budding Yeast,” Science, vol. 282, no. 5389, pp. 699-705, 1998.
  14. 14.D.R. Bickel, “Robust Cluster Analysis of DNA Microarray Data: An Application of Nonparametric Correlation Dissimilarity,” Proc. Joint Statistical Meetings of the Am. Statistical Assoc., (Biometrics Section), 2001.
  15. 15.J.L. DeRisi, V.R. Iyer, and P.O. Brown, “Exploring the Metabolic and Genetic Control of Gene Expression on a Genomic Scale,” Science, pp. 680-686, 1997.
  16. 16.P. D’haeseleer, X. Wen, S. Fuhrman, and R. Somogyi, “Mining the Gene Expression Matrix: Inferring Gene Relationships From Large Scale Gene Expression Data,” Information Processing in Cells and Tissues, pp. 203-212, 1998.
  17. 17.C. Ding, “Analysis of Gene Expression Profiles: Class Discovery and Leaf Ordering,” Proc. Int’l Conf. Computational Molecular Biology (RECOMB), pp. 27-136, Apr. 2002.
  18. 18.R. Dubes and A. Jain, Algorithms for Clustering Data. Prentice Hall, 1988.
  19. 19.B. Efron, “The Jackknife, the Bootstrap, and Other Resampling Plans,” Proc. CBMS-NSF Regional Conf. Series in Applied Math., vol. 38, 1982.
  20. 20.M.B. Eisen, P.T. Spellman, P.O. Brown, and D. Botstein, “Cluster Analysis and Display of Genome-Wide Expression Patterns,” Proc. Nat’l Academy of Science, vol. 95, no. 25, pp. 14863-14868, Dec. 1998.
  21. 21.C. Fraley and A.E. Raftery, “How Many Clusters? Which Clustering Method? Answers Via Model-Based Cluster Analysis,” The Computer J., vol. 41, no. 8, pp. 578-588, 1998.
  22. 22.G. Getz, E. Levine, and E. Domany, “Coupled Two-Way Clustering Analysis of Gene Microarray Data,” Proc. Nat’l Academy of Science, vol. 97, no. 22, pp. 12079-12084, Oct. 2000.
  23. 23.D. Ghosh and A.M. Chinnaiyan, “Mixture Modelling of Gene Expression Data from Microarray Experiments,” Bioinformatics, vol. 18, pp. 275-286, 2002.
  24. 24.T.R. Golub, D.K. Slonim, P. Tamayo, C. Huard, M. Gassenbeek, J.P. Mesirov, H. Coller, M.L. Loh, J.R. Downing, M.A. Caligiuri, D.D. Bloomfield, and E.S. Lander, “Molecular Classification of Cancer: Class Discovery and Class Prediction by Gene Expression Monitoring,” Science, vol. 286, no. 15, pp. 531-537, Oct. 1999.
  25. 25.M. Halkidi, Y. Batistakis, and M. Vazirgiannis, “On Clustering Validation Techniques,” Intelligent Information Systems J., 2001.
  26. 26.E. Hartuv and R. Shamir, “A Clustering Algorithm Based on Graph Connectivity,” Information Processing Letters, vol. 76, nos. 4-6, pp. 175-181, 2000.
  27. 27.T. Hastie, R. Tibshirani, D. Boststein, and P. Brown, “Supervised Harvesting of Expression Trees,” Genome Biology, vol. 2, no. 1, pp. 0003.1-0003.12, Jan. 2001.
  28. 28.I. Hedenfalk, D. Duggan, Y.D. Chen, M. Radmacher, M. Bittner, R. Simon, P. Meltzer, B. Gusterson, M. Esteller, O.P. Kallioniemi, B. Wilfond, A. Borg, and J. Trent, “Gene-Expression Profiles in Hereditary Breast Cancer,” The New England J. Medicine, vol. 344, no. 8, pp. 539-548, Feb. 2001.
  29. 29.J. Herrero, A. Valencia, and J. Dopazo, “A Hierarchical Unsupervised Growing Neural Network for Clustering Gene Expression Patterns,” Bioinformatics, vol. 17, pp. 126-136, 2001.
  30. 30.L.J. Heyer, S. Kruglyak, and S. Yooseph, “Exploring Expression Data: Identification and Analysis of Coexpressed Genes,” Genome Research, 1999.
  31. 31.L.J. Heyer, S. Kruglyak, and S. Yooseph, “Exploring Expression Data: Identification and Analysis of Coexpressed Genes,” Genome Research, vol. 9, no. 11, pp. 1106-1115, 1999.
  32. 32.A. Hill, E. Brown, M. Whitley, G. Tucker-Kellogg, C. Hunter, and D. Slonim, “Evaluation of Normalization Procedures for Oligonucleotide Array Data Based on Spiked cRNA Contros,” Genome Biology, vol. 2, no. 12, pp. research0055.-1-0055.13, 2001.
  33. 33.V.R. Iyer, M.B. Eisen, D.T. Ross, G. Schuler, T. Moore, J.C.F. Lee, J.M. Trent, L.M. Staudt, J. Hudson Jr., M.S. Boguski, D. Lashkari, D. Shalon, D. Botstein, and P.O. Brown, “The Transcriptional Program in the Response of Human Fibroblasts to Serum,” Science, vol. 283, pp. 83-87, 1999.
  34. 34.A.K. Jain, M.N. Murty, and P.J. Flynn, “Data Clustering: A Review,” ACM Computing Surveys, vol. 31, no. 3, pp. 254-323, Sept. 1999.
  35. 35.L.M. Jakt, L. Cao, K.S.E. Cheah, and D.K. Smith, “Assessing Clusters and Motifs from Gene Expression Data,” Genome Research, vol. 11, pp. 1112-123, 2001.
  36. 36.D. Jiang, J. Pei, and A. Zhang, “DHC: A Density-Based Hierarchical Clustering Method for Time-Series Gene Expression Data,” Proc. BIBE2003: Third IEEE Int’l Symp. Bioinformatics and Bioeng., 2003.
  37. 37.D. Jiang, J. Pei, and A. Zhang, “Interactive Exploration of Coherent Patterns in Time-Series Gene Expression Data,” Proc. Ninth ACM SIGKDD Int’l Conf. Knowledge Discovery and Data Mining (SIGKDD ’03), 2003.
  38. 38.L. Kaufman and P.J. Rousseeuw, Finding Groups in Data: An Introduction to Cluster Analysis. John Wiley and Sons, 1990.
  39. 39.T. Kohonen, Self-Organization and Associative Memory. Berlin: Spring-Verlag, 1984.
  40. 40.L. Lazzeroni and A. Owen, “Plaid Models for Gene Expression Data,” Statistica Sinica, vol. 12, no. 1, pp. 61-86, 2002.
  41. 41.E. Levine and E. Domany, “Resampling Methods for Unsupervised Estimation of Cluster Validity,” Neural Computation, vol. 13, pp. 2573-2593, 2001.
  42. 42.L. Li, W. Leping, C.R. Weinberg, T.A. Darden, and L.G. Pedersen, “Gene Selection for Sample Classification Based on Gene Expression Data: Study of Sensitivity to Choice of Parameters of the ga/knn Method,” Bioinformatics, vol. 17, pp. 1131-1142, 2001.
  43. 43.W. Li, “Zipf’s Law in Importance of Genes for Cancer Classification Using Microarray Data,” Lab of Statistical Genetics, Rockefeller Univ., Apr. 2001.
  44. 44.D. Lockhart et al., “Expression Monitoring by Hybridization to High-Density Oligonucleotide Arrays,” Nature Biotechnology, vol. 14, pp. 1675-1680, 1996.
  45. 45.G.J. McLachlan, R.W. Bean, and D. Peel, “A Mixture Model-Based Approach to the Clustering of Microarray Expression Data,” Bioinformatics, vol. 18, 413-422, 2002.
  46. 46.J.B. McQueen, “Some Methods for Classification and Analysis of Multivariate Observations,” Proc. Fifth Berkeley Symp. Math. Statistics and Probability, vol. 1, pp. 281-297, 1967.
  47. 47.E.J. Moler, M.L. Chow, and I.S. Mian, “Analysis of Molecular Profile Data Using Generative and Discriminative Methods.” Physiological Genomics, vol. 4, no. 2, pp. 109-126, 2000.
  48. 48.L.T. Nguyen et al., “Flow Cytometric Analysis of in Vitro Proinflammatory Cytokine Secretion in Peripheral Blood from Multiple Sclerosis Patients,” J. Clinical Immunology, vol. 19, no. 3, pp. 179-185, 1999.
  49. 49.P.J. Park, M. Pagano, and M. Bonetti, “A Nonparametric Scoring Algorithm for Identifying Informative Genes from Microarray Data,” Proc. Pacific Symp. Biocomputing, pp. 52-63, 2001.
  50. 50.C.M. Perou, S.S. Jeffrey, M.V.D. Rijn, C.A. Rees, M.B. Eisen, D.T. Ross, A. Pergamenschikov, C.F. Williams, S.X. Zhu, J.C.F. Lee, D. Lashkari, D. Shalon, P.O. Brown, and D. Bostein, “Distinctive Gene Expression Patterns in Human Mammary Epithelial Cells and Breast Cancers,” Proc. Nat’l Academy of Science, vol. 96, no. 16, pp. 9212-9217, Aug. 1999.
  51. 51.P.A. Ralf-Herwig, C. Muller, C. Bull, H. Lehrach, and J. O’Brien, “Large-Scale Clustering of cDNA-Fingerprinting Data,” Genome Research, vol. 9, pp. 1093-1105, 1999.
  52. 52.K. Rose, “Deterministic Annealing for Clustering, Compression, Classification, Regression, and Related Optimization Problems,” Proc. IEEE, vol. 96, pp. 2210-2239, 1998.
  53. 53.K. Rose, E. Gurewitz, and G. Fox, Physical Rev. Letters, vol. 65, pp. 945-948, 1990.
  54. 54.M.D. Schena, R. Shalon, R. Davis, and P. Brown, “Quantitative Monitoring of Gene Expression Patterns with a Compolementatry DNA Microarray,” Science, vol. 270, pp. 467-470, 1995.
  55. 55.J. Schuchhardt, D. Beule, A. Malik, E. Wolski, H. Eickhoff, H. Lehrach, and H. Herzel, “Normalization Strategies for cDNA Microarrays,” Nucleic Acids Research, vol. 28, no. 10, 2000.
  56. 56.R. Shamir and R. Sharan, “Click: A Clustering Algorithm for Gene Expression Analysis,” Proc. Eighth Int’l Conf. Intelligent Systems for Molecular Biology (ISMB ’00), 2000.
  57. 57.G. Sherlock, “Analysis of Large-Scale Gene Expression Data,” Current Opinion in Immunology, vol. 12, no. 2, pp. 201-205, 2000.
  58. 58.J.N. Siedow, “Meeting Report: Making Sense of Microarrays,” Genome Biology, vol. 2, no. 2, pp. reports 4003.1-4003.2, 2001.
  59. 59.F.D. Smet, J. Mathys, K. Marchal, G. Thijs, M. Moor, D. Bart, and Y. Moreau, “Adaptive Quality-Based Clustering of Gene Expression Profiles,” Bioinformatics, vol. 18, pp. 735-746, 2002.
  60. 60.R.R. Sokal, “Clustering and Classification: Background and Current Directions,” Classifincation and Clustering, J. Van Ryzin, ed., Academic Press, 1977.
  61. 61.P.T. Spellman et al., “Comprehensive Identification of Cell Cycle-Regulated Genes of the Yeast Saccharomyces Cerevisiae by Microarray Hybridization,” Molecular Biology of the Cell, vol. 9, no. 12, pp. 3273-3297, 1998.
  62. 62.P. Tamayo, D. Solni, J. Mesirov, Q. Zhu, S. Kitareewan, E. Dmitrovsky, E.S. Lander, and T.R. Golub, “Interpreting Patterns of Gene Expression with Self-Organizing Maps: Methods and Application to Hematopoietic Differentiation,” Proc. Nat’l Academy of Science, vol. 96, no. 6, pp. 2907-2912, Mar. 1999.
  63. 63.C. Tang, A. Zhang, and J. Pei, “Mining Phenotypes and Informative Genes from Gene Expression Data,” Proc. Ninth ACM SIGKDD Int’l Conf. Knowledge Discovery and Data Mining (SIGKDD ’03), 2003.
  64. 64.C. Tang, L. Zhang, A. Zhang, and M. Ramanathan, “Interrelated Two-Way Clustering: An Unsupervised Approach for Gene Expression Data Analysis,” Proc. BIBE2001: Second IEEE Int’l Symp. Bioinformatics and Bioeng., pp. 41-48, 2001.
  65. 65.C. Tang and A. Zhang, “An Iterative Strategy for Pattern Discovery in High-Dimensional Data Sets,” Proc. 11th Int’l Conf. Information and Knowledge Management (CIKM ’02), 2002.
  66. 66.S. Tavazoie, D. Hughes, M.J. Campbell, R.J. Cho, and G.M. Church, “Systematic Determination of Genetic Network Architecture,” Nature Genetics, pp. 281-285, 1999.
  67. 67.A. Tefferi, E. Bolander, M. Ansell, D. Wieben, and C. Spelsberg, “Primer on Medical Genomics Part III: Microarray Experiments and Data Analysis,” Mayo Clinic Proc., vol. 77, pp. 927-940, 2002.
  68. 68.J.G. Thomas, J.M. Olson, S.J. Tapscott, and L.P. Zhao, “An Efficient and Robust Statistical Modeling Approach to Discover Differentially Expressed Genes Using Genomic Expression Profiles,” Genome Research, vol. 11, no. 7, pp. 1227-1236, 2001.
  69. 69.O. Troyanskaya, M. Cantor, G. Sherlock, P. Brown, T. Hastie, R. Tibshirani, D. Botstein, and R. Altman, “Missing Value Estimation Methods for Dna Microarrays,” Bioinformatics, in press.
  70. 70.V.G. Tusher, R. Tibshirani, and G. Chu, “Significance Analysis of Microarrays Applied to the Ionizing Radiation Response,” Proc. Nat’l Academy of Science, vol. 98, no. 9, pp. 5116-5121, Apr. 2001.
  71. 71.H. Wang, W. Wang, Y. Wei, J. Yang, and P.S. Yu, “Clustering by Pattern Similarity in Large Data Sets,” SIGMOD 2002, Proc. ACM SIGMOD Int’l Conf. Management of Data, pp. 394-405, 2002.
  72. 72.X. Wen, S. Fuhrman, G.S. Michaels, D.B. Carr, S. Smith, J.L. Barker, and R. Smomgyi, “Large-Scale Temporal Gene Expression Mapping of Central Nervous System Development,” Proc. Nat’l Academy of Science, vol. 95, pp. 334-339, Jan. 1998.
  73. 73.E.P. Xing and R.M. Karp, “Cliff: Clustering of High-Dimensional Microarray Data via Iterative Feature Filtering Using Normalized Cuts,” Bioinformatics, vol. 17, no. 1, pp. 306-315, 2001.
  74. 74.J. Yang, W. Wang, H. Wang, and P.S. Yu, “̂-Cluster: Capturing Subspace Correlation in a Large Data Set,” Proc. 18th Int’l Conf. Data Eng. (ICDE 2002), pp. 517-528, 2002.
  75. 75.K.Y. Yeung and W.L. Ruzzo, “An Empirical Study on Principal Component Analysis for Clustering Gene Expression Data,” Technical Report UW-CSE-2000-11-03, Dept. of Computer Science & Eng., Univ. of Washington, 2000.
  76. 76.K.Y. Yeung, C. Fraley, A. Murua, A.E. Raftery, and W.L. Ruzz, “Model-Based Clustering and Data Transformations for Gene Expression Data,” Bioinformatics, vol. 17, pp. 977-987, 2001.
  77. 77.K.Y. Yeung, D.R. Haynor, and W.L. Ruzzo, “Validating Clustering for Gene Expression Data,” Bioinformatics, vol. 17, no. 4, pp. 309-318, 2001.

Citation

MLA
Daxin Jiang, et al. “Cluster Analysis for Gene Expression Data: A Survey”. IEEE Transactions on Knowledge and Data Engineering, vol. 16, no. 11, 2004, pp. 1370–86, https://doi.org/10.1109/TKDE.2004.68.
APA
Daxin Jiang, Chun Tang, & Aidong Zhang. (2004). Cluster analysis for gene expression data: a survey. IEEE Transactions on Knowledge and Data Engineering, 16(11), 1370–1386. https://doi.org/10.1109/TKDE.2004.68
Chicago
Daxin Jiang, Chun Tang, and Aidong Zhang. 2004. “Cluster Analysis for Gene Expression Data: A Survey”. IEEE Transactions on Knowledge and Data Engineering 16 (11): 1370–86. https://doi.org/10.1109/TKDE.2004.68.
Harvard
Daxin Jiang, Chun Tang and Aidong Zhang (2004) “Cluster analysis for gene expression data: a survey”, IEEE Transactions on Knowledge and Data Engineering, 16(11), pp. 1370–1386. Available at: https://doi.org/10.1109/TKDE.2004.68.
Vancouver
1. Daxin Jiang, Chun Tang, Aidong Zhang (2004) Cluster analysis for gene expression data: a survey. IEEE Transactions on Knowledge and Data Engineering 16:1370–1386

BibTeX

@article{Daxin_Jiang_2004, title={Cluster analysis for gene expression data: a survey}, volume={16}, ISSN={1041-4347}, url={http://dx.doi.org/10.1109/TKDE.2004.68}, DOI={10.1109/tkde.2004.68}, number={11}, journal={IEEE Transactions on Knowledge and Data Engineering}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Daxin Jiang and Chun Tang and Aidong Zhang}, year={2004}, month=Nov, pages={1370–1386} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF