Consensus Clustering: A Resampling-Based Method for Class Discovery and Visualization of Gene Expression Microarray Data

Stefano MontiPablo TamayoJill MesirovTodd Golub

article2003Machine-mediated learning2,112 citations
Cover for Consensus Clustering: A Resampling-Based Method for Class Discovery and Visualization of Gene Expression Microarray Data

Abstract

In this paper we present a new methodology of class discovery and clustering validation tailored to the task of analyzing gene expression data. The method can be best thought of as an analysis approach, to guide and assist in the use of any of a wide range of available clustering algorithms. We call the new methodology consensus clustering, and in conjunction with resampling techniques, it provides for a method to represent the consensus across multiple runs of a clustering algorithm and to assess the stability of the discovered clusters. The method can also be used to represent the consensus over multiple runs of a clustering algorithm with random restart (such as K-means, model-based Bayesian clustering, SOM, etc.), so as to account for its sensitivity to the initial conditions. Finally, it provides for a visualization tool to inspect cluster number, membership, and boundaries. We present the results of our experiments on both simulated data and real gene expression data aimed at evaluating the effectiveness of the methodology in discovering biologically meaningful clusters.

Table of Contents

  • 1. Introduction
  • 2. Background
  • 3. Methodology
  • 3.1. Notation
  • 3.2. Measuring consensus
  • 3.2.1. Consensus matrix reordering and visualization
  • 3.2.2. Consensus' summary statistics
  • 3.3. Determining the number of clusters
  • 3.3.1. Consensus distribution
  • 3.4. Resampling schemes
  • 4. Experimental evaluation
  • 4.1. Evaluation methodology
  • 4.1.1. Evaluation metrics
  • 4.1.2. Experimental design
  • 4.1.3. Consensus clustering settings
  • 4.2. Datasets
  • 4.2.1. Simulated data
  • 4.2.2. Gene-expression microarray data
  • 4.3. Results
  • 4.3.1. Simulated data
  • 4.3.2. Gene-expression data
  • 4.4. Discussion
  • 5. Conclusions
  • Appendix: The adjusted Rand index
  • Acknowledgments
  • Notes
  • References

Knowls

  1. Knowl 1 — Consensus Clustering Algorithm

    algorithm

    Consensus clustering is a resampling-based framework designed to estimate the number of clusters in a dataset, assess cluster stability, and discover robust cluster assignments by measuring agreement across multiple perturbed runs of an underlying clustering algorithm.

    Input: Dataset D={e1,e2,…,eN}D = \{e_1, e_2, \dots, e_N\}, clustering algorithm Cluster\text{Cluster}, resampling scheme Resample\text{Resample}, number of resampling iterations HH, set of candidate cluster numbers K={K1,…,Kmax⁡}\mathcal{K} = \{K_1, \dots, K_{\max}\}
    Output: Estimated optimal number of clusters K^\hat{K}, consensus partitions PP, and consensus matrices {M(K):K∈K}\{\mathcal{M}^{(K)} : K \in \mathcal{K}\}
    for K∈KK \in \mathcal{K} do
        Initialize M=∅M = \emptyset (set of connectivity matrices)
        for h=1h = 1 to HH do
            D(h)←Resample(D)D^{(h)} \leftarrow \text{Resample}(D) # Generate perturbed subset/bootstrap of DD
            M(h)←Cluster(D(h),K)M^{(h)} \leftarrow \text{Cluster}(D^{(h)}, K) # Cluster perturbed data into KK clusters
            M←M∪{M(h)}M \leftarrow M \cup \{M^{(h)}\}
        end for
        Compute consensus matrix M(K)\mathcal{M}^{(K)} by averaging connectivity across all hh where sample pairs co-occur
    end for
    Select best K^∈K\hat{K} \in \mathcal{K} based on consensus distribution / CDF progression of M(K)\mathcal{M}^{(K)}
    Partition DD into K^\hat{K} clusters based on M(K^)\mathcal{M}^{(\hat{K})}
    return K^\hat{K}, partition PP, and {M(K):K∈K}\{\mathcal{M}^{(K)} : K \in \mathcal{K}\}

    To construct the final partition from M(K^)\mathcal{M}^{(\hat{K})}, the distance matrix 1−M(K^)1 - \mathcal{M}^{(\hat{K})} is supplied as an empirical pairwise dissimilarity matrix to agglomerative hierarchical clustering (typically with average linkage), cutting the dendrogram when K^\hat{K} branches remain.

  2. Knowl 2 — Consensus Matrix and Connectivity Formulation

    equation

    Given a dataset D={e1,e2,…,eN}D = \{e_1, e_2, \dots, e_N\} of NN items, let D(1),D(2),…,D(H)D^{(1)}, D^{(2)}, \dots, D^{(H)} be HH perturbed datasets generated via resampling. For the hh-th perturbation, let M(h)M^{(h)} be the N×NN \times N connectivity matrix resulting from clustering D(h)D^{(h)} into KK clusters:

    M(h)(i,j)={1if items i and j belong to the same cluster in run h,0otherwise.M^{(h)}(i, j) = \begin{cases} 1 & \text{if items } i \text{ and } j \text{ belong to the same cluster in run } h, \\ 0 & \text{otherwise.} \end{cases}

    Let I(h)I^{(h)} be an N×NN \times N indicator matrix where I(h)(i,j)=1I^{(h)}(i, j) = 1 if both items ii and jj were present in the resampled subset D(h)D^{(h)}, and 00 otherwise. The consensus matrix M\mathcal{M} (or M(K)\mathcal{M}^{(K)} for a given number of clusters KK) is defined by normalizing over all joint resamplings:

    M(i,j)=∑h=1HM(h)(i,j)∑h=1HI(h)(i,j)\mathcal{M}(i, j) = \frac{\sum_{h=1}^H M^{(h)}(i, j)}{\sum_{h=1}^H I^{(h)}(i, j)}

    Each entry M(i,j)∈[0,1]\mathcal{M}(i, j) \in [0, 1] represents the consensus index between items ii and jj. For probabilistic soft-clustering algorithms outputting posterior membership probabilities P(i∈Ck∣D)P(i \in C_k \mid D), the connectivity matrix generalizes to:

    M(h)(i,j)=∑k=1KP(i∈Ck∣D(h))P(j∈Ck∣D(h))M^{(h)}(i, j) = \sum_{k=1}^K P(i \in C_k \mid D^{(h)}) P(j \in C_k \mid D^{(h)})

  3. Knowl 3 — Cluster Consensus and Item Consensus Statistics

    definition

    For a dataset D={e1,…,eN}D = \{e_1, \dots, e_N\} partitioned into KK clusters, let Ik={j:ej∈k}I_k = \{j : e_j \in k\} denote the set of indices belonging to cluster kk, and let Nk=∣Ik∣N_k = |I_k| be the cluster size. From the consensus matrix M\mathcal{M}:

    1. Cluster Consensus m(k)m(k): The mean consensus index across all distinct pairs belonging to cluster kk:

    m(k)=1Nk(Nk−1)/2∑i,j∈Iki<jM(i,j)m(k) = \frac{1}{N_k(N_k - 1)/2} \sum_{\substack{i, j \in I_k \\ i < j}} \mathcal{M}(i, j)

    A value close to 11 reflects a highly stable and homogeneous cluster.

    1. Item Consensus mi(k)m_i(k): The average consensus index between item eie_i and all other items belonging to cluster kk:

    mi(k)=1Nk−1{ei∈Ik}∑j∈Ikj≠iM(i,j)m_i(k) = \frac{1}{N_k - \mathbf{1}\{e_i \in I_k\}} \sum_{\substack{j \in I_k \\ j \neq i}} \mathcal{M}(i, j)

    where 1{⋅}\mathbf{1}\{\cdot\} is the indicator function. Item consensus measures how prototypically an item belongs to a target cluster kk, enabling sample ranking within clusters and identification of borderline or outlier samples.

  4. Knowl 4 — Model Selection via Consensus Distribution and CDF Area Increase

    model/method

    To identify the true number of clusters KK, the distribution of entries in the consensus matrix M(K)\mathcal{M}^{(K)} is evaluated across candidate values K∈{2,3,…,Kmax⁡}K \in \{2, 3, \dots, K_{\max}\}. For a well-fitted, stable clustering, consensus entries concentrate at 00 (items never co-clustered) and 11 (items always co-clustered).

    The empirical cumulative distribution function (CDF) over consensus values c∈[0,1]c \in [0, 1] is computed as:

    CDF(c)=∑i<j1{M(i,j)≤c}N(N−1)/2\text{CDF}(c) = \frac{\sum_{i < j} \mathbf{1}\{\mathcal{M}(i, j) \le c\}}{N(N - 1)/2}

    The area under the consensus CDF curve, denoted A(K)A(K), is calculated over the sorted unique entries {x1,x2,…,xm}\{x_1, x_2, \dots, x_m\} of M(K)\mathcal{M}^{(K)} (m=N(N−1)/2m = N(N-1)/2):

    A(K)=∑i=2m[xi−xi−1]CDF(xi)A(K) = \sum_{i=2}^m [x_i - x_{i-1}] \text{CDF}(x_i)

    The proportional increase in CDF area Δ(K)\Delta(K) measures the gain in consensus concentration:

    Δ(K)={A(K)if K=2A(K+1)−A(K)A(K)if K>2\Delta(K) = \begin{cases} A(K) & \text{if } K = 2 \\ \frac{A(K+1) - A(K)}{A(K)} & \text{if } K > 2 \end{cases}

    For algorithms where cluster boundaries are not strictly nested across increments of KK (e.g., self-organizing maps), A(K)A(K) is replaced by A^(K)=max⁡K′∈{2,…,K}A(K′)\hat{A}(K) = \max_{K' \in \{2, \dots, K\}} A(K'). The optimal cluster number K^\hat{K} corresponds to the point where the CDF progression ceases to exhibit significant increases, indicated by Δ(K)\Delta(K) leveling off close to zero.

  5. Knowl 5 — Consensus Matrix Visualization and Leaf-Ordered Heatmaps

    model/method

    Consensus matrices are visualized as color-coded heatmaps on a [0,1][0, 1] color gradient (white for 00, dark red for 11). A matrix exhibiting perfect consensus across iterations presents sharply defined block-diagonal red squares along the diagonal corresponding to clusters, surrounded by pure white background.

    To display an unlabelled dataset, the pairwise dissimilarity matrix 1−M1 - \mathcal{M} is clustered using agglomerative hierarchical clustering. The leaves of the resulting dendrogram are reordered using an optimal leaf-ordering algorithm that places items with the highest pairwise consensus index adjacent to one another. Applying this identical optimal ordering to both rows and columns aligns contiguous cluster blocks along the diagonal without altering the underlying distance topology.

  6. Knowl 6 — Consensus Clustering Experimental Protocol and Resampling Settings

    experimental setup

    The standard evaluation settings for consensus clustering in gene expression and simulated benchmark evaluations are:

    1. Resampling Scheme: Subsampling without replacement of 80%80\% of data items per perturbation over H=500H = 500 iterations when using hierarchical clustering (average linkage) and H=200H = 200 iterations when using Self-Organizing Maps (SOM).
    2. Feature Selection: In phenotype validation datasets, candidate genes are ranked using a signal-to-noise ratio (SNR) metric against the phenotype classes. The top nn up-regulated genes for each class distinction with permutation test significance p≤0.05p \le 0.05 are selected.
    3. Data Normalization: Dataset expression matrices are row- and column-standardized (mean 00, standard deviation 11) to ensure balanced branch splits during inner-loop hierarchical clustering.
    4. Validation Metric: Partition concordance against known biological class labels is scored using the adjusted Rand index, ranging between 00 (random partitioning under a generalized hypergeometric model) and 11 (exact partition match).
  7. Knowl 7 — Cluster Number Estimation and Partition Accuracy on Simulated Datasets

    data/table

    Consensus clustering (CC) using hierarchical clustering (HC) and SOM was evaluated on simulated datasets ranging from null distributions (Uniform1, Gaussian1) to overlapping mixtures (Gaussian3, Gaussian4, Gaussian5 with inter-center distances λ=2,3\lambda = 2, 3, Simulated6, Simulated4), and compared to the Gap statistic and supervised Naive-Bayes (NB) baselines.

    Dataset KtrueK_{\text{true}} CCHC\text{CC}_{\text{HC}} CCSOM\text{CC}_{\text{SOM}} GapHC\text{Gap}_{\text{HC}} GapSOM\text{Gap}_{\text{SOM}} Rand CCHC\text{Rand } \text{CC}_{\text{HC}}
    Uniform1 1 1 1 1 1 –
    Gaussian1 1 1 1 3 1 –
    Gaussian3 3 3 3 3 3 1.000
    Gaussian4 4 4 4 1 1 (4) 0.915
    Gaussian5 (λ=3\lambda = 3) 5 5 5 1 (5) 1 (5) 0.932
    Gaussian5 (λ=2\lambda = 2) 5 4–5 4 1 1 0.589
    Simulated6 6–7 7 6 7 3 0.986
    Simulated4 4 4 4 4 2 1.000

    The results show that consensus clustering accurately recovers KtrueK_{\text{true}} on multi-cluster datasets where the standard Gap statistic criteria frequently select K=1K = 1. In datasets with heavy overlap (Gaussian5, λ=2\lambda=2), Rand accuracy drops predictably as points fall within adjacent distribution tails.

  8. Knowl 8 — Evaluation on Gene Expression Microarray Datasets

    data/table

    Consensus clustering was applied to six published cancer microarray datasets (Golub et al. Leukemia, Su et al. Novartis multi-tissue, Yeoh et al. St. Jude leukemia, Bhattacharjee et al. Lung cancer, Pomeroy et al. CNS tumors, and Ramaswamy et al. Normal tissues).

    Dataset KtrueK_{\text{true}} CCHC\text{CC}_{\text{HC}} CCSOM\text{CC}_{\text{SOM}} GapHC\text{Gap}_{\text{HC}} Rand HC Rand CCHC\text{Rand } \text{CC}_{\text{HC}} Rand NB
    Leukemia 3 5 4 5 0.648 (0.46) 0.648 (1.00) 1.000
    Novartis 4 4 4 4 0.830 0.921 0.946
    St. Jude 6 5 (6) 5/7 (6) 5 0.949 0.948 0.971
    Lung cancer 4+ 5 5 (7) 5 0.307 (0.28) 0.310 (0.28) 0.904
    CNS tumors 5 5 5/6 6 0.628 0.549 0.632
    Normal tissues 13 7 4/5 12 0.457 (0.572) 0.457 (0.572) 0.655

    Consensus clustering using distance 1−M1 - \mathcal{M} generally achieves higher or equal Rand accuracy compared to standard hierarchical clustering applied directly to Euclidean raw data distances (e.g., Novartis improves from 0.8300.830 to 0.9210.921). Numbers in parentheses for estimated KK indicate consensus values confirmed by visual heatmap inspection, and for Rand index indicate accuracy when forced to evaluate at KtrueK_{\text{true}}.

  9. Knowl 9 — Discrepancy Resolution Between Quantitative Consensus and Outlier Substructures

    empirical result

    When a dataset contains a single outlier or a hybrid sample expressing markers from multiple classes (such as sample 8 in Simulated6, or a single hyperdiploid >50 sample in the St. Jude leukemia dataset), automated selection via Δ(K)\Delta(K) can suggest Ktrue+1K_{\text{true}} + 1 clusters because the outlier isolates itself into a stable singleton cluster.

    In such cases, quantitative distribution curves alone may obscure the underlying structure, but visual examination of the ordered consensus heatmap immediately reveals clean block-diagonal forms for the KtrueK_{\text{true}} biological classes alongside an isolated 11-sample block. Visual inspection of heatmap progression thus resolves ambiguities in automated cluster estimation.

  10. Knowl 10 — Limitations in Cluster Disambiguation and High-Class-to-Sample Scenarios

    limitation

    Consensus clustering exhibits specific operational constraints:

    1. K=1K=1 vs. K=2K=2 Distinction: Because the consensus matrix area for K=1K=1 is identically zero (A(1)=0A(1) = 0), the proportional area metric Δ(K)\Delta(K) cannot be computed between 11 and 22 clusters; distinguishing a single homogeneous population from two clusters requires manual inspection of consensus CDF bimodality.
    2. Small Sample Size with Large Class Counts: Performance deteriorates significantly when the number of true classes is high relative to the sample size (e.g., the Normal tissues dataset containing 1313 classes across 9090 samples, where consensus clustering resolved only 77 broader cluster aggregates).
    3. Sensitivity to Base Clustering and Normalization: Results depend directly on the inner-loop clustering algorithm (HC generally outperforms SOM) and preprocessing, requiring row- and column-normalization to prevent hierarchical linkage from producing trivial singleton branches.

Coverage note — None was omitted.

References

  1. 1.Banfield, J., & Raftery, A. E. (1993). Model-based Gaussian and non-Gaussian clustering. Biometrics, 49, 803–821.
  2. 2.Bar-Joseph, Z., Demaine, E. D., Gifford, D. K., Hamel, A. M., Jaakkola, T. S., & Srebro, N. (2002). K-ary clustering with optimal leaf ordering for gene expression data. Bioinformatics, to appear.
  3. 3.Ben-Hur, A., Elisseeff, A., & Guyon, I. (2002). A stability based method for discovering structure in clustered data. In Pacific Symposium on Biocomputing 2002, vol. 7, pp. 6–17, Lihue, Hawaii.
  4. 4.Bhattacharjee, A., Richards, W. G., Staunton, J., Li, C., Monti, S., Vasa, P., Ladd, C., Beheshti, J., Bueno, R., Gillette, M., Loda, M., Weber, G., Mark, E. J., Lander, E. S., Wong, W., Johnson, B. E., Golub, T. R., Sugarbaker, D. J., & Meyerson, M. (2001). Classification of human lung carcinomas by mRNA expression profiling reveals distinct adenocarcinomas sub-classes. In Proceedings of the National Academy of Sciences, 98:24, 13790–13795.
  5. 5.Bock, H. (1985). On some significance tests in cluster analysis. Journal of Classification, 2, 77–108.
  6. 6.Cheeseman, P., & Stutz, J. (1996), Bayesian classification (AutoClass): Theory and results. In U. M. Fayyad, G. Piatetsky-Shapiro, P. Smyth, & R. Uthurasamy (Eds.), Advances in Knowledge Discovery and Data Mining, pp. 153–180, MIT Press.
  7. 7.Chickering, D. M., & Heckerman, D. (1997). Efficient approximation for the marginal likelihood of Bayesian networks with hidden variables. Machine Learning, 29, 181–212.
  8. 8.Cowell, F. A. (1995). Measuring Inequality. New York: Prentice Hall.
  9. 9.Duda, R. O., & Hart, P. E. (1973). Pattern Classification and Scene Analysis. John Wiley & Sons.
  10. 10.Dudoit, S., & Fridlyand, J. (2002). A prediction-based resampling method for estimating the number of clusters in a dataset. Genome Biology, 3:7, 1–21.
  11. 11.Efron, B., & Tibshirani, R. J. (1994). An Introduction to the Bootstrap, No. 57 in Monographs on Statistics and Applied Probability. CRC Press.
  12. 12.Eisen, M. B., Spellman, P. T., Brown, P. O., & Botstein, D. (1998). Cluster analysis and display of genome-wide expression patterns. Proceedings of the National Academy of Sciences, 95, 14863–14868.
  13. 13.Golub, T. R., Slonim, D. K., Tamayo, P., Huard, C., Gaasenbeek, M., Mesirov, J. P., Coller, H., Loh, M., Downing, J., Caligiuri, M., Bloomfield, C., & Lander, E. (1999). Molecular classification of cancer: Class discovery and class prediction by gene expression. Science, 286:5439, 531–537.
  14. 14.Hartigan, J. A. (1978). Asymptotic distributions for clustering criteria. Annals of Statistics, 6:1, 117–131.
  15. 15.Hastie, T., Tibshirani, R., & Friedman, J. (2001). The Elements of Statistical Learning, Statistics. New York: Springer.
  16. 16.Hubert, L., & Arabie, P. (1985). Comparing partitions. Journal of Classification, 2, 193–218.
  17. 17.Jain, A. K., & Dubes, R. C. (1988). Algorithms for Clustering Data. Englewood Cliffs, NJ: Prentice Hall.
  18. 18.Jain, A. K., & Moreau, J. (1988). Bootstrap techniques in cluster analysis. Pattern Recognition, 20, 547–568.
  19. 19.Kass, R. E. & Raftery, A. E. (1995). Bayes factors. Journal of the American Statistical Association, 90, 773–795.
  20. 20.Kohonen, T. (1990). The self-organizing map. Proceedings of the IEEE, 78:9, 1464–1480.
  21. 21.Kohonen, T. (1997). Self-Organizing Maps, Information Sciences. Springer.
  22. 22.Levine, E., & Domany, E. (2001). Resampling method for unsupervised estimation of cluster validity. Neural Computation, 13:11, 2573–2593.
  23. 23.Milligan, G., & Cooper, M. (1985). An examination of procedures for determining the number of clusters in a data set. Psyochometrika, 50, 159–179.
  24. 24.Milligan, G. & Cooper, M. (1986). A study of the comparability of external criteria for hierarchical cluster analysis. Multivariate Behavioral Research, 21, 441–458.
  25. 25.Pomeroy, S., Tamayo, P., Gaasenbeek, M., Angelo, L. M. S. M., McLaughlin, M. E., Kim, J. Y., Goumnerova, L. C., Black, P. M., Lau, C., Allen, J. C., Zagzag, D., Olson, J. M., Curran, T., Wetmore, C., Biegel, J. A., Poggio, T., Mukherjee, S., Rifkin, A., Califano, G., Stolovitzky, D. N., Louis, J. P., Mesirov, E. S., Lander, R., & Golub, T. R. (2002). Gene expression-based classification and outcome prediction of central nervous system embryonal tumors. Nature, 415:6870, 436–442.
  26. 26.Ramaswamy, S., Tamayo, P., Rifkin, R., Mukherjee, S., Yeang, C.-H., Angelo, M., Ladd, C., Reich, M., Latulippe, E., Mesirov, J. P., Poggio, T., Gerald, W., Loda, M., Lander, E. S., & Golub, T. R. (2001). Multi-class cancer diagnosis using tumor gene expression signatures. Proceedings of the National Academy of Sciences, 98:26, 15149–15154.
  27. 27.Ramoni, M., Sebastiani, P., & Kohane, I. S. (2002). Cluster analysis of gene expression dynamics. In Proceedings of the National Academy of Sciences, 99:14, 9121–9126.
  28. 28.Slonim, D. K., Tamayo, P., Mesirov, J. P., Golub, T. R., & Lander, E. S. (2000). Class prediction and discovery using gene expression data. In RECOMB 2000: The Fourth Annual International Conference on Research in Computational Molecular Biology (pp. 263–272), Tokyo, Japan.
  29. 29.Su, A. I., Cooke, M. P., Ching, K. A., Hakak, Y., Walker, J. R., Wiltshire, T., Orth, A. P., Vega, R. G., Sapinoso, L. M., Moqrich, A., Patapoutian, A., Hampton, G. M., Schultz, P. G., & Hogenesch, J. B. (2002). Large-scale analysis of the human and mouse transcriptomes. Proceedings of the National Academy of Sciences, 99:7, 4465–447.
  30. 30.Tamayo, P., Slonim, D., Mesirov, J., Zhu, Q., Kitareewan, S., Dmitrovsky, E., Lander, E. S., & Godlub, T. R. (1999), Interpreting patterns of gene expression with self-organizing maps: Methods and application to hematopoietic differentiation. Proceedings of the National Academy of Sciences, 96, 2907–2912.
  31. 31.Tibshirani, R., Walther, G., Botstein, D., & Brown, P. (2001a). Cluster validation by prediction strength. Unpublished manuscript (http://www-stat.stanford.edu/~tibs/ftp/predstr.pdf).
  32. 32.Tibshirani, R., Walther, G., & Hastie, T. (2001b). Estimating the number of clusters in a dataset via the gap statistic. Journal of the Royal Statistical Society B, 63:2, 411–423.
  33. 33.Titterington, D., Smith, A., & Makov, U. (1985). Statistical Analysis of Finite Mixture Distributions. New York: Wiley.
  34. 34.Todd Golub et. al. (2002). GeneCluster 2.0. http://www-genome.wi.mit.edu/cancer/software/genecluster2/ gc2.html.
  35. 35.West, M. (2002). Bayesian factor regression models in the Large p, Small n Paradigm. Bayesian Statistics, 7, to appear.
  36. 36.West, M., Blanchette, C., Dressman, H., Huang, E., Ishida, S., Spang, R., Zuzan, H., Olson Jr., J. A., Marks, J. R., & Nevins, J. R. (2001). Predicting the clinical status of human breast cancer by using gene expression profiles. Proceedings of the National Academy of Sciences, 98:20, 11462–11467.
  37. 37.Yeoh, E.-J., Ross, M. E., Shurtleff, S. A., Williams, W. K., Patel, D., Mahfouz, R., Behm, F. G., Raimondi, S. C., Relling, M. V., Patel, A., Cheng, C., Campana, D., Wilkins, D., Zhou, X., Li, J., Liu, H., Pui, C.-H., Evans, W. E., Naeve, C., Wong, L., & Downing, J. R. (2002). Classification, subtype discovery, and prediction of outcome in pediatric acute lymphoblastic leukemia by gene expression profiling. Cancer Cell, 1:2.
  38. 38.Yeung, K. Y., Fraley, C., Murua, A., Raftery, A. E., & Ruzzo, W. L. (2001a). Model-based clustering and data transformations for gene expression data. Bioinformatics, 17:10, 977–987.
  39. 39.Yeung, K. Y., Haynor, D. R., & Ruzzo, W. L. (2001b) Validating clustering for gene expression data. Bioinformatics, 17:4.

Citation

MLA
Monti, S., et al. “Consensus Clustering: A Resampling-Based Method for Class Discovery and Visualization of Gene Expression Microarray Data”. Machine Learning, vol. 52, nos. 1-2, 2003, pp. 91–118, https://doi.org/10.1023/A:1023949509487.
APA
Monti, S., Tamayo, P., Mesirov, J., & Golub, T. (2003). Consensus Clustering: A Resampling-Based Method for Class Discovery and Visualization of Gene Expression Microarray Data. Machine Learning, 52(1-2), 91–118. https://doi.org/10.1023/A:1023949509487
Chicago
Monti, S., P. Tamayo, J. Mesirov, and T. Golub. 2003. “Consensus Clustering: A Resampling-Based Method for Class Discovery and Visualization of Gene Expression Microarray Data”. Machine Learning 52 (1-2): 91–118. https://doi.org/10.1023/A:1023949509487.
Harvard
Monti, S. et al. (2003) “Consensus Clustering: A Resampling-Based Method for Class Discovery and Visualization of Gene Expression Microarray Data”, Machine Learning, 52(1-2), pp. 91–118. Available at: https://doi.org/10.1023/A:1023949509487.
Vancouver
1. Monti S, Tamayo P, Mesirov J, Golub T (2003) Consensus Clustering: A Resampling-Based Method for Class Discovery and Visualization of Gene Expression Microarray Data. Machine Learning 52:91–118

BibTeX

@article{Monti_2003, title={Consensus Clustering: A Resampling-Based Method for Class Discovery and Visualization of Gene Expression Microarray Data}, volume={52}, ISSN={1573-0565}, url={http://dx.doi.org/10.1023/A:1023949509487}, DOI={10.1023/a:1023949509487}, number={1-2}, journal={Machine Learning}, publisher={Springer Science and Business Media LLC}, author={Monti, Stefano and Tamayo, Pablo and Mesirov, Jill and Golub, Todd}, year={2003}, month=July, pages={91–118} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF