keyword
cluster analysis
Cluster analysis is an unsupervised data analysis and machine learning technique that partitions a set of objects or data points into distinct groups, known as clusters, such that items within the same group share greater similarity with one another than with items in different groups. Operating without predefined category labels, it discovers inherent structures, distributions, and patterns directly from the underlying data based on specified feature measurements, distance metrics, or density functions. Common methodologies include partitioning, hierarchical, mode-seeking, and representation-learning approaches, which are typically assessed using specialized validation measures to determine the quality, stability, and separation of the resulting groupings. The technique is widely utilized across scientific and technical domains, including bioinformatics, image analysis, text mining, and pattern recognition, to organize and interpret complex, high-dimensional datasets.
4 items

Cluster analysis for gene expression data: a survey
Daxin Jiang, Chun Tang, Aidong Zhang
Why you should read this
Categorizes clustering methods for microarray gene expression data into gene-based, sample-based, and subspace approaches while reviewing specific algorithms, proximity measures, and validation techniques to guide functional genomics research.
DNA microarray technology has now made it possible to simultaneously monitor the expression levels of thousands of genes during important biological processes and across collections of related samples. Elucidating the patterns hidden in gene expression data offers a tremendous opportunity for an enhanced understanding of functional genomics. However, the large number of genes and the complexity of biological networks greatly increases the challenges of comprehending and interpreting the resulting mass of data, which often consists of millions of measurements. A first step toward addressing this challenge is the use of clustering techniques, which is essential in the data mining process to reveal natural structures and identify interesting patterns in the underlying data. Cluster analysis seeks to partition a given data set into groups based on specified features so that the data points within a group are more similar to each other than the points in different groups. A very rich literature on cluster analysis has developed over the past three decades. Many conventional clustering algorithms have been adapted or directly applied to gene expression data, and also new algorithms have recently been proposed specifically aiming at gene expression data. These clustering algorithms have been proven useful for identifying biologically relevant groups of genes and samples. In this paper, we first briefly introduce the concepts of microarray technology and discuss the basic elements of clustering on gene expression data. In particular, we divide cluster analysis for gene expression data into three categories. Then, we present specific challenges pertinent to each clustering category and introduce several representative approaches. We also discuss the problem of cluster validation in three aspects and review various methods to assess the quality and reliability of clustering results. Finally, we conclude this paper and suggest the promising trends in this field.
Added
2026-09-25

Information Theoretic Measures for Clusterings Comparison: Variants, Properties, Normalization and Correction for Chance
X. Nguyen, Julien Epps, James Bailey
Why you should read this
Establishes which information-theoretic clustering comparison measures satisfy metric, normalization, and chance-correction properties, motivating normalized information distance as a principled default.
Information theoretic measures form a fundamental class of measures for comparing clusterings, and have recently received increasing interest. Nevertheless, a number of questions concerning their properties and inter-relationships remain unresolved. In this paper, we perform an organized study of information theoretic measures for clustering comparison, including several existing popular measures in the literature, as well as some newly proposed ones. We discuss and prove their important properties, such as the metric property and the normalization property. We then highlight to the clustering community the importance of correcting information theoretic measures for chance, especially when the data size is small compared to the number of clusters present therein. Of the available information theoretic based measures, we advocate the normalized information distance (NID) as a general measure of choice, for it possesses concurrently several important properties, such as being both a metric and a normalized measure, admitting an exact analytical adjusted-for-chance form, and using the nominal [0, 1] range better than other normalized variants.
Added
2026-09-14

Unsupervised Deep Embedding for Clustering Analysis
Junyuan Xie, Ross B. Girshick, Ali Farhadi
Why you should read this
Proposes Deep Embedded Clustering (DEC), an unsupervised method that simultaneously learns low-dimensional feature representations and cluster assignments within deep neural networks to outperform traditional clustering pipelines on image and text benchmarks.
Clustering is central to many data-driven application domains and has been studied extensively in terms of distance functions and grouping algorithms. Relatively little work has focused on learning representations for clustering. In this paper, we propose Deep Embedded Clustering (DEC), a method that simultaneously learns feature representations and cluster assignments using deep neural networks. DEC learns a mapping from the data space to a lower-dimensional feature space in which it iteratively optimizes a clustering objective. Our experimental evaluations on image and text corpora show significant improvement over state-of-the-art methods.
Added
2026-09-11

Mean Shift, Mode Seeking, and Clustering
Yizong Cheng
Why you should read this
Establishes a rigorous theoretical foundation for the generalized mean shift algorithm by proving it acts as adaptive-step gradient ascent on kernel density surfaces, analyzing its convergence, and showing how k-means clustering emerges as a limiting case.
Mean shift, a simple iterative procedure that shifts each data point to the average of data points in its neighborhood, is generalized and analyzed in this paper. This generalization makes some k-means like clustering algorithms its special cases. It is shown that mean shift is a mode-seeking process on a surface constructed with a “shadow” kernel. For Gaussian kernels, mean shift is a gradient mapping. Convergence is studied for mean shift iterations. Cluster analysis is treated as a deterministic problem of finding a fixed point of mean shift that characterizes the data. Applications in clustering and Hough transform are demonstrated. Mean shift is also considered as an evolutionary strategy that performs multistart global optimization.
Added
2026-09-10
