keyword
subspace clustering
Subspace clustering is an unsupervised learning technique that partitions a collection of high-dimensional data points into distinct clusters such that points within each cluster lie in or near a common low-dimensional linear or affine subspace. Unlike traditional clustering algorithms that evaluate point-to-point distances across the entire ambient feature space, subspace clustering models complex high-dimensional datasets as a union of multiple lower-dimensional subspaces. Common methodologies for solving this problem include iterative geometric fitting, statistical models, and self-representation frameworks that utilize sparse or low-rank optimization to build affinity matrices followed by spectral graph partitioning. This technique is especially valuable in high-dimensional signal processing, computer vision, and data analysis tasks—such as motion segmentation, facial image grouping, and genomics—where data samples from different categories naturally conform to distinct low-rank linear structures.
7 items

Semantic-Enhanced Image Clustering
Shaotian Cai, Liping Qiu, Xiaojun Chen, Qin Zhang, Longteng Chen
Why you should read this
Proposes a CLIP-guided image clustering framework that extracts semantic WordNet spaces and enforces dual-space consistency to accurately separate visually similar yet semantically distinct images without predefined class names.
Image clustering is an important and open-challenging task in computer vision. Although many methods have been proposed to solve the image clustering task, they only explore images and uncover clusters according to the image features, thus being unable to distinguish visually similar but semantically different images. In this paper, we propose to investigate the task of image clustering with the help of a visual-language pre-training model. Different from the zero-shot setting, in which the class names are known, we only know the number of clusters in this setting. Therefore, how to map images to a proper semantic space and how to cluster images from both image and semantic spaces are two key problems. To solve the above problems, we propose a novel image clustering method guided by the visual-language pre-training model CLIP, named Semantic-Enhanced Image Clustering (SIC). In this new method, we propose a method to map the given images to a proper semantic space first and efficient methods to generate pseudo-labels according to the relationships between images and semantics. Finally, we propose performing clustering with consistency learning in both image space and semantic space, in a self-supervised learning fashion. The theoretical result of convergence analysis shows that our proposed method can converge at a sublinear speed. Theoretical analysis of expectation risk also shows that we can reduce the expected risk by improving neighborhood consistency, increasing prediction confidence, or reducing neighborhood imbalance. Experimental results on five benchmark datasets clearly show the superiority of our new method.
Added
2026-09-26

XAI Beyond Classification: Interpretable Neural Clustering
Xi Peng, Yunfan Li, Ivor W. Tsang, Hongyuan Zhu, Jiancheng Lv, Joey Tianyi Zhou
Why you should read this
Proposes an intrinsically explainable neural network that reformulates discrete k-means into a differentiable layer, enabling end-to-end parallel optimization, online clustering on data streams, and provable convergence without relying on post-hoc interpretations.
In this paper, we study two challenging problems in explainable AI (XAI) and data clustering. The first is how to directly design a neural network with inherent interpretability, rather than giving post-hoc explanations of a black-box model. The second is implementing discrete k-means with a differentiable neural network that embraces the advantages of parallel computing, online clustering, and clustering-favorable representation learning. To address these two challenges, we design a novel neural network, which is a differentiable reformulation of the vanilla k-means, called inTerpretable nEuraL cLustering (TELL). Our contributions are threefold. First, to the best of our knowledge, most existing XAI works focus on supervised learning paradigms. This work is one of the few XAI studies on unsupervised learning, in particular, data clustering. Second, TELL is an interpretable, or the so-called intrinsically explainable and transparent model. In contrast, most existing XAI studies resort to various means for understanding a black-box model with post-hoc explanations. Third, from the view of data clustering, TELL possesses many properties highly desired by k-means, including but not limited to online clustering, plug-and-play module, parallel computing, and provable convergence. Extensive experiments show that our method achieves superior performance comparing with 14 clustering approaches on three challenging data sets. The source code could be accessed at www.pengxi.me.
Added
2026-09-26

GCFAgg: Global and Cross-View Feature Aggregation for Multi-View Clustering
Weiqing Yan, Yuanyang Zhang, Chenlei Lv, Chang Tang, Guanghui Yue, Liang Liao, Weisi Lin
Why you should read this
Proposes a global and cross-view feature aggregation framework that integrates transformer-based sample relationships with structure-guided contrastive learning to boost multi-view clustering performance on both complete and incomplete datasets.
Multi-view clustering can partition data samples into their categories by learning a consensus representation in unsupervised way and has received more and more attention in recent years. However, most existing deep clustering methods learn consensus representation or view-specific representations from multiple views via view-wise aggregation way, where they ignore structure relationship of all samples. In this paper, we propose a novel multi-view clustering network to address these problems, called Global and Cross-view Feature Aggregation for Multi-View Clustering (GCFAggMVC). Specifically, the consensus data presentation from multiple views is obtained via cross-sample and cross-view feature aggregation, which fully explores the complementary of similar samples. Moreover, we align the consensus representation and the view-specific representation by the structure-guided contrastive learning module, which makes the view-specific representations from different samples with high structure relationship similar. The proposed module is a flexible multi-view data representation module, which can be also embedded to the incomplete multi-view data clustering task via plugging our module into other frameworks. Extensive experiments show that the proposed method achieves excellent performance in both complete multi-view data clustering tasks and incomplete multi-view data clustering tasks.
Added
2026-09-26

Cluster analysis for gene expression data: a survey
Daxin Jiang, Chun Tang, Aidong Zhang
Why you should read this
Categorizes clustering methods for microarray gene expression data into gene-based, sample-based, and subspace approaches while reviewing specific algorithms, proximity measures, and validation techniques to guide functional genomics research.
DNA microarray technology has now made it possible to simultaneously monitor the expression levels of thousands of genes during important biological processes and across collections of related samples. Elucidating the patterns hidden in gene expression data offers a tremendous opportunity for an enhanced understanding of functional genomics. However, the large number of genes and the complexity of biological networks greatly increases the challenges of comprehending and interpreting the resulting mass of data, which often consists of millions of measurements. A first step toward addressing this challenge is the use of clustering techniques, which is essential in the data mining process to reveal natural structures and identify interesting patterns in the underlying data. Cluster analysis seeks to partition a given data set into groups based on specified features so that the data points within a group are more similar to each other than the points in different groups. A very rich literature on cluster analysis has developed over the past three decades. Many conventional clustering algorithms have been adapted or directly applied to gene expression data, and also new algorithms have recently been proposed specifically aiming at gene expression data. These clustering algorithms have been proven useful for identifying biologically relevant groups of genes and samples. In this paper, we first briefly introduce the concepts of microarray technology and discuss the basic elements of clustering on gene expression data. In particular, we divide cluster analysis for gene expression data into three categories. Then, we present specific challenges pertinent to each clustering category and introduce several representative approaches. We also discuss the problem of cluster validation in three aspects and review various methods to assess the quality and reliability of clustering results. Finally, we conclude this paper and suggest the promising trends in this field.
Added
2026-09-25

Abnormal Event Detection at 150 FPS in MATLAB
Cewu Lu, Jianping Shi, Jiaya Jia
Why you should read this
Proposes a sparse combination learning framework that replaces costly sparse coding with small-scale least-squares projections, enabling abnormal event detection in surveillance video at 150 frames per second in MATLAB without sacrificing accuracy.
Speedy abnormal event detection meets the growing demand to process an enormous number of surveillance videos. Based on inherent redundancy of video structures, we propose an efficient sparse combination learning framework. It achieves decent performance in the detection phase without compromising result quality. The short running time is guaranteed because the new method effectively turns the original complicated problem to one in which only a few costless small-scale least square optimization steps are involved. Our method reaches high detection rates on benchmark datasets at a speed of 140~150 frames per second on average when computing on an ordinary desktop PC using MATLAB.
Added
2026-09-25

Robust Subspace Segmentation by Low-Rank Representation
Guangcan Liu, Zhouchen Lin, Yong Yu
Why you should read this
Proposes a low-rank representation framework that jointly captures global data structures via nuclear norm minimization to accurately segment data lying across multiple subspaces even when severely corrupted by noise and outliers.
We propose low-rank representation (LRR) to segment data drawn from a union of multiple linear (or affine) subspaces. Given a set of data vectors, LRR seeks the lowest-rank representation among all the candidates that represent all vectors as the linear combination of the bases in a dictionary. Unlike the well-known sparse representation (SR), which computes the sparsest representation of each data vector individually, LRR aims at finding the lowest-rank representation of a collection of vectors jointly. LRR better captures the global structure of data, giving a more effective tool for robust subspace segmentation from corrupted data. Both theoretical and experimental results show that LRR is a promising tool for subspace segmentation.
Added
2026-09-24

Sparse Subspace Clustering: Algorithm, Theory, and Applications
Ehsan Elhamifar, Rene Vidal
Why you should read this
Introduces Sparse Subspace Clustering, a principled optimization framework that uses sparse representation and spectral clustering to group high-dimensional data across intersecting subspaces while effectively handling noise, outliers, and missing entries.
In many real-world problems, we are dealing with collections of high-dimensional data, such as images, videos, text and web documents, DNA microarray data, and more. Often, high-dimensional data lie close to low-dimensional structures corresponding to several classes or categories the data belongs to. In this paper, we propose and study an algorithm, called Sparse Subspace Clustering (SSC), to cluster data points that lie in a union of low-dimensional subspaces. The key idea is that, among infinitely many possible representations of a data point in terms of other points, a sparse representation corresponds to selecting a few points from the same subspace. This motivates solving a sparse optimization program whose solution is used in a spectral clustering framework to infer the clustering of data into subspaces. Since solving the sparse optimization program is in general NP-hard, we consider a convex relaxation and show that, under appropriate conditions on the arrangement of subspaces and the distribution of data, the proposed minimization program succeeds in recovering the desired sparse representations. The proposed algorithm can be solved efficiently and can handle data points near the intersections of subspaces. Another key advantage of the proposed algorithm with respect to the state of the art is that it can deal with data nuisances, such as noise, sparse outlying entries, and missing entries, directly by incorporating the model of the data into the sparse optimization program. We demonstrate the effectiveness of the proposed algorithm through experiments on synthetic data as well as the two real-world problems of motion segmentation and face clustering.
Added
2026-09-14
