keyword
Document clustering
Document clustering is an unsupervised machine learning and natural language processing technique that automatically organizes a collection of text documents into distinct groups based on their content similarity, without relying on predefined category labels. To perform this grouping, documents are typically converted into numerical representations, such as high-dimensional vectors or term-document matrices based on vocabulary frequency, term weighting schemes, or latent semantic features. Various algorithms, including partitioning methods, hierarchical clustering, and matrix factorization, analyze these mathematical representations to discover underlying thematic patterns. By maximizing the similarity of documents within the same cluster while minimizing the similarity across different clusters, document clustering facilitates large-scale text organization, information retrieval, automated topic discovery, and text mining.
4 items

Orthogonal nonnegative matrix t-factorizations for clustering
C. Ding, Tao Li, Wei Peng, Haesun Park
Why you should read this
Establishes a rigorous mathematical foundation and convergent update algorithms for orthogonal three-factor nonnegative matrix factorization, enabling simultaneous, interpretable co-clustering of rows and columns in complex data matrices.
Currently, most research on nonnegative matrix factorization (NMF) focus on 2-factor X = FG^T factorization. We provide a systematic analysis of 3-factor X = FSG^T NMF. While unconstrained 3-factor NMF is equivalent to unconstrained 2-factor NMF, constrained 3-factor NMF brings new features to constrained 2-factor NMF. We study the orthogonality constraint because it leads to rigorous clustering interpretation. We provide new rules for updating F,S,G and prove the convergence of these algorithms. Experiments on 5 datasets and a real world case study are performed to show the capability of bi-orthogonal 3-factor NMF on simultaneously clustering rows and columns of the input data matrix. We provide a new approach of evaluating the quality of clustering on words using class aggregate distribution and multi-peak distribution. We also provide an overview of various NMF extensions and examine their relationships.
Added
2026-09-25

Concept Decompositions for Large Sparse Text Data Using Clustering
I. Dhillon, D. Modha
Why you should read this
Demonstrates that spherical k-means clustering yields sparse, localized concept decompositions that rival traditional singular value decomposition in matrix approximation accuracy while providing superior interpretability for high-dimensional text data.
Unlabeled document collections are becoming increasingly common and available; mining such data sets represents a major contemporary challenge. Using words as features, text documents are often represented as high-dimensional and sparse vectors–a few thousand dimensions and a sparsity of 95 to 99% is typical. In this paper, we study a certain spherical k-means algorithm for clustering such document vectors. The algorithm outputs k disjoint clusters each with a concept vector that is the centroid of the cluster normalized to have unit Euclidean norm. As our first contribution, we empirically demonstrate that, owing to the high-dimensionality and sparsity of the text data, the clusters produced by the algorithm have a certain “fractal-like” and “self-similar” behavior. As our second contribution, we introduce concept decompositions to approximate the matrix of document vectors; these decompositions are obtained by taking the least-squares approximation onto the linear subspace spanned by all the concept vectors. We empirically establish that the approximation errors of the concept decompositions are close to the best possible, namely, to truncated singular value decompositions. As our third contribution, we show that the concept vectors are localized in the word space, are sparse, and tend towards orthonormality. In contrast, the singular vectors are global in the word space and are dense. Nonetheless, we observe the surprising fact that the linear subspaces spanned by the concept vectors and the leading singular vectors are quite close in the sense of small principal angles between them. In conclusion, the concept vectors produced by the spherical k- means algorithm constitute a powerful sparse and localized “basis” for text data sets.
Added
2026-09-24

V-Measure: A Conditional Entropy-Based External Cluster Evaluation Measure
Andrew Rosenberg, Julia Hirschberg
Why you should read this
Introduces V-measure, an entropy-based external evaluation metric that transparently balances clustering homogeneity and completeness without depending on dataset size or cluster counts.
We present V-measure, an external entropy-based cluster evaluation measure. V-measure provides an elegant solution to many problems that affect previously defined cluster evaluation measures including 1) dependence on clustering algorithm or data set, 2) the “problem of matching”, where the clustering of only a portion of data points are evaluated and 3) accurate evaluation and combination of two desirable aspects of clustering, homogeneity and completeness. We compare V-measure to a number of popular cluster evaluation measures and demonstrate that it satisfies several desirable properties of clustering solutions, using simulated clustering results. Finally, we use V-measure to evaluate two clustering tasks: document clustering and pitch accent type clustering.
Added
2026-09-18

Document clustering based on non-negative matrix factorization
Wei Xu, Xin Liu, Yihong Gong
In this paper, we propose a novel document clustering method based on the non-negative factorization of the term-document matrix of the given document corpus. In the latent semantic space derived by the non-negative matrix factorization (NMF), each axis captures the base topic of a particular document cluster, and each document is represented as an additive combination of the base topics. The cluster membership of each document can be easily determined by finding the base topic (the axis) with which the document has the largest projection value. Our experimental evaluations show that the proposed document clustering method surpasses the latent semantic indexing and the spectral clustering methods not only in the easy and reliable derivation of document clustering results, but also in document clustering accuracies.
Source
https://courses.cs.umbc.edu/graduate/CMSC601/Spring11/HW/hw2/HW2_papers/document_clustering.pdfAdded
2026-09-16
