ClusterLLM: Large Language Models as a Guide for Text Clustering
Yuwei ZhangZihan WangJingbo Shang
Proposes a cost-effective framework that queries black-box large language models through informative triplet and pairwise comparisons to guide small embedders and determine cluster granularity based on user preferences.
Text clustering is a critical capability for organizing large volumes of unstructured textual data into meaningful groups, supporting applications such as customer intent discovery, topic tracking, and risk analysis. Traditional clustering relies on smaller embedding models that map text into geometric space, but these models often misinterpret human nuances or fail to group items according to specific business goals. While state-of-the-art large language models offer superior semantic understanding, they cannot be directly used for text clustering because their internal vector embeddings are kept private and inaccessible behind commercial application programming interfaces.
The article introduces and evaluates CLUSTERLLM, an automated framework designed to harness the advanced reasoning of large language models to guide smaller, accessible embedding models. The primary objective is to demonstrate that large models can steer clustering toward user-preferred perspectives and identify the optimal number of clusters at an exceptionally low operational cost.
To achieve this without expensive exhaustive searches, the approach splits clustering into two distinct, cost-effective stages. In the first stage, the framework identifies the most ambiguous data points using an entropy-based sampling strategy and asks the large language model targeted comparison questions—specifically, which of two candidate texts is closer in meaning to an anchor text based on a given instruction. These answers are then used to fine-tune the smaller embedding model. In the second stage, the framework samples pairs across different levels of a cluster hierarchy and prompts the large model with a few annotated examples to determine whether pairs belong in the same group, thereby pinpointing the ideal cluster granularity. The authors evaluated this system across 14 diverse benchmark datasets spanning intent recognition, domain discovery, topic mining, and emotion detection, testing both small datasets and large collections of up to 50,000 texts.
The evaluation produced several key findings. First, guiding smaller models with large language model feedback consistently boosted clustering accuracy and alignment metrics across 14 benchmarks, outperforming traditional unsupervised deep clustering and self-supervised baselines. For instance, clustering accuracy improved by over 9 percentage points on banking intent data and nearly 7 percentage points on relation type data. Second, the entropy-based sampling strategy proved vital; querying the large model on only 1,024 strategically selected ambiguous examples successfully fine-tuned models on datasets as large as 50,000 records, whereas random sampling degraded performance. Third, the pairwise hierarchy method accurately determined the correct number of clusters, distinguishing coarse domains from fine-grained intents far better than standard statistical criteria. Finally, the framework achieved these gains at a negligible average cost of approximately $0.61 per dataset using standard commercial interfaces.
These findings demonstrate that organizations can achieve near-frontier language model accuracy on unsupervised grouping tasks without bearing the prohibitive financial costs of processing every data pair through commercial interfaces. It minimizes human labeling effort by relying only on high-level instructions and a handful of demonstration examples. Additionally, it enables teams to flexibly adapt the same text repository to multiple business use cases—such as sorting customer inquiries either by overarching department or by precise root cause—by merely adjusting prompt instructions.
Organizations handling large-scale text categorization should consider piloting this guided clustering framework before committing resources to extensive manual annotation or large-scale internal model training. When implementing, practitioners should prioritize using well-structured few-shot examples with brief rationales during the granularity stage, as demonstrations significantly improve cluster count estimation. If higher accuracy is needed, teams can run the framework iteratively, using fine-tuned models to sample progressively more challenging triplets.
The findings are supported by consistent results across diverse domains, yet some operational limitations remain. The method currently experiences sub-optimal performance on very coarse groupings, such as broad domain discovery, where fine-tuning creates overly compact clusters that struggle with broad aggregation. Furthermore, the framework still requires local computational capacity to fine-tune the underlying small embedding model and relies on external cloud interfaces, which necessitates careful consideration of data privacy when processing sensitive records.
- Paper: Text Embeddings by Weakly-Supervised Contrastive Pre-training, Liang Wang et al. (2022). E5 establishes how contrastively trained text embeddings support clustering and related tasks, clarifying the small-embedder baseline that ClusterLLM guides.
- Paper: MetaICL: Learning to Learn In Context, Sewon Min et al. (2022). MetaICL explains how models learn task behavior from in-context examples, a foundation for understanding ClusterLLM’s use of prompted LLM judgments.
- Paper: TopicGPT: A Prompt-based Topic Modeling Framework, Chau Pham et al. (2024). TopicGPT shows how prompted LLMs can structure and categorize document collections, providing a direct precursor to LLM-guided text clustering.
- Paper: Scalable Model-Based Clustering with Sequential Monte Carlo, Connie Trojan et al. (2026). This later work carries model-guided clustering into scalable online inference, extending ClusterLLM’s use of learned models to support clustering decisions.
