ClusterLLM: Large Language Models as a Guide for Text Clustering

Yuwei ZhangZihan WangJingbo Shang

article2023EMNLP142 citations

Proposes a cost-effective framework that queries black-box large language models through informative triplet and pairwise comparisons to guide small embedders and determine cluster granularity based on user preferences.

Listen

Text clustering is a critical capability for organizing large volumes of unstructured textual data into meaningful groups, supporting applications such as customer intent discovery, topic tracking, and risk analysis. Traditional clustering relies on smaller embedding models that map text into geometric space, but these models often misinterpret human nuances or fail to group items according to specific business goals. While state-of-the-art large language models offer superior semantic understanding, they cannot be directly used for text clustering because their internal vector embeddings are kept private and inaccessible behind commercial application programming interfaces.

The article introduces and evaluates CLUSTERLLM, an automated framework designed to harness the advanced reasoning of large language models to guide smaller, accessible embedding models. The primary objective is to demonstrate that large models can steer clustering toward user-preferred perspectives and identify the optimal number of clusters at an exceptionally low operational cost.

To achieve this without expensive exhaustive searches, the approach splits clustering into two distinct, cost-effective stages. In the first stage, the framework identifies the most ambiguous data points using an entropy-based sampling strategy and asks the large language model targeted comparison questions—specifically, which of two candidate texts is closer in meaning to an anchor text based on a given instruction. These answers are then used to fine-tune the smaller embedding model. In the second stage, the framework samples pairs across different levels of a cluster hierarchy and prompts the large model with a few annotated examples to determine whether pairs belong in the same group, thereby pinpointing the ideal cluster granularity. The authors evaluated this system across 14 diverse benchmark datasets spanning intent recognition, domain discovery, topic mining, and emotion detection, testing both small datasets and large collections of up to 50,000 texts.

The evaluation produced several key findings. First, guiding smaller models with large language model feedback consistently boosted clustering accuracy and alignment metrics across 14 benchmarks, outperforming traditional unsupervised deep clustering and self-supervised baselines. For instance, clustering accuracy improved by over 9 percentage points on banking intent data and nearly 7 percentage points on relation type data. Second, the entropy-based sampling strategy proved vital; querying the large model on only 1,024 strategically selected ambiguous examples successfully fine-tuned models on datasets as large as 50,000 records, whereas random sampling degraded performance. Third, the pairwise hierarchy method accurately determined the correct number of clusters, distinguishing coarse domains from fine-grained intents far better than standard statistical criteria. Finally, the framework achieved these gains at a negligible average cost of approximately $0.61 per dataset using standard commercial interfaces.

These findings demonstrate that organizations can achieve near-frontier language model accuracy on unsupervised grouping tasks without bearing the prohibitive financial costs of processing every data pair through commercial interfaces. It minimizes human labeling effort by relying only on high-level instructions and a handful of demonstration examples. Additionally, it enables teams to flexibly adapt the same text repository to multiple business use cases—such as sorting customer inquiries either by overarching department or by precise root cause—by merely adjusting prompt instructions.

Organizations handling large-scale text categorization should consider piloting this guided clustering framework before committing resources to extensive manual annotation or large-scale internal model training. When implementing, practitioners should prioritize using well-structured few-shot examples with brief rationales during the granularity stage, as demonstrations significantly improve cluster count estimation. If higher accuracy is needed, teams can run the framework iteratively, using fine-tuned models to sample progressively more challenging triplets.

The findings are supported by consistent results across diverse domains, yet some operational limitations remain. The method currently experiences sub-optimal performance on very coarse groupings, such as broad domain discovery, where fine-tuning creates overly compact clusters that struggle with broad aggregation. Furthermore, the framework still requires local computational capacity to fine-tune the underlying small embedding model and relies on external cloud interfaces, which necessitates careful consideration of data privacy when processing sensitive records.

Cover for ClusterLLM: Large Language Models as a Guide for Text Clustering

Abstract

We introduce ClusterLLM, a novel text clustering framework that leverages feedback from an instruction-tuned large language model, such as ChatGPT. Compared with traditional unsupervised methods that builds upon “small” embedders, ClusterLLM exhibits two intriguing advantages: (1) it enjoys the emergent capability of LLM even if its embeddings are inaccessible; and (2) it understands the user’s preference on clustering through textual instruction and/or a few annotated data. First, we prompt ChatGPT for insights on clustering perspective by constructing hard triplet questions <does A better correspond to B than C>, where A, B and C are similar data points that belong to different clusters according to small embedder. We empirically show that this strategy is both effective for fine-tuning small embedder and cost-efficient to guide ChatGPT. Second, we prompt ChatGPT for helps on clustering granularity by carefully designed pairwise questions <do A and B belong to the same category>, and tune the granularity from cluster hierarchies that is the most consistent with the ChatGPT answers. Extensive experiments on 14 datasets show that ClusterLLM consistently improves clustering quality, at an average cost of ~$0.6¹ per dataset. The code will be available at https://github.com/zhang-yu-wei/ClusterLLM.

Table of Contents

  • 1 Introduction
  • 2 Preliminary
  • 3 Our CLUSTERLLM
  • 3.1 Triplet Task for Perspective
  • 3.1.1 Entropy-based Triplet Sampling
  • 3.1.2 Finetuning Embedder
  • 3.2 Pairwise Task for Granularity
  • 3.2.1 Determine Granularity with Pairwise Hierarchical Sampling
  • 4 Experiments
  • 4.1 Datasets
  • 4.2 Experiment Details
  • 4.3 Compared Methods
  • 4.4 Main Results
  • 4.5 Analysis on Triplet Prediction Accuracy
  • 4.6 Ablation Study
  • 4.7 Determining Cluster Granularity
  • 5 Related Works
  • 6 Conclusion
  • Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • A Details of Scaling up Hierarchical Clustering
  • B More Details about Determining Cluster Granularity
  • C Analysis for Determining Granularity
  • D Details of Embedders and Fine-tuning
  • E Description of Datasets
  • F Results of More Iterations
  • G More Related Works
  • H Sub-optimal Performance on Domain Discovery
  • I Dataset Leakage

Knowls

  1. Knowl 1 — Two-stage LLM-guided text clustering

    model/method

    CLUSTERLLM uses an API-accessible instruction-tuned large language model (LLM) to guide a smaller pretrained text embedder whose vectors can be used for clustering. It addresses two distinct choices in clustering: the perspective, or criterion for grouping texts, and the granularity, or number and scope of groups. For perspective, the LLM compares an anchor text with two candidates under a user-provided task instruction; its preferences become triplet supervision for fine-tuning the embedder. For granularity, the LLM judges whether text pairs belong to the same category, with user-annotated examples as demonstrations; those judgments are compared with partitions in a cluster hierarchy to select a granularity. The approach therefore uses LLM judgments without requiring access to LLM embeddings, and can incorporate textual instructions and a few annotated examples.

  2. Knowl 2 — Entropy-guided selection of informative triplets

    algorithm

    For a corpus of NN texts with embeddings zi=f(xi)z_i=f(x_i), CLUSTERLLM selects ambiguous texts as triplet anchors, then draws candidate texts from their nearby clusters. Let KK be the number of clusters, μk\mu_k the mean embedding of cluster kk, and α=1\alpha=1. For text ii, the soft assignment to cluster kk is computed with a Student tt distribution:

    pik=(1+∥zi−μk∥2/α)−(α+1)/2∑k′(1+∥zi−μk′∥2/α)−(α+1)/2.p_{ik}=\frac{(1+\lVert z_i-\mu_k\rVert^2/\alpha)^{-(\alpha+1)/2}}{\sum_{k'}(1+\lVert z_i-\mu_{k'}\rVert^2/\alpha)^{-(\alpha+1)/2}}.

    The number of candidate clusters is Kclosest=max⁡(ϵK,2)K_{\mathrm{closest}}=\max(\epsilon K,2), where ϵ\epsilon is a small fraction. For each text, retain its KclosestK_{\mathrm{closest}} clusters with the largest assignments, renormalize their probabilities to pik′p'_{ik}, and compute entropy hi=−∑k=1Kclosestpik′log⁡pik′h_i=-\sum_{k=1}^{K_{\mathrm{closest}}}p'_{ik}\log p'_{ik}. Sort texts by decreasing entropy and retain ranks from γhighN\gamma_{\mathrm{high}}N through γlowN\gamma_{\mathrm{low}}N. For each retained anchor, sample two candidate clusters from its closest clusters and one text from each; discard a triplet if either candidate is the anchor or if the triplet is already present. Continue until the query budget QQ is met. In the experiments, ϵ=2%\epsilon=2\%, γhigh=0\gamma_{\mathrm{high}}=0, γlow=20%\gamma_{\mathrm{low}}=20\%, and Q=1,024Q=1{,}024.

    Input: Text embeddings, cluster assignments, interval boundaries γhigh\gamma_{\mathrm{high}} and γlow\gamma_{\mathrm{low}}, closest-cluster fraction ϵ\epsilon, query budget QQ
    Compute each cluster center as the mean of its assigned embeddings
    For each text, compute Student t soft assignments to all cluster centers
    For each text, retain its Kclosest=max⁡(ϵK,2)K_{\mathrm{closest}}=\max(\epsilon K,2) highest-scoring clusters
    Renormalize assignments over the retained clusters and compute the text's entropy
    Sort texts by decreasing entropy and keep ranks from γhighN\gamma_{\mathrm{high}}N to γlowN\gamma_{\mathrm{low}}N
    Initialize an empty set of triplets
    While fewer than QQ unique triplets have been collected:
        For each retained text as anchor:
            Sample two candidate clusters from its retained closest clusters
            Sample one text from each candidate cluster
            Add the anchor and candidates as a triplet if neither candidate equals the anchor and the triplet is new
    Return the collected triplets

    For the experiments, the initial clustering used agglomerative clustering with fixed distance thresholds for small datasets and mini-batch KK-means with 100 clusters for large datasets.

  3. Knowl 3 — Triplet preferences fine-tune the text embedder

    equation

    A triplet consists of an anchor text aa, a positive choice c+c^+ selected by the LLM, and a negative choice c−c^-. The LLM selects between the two choices according to a task instruction specifying the desired clustering perspective. Fine-tuning increases the probability assigned to the positive choice relative to a candidate set BB containing the positive, the hard negative, and other in-batch negatives:

    ℓj=exp⁡(s(a,c+)/τ)∑cl∈Bexp⁡(s(a,cl)/τ).\ell_j=\frac{\exp(s(a,c^+)/\tau)}{\sum_{c_l\in B}\exp(s(a,c_l)/\tau)}.

    Here s(u,v)s(u,v) is the embedder's similarity score for texts uu and vv, τ\tau is a temperature parameter, and jj indexes the triplet. The training procedure also computes the objective with aa and c+c^+ exchanged. Clustering assignments are obtained by applying a clustering algorithm to the fine-tuned embeddings; the fine-tuned embedder can also be used to sample new triplets for another training iteration.

  4. Knowl 4 — Pairwise judgments select a granularity from a cluster hierarchy

    model/method

    To choose cluster granularity, CLUSTERLLM constructs a hierarchy by starting with individual texts as clusters and repeatedly merging the closest clusters. A pairwise LLM prompt asks whether two texts belong to the same category, using typically four user-annotated demonstration pairs and a brief justification for each. The user-provided minimum and maximum cluster counts, kmin⁡k_{\min} and kmax⁡k_{\max}, bound the range of granularities to consider. At each merge step in this range, sample λ\lambda text pairs from the two clusters being merged, yielding Np=λ(kmax⁡−kmin⁡)N_p=\lambda(k_{\max}-k_{\min}) candidate pairs. Compare the LLM's same/different judgments WpW^p with the same/different labels WkW^k implied for those pairs by each hierarchy level kk, and choose

    k∗=arg⁡max⁡kM(Wp,Wk),k^*=\arg\max_k M(W^p,W^k),

    where MM is an FβF_\beta score, with LLM judgments treated as labels and hierarchy assignments as predictions. The experiments use λ∈{1,3}\lambda\in\{1,3\} and set the FβF_\beta weight to 0.920.92. Sampling pairs at merge steps exposes the LLM to examples spanning the candidate granularities. The method assumes that the hierarchy contains meaningful cluster structure: if its clusters are essentially random, accurate pairwise judgments alone cannot recover a useful granularity.

  5. Knowl 5 — A two-step hierarchy construction scales to larger datasets

    model/method

    Because full hierarchical clustering is described as having O(N3)O(N^3) time complexity, CLUSTERLLM avoids applying it directly to every instance in large datasets. It first runs mini-batch KK-means with kmax⁡k_{\max} clusters, then performs agglomerative clustering on those cluster assignments using Ward's method. It computes distances between cluster pairs and supplies them to a nearest-neighbor-chain algorithm to obtain the hierarchy, which is combined with the initial KK-means assignments. This construction is used to supply the range of hierarchy levels needed for pairwise granularity selection.

  6. Knowl 6 — Evaluation covers 14 datasets and a fixed LLM query budget

    experimental setup

    The evaluation spans 14 datasets covering intent discovery (Bank77, CLINC(I), MTOP(I), Massive(I)), type discovery (FewRel, FewNerd, FewEvent), topic mining (StackEx, ArxivS2S, Reddit), emotion detection (GoEmo), and domain discovery (CLINC(D), MTOP(D), Massive(D)). Each dataset has small- and large-scale versions with the same number of clusters; the datasets range from 10 to 150 clusters. The study uses the large versions of the Instructor and E5 embedders and GPT-3.5 for LLM predictions. Perspective tuning uses 1,024 triplet queries per dataset; reported generation settings include temperature 0.5 and a 10-token maximum. Clustering is evaluated with accuracy after Hungarian label alignment (ACC) and normalized mutual information (NMI), averaging over five KK-means seeds with the ground-truth cluster count. The paper reports an average GPT-3.5 cost of about 0.20perdatasetforperspectivetuningand0.20 per dataset for perspective tuning and 0.40 for granularity selection, or about $0.60 combined.

  7. Knowl 7 — LLM-guided tuning improves average clustering quality

    empirical result

    Across the 14 datasets, the LLM-guided methods improve average ACC and NMI over the pretrained embedders and the self-supervised triplet baseline. The table reports the averages from the known-granularity clustering evaluation; ACC is computed after Hungarian alignment, and higher values are better.

    Method Average ACC Average NMI
    E5 47.72 61.35
    Self-supervised E5 49.87 63.60
    CLUSTERLLM-E 50.77 64.77
    CLUSTERLLM-E, iterative 52.40 66.18
    Instructor 49.90 63.85
    Self-supervised Instructor 51.39 65.33
    CLUSTERLLM-Instructor 53.09 66.58
    CLUSTERLLM-Instructor, iterative 53.99 67.53

    The iterative variants repeat the framework using the previous fine-tuned embedder to sample new triplets and initialize the next fine-tuning round. Relative to the original Instructor, the iterative method raises average ACC by 4.09 points and average NMI by 3.68 points. Improvements are not universal: domain discovery, particularly Massive(D) and CLINC(D), is a reported weak area.

  8. Knowl 8 — Entropy sampling outperforms random triplet selection

    empirical result

    On the small-scale evaluation with Instructor as the embedder, replacing entropy-based triplet sampling with random triplet sampling lowers the 14-dataset average from 53.09 to 48.27 ACC and from 66.58 to 61.96 NMI. The compared methods use GPT-3.5 triplet predictions; the result supports the value of selecting informative examples rather than querying arbitrary triplets. The triplet-prediction analysis likewise finds that GPT-3.5 has higher accuracy than Instructor on entropy-selected triplets with ground-truth preferences on 13 of 14 datasets. For example, on FewRel, accuracy is 76.69% for GPT-3.5 versus 62.41% for Instructor over 266 ground-truth triplets. Random sampling yields far fewer ground-truth triplets in that dataset: 41.

  9. Knowl 9 — Pairwise judgments distinguish intent and domain granularities

    empirical result

    On eight small-scale datasets, granularity selection with GPT-3.5 and λ=3\lambda=3 achieves an aggregate rank of 2 among the evaluated methods, compared with rank 7 for BIC. The following results use kmax⁡=200k_{\max}=200 and kmin⁡=2k_{\min}=2. Each inferred cluster count is followed by its relative error in percent, as reported by the paper.

    Dataset Ground-truth clusters GPT-3.5, λ=3\lambda=3 BIC
    Bank77 77 64 (16.88) 123 (59.74)
    FewRel 64 46 (28.13) 58 (9.38)
    Massive(I) 59 52 (11.86) 56 (5.08)
    Massive(D) 18 37 (105.6) 60 (233.3)
    MTOP(I) 102 92 (9.80) 69 (32.35)
    MTOP(D) 11 18 (63.63) 64 (481.8)
    CLINC(I) 150 142 (5.33) 167 (11.33)
    CLINC(D) 10 107 (970.0) 176 (1660)

    A notable distinction is MTOP intent versus domain discovery: GPT-3.5 with λ=3\lambda=3 infers 92 and 18 clusters, respectively, while BIC infers 69 and 64. Granularity estimates can nevertheless remain far from the ground truth on some datasets, including CLINC(D).

  10. Knowl 10 — The framework depends on embedding geometry and has coarse-domain weaknesses

    limitation

    CLUSTERLLM relies on a pretrained embedder both to identify ambiguous anchors for triplet queries and to provide the embedding space used for clustering. It is therefore inapplicable to black-box embedding models when their vectors cannot be used, and the authors identify fine-tuning as a computational cost. The paper also reports weak or negative results on some domain-discovery datasets. It offers, without claiming a rigorous explanation, the hypothesis that fine-tuning can make embedding clusters more compact and separated in ways that help fine-grained clustering but hurt coarser groupings. Separately, sending text to an external LLM API can pose a privacy risk when the data are sensitive.

Coverage note — Detailed dataset-construction procedures, prompt examples, and per-dataset ablation tables are omitted because the selected knowls retain the main method, evaluation scope, aggregate results, and principal limitations without reproducing supplementary detail.

References

  1. 1.Wenbin An, Feng Tian, Qinghua Zheng, Wei Ding, QianYing Wang, and Ping Chen. 2022. Generalized category discovery with decoupled prototypical network. arXiv preprint arXiv:2211.15115.
  2. 2.Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. 2022. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073.
  3. 3.Sugato Basu, Arindam Banerjee, and Raymond J Mooney. 2004. Active semi-supervision for pairwise constrained clustering. In Proceedings of the 2004 SIAM international conference on data mining, pages 333–344. SIAM.
  4. 4.Sugato Basu, Ian Davidson, and Kiri Wagstaff. 2008. Constrained clustering: Advances in algorithms, theory, and applications. CRC Press.
  5. 5.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  6. 6.Kaidi Cao, Maria Brbic, and Jure Leskovec. 2022. Open-world semi-supervised learning. In International Conference on Learning Representations.
  7. 7.Mathilde Caron, Piotr Bojanowski, Armand Joulin, and Matthijs Douze. 2018. Deep clustering for unsupervised learning of visual features. In Proceedings of the European conference on computer vision (ECCV), pages 132–149.
  8. 8.Iñigo Casanueva, Tadas Temcinas, Daniela Gerz, Matthew Henderson, and Ivan Vulic. 2020. Efficient intent detection with dual sentence encoders. In Proceedings of the 2nd Workshop on NLP for ConvAI - ACL 2020. Data available at https://github.com/PolyAI-LDN/task-specific-datasets.
  9. 9.Jianlong Chang, Lingfeng Wang, Gaofeng Meng, Shiming Xiang, and Chunhong Pan. 2017. Deep adaptive image clustering. In Proceedings of the IEEE international conference on computer vision, pages 5879–5887.
  10. 10.Jeff Cheeger. 1970. A lower bound for the smallest eigenvalue of the laplacian, problems in analysis (papers dedicated to salomon bochner, 1969).
  11. 11.Qinyuan Cheng, Xiaogui Yang, Tianxiang Sun, Linyang Li, and Xipeng Qiu. 2023. Improving contrastive learning of sentence embeddings from ai feedback.
  12. 12.Wei-Lin Chiang, Xuanqing Liu, Si Si, Yang Li, Samy Bengio, and Cho-Jui Hsieh. 2019. Cluster-gcn: An efficient algorithm for training deep and large graph convolutional networks. In Proceedings of the 25th ACM SIGKDD international conference on knowledge discovery & data mining, pages 257–266.
  13. 13.Pierre Colombo, Nathan Noiry, Ekhine Irurozki, and Stephan Clemencon. 2022. What are the best systems? new perspectives on nlp benchmarking. arXiv preprint arXiv:2202.03799.
  14. 14.Dorottya Demszky, Dana Movshovitz-Attias, Jeongwoo Ko, Alan Cowen, Gaurav Nemade, and Sujith Ravi. 2020. GoEmotions: A dataset of fine-grained emotions. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4040–4054, Online. Association for Computational Linguistics.
  15. 15.Shumin Deng, Ningyu Zhang, Jiaojian Kang, Yichi Zhang, Wei Zhang, and Huajun Chen. 2020. Meta-learning with dynamic-memory-based prototypical network for few-shot event detection. In Proceedings of the 13th International Conference on Web Search and Data Mining, WSDM ’20, page 151–159, New York, NY, USA. Association for Computing Machinery.
  16. 16.Nat Dilokthanakul, Pedro AM Mediano, Marta Garnelo, Matthew CH Lee, Hugh Salimbeni, Kai Arulkumaran, and Murray Shanahan. 2016. Deep unsupervised clustering with gaussian mixture variational autoencoders. arXiv preprint arXiv:1611.02648.
  17. 17.Ning Ding, Guangwei Xu, Yulin Chen, Xiaobin Wang, Xu Han, Pengjun Xie, Haitao Zheng, and Zhiyuan Liu. 2021. Few-NERD: A few-shot named entity recognition dataset. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3198–3213, Online. Association for Computational Linguistics.
  18. 18.William E Donath and Alan J Hoffman. 1972. Algorithms for partitioning of graphs and computer logic based on eigenvectors of connection matrices. IBM Technical Disclosure Bulletin, 15(3):938–944.
  19. 19.Jack FitzGerald, Christopher Hench, Charith Peris, Scott Mackie, Kay Rottmann, Ana Sanchez, Aaron Nash, Liam Urbach, Vishesh Kakarala, Richa Singh, Swetha Ranganath, Laurie Crist, Misha Britan, Wouter Leeuwis, Gokhan Tur, and Prem Natarajan. 2022. Massive: A 1m-example multilingual natural language understanding dataset with 51 typologically-diverse languages.
  20. 20.Tianyu Gao, Xu Han, Hao Zhu, Zhiyuan Liu, Peng Li, Maosong Sun, and Jie Zhou. 2019. FewRel 2.0: Towards more challenging few-shot relation classification. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 6251–6256, Hong Kong, China. Association for Computational Linguistics.
  21. 21.Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021. SimCSE: Simple contrastive learning of sentence embeddings. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6894–6910, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  22. 22.Gregor Geigle, Nils Reimers, Andreas Rücklé, and Iryna Gurevych. 2021. TWEAC: transformer with extendable QA agent classifiers. CoRR, abs/2104.07081.
  23. 23.Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. 2023. Chatgpt outperforms crowd-workers for text-annotation tasks. arXiv preprint arXiv:2303.15056.
  24. 24.Cyril Goutte, Lars Kai Hansen, Matthew G Liptrot, and Egill Rostrup. 2001. Feature-space clustering for fmri meta-analysis. Human brain mapping, 13(3):165–183.
  25. 25.Amir Hadifar, Lucas Sterckx, Thomas Demeester, and Chris Develder. 2019. A self-training approach for short text clustering. In Proceedings of the 4th Workshop on Representation Learning for NLP (RepL4NLP-2019), pages 194–199, Florence, Italy. Association for Computational Linguistics.
  26. 26.Xingwei He, Zhenghao Lin, Yeyun Gong, A Jin, Hang Zhang, Chen Lin, Jian Jiao, Siu Ming Yiu, Nan Duan, Weizhu Chen, et al. 2023. Annollm: Making large language models to be better crowdsourced annotators. arXiv preprint arXiv:2303.16854.
  27. 27.William Hogan, Jiacheng Li, and Jingbo Shang. 2023. Open-world semi-supervised generalized relation discovery aligned in a real-world setting.
  28. 28.Peihao Huang, Yan Huang, Wei Wang, and Liang Wang. 2014. Deep embedding network for clustering. In 2014 22nd International conference on pattern recognition, pages 1532–1537. IEEE.
  29. 29.Harold W Kuhn. 1955. The hungarian method for the assignment problem. Naval research logistics quarterly, 2(1-2):83–97.
  30. 30.Stefan Larson, Anish Mahendran, Joseph J. Peper, Christopher Clarke, Andrew Lee, Parker Hill, Jonathan K. Kummerfeld, Kevin Leach, Michael A. Laurenzano, Lingjia Tang, and Jason Mars. 2019. An evaluation dataset for intent classification and out-of-scope prediction. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP).
  31. 31.Haoran Li, Abhinav Arora, Shuohui Chen, Anchit Gupta, Sonal Gupta, and Yashar Mehdad. 2021. MTOP: A comprehensive multilingual task-oriented semantic parsing benchmark. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 2950–2962, Online. Association for Computational Linguistics.
  32. 32.Sha Li, Heng Ji, and Jiawei Han. 2022. Open relation and event type discovery with type abstraction. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 6864–6877, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  33. 33.Ting-En Lin, Hua Xu, and Hanlei Zhang. 2020. Discovering new intents via constrained deep adaptive clustering with cluster refinement. In Thirty-Fourth AAAI Conference on Artificial Intelligence.
  34. 34.J MacQueen. 1967. Classification and analysis of multivariate observations. In 5th Berkeley Symp. Math. Statist. Probability, pages 281–297. University of California Los Angeles LA USA.
  35. 35.Laura Manduchi, Kieran Chin-Cheong, Holger Michel, Sven Wellmann, and Julia E Vogt. 2021. Deep conditional gaussian mixture model for constrained clustering. In Advances in Neural Information Processing Systems.
  36. 36.José Manuel Guaita Martínez, Patricia Carracedo, Dolores Gorgues Comas, and Carlos H Siemens. 2022. An analysis of the blockchain and covid-19 research landscape using a bibliometric study. Sustainable Technology and Entrepreneurship, 1(1):100006.
  37. 37.Yutao Mou, Keqing He, Yanan Wu, Zhiyuan Zeng, Hong Xu, Huixing Jiang, Wei Wu, and Weiran Xu. 2022. Disentangled knowledge transfer for OOD intent discovery with unified contrastive learning. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 46–53, Dublin, Ireland. Association for Computational Linguistics.
  38. 38.Niklas Muennighoff, Nouamane Tazi, Loïc Magne, and Nils Reimers. 2022. Mteb: Massive text embedding benchmark. arXiv preprint arXiv:2210.07316.
  39. 39.Fionn Murtagh and Pedro Contreras. 2011. Methods of hierarchical clustering. arXiv preprint arXiv:1105.0121.
  40. 40.Jianmo Ni, Gustavo Hernandez Abrego, Noah Constant, Ji Ma, Keith Hall, Daniel Cer, and Yinfei Yang. 2022a. Sentence-t5: Scalable sentence encoders from pre-trained text-to-text models. In Findings of the Association for Computational Linguistics: ACL 2022, pages 1864–1874, Dublin, Ireland. Association for Computational Linguistics.
  41. 41.Jianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai, Gustavo Hernandez Abrego, Ji Ma, Vincent Zhao, Yi Luan, Keith Hall, Ming-Wei Chang, and Yinfei Yang. 2022b. Large dual encoders are generalizable retrievers. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 9844–9855, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  42. 42.Chuang Niu, Hongming Shan, and Ge Wang. 2022. Spice: Semantic pseudo-labeling for image clustering. IEEE Transactions on Image Processing, 31:7264–7278.
  43. 43.Chuang Niu, Jun Zhang, Ge Wang, and Jimin Liang. 2020. Gatcluster: Self-supervised gaussian-attention network for image clustering. In European Conference on Computer Vision (ECCV).
  44. 44.Robert M Nosofsky. 2011. The generalized context model: An exemplar model of classification. Formal approaches in categorization, pages 18–39.
  45. 45.OpenAI. 2023. Gpt-4 technical report. ArXiv, abs/2303.08774.
  46. 46.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744.
  47. 47.June Young Park, Evan Mistur, Donghwan Kim, Yunjeong Mo, and Richard Hoefer. 2022. Toward human-centric urban infrastructure: Text mining for social media data to identify the public perception of covid-19 policy in transportation hubs. Sustainable Cities and Society, 76:103524.
  48. 48.F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. 2011. Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12:2825–2830.
  49. 49.Dan Pelleg, Andrew W Moore, et al. 2000. X-means: Extending k-means with efficient estimation of the number of clusters. In Icml, volume 1, pages 727–734.
  50. 50.Maarten De Raedt, Fréderic Godin, Thomas Demeester, and Chris Develder. 2023. Idas: Intent discovery with abstractive summarization.
  51. 51.Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3982–3992, Hong Kong, China. Association for Computational Linguistics.
  52. 52.Mamshad Nayeem Rizve, Navid Kardan, Salman Khan, Fahad Shahbaz Khan, and Mubarak Shah. 2022a. Openldn: Learning to discover novel classes for open-world semi-supervised learning. In European Conference on Computer Vision, pages 382–401. Springer.
  53. 53.Mamshad Nayeem Rizve, Navid Kardan, and Mubarak Shah. 2022b. Towards realistic semi-supervised learning. In European Conference on Computer Vision, pages 437–455. Springer.
  54. 54.Peter J Rousseeuw. 1987. Silhouettes: a graphical aid to the interpretation and validation of cluster analysis. Journal of computational and applied mathematics, 20:53–65.
  55. 55.Hongjin Su, Weijia Shi, Jungo Kasai, Yizhong Wang, Yushi Hu, Mari Ostendorf, Wen-tau Yih, Noah A. Smith, Luke Zettlemoyer, and Tao Yu. 2022. One embedder, any task: Instruction-finetuned text embeddings.
  56. 56.Yiyou Sun and Yixuan Li. 2022. Opencon: Open-world contrastive learning. In Transactions on Machine Learning Research.
  57. 57.Robert Thorndike. 1953. Who belongs in the family? Psychometrika, 18(4):267–276.
  58. 58.Laurens Van der Maaten and Geoffrey Hinton. 2008. Visualizing data using t-sne. Journal of machine learning research, 9(11).
  59. 59.Wouter Van Gansbeke, Simon Vandenhende, Stam atios Georgoulis, Marc Proesmans, and Luc Van Gool. 2020. Scan: Learning to classify images without labels. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part X, pages 268–285. Springer.
  60. 60.Sagar Vaze, Kai Han, Andrea Vedaldi, and Andrew Zisserman. 2022. Generalized category discovery. In IEEE Conference on Computer Vision and Pattern Recognition.
  61. 61.Kiri Wagstaff, Claire Cardie, Seth Rogers, Stefan Schrödl, et al. 2001. Constrained k-means clustering with background knowledge. In Icml, volume 1, pages 577–584.
  62. 62.Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei. 2022. Text embeddings by weakly-supervised contrastive pre-training. arXiv preprint arXiv:2212.03533.
  63. 63.Zihan Wang, Jingbo Shang, and Ruiqi Zhong. 2023. Goal-driven explainable clustering via language descriptions.
  64. 64.Joe H Ward Jr. 1963. Hierarchical grouping to optimize an objective function. Journal of the American statistical association, 58(301):236–244.
  65. 65.Junyuan Xie, Ross Girshick, and Ali Farhadi. 2016. Unsupervised deep embedding for clustering analysis. In International conference on machine learning, pages 478–487. PMLR.
  66. 66.Hui Xu, Yi Liu, Chi-Min Shu, Mingqi Bai, Mailidan Motalifu, Zhongxu He, Shuncheng Wu, Penggang Zhou, and Bing Li. 2022. Cause analysis of hot work accidents based on text mining and deep learning. Journal of Loss Prevention in the Process Industries, 76:104747.
  67. 67.Jiaming Xu, Peng Wang, Guanhua Tian, Bo Xu, Jun Zhao, Fangyuan Wang, and Hongwei Hao. 2015. Short text clustering via convolutional neural networks. In Proceedings of the 1st Workshop on Vector Space Modeling for Natural Language Processing, pages 62–69, Denver, Colorado. Association for Computational Linguistics.
  68. 68.Kouta Nakata Yaling Tao, Kentaro Takagi. 2021. Clustering-friendly representation learning via instance discrimination and feature decorrelation. Proceedings of ICLR 2021.
  69. 69.Jianwei Yang, Devi Parikh, and Dhruv Batra. 2016. Joint unsupervised learning of deep representations and image clusters. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5147–5156.
  70. 70.Muli Yang, Yuehua Zhu, Jiaping Yu, Aming Wu, and Cheng Deng. 2022. Divide and conquer: Compositional experts for generalized novel class discovery. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14268–14277.
  71. 71.Dejiao Zhang, Feng Nan, Xiaokai Wei, Shang-Wen Li, Henghui Zhu, Kathleen McKeown, Ramesh Nallapati, Andrew O. Arnold, and Bing Xiang. 2021a. Supporting clustering with contrastive learning. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5419–5430, Online. Association for Computational Linguistics.
  72. 72.Hanlei Zhang, Hua Xu, Ting-En Lin, and Rui Lyu. 2021b. Discovering new intents with deep aligned clustering. Proceedings of the AAAI Conference on Artificial Intelligence, 35(16):14365–14373.
  73. 73.Hanlei Zhang, Hua Xu, Xin Wang, Fei Long, and Kai Gao. 2023. Usnid: A framework for unsupervised and semi-supervised new intent discovery. arXiv preprint arXiv:2304.07699.
  74. 74.Hongjing Zhang, Sugato Basu, and Ian Davidson. 2020. A framework for deep constrained clustering-algorithms and advances. In Machine Learning and Knowledge Discovery in Databases: European Conference, ECML PKDD 2019, Würzburg, Germany, September 16–20, 2019, Proceedings, Part I, pages 57–72. Springer.
  75. 75.Yuwei Zhang, Haode Zhang, Li-Ming Zhan, Xiao-Ming Wu, and Albert Lam. 2022. New intent discovery with pre-training and contrastive learning. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 256–269, Dublin, Ireland. Association for Computational Linguistics.
  76. 76.Sheng Zhou, Hongjia Xu, Zhuonan Zheng, Jiawei Chen, Jiajun Bu, Jia Wu, Xin Wang, Wenwu Zhu, Martin Ester, et al. 2022. A comprehensive survey on deep clustering: Taxonomy, challenges, and future directions. arXiv preprint arXiv:2206.07579.
  77. 77.Yiming Zhu, Peixian Zhang, Ehsan-Ul Haq, Pan Hui, and Gareth Tyson. 2023. Can chatgpt reproduce human-generated labels? a study of social computing tasks. arXiv preprint arXiv:2304.10145.

Citation

MLA
Zhang, Y., et al. “ClusterLLM: Large Language Models as a Guide for Text Clustering”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 13903–20, https://doi.org/10.18653/v1/2023.emnlp-main.858.
APA
Zhang, Y., Wang, Z., & Shang, J. (2023). ClusterLLM: Large Language Models as a Guide for Text Clustering. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 13903–13920. https://doi.org/10.18653/v1/2023.emnlp-main.858
Chicago
Zhang, Y., Z. Wang, and J. Shang. 2023. “ClusterLLM: Large Language Models as a Guide for Text Clustering”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 13903–20. https://doi.org/10.18653/v1/2023.emnlp-main.858.
Harvard
Zhang, Y., Wang, Z. and Shang, J. (2023) “ClusterLLM: Large Language Models as a Guide for Text Clustering”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 13903–13920. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.858.
Vancouver
1. Zhang Y, Wang Z, Shang J (2023) ClusterLLM: Large Language Models as a Guide for Text Clustering. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 13903–13920

BibTeX

@inproceedings{zhang-etal-2023-clusterllm,
    title = "{C}luster{LLM}: Large Language Models as a Guide for Text Clustering",
    author = "Zhang, Yuwei  and
      Wang, Zihan  and
      Shang, Jingbo",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.858/",
    doi = "10.18653/v1/2023.emnlp-main.858",
    pages = "13903--13920"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/