keyword
topic modeling
Topic modeling is an unsupervised natural language processing and machine learning technique used to automatically discover latent themes, concepts, and semantic patterns across a collection of text documents. By analyzing statistical associations and co-occurrences of words and phrases, topic modeling algorithms group related vocabulary into distinct topics and characterize individual documents by the distribution of topics they contain. Traditional implementations rely on statistical and probabilistic frameworks such as Latent Dirichlet Allocation to identify clusters of co-occurring words, while modern approaches incorporate contextual embeddings and neural language models to capture more nuanced semantic relationships. The technique is widely applied in text mining, information retrieval, and exploratory data analysis to organize, summarize, and navigate large volumes of unstructured text without requiring manual annotation.
4 items

TopicGPT: A Prompt-based Topic Modeling Framework
Chau Pham, Alexander Miserlis Hoyle, Simeng Sun, Philip Resnik, Mohit Iyyer
Why you should read this
Proposes a prompt-based framework using large language models to generate interpretable natural language topic labels, textual descriptions, and verifiable document quotes, significantly outperforming traditional methods like LDA and BERTopic on human alignment benchmarks without requiring model retraining.
Topic modeling is a well-established technique for exploring text corpora. Conventional topic models (e.g., LDA) represent topics as bags of words that often require “reading the tea leaves” to interpret; additionally, they offer users minimal control over the formatting and specificity of resulting topics. To tackle these issues, we introduce TopicGPT, a prompt-based framework that uses large language models (LLMs) to uncover latent topics in a text collection. TopicGPT produces topics that align better with human categorizations compared to competing methods: it achieves a harmonic mean purity of 0.74 against human-annotated Wikipedia topics compared to 0.64 for the strongest baseline. Its topics are also interpretable, dispensing with ambiguous bags of words in favor of topics with natural language labels and associated free-form descriptions. Moreover, the framework is highly adaptable, allowing users to specify constraints and modify topics without the need for model retraining. By streamlining access to high-quality and interpretable topics, TopicGPT represents a compelling, human-centered approach to topic modeling.
Added
2026-09-28

UCTopic: Unsupervised Contrastive Learning for Phrase Representations and Topic Mining
Jiacheng Li, Jingbo Shang, Julian J. McAuley
Why you should read this
Presents an unsupervised contrastive learning framework that pairs masked-phrase contexts and applies cluster-assisted negative sampling to learn context-aware phrase representations and extract coherent topical terms without labeled data.
High-quality phrase representations are essential to finding topics and related terms in documents (a.k.a. topic mining). Existing phrase representation learning methods either simply combine unigram representations in a context-free manner or rely on extensive annotations to learn context-aware knowledge. In this paper, we propose UCTOPIC, a novel unsupervised contrastive learning framework for context-aware phrase representations and topic mining. UCTOPIC is pretrained in a large scale to distinguish if the contexts of two phrase mentions have the same semantics. The key to pretraining is positive pair construction from our phrase-oriented assumptions. However, we find traditional in-batch negatives cause performance decay when finetuning on a dataset with small topic numbers. Hence, we propose cluster-assisted contrastive learning (CCL) which largely reduces noisy negatives by selecting negatives from clusters and further improves phrase representations for topics accordingly. UCTOPIC outperforms the state-of-the-art phrase representation model by 38.2% NMI in average on four entity clustering tasks. Comprehensive evaluation on topic mining shows that UCTOPIC can extract coherent and diverse topical phrases.
Added
2026-09-26

Latent Dirichlet allocation (LDA) and topic modeling: models, applications, a survey
Hamed Jelodar, Yongli Wang, Chi Yuan, Xia Feng, Xia-Hui Jiang, Yanchao Li, Liang Zhao
Why you should read this
Synthesizes over a decade of Latent Dirichlet Allocation research by categorizing major model extensions, cross-disciplinary applications, benchmark datasets, and practical software tools for text mining.
Topic modeling is one of the most powerful techniques in text mining for data mining, latent data discovery, and finding relationships among data, text documents. Researchers have published many articles in the field of topic modeling and applied in various fields such as software engineering, political science, medical and linguistic science, etc. There are various methods for topic modeling, which Latent Dirichlet allocation (LDA) is one of the most popular methods in this field. Researchers have proposed various models based on the LDA in topic modeling. According to previous work, this paper can be very useful and valuable for introducing LDA approaches in topic modeling. In this paper, we investigated scholarly articles highly (between 2003 to 2016) related to Topic Modeling based on LDA to discover the research development, current trends and intellectual structure of topic modeling. Also, we summarize challenges and introduce famous tools and datasets in topic modeling based on LDA.
Added
2026-09-20

The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, Connor Leahy
Why you should read this
Introduces an 825 GiB open-source dataset combining 22 specialized text sources to address the limitations of standard web crawls and improve the cross-domain generalization of large language models.
Recent work has demonstrated that increased training dataset diversity improves general cross-domain knowledge and downstream generalization capability for large-scale language models. With this in mind, we present \textit{the Pile}: an 825 GiB English text corpus targeted at training large-scale language models. The Pile is constructed from 22 diverse high-quality subsets -- both existing and newly constructed -- many of which derive from academic or professional sources. Our evaluation of the untuned performance of GPT-2 and GPT-3 on the Pile shows that these models struggle on many of its components, such as academic writing. Conversely, models trained on the Pile improve significantly over both Raw CC and CC-100 on all components of the Pile, while improving performance on downstream evaluations. Through an in-depth exploratory analysis, we document potentially concerning aspects of the data for prospective users. We make publicly available the code used in its construction.
Added
2026-09-13

