TopicGPT: A Prompt-based Topic Modeling Framework
Chau PhamAlexander Miserlis HoyleSimeng SunPhilip ResnikMohit Iyyer
Proposes a prompt-based framework using large language models to generate interpretable natural language topic labels, textual descriptions, and verifiable document quotes, significantly outperforming traditional methods like LDA and BERTopic on human alignment benchmarks without requiring model retraining.
Organizations routinely rely on automated topic modeling to analyze and categorize large volumes of unstructured text data. However, traditional approaches such as Latent Dirichlet Allocation generate topics as ambiguous collections of words that require time-consuming manual interpretation and offer little direct control over topic structure. The article introduces and evaluates TopicGPT, a prompt-based framework that uses large language models to generate intuitive topic labels with natural language descriptions and assign them to documents with verifiable text evidence.
The framework operates in two core stages: topic generation and topic assignment. During topic generation, an advanced language model processes a representative sample of documents alongside a small set of example topics to produce high-level candidate topics. The framework then merges near duplicates using sentence similarity and removes infrequent topics before using a secondary model to assign the final topics and supporting quotes to individual documents. The authors benchmarked this framework against standard methods, including traditional and neural topic models, across two diverse datasets: a corpus of 14,290 Wikipedia articles and a collection of 32,661 United States Congressional bill summaries.
The evaluation revealed several clear findings. First, TopicGPT significantly outperformed all baseline topic models in aligning with human-annotated ground-truth categories. For example, it achieved a harmonic mean purity of 0.74 on Wikipedia articles compared to 0.64 for the strongest baseline, Latent Dirichlet Allocation, and 0.57 versus 0.52 on congressional bills. Second, human evaluations showed that TopicGPT generated topics with far fewer semantic errors, producing only 30.3% misaligned topics on Wikipedia data compared to 62.4% for traditional modeling. Third, the framework demonstrated high operational stability across different document subsets, prompt variations, and document orderings. Finally, tests with open-source models showed that while smaller models can effectively perform topic assignment, sophisticated models like GPT-4 remain necessary to successfully generate coherent, well-structured topic lists.
These results demonstrate that using language models for topic discovery can significantly reduce the manual effort required to decipher statistical topic clusters while enhancing auditability through cited textual evidence. Organizations can deploy this framework directly for content analysis without custom model training. The article recommends providing a concise seed list of two to three high-quality example topics to steer topic scope, using multi-label assignment prompts for documents with overlapping themes, and tuning refinement thresholds to avoid filtering out rare but relevant topics. Teams should sample approximately 600 to 1,000 documents for topic generation to minimize processing costs before running full assignment.
Decision-makers should nevertheless consider specific operational constraints. Running commercial language model APIs over large collections incurs variable costs—ranging from approximately 155 per dataset evaluated—and relies on proprietary models with limited architectural transparency. Additionally, long documents may require truncation to fit model context windows, and performance has not yet been established on non-English corpora. Despite these boundaries, confidence in TopicGPT's ability to produce robust, interpretable, and human-aligned topic hierarchies remains high for English-language textual analysis.
- Paper: Latent Dirichlet Allocation, David M. Blei et al. (2003). Introduces Latent Dirichlet Allocation (LDA), the foundational bag-of-words topic modeling framework whose interpretability and control limitations TopicGPT directly aims to overcome.
- Paper: Reading Tea Leaves: How Humans Interpret Topic Models, Jonathan D. Chang et al. (2009). Pioneered the study of human interpretation and evaluation of topic models, formulating the 'reading the tea leaves' problem cited by TopicGPT.
- Paper: BERTopic: Neural topic modeling with a class-based TF-IDF procedure, Maarten Grootendorst (2022). Presents BERTopic, a prominent neural topic modeling baseline utilizing transformer embeddings and clustering that serves as a key contemporary point of comparison.
- Paper: Exploring the Space of Topic Coherence Measures, Michael Röder et al. (2015). Systematizes automated topic coherence and evaluation metrics used to quantify topic interpretability against human judgment.
- Paper: Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing, Pengfei Liu et al. (2021). Provides a comprehensive taxonomy of prompt-based learning methods that underpin TopicGPT's prompt-driven paradigm.
- Paper: Language Models are Few-Shot Learners, T. B. Brown et al. (2020). Establishes in-context and few-shot prompting capabilities in large language models, enabling prompt-based extraction of latent concepts without parameter updates.
No sufficiently relevant recommendations were found.
