TopicGPT: A Prompt-based Topic Modeling Framework

Chau PhamAlexander Miserlis HoyleSimeng SunPhilip ResnikMohit Iyyer

article2024NAACL260 citations

Proposes a prompt-based framework using large language models to generate interpretable natural language topic labels, textual descriptions, and verifiable document quotes, significantly outperforming traditional methods like LDA and BERTopic on human alignment benchmarks without requiring model retraining.

Listen

Organizations routinely rely on automated topic modeling to analyze and categorize large volumes of unstructured text data. However, traditional approaches such as Latent Dirichlet Allocation generate topics as ambiguous collections of words that require time-consuming manual interpretation and offer little direct control over topic structure. The article introduces and evaluates TopicGPT, a prompt-based framework that uses large language models to generate intuitive topic labels with natural language descriptions and assign them to documents with verifiable text evidence.

The framework operates in two core stages: topic generation and topic assignment. During topic generation, an advanced language model processes a representative sample of documents alongside a small set of example topics to produce high-level candidate topics. The framework then merges near duplicates using sentence similarity and removes infrequent topics before using a secondary model to assign the final topics and supporting quotes to individual documents. The authors benchmarked this framework against standard methods, including traditional and neural topic models, across two diverse datasets: a corpus of 14,290 Wikipedia articles and a collection of 32,661 United States Congressional bill summaries.

The evaluation revealed several clear findings. First, TopicGPT significantly outperformed all baseline topic models in aligning with human-annotated ground-truth categories. For example, it achieved a harmonic mean purity of 0.74 on Wikipedia articles compared to 0.64 for the strongest baseline, Latent Dirichlet Allocation, and 0.57 versus 0.52 on congressional bills. Second, human evaluations showed that TopicGPT generated topics with far fewer semantic errors, producing only 30.3% misaligned topics on Wikipedia data compared to 62.4% for traditional modeling. Third, the framework demonstrated high operational stability across different document subsets, prompt variations, and document orderings. Finally, tests with open-source models showed that while smaller models can effectively perform topic assignment, sophisticated models like GPT-4 remain necessary to successfully generate coherent, well-structured topic lists.

These results demonstrate that using language models for topic discovery can significantly reduce the manual effort required to decipher statistical topic clusters while enhancing auditability through cited textual evidence. Organizations can deploy this framework directly for content analysis without custom model training. The article recommends providing a concise seed list of two to three high-quality example topics to steer topic scope, using multi-label assignment prompts for documents with overlapping themes, and tuning refinement thresholds to avoid filtering out rare but relevant topics. Teams should sample approximately 600 to 1,000 documents for topic generation to minimize processing costs before running full assignment.

Decision-makers should nevertheless consider specific operational constraints. Running commercial language model APIs over large collections incurs variable costs—ranging from approximately 88to88 to 155 per dataset evaluated—and relies on proprietary models with limited architectural transparency. Additionally, long documents may require truncation to fit model context windows, and performance has not yet been established on non-English corpora. Despite these boundaries, confidence in TopicGPT's ability to produce robust, interpretable, and human-aligned topic hierarchies remains high for English-language textual analysis.

arXiv: 2311.01449chtmp223/topicGPT

No sufficiently relevant recommendations were found.

Cover for TopicGPT: A Prompt-based Topic Modeling Framework

Abstract

Topic modeling is a well-established technique for exploring text corpora. Conventional topic models (e.g., LDA) represent topics as bags of words that often require “reading the tea leaves” to interpret; additionally, they offer users minimal control over the formatting and specificity of resulting topics. To tackle these issues, we introduce TopicGPT, a prompt-based framework that uses large language models (LLMs) to uncover latent topics in a text collection. TopicGPT produces topics that align better with human categorizations compared to competing methods: it achieves a harmonic mean purity of 0.74 against human-annotated Wikipedia topics compared to 0.64 for the strongest baseline. Its topics are also interpretable, dispensing with ambiguous bags of words in favor of topics with natural language labels and associated free-form descriptions. Moreover, the framework is highly adaptable, allowing users to specify constraints and modify topics without the need for model retraining. By streamlining access to high-quality and interpretable topics, TopicGPT represents a compelling, human-centered approach to topic modeling.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methodology
  • 3.1 Stage 1: Topic Generation
  • 3.2 Stage 2: Topic Assignment
  • 4 Experiments
  • 4.1 Datasets
  • 4.2 Baselines
  • 4.3 Sampling documents for TopicGPT
  • 4.4 TopicGPT implementation details
  • 4.5 Evaluation Setup
  • 4.5.1 Topical alignment
  • 4.5.2 Stability
  • 5 Results
  • 5.1 TopicGPT is strongly aligned to ground truth labels
  • 5.2 TopicGPT is stable
  • 5.3 TopicGPT topics are semantically close to ground truth
  • 5.4 Implementing TopicGPT with open-source LLMs
  • 6 Future Work
  • 7 Conclusion
  • Limitations
  • Ethical Considerations
  • Acknowledgements
  • References
  • A Hierarchical Implementation
  • B Comparison of TopicGPT topics with human-curated qualitative categories in a different domain
  • C Probabilistic Justification for Data Sampling
  • D Example Topics
  • E Varying number of topics k for LDA baseline
  • F Human annotations on topical alignment between TopicGPT and ground truth of Bills dataset
  • G Topical alignment metrics
  • H Experiment datasets
  • I Additional seeded/anchored topic models
  • J Prompts

Knowls

  1. Knowl 1 — TopicGPT Prompt-Based Topic Modeling Framework

    model/method

    TopicGPT is a prompt-based topic modeling framework designed for automated content analysis that discovers latent thematic structures using large language models (LLMs) rather than probabilistic word distributions. In TopicGPT, a topic is defined as a concise, human-interpretable natural language label paired with a broad one-sentence descriptive definition (e.g., Trade: Mentions the exchange of capital, goods, and services).

    The framework executes in two main stages:

    1. Topic Generation and Refinement: An LLM (such as GPT-4) is prompted sequentially across a sampled subset of documents from the corpus along with a small set of demonstration seed topics (2--3 topics). For each document, the model assigns an existing topic from the running topic list or generates a new label and definition. The accumulated topics are subsequently refined by merging semantically duplicate topic pairs via embedding cosine similarity prompts and removing topics whose generation frequency falls below a threshold.

    2. Topic Assignment and Self-Correction: An LLM (such as GPT-3.5-turbo or Mistral-7B) is prompted with the full refined topic taxonomy, few-shot demonstration examples, and the document to be coded. The model outputs the assigned topic label, an explanation of the assignment, and a verbatim quote from the text that provides auditable evidence. An automated parsing and retry mechanism detects hallucinated labels or formatting anomalies and prompts the LLM to self-correct.

  2. Knowl 2 — Iterative Topic Generation and Refinement Algorithm

    algorithm

    The topic generation and refinement algorithm generates a coherent, non-redundant taxonomy of natural language topics from a document sample and a few seed demonstration topics.

    Input: Document sample Dsample={d1,d2,…,dm}D_{sample} = \{d_1, d_2, \dots, d_m\}, Initial topic set S={(l1,desc1),… }S = \{(l_1, desc_1), \dots\}, Similarity threshold τsim=0.5\tau_{sim} = 0.5, Frequency threshold τfreq\tau_{freq}
    Output: Refined topic set SrefinedS_{refined}
    Initialize frequency counter freq[t]←0freq[t] \leftarrow 0 for all t∈St \in S
    for each document d∈Dsampled \in D_{sample} do
        Prompt LLM with topic set SS, instructions, few-shot examples, and document dd
        Parse response rr
        if rr proposes new topic tnew=(label,description)t_{new} = (label, description) then
            S←S∪{tnew}S \leftarrow S \cup \{t_{new}\}
            freq[tnew]←1freq[t_{new}] \leftarrow 1
        else if rr assigns existing topic texisting∈St_{existing} \in S then
            freq[texisting]←freq[texisting]+1freq[t_{existing}] \leftarrow freq[t_{existing}] + 1
        end if
    end for
    Compute Sentence-Transformer embeddings E(t)E(t) for each topic t∈St \in S
    Identify topic pairs (ti,tj)(t_i, t_j) with cosine similarity cos⁡(E(ti),E(tj))≥τsim\cos(E(t_i), E(t_j)) \ge \tau_{sim}
    Group identified pairs into batches of 5 and prompt LLM to merge near-duplicates into unified topics
    Update SS by substituting merged pairs with their merged definitions and summing their frequencies
    Srefined←∅S_{refined} \leftarrow \emptyset
    for each topic t∈St \in S do
        if freq[t]≥τfreqfreq[t] \ge \tau_{freq} then
            Srefined←Srefined∪{t}S_{refined} \leftarrow S_{refined} \cup \{t\}
        end if
    end for
    return SrefinedS_{refined}
  3. Knowl 3 — Topic Assignment with Grounded Quotations and Self-Correction

    model/method

    In the assignment stage of TopicGPT, documents are mapped to the generated topic set while providing contextual textual proof.

    Given a target document dd, the assignment prompt supplies:

    1. The generated list of topic labels and their descriptions.
    2. Two to three few-shot examples demonstrating formatting.
    3. The text of document dd.

    The LLM is instructed to output the assigned topic label, a document-specific description justifying the choice, and an exact excerpt from the document enclosed in parentheses (e.g., Agriculture: Mentions changes in agricultural export requirements ("...repeal of the agricultural export requirements...")). The excerpt allows users to verify model decisions.

    To ensure formatting validity and eliminate hallucinations, a parser inspects each assignment. If an output contains an unlisted topic label or error strings (such as None or Error), the self-correction module re-prompts the LLM with the document and specific error diagnostic, iterating until a valid assignment is returned (up to a 10-retry limit).

  4. Knowl 4 — Probabilistic Sample Sizing for Topic Generation

    theoretical result

    To minimize LLM API expenditure while ensuring coverage of rare topics, determining the generation sample size nsn_s is formulated as an occupancy problem under a multinomial distribution.

    Let NN be the total number of documents in the corpus and let ndn_d be a user-specified lower bound on the number of documents belonging to the least-prevalent topic of interest. This induces an upper bound KuK_u on the number of discoverable topics: Ku=⌊Nnd⌋K_u = \left\lfloor \frac{N}{n_d} \right\rfloor

    Assuming a uniform topic distribution, sampling nsn_s documents uniformly at random from the corpus is modeled as a draw c∼Multi(ns,Ku)\mathbf{c} \sim \text{Multi}(n_s, K_u), where c=(c1,…,cKu)\mathbf{c} = (c_1, \dots, c_{K_u}) represents the number of sampled documents in each topic. The expected number of undetected topics (zeros in c\mathbf{c}) is: E[zeros]=(Ku−1)nsKuns−1\mathbb{E}[\text{zeros}] = \frac{(K_u - 1)^{n_s}}{K_u^{n_s - 1}}

    The sample size nsn_s is selected via simulation by minimizing: ns=arg⁡min⁡n∣P(min⁡kck=0)−ε∣n_s = \arg\min_{n} \left| P\left(\min_k c_k = 0\right) - \varepsilon \right| where ε\varepsilon is a user-defined maximum failure probability of failing to sample the least-prevalent topic. For example, for N=14,290N = 14{,}290 and nd=140n_d = 140 (1%1\% corpus prevalence), setting ns=1,100n_s = 1{,}100 guarantees a failure probability p∗≈0.005p^* \approx 0.005.

  5. Knowl 5 — Topical Alignment Performance of TopicGPT versus Baseline Topic Models

    data/table

    Topical alignment with human ground-truth labels was evaluated on Wikipedia articles (Wiki, 14,290 documents, 15 high-level classes) and Congressional Bill summaries (Bills, 32,661 documents, 21 high-level classes). Metrics include the harmonic mean of purity and inverse purity (P1P_1), Adjusted Rand Index (ARI), and Normalized Mutual Information (NMI). Baselines were evaluated with the number of topics kk set equal to the number of topics discovered by TopicGPT.

    Dataset Setting TopicGPT LDA BERTopic SeededLDA
    P1P_1 ARI NMI P1P_1 ARI NMI P1P_1 ARI NMI P1P_1 ARI NMI
    Wiki Default (k=31k=31) 0.73 0.58 0.71 0.59 0.44 0.65 0.54 0.24 0.50 0.61 0.47 0.65
    Wiki Refined (k=22k=22) 0.74 0.60 0.70 0.64 0.52 0.67 0.58 0.28 0.50 0.62 0.51 0.65
    Bills Default (k=79k=79) 0.57 0.42 0.52 0.39 0.21 0.47 0.42 0.10 0.40 0.50 0.28 0.43
    Bills Refined (k=24k=24) 0.57 0.40 0.49 0.52 0.32 0.46 0.39 0.12 0.34 0.52 0.31 0.45

    TopicGPT substantially outperforms traditional LDA, SeededLDA, and neural BERTopic across all metrics and datasets. On the refined Wiki benchmark, TopicGPT achieves P1=0.74P_1 = 0.74 versus 0.640.64 for LDA, 0.620.62 for SeededLDA, and 0.580.58 for BERTopic. Additional seeded/anchored baselines also underperform TopicGPT: CorEx achieves P1=0.52P_1 = 0.52 on Wiki (k=31k=31) and P1=0.38P_1 = 0.38 on Bills (k=79k=79), while BERTopic-guided achieves P1=0.26P_1 = 0.26 on Wiki and P1=0.30P_1 = 0.30 on Bills.

  6. Knowl 6 — Proportion of Misaligned Topics in Qualitative Human Evaluation

    data/table

    To evaluate semantic topic quality independent of document assignments, three human annotators mapped model-generated topics to gold-standard labels. Unmatched topics were classified into three misalignment types: Out-of-scope (too specific or too broad), Missing (ground-truth categories absent in generated outputs), and Repeated (redundant duplicate topics).

    Dataset Setting Out-of-scope (%) Missing (%) Repeated (%) Total Misaligned (%)
    Wiki LDA (k=31k=31) 46.3 4.3 11.9 62.4
    Wiki TopicGPT Unrefined (k=31k=31) 38.7 0.0 1.1 39.8
    Wiki TopicGPT Refined (k=22k=22) 30.3 0.0 0.0 30.3
    Bills LDA (k=79k=79) 56.1 2.1 22.0 80.2
    Bills TopicGPT Unrefined (k=79k=79) 65.0 1.3 3.8 70.1
    Bills TopicGPT Refined (k=24k=24) 27.8 4.2 0.0 31.9

    TopicGPT's refined topics reduce total semantic misalignment to 30.3% on Wiki and 31.9% on Bills, compared to 62.4% and 80.2% for LDA. Refinement completely eliminates repeated topics (0.0% on both datasets) and reduces out-of-scope topics.

  7. Knowl 7 — Stability and Sensitivity Analysis of TopicGPT

    empirical result

    TopicGPT topic assignments show high stability under variations in prompts, generation data subsets, and document order on the Congressional Bills dataset:

    1. Replicate Consistency: Running the identical TopicGPT pipeline twice on Bills (k=79k=79) achieved internal alignment scores of P1=0.95P_1 = 0.95, ARI=0.92\text{ARI} = 0.92, and NMI=0.92\text{NMI} = 0.92 between runs, compared to an average internal agreement across 10 LDA runs (k=79k=79) of P1=0.64P_1 = 0.64, ARI=0.55\text{ARI} = 0.55, and NMI=0.71\text{NMI} = 0.71.
    2. Sampling Variations:
      • Shuffling the document presentation order during generation (k=118k=118) yielded ground-truth alignment of P1=0.55P_1 = 0.55, ARI=0.40\text{ARI} = 0.40, NMI=0.52\text{NMI} = 0.52.
      • Using an independent random generation sample (k=73k=73) maintained alignment of P1=0.57P_1 = 0.57, ARI=0.40\text{ARI} = 0.40, NMI=0.51\text{NMI} = 0.51.
    3. Out-of-Domain Prompts: Applying prompt instructions and seed topics crafted for Wikipedia directly to Bills (k=147k=147) achieved P1=0.55P_1 = 0.55, ARI=0.39\text{ARI} = 0.39, NMI=0.51\text{NMI} = 0.51.
    4. Example Topic Count Sensitivity: Expanding the seed prompt from 2 examples to 5 examples (k=123k=123) resulted in lower alignment (P1=0.50P_1 = 0.50, ARI=0.33\text{ARI} = 0.33, NMI=0.49\text{NMI} = 0.49), indicating that overloading the prompt with examples harms model focus and that 2--3 seed examples are optimal.
  8. Knowl 8 — Asymmetric Capability of Open-Source LLMs for Topic Assignment versus Generation

    empirical result

    Evaluating open-source instruction-tuned LLMs (specifically Mistral-7B-Instruct) within TopicGPT reveals a split in capability between assignment and generation:

    • Topic Assignment: Mistral-7B-Instruct performs competently as a topic assigner. On the Congressional Bills dataset (k=79k=79), Mistral-7B achieved P1=0.51P_1 = 0.51, ARI=0.37\text{ARI} = 0.37, and NMI=0.46\text{NMI} = 0.46. Although lower than GPT-3.5-turbo (P1=0.57P_1 = 0.57, ARI=0.42\text{ARI} = 0.42), Mistral-7B still outperforms baseline LDA (P1=0.39P_1 = 0.39, ARI=0.21\text{ARI} = 0.21) and BERTopic (P1=0.42P_1 = 0.42, ARI=0.10\text{ARI} = 0.10).
    • Topic Generation: Both Mistral-7B-Instruct and GPT-3.5-turbo fail at iterative topic generation because they struggle with complex instructions (e.g., maintaining generalizability, level formatting, and avoiding document-specific labels). Mistral-7B and GPT-3.5-turbo generated 1,418 and 151 overly specific topics, respectively, making the output unsuitable for refinement and subsequent assignment prompts. Thus, high-capacity models such as GPT-4 are necessary for the topic generation phase.
  9. Knowl 9 — Hierarchical Topic Taxonomy Generation with TopicGPT

    model/method

    TopicGPT extends to hierarchical topic modeling by generating multi-tier topic structures recursively:

    1. First-Level Topics: The primary generation and refinement pipeline produces broad, high-level root topics (e.g., Architecture & Design, Animal Breeds & Husbandry).
    2. Branch Prompting: For each top-level topic tt, the subset of corpus documents dtd_t assigned to tt is extracted. An LLM is prompted with the top-level topic label, demonstration subtopic examples, and dtd_t.
    3. Document-Grounded Subtopic Induction: The model generates second-level subtopics along with the specific document IDs supporting each subtopic. Grounding each proposed subtopic in specific document references prevents hallucinated categories. If documents exceed the context window, they are split across prompt batches, carrying forward previously generated subtopics into subsequent calls.

    On the Wikipedia corpus, this generates nuanced subtopic structures (e.g., expanding Architecture & Design into Religious Architecture, Parks and Public Spaces, and Historic Structures).

  10. Knowl 10 — Limitations of the TopicGPT Framework

    limitation

    The TopicGPT framework has four main limitations:

    1. Context Window Truncation: Long documents must be truncated to fit within LLM context window limits, risking the loss of context and mischaracterization of full document contents.
    2. Dependence on Closed-Source Models: Reliable topic generation currently depends on proprietary models (GPT-4) with opaque training corpora and architectures.
    3. Inference Cost: Relying on commercial LLM APIs incurs monetary expense. In benchmark evaluations, coding 15,242 test documents in Bills cost approximately $88\$88, while 8,024 test documents in Wiki cost $155\$155 due to longer document lengths (averaging 3,412 tokens/document for Wiki vs. 261 tokens/document for Bills).
    4. English-Centric Evaluation: The framework has only been evaluated on English-language corpora; degraded instruction-following in non-English languages presents challenges for multilingual transfer.

Coverage note — Omitted the Reddit 'Reasons to Live' case study qualitative coding details (Tables 5-7) and the varying-k LDA baseline ablation (Table 8), as their core insights (qualitative alignment and TopicGPT outperforming LDA across all k) are captured in the included knowls.

References

  1. 1.E Scott Adler and John Wilkerson. 2018. Congressional bills project: 1995-2018.
  2. 2.Enrique Amigó, Julio Gonzalo, Javier Artiles, and Felisa Verdejo. 2009. A comparison of extrinsic clustering evaluation metrics based on formal constraints. Information retrieval, 12:461–486.
  3. 3.David Andrzejewski and Xiaojin Zhu. 2009. Latent Dirichlet Allocation with Topic-in-Set Knowledge. In Proceedings of the NAACL HLT 2009 Workshop on Semi-supervised Learning for Natural Language Processing, pages 43–48, Boulder, Colorado. Association for Computational Linguistics.
  4. 4.Christian Baden, Christian Pipal, Martijn Schoonvelde, and Mariken A.C.G. van der Velden. 2021. Three gaps in computational text analysis methods for social sciences: A research agenda. Communication Methods and Measures, 16:1 – 18.
  5. 5.Federico Bianchi, Silvia Terragni, and Dirk Hovy. 2021. Pre-training is a hot topic: Contextualized document embeddings improve topic coherence.
  6. 6.David M Blei, Andrew Y Ng, and Michael I Jordan. 2003. Latent dirichlet allocation. Journal of machine Learning research, 3(Jan):993–1022.
  7. 7.Jonathan Chang, Sean Gerrish, Chong Wang, Jordan Boyd-Graber, and David Blei. 2009. Reading tea leaves: How humans interpret topic models. Advances in neural information processing systems, 22.
  8. 8.Robert Chew, John Bollenbacher, Michael Wenger, Jessica Speer, and Annice Kim. 2023. Llm-assisted content analysis: Using large language models to support deductive coding.
  9. 9.Jason Chuang, Sonal Gupta, Christopher Manning, and Jeffrey Heer. 2013. Topic model diagnostics: Assessing domain relevance via topical alignment. In Proceedings of the 30th International Conference on Machine Learning, volume 28 of Proceedings of Machine Learning Research, pages 612–620, Atlanta, Georgia, USA. PMLR.
  10. 10.Caitlin Doogan and Wray Buntine. 2021. Topic model or topic twaddle? re-evaluating semantic interpretability measures. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3824–3848, Online. Association for Computational Linguistics.
  11. 11.Ryan J. Gallagher, Kyle Reing, David Kale, and Greg Ver Steeg. 2017. Anchored Correlation Explanation: Topic Modeling with Minimal Domain Knowledge. Transactions of the Association for Computational Linguistics, 5:529–542.
  12. 12.Thomas Griffiths, Michael Jordan, Joshua Tenenbaum, and David Blei. 2003. Hierarchical topic models and the nested chinese restaurant process. Advances in neural information processing systems, 16.
  13. 13.Tom Griffiths. 2002. Gibbs sampling in the generative model of latent dirichlet allocation. Standford University, 518(11):1–3.
  14. 14.Maarten Grootendorst. 2022. Bertopic: Neural topic modeling with a class-based tf-idf procedure.
  15. 15.Alexander Hoyle, Pranav Goel, Andrew Hian-Cheong, Denis Peskov, Jordan Boyd-Graber, and Philip Resnik. 2021. Is Automated Topic Model Evaluation Broken? The Incoherence of Coherence. In Advances in Neural Information Processing Systems, volume 34, pages 2018–2033. Curran Associates, Inc.
  16. 16.Alexander Hoyle, Rupak Sarkar, Pranav Goel, and Philip Resnik. 2023. Natural language decompositions of implicit content enable better text representations.
  17. 17.Alexander Miserlis Hoyle, Pranav Goel, Rupak Sarkar, and Philip Resnik. 2022. Are neural topic models broken? In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 5321–5344, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  18. 18.Hsiu-Fang Hsieh and Sarah E Shannon. 2005. Three approaches to qualitative content analysis. Qualitative health research, 15(9):1277–1288.
  19. 19.Yuening Hu, Jordan Boyd-Graber, Brianna Satinoff, and Alison Smith. 2014. Interactive topic modeling. Machine learning, 95:423–469.
  20. 20.Haoyang Huang, Tianyi Tang, Dongdong Zhang, Wayne Xin Zhao, Ting Song, Yan Xia, and Furu Wei. 2023. Not all languages are created equal in llms: Improving multilingual capability by crosslingual-thought prompting. In Findings of Empirical Methods in Natural Language Processing.
  21. 21.Lawrence Hubert and Phipps Arabie. 1985. Comparing partitions. Journal of classification, 2:193–218.
  22. 22.Jagadeesh Jagarlamudi, Hal Daumé III, and Raghavendra Udupa. 2012. Incorporating Lexical Priors into Topic Models. In Proceedings of the 13th Conference of the European Chapter of the Association for Computational Linguistics, pages 204–213, Avignon, France. Association for Computational Linguistics.
  23. 23.Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023. Mistral 7b.
  24. 24.Damir Korenčić, Strahil Ristov, Jelena Repar, and Jan Šnajder. 2021. A Topic Coverage Approach to Evaluation of Topic Models. IEEE Access, 9:123280–123312. ArXiv:2012.06274 [cs].
  25. 25.Helvi Kyngäs. 2020. Inductive content analysis. The application of content analysis in nursing science research, pages 13–21.
  26. 26.Jey Han Lau, Karl Grieser, David Newman, and Timothy Baldwin. 2011. Automatic labelling of topic models. In Proceedings of the 49th annual meeting of the association for computational linguistics: human language technologies, pages 1536–1545.
  27. 27.Haonan Li, Fajri Koto, Minghao Wu, Alham Fikri Aji, and Timothy Baldwin. 2023. Bactrian-x : A multilingual replicable instruction-following model with low-rank adaptation.
  28. 28.Marsha M Linehan, Judith L Goodstein, Stevan L Nielsen, and John A Chiles. 1983. Reasons for staying alive when you are thinking of killing yourself: the reasons for living inventory. Journal of consulting and clinical psychology, 51(2):276.
  29. 29.Sengjie Liu and Christopher G Healey. 2023. Abstractive summarization of large document collections using gpt. arXiv preprint arXiv:2310.05690.
  30. 30.Andrew Kachites McCallum. 2002. Mallet: A machine learning for language toolkit. Http://www.cs.umass.edu/ mccallum/mallet.
  31. 31.Qiaozhu Mei, Xuehua Shen, and ChengXiang Zhai. 2007. Automatic labeling of multinomial topic models. In Knowledge Discovery and Data Mining.
  32. 32.Marina Meilă. 2007. Comparing clusterings—an information based distance. Journal of multivariate analysis, 98(5):873–895.
  33. 33.Yu Meng, Jiaxin Huang, Guangyuan Wang, Zihan Wang, Chao Zhang, Yu Zhang, and Jiawei Han. 2020. Discriminative Topic Mining via Category-Name Guided Text Embedding. In Proceedings of The Web Conference 2020, pages 2121–2132. ArXiv:1908.07162 [cs].
  34. 34.Stephen Merity, Nitish Shirish Keskar, and Richard Socher. 2018. Regularizing and optimizing LSTM language models. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings. OpenReview.net.
  35. 35.David Mimno, Wei Li, and Andrew McCallum. 2007. Mixtures of hierarchical topics with pachinko allocation. In Proceedings of the 24th international conference on Machine learning, pages 633–640.
  36. 36.David Newman, Jey Han Lau, Karl Grieser, and Timothy Baldwin. 2010. Automatic evaluation of topic coherence. In Human language technologies: The 2010 annual conference of the North American chapter of the association for computational linguistics, pages 100–108.
  37. 37.Sergey I Nikolenko, Sergei Koltcov, and Olessia Koltsova. 2017. Topic modelling for qualitative studies. Journal of Information Science, 43(1):88–102.
  38. 38.OpenAI. 2023. Gpt-4 technical report.
  39. 39.John Paisley, Chong Wang, David M Blei, and Michael I Jordan. 2014. Nested hierarchical dirichlet processes. IEEE transactions on pattern analysis and machine intelligence, 37(2):256–270.
  40. 40.Forough Poursabzi-Sangdeh, Jordan Boyd-Graber, Leah Findlater, and Kevin Seppi. 2016. Alto: Active learning with topic overviews for speeding label induction and document labeling. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1158–1169.
  41. 41.Daniel Ramage, Christopher D. Manning, and Susan Dumais. 2011. Partially labeled topic models for interpretable text mining. In Proceedings of the 17th ACM SIGKDD international conference on Knowledge discovery and data mining, pages 457–465, San Diego California USA. ACM.
  42. 42.William M Rand. 1971. Objective criteria for the evaluation of clustering methods. Journal of the American Statistical association, 66(336):846–850.
  43. 43.Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence embeddings using Siamese BERTnetworks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3982–3992, Hong Kong, China. Association for Computational Linguistics.
  44. 44.Philip Resnik, Pranav Goel, Alexander Hoyle, Rupak Sarkar, Josh Hagedorn, Maeve Gearing, and Carol Bruce. 2022. A step-by-step protocol for curation of topic models by subject matter experts. In New Directions in Analyzing Text as Data.
  45. 45.Philip Resnik, Bolei Ma, Alexander Hoyle, Pranav Goel, Rupak Sarkar, Maeve Gearing, Anna-Carolina Haensch, and Frauke Kreuter. 2024a. A topic-oriented protocol for content analysis of text. Manuscript in preparation.
  46. 46.Philip Resnik, Katherine Musacchio Schafer, Rebecca Resnik, Josh Hagedorn, and Jonathan Singer. 2024b. Reasons to live instead of dying by suicide: New insights from a computer assisted content analysis. Manuscript in preparation.
  47. 47.Emil Rijcken, Floortje Scheepers, Kalliopi Zervanou, Marco Spruit, Pablo Mosteiro, and Uzay Kaymak. 2023. Towards Interpreting Topic Models with ChatGPT: The 20th World Congress of the International Fuzzy Systems Association.
  48. 48.Claude Elwood Shannon. 1948. A mathematical theory of communication. The Bell system technical journal, 27(3):379–423.
  49. 49.Noah Shinn, Federico Cassano, Beck Labash, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2023. Reflexion: Language agents with verbal reinforcement learning.
  50. 50.Suzanna Sia, Ayush Dalmia, and Sabrina J. Mielke. 2020. Tired of topic models? clusters of pretrained word embeddings make for fast and good topics too! In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1728–1736, Online. Association for Computational Linguistics.
  51. 51.Akash Srivastava and Charles Sutton. 2017. Autoencoding variational inference for topic models. arXiv preprint arXiv:1703.01488.
  52. 52.Dominik Stammbach, Vilém Zouhar, Alexander Hoyle, Mrinmaya Sachan, and Elliott Ash. 2023. Revisiting automated topic model evaluation with large language models.
  53. 53.Alexander Strehl and Joydeep Ghosh. 2002. Cluster ensembles—a knowledge reuse framework for combining multiple partitions. Journal of machine learning research, 3(Dec):583–617.
  54. 54.Simeng Sun, Yang Liu, Shuohang Wang, Chenguang Zhu, and Mohit Iyyer. 2023. Pearl: Prompting large language models to plan and execute actions over long documents.
  55. 55.Robert H Tai, Lillian R Bentley, Xin Xia, Jason M Sitt, Sarah C Fankhauser, Ana M Chicas-Mosier, and Barnas G Monteith. 2023. Use of large language models to aid analysis of textual data. bioRxiv, pages 2023–07.
  56. 56.Yee Whye Teh, Michael I Jordan, Matthew J Beal, and David M Blei. 2006. Hierarchical Dirichlet Processes. Journal of the American Statistical Association, 101(476):1566–1581.
  57. 57.Laure Thompson and David Mimno. 2020. Topic Modeling with Contextualized Word Representation Clusters. ArXiv:2010.12626 [cs].
  58. 58.Danya F Vears and Lynn Gillam. 2022. Inductive content analysis: A guide for beginning qualitative researchers. Focus on Health Professional Education: A Multi-disciplinary Journal, 23(1):111–127.
  59. 59.Nguyen Xuan Vinh, Julien Epps, and James Bailey. 2009. Information theoretic measures for clusterings comparison: is a correction for chance necessary? In Proceedings of the 26th annual international conference on machine learning, pages 1073–1080.
  60. 60.Vijay Viswanathan, Kiril Gashteovski, Carolin Lawrence, Tongshuang Wu, and Graham Neubig. 2023. Large language models enable few-shot clustering.
  61. 61.Hanna Wallach, David Mimno, and Andrew McCallum. 2009. Rethinking lda: Why priors matter. Advances in neural information processing systems, 22.
  62. 62.Xiaojun Wan and Tianming Wang. 2016. Automatic labeling of topic models using text summaries. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2297–2305, Berlin, Germany. Association for Computational Linguistics.
  63. 63.Zihan Wang, Jingbo Shang, and Ruiqi Zhong. 2023. Goal-driven explainable clustering via language descriptions.
  64. 64.Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, et al. 2021. Ethical and social risks of harm from language models. arXiv preprint arXiv:2112.04359.
  65. 65.Yuwei Zhang, Zihan Wang, and Jingbo Shang. 2023. ClusterLLM: Large Language Models as a Guide for Text Clustering. ArXiv:2305.14871 [cs].
  66. 66.Ying Zhao. 2005. Criterion Functions for Document Clustering. Ph.D. thesis, University of Minnesota, USA. AAI3180039.

Citation

MLA
Pham, C. M., et al. “TopicGPT: A Prompt-based Topic Modeling Framework”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2024, pp. 2956–84, https://doi.org/10.18653/v1/2024.naacl-long.164.
APA
Pham, C. M., Hoyle, A. M., Sun, S., Resnik, P., & Iyyer, M. (2024). TopicGPT: A Prompt-based Topic Modeling Framework. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2956–2984. https://doi.org/10.18653/v1/2024.naacl-long.164
Chicago
Pham, C. M., A. M. Hoyle, S. Sun, P. Resnik, and M. Iyyer. 2024. “TopicGPT: A Prompt-based Topic Modeling Framework”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2956–84. https://doi.org/10.18653/v1/2024.naacl-long.164.
Harvard
Pham, C.M. et al. (2024) “TopicGPT: A Prompt-based Topic Modeling Framework”, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 2956–2984. Available at: https://doi.org/10.18653/v1/2024.naacl-long.164.
Vancouver
1. Pham CM, Hoyle AM, Sun S, Resnik P, Iyyer M (2024) TopicGPT: A Prompt-based Topic Modeling Framework. In: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 2956–2984

BibTeX

@inproceedings{pham-etal-2024-topicgpt,
    title = "{T}opic{GPT}: A Prompt-based Topic Modeling Framework",
    author = "Pham, Chau Minh  and
      Hoyle, Alexander  and
      Sun, Simeng  and
      Resnik, Philip  and
      Iyyer, Mohit",
    editor = "Duh, Kevin  and
      Gomez, Helena  and
      Bethard, Steven",
    booktitle = "Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = jun,
    year = "2024",
    address = "Mexico City, Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.naacl-long.164/",
    doi = "10.18653/v1/2024.naacl-long.164",
    pages = "2956--2984"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/