Optimizing Semantic Coherence in Topic Models

David MimnoHanna M. WallachEdmund TalleyMiriam LeendersAndrew McCallum

article2011EMNLP1,995 citations

Proposes an intrinsic semantic coherence metric and a generalized Pólya urn topic model to automatically detect and eliminate low-quality, nonsensical latent topics without requiring human evaluation or external corpora.

Listen

Statistical topic models help organizations automatically summarize and explore large document collections. However, standard methods like latent Dirichlet allocation frequently generate low-quality topics that mix unrelated concepts or arbitrarily chain disparate ideas together. When domain experts encounter these flawed topics—which can constitute up to 10% or more of total outputs—they lose trust in the automated analysis. In response, the article set out to identify the structural flaws in generated topics, develop an automated metric that accurately scores semantic quality without human annotators or external corpora, and demonstrate a new statistical model that directly improves topic quality.

The authors conducted an expert-driven annotation study using over 300,000 research abstracts from the National Institutes of Health. Medical experts categorized and labeled specific failure modes in topics, including chained, intruded, random, and unbalanced concepts. Building on these observations, the authors developed an automated topic coherence metric based purely on internal word co-occurrence statistics. They subsequently designed a generalized statistical topic model that uses an urn-based sampling schema to promote related words during training, comparing its performance and speed against standard topic modeling approaches across multiple cross-validation runs.

The analysis yielded several key findings. First, the automated topic coherence metric closely matched human domain evaluations, achieving high ranking accuracy (area under the curve of 0.87) and outperforming existing measures like topic size and mutual information without requiring outside reference text. Second, standard intruder-word evaluations failed to reliably detect certain flawed topics, such as chained concepts, which the new coherence metric successfully caught. Third, the new generalized model reached high average coherence levels within just 50 sampling iterations—far faster than standard methods—and significantly improved the quality of the ten lowest-scoring topics. Finally, while the model made the worst topics noticeably less flawed, it only modestly reduced the overall proportion of bad topics from 16.5% to 13.5%, a difference that was not statistically significant.

These findings indicate that internal document statistics contain enough unexploited information to automatically evaluate and improve topic interpretability. Organizations can filter out or flag low-quality outputs before presenting them to end-users, protecting user trust and reducing the manual oversight traditionally needed for quality control. Furthermore, the ability of the new model to converge in substantially fewer iterations offers opportunities to lower computing runtimes during training, though the sampling process per iteration requires two to three times more computational overhead due to expanded bookkeeping.

For practical implementation, organizations deploying topic models should adopt the automated coherence metric as a quality-control filter to automatically suppress or review low-scoring topics. Teams seeking higher semantic quality and faster model convergence should consider generalized urn-based topic models, balancing their faster convergence against the added per-iteration processing requirements. Future work should focus on refining inference algorithms to further eliminate the remaining baseline of bad topics and applying these word co-occurrence principles to online, large-scale streaming systems.

The study's primary limitations stem from its evaluation on a single specialized corpus of biomedical grant abstracts and the approximate nature of the sampling distribution used to maintain tractability. While decision-makers can have high confidence in the coherence metric's ability to rank and identify incoherent topics, they should exercise caution before assuming the modeling framework will completely eliminate flawed outputs without an automated filtering threshold.

Mimno et al (2011).pdf
  • Paper: Reading Tea Leaves: How Humans Interpret Topic Models, Jonathan D. Chang et al. (2009). Its human-intrusion tests establish the interpretability evaluation that this paper refines by showing which incoherent topics those tests miss.
  • Paper: Latent Dirichlet Allocation, David M. Blei et al. (2003). LDA is the standard topic-model foundation whose flawed topics and sampling behavior this paper evaluates and seeks to improve.
  • Paper: Exploring the Space of Topic Coherence Measures, Michael Röder et al. (2015). This later study systematically searches coherence-measure designs and benchmarks them against human judgments, extending the source’s internal-statistics approach to topic evaluation.
Cover for Optimizing Semantic Coherence in Topic Models

Abstract

Latent variable models have the potential to add value to large document collections by discovering interpretable, low-dimensional subspaces. In order for people to use such models, however, they must trust them. Unfortunately, typical dimensionality reduction methods for text, such as latent Dirichlet allocation, often produce low-dimensional subspaces (topics) that are obviously flawed to human domain experts. The contributions of this paper are threefold: (1) An analysis of the ways in which topics can be flawed; (2) an automated evaluation metric for identifying such topics that does not rely on human annotators or reference collections outside the training data; (3) a novel statistical topic model based on this metric that significantly improves topic quality in a large-scale document collection from the National Institutes of Health (NIH).

Table of Contents

  • 1 Introduction
  • 2 Latent Dirichlet Allocation
  • 3 Expert Opinions of Topic Quality
  • 3.1 Expert-Driven Annotation Protocol
  • 3.2 Annotation Results
  • 4 Automated Metrics for Predicting Expert Annotations
  • 4.1 Topic Size
  • 4.2 Topic Coherence
  • 4.3 Comparison to Word Intrusion
  • 5 Generalized Pólya Urn Models
  • 5.1 Setting the Schema
  • 5.2 Experimental Results
  • 6 Discussion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Intrinsic Topic Coherence Metric Based on Document Co-occurrences

    equation

    Let V(t)=(v1(t),…,vM(t))V^{(t)} = (v_1^{(t)}, \dots, v_M^{(t)}) be the ordered list of the MM most probable words in topic tt, ranked in descending order of their topic-specific collapsed probabilities P(w∣t)=Nw∣t+βNt+∣V∣βP(w \mid t) = \frac{N_{w|t} + \beta}{N_t + |V|\beta} (where Nw∣tN_{w|t} is the number of tokens of type ww assigned to topic tt, Nt=∑wNw∣tN_t = \sum_w N_{w|t}, ∣V∣|V| is the vocabulary size, and eta is the symmetric Dirichlet prior hyperparameter). Topic coherence C(t;V(t))C(t; V^{(t)}) is defined as:

    C(t;V(t))=∑m=2M∑l=1m−1log⁡D(vm(t),vl(t))+1D(vl(t))C(t; V^{(t)}) = \sum_{m=2}^M \sum_{l=1}^{m-1} \log \frac{D(v_m^{(t)}, v_l^{(t)}) + 1}{D(v_l^{(t)})}

    where D(v)D(v) is the document frequency of word type vv in the training corpus (the count of documents containing at least one token of type vv), and D(v,v′)D(v, v') is the co-document frequency of word types vv and v′v' (the count of documents containing at least one token of type vv and at least one token of type v′v'). The +1+1 term is a smoothing constant to avoid the logarithm of zero.

    This metric measures the empirical conditional probability of seeing lower-ranked topic words given higher-ranked topic words entirely within the training corpus, without requiring an external reference corpus.

  2. Knowl 2 — Generalized Pólya Urn Topic Model Formulation

    model/method

    In the Generalized Pólya Urn (GPU) topic model, the standard Dirichlet compound multinomial (simple Pólya urn) distribution over words in Latent Dirichlet Allocation (LDA) is replaced by a generalized Pólya urn framework. When a word token of type ww is drawn in topic tt, it returns AvwA_{vw} additional weight for every word type v∈{1,…,∣V∣}v \in \{1, \dots, |V|\}, where A∈R∣V∣×∣V∣\mathbf{A} \in \mathbb{R}^{|V| \times |V|} is a predefined addition schema matrix and ∣V∣|V| is the vocabulary size.

    Given the observed word tokens W\mathbf{W}, latent topic assignments Z\mathbf{Z}, Dirichlet symmetric hyperparameter β\beta, and schema matrix A\mathbf{A}, the conditional posterior probability of word type ww given topic tt is:

    P(w∣t,W,Z,β,A)=∑v=1∣V∣Nv∣tAvw+βNt+∣V∣βP(w \mid t, \mathbf{W}, \mathbf{Z}, \beta, \mathbf{A}) = \frac{\sum_{v=1}^{|V|} N_{v|t} A_{vw} + \beta}{N_t + |V|\beta}

    where Nv∣tN_{v|t} is the count of tokens of type vv currently assigned to topic tt, and Nt=∑v=1∣V∣Nv∣tN_t = \sum_{v=1}^{|V|} N_{v|t}. Standard LDA is recovered when A=I\mathbf{A} = \mathbf{I} (the identity matrix).

  3. Knowl 3 — Generalized Pólya Urn Schema Matrix Construction with IDF Weighting and Stopword Filtering

    model/method

    For a corpus of DD documents and a vocabulary of size ∣V∣|V|, the addition matrix A∈R∣V∣×∣V∣\mathbf{A} \in \mathbb{R}^{|V| \times |V|} of the Generalized Pólya Urn topic model is computed from corpus-internal document frequencies:

    Avv∝λvD(v)Avw∝λvD(w,v)(v≠w)\begin{aligned} A_{vv} &\propto \lambda_v D(v) \\ A_{vw} &\propto \lambda_v D(w, v) \quad (v \neq w) \end{aligned}

    where D(v)D(v) is the document frequency of word vv, D(w,v)D(w, v) is the co-document frequency of words ww and vv, and λv=log⁡(D/D(v))\lambda_v = \log(D / D(v)) is the inverse document frequency (IDF) weight of word type vv, boosting the association weights of rare words. Each column of A\mathbf{A} is normalized such that ∑v=1∣V∣Avw=1\sum_{v=1}^{|V|} A_{vw} = 1.

    To prevent high-frequency words from diffusing across all topics and destroying topic specificity:

    1. Off-diagonal entries AvwA_{vw} (v≠wv \neq w) are set to zero for all high-frequency word types occurring in more than 5% of documents (types with IDF<3.0\text{IDF} < 3.0).
    2. The resulting matrix A\mathbf{A} is sparse, containing non-zero off-diagonal associations only between semantically informative, lower-frequency words.
  4. Knowl 4 — Approximate Gibbs Sampling for Generalized Pólya Urn Topic Models

    algorithm

    Because generalized Pólya urn distributions are non-exchangeable (the joint token probability depends on token sequence order), exact Gibbs sampling requires computing predictive distributions that account for all subsequent token assignments. This process is approximated by treating each token wn(d)w_n^{(d)} as if it were the last drawn token during inference.

    Input: Corpus documents D\mathcal{D}, topic assignments Z\mathbf{Z}, counts Nz∣dN_{z|d} and Nw∣zN_{w|z}, Dirichlet hyperparameters α,β\boldsymbol{\alpha}, \beta, schema matrix A\mathbf{A}
    Output: Updated topic assignments Z\mathbf{Z} and count matrices
    for each document d∈Dd \in \mathcal{D} do
        for each token position nn with word type wn∈w(d)w_n \in \mathbf{w}^{(d)} and current topic assignment zz do
            Nz∣d←Nz∣d−1N_{z|d} \leftarrow N_{z|d} - 1
            for each word type vv where Avwn≠0A_{v w_n} \neq 0 do
                Nv∣z←Nv∣z−AvwnN_{v|z} \leftarrow N_{v|z} - A_{v w_n}
            end for
            Sample new topic assignment znewz_{\text{new}} with probability proportional to:
                (Nznew∣d+αznew)⋅Nwn∣znew+β∑z′(Nwn∣z′+β)(N_{z_{\text{new}}|d} + \alpha_{z_{\text{new}}}) \cdot \frac{N_{w_n|z_{\text{new}}} + \beta}{\sum_{z'} (N_{w_n|z'} + \beta)}
            Nznew∣d←Nznew∣d+1N_{z_{\text{new}}|d} \leftarrow N_{z_{\text{new}}|d} + 1
            for each word type vv where Avwn≠0A_{v w_n} \neq 0 do
                Nv∣znew←Nv∣znew+AvwnN_{v|z_{\text{new}}} \leftarrow N_{v|z_{\text{new}}} + A_{v w_n}
            end for
            zn(d)←znewz_n^{(d)} \leftarrow z_{\text{new}}
        end for
    end for

    When A\mathbf{A} is sparse, updating counts requires iterating only over the non-zero elements of column wnw_n, scaling the per-iteration computational cost to roughly 2–3 times that of standard LDA Gibbs sampling.

  5. Knowl 5 — Qualitative Taxonomy of Semantic Flaws in Inferred Topics

    definition

    Expert evaluation of latent topic models identifies four distinct failure modes that characterize low-quality or nonsensical topics:

    1. Chained Topics: Topics where every word is connected to every other word through a pairwise associative chain, but the overall topic combines disparate concepts (e.g., top words "acids", "fatty", and "nucleic" linking fatty acids and nucleic acids solely through the broad bridge word "acids").
    2. Intruded Topics: Topics containing two or more unrelated sub-vocabularies merged arbitrarily, or an otherwise coherent topic corrupted by several unrelated "intruder" words.
    3. Random Topics: Topics with no discernible semantic connection among more than a few words.
    4. Unbalanced Topics: Topics where words are logically related, but combine overly general concepts with highly specific subfield concepts (e.g., combining "signal transduction" with "notch signaling").

    In an expert study of 148 LDA topics from NIH grant abstracts, 58 topics were labeled as non-good (21 intermediate, 37 bad), consisting of 23 chained, 21 intruded, 15 unbalanced, and 3 random assignments (multiple flaw labels were allowed per topic).

  6. Knowl 6 — Topic Coherence Metric Performance Against Human Quality Judgments

    empirical result

    In an evaluation of 148 topics generated by Latent Dirichlet Allocation against expert consensus labels (90 good, 21 intermediate, 37 bad) from the National Institutes of Health (NIH):

    • Topic Coherence Metric: Achieved an Average Precision (AP) of 0.94 and Area Under the ROC Curve (AUC) of 0.87 when classifying good topics versus intermediate/bad topics. In a logistic regression model predicting whether a topic was rated bad, topic coherence alone achieved an Akaike Information Criterion (AIC) of 113.8.
    • Topic Size Baseline (number of tokens assigned to the topic): Achieved an AP of 0.89, AUC of 0.79, and logistic regression AIC of 152.5. Combining topic size with coherence yielded an AIC of 115.8, indicating that coherence alone provides the best model fit.
    • Pointwise Mutual Information (PMI): Evaluated using training set co-occurrences, PMI achieved an AUC of 0.64 and logistic regression AIC of 170.51.
    • Qualitative Ranking: Among the 20 highest-scoring topics by coherence, 18 were expert-labeled good, 1 intermediate, and 1 bad. Among the 20 lowest-scoring topics, 15 were bad, 4 intermediate, and only 1 good.
  7. Knowl 7 — Limitations of Word Intrusion Detection on Chained Semantic Topics

    empirical result

    In a word intrusion evaluation with 10 domain experts (where an intruder word was randomly drawn from the corpus vocabulary and inserted into the top 10 words of a topic):

    1. Predictive Performance: Intruder detection accuracy performed poorly at predicting expert topic quality, yielding a logistic regression AIC of 163.18 (worse than topic size at 152.5 and topic coherence at 113.8).
    2. Failure on Chained Topics: Domain experts easily identified random intruder words in bad topics characterized by semantic chaining. For instance, in the chained topic receptors, cannabinoid, cannabinoids, ligands, cannabis, endocannabinoid, cxcr4, [virus], receptor, sdf1 (where CXCR4 is linked to cannabinoids only through general receptor terms), 8 out of 10 annotators correctly identified [virus] as the intruder despite the topic being semantically flawed, leading to high intrusion accuracy for a bad topic.
  8. Knowl 8 — Topic Separation Preservation via High-Frequency Truncation in Pólya Urns

    empirical result

    In Generalized Pólya Urn topic models, retaining non-zero off-diagonal co-occurrence values for high-frequency word types in schema matrix A\mathbf{A} causes frequent words to disperse across all topics, collapsing topic specificity and producing highly redundant topics.

    Evaluating topic separation via the Jensen-Shannon (JS) divergence between topic-word distributions at T=200T = 200 topics:

    • Untruncated Schema: When non-zero off-diagonal weights are used for all word types, the mean JS divergence among the 100 most similar topic pairs drops to 0.29±0.050.29 \pm 0.05 (where 1.0 represents distributions with disjoint support).
    • Standard LDA Baseline (A=I\mathbf{A} = \mathbf{I}): The mean JS divergence among the 100 most similar topic pairs is 0.67±0.050.67 \pm 0.05.
    • Truncated Schema (zeroing off-diagonal elements for words with IDF<3.0\text{IDF} < 3.0 / document frequency >5%> 5\%): The mean JS divergence among the 100 most similar topic pairs increases to 0.822±0.090.822 \pm 0.09, preventing topic collapse and exceeding standard LDA topic distinction.
  9. Knowl 9 — Topic Coherence and Convergence Rates in Generalized Pólya Urn Topic Models

    empirical result

    Across 5-fold cross-validation on 18,756 NIH grant abstracts (2,150,172 tokens, vocabulary size 28,702) evaluated at T∈{100,200,300,400}T \in \{100, 200, 300, 400\} topics:

    • Average Topic Coherence and Convergence: The Generalized Pólya Urn (GPU) model reaches its final average topic coherence score within the first 50 Gibbs sampling iterations, consistently outperforming standard LDA (which requires hundreds of iterations).
    • Improvement on Worst Topics: GPU-LDA consistently improves the average coherence score of the 10 lowest-scoring topics across all tested topic counts TT.
    • Blind Expert Annotation: In a double-blind evaluation of 200 topics (T=200T = 200, 103 GPU topics and 97 LDA topics), the topic coherence metric predicted bad topics with an AUC of 0.83 for GPU and 0.80 for LDA. Topics marked as bad decreased from 16.5% in LDA to 13.5% in GPU (not statistically significant at p=0.05p = 0.05).
    • Held-Out Log Probability Discrepancy: GPU-LDA achieved higher held-out document log probability during early iterations (iterations 1–100), but was eventually overtaken by standard LDA, illustrating that held-out probability does not track human semantic coherence judgments.

Coverage note — None was omitted; all key theoretical definitions, model equations, Gibbs sampling algorithms, qualitative failure taxonomies, and empirical evaluation results from the paper were extracted.

References

  1. 1.Loulwah AlSumait, Daniel Barbara, James Gentle, and Carlotta Domeniconi. 2009. Topic significance ranking of LDA generative models. In ECML.
  2. 2.David Andrzejewski, Xiaojin Zhu, and Mark Craven. 2009. Incorporating domain knowledge into topic modeling via Dirichlet forest priors. In Proceedings of the 26th Annual International Conference on Machine Learning, pages 25–32.
  3. 3.David M. Blei, Andrew Y. Ng, and Michael I. Jordan. 2003. Latent Dirichlet allocation. Journal of Machine Learning Research, 3:993–1022, January.
  4. 4.K.R. Canini, L. Shi, and T.L. Griffiths. 2009. Online inference of topics with latent Dirichlet allocation. In Proceedings of the 12th International Conference on Artificial Intelligence and Statistics.
  5. 5.Jonathan Chang, Jordan Boyd-Graber, Chong Wang, Sean Gerrish, and David M. Blei. 2009. Reading tea leaves: How humans interpret topic models. In Advances in Neural Information Processing Systems 22, pages 288–296.
  6. 6.Kenneth Church and Patrick Hanks. 1990. Word association norms, mutual information, and lexicography. Computational Linguistics, 6(1):22–29.
  7. 7.Gabriel Doyle and Charles Elkan. 2009. Accounting for burstiness in topic models. In ICML.
  8. 8.S. Geman and D. Geman. 1984. Stochastic relaxation, Gibbs distributions, and the Bayesian restoration of images. IEEE Transaction on Pattern Analysis and Machine Intelligence 6, pages 721–741.
  9. 9.Thomas L. Griffiths and Mark Steyvers. 2004. Finding scientific topics. Proceedings of the National Academy of Sciences, 101(suppl. 1):5228–5235.
  10. 10.Matthew Hoffman, David Blei, and Francis Bach. 2010. Online learning for latent dirichlet allocation. In NIPS.
  11. 11.Hosan Mahmoud. 2008. Pólya Urn Models. Chapman & Hall/CRC Texts in Statistical Science.
  12. 12.Qiaozhu Mei, Xuehua Shen, and ChengXiang Zhai. 2007. Automatic labeling of multinomial topic models. In Proceedings of the 13th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pages 490–499.
  13. 13.David Newman, Jey Han Lau, Karl Grieser, and Timothy Baldwin. 2010. Automatic evaluation of topic coherence. In Human Language Technologies: The Annual Conference of the North American Chapter of the Association for Computational Linguistics.
  14. 14.Yee Whye Teh, Dave Newman, and Max Welling. 2006. A collapsed variational Bayesian inference algorithm for lat ent Dirichlet allocation. In Advances in Neural Information Processing Systems 18.
  15. 15.Hanna Wallach, Iain Murray, Ruslan Salakhutdinov, and David Mimno. 2009. Evaluation methods for topic models. In Proceedings of the 26th Interational Conference on Machine Learning.
  16. 16.Xing Wei and Bruce Croft. 2006. LDA-based document models for ad-hoc retrival. In Proceedings of the 29th Annual International SIGIR Conference.

Citation

MLA
Mimno, D., et al. “Optimizing Semantic Coherence in Topic Models”. ScholarWorks@UMassAmherst (University of Massachusetts Amherst), 2011, pp. 262–72, https://works.bepress.com/andrew_mccallum/75.
APA
Mimno, D., Wallach, H., Talley, E. M., Leenders, M., & McCallum, A. (2011). Optimizing Semantic Coherence in Topic Models. ScholarWorks@UMassAmherst (University of Massachusetts Amherst), 262–272. https://works.bepress.com/andrew_mccallum/75
Chicago
Mimno, D., H. Wallach, E. M. Talley, M. Leenders, and A. McCallum. 2011. “Optimizing Semantic Coherence in Topic Models”. ScholarWorks@UMassAmherst (University of Massachusetts Amherst), 262–72. https://works.bepress.com/andrew_mccallum/75.
Harvard
Mimno, D. et al. (2011) “Optimizing Semantic Coherence in Topic Models”, ScholarWorks@UMassAmherst (University of Massachusetts Amherst), pp. 262–272. Available at: https://works.bepress.com/andrew_mccallum/75.
Vancouver
1. Mimno D, Wallach H, Talley EM, Leenders M, McCallum A (2011) Optimizing Semantic Coherence in Topic Models. ScholarWorks@UMassAmherst (University of Massachusetts Amherst) 262–272

BibTeX

@article{mimno2011optimizing,
  title = {Optimizing Semantic Coherence in Topic Models},
  author = {Mimno, David and Wallach, Hanna and Talley, Edmund M. and Leenders, Miriam and McCallum, Andrew},
  year = {2011},
  journal = {ScholarWorks@UMassAmherst (University of Massachusetts Amherst)},
  pages = {262-272},
  url = {https://works.bepress.com/andrew_mccallum/75}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by-nc-sa/4.0/