Optimizing Semantic Coherence in Topic Models
David MimnoHanna M. WallachEdmund TalleyMiriam LeendersAndrew McCallum
Proposes an intrinsic semantic coherence metric and a generalized Pólya urn topic model to automatically detect and eliminate low-quality, nonsensical latent topics without requiring human evaluation or external corpora.
Statistical topic models help organizations automatically summarize and explore large document collections. However, standard methods like latent Dirichlet allocation frequently generate low-quality topics that mix unrelated concepts or arbitrarily chain disparate ideas together. When domain experts encounter these flawed topics—which can constitute up to 10% or more of total outputs—they lose trust in the automated analysis. In response, the article set out to identify the structural flaws in generated topics, develop an automated metric that accurately scores semantic quality without human annotators or external corpora, and demonstrate a new statistical model that directly improves topic quality.
The authors conducted an expert-driven annotation study using over 300,000 research abstracts from the National Institutes of Health. Medical experts categorized and labeled specific failure modes in topics, including chained, intruded, random, and unbalanced concepts. Building on these observations, the authors developed an automated topic coherence metric based purely on internal word co-occurrence statistics. They subsequently designed a generalized statistical topic model that uses an urn-based sampling schema to promote related words during training, comparing its performance and speed against standard topic modeling approaches across multiple cross-validation runs.
The analysis yielded several key findings. First, the automated topic coherence metric closely matched human domain evaluations, achieving high ranking accuracy (area under the curve of 0.87) and outperforming existing measures like topic size and mutual information without requiring outside reference text. Second, standard intruder-word evaluations failed to reliably detect certain flawed topics, such as chained concepts, which the new coherence metric successfully caught. Third, the new generalized model reached high average coherence levels within just 50 sampling iterations—far faster than standard methods—and significantly improved the quality of the ten lowest-scoring topics. Finally, while the model made the worst topics noticeably less flawed, it only modestly reduced the overall proportion of bad topics from 16.5% to 13.5%, a difference that was not statistically significant.
These findings indicate that internal document statistics contain enough unexploited information to automatically evaluate and improve topic interpretability. Organizations can filter out or flag low-quality outputs before presenting them to end-users, protecting user trust and reducing the manual oversight traditionally needed for quality control. Furthermore, the ability of the new model to converge in substantially fewer iterations offers opportunities to lower computing runtimes during training, though the sampling process per iteration requires two to three times more computational overhead due to expanded bookkeeping.
For practical implementation, organizations deploying topic models should adopt the automated coherence metric as a quality-control filter to automatically suppress or review low-scoring topics. Teams seeking higher semantic quality and faster model convergence should consider generalized urn-based topic models, balancing their faster convergence against the added per-iteration processing requirements. Future work should focus on refining inference algorithms to further eliminate the remaining baseline of bad topics and applying these word co-occurrence principles to online, large-scale streaming systems.
The study's primary limitations stem from its evaluation on a single specialized corpus of biomedical grant abstracts and the approximate nature of the sampling distribution used to maintain tractability. While decision-makers can have high confidence in the coherence metric's ability to rank and identify incoherent topics, they should exercise caution before assuming the modeling framework will completely eliminate flawed outputs without an automated filtering threshold.
- Paper: Reading Tea Leaves: How Humans Interpret Topic Models, Jonathan D. Chang et al. (2009). Its human-intrusion tests establish the interpretability evaluation that this paper refines by showing which incoherent topics those tests miss.
- Paper: Latent Dirichlet Allocation, David M. Blei et al. (2003). LDA is the standard topic-model foundation whose flawed topics and sampling behavior this paper evaluates and seeks to improve.
- Paper: Exploring the Space of Topic Coherence Measures, Michael Röder et al. (2015). This later study systematically searches coherence-measure designs and benchmarks them against human judgments, extending the source’s internal-statistics approach to topic evaluation.
