Latent Dirichlet allocation (LDA) and topic modeling: models, applications, a survey

Hamed JelodarYongli WangChi YuanXia FengXia-Hui JiangYanchao LiLiang Zhao

article2017Multimedia tools and applications1,813 citations

Synthesizes over a decade of Latent Dirichlet Allocation research by categorizing major model extensions, cross-disciplinary applications, benchmark datasets, and practical software tools for text mining.

Listen

Modern organizations and researchers face significant challenges in analyzing, structuring, and extracting meaningful patterns from enormous collections of unstructured electronic text. Topic modeling has emerged as a critical automated technique within machine learning to discover hidden themes across vast archives without requiring manual labeling. Latent Dirichlet Allocation (LDA) serves as the foundational, widely adopted framework for these efforts. The article provides a comprehensive survey of LDA-based research from 2003 through 2016, charting the development, core inference mechanisms, domain-specific adaptations, tools, datasets, and open technical challenges across the field.

To conduct this evaluation, the authors performed an extensive literature review of influential academic studies and developments over the 14-year period following the initial formulation of LDA. The survey reviews generative probability principles, parameter estimation and training methods—such as Gibbs sampling, Expectation-Maximization, and Variational Bayes inference—and analyzes empirical implementations spanning text corpora, source code repositories, social networks, and multimedia datasets.

Several key findings emerge from the review. First, Gibbs sampling is the most widely adopted inference method in practice, though variational methods are increasingly used to scale algorithms across distributed frameworks. Second, LDA has evolved far beyond basic text categorization into highly specialized extensions across multiple disciplines. Key application areas include software engineering for bug localization and code refactoring, healthcare and biomedicine for adverse drug reaction discovery, and political science for speech analysis. Third, topic modeling has proven effective in social network analysis, notably on microblogging platforms for hashtag recommendation, user behavior modeling, and real-time detection of sudden events. Finally, the ecosystem is supported by mature open-source tools such as Mallet, Gensim, and Stanford TMT, alongside established multi-domain datasets.

The findings demonstrate that LDA-based frameworks significantly reduce the time, labor, and costs associated with organizing massive textual and relational data. By automatically identifying underlying thematic structures, organizations can improve recommendation systems, enhance risk detection in public health and cybersecurity, and streamline information retrieval across complex digital environments.

For future development, the article highlights seven open research areas that require further work before systems can reach optimal performance. Practitioners and researchers should focus on extending topic models to image classification, audio and music information retrieval, drug safety assessment, user behavior modeling in mobile and social networks, group discovery in graph analytics, and enhanced interactive visualization tools like LDAvis and Termite. While the surveyed models demonstrate high utility across diverse domains, the article notes limitations around handling short texts, high-velocity streaming data, and complex multi-modal inputs, suggesting that implementers carefully evaluate topic quality and scaling trade-offs when deploying these techniques in production.

arXiv: 1711.04305
  • Paper: Latent Dirichlet Allocation, David M. Blei et al. (2003). This seminal paper introduces Latent Dirichlet Allocation (LDA), providing the foundational probabilistic model that the survey reviews, categorizes, and evaluates.
  • Paper: Probabilistic Latent Semantic Analysis, Thomas Hofmann (1999). This paper establishes Probabilistic Latent Semantic Analysis (PLSA), the direct generative predecessor to LDA that provides essential conceptual background for statistical topic modeling.
  • Paper: Indexing By Latent Semantic Analysis, Scott Deerwester et al. (1990). This classic paper introduces Latent Semantic Indexing, laying the groundwork for algebraic dimensionality reduction and latent topic discovery in text collections.
  • Paper: Supervised Topic Models, David M. Blei et al. (2007). This work introduces supervised LDA, demonstrating a major structural extension of topic models to incorporate document-level response variables surveyed in the source.
  • Paper: Stochastic variational inference, Matt Hoffman et al. (2012). This article develops stochastic variational inference, a crucial computational advance that enables LDA and related Bayesian topic models to scale to massive document corpora.
  • Paper: Exploring the Space of Topic Coherence Measures, Michael Röder et al. (2015). This paper provides a systematic evaluation of automated topic coherence measures, which are essential evaluation metrics discussed in the survey's intellectual landscape of topic modeling.
  • Paper: Optimizing Semantic Coherence in Topic Models, David Mimno et al. (2011). This work introduces intrinsic semantic coherence metrics for topic models, addressing key interpretability and evaluation challenges synthesized in the survey.
  • Paper: Reading Tea Leaves: How Humans Interpret Topic Models, Jonathan D. Chang et al. (2009). This study establishes human interpretability benchmarks for topic models, detailing critical evaluation methodologies and pitfalls reviewed in the survey.
  • Paper: From Frequency to Meaning: Vector Space Models of Semantics, Peter D. Turney et al. (2010). This comprehensive survey of vector space models provides the broader mathematical and semantic foundation underlying bag-of-words document representations and matrix factorization techniques.
Cover for Latent Dirichlet allocation (LDA) and topic modeling: models, applications, a survey

Abstract

Topic modeling is one of the most powerful techniques in text mining for data mining, latent data discovery, and finding relationships among data, text documents. Researchers have published many articles in the field of topic modeling and applied in various fields such as software engineering, political science, medical and linguistic science, etc. There are various methods for topic modeling, which Latent Dirichlet allocation (LDA) is one of the most popular methods in this field. Researchers have proposed various models based on the LDA in topic modeling. According to previous work, this paper can be very useful and valuable for introducing LDA approaches in topic modeling. In this paper, we investigated scholarly articles highly (between 2003 to 2016) related to Topic Modeling based on LDA to discover the research development, current trends and intellectual structure of topic modeling. Also, we summarize challenges and introduce famous tools and datasets in topic modeling based on LDA.

Table of Contents

  • 1. Introduction
  • 2.2.2. A brief look at past work: Research between 2010 to 2011
  • 2.2.3. A brief look at past work: Research between 2012 to 2013
  • 2.2.4. A brief look from some impressive past works: Research in 2016
  • 2.2 Topic Modeling for which the area is used?
  • A. Topic modeling in Linguistic science
  • B. Topic modeling in political science
  • F. Topic modeling in Social Network / Microblogs (such as Twitter)
  • 4. Open source library and tools / datasets / Software packages and tools for the analysis
  • 4.1 library/tools
  • Reference

Knowls

  1. Knowl 1 — Latent Dirichlet Allocation Generative Model and Marginal Likelihood

    model/method

    Latent Dirichlet Allocation (LDA) is an unsupervised generative probabilistic model for collections of discrete data such as text corpora. In LDA, each document is represented as a finite mixture over an underlying set of latent topics, and each topic is characterized by a discrete probability distribution over the vocabulary. Given a corpus DD consisting of MM documents, where document d∈{1,…,M}d \in \{1, \dots, M\} contains NdN_d words, the generative process is defined as follows:

    1. For each topic t∈{1,…,T}t \in \{1, \dots, T\}, draw a multinomial word distribution ϕt∼Dirichlet(β)\phi_t \sim \text{Dirichlet}(\beta).
    2. For each document d∈{1,…,M}d \in \{1, \dots, M\}, draw a multinomial topic mixture distribution θd∼Dirichlet(α)\theta_d \sim \text{Dirichlet}(\alpha).
    3. For each word position n∈{1,…,Nd}n \in \{1, \dots, N_d\} in document dd: a. Draw a specific topic assignment zd,n∼Multinomial(θd)z_{d,n} \sim \text{Multinomial}(\theta_d). b. Draw an observed word wd,n∼Multinomial(ϕzd,n)w_{d,n} \sim \text{Multinomial}(\phi_{z_{d,n}}).

    The marginal probability of the observed corpus DD given the Dirichlet prior hyperparameters α\alpha and β\beta is obtained by integrating over the latent document-topic proportions θd\theta_d and topic-word distributions Φ={ϕ1,…,ϕT}\Phi = \{\phi_1, \dots, \phi_T\}, and summing over all possible latent topic assignments zd,nz_{d,n}:

    p(D∣α,β)=∏d=1M∫p(θd∣α)(∏n=1Nd∑zd,n=1Tp(zd,n∣θd)p(wd,n∣zd,n,Φ))dθdp(D \mid \alpha, \beta) = \prod_{d=1}^M \int p(\theta_d \mid \alpha) \left( \prod_{n=1}^{N_d} \sum_{z_{d,n}=1}^T p(z_{d,n} \mid \theta_d) p(w_{d,n} \mid z_{d,n}, \Phi) \right) d\theta_d

  2. Knowl 2 — Parameter Estimation and Inference Approaches for LDA

    model/method

    Parameter estimation and posterior inference in Latent Dirichlet Allocation involve estimating the latent distributions θ\theta, ϕ\phi, and latent topic assignments zz, as well as optimizing hyperparameters α\alpha and β\beta. Three primary algorithmic paradigms are used:

    • Gibbs Sampling: A Markov Chain Monte Carlo (MCMC) algorithm that generates samples from the joint distribution by iteratively drawing each latent variable conditioned on the current values of all other variables. Collapsed Gibbs sampling integrates out the multinomial parameters θ\theta and ϕ\phi, sampling only the discrete topic indicator variables zz.
    • Expectation-Maximization (EM): An optimization framework for discovering maximum likelihood or maximum a posteriori estimates in the presence of latent variables. It alternates between an Expectation step (E-step), which calculates the expected value of the latent topic states, and a Maximization step (M-step), which optimizes parameter estimates given those expectations.
    • Variational Bayes (VB) Inference: A deterministic approximation method that posits a tractable parametric family of distributions over latent variables and optimizes the variational parameters to minimize the Kullback-Leibler (KL) divergence to the true intractable posterior distribution.
  3. Knowl 3 — Word Velocity and Acceleration for Real-Time Bursty Topic Detection

    equation

    In streaming social text analysis, bursty topic detection identifies sudden spikes in topic volume by modeling the rate of arrival and acceleration of words or word n-grams over time. The time-decayed velocity v^ΔT(t)\hat{v}_{\Delta T}(t) of a word at timestamp tt over a time window ΔT\Delta T is defined as:

    v^ΔT(t)=∑ti≤tXi⋅exp⁡(ti−tΔT)ΔT\hat{v}_{\Delta T}(t) = \sum_{t_i \le t} X_i \cdot \frac{\exp\left(\frac{t_i - t}{\Delta T}\right)}{\Delta T}

    where XiX_i is the frequency of the word in the ii-th text message, tit_i is the timestamp of the message, and ΔT\Delta T acts as a soft moving window size giving higher weight to recent tokens.

    The word acceleration α^(t)\hat{\alpha}(t), capturing the rate of change between two distinct window scales ΔT1\Delta T_1 and ΔT2\Delta T_2 (where ΔT1≠ΔT2\Delta T_1 \ne \Delta T_2), is given by:

    α^(t)=v^ΔT2(t)−v^ΔT1(t)ΔT1−ΔT2\hat{\alpha}(t) = \frac{\hat{v}_{\Delta T_2}(t) - \hat{v}_{\Delta T_1}(t)}{\Delta T_1 - \Delta T_2}

    Bursty topics exhibit strong positive acceleration (α^(t)≫0 \hat{\alpha}(t) \gg 0), enabling their detection while filtering out general or persistent background topics that have near-zero acceleration.

  4. Knowl 4 — Software Tools and Toolboxes for Topic Modeling

    data/table
    Tool Language Inference Method Source Availability / URL
    Mallet Java Gibbs sampling http://mallet.cs.umass.edu/topics.php
    Stanford TMT Java Gibbs sampling https://nlp.stanford.edu/software/tmt/tmt-0.4/
    Mr.LDA Java Variational Bayes https://github.com/lintool/Mr.LDA
    JGibbLDA Java Gibbs sampling http://jgibblda.sourceforge.net/
    Gensim Python Gibbs sampling / Online VB https://radimrehurek.com/gensim
    TopicXP Java (Eclipse plugin) Gibbs sampling http://www.cs.wm.edu/semeru/TopicXP/
    Matlab Topic Modeling Toolbox Matlab Gibbs sampling http://psiexp.ss.uci.edu/research/programs_data/toolbox.htm
    Yahoo_LDA C++ Gibbs sampling https://github.com/shravanmn/Yahoo_LDA
    lda in R R Gibbs sampling https://cran.r-project.org/web/packages/lda/

    Open-source libraries provide implementations of LDA and related models. While most toolkits (including Mallet, JGibbLDA, Matlab Toolbox, and R packages) rely on MCMC Gibbs sampling for parameter inference, distributed frameworks designed for large document clusters (such as Mr.LDA over MapReduce) utilize Variational Bayesian inference.

  5. Knowl 5 — Benchmark Corpora for Topic Modeling Evaluation

    data/table
    Dataset Language Release Year Domain / Content Description
    Reuters-21578 English 1997 Categorized financial and news articles
    Reuters-Volume I (RCV1) English 2004 Large-scale categorized newswire articles
    UDI-TwitterCrawl-Aug2012 English 2012 Multi-million microblog tweet corpus
    SemEval-2013 English 2013 Twitter messages annotated for sentiment and semantics
    Wiki10 English 2009 Wikipedia documents with user-contributed social tags
    Weibo Dataset Chinese 2013 Microblog post dataset from Sina Weibo
    Bag of Words (UCI) English 2008 Multi-domain texts (PubMed, KOS blog, NYTimes, NIPS, Enron)
    CiteULike English 2011 Academic reference sharing metadata and abstracts
    DBLP Dataset English 2011 Computer science bibliographic metadata
    HowNet Lexicon Chinese 2000–2013 Machine-readable bilingual lexical knowledge base
    Virastyar Persian 2013 Persian electronic lexical corpus and poems
    NIPS Conference Abstracts English 2016 Full text and abstracts from NIPS papers (1987–2015)
    20 Newsgroups English 2008 Usenet messages across 20 distinct discussion groups

    Standard benchmark datasets spanning diverse languages, document lengths, and structural formats are used to evaluate topic model performance in classification, document clustering, and semantic retrieval.

  6. Knowl 6 — LDA Applications in Software Engineering

    model/method

    Latent Dirichlet Allocation has been adapted for source code analysis and software maintenance by treating source files, commit logs, and defect logs as textual documents:

    • Software Categorization and Similarity: Models like LACT extract topic distributions from software repositories (e.g., Apache, SourceForge) to calculate inter-project similarity and categorize software by programming language or domain.
    • Software Traceability Link Recovery: Combining LDA with automated link capture connects requirements documents, architectural models, and source code classes based on topic overlap.
    • Coupling Metrics: Relational Topic Models (RTM) form the basis for Relational Topic based Coupling (RTC), which quantifies structural and conceptual coupling between object-oriented classes using latent topic distributions.
    • Bug Localization: LDA models source code repositories and bug reports to retrieve relevant source entities and locate defect locations, outperforming traditional Latent Semantic Indexing (LSI).
    • Malicious Behavior Identification: Topic modeling coupled with Genetic Algorithms characterizes Android applications by pairing user descriptions with sensitive API data-flow signatures to identify malicious patterns.
  7. Knowl 7 — LDA Applications in Biomedical Informatics and Drug Safety

    model/method

    In medical and biological data mining, LDA extensions process literature and clinical records for discovery tasks:

    • Adverse Drug Reaction (ADR) Discovery: Symbolic and probabilistic LDA models extract patterns from pharmacovigilance databases (such as ADReCS and SIDER) to predict unknown drug side effects and support drug repositioning.
    • Pharmacogenomics (PGx): Models like LDA-PKL utilize Kullback-Leibler divergence between topic distributions over biomedical literature (PubMed) to rank and detect gene-drug associations.
    • Functional Genomic Modules: Correspondence Latent Dirichlet Allocation (Corr-LDA) integrates heterogeneous miRNA and mRNA expression datasets to discover functional regulatory modules.
    • Clinical Pathway Extraction and Traditional Medicine: The Symptom-Herb-Diagnosis Topic (SHDT) model applies Author-Topic architectures to clinical records of diabetic patients to link symptoms with herbal prescriptions and disease comorbidities.
    • Healthcare Recommendation: Recommender systems such as iDoctor combine hybrid matrix factorization and LDA to extract physician feature topics from patient review corpora.
  8. Knowl 8 — Geographical and Multimodal Topic Modeling

    model/method

    Geographical topic modeling integrates spatial metadata (GPS coordinates, geotags) and visual data with textual content:

    • Spatial-Textual Integration: GeoFolk uses multi-modal Bayesian models to combine spatial coordinates and textual tags from social media (such as Flickr), enabling location-aware content retrieval.
    • Latent Geographical Topic Analysis (LGTA): Merges text-driven topic extraction with geographical clustering algorithms to model spatial distributions of topics.
    • Geographic Lexical Variation: Decomposes text into geographic and topical components to discover regional linguistic variations and infer an author's location from text.
    • Remote Sensing and Image Clustering: Multiscale LDA models map high-resolution satellite images into visual words and latent topics to perform semantic object-oriented clustering and land-cover segmentation.
  9. Knowl 9 — Topic Modeling in Political Science and Social Networks

    model/method

    In political science and microblog analytics, LDA extensions model opinions, event dynamics, and user preferences:

    • Political Speech and Stance Modeling: Cross-Perspective Topic (CPT) models and two-layer matrix factorization analyze plenary speeches (e.g., European Parliament and U.S. Senate) to detect ideological stances, political agendas, and contrastive opinions across parties.
    • Hashtag and Follower Recommendation: Models such as Hashtag-LDA, TWILITE, and TOT-MMM (Topic-Over-Time Mixed Membership Model) combine temporal clustering and topic distributions to recommend relevant tags and users.
    • Event Tracking and Sentiment Dynamics: Frameworks such as Event and Tweets LDA (ET-LDA) and Hierarchical Dirichlet Processes (HDP) isolate distinct sub-events and track shifts in public sentiment from social media posts.
    • Community and Anomaly Detection: Models like the Author-Topic-Community (ATC) model and Group Latent Anomaly Detection (GLAD) jointly capture user interactions and content to identify latent social circles and behavioral anomalies.
  10. Knowl 10 — Open Research Challenges and Gaps in Topic Modeling

    limitation

    Seven major open challenges and research gaps remain in probabilistic topic modeling:

    1. Joint Image Classification and Annotation: Achieving robust multi-modal integration of visual region descriptors, object geometry, and class labels under extreme viewpoint variations.
    2. Continuous Audio and Music Processing: Formulating continuous distributions (e.g., Gaussian mixtures) over acoustic feature spaces to model singing timbre, instrument mixtures, and audio events.
    3. Drug Safety Mining from Unstructured Social Text: Extracting adverse drug reactions and consumer health expressions from noisy social media to complement formal biomedical databases.
    4. Cross-Cultural Narrative and Demographic Modeling: Bridging multilingual gaps in public comment streams to extract common semantics across languages and demographic groups.
    5. Large-Scale Graph Group Discovery: Scaling Bayesian topic models to discover overlapping communities and anomalous groupings in massive relational networks.
    6. User Behavior and Mobility Modeling: Synergistically combining heterogeneous mobile web logs, location trajectories, and interaction histories into unified topic representations.
    7. Interactive Topic Model Visualization: Designing visual diagnostic tools (such as Termite, LDA-SOM, and LDAvis) that leverage term saliency, distinctiveness, and interactive projections to interpret complex fitted models.

Coverage note — Specific sub-model citations and minor one-off application case studies mentioned in the survey's extensive literature tables were synthesized into the respective domain-level knowls rather than listed individually.

References

  1. 1.AHMED, A., ALY, M., GONZALEZ, J., NARAYANAMURTHY, S. & SMOLA, A. J. Scalable inference in latent variable models. Proceedings of the fifth ACM international conference on Web search and data mining, 2012. ACM, 123-132.
  2. 2.ALAM, M. H., RYU, W.-J. & LEE, S. 2016. Joint multi-grain topic sentiment: modeling semantic aspects for online reviews. Information Sciences, 339, 206-223.
  3. 3.ALASHRI, S., KANDALA, S. S., BAJAJ, V., RAVI, R., SMITH, K. L. & DESOUZA, K. C. An analysis of sentiments on facebook during the 2016 US presidential election. Advances in Social Networks Analysis and Mining (ASONAM), 2016 IEEE/ACM International Conference on, 2016. IEEE, 795-802.
  4. 4.ALSUMAIT, L., BARBARÁ, D. & DOMENICONI, C. On-line lda: Adaptive topic models for mining text streams with applications to topic detection and tracking. Data Mining, 2008. ICDM'08. Eighth IEEE International Conference on, 2008. IEEE, 3-12.
  5. 5.ASGARI, E. & CHAPPELIER, J.-C. Linguistic Resources and Topic Models for the Analysis of Persian Poems. CLfL@ NAACL-HLT, 2013. 23-31.
  6. 6.ASUNCION, H. U., ASUNCION, A. U. & TAYLOR, R. N. Software traceability with topic modeling. Proceedings of the 32nd ACM/IEEE International Conference on Software Engineering-Volume 1, 2010. ACM, 95-104.
  7. 7.BAGHERI, A., SARAEE, M. & DE JONG, F. 2014. ADM-LDA: An aspect detection model based on topic modelling using the structure of review sentences. Journal of Information Science, 40, 621-636.
  8. 8.BALASUBRAMANYAN, R., COHEN, W. W., PIERCE, D. & REDLAWSK, D. P. Modeling polarizing topics: When do different political communities respond differently to the same news? ICWSM, 2012.
  9. 9.BAUER, S., NOULAS, A., SÉAGHDHA, D. O., CLARK, S. & MASCOLO, C. Talking places: Modelling and analysing linguistic content in foursquare. Privacy, Security, Risk and Trust (PASSAT), 2012 International Conference on and 2012 International Confernece on Social Computing (SocialCom), 2012. IEEE, 348-357.
  10. 10.BHATTACHARYA, P., ZAFAR, M. B., GANGULY, N., GHOSH, S. & GUMMADI, K. P. Inferring user interests in the twitter social network. Proceedings of the 8th ACM Conference on Recommender systems, 2014. ACM, 357-360.
  11. 11.BISGIN, H., LIU, Z., FANG, H., KELLY, R., XU, X. & TONG, W. 2014. A phenome-guided drug repositioning through a latent variable model. BMC bioinformatics, 15, 267.
  12. 12.BLEI, D. M. & JORDAN, M. I. Modeling annotated data. Proceedings of the 26th annual international ACM SIGIR conference on Research and development in informaion retrieval, 2003. ACM, 127-134.
  13. 13.BLEI, D. M. & LAFFERTY, J. D. Dynamic topic models. Proceedings of the 23rd international conference on Machine learning, 2006. ACM, 113-120.
  14. 14.BLEI, D. M., NG, A. Y. & JORDAN, M. I. 2003. Latent dirichlet allocation. Journal of machine Learning research, 3, 993-1022.
  15. 15.CHANEY, A. J.-B. & BLEI, D. M. Visualizing Topic Models. ICWSM, 2012.
  16. 16.CHANG, J. 2011. lda: Collapsed Gibbs sampling methods for topic models. R.
  17. 17.CHANG, J. & BLEI, D. M. Relational topic models for document networks. International conference on artificial intelligence and statistics, 2009. 81-88.
  18. 18.CHEN, B., ZHU, L., KIFER, D. & LEE, D. What Is an Opinion About? Exploring Political Standpoints Using Opinion Scoring Model. AAAI, 2010.
  19. 19.CHEN, L., WANG, Y., YU, Q., ZHENG, Z. & WU, J. WT-LDA: user tagging augmented LDA for web service clustering. International Conference on Service-Oriented Computing, 2013. Springer, 162-176.
  20. 20.CHEN, S.-H., SANTOSO, A., LEE, Y.-S. & WANG, J.-C. Latent dirichlet allocation based blog analysis for criminal intention detection system. Security Technology (ICCST), 2015 International Carnahan Conference on, 2015. IEEE, 73-76.
  21. 21.CHEN, T.-H., THOMAS, S. W., NAGAPPAN, M. & HASSAN, A. E. Explaining software defects using topic models. Mining Software Repositories (MSR), 2012 9th IEEE Working Conference on, 2012. IEEE, 189-198.
  22. 22.CHENG, V. C., LEUNG, C. H., LIU, J. & MILANI, A. 2014. Probabilistic aspect mining model for drug reviews. IEEE transactions on knowledge and data engineering, 26, 2002-2013.
  23. 23.CHENG, Z. & SHEN, J. 2016. On effective location-aware music recommendation. ACM Transactions on Information Systems (TOIS), 34, 13.
  24. 24.CHIEN, J.-T. & CHUEH, C.-H. 2011. Dirichlet class language models for speech recognition. IEEE Transactions on Audio, Speech, and Language Processing, 19, 482-495.
  25. 25.CHONG, W., BLEI, D. & LI, F.-F. Simultaneous image classification and annotation. Computer Vision and Pattern Recognition, 2009. CVPR 2009. IEEE Conference on, 2009. IEEE, 1903-1910.
  26. 26.CHOO, J., LEE, C., REDDY, C. K. & PARK, H. 2013. Utopian: User-driven topic modeling based on interactive nonnegative matrix factorization. IEEE transactions on visualization and computer graphics, 19, 1992-2001.
  27. 27.CHUANG, J., MANNING, C. D. & HEER, J. Termite: Visualization techniques for assessing textual topic models. Proceedings of the international working conference on advanced visual interfaces, 2012. ACM, 74-77.
  28. 28.COHEN, R., AVIRAM, I., ELHADAD, M. & ELHADAD, N. 2014. Redundancy-aware topic modeling for patient record notes. PloS one, 9, e87555.
  29. 29.COHEN, R. & RUTHS, D. Classifying political orientation on Twitter: It's not easy! ICWSM, 2013.
  30. 30.CONG, Y., QIN, Z., YU, J. & WAN, T. Cross-Modal Information Retrieval–A Case Study on Chinese Wikipedia. International Conference on Advanced Data Mining and Applications, 2012. Springer, 15-26.
  31. 31.CORDEIRO, M. Twitter event detection: combining wavelet analysis and topic inference summarization. Doctoral symposium on informatics engineering, 2012. 11-16.
  32. 32.CRISTANI, M., PERINA, A., CASTELLANI, U. & MURINO, V. Geo-located image analysis using latent representations. Computer Vision and Pattern Recognition, 2008. CVPR 2008. IEEE Conference on, 2008. IEEE, 1-8.
  33. 33.DIAO, Q., JIANG, J., ZHU, F. & LIM, E.-P. Finding bursty topics from microblogs. Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics: Long Papers-Volume 1, 2012. Association for Computational Linguistics, 536-544.
  34. 34.EIDELMAN, V., BOYD-GRABER, J. & RESNIK, P. Topic models for dynamic translation model adaptation. Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics: Short Papers-Volume 2, 2012. Association for Computational Linguistics, 115-119.
  35. 35.EISENSTEIN, J., O'CONNOR, B., SMITH, N. A. & XING, E. P. A latent variable model for geographic lexical variation. Proceedings of the 2010 Conference on Empirical Methods in Natural Language Processing, 2010. Association for Computational Linguistics, 1277-1287.
  36. 36.EVERINGHAM, M., VAN GOOL, L., WILLIAMS, C. K., WINN, J. & ZISSERMAN, A. 2008. The pascal visual object classes challenge 2007 (voc 2007) results (2007).
  37. 37.EVERINGHAM, M., VAN GOOL, L., WILLIAMS, C. K., WINN, J. & ZISSERMAN, A. 2010. The pascal visual object classes (voc) challenge. International journal of computer vision, 88, 303-338.
  38. 38.FANG, Y., SI, L., SOMASUNDARAM, N. & YU, Z. Mining contrastive opinions on political texts using cross-perspective topic model. Proceedings of the fifth ACM international conference on Web search and data mining, 2012. ACM, 63-72.
  39. 39.FU, X., LI, J., YANG, K., CUI, L. & YANG, L. 2016. Dynamic online HDP model for discovering evolutionary topics from Chinese social texts. Neurocomputing, 171, 412-424.
  40. 40.FU, X., YANG, K., HUANG, J. Z. & CUI, L. 2015. Dynamic non-parametric joint sentiment topic mixture model. Knowledge-Based Systems, 82, 102-114.
  41. 41.GERBER, M. S. 2014. Predicting crime using Twitter and kernel density estimation. Decision Support Systems, 61, 115-125.
  42. 42.GETHERS, M. & POSHYVANYK, D. Using relational topic models to capture coupling among classes in object-oriented software systems. Software Maintenance (ICSM), 2010 IEEE International Conference on, 2010. IEEE, 1-10.
  43. 43.GIRI, R., CHOI, H., HOO, K. S. & RAO, B. D. User behavior modeling in a cellular network using latent dirichlet allocation. International Conference on Intelligent Data Engineering and Automated Learning, 2014. Springer, 36-44.
  44. 44.GODIN, F., SLAVKOVIKJ, V., DE NEVE, W., SCHRAUWEN, B. & VAN DE WALLE, R. Using topic models for twitter hashtag recommendation. Proceedings of the 22nd International Conference on World Wide Web, 2013. ACM, 593-596.
  45. 45.GREENE, D. & CROSS, J. P. Unveiling the Political Agenda of the European Parliament Plenary: A Topical Analysis. Proceedings of the ACM Web Science Conference, 2015. ACM, 2.
  46. 46.GRETARSSON, B., O’DONOVAN, J., BOSTANDJIEV, S., HÖLLERER, T., ASUNCION, A., NEWMAN, D. & SMYTH, P. 2012. Topicnets: Visual analysis of large text corpora with topic modeling. ACM Transactions on Intelligent Systems and Technology (TIST), 3, 23.
  47. 47.GRIFFITHS, T. L. & STEYVERS, M. 2004. Finding scientific topics. Proceedings of the National academy of Sciences, 101, 5228-5235.
  48. 48.GUO, J., XU, G., CHENG, X. & LI, H. Named entity recognition in query. Proceedings of the 32nd international ACM SIGIR conference on Research and development in information retrieval, 2009. ACM, 267-274.
  49. 49.HEINTZ, I., GABBARD, R., SRINIVASAN, M., BARNER, D., BLACK, D. S., FREEDMAN, M. & WEISCHEDEL, R. Automatic extraction of linguistic metaphor with lda topic modeling. Proceedings of the First Workshop on Metaphor in NLP, 2013. 58-66.
  50. 50.HENDERSON, K. & ELIASSI-RAD, T. Applying latent dirichlet allocation to group discovery in large graphs. Proceedings of the 2009 ACM symposium on Applied Computing, 2009. ACM, 1456-1461.
  51. 51.HONG, L., DAN, O. & DAVISON, B. D. Predicting popular messages in twitter. Proceedings of the 20th international conference companion on World wide web, 2011. ACM, 57-58.
  52. 52.HONG, L., FRIAS-MARTINEZ, E. & FRIAS-MARTINEZ, V. Topic Models to Infer Socio-Economic Maps. AAAI, 2016. 3835-3841.
  53. 53.HOU, L., LI, J., WANG, Z., TANG, J., ZHANG, P., YANG, R. & ZHENG, Q. 2015. Newsminer: Multifaceted news analysis for event search. Knowledge-Based Systems, 76, 17-29.
  54. 54.HU, P., LIU, W., JIANG, W. & YANG, Z. 2014. Latent topic model for audio retrieval. Pattern Recognition, 47, 1138-1143.
  55. 55.HU, Y., JOHN, A., WANG, F. & KAMBHAMPATI, S. ET-LDA: Joint Topic Modeling for Aligning Events and their Twitter Feedback. AAAI, 2012. 59-65.
  56. 56.HUANG, Z., LU, X. & DUAN, H. 2013. Latent treatment pattern discovery for clinical processes. Journal of medical systems, 37, 9915.
  57. 57.JAGARLAMUDI, J. & DAUMÉ III, H. Extracting Multilingual Topics from Unaligned Comparable Corpora. ECIR, 2010. Springer, 444-456.
  58. 58.JIANG, D., VOSECKY, J., LEUNG, K. W.-T., YANG, L. & NG, W. 2015. SG-WSTD: A framework for scalable geographic web search topic discovery. Knowledge-Based Systems, 84, 18-33.
  59. 59.JIANG, Z., ZHOU, X., ZHANG, X. & CHEN, S. Using link topic model to analyze traditional chinese medicine clinical symptom-herb regularities. e-Health Networking, Applications and Services (Healthcom), 2012 IEEE 14th International Conference on, 2012. IEEE, 15-18.
  60. 60.JO, Y. & OH, A. H. Aspect and sentiment unification model for online review analysis. Proceedings of the fourth ACM international conference on Web search and data mining, 2011. ACM, 815-824.
  61. 61.KIM, M., KANG, K., PARK, D., CHOO, J. & ELMQVIST, N. 2017. Topiclens: Efficient multi-level visual topic exploration of large-scale document collections. IEEE transactions on visualization and computer graphics, 23, 151-160.
  62. 62.KIM, Y. & SHIM, K. 2014. TWILITE: A recommendation system for Twitter using a probabilistic model based on latent Dirichlet allocation. Information Systems, 42, 59-77.
  63. 63.LACOSTE-JULIEN, S., SHA, F. & JORDAN, M. I. DiscLDA: Discriminative learning for dimensionality reduction and classification. Advances in neural information processing systems, 2009. 897-904.
  64. 64.LANGE, D. & NAUMANN, F. Frequency-aware similarity measures: why Arnold Schwarzenegger is always a duplicate. Proceedings of the 20th ACM international conference on Information and knowledge management, 2011. ACM, 243-248.
  65. 65.LARKEY, L. S. & CONNELL, M. E. Arabic Information Retrieval at UMass in TREC-10. TREC, 2001.
  66. 66.LEE, S., KIM, S., LEE, S., YOON, H., LEE, D., CHOI, J. & LEE, J.-R. 2016. LARGen: automatic signature generation for Malwares using latent Dirichlet allocation. IEEE Transactions on Dependable and Secure Computing.
  67. 67.LEVY, K. E. & FRANKLIN, M. 2014. Driving regulation: using topic models to examine political contention in the US trucking industry. Social Science Computer Review, 32, 182-194.
  68. 68.LEWIS, D. D. 1997. Reuters-21578 text categorization collection.
  69. 69.LEWIS, D. D., YANG, Y., ROSE, T. G. & LI, F. 2004. Rcv1: A new benchmark collection for text categorization research. Journal of machine learning research, 5, 361-397.
  70. 70.LI, C., CHEUNG, W. K., YE, Y., ZHANG, X., CHU, D. & LI, X. 2015a. The author-topic-community model for author interest profiling and community discovery. Knowledge and Information Systems, 44, 359-383.
  71. 71.LI, C., RANA, S., PHUNG, D. & VENKATESH, S. 2016a. Hierarchical Bayesian nonparametric models for knowledge discovery from electronic medical records. Knowledge-Based Systems, 99, 168-182.
  72. 72.LI, F., HUANG, M. & ZHU, X. Sentiment Analysis with Global Topics and Local Dependency. AAAI, 2010. 1371-1376.
  73. 73.LI, J., CARDIE, C. & LI, S. TopicSpam: a Topic-Model based approach for spam detection. ACL (2), 2013. 217-221.
  74. 74.LI, R., WANG, S., DENG, H., WANG, R. & CHANG, K. C.-C. Towards social user profiling: unified and discriminative influence model for inferring home locations. Proceedings of the 18th ACM SIGKDD international conference on Knowledge discovery and data mining, 2012. ACM, 1023-1031.
  75. 75.LI, W. & MCCALLUM, A. Pachinko allocation: DAG-structured mixture models of topic correlations. Proceedings of the 23rd international conference on Machine learning, 2006. ACM, 577-584.
  76. 76.LI, X., OUYANG, J. & ZHOU, X. 2015b. Supervised topic models for multi-label classification. Neurocomputing, 149, 811-819.
  77. 77.LI, Y., ZHOU, X., SUN, Y. & ZHANG, H. 2016b. Design and implementation of Weibo sentiment analysis based on LDA and dependency parsing. China Communications, 13, 91-105.
  78. 78.LIENOU, M., MAITRE, H. & DATCU, M. 2010. Semantic annotation of satellite images using latent Dirichlet allocation. IEEE Geoscience and Remote Sensing Letters, 7, 28-32.
  79. 79.LIN, C. X., ZHAO, B., MEI, Q. & HAN, J. PET: a statistical model for popular events tracking in social communities. Proceedings of the 16th ACM SIGKDD international conference on Knowledge discovery and data mining, 2010. ACM, 929-938.
  80. 80.LIN, J., SUGIYAMA, K., KAN, M.-Y. & CHUA, T.-S. Addressing cold-start in app recommendation: latent user models constructed from twitter followers. Proceedings of the 36th international ACM SIGIR conference on Research and development in information retrieval, 2013. ACM, 283-292.
  81. 81.LINSTEAD, E., LOPES, C. & BALDI, P. An application of latent Dirichlet allocation to analyzing software evolution. Machine Learning and Applications, 2008. ICMLA'08. Seventh International Conference on, 2008. IEEE, 813-818.
  82. 82.LINSTEAD, E., RIGOR, P., BAJRACHARYA, S., LOPES, C. & BALDI, P. Mining concepts from code with probabilistic topic models. Proceedings of the twenty-second IEEE/ACM international conference on Automated software engineering, 2007. ACM, 461-464.
  83. 83.LIU, B., LIU, L., TSYKIN, A., GOODALL, G. J., GREEN, J. E., ZHU, M., KIM, C. H. & LI, J. 2010. Identifying functional miRNA–mRNA regulatory modules with correspondence latent dirichlet allocation. Bioinformatics, 26, 3105-3111.
  84. 84.LIU, Y., WANG, J. & JIANG, Y. 2016. PT-LDA: A latent variable model to predict personality traits of social network users. Neurocomputing, 210, 155-163.
  85. 85.LIU, Z., ZHANG, Y., CHANG, E. Y. & SUN, M. 2011. Plda+: Parallel latent dirichlet allocation with data placement and pipeline processing. ACM Transactions on Intelligent Systems and Technology (TIST), 2, 26.
  86. 86.LU, H.-M. & LEE, C.-H. 2015. The Topic-Over-Time Mixed Membership Model (TOT-MMM): A Twitter Hashtag Recommendation Model that Accommodates for Temporal Clustering Effects. IEEE Intelligent Systems, 1-1.
  87. 87.LU, H.-M., WEI, C.-P. & HSIAO, F.-Y. 2016. Modeling healthcare data using multiple-channel latent Dirichlet allocation. Journal of biomedical informatics, 60, 210-223.
  88. 88.LUI, M., LAU, J. H. & BALDWIN, T. 2014. Automatic detection and language identification of multilingual documents. Transactions of the Association for Computational Linguistics, 2, 27-40.
  89. 89.LUKINS, S. K., KRAFT, N. A. & ETZKORN, L. H. Source code retrieval for bug localization using latent dirichlet allocation. Reverse Engineering, 2008. WCRE'08. 15th Working Conference on, 2008. IEEE, 155-164.
  90. 90.LUKINS, S. K., KRAFT, N. A. & ETZKORN, L. H. 2010. Bug localization using latent dirichlet allocation. Information and Software Technology, 52, 972-990.
  91. 91.MADAN, A., FARRAHI, K., GATICA-PEREZ, D. & PENTLAND, A. S. Pervasive sensing to model political opinions in face-to-face networks. International Conference on Pervasive Computing, 2011. Springer, 214-231.
  92. 92.MANANDHAR, S. & YURET, D. Second joint conference on lexical and computational semantics (* sem), volume 2: Proceedings of the seventh international workshop on semantic evaluation (semeval 2013). Second Joint Conference on Lexical and Computational Semantics (* SEM), Volume 2: Proceedings of the Seventh International Workshop on Semantic Evaluation (SemEval 2013), 2013.
  93. 93.MAO, X.-L., MING, Z.-Y., CHUA, T.-S., LI, S., YAN, H. & LI, X. SSHLDA: a semi-supervised hierarchical topic model. Proceedings of the 2012 joint conference on empirical methods in natural language processing and computational natural language learning, 2012. Association for Computational Linguistics, 800-809.
  94. 94.MCCALLUM, A., CORRADA-EMMANUEL, A. & WANG, X. 2005. Topic and role discovery in social networks. Computer Science Department Faculty Publication Series, 3.
  95. 95.MCCALLUM, A. K. 2002. Mallet: A machine learning for language toolkit.
  96. 96.MCFARLAND, D. A., RAMAGE, D., CHUANG, J., HEER, J., MANNING, C. D. & JURAFSKY, D. 2013. Differentiating language usage through topic models. Poetics, 41, 607-625.
  97. 97.MCINERNEY, J. & BLEI, D. M. Discovering newsworthy tweets with a geographical topic model. NewsKDD: Data Science for News Publishing workshop Workshop in conjunction with KDD2014 the 20th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2014.
  98. 98.MIAO, J., HUANG, J. X. & ZHAO, J. 2016. TopPRF: A probabilistic framework for integrating topic space into pseudo relevance feedback. ACM Transactions on Information Systems (TOIS), 34, 22.
  99. 99.MILLAR, J. R., PETERSON, G. L. & MENDENHALL, M. J. Document Clustering and Visualization with Latent Dirichlet Allocation and Self-Organizing Maps. FLAIRS Conference, 2009. 69-74.
  100. 100.MINKA, T. & LAFFERTY, J. Expectation-propagation for the generative aspect model. Proceedings of the Eighteenth conference on Uncertainty in artificial intelligence, 2002. Morgan Kaufmann Publishers Inc., 352-359.
  101. 101.MURDOCK, J. & ALLEN, C. Visualization Techniques for Topic Model Checking. AAAI, 2015. 4284-4285.
  102. 102.NAKANO, T., YOSHII, K. & GOTO, M. Vocal timbre analysis using latent Dirichlet allocation and cross-gender vocal timbre similarity. Acoustics, Speech and Signal Processing (ICASSP), 2014 IEEE International Conference on, 2014. IEEE, 5202-5206.
  103. 103.NGUYEN, D. Q., BILLINGSLEY, R., DU, L. & JOHNSON, M. 2015. Improving topic models with latent feature word representations. Transactions of the Association for Computational Linguistics, 3, 299-313.
  104. 104.PANICHELLA, A., DIT, B., OLIVETO, R., DI PENTA, M., POSHYVANYK, D. & DE LUCIA, A. How to effectively use topic models for software engineering tasks? an approach based on genetic algorithms. Proceedings of the 2013 International Conference on Software Engineering, 2013. IEEE Press, 522-531.
  105. 105.PAUL, M. & DREDZE, M. Factorial LDA: Sparse multi-dimensional text models. Advances in Neural Information Processing Systems, 2012. 2582-2590.
  106. 106.PAUL, M. & GIRJU, R. 2010. A two-dimensional topic-aspect model for discovering multi-faceted topics. Urbana, 51, 36.
  107. 107.PAUL, M. J. & DREDZE, M. 2011. You are what you Tweet: Analyzing Twitter for public health. Icwsm, 20, 265-272.
  108. 108.PHAN, X.-H. & NGUYEN, C.-T. 2006. Jgibblda: A java implementation of latent dirichlet allocation (lda) using gibbs sampling for parameter estimation and inference.
  109. 109.PHILBIN, J., SIVIC, J. & ZISSERMAN, A. 2011. Geometric latent dirichlet allocation on a matching graph for large-scale image datasets. International journal of computer vision, 95, 138-153.
  110. 110.PRIER, K. W., SMITH, M. S., GIRAUD-CARRIER, C. & HANSON, C. L. Identifying health-related topics on twitter. International Conference on Social Computing, Behavioral-Cultural Modeling, and Prediction, 2011. Springer, 18-25.
  111. 111.QIAN, S., ZHANG, T., XU, C. & SHAO, J. 2016. Multi-modal event topic model for social event analysis. IEEE Transactions on Multimedia, 18, 233-246.
  112. 112.QIN, Z., CONG, Y. & WAN, T. 2016. Topic modeling of Chinese language beyond a bag-of-words. Computer Speech & Language, 40, 60-78.
  113. 113.RAMAGE, D., HALL, D., NALLAPATI, R. & MANNING, C. D. Labeled LDA: A supervised topic model for credit attribution in multi-labeled corpora. Proceedings of the 2009 Conference on Empirical Methods in Natural Language Processing: Volume 1-Volume 1, 2009. Association for Computational Linguistics, 248-256.
  114. 114.RAMAGE, D. & ROSEN, E. 2011. Stanford topic modeling toolbox. Dec.
  115. 115.RAO, Y. 2016. Contextual sentiment topic model for adaptive social emotion classification. IEEE Intelligent Systems, 31, 41-47.
  116. 116.RAO, Y., LEI, J., WENYIN, L., LI, Q. & CHEN, M. 2014. Building emotional dictionary for sentiment analysis of online news. World Wide Web, 17, 723-742.
  117. 117.ŘEHŮŘEK, R. & SOJKA, P. 2011. Gensim—Statistical Semantics in Python.
  118. 118.REN, Y., WANG, R. & JI, D. 2016. A topic-enhanced word embedding for Twitter sentiment classification. Information Sciences, 369, 188-198.
  119. 119.RENNIE, J. 2008. The 20 Newsgroups data set. http.
  120. 120.ROBERTS, K., ROACH, M. A., JOHNSON, J., GUTHRIE, J. & HARABAGIU, S. M. EmpaTweet: Annotating and Detecting Emotions on Twitter. LREC, 2012. 3806-3813.
  121. 121.ROSEN-ZVI, M., GRIFFITHS, T., STEYVERS, M. & SMYTH, P. The author-topic model for authors and documents. Proceedings of the 20th conference on Uncertainty in artificial intelligence, 2004. AUAI Press, 487-494.
  122. 122.SANDHAUS, E. 2008. The New York Times Annotated Corpus. Philadelphia, PA: Linguistic Data Consortium.
  123. 123.SAVAGE, T., DIT, B., GETHERS, M. & POSHYVANYK, D. Topic XP: Exploring topics in source code using Latent Dirichlet Allocation. Software Maintenance (ICSM), 2010 IEEE International Conference on, 2010. IEEE, 1-6.
  124. 124.SHARMA, V., KULSHRESHTHA, R., SINGH, P., AGRAWAL, N. & KUMAR, A. Analyzing Newspaper Crime Reports for Identification of Safe Transit Paths. HLT-NAACL, 2015. 17-24.
  125. 125.SHI, B., LAM, W., BING, L. & XU, Y. Detecting Common Discussion Topics Across Culture From News Reader Comments. ACL (1), 2016.
  126. 126.SIERSDORFER, S., CHELARU, S., PEDRO, J. S., ALTINGOVDE, I. S. & NEJDL, W. 2014. Analyzing and mining comments and comment ratings on the social web. ACM Transactions on the Web (TWEB), 8, 17.
  127. 127.SIZOV, S. Geofolk: latent spatial semantics in web 2.0 social media. Proceedings of the third ACM international conference on Web search and data mining, 2010. ACM, 281-290.
  128. 128.SONG, M., KIM, M. C. & JEONG, Y. K. 2014. Analyzing the political landscape of 2012 korean presidential election in twitter. IEEE Intelligent Systems, 29, 18-26.
  129. 129.SRIJITH, P., HEPPLE, M., BONTCHEVA, K. & PREOTIUC-PIETRO, D. 2017. Sub-story detection in Twitter with hierarchical Dirichlet processes. Information Processing & Management, 53, 989-1003.
  130. 130.STEYVERS, M. & GRIFFITHS, T. 2007. Probabilistic topic models. Handbook of latent semantic analysis, 427, 424-440.
  131. 131.STEYVERS, M. & GRIFFITHS, T. 2011. Matlab topic modeling toolbox 1.4. URL http://psiexp. ss. uci. edu/research/programs_data/toolbox. htm.
  132. 132.TAN, S., LI, Y., SUN, H., GUAN, Z., YAN, X., BU, J., CHEN, C. & HE, X. 2014. Interpreting the public sentiment variations on twitter. IEEE transactions on knowledge and data engineering, 26, 1158-1170.
  133. 133.TANG, H., SHEN, L., QI, Y., CHEN, Y., SHU, Y., LI, J. & CLAUSI, D. A. 2013. A multiscale latent Dirichlet allocation model for object-oriented clustering of VHR panchromatic satellite images. IEEE Transactions on Geoscience and Remote Sensing, 51, 1680-1692.
  134. 134.THOMAS, S. W. Mining software repositories using topic models. Proceedings of the 33rd International Conference on Software Engineering, 2011. ACM, 1138-1139.
  135. 135.THOMAS, S. W., ADAMS, B., HASSAN, A. E. & BLOSTEIN, D. Modeling the evolution of topics in source code histories. Proceedings of the 8th working conference on mining software repositories, 2011. ACM, 173-182.
  136. 136.TIAN, K., REVELLE, M. & POSHYVANYK, D. Using latent dirichlet allocation for automatic categorization of software. Mining Software Repositories, 2009. MSR'09. 6th IEEE International Working Conference on, 2009. IEEE, 163-166.
  137. 137.TITOV, I. & MCDONALD, R. Modeling online reviews with multi-grain topic models. Proceedings of the 17th international conference on World Wide Web, 2008. ACM, 111-120.
  138. 138.VADUVA, C., GAVAT, I. & DATCU, M. 2013. Latent Dirichlet allocation for spatial analysis of satellite images. IEEE Transactions on Geoscience and Remote Sensing, 51, 2770-2786.
  139. 139.VULIĆ, I., DE SMET, W. & MOENS, M.-F. Identifying word translations from comparable corpora using latent topic models. Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies: short papers-Volume 2, 2011. Association for Computational Linguistics, 479-484.
  140. 140.WALLACH, H. M., MIMNO, D. M. & MCCALLUM, A. Rethinking LDA: Why priors matter. Advances in neural information processing systems, 2009. 1973-1981.
  141. 141.WANG, C. & BLEI, D. M. Decoupling sparsity and smoothness in the discrete hierarchical dirichlet process. Advances in neural information processing systems, 2009. 1982-1989.
  142. 142.WANG, C. & BLEI, D. M. Collaborative topic modeling for recommending scientific articles. Proceedings of the 17th ACM SIGKDD international conference on Knowledge discovery and data mining, 2011. ACM, 448-456.
  143. 143.WANG, H., DING, Y., TANG, J., DONG, X., HE, B., QIU, J. & WILD, D. J. 2011. Finding complex biological relationships in recent PubMed articles using Bio-LDA. PloS one, 6, e17243.
  144. 144.WANG, J., ZHOU, J., XU, H., MEI, T., HUA, X.-S. & LI, S. 2014a. Image tag refinement by regularized latent Dirichlet allocation. Computer Vision and Image Understanding, 124, 61-70.
  145. 145.WANG, S., WANG, Z., JIANG, S. & HUANG, Q. Cross media topic analytics based on synergetic content and user behavior modeling. Multimedia and Expo (ICME), 2014 IEEE International Conference on, 2014b. IEEE, 1-6.
  146. 146.WANG, T., CAI, Y., LEUNG, H.-F., LAU, R. Y., LI, Q. & MIN, H. 2014c. Product aspect extraction supervised with online domain knowledge. Knowledge-Based Systems, 71, 86-100.
  147. 147.WANG, X., GERBER, M. S. & BROWN, D. E. 2012. Automatic Crime Prediction Using Events Extracted from Twitter Posts. SBP, 12, 231-238.
  148. 148.WANG, X. & MCCALLUM, A. Topics over time: a non-Markov continuous-time model of topical trends. Proceedings of the 12th ACM SIGKDD international conference on Knowledge discovery and data mining, 2006. ACM, 424-433.
  149. 149.WANG, Y.-C., BURKE, M. & KRAUT, R. E. Gender, topic, and audience response: an analysis of user-generated content on facebook. Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, 2013. ACM, 31-34.
  150. 150.WANG, Y., LUO, J., NIEMI, R., LI, Y. & HU, T. Catching Fire via" Likes": Inferring Topic Preferences of Trump Followers on Twitter. ICWSM, 2016. 719-722.
  151. 151.WANG, Y. & MORI, G. Max-margin Latent Dirichlet Allocation for Image Classification and Annotation. BMVC, 2011. 7.
  152. 152.WENG, J. & LEE, B.-S. 2011. Event detection in twitter. ICWSM, 11, 401-408.
  153. 153.WENG, J., LIM, E.-P., JIANG, J. & HE, Q. Twitterrank: finding topic-sensitive influential twitterers. Proceedings of the third ACM international conference on Web search and data mining, 2010. ACM, 261-270.
  154. 154.WICK, M., ROSS, M. & LEARNED-MILLER, E. Context-sensitive error correction: Using topic models to improve OCR. Document Analysis and Recognition, 2007. ICDAR 2007. Ninth International Conference on, 2007. IEEE, 1168-1172.
  155. 155.WILSON, A. T. & CHEW, P. A. Term weighting schemes for latent dirichlet allocation. human language technologies: The 2010 annual conference of the North American Chapter of the Association for Computational Linguistics, 2010. Association for Computational Linguistics, 465-473.
  156. 156.WU, H., BU, J., CHEN, C., ZHU, J., ZHANG, L., LIU, H., WANG, C. & CAI, D. 2012a. Locally discriminative topic modeling. Pattern Recognition, 45, 617-625.
  157. 157.WU, Y., LIU, M., ZHENG, W. J., ZHAO, Z. & XU, H. Ranking gene-drug relationships in biomedical literature using latent dirichlet allocation. Pacific Symposium on Biocomputing. Pacific Symposium on Biocomputing, 2012b. NIH Public Access, 422.
  158. 158.XIANGHUA, F., GUO, L., YANYAN, G. & ZHIQIANG, W. 2013. Multi-aspect sentiment analysis for Chinese online social reviews based on topic modeling and HowNet lexicon. Knowledge-Based Systems, 37, 186-195.
  159. 159.XIAO, C., ZHANG, P., CHAOVALITWONGSE, W. A., HU, J. & WANG, F. Adverse Drug Reaction Prediction with Symbolic Latent Dirichlet Allocation. AAAI, 2017. 1590-1596.
  160. 160.XIE, P., YANG, D. & XING, E. P. Incorporating Word Correlation Knowledge into Topic Modeling. HLT-NAACL, 2015. 725-734.
  161. 161.XIE, W., ZHU, F., JIANG, J., LIM, E.-P. & WANG, K. 2016. Topicsketch: Real-time bursty topic detection from twitter. IEEE Transactions on Knowledge and Data Engineering, 28, 2216-2229.
  162. 162.YANG, M.-C. & RIM, H.-C. 2014. Identifying interesting Twitter contents using topical analysis. Expert Systems with Applications, 41, 4330-4336.
  163. 163.YANG, M. & KIANG, M. Extracting Consumer Health Expressions of Drug Safety from Web Forum. System Sciences (HICSS), 2015 48th Hawaii International Conference on, 2015. IEEE, 2896-2905.
  164. 164.YANG, X., LO, D., LI, L., XIA, X., BISSYANDÉ, T. F. & KLEIN, J. 2017. Characterizing malicious Android apps by mining topic-specific data flow signatures. Information and Software Technology.
  165. 165.YANO, T., COHEN, W. W. & SMITH, N. A. Predicting response to political blog posts with topic models. Proceedings of Human Language Technologies: The 2009 Annual Conference of the North American Chapter of the Association for Computational Linguistics, 2009. Association for Computational Linguistics, 477-485.
  166. 166.YANO, T. & SMITH, N. A. What's Worthy of Comment? Content and Comment Volume in Political Blogs. ICWSM, 2010.
  167. 167.YEH, J.-F., TAN, Y.-S. & LEE, C.-H. 2016. Topic detection and tracking for conversational content by using conceptual dynamic latent Dirichlet allocation. Neurocomputing, 216, 310-318.
  168. 168.YIN, H., CUI, B., CHEN, L., HU, Z. & HUANG, Z. A temporal context-aware model for user behavior modeling in social media systems. Proceedings of the 2014 ACM SIGMOD international conference on Management of data, 2014. ACM, 1543-1554.
  169. 169.YIN, Z., CAO, L., HAN, J., ZHAI, C. & HUANG, T. Geographical topic discovery and comparison. Proceedings of the 20th international conference on World wide web, 2011. ACM, 247-256.
  170. 170.YOSHII, K. & GOTO, M. 2012. A nonparametric Bayesian multipitch analyzer based on infinite latent harmonic allocation. IEEE Transactions on Audio, Speech, and Language Processing, 20, 717-730.
  171. 171.YU, K., ZHANG, J., CHEN, M., XU, X., SUZUKI, A., ILIC, K. & TONG, W. 2014. Mining hidden knowledge for drug safety assessment: topic modeling of LiverTox as a case study. BMC bioinformatics, 15, S6.
  172. 172.YU, R., HE, X. & LIU, Y. 2015a. Glad: group anomaly detection in social media analysis. ACM Transactions on Knowledge Discovery from Data (TKDD), 10, 18.
  173. 173.YU, X., YANG, J. & XIE, Z.-Q. 2015b. A semantic overlapping community detection algorithm based on field sampling. Expert Systems with Applications, 42, 366-375.
  174. 174.YUAN, B., XU, B., WU, C. & MA, Y. Mobile Web User Behavior Modeling. International Conference on Web Information Systems Engineering, 2014. Springer, 388-397.
  175. 175.YUAN, J., GAO, F., HO, Q., DAI, W., WEI, J., ZHENG, X., XING, E. P., LIU, T.-Y. & MA, W.-Y. Lightlda: Big topic models on modest computer clusters. Proceedings of the 24th International Conference on World Wide Web, 2015. International World Wide Web Conferences Steering Committee, 1351-1361.
  176. 176.ZENG, J., LIU, Z.-Q. & CAO, X.-Q. 2016. Fast online EM for big topic modeling. IEEE Transactions on Knowledge and Data Engineering, 28, 675-688.
  177. 177.ZHAI, K., BOYD-GRABER, J., ASADI, N. & ALKHOUJA, M. L. Mr. LDA: A flexible large scale topic modeling package using variational inference in mapreduce. Proceedings of the 21st international conference on World Wide Web, 2012. ACM, 879-888.
  178. 178.ZHAI, Z., LIU, B., XU, H. & JIA, P. 2011. Constrained LDA for grouping product features in opinion mining. Advances in knowledge discovery and data mining, 448-459.
  179. 179.ZHANG, H., GILES, C. L., FOLEY, H. C. & YEN, J. Probabilistic community discovery using hierarchical latent gaussian mixture model. AAAI, 2007. 663-668.
  180. 180.ZHANG, J., LIU, B., TANG, J., CHEN, T. & LI, J. Social Influence Locality for Modeling Retweeting Behaviors. IJCAI, 2013. 2761-2767.
  181. 181.ZHANG, L., SUN, X. & ZHUGE, H. 2015. Topic discovery of clusters from documents with geographical location. Concurrency and Computation: Practice and Experience, 27, 4015-4038.
  182. 182.ZHANG, X.-P., ZHOU, X.-Z., HUANG, H.-K., FENG, Q., CHEN, S.-B. & LIU, B.-Y. 2011. Topic model for chinese medicine diagnosis and prescription regularities analysis: case on diabetes. Chinese journal of integrative medicine, 17, 307-313.
  183. 183.ZHANG, Y., CHEN, M., HUANG, D., WU, D. & LI, Y. 2017. iDoctor: Personalized and professionalized medical recommendations based on hybrid matrix factorization. Future Generation Computer Systems, 66, 30-35.
  184. 184.ZHAO, F., ZHU, Y., JIN, H. & YANG, L. T. 2016. A personalized hashtag recommendation approach using LDA-based topic model in microblog environment. Future Generation Computer Systems, 65, 196-206.
  185. 185.ZHAO, W. X., JIANG, J., WENG, J., HE, J., LIM, E.-P., YAN, H. & LI, X. Comparing twitter and traditional media using topic models. European Conference on Information Retrieval, 2011. Springer, 338-349.
  186. 186.ZHENG, X., LIN, Z., WANG, X., LIN, K.-J. & SONG, M. 2014. Incorporating appraisal expression patterns into topic modeling for aspect and sentiment word identification. Knowledge-Based Systems, 61, 29-47.
  187. 187.ZHU, J., AHMED, A. & XING, E. P. MedLDA: maximum margin supervised topic models for regression and classification. Proceedings of the 26th annual international conference on machine learning, 2009. ACM, 1257-1264.
  188. 188.ZIRN, C. & STUCKENSCHMIDT, H. 2014. Multidimensional topic analysis in political texts. Data & Knowledge Engineering, 90, 38-53.
  189. 189.ZOGHBI, S., VULIĆ, I. & MOENS, M.-F. 2016. Latent Dirichlet allocation for linking user-generated content and e-commerce data. Information Sciences, 367, 573-599.

Citation

MLA
Jelodar, H., et al. “Latent Dirichlet Allocation (LDA) and Topic Modeling: Models, Applications, a Survey”. Multimedia Tools and Applications, vol. 78, no. 11, 2018, pp. 15169–211, https://doi.org/10.1007/s11042-018-6894-4.
APA
Jelodar, H., Wang, Y., Yuan, C., Feng, X., Jiang, X., Li, Y., & Zhao, L. (2018). Latent Dirichlet allocation (LDA) and topic modeling: models, applications, a survey. Multimedia Tools and Applications, 78(11), 15169–15211. https://doi.org/10.1007/s11042-018-6894-4
Chicago
Jelodar, H., Y. Wang, C. Yuan, et al. 2018. “Latent Dirichlet Allocation (LDA) and Topic Modeling: Models, Applications, a Survey”. Multimedia Tools and Applications 78 (11): 15169–211. https://doi.org/10.1007/s11042-018-6894-4.
Harvard
Jelodar, H. et al. (2018) “Latent Dirichlet allocation (LDA) and topic modeling: models, applications, a survey”, Multimedia Tools and Applications, 78(11), pp. 15169–15211. Available at: https://doi.org/10.1007/s11042-018-6894-4.
Vancouver
1. Jelodar H, Wang Y, Yuan C, Feng X, Jiang X, Li Y, Zhao L (2018) Latent Dirichlet allocation (LDA) and topic modeling: models, applications, a survey. Multimedia Tools and Applications 78:15169–15211

BibTeX

@article{Jelodar_2018, title={Latent Dirichlet allocation (LDA) and topic modeling: models, applications, a survey}, volume={78}, ISSN={1573-7721}, url={http://dx.doi.org/10.1007/s11042-018-6894-4}, DOI={10.1007/s11042-018-6894-4}, number={11}, journal={Multimedia Tools and Applications}, publisher={Springer Science and Business Media LLC}, author={Jelodar, Hamed and Wang, Yongli and Yuan, Chi and Feng, Xia and Jiang, Xiahui and Li, Yanchao and Zhao, Liang}, year={2018}, month=Nov, pages={15169–15211} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF