Labeled LDA: A supervised topic model for credit attribution in multi-labeled corpora

Daniel RamageDavid Leo Wright HallRamesh NallapatiChristopher D. Manning

article2009EMNLP1,480 citations

Introduces Labeled LDA, a supervised extension of Latent Dirichlet Allocation that constrains latent topics to observed document tags to directly learn word-tag correspondences and extract tag-specific snippets.

Listen

Modern digital collections, from collaborative web portals to enterprise knowledge bases, rely heavily on multi-tagging to organize content. However, human-provided tags typically apply with uneven specificity across different parts of a document rather than describing the entire text uniformly. Standard unsupervised topic models like Latent Dirichlet Allocation (LDA) fail to address this credit attribution challenge because they do not align learned latent themes with predefined human labels. Meanwhile, traditional classification algorithms such as Support Vector Machines (SVMs) lack the internal word-level attribution necessary to pinpoint precisely which portions of a document correspond to which tags.

The article introduces and evaluates Labeled LDA (L-LDA), a supervised probabilistic topic model designed to solve the credit attribution problem by constraining each document’s latent topics to match only its assigned user tags in a direct, one-to-one correspondence.

To evaluate the model, the researchers conducted empirical analyses across multiple tagging and classification benchmarks, including 4,000 tagged web pages from del.icio.us covering 20 distinct topics and multi-category datasets from the Yahoo directory. The approach uses collapsed Gibbs sampling for parameter learning and inference, comparing performance against standard unsupervised LDA and baseline one-versus-rest SVM classifiers across visualization, snippet extraction, and document-level multi-label classification tasks.

The analysis produced three primary findings. First, for tag-specific snippet extraction, human evaluators preferred L-LDA's extracted text passages over SVM outputs by more than three to one (72 preferred cases versus 21 out of 149 evaluated pairs, with 24 unanimous approvals versus only 2 for SVMs). Second, unlike standard LDA, which over-allocates topics to dominant themes and misses rare tags entirely, L-LDA guarantees full coverage and direct human interpretability for every assigned tag. Third, in multi-label classification on naturally multi-tagged web corpora, L-LDA outperformed SVM baselines, raising overall Micro-F1 accuracy from 39.33% to 52.12% while maintaining comparable Macro-F1 performance.

These findings indicate that directly integrating human label supervision into generative topic models delivers significant utility for enterprise search and document retrieval interfaces. By accurately mapping specific words and passages to defined tags, organizations can provide richer visual overviews, automated content summarization, and highly targeted snippet previews that reduce the time users spend searching for relevant information.

Organizations handling multi-labeled document repositories should consider deploying L-LDA for tag-focused passage extraction, contextual visualization, and classification. However, decision-makers should note that test-time classification in the current implementation relies on an unconstrained sampling approximation across all possible categories, and the model does not yet natively model label correlations. Prior to large-scale production deployment on complex multi-label classification pipelines, teams should validate performance on their domain-specific taxonomies and monitor ongoing developments in exact sampling algorithms and correlated topic modeling.

  • Paper: Latent Dirichlet Allocation, David M. Blei et al. (2003). It introduces Latent Dirichlet Allocation, the core generative probabilistic framework that Labeled LDA directly constrains and builds upon for credit attribution.
  • Paper: Supervised Topic Models, David M. Blei et al. (2007). It establishes the foundational supervised topic modeling paradigm (sLDA), providing key background on incorporating document-level supervisory signals into LDA.
  • Paper: The Author-Topic Model for Authors and Documents, Michal Rosen-Zvi et al. (2004). It demonstrates how to constrain topic distributions to specific document metadata (authors), serving as conceptual precursor to mapping topics directly to user-defined document labels.
  • Paper: A kernel method for multi-labelled classification, A. Elisseeff et al. (2001). It formulates the problem of multi-label classification and ranking baselines against which generative multi-label models like Labeled LDA are benchmarked.
  • Paper: Probabilistic latent semantic indexing, Thomas Hofmann (1999). It introduces probabilistic latent semantic indexing, the foundational aspect model predecessor to LDA and topic-based credit attribution.
Cover for Labeled LDA: A supervised topic model for credit attribution in multi-labeled corpora

Abstract

A significant portion of the world's text is tagged by readers on social bookmarking websites. Credit attribution is an inherent problem in these corpora because most pages have multiple tags, but the tags do not always apply with equal specificity across the whole document. Solving the credit attribution problem requires associating each word in a document with the most appropriate tags and vice versa. This paper introduces Labeled LDA, a topic model that constrains Latent Dirichlet Allocation by defining a one-to-one correspondence between LDA's latent topics and user tags. This allows Labeled LDA to directly learn word-tag correspondences. We demonstrate Labeled LDA's improved expressiveness over traditional LDA with visualizations of a corpus of tagged web pages from del.icio.us. Labeled LDA outperforms SVMs by more than 3 to 1 when extracting tag-specific document snippets. As a multi-label text classifier, our model is competitive with a discriminative baseline on a variety of datasets.

Table of Contents

  • 1 Introduction
  • 2 Labeled LDA
  • 2.1 Learning and inference
  • 2.2 Relationship to Naive Bayes
  • 3 Credit attribution within tagged documents
  • 4 Topic Visualization
  • 5 Tagged document visualization
  • 6 Snippet Extraction
  • 7 Multilabeled Text Classification
  • 7.1 Yahoo
  • 7.2 Tagged Web Pages
  • 8 Discussion
  • 9 Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Generative Process of Labeled LDA

    model/method

    Labeled Latent Dirichlet Allocation (L-LDA) is a supervised generative topic model for multi-labeled document collections that constrains Latent Dirichlet Allocation by defining a one-to-one correspondence between latent topics and human-provided labels. Let KK denote the total number of unique labels in the corpus and VV the vocabulary size.

    Each document dd consists of a list of word tokens w(d)=(w1,…,wNd)\mathbf{w}^{(d)} = (w_1, \dots, w_{N_d}) with wi∈{1,…,V}w_i \in \{1, \dots, V\}, and an observed binary label presence vector Λ(d)=(l1,…,lK)∈{0,1}K\mathbf{\Lambda}^{(d)} = (l_1, \dots, l_K) \in \{0, 1\}^K. Let λ(d)={k∈{1,…,K}∣lk=1}\lambda^{(d)} = \{k \in \{1, \dots, K\} \mid l_k = 1\} be the index set of active labels for document dd, and let Md=∣λ(d)∣M_d = |\lambda^{(d)}| be the number of active labels.

    The generative process proceeds as follows:

    1. For each topic/label k∈{1,…,K}k \in \{1, \dots, K\}, draw a multinomial word distribution: βk∼Dirichlet(η)\beta_k \sim \text{Dirichlet}(\eta) where η∈R>0V\eta \in \mathbb{R}^V_{>0} is the symmetric or asymmetric Dirichlet word prior.

    2. For each document d∈{1,…,D}d \in \{1, \dots, D\}:

    a. For each topic k∈{1,…,K}k \in \{1, \dots, K\}, generate label indicator Λk(d)∼Bernoulli(Φk)\Lambda_k^{(d)} \sim \text{Bernoulli}(\Phi_k), where Φk\Phi_k is the prior probability of label kk.

    b. Construct an Md×KM_d \times K label projection matrix L(d)L^{(d)} whose entries are: Lij(d)={1if λi(d)=j0otherwiseL_{ij}^{(d)} = \begin{cases} 1 & \text{if } \lambda_i^{(d)} = j \\ 0 & \text{otherwise} \end{cases} for row i∈{1,…,Md}i \in \{1, \dots, M_d\} and column j∈{1,…,K}j \in \{1, \dots, K\}.

    c. Compute the document-specific topic prior parameter vector α(d)∈RMd\alpha^{(d)} \in \mathbb{R}^{M_d} by projecting the global corpus topic prior α=(α1,…,αK)T\alpha = (\alpha_1, \dots, \alpha_K)^T: α(d)=L(d)α=(αλ1(d),…,αλMd(d))T\alpha^{(d)} = L^{(d)} \alpha = (\alpha_{\lambda_1^{(d)}}, \dots, \alpha_{\lambda_{M_d}^{(d)}})^T

    d. Draw the document topic mixture proportion vector restricted to active labels: θ(d)∼Dirichlet(α(d))\theta^{(d)} \sim \text{Dirichlet}(\alpha^{(d)})

    e. For each word token position i∈{1,…,Nd}i \in \{1, \dots, N_d\} in document dd:

    • Draw a topic assignment zi∈λ(d)z_i \in \lambda^{(d)} from the restricted mixture: zi∼Multinomial(θ(d))z_i \sim \text{Multinomial}(\theta^{(d)})
    • Draw word token wi∈{1,…,V}w_i \in \{1, \dots, V\} from the corresponding topic distribution: wi∼Multinomial(βzi)w_i \sim \text{Multinomial}(\beta_{z_i})
  2. Knowl 2 — Collapsed Gibbs Sampling for Labeled LDA Training

    algorithm

    When document labels Λ(d)\mathbf{\Lambda}^{(d)} are observed during training, the Bernoulli label priors Φ\Phi are d-separated from the rest of the network. Topic inference over word assignments z\mathbf{z} is performed using collapsed Gibbs sampling, where the latent topic ziz_i at word position ii in document dd is restricted to the document's active label set λ(d)={k∣Λk(d)=1}\lambda^{(d)} = \{k \mid \Lambda_k^{(d)} = 1\}.

    The conditional posterior distribution for assigning topic j∈λ(d)j \in \lambda^{(d)} to token position ii with observed word type wiw_i, conditioned on all other assignments z−i\mathbf{z}_{-i}, is:

    P(zi=j∣z−i,w,Λ(d))∝n−i,jwi+ηwin−i,j(⋅)+∑v=1Vηv×n−i,j(d)+αjn−i,⋅(d)+∑k∈λ(d)αkP(z_i = j \mid \mathbf{z}_{-i}, \mathbf{w}, \mathbf{\Lambda}^{(d)}) \propto \frac{n_{-i, j}^{w_i} + \eta_{w_i}}{n_{-i, j}^{(\cdot)} + \sum_{v=1}^V \eta_v} \times \frac{n_{-i, j}^{(d)} + \alpha_j}{n_{-i, \cdot}^{(d)} + \sum_{k \in \lambda^{(d)}} \alpha_k}

    where n−i,jwin_{-i, j}^{w_i} is the count of word type wiw_i assigned to topic jj corpus-wide excluding current position ii, n−i,j(⋅)=∑v=1Vn−i,jvn_{-i, j}^{(\cdot)} = \sum_{v=1}^V n_{-i, j}^v is the total count of all words assigned to topic jj excluding position ii, n−i,j(d)n_{-i, j}^{(d)} is the count of words in document dd assigned to topic jj excluding position ii, and n−i,⋅(d)=∑k∈λ(d)n−i,k(d)n_{-i, \cdot}^{(d)} = \sum_{k \in \lambda^{(d)}} n_{-i, k}^{(d)} is the total number of words in document dd excluding position ii.

    Input: Word tokens w(d)\mathbf{w}^{(d)} and label indicator sets λ(d)\lambda^{(d)} for documents d∈{1,…,D}d \in \{1, \dots, D\}; Dirichlet hyperparameters α,η\alpha, \eta; number of iterations TT
    Output: Word-topic counts njwn_j^w, document-topic counts nj(d)n_j^{(d)}, and word topic assignments z\mathbf{z}
    Initialize zi(d)z_i^{(d)} uniformly at random from λ(d)\lambda^{(d)} for each document dd and token i∈{1,…,Nd}i \in \{1, \dots, N_d\}
    Initialize count matrices njwn_j^w and nj(d)n_j^{(d)} based on initial assignments
    for iter = 1 to TT do
        for each document d∈{1,…,D}d \in \{1, \dots, D\} do
            for each token index i∈{1,…,Nd}i \in \{1, \dots, N_d\} do
                w←wi(d)w \leftarrow w_i^{(d)}
                jold←zi(d)j_{\text{old}} \leftarrow z_i^{(d)}
                Decrement njoldwn_{j_{\text{old}}}^w and njold(d)n_{j_{\text{old}}}^{(d)}
                
                for each candidate label j∈λ(d)j \in \lambda^{(d)} do
                    P(j)←njw+ηw∑v=1V(njv+ηv)×nj(d)+αj∑k∈λ(d)(nk(d)+αk)P(j) \leftarrow \frac{n_j^w + \eta_w}{\sum_{v=1}^V (n_j^v + \eta_v)} \times \frac{n_j^{(d)} + \alpha_j}{\sum_{k \in \lambda^{(d)}} (n_k^{(d)} + \alpha_k)}
                end for
                
                Normalize P(j)P(j) across all j∈λ(d)j \in \lambda^{(d)}
                Sample new topic assignment jnew∼P(⋅)j_{\text{new}} \sim P(\cdot)
                zi(d)←jnewz_i^{(d)} \leftarrow j_{\text{new}}
                Increment njnewwn_{j_{\text{new}}}^w and njnew(d)n_{j_{\text{new}}}^{(d)}
            end for
        end for
    end for
  3. Knowl 3 — Equivalence and Differences Between Labeled LDA and Multinomial Naive Bayes

    theoretical result

    Labeled LDA generalizes the generative event model of a Multinomial Naive Bayes classifier by introducing latent topic mixtures over observed labels:

    1. Singly Labeled Documents (Md=1M_d = 1): If every document dd in the corpus has exactly one label ldl_d, the label set constraint forces zi=ldz_i = l_d for all word positions i∈{1,…,Nd}i \in \{1, \dots, N_d\}. In this setting, collapsed Gibbs sampling increments the count of word wiw_i for topic ldl_d deterministically. The probability of document words P(w(d)∣ld)P(\mathbf{w}^{(d)} \mid l_d) under Labeled LDA is mathematically identical to the document likelihood under a standard Multinomial Naive Bayes event model trained on the same instances.

    2. Multiply Labeled Documents (Md>1M_d > 1): In a standard one-versus-rest Multinomial Naive Bayes model, a document labeled with multiple tags contributes a full count of 1 for each of its words to every label classifier independently. In contrast, Labeled LDA treats each document as a Dirichlet-distributed mixture over its observed labels λ(d)\lambda^{(d)}, distributing the credit mass of each individual word occurrence across only one topic assignment zi∈λ(d)z_i \in \lambda^{(d)} at a time via latent mixture sampling.

  4. Knowl 4 — Unlabeled Document Inference in Labeled LDA for Multi-Label Classification

    model/method

    For multi-label document classification at test time, the true label set Λ(d)\mathbf{\Lambda}^{(d)} is unobserved. Exact maximum a posteriori label set prediction requires evaluating posterior probabilities over all 2K2^K binary configurations of Λ(d)\mathbf{\Lambda}^{(d)}, where the parameter vector α(d)\alpha^{(d)} changes support with each subset.

    To perform scalable test-time inference, the model is relaxed to standard Latent Dirichlet Allocation over the full KK-dimensional topic space using the vocabulary topic distributions β1,…,βK\beta_1, \dots, \beta_K learned during supervised training. Each test word wiw_i can sample any topic zi∈{1,…,K}z_i \in \{1, \dots, K\} using standard LDA collapsed Gibbs sampling. The document's posterior topic distribution θ(d)=(θ1(d),…,θK(d))\theta^{(d)} = (\theta_1^{(d)}, \dots, \theta_K^{(d)}) is estimated by normalizing its sampled topic assignments:

    θk(d)=nk(d)+αk∑j=1K(nj(d)+αj)\theta_k^{(d)} = \frac{n_k^{(d)} + \alpha_k}{\sum_{j=1}^K (n_j^{(d)} + \alpha_j)}

    Predicted labels for the document are then determined by assigning label kk if its posterior topic probability θk(d)\theta_k^{(d)} exceeds a tuned global threshold.

  5. Knowl 5 — Tag-Specific Snippet Extraction Methodology

    experimental setup

    Tag-specific snippet extraction aims to select a passage of a document that best reflects the document's content from the perspective of a particular assigned tag kk.

    Given a document dd and a target tag k∈λ(d)k \in \lambda^{(d)}, all candidate 15-word contiguous text windows are evaluated:

    • Labeled LDA Scoring: L-LDA scores a candidate window W=(w1,…,w15)W = (w_1, \dots, w_{15}) by the expected probability that the target tag kk generated each word token in the window: ScoreL-LDA(W,k)=∏t=115P(wt∣zt=k)=∏t=115βk,wt\text{Score}_{\text{L-LDA}}(W, k) = \prod_{t=1}^{15} P(w_t \mid z_t = k) = \prod_{t=1}^{15} \beta_{k, w_t}

    • SVM Baseline Scoring: A collection of one-versus-rest linear Support Vector Machines (SVMs) is trained, one for each tag. Each 15-word window WW is represented as an independent bag-of-words pseudo-document and scored using the un-thresholded continuous decision value output by the target tag's SVM classifier.

    For each method, the 15-word window achieving the maximum score for tag kk in document dd is extracted.

  6. Knowl 6 — Human Evaluation of Tag-Specific Snippet Extraction

    empirical result

    Tag-specific snippet extraction was evaluated on 29 web pages from del.icio.us with at least two tags from a 20-tag vocabulary, forming 149 (document, tag) evaluation pairs. Three human judges rated 15-word snippets generated by Labeled LDA versus one-versus-rest SVMs in a randomized, blinded study.

    Model Best Snippet Unanimous Agreement
    Labeled LDA 72 / 149 24 / 51
    SVM 21 / 149 2 / 51

    Out of 149 total evaluations, L-LDA was judged superior in 72 cases, compared to 21 cases for SVM (with remaining cases judged as ties). The performance difference is statistically significant (p<0.001p < 0.001, sign test).

    On the subset of 51 document-tag pairs evaluated independently by all three judges, inter-annotator agreement reached Fleiss' κ=0.63\kappa = 0.63. Within this 51-pair subset:

    • L-LDA was selected as superior by at least one annotator in 33 of 51 cases, and unanimously selected by all three annotators in 24 of 51 cases.
    • SVM was selected by at least one annotator in 10 of 51 cases, and unanimously selected in only 2 of 51 cases.
  7. Knowl 7 — Multi-Label Classification Performance on Tagged del.icio.us Web Pages

    empirical result

    Multi-label document classification was evaluated on 4,000 web pages crawled from del.icio.us labeled across 20 medium-to-high frequency tags. The corpus contains an average of 781 non-stop words and 4 tags per document, with over 89% of documents possessing multiple labels.

    Models were trained on 80% of the documents and evaluated on the remaining 20% across 10 random train/test splits. L-LDA parameters (classification threshold and Dirichlet scaling constants) were tuned via validation, and one-versus-rest SVMs used term frequency features and cost parameter C=10.0C = 10.0 tuned via 4-fold cross-validation.

    Model Macro-F1 (%) Micro-F1 (%)
    Labeled LDA 39.85 ( 0.989) 52.12 ( 0.434)
    SVM 39.00 ( 0.423) 39.33 ( 0.574)

    Labeled LDA achieves a comparable Macro-F1 score to one-versus-rest SVMs (39.85%39.85\% vs 39.00%39.00\%) while outperforming SVMs on Micro-F1 by 12.79%12.79\% (52.12%52.12\% vs 39.33%39.33\%). Both improvements are statistically significant under a two-tailed paired t-test at the 95% confidence level.

  8. Knowl 8 — Multi-Label Text Classification Performance on Yahoo Directory Datasets

    data/table

    Multi-label classification performance was evaluated across 8 top-level categories from the Yahoo directory, where each document is assigned to one or more subcategories within the domain. Performance represents the mean (and standard deviation) over 10 independent runs.

    Dataset % Macro-F1 % Micro-F1
    L-LDA SVM L-LDA SVM
    Arts 30.70 (1.62) 23.23 (0.67) 39.81 (1.85) 48.42 (0.45)
    Business 30.81 (0.75) 22.82 (1.60) 67.00 (1.29) 72.15 (0.62)
    Computers 27.55 (1.98) 18.29 (1.53) 48.95 (0.76) 61.97 (0.54)
    Education 33.78 (1.70) 36.03 (1.30) 41.19 (1.48) 59.45 (0.56)
    Entertainment 39.42 (1.38) 43.22 (0.49) 47.71 (0.61) 62.89 (0.50)
    Health 45.36 (2.00) 47.86 (1.72) 58.13 (0.43) 72.21 (0.26)
    Recreation 37.63 (1.00) 33.77 (1.17) 43.71 (0.31) 59.15 (0.71)
    Society 27.32 (1.24) 23.89 (0.74) 42.98 (0.28) 52.29 (0.67)

    On the full Yahoo dataset—where only 33% of documents are multiply labeled—L-LDA outperforms one-versus-rest SVMs on Macro-F1 on 5 out of 8 categories (Arts, Business, Computers, Recreation, Society), but SVMs achieve higher Micro-F1 across all 8 datasets.

    When evaluated specifically on the subset of Yahoo documents having more than one label in training, L-LDA beats SVMs on Macro-F1 on 4 datasets (1 SVM win, 3 ties) and on Micro-F1 on 4 datasets (with 4 ties), reflecting L-LDA's credit assignment advantage on multi-labeled documents.

  9. Knowl 9 — Topic Specialization and Label Alignment in Labeled LDA versus Unsupervised LDA

    empirical result

    When trained on the 4,000-document del.icio.us corpus containing 20 distinct tags, standard unsupervised Latent Dirichlet Allocation (with K=20K = 20) and Labeled LDA exhibit contrasting topic allocations:

    • Unsupervised LDA: Allocates multiple redundant topic vectors to high-frequency semantic areas in the corpus (learning multiple distinct topics mapped by cosine similarity to web, culture, computer, reference, and politics), while failing to discover topics corresponding to lower-frequency tags (discovering zero topics for books, english, science, history, grammar, java, and philosophy).
    • Labeled LDA: Explicitly constrains each of the 20 topics to correspond directly to one user tag in a one-to-one mapping, ensuring that low-frequency tags maintain dedicated word multinomial distributions βk\beta_k and eliminating the need for post-hoc topic labeling.

Coverage note — No substantial contributed material was omitted.

References

  1. 1.D. M. Blei and J. Lafferty. 2006. Correlated Topic Models. NIPS, 18:147.
  2. 2.D. Blei and J McAuliffe. 2007. Supervised Topic Models. In NIPS, volume 21.
  3. 3.D. M. Blei, A.Y. Ng, and M.I. Jordan. 2003. Latent Dirichlet allocation. JMLR.
  4. 4.T. L. Griffiths and M. Steyvers. 2004. Finding scientific topics. PNAS, 1:5228–35.
  5. 5.P. Heymann, G. Koutrika, and H. Garcia-Molina. 2008. Can social bookmarking improve web search. In WSDM.
  6. 6.S. Ji, L. Tang, S. Yu, and J. Ye. 2008. Extracting shared subspace for multi-label classification. In KDD, pages 381–389, New York, NY, USA. ACM.
  7. 7.H. Kazawa, H. Taira T. Izumitani, and E. Maeda. 2004. Maximal margin labeling for multi-topic text categorization. In NIPS.
  8. 8.S. Lacoste-Julien, F. Sha, and M. I. Jordan. 2008. DiscLDA: Discriminative learning for dimensionality reduction and classification. In NIPS, volume 22.
  9. 9.D. D. Lewis, Y. Yang, T. G. Rose, G. Dietterich, F. Li, and F. Li. 2004. RCV1: A new benchmark collection for text categorization research. JMLR, 5:361–397.
  10. 10.Wei Li and Andrew McCallum. 2006. Pachinko allocation: Dag-structured mixture models of topic correlations. In International conference on Machine learning, pages 577–584.
  11. 11.A. McCallum and K. Nigam. 1998. A comparison of event models for naive bayes text classification. In AAAI-98 workshop on learning for text categorization, volume 7.
  12. 12.Q. Mei, X. Shen, and C Zhai. 2007. Automatic labeling of multinomial topic models. In KDD.
  13. 13.D. Ramage, P. Heymann, C. D. Manning, and H. Garcia-Molina. 2009. Clustering the tagged web. In WSDM.
  14. 14.N. Ueda and K. Saito. 2003. Parametric mixture models for multi-labeled text includes models that can be seen to fit within a dimensionality reduction framework. In NIPS.

Citation

MLA
Ramage, D., et al. “Labeled LDA”. Proceedings of the 2009 Conference on Empirical Methods in Natural Language Processing Volume 1 - EMNLP '09, vol. 1, 2009, p. 248, https://doi.org/10.3115/1699510.1699543.
APA
Ramage, D., Hall, D., Nallapati, R., & Manning, C. D. (2009). Labeled LDA. Proceedings of the 2009 Conference on Empirical Methods in Natural Language Processing Volume 1 - EMNLP '09, 1, 248. https://doi.org/10.3115/1699510.1699543
Chicago
Ramage, D., D. Hall, R. Nallapati, and C. D. Manning. 2009. “Labeled LDA”. Proceedings of the 2009 Conference on Empirical Methods in Natural Language Processing Volume 1 - EMNLP '09 1: 248. https://doi.org/10.3115/1699510.1699543.
Harvard
Ramage, D. et al. (2009) “Labeled LDA”, Proceedings of the 2009 Conference on Empirical Methods in Natural Language Processing Volume 1 - EMNLP '09. Association for Computational Linguistics, p. 248. Available at: https://doi.org/10.3115/1699510.1699543.
Vancouver
1. Ramage D, Hall D, Nallapati R, Manning CD (2009) Labeled LDA. In: Proceedings of the 2009 Conference on Empirical Methods in Natural Language Processing Volume 1 - EMNLP '09. Association for Computational Linguistics, p 248

BibTeX

@inproceedings{Ramage_2009, series={EMNLP ’09}, title={Labeled LDA: a supervised topic model for credit attribution in multi-labeled corpora}, volume={1}, url={http://dx.doi.org/10.3115/1699510.1699543}, DOI={10.3115/1699510.1699543}, booktitle={Proceedings of the 2009 Conference on Empirical Methods in Natural Language Processing Volume 1 - EMNLP ’09}, publisher={Association for Computational Linguistics}, author={Ramage, Daniel and Hall, David and Nallapati, Ramesh and Manning, Christopher D.}, year={2009}, pages={248}, collection={EMNLP ’09} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by-nc-sa/4.0/