Built independently by an author, for readers. Read the story and support ChapterPal

keyword

multi-labeled corpora

Multi-labeled corpora are structured collections of texts or documents in which individual entries are annotated with multiple, nonexclusive labels, categories, or tags simultaneously, rather than being restricted to a single mutually exclusive classification. In machine learning and natural language processing, these corpora reflect complex data environments—such as tagged web pages, multi-topic articles, and indexed legal or academic documents—where a single text naturally spans several subjects, themes, or metadata fields. Because different labels may correspond to distinct sections or varying levels of specificity within a document, multi-labeled corpora are essential for training and evaluating algorithms designed for multi-label text classification, fine-grained semantic analysis, and supervised topic modeling.

1 item

Labeled LDA: A supervised topic model for credit attribution in multi-labeled corpora

Labeled LDA: A supervised topic model for credit attribution in multi-labeled corpora

Daniel Ramage, David Leo Wright Hall, Ramesh Nallapati, Christopher D. Manning

OrganizationsStanford University

Why you should read this

Introduces Labeled LDA, a supervised extension of Latent Dirichlet Allocation that constrains latent topics to observed document tags to directly learn word-tag correspondences and extract tag-specific snippets.

A significant portion of the world's text is tagged by readers on social bookmarking websites. Credit attribution is an inherent problem in these corpora because most pages have multiple tags, but the tags do not always apply with equal specificity across the whole document. Solving the credit attribution problem requires associating each word in a document with the most appropriate tags and vice versa. This paper introduces Labeled LDA, a topic model that constrains Latent Dirichlet Allocation by defining a one-to-one correspondence between LDA's latent topics and user tags. This allows Labeled LDA to directly learn word-tag correspondences. We demonstrate Labeled LDA's improved expressiveness over traditional LDA with visualizations of a corpus of tagged web pages from del.icio.us. Labeled LDA outperforms SVMs by more than 3 to 1 when extracting tag-specific document snippets. As a multi-label text classifier, our model is competitive with a discriminative baseline on a variety of datasets.

Added

2026-09-25