keyword
text categorization
Text categorization, also known as text classification, is the automated process in natural language processing and machine learning of assigning predefined labels or categories to unstructured text documents based on their content. By analyzing linguistic features, statistical word patterns, or semantic representations, algorithms map textual data—such as articles, emails, web pages, or customer reviews—into one or more designated target classes. The process can be structured as binary, multi-class, or multi-label classification and serves as a foundational technique for applications including topic identification, spam filtering, sentiment analysis, and document organization.
16 items

Learning and Revising User Profiles: The Identification of Interesting Web Sites
MICHAEL PAZZANI, DANIEL BILLSUS
Why you should read this
Presents Syskill & Webert, an intelligent agent that uses naive Bayesian classification combined with user background knowledge and lexical feature selection to accurately predict and identify web pages matching a user's long-term interests.
We discuss algorithms for learning and revising user profiles that can determine which World Wide Web sites on a given topic would be interesting to a user. We describe the use of a naive Bayesian classifier for this task, and demonstrate that it can incrementally learn profiles from user feedback on the interestingness of Web sites. Furthermore, the Bayesian classifier may easily be extended to revise user provided profiles. In an experimental evaluation we compare the Bayesian classifier to computationally more intensive alternatives, and show that it performs at least as well as these approaches throughout a range of different domains. In addition, we empirically analyze the effects of providing the classifier with background knowledge in form of user defined profiles and examine the use of lexical knowledge for feature selection. We find that both approaches can substantially increase the prediction accuracy.
Added
2026-09-25

Heterogeneous Uncertainty Sampling for Supervised Learning
David D. Lewis, Jason Catlett
Why you should read this
Demonstrates that using a fast probabilistic classifier to actively select training instances for a more complex C4.5 rule induction model achieves lower error rates on text categorization tasks than random sampling sets ten times larger.
Uncertainty sampling methods iteratively request class labels for training instances whose classes are uncertain despite the previous labeled instances. These methods can greatly reduce the number of instances that an expert need label. One problem with this approach is that the classifier best suited for an application may be too expensive to train or use during the selection of instances. We test the use of one classifier (a highly efficient probabilistic one) to select examples for training another (the C4.5 rule induction program). Despite being chosen by this heterogeneous approach, the uncertainty samples yielded classifiers with lower error rates than random samples ten times larger.
Added
2026-09-25

Finding Deceptive Opinion Spam by Any Stretch of the Imagination
Myle Ott, Yejin Choi, Claire Cardie, Jeffrey T. Hancock
Why you should read this
Develops a computational approach to detect deceptive online reviews with nearly 90% accuracy, establishing a key linguistic link between fictitious consumer opinions and imaginative writing.
Consumers increasingly rate, review and research products online. Consequently, websites containing consumer reviews are becoming targets of opinion spam. While recent work has focused primarily on manually identifiable instances of opinion spam, in this work we study deceptive opinion spam---fictitious opinions that have been deliberately written to sound authentic. Integrating work from psychology and computational linguistics, we develop and compare three approaches to detecting deceptive opinion spam, and ultimately develop a classifier that is nearly 90% accurate on our gold-standard opinion spam dataset. Based on feature analysis of our learned models, we additionally make several theoretical contributions, including revealing a relationship between deceptive opinions and imaginative writing.
Added
2026-09-25

A kernel method for multi-labelled classification
A. Elisseeff, J. Weston
Why you should read this
Presents a large-margin kernel ranking framework for multi-label classification that directly minimizes ranking loss while capturing correlations between labels better than standard binary decomposition methods.
This article presents a Support Vector Machine (SVM) like learning system to handle multi-label problems. Such problems are usually decomposed into many two-class problems but the expressive power of such a system can be weak [5, 7]. We explore a new direct approach. It is based on a large margin ranking system that shares a lot of common properties with SVMs. We tested it on a Yeast gene functional classification problem with positive results.
Added
2026-09-25

One-Class SVMs for Document Classification
Larry M. Manevitz, Malik Yousef
Why you should read this
Evaluates one-class support vector machines alongside alternative single-class algorithms on the Reuters benchmark, revealing that while one-class SVMs achieve competitive document classification accuracy, their extreme sensitivity to feature representations and kernel choices makes compression neural networks a more reliable alternative for positive-only training.
We implemented versions of the SVM appropriate for one-class classification in the context of information retrieval. The experiments were conducted on the standard Reuters data set. For the SVM implementation we used both a version of Schölkopf et al. and a somewhat different version of one-class SVM based on identifying "outlier" data as representative of the second-class. We report on experiments with different kernels for both of these implementations and with different representations of the data, including binary vectors, tf-idf representation and a modification called "Hadamard" representation. Then we compared it with one-class versions of the algorithms prototype (Rocchio), nearest neighbor, naive Bayes, and finally a natural one-class neural network classification method based on "bottleneck" compression generated filters. The SVM approach as represented by Schölkopf was superior to all the methods except the neural network one, where it was, although occasionally worse, essentially comparable. However, the SVM methods turned out to be quite sensitive to the choice of representation and kernel in ways which are not well understood; therefore, for the time being leaving the neural network approach as the most robust.
Added
2026-09-25


Text Classification Algorithms: A Survey
Kamran Kowsari, Kiana Jafari Meimandi, Mojtaba Heidarysafa, Sanjana Mendu, Laura E. Barnes, Donald E. Brown
Why you should read this
Synthesizes the complete text classification pipeline by systematically comparing feature extraction techniques, dimensionality reduction approaches, machine learning algorithms, and evaluation metrics alongside their practical limitations in real-world applications.
In recent years, there has been an exponential growth in the number of complex documents and texts that require a deeper understanding of machine learning methods to be able to accurately classify texts in many applications. Many machine learning approaches have achieved surpassing results in natural language processing. The success of these learning algorithms relies on their capacity to understand complex models and non-linear relationships within data. However, finding suitable structures, architectures, and techniques for text classification is a challenge for researchers. In this paper, a brief overview of text classification algorithms is discussed. This overview covers different text feature extractions, dimensionality reduction methods, existing algorithms and techniques, and evaluations methods. Finally, the limitations of each technique and their application in the real-world problem are discussed.
Added
2026-09-24

A Brief Introduction to Boosting
R. Schapire
Why you should read this
Explains the theoretical foundations and mechanics of AdaBoost, providing clear proofs of exponential training error reduction alongside a margin-based explanation for why boosting resists overfitting even after achieving zero training error.
Boosting is a general method for improving the accuracy of any given learning algorithm. This short paper introduces the boosting algorithm AdaBoost, and explains the underlying theory of boosting, including an explanation of why boosting often does not suffer from overfitting. Some examples of recent applications of boosting are also described.
Added
2026-09-24

Support Vector Machines for Multiple-Instance Learning
S. Andrews, Ioannis Tsochantaridis, Thomas Hofmann
Why you should read this
Extends Support Vector Machines to multiple-instance learning by formulating pattern-level and bag-level maximum margin optimizations as mixed integer quadratic programs, enabling kernel-based classification on weakly annotated data across drug design, image indexing, and text categorization tasks.
This paper presents two new formulations of multiple-instance learning as a maximum margin problem. The proposed extensions of the Support Vector Machine (SVM) learning approach lead to mixed integer quadratic programs that can be solved heuristically. Our generalization of SVMs makes a state-of-the-art classification technique, including non-linear classification via kernels, available to an area that up to now has been largely dominated by special purpose methods. We present experimental results on a pharmaceutical data set and on applications in automated image indexing and document categorization.
Added
2026-09-24

A Bayesian Approach to Filtering Junk E-Mail
M. Sahami, S. Dumais, D. Heckerman, E. Horvitz
Why you should read this
Presents a decision-theoretic Bayesian approach that combines text classification with domain-specific email features and asymmetric misclassification costs to filter spam accurately in real-world scenarios.
In addressing the growing problem of junk E-mail on the Internet, we examine methods for the automated construction of filters to eliminate such unwanted messages from a user's mail stream. By casting this problem in a decision theoretic framework, we are able to make use of probabilistic learning methods in conjunction with a notion of differential misclassification cost to produce filters which are especially appropriate for the nuances of this task. While this may appear, at first, to be a straight-forward text classification problem, we show that by considering domain-specific features of this problem in addition to the raw text of E-mail messages, we can produce much more accurate filters. Finally, we show the efficacy of such filters in a real world usage scenario, arguing that this technology is mature enough for deployment.
Added
2026-09-24

A Probabilistic Analysis of the Rocchio Algorithm with TFIDF for Text Categorization
T. Joachims
Why you should read this
Presents a theoretical probabilistic foundation for the heuristic Rocchio and TFIDF classifier, introducing an optimized probabilistic variant, PrTFIDF, that improves text categorization accuracy without adding computational complexity.
The Rocchio relevance feedback algorithm is one of the most popular and widely applied learning methods from information retrieval. Here, a probabilistic analysis of this algorithm is presented in a text categorization framework. The analysis gives theoretical insight into the heuristics used in the Rocchio algorithm, particularly the word weighting scheme and the similarity metric. It also suggests improvements which lead to a probabilistic variant of the Rocchio classifier. The Rocchio classifier, its probabilistic variant, and a naive Bayes classifier are compared on six text categorization tasks. The results show that the probabilistic algorithms are preferable to the heuristic Rocchio classifier not only because they are more well-founded, but also because they achieve better performance.
Added
2026-09-24

Computing Semantic Relatedness Using Wikipedia-based Explicit Semantic Analysis
Evgeniy Gabrilovich, Shaul Markovitch
Why you should read this
Shows how Explicit Semantic Analysis turns Wikipedia concepts into interpretable text vectors and substantially improves agreement with human semantic-relatedness judgments.
Computing semantic relatedness of natural language texts requires access to vast amounts of common-sense and domain-specific world knowledge. We propose Explicit Semantic Analysis (ESA), a novel method that represents the meaning of texts in a high-dimensional space of concepts derived from Wikipedia. We use machine learning techniques to explicitly represent the meaning of any text as a weighted vector of Wikipedia-based concepts. Assessing the relatedness of texts in this space amounts to comparing the corresponding vectors using conventional metrics (e.g., cosine). Compared with the previous state of the art, using ESA results in substantial improvements in correlation of computed relatedness scores with human judgments: from r = 0.56 to 0.75 for individual words and from r = 0.60 to 0.72 for texts. Importantly, due to the use of natural concepts, the ESA model is easy to explain to human users.
Added
2026-09-14

Classifier chains for multi-label classification
Jesse Read, Bernhard Pfahringer, Geoff Holmes, Eibe Frank
Why you should read this
Proposes classifier chains and ensemble extensions that model label correlations in multi-label classification while retaining the computational efficiency and linear scalability of binary relevance across large datasets.
The widely known binary relevance method for multi-label classification, which considers each label as an independent binary problem, has often been overlooked in the literature due to the perceived inadequacy of not directly modelling label correlations. Most current methods invest considerable complexity to model interdependencies between labels. This paper shows that binary relevance-based methods have much to offer, and that high predictive performance can be obtained without impeding scalability to large datasets. We exemplify this with a novel classifier chains method that can model label correlations while maintaining acceptable computational complexity. We extend this approach further in an ensemble framework. An extensive empirical evaluation covers a broad range of multi-label datasets with a variety of evaluation metrics. The results illustrate the competitiveness of the chaining method against related and state-of-the-art methods, both in terms of predictive performance and time complexity.
Added
2026-09-14

A sequential algorithm for training text classifiers
David D. Lewis, William A. Gale
Why you should read this
Introduces uncertainty sampling, a sequential learning method that reduces the volume of manually labeled training data needed for accurate text classification by up to 500-fold.
The ability to cheaply train text classifiers is critical to their use in information retrieval, content analysis, natural language processing, and other tasks involving data which is partly or fully textual. An algorithm for sequential sampling during machine learning of statistical classifiers was developed and tested on a newswire text categorization task. This method, which we call uncertainty sampling, reduced by as much as 500-fold the amount of training data that would have to be manually classified to achieve a given level of effectiveness.
Added
2026-09-14

Fast Effective Rule Induction
William W. Cohen
Why you should read this
Introduces RIPPERk, a fast rule learning algorithm that matches the classification accuracy of C4.5rules on benchmark problems while scaling nearly linearly to efficiently handle hundreds of thousands of noisy training examples.
Many existing rule learning systems are computationally expensive on large noisy datasets. In this paper we evaluate the recently-proposed rule learning algorithm IREP on a large and diverse collection of benchmark problems. We show that while IREP is extremely efficient, it frequently gives error rates higher than those of C4.5 and C4.5rules. We then propose a number of modifications resulting in an algorithm RIPPERk that is very competitive with C4.5rules with respect to error rates, but much more efficient on large samples. RIPPERk obtains error rates lower than or equivalent to C4.5rules on 22 of 37 benchmark problems, scales nearly linearly with the number of training examples, and can efficiently process noisy datasets containing hundreds of thousands of examples.
Added
2026-09-11

A comparison of event models for naive bayes text classification
Andrew McCallum, Kamal Nigam
Why you should read this
Demonstrates why the multinomial event model consistently outperforms the multivariate Bernoulli model in naive Bayes text classification, achieving an average 27% reduction in error across standard text corpora.
Recent approaches to text classification have used two different first-order probabilistic models for classification, both of which make the naive Bayes assumption. Some use a multi-variate Bernoulli model, that is, a Bayesian Network with no dependencies between words and binary word features (e.g. Larkey and Croft 1996; Koller and Sahami 1997). Others use a multinomial model, that is, a uni-gram language model with integer word counts (e.g. Lewis and Gale 1994; Mitchell 1997). This paper aims to clarify the confusion by describing the differences and details of these two models, and by empirically comparing their classification performance on five text corpora. We find that the multi-variate Bernoulli performs well with small vocabulary sizes, but that the multinomial performs usually performs even better at larger vocabulary sizes—providing on average a 27% reduction in error over the multi-variate Bernoulli model at any vocabulary size.
Added
2026-09-10

Thumbs up? Sentiment Classification using Machine Learning Techniques
Bo Pang, Lillian Lee, Shivakumar Vaithyanathan
Why you should read this
Establishes machine learning as a viable framework for sentiment classification by demonstrating that standard classifiers outperform human baselines on movie reviews while identifying key challenges that make sentiment harder than topic-based categorization.
We consider the problem of classifying documents not by topic, but by overall sentiment, e.g., determining whether a review is positive or negative. Using movie reviews as data, we find that standard machine learning techniques definitively outperform human-produced baselines. However, the three machine learning methods we employed (Naive Bayes, maximum entropy classification, and support vector machines) do not perform as well on sentiment classification as on traditional topic-based categorization. We conclude by examining factors that make the sentiment classification problem more challenging.
Added
2026-09-07
