keyword
naive Bayesian classifier
A naive Bayesian classifier is a supervised machine learning algorithm that predicts the category of a given data sample using probabilities derived from Bayes theorem. It operates under the simplified assumption that all input features are conditionally independent of one another given the target class. To classify an instance, the algorithm calculates the posterior probability for each possible class based on the observed feature values and assigns the instance to the class with the highest probability. Despite the assumption of feature independence rarely holding true in real-world scenarios, the classifier is computationally efficient, requires relatively little training data, scales well to high-dimensional feature spaces, and is widely applied in tasks such as text categorization, spam filtering, and sentiment analysis.
2 items

Learning and Revising User Profiles: The Identification of Interesting Web Sites
MICHAEL PAZZANI, DANIEL BILLSUS
Why you should read this
Presents Syskill & Webert, an intelligent agent that uses naive Bayesian classification combined with user background knowledge and lexical feature selection to accurately predict and identify web pages matching a user's long-term interests.
We discuss algorithms for learning and revising user profiles that can determine which World Wide Web sites on a given topic would be interesting to a user. We describe the use of a naive Bayesian classifier for this task, and demonstrate that it can incrementally learn profiles from user feedback on the interestingness of Web sites. Furthermore, the Bayesian classifier may easily be extended to revise user provided profiles. In an experimental evaluation we compare the Bayesian classifier to computationally more intensive alternatives, and show that it performs at least as well as these approaches throughout a range of different domains. In addition, we empirically analyze the effects of providing the classifier with background knowledge in form of user defined profiles and examine the use of lexical knowledge for feature selection. We find that both approaches can substantially increase the prediction accuracy.
Added
2026-09-25

Statistical Comparisons of Classifiers over Multiple Data Sets
Janez Demšar
Why you should read this
Establishes foundational guidelines for comparing machine learning algorithms across multiple datasets using non-parametric statistical tests and introduces critical difference diagrams for visual performance analysis.
While methods for comparing two learning algorithms on a single data set have been scrutinized for quite some time already, the issue of statistical tests for comparisons of more algorithms on multiple data sets, which is even more essential to typical machine learning studies, has been all but ignored. This article reviews the current practice and then theoretically and empirically examines several suitable tests. Based on that, we recommend a set of simple, yet safe and robust non-parametric tests for statistical comparisons of classifiers: the Wilcoxon signed ranks test for comparison of two classifiers and the Friedman test with the corresponding post-hoc tests for comparison of more classifiers over multiple data sets. Results of the latter can also be neatly presented with the newly introduced CD (critical difference) diagrams.
Added
2026-09-07

