Learning and Revising User Profiles: The Identification of Interesting Web Sites
MICHAEL PAZZANIDANIEL BILLSUS
Presents Syskill & Webert, an intelligent agent that uses naive Bayesian classification combined with user background knowledge and lexical feature selection to accurately predict and identify web pages matching a user's long-term interests.
The rapid growth of the World Wide Web presents significant challenges for users attempting to locate information tailored to recurring, long-term interests. Standard search engines return voluminous results, but users frequently find only a small fraction relevant to their specific preferences. The article evaluates automated methods for learning and revising individualized user profiles to predict whether unseen web pages on a given topic will be interesting or uninteresting.
The authors evaluate an intelligent agent named Syskill & Webert across nine topic domains using feedback from four users. Web pages are represented as Boolean feature vectors based on the presence or absence of informative words. The evaluation assesses a naive Bayesian classifier against computationally intensive machine learning alternatives—including decision trees, neural networks, nearest neighbor approaches, and information retrieval methods. It also tests techniques for incorporating initial user-provided keywords and lexical background knowledge from WordNet to improve classification accuracy when training data is scarce.
The analysis reveals several key findings. First, the naive Bayesian classifier performs at least as well as more complex and resource-intensive algorithms, such as multi-layer neural networks and Rocchio's algorithm, while offering substantial computational speed and linear scaling. Second, incorporating an initial user profile and incrementally revising it using Bayesian conjugate priors substantially improves accuracy on small training sets; for example, accuracy on a sample topic rose from roughly 70% to about 85%. Third, experiments demonstrate that the primary benefit of initial user profiles stems from identifying a concise set of 10 to 15 relevant keyword features rather than precise probability estimates. Fourth, filtering informative words using the WordNet lexical database to remove terms unrelated to the topic noticeably boosts accuracy when few rated examples are available.
These results indicate that complex nonlinear algorithms are unnecessary for effective web page filtering, as linear models that aggregate evidence across multiple features perform reliably. Furthermore, the bottleneck in early classification accuracy is the selection of irrelevant features rather than algorithm sophistication. Systems can reduce training burdens and improve user adoption by eliciting small sets of initial keywords or applying lexical filters to prune unrelated words.
Organizations developing information filtering systems should implement lightweight Bayesian classifiers augmented by initial user-defined keywords rather than deploying complex, high-overhead models. System developers should also incorporate lexical databases to pre-filter candidate features automatically. Future initiatives should focus on improving feature representation and language modeling rather than refining core classification algorithms, as linguistic enhancements deliver greater performance gains.
The findings are subject to certain limitations, including relatively small evaluation datasets consisting of four users and nine topics, as well as degraded performance in noisy domains where text alone is insufficient to capture user interest, such as music evaluation. Confidence in the relative performance of the Bayesian classifier and the feature selection techniques is high across text-heavy domains, but caution is warranted when generalizing these methods to non-textual or highly subjective domains.
- Paper: On the Optimality of the Simple Bayesian Classifier under Zero-One Loss, Pedro Domingos et al. (1997). This paper establishes the theoretical foundations for why simple naive Bayesian classifiers perform well even when feature independence assumptions are violated, providing essential justification for the source's profiling model.
- Paper: An Analysis of Bayesian Classifiers, Pat Langley et al. (1992). It provides foundational mathematical and empirical analysis of simple Bayesian classifiers against decision trees, framing the baseline learning behaviors leveraged in user profile induction.
- Paper: Letizia: An Agent That Assists Web Browsing, Henry Lieberman (1995). This work introduces the foundational paradigm of autonomous agents tracking browsing actions to infer user interests without manual rule configuration.
- Paper: A sequential algorithm for training text classifiers, David D. Lewis et al. (1994). It establishes sequential and active learning mechanisms for probabilistic text classifiers from iterative user feedback, directly informing how user profiles are revised incrementally.
- Paper: Learning Bayesian networks: The combination of knowledge and statistical data, David Heckerman et al. (1994). It details the core Bayesian methodology for integrating prior background knowledge and expert rules directly with empirical statistical data during model learning.
- Paper: Instance-Based Learning Algorithms, David W. Aha et al. (1991). This paper establishes the standard instance-based learning framework used as a primary comparative alternative against Bayesian classifiers in document filtering tasks.
- Paper: A comparison of event models for naive bayes text classification, Andrew McCallum et al. (1998). This study refines naive Bayes classification for web documents by comparing multinomial and multi-variate Bernoulli event models across diverse text collections.
- Paper: Text Classification from Labeled and Unlabeled Documents using EM, Kamal Nigam et al. (2000). This paper extends naive Bayes web text classification by incorporating large pools of unlabeled documents via Expectation-Maximization to alleviate labeled data scarcity.
- Paper: A Bayesian Approach to Filtering Junk E-Mail, M. Sahami et al. (1998). This work applies naive Bayesian classification and feature selection specifically to personal message filtering while managing asymmetric misclassification costs.
- Paper: Empirical Analysis of Predictive Algorithms for Collaborative Filtering, John S. Breese et al. (1998). It expands upon personalized user preference modeling by systematically comparing model-based Bayesian approaches and memory-based collaborative filtering across web browsing datasets.
- Paper: Optimizing search engines using clickthrough data, Thorsten Joachims (2002). This paper advances the automated learning of user preferences from web interaction data by formalizing implicit clickthrough feedback into a ranking optimization framework.
- Paper: Web mining research: a survey, Raymond Kosala et al. (2000). This survey synthesizes web personalization, content categorization, and usage profiling techniques into a comprehensive taxonomy of web mining research.
- Paper: One-Class SVMs for Document Classification, Larry M. Manevitz et al. (2002). This work addresses document classification in personalization scenarios where only positive browsing feedback is available, comparing one-class methods with naive Bayes.
- Paper: Machine learning in automated text categorization, Fabrizio Sebastiani (2001). This comprehensive survey contextualizes machine learning approaches, including naive Bayes and feature selection techniques, for automated document and web page categorization.
