Learning and Revising User Profiles: The Identification of Interesting Web Sites

MICHAEL PAZZANIDANIEL BILLSUS

article1997Machine-mediated learning1,345 citations

Presents Syskill & Webert, an intelligent agent that uses naive Bayesian classification combined with user background knowledge and lexical feature selection to accurately predict and identify web pages matching a user's long-term interests.

Listen

The rapid growth of the World Wide Web presents significant challenges for users attempting to locate information tailored to recurring, long-term interests. Standard search engines return voluminous results, but users frequently find only a small fraction relevant to their specific preferences. The article evaluates automated methods for learning and revising individualized user profiles to predict whether unseen web pages on a given topic will be interesting or uninteresting.

The authors evaluate an intelligent agent named Syskill & Webert across nine topic domains using feedback from four users. Web pages are represented as Boolean feature vectors based on the presence or absence of informative words. The evaluation assesses a naive Bayesian classifier against computationally intensive machine learning alternatives—including decision trees, neural networks, nearest neighbor approaches, and information retrieval methods. It also tests techniques for incorporating initial user-provided keywords and lexical background knowledge from WordNet to improve classification accuracy when training data is scarce.

The analysis reveals several key findings. First, the naive Bayesian classifier performs at least as well as more complex and resource-intensive algorithms, such as multi-layer neural networks and Rocchio's algorithm, while offering substantial computational speed and linear scaling. Second, incorporating an initial user profile and incrementally revising it using Bayesian conjugate priors substantially improves accuracy on small training sets; for example, accuracy on a sample topic rose from roughly 70% to about 85%. Third, experiments demonstrate that the primary benefit of initial user profiles stems from identifying a concise set of 10 to 15 relevant keyword features rather than precise probability estimates. Fourth, filtering informative words using the WordNet lexical database to remove terms unrelated to the topic noticeably boosts accuracy when few rated examples are available.

These results indicate that complex nonlinear algorithms are unnecessary for effective web page filtering, as linear models that aggregate evidence across multiple features perform reliably. Furthermore, the bottleneck in early classification accuracy is the selection of irrelevant features rather than algorithm sophistication. Systems can reduce training burdens and improve user adoption by eliciting small sets of initial keywords or applying lexical filters to prune unrelated words.

Organizations developing information filtering systems should implement lightweight Bayesian classifiers augmented by initial user-defined keywords rather than deploying complex, high-overhead models. System developers should also incorporate lexical databases to pre-filter candidate features automatically. Future initiatives should focus on improving feature representation and language modeling rather than refining core classification algorithms, as linguistic enhancements deliver greater performance gains.

The findings are subject to certain limitations, including relatively small evaluation datasets consisting of four users and nine topics, as well as degraded performance in noisy domains where text alone is insufficient to capture user interest, such as music evaluation. Confidence in the relative performance of the Bayesian classifier and the feature selection techniques is high across text-heavy domains, but caution is warranted when generalizing these methods to non-textual or highly subjective domains.

  • Paper: On the Optimality of the Simple Bayesian Classifier under Zero-One Loss, Pedro Domingos et al. (1997). This paper establishes the theoretical foundations for why simple naive Bayesian classifiers perform well even when feature independence assumptions are violated, providing essential justification for the source's profiling model.
  • Paper: An Analysis of Bayesian Classifiers, Pat Langley et al. (1992). It provides foundational mathematical and empirical analysis of simple Bayesian classifiers against decision trees, framing the baseline learning behaviors leveraged in user profile induction.
  • Paper: Letizia: An Agent That Assists Web Browsing, Henry Lieberman (1995). This work introduces the foundational paradigm of autonomous agents tracking browsing actions to infer user interests without manual rule configuration.
  • Paper: A sequential algorithm for training text classifiers, David D. Lewis et al. (1994). It establishes sequential and active learning mechanisms for probabilistic text classifiers from iterative user feedback, directly informing how user profiles are revised incrementally.
  • Paper: Learning Bayesian networks: The combination of knowledge and statistical data, David Heckerman et al. (1994). It details the core Bayesian methodology for integrating prior background knowledge and expert rules directly with empirical statistical data during model learning.
  • Paper: Instance-Based Learning Algorithms, David W. Aha et al. (1991). This paper establishes the standard instance-based learning framework used as a primary comparative alternative against Bayesian classifiers in document filtering tasks.
  • Paper: A comparison of event models for naive bayes text classification, Andrew McCallum et al. (1998). This study refines naive Bayes classification for web documents by comparing multinomial and multi-variate Bernoulli event models across diverse text collections.
  • Paper: Text Classification from Labeled and Unlabeled Documents using EM, Kamal Nigam et al. (2000). This paper extends naive Bayes web text classification by incorporating large pools of unlabeled documents via Expectation-Maximization to alleviate labeled data scarcity.
  • Paper: A Bayesian Approach to Filtering Junk E-Mail, M. Sahami et al. (1998). This work applies naive Bayesian classification and feature selection specifically to personal message filtering while managing asymmetric misclassification costs.
  • Paper: Empirical Analysis of Predictive Algorithms for Collaborative Filtering, John S. Breese et al. (1998). It expands upon personalized user preference modeling by systematically comparing model-based Bayesian approaches and memory-based collaborative filtering across web browsing datasets.
  • Paper: Optimizing search engines using clickthrough data, Thorsten Joachims (2002). This paper advances the automated learning of user preferences from web interaction data by formalizing implicit clickthrough feedback into a ranking optimization framework.
  • Paper: Web mining research: a survey, Raymond Kosala et al. (2000). This survey synthesizes web personalization, content categorization, and usage profiling techniques into a comprehensive taxonomy of web mining research.
  • Paper: One-Class SVMs for Document Classification, Larry M. Manevitz et al. (2002). This work addresses document classification in personalization scenarios where only positive browsing feedback is available, comparing one-class methods with naive Bayes.
  • Paper: Machine learning in automated text categorization, Fabrizio Sebastiani (2001). This comprehensive survey contextualizes machine learning approaches, including naive Bayes and feature selection techniques, for automated document and web page categorization.
Cover for Learning and Revising User Profiles: The Identification of Interesting Web Sites

Abstract

We discuss algorithms for learning and revising user profiles that can determine which World Wide Web sites on a given topic would be interesting to a user. We describe the use of a naive Bayesian classifier for this task, and demonstrate that it can incrementally learn profiles from user feedback on the interestingness of Web sites. Furthermore, the Bayesian classifier may easily be extended to revise user provided profiles. In an experimental evaluation we compare the Bayesian classifier to computationally more intensive alternatives, and show that it performs at least as well as these approaches throughout a range of different domains. In addition, we empirically analyze the effects of providing the classifier with background knowledge in form of user defined profiles and examine the use of lexical knowledge for feature selection. We find that both approaches can substantially increase the prediction accuracy.

Table of Contents

  • 1. Introduction
  • 2. Syskill and Webert
  • 2.1. User interface
  • 2.2. Learning user profiles
  • 2.3. Initial experiments
  • 3. Experimental comparisons
  • 3.1. Nearest neighbor
  • 3.2. PEBLS
  • 3.3. Decision trees
  • 3.4. Rocchio's algorithm
  • 3.5. Neural nets
  • 3.6. Results
  • 4. Using predefined user profiles
  • 5. Using lexical knowledge
  • 6. Related work
  • 7. Future directions
  • 8. Conclusions
  • Notes
  • References

Knowls

  1. Knowl 1 — User Profile Revision Using Equivalent Sample Size Conjugate Priors

    model/method

    User profile revision combines prior human knowledge with empirical data from rated web pages within a naive Bayesian framework. A user specifies a set of indicator words for a topic and optionally provides conditional probabilities p0(wordi=present∣hot)p_0(\text{word}_i = \text{present} \mid \text{hot}) and p0(wordi=present∣cold)p_0(\text{word}_i = \text{present} \mid \text{cold}). If the user provides no probabilities, defaults of 0.70.7 for hot and 0.30.3 for cold are assigned. The complement probabilities are computed as p0(wordi=absent∣classj)=1−p0(wordi=present∣classj)p_0(\text{word}_i = \text{absent} \mid \text{class}_j) = 1 - p_0(\text{word}_i = \text{present} \mid \text{class}_j).

    To update these prior estimates as training pages are rated, conjugate priors are implemented via the equivalent sample size approach. The user's prior probability p0p_0 is weighted as equivalent to having observed N0N_0 pseudo-examples (by default N0=50N_0 = 50), corresponding to k0=p0⋅N0k_0 = p_0 \cdot N_0 occurrences. When NDN_D training pages of class Cj∈{hot,cold}C_j \in \{\text{hot}, \text{cold}\} are observed, of which kDk_D contain the word, the revised conditional probability is computed as:

    P(wordi=present∣Cj)=p0N0+kDN0+ND=k0+kDN0+NDP(\text{word}_i = \text{present} \mid C_j) = \frac{p_0 N_0 + k_D}{N_0 + N_D} = \frac{k_0 + k_D}{N_0 + N_D}

    When revising a profile, the feature vector includes all user-specified words supplemented with the top information-gain words from the training data up to a fixed total number of features (e.g., 96 words).

  2. Knowl 2 — Dominance of Keyword Selection Over Probability Calibration in User Profiles

    empirical result

    An empirical dissection comparing four naive Bayesian profile variants reveals that the predictive benefit of user-provided background knowledge stems almost entirely from selecting relevant keywords rather than from supplying accurate probability estimates:

    1. Data: Uses the top 96 information-gain features selected from training data alone, with probabilities estimated purely from data.
    2. Revision: Uses user-provided words supplemented by top information-gain features (total 96), updated using conjugate priors (N0=50N_0 = 50).
    3. Fixed: Uses only user-provided words and their fixed prior probabilities without updating.
    4. ProfileFeatures: Uses only the user-provided words (typically 10–15 features), but estimates all conditional probabilities purely from training data.

    Across multiple domains (Biomedical, Bands, and Goats), the Revision strategy achieved significantly higher accuracy than Data alone on small training sets (e.g., ~85% vs ~70% on Goats). However, the ProfileFeatures strategy achieved classification accuracy equal to or higher than the conjugate priors Revision strategy throughout all training set sizes, while using only 10–15 user-selected words. This demonstrates that human users are highly effective at identifying discriminative vocabulary for a domain even when unreliable at estimating numerical conditional probabilities.

  3. Knowl 3 — WordNet Semantic Link Filtering for Feature Selection in Small Samples

    model/method

    When learning from small web page sample sizes (e.g., 10–20 pages), information-gain feature selection frequently selects spurious words that correlate with the target class by chance (e.g., "took", "other", "however"). To eliminate non-topic-related words, candidate informative words selected by information gain are filtered through WordNet, a lexical database connecting approximately 30,000 English words.

    A word candidate is retained as a feature if and only if a semantic path exists between the word and the domain topic keyword (e.g., "goat") via any of the following lexical relationships in WordNet:

    • Hypernym
    • Antonym
    • Member-Holonym
    • Part-Holonym
    • Similar-to
    • Pertainnym
    • Derived-from

    Words lacking any such relational link to the topic are discarded from the feature set, in addition to standard stoplist filtering (~600 common English words and HTML tags).

  4. Knowl 4 — Classification Algorithm Comparison for Web Page Filtering

    data/table

    Seven machine learning classification algorithms were evaluated across nine web page recommendation domains using 40 paired trials. Each trial used a training set of 20 randomly sampled pages, selected the 128 most informative Boolean word features by expected information gain, and evaluated test accuracy on the remaining pages.

    Domain N Neigh ID3 Percept BackP PEBLS Bayes Rocchio
    Bio(Topic) 74.5 70.2 73.2 76.0+ 74.6 77.3+ 77.5+
    Goats 62.0 64.7 66.3+ 67.0+ 62.7 62.9 69.4+
    Bio(lycos) 80.0+ 75.9 80.2+ 80.9+ 79.9+ 78.2 78.1
    Mail 63.1 62.4 62.8 64.2 63.3 66.9+ 67.9+
    Movies 58.6 69.4+ 70.5+ 67.4 60.0 69.3+ 69.0+
    Protein 77.0+ 73.7 70.4 74.1 77.0+ 77.0+ 74.6
    Sheep 79.3 78.4 78.9 80.5+ 79.3 81.5+ 78.8
    Bands(s) 74.4+ 70.7 71.4 73.1+ 74.5+ 73.4+ 73.7+
    Bands(t) 75.0+ 68.6 69.6 73.9+ 75.0+ 74.6+ 74.5+

    A + indicates no statistically significant difference from the highest-accuracy algorithm on that domain (paired two-tailed tt-test at the α=0.05\alpha = 0.05 level). Linear classifiers that aggregate evidence across many features (Naive Bayes and Rocchio's TF-IDF prototype classifier) and multi-layer neural networks (Backpropagation with 12 hidden units) consistently achieve top performance. Decision tree induction (ID3) and standard 1-Nearest Neighbor perform significantly worse because ID3 tests as few features as possible rather than combining evidence across many words, while nearest neighbor is sensitive to high dimensionality.

  5. Knowl 5 — Feature Set Size Optimization for Naive Bayesian Web Classification

    data/table

    The classification accuracy of a naive Bayesian classifier varies with the number of informative word features selected via expected information gain. Experiments across six domains using 24 trials with 20 training examples per trial evaluated feature set sizes ranging from 16 to 400 words:

    Features Bands (Reading) Bands (Sound) Biomed (Topic) Biomed (Lycos) Movies Protein Average
    16 70.4 73.0 74.5 77.1 74.6 71.6 73.9
    32 72.0 71.6 73.6 79.2 76.8 73.8 74.6
    64 74.3 71.7 74.1 80.9 73.1 75.0 74.8
    96 74.3 75.1 76.9 78.7 72.5 78.8 75.5
    128 73.8 73.4 77.3 77.0 74.2 75.1 75.1
    200 74.3 73.3 77.1 77.4 71.4 75.0 74.7
    256 74.5 73.4 76.9 77.0 71.3 76.5 74.6
    400 74.8 73.3 75.5 76.9 69.2 70.6 73.9

    Intermediate feature counts yield peak performance, with 96 features achieving the highest average accuracy (75.5%). Selecting too few features (e.g., 16) discards important discriminating words, while selecting too many features (e.g., 400) degrades performance due to the inclusion of noisy, uninformative words that overfit the small training sample.

  6. Knowl 6 — Expected Information Gain for Binary Word Features

    equation

    Informative word features for web page classification are identified by computing the expected information gain E(W,S)E(W, S) for the presence or absence of each candidate word WW across the training set of web pages SS:

    E(W,S)=I(S)−[P(W=present)I(SW=present)+P(W=absent)I(SW=absent)]E(W, S) = I(S) - \left[ P(W = \text{present}) I(S_{W = \text{present}}) + P(W = \text{absent}) I(S_{W = \text{absent}}) \right]

    where the class entropy I(S)I(S) of a set of pages SS is defined over binary user preference classes c∈{hot,cold}c \in \{\text{hot}, \text{cold}\} as:

    I(S)=−∑c∈{hot,cold}p(Sc)log⁡2(p(Sc))I(S) = - \sum_{c \in \{\text{hot}, \text{cold}\}} p(S_c) \log_2(p(S_c))

    Here, ScS_c denotes the subset of pages in SS belonging to class cc, p(Sc)=∣Sc∣/∣S∣p(S_c) = |S_c| / |S| is the empirical proportion of class cc in SS, P(W=present)P(W = \text{present}) is the fraction of pages in SS that contain at least one occurrence of word WW, SW=presentS_{W = \text{present}} is the subset of pages containing WW, and SW=absentS_{W = \text{absent}} is the subset of pages from which WW is absent. Words appearing on a ~600-word stop list of common English words and HTML tags are excluded before ranking by E(W,S)E(W, S).

  7. Knowl 7 — Syskill and Webert Naive Bayes Document Representation and Classification

    model/method

    In the Syskill & Webert agent, an HTML document is parsed into tokens delimited by non-alphabetic characters, converted to uppercase, filtered against a ~600-word stoplist, and represented as a Boolean feature vector (A1,A2,…,An)(A_1, A_2, \dots, A_n) where Ak∈{present,absent}A_k \in \{\text{present}, \text{absent}\} indicates whether the kk-th informative word occurs at least once in the document.

    Assuming conditional independence of attribute values given class Ci∈{hot,cold}C_i \in \{\text{hot}, \text{cold}\}, the posterior class probability for a page instance jj is proportional to:

    P(Ci∣A1=V1j∧⋯∧An=Vnj)∝P(Ci)∏k=1nP(Ak=Vkj∣Ci)P(C_i \mid A_1 = V_{1j} \land \dots \land A_n = V_{nj}) \propto P(C_i) \prod_{k=1}^n P(A_k = V_{kj} \mid C_i)

    where P(Ci)P(C_i) is estimated by the fraction of training pages belonging to class CiC_i, and P(Ak=Vkj∣Ci)P(A_k = V_{kj} \mid C_i) is estimated from the frequency of feature value VkjV_{kj} among training examples of class CiC_i. An unrated page is assigned to the class with the higher probability, and the continuous value P(hot∣A1=V1j,…,An=Vnj)∈[0,1]P(\text{hot} \mid A_1 = V_{1j}, \dots, A_n = V_{nj}) \in [0, 1] is used to rank-order candidate hyperlinks for user exploration.

  8. Knowl 8 — Empirical Sample-Efficiency Improvement from WordNet Lexical Filtering

    empirical result

    Filtering candidate features using WordNet semantic links to the topic name significantly improves classification accuracy on small training sets. In 25 trials on the Goats domain, retaining only informative words semantically linked to "goat" yielded a large accuracy advantage over using all information-gain features when training on 10 to 30 examples:

    • At 10 training examples: WordNet-filtered features achieved ~68% accuracy versus ~57% for all informative features (~11 percentage point gain).
    • At 20 training examples: WordNet-filtered features achieved ~72% accuracy versus ~65% for all features.
    • At 50+ training examples: The performance difference diminished, with both approaches converging near ~74–76% accuracy.

    This confirms that lexical semantic filtering acts as an effective regularizer against spurious statistical associations when empirical training data is sparse.

Coverage note — None omitted; all core methods (Syskill & Webert naive Bayes, information gain feature selection, conjugate prior profile revision, WordNet lexical filtering), algorithm comparisons (Table 4), feature size experiments (Table 3), and empirical analyses are fully covered.

References

  1. 1.Armstrong, R., Freitag, D., Joachims, T., & Mitchell, T. (1995). WebWatcher: A learning apprentice for the World Wide Web. Working Notes of the AAAI Spring Symposium Series on Information Gathering from Distributed, Heterogeneous Environments (pp. 6–12). Palo Alto, CA.
  2. 2.Balabanovic, Shoham, & Yun. (1995). An adaptive agent for automated web browsing (Technical Report CS-TN-97-52). Stanford University, Palo Alto, CA.
  3. 3.Cost, S., & Salzberg, S. (1993). A weighted nearest neighbor algorithm for learning with symbolic features. Machine Learning, 10:57–78.
  4. 4.Croft, W.B., & Harper, D. (1979). Using probabilistic models of document retrieval without relevance. Journal of Documentation, 35:285–295.
  5. 5.Domingos, P., & Pazzani, M. (1996). Beyond independence: Conditions for the optimality of the Simple Bayesian Classifier. Proceedings of the Thirteenth International Conference on Machine Learning (pp. 105–112). Morgan Kaufmann, San Fransico, CA.
  6. 6.Duda, R., & Hart, P. (1973). Pattern Classification and Scene Analysis. John Wiley & Sons, New York.
  7. 7.Harman, D.K. (1994). Overview of the second Text Retrieval Conference (TREC-2). Proceedings of the Second Text Retrieval Conference (TREC-2), NIST Special Publication.
  8. 8.Heckerman, D. (1995). A Tutorial on Learning with Bayesian Networks (Technical Report MSR-TR-95-06). Microsoft Corporation.
  9. 9.Ittner, D., Lewis, D., & Ahn, D. (1995). Text categorization of low quality images. Symposium on Document Analysis and Information Retrieval (pp. 301–315). UNLV, Las Vegas, NV, ISRI.
  10. 10.John, G., Kohavi, R., & Pfleger, K. (1994). Irrelevant features and the subset selection problem. Proceedings of the Eleventh International Conference on Machine Learning (pp. 121–138). New Brunswick, NJ.
  11. 11.Kittler, J. (1986). Feature selection and extraction. In Young, & Fu, (Eds.), Handbook of Pattern Recognition and Image Processing. Academic Press, New York.
  12. 12.Kononenko, I. (1990). Comparison of inductive and naive Bayesian learning approaches to automatic knowledge acquisition. In B. Wielinga (Ed.), Current Trends in Knowledge Acquisition. IOS Press, Amsterdam.
  13. 13.Lang, K. (1995). NewsWeeder: Learning to filter news. Proceedings of the Twelfth International Conference on Machine Learning (pp. 331–339). Lake Tahoe, CA.
  14. 14.Lashkari, Y. (1995). The WebHound Personalized Document Filtering System. http://rg.media.mit.edu/projects/webhound/
  15. 15.Lewis, D. (1992). Representation and learning in information retrieval. Doctoral dissertation, Department of Computer and Information Science, University of Massachusetts.
  16. 16.Lieberman, H. (1995). Letizia: An agent that assists web browsing. Proceedings of the International Joint Conference on Artificial Intelligence (pp. 924–929), Montreal, August 1995.
  17. 17.Maron, M. (1961). Automatic indexing: An experimental inquiry. Journal of the Association for Computing Machinery, 8:404–417.
  18. 18.Mauldin, M., & Leavitt, J. (1994). Web agent related research at the center for machine translation. Proceedings of the ACM Special Interest Group on Networked Information Discovery and Retrieval. The MITRE Corporation, McLean, Virgiana.
  19. 19.Miller, G. (1991). WordNet: An on-line lexical database. International Journal of Lexicography, 3(4), 235–312.
  20. 20.Minsky, M., & Papert, S. (1969). Perceptrons. MIT Press, Cambridge, MA.
  21. 21.Pazzani, M., Muramatsu J., and Billsus, D. (1996). Syskill & Webert: Identifying interesting web sites.Proceedings of the National Conference on Artificial Intelligence (pp. 54–61). Portland, OR.
  22. 22.Quinlan, J.R. (1986). Induction of decision trees. Machine Learning, 1:81–106.
  23. 23.Rachlin, Kasif, Salzberg, & Aha, (1994). Towards a better understanding of memory-based reasoning systems. Proceedings of the Eleventh International Conference on Machine Learning (pp. 242–250). New Brunswick, NJ.
  24. 24.Rocchio, J. (1971). Relevance feedback information retrieval. In Gerald Salton (Ed.), The SMART Retrieval System—Experiments in Automated Document Processing (pp. 313–323). Prentice-Hall, Englewood Cliffs, NJ.
  25. 25.Rumelhart, D., Hinton, G., & Williams, R. (1986). Learning internal representations by error propagation. In D. Rumelhart & J. McClelland (Eds.), Parallel Distributed Processing: Explorations in the Microstructure of Cognition. Volume 1: Foundations, (pp. 318–362). MIT Press, Cambridge, MA.
  26. 26.Salton, G. (1989). Automatic Text Processing. Addison-Wesley.
  27. 27.Salton, G., & Buckley, C. (1990). Improving retrieval performance by relevance feedback. Journal of the American Society for Information Science, 41:288–297.
  28. 28.Skalak, D. (1994). Prototype and feature selection by sampling and random mutation hill climbing algorithms. Proceedings of the Eleventh International Conference on Machine Learning (pp. 293–301). New Brunswick, NJ.
  29. 29.Stanfill, C., & Waltz, D. (1986). Towards memory-based reasoning. Communications of the ACM, 29:1213–1228.
  30. 30.Widrow, G., & Hoff, M. (1960). Adaptive switching circuits. Institute of Radio Engineers, Western Electronic Show and Convention, Convention Record, Part 4.

Citation

MLA
Pazzani, M., and D. Billsus. “Learning and Revising User Profiles: The Identification of Interesting Web Sites”. Machine Learning, vol. 27, no. 3, 1997, pp. 313–31, https://doi.org/10.1023/A:1007369909943.
APA
Pazzani, M., & Billsus, D. (1997). Learning and Revising User Profiles: The Identification of Interesting Web Sites. Machine Learning, 27(3), 313–331. https://doi.org/10.1023/A:1007369909943
Chicago
Pazzani, M., and D. Billsus. 1997. “Learning and Revising User Profiles: The Identification of Interesting Web Sites”. Machine Learning 27 (3): 313–31. https://doi.org/10.1023/A:1007369909943.
Harvard
Pazzani, M. and Billsus, D. (1997) “Learning and Revising User Profiles: The Identification of Interesting Web Sites”, Machine Learning, 27(3), pp. 313–331. Available at: https://doi.org/10.1023/A:1007369909943.
Vancouver
1. Pazzani M, Billsus D (1997) Learning and Revising User Profiles: The Identification of Interesting Web Sites. Machine Learning 27:313–331

BibTeX

@article{Pazzani_1997, title={Learning and Revising User Profiles: The Identification of Interesting Web Sites}, volume={27}, ISSN={1573-0565}, url={http://dx.doi.org/10.1023/A:1007369909943}, DOI={10.1023/a:1007369909943}, number={3}, journal={Machine Learning}, publisher={Springer Science and Business Media LLC}, author={Pazzani, Michael and Billsus, Daniel}, year={1997}, month=June, pages={313–331} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF