Web mining research: a survey

Raymond KosalaHendrik Blockeel

article2000SKDD1,912 citations

Establishes a clear three-part taxonomy of web content, structure, and usage mining to resolve conceptual confusion across database systems, information retrieval, and machine learning research.

Listen

Rapid expansion of online data and electronic services has led to acute information overload, making it difficult for users to find relevant content, synthesize new knowledge, and obtain personalized experiences, while site operators struggle to understand user behavior and optimize online structures. Although data mining offers powerful mechanisms to extract value from large volumes of data, conflicting definitions and overlapping terminology from artificial intelligence, database systems, and information retrieval have created confusion in the field.

The article establishes a structured taxonomy for web mining, clarifying conceptual boundaries and classifying emerging research based on data representation, processing workflows, algorithms, and practical applications.

The authors conduct a comprehensive literature survey, evaluating standard four-stage knowledge discovery workflows (resource finding, information selection and pre-processing, pattern generalization, and pattern analysis) across diverse methodologies and user-agent paradigms.

The survey yields four key findings. First, web mining divides cleanly into three operational categories: content mining (extracting information from text, multimedia, or semi-structured records), structure mining (analyzing hyperlinks to identify authoritative hubs and community networks), and usage mining (evaluating server and interaction logs to identify behavioral trends). Second, web content mining operates through two distinct perspectives: an information retrieval view focused on categorization and extraction across unstructured and semi-structured documents, and a database view that models sites via graph schemas to support complex querying. Third, while traditional text mining relies heavily on single-word representations, no single data representation—whether phrases, conceptual hierarchies, or relational logic—consistently outperforms others across diverse text categorization domains. Fourth, graph representations are pervasive across structural and semi-structured web tasks, exposing significant limitations in standard machine learning algorithms designed exclusively for flat, tabular data.

These findings indicate that addressing web scale and complexity requires multi-disciplinary techniques combining machine learning, natural language processing, and database schema management. Relying purely on traditional keyword retrieval or basic data analysis creates operational risks, including poor search precision and lost insights into customer needs. Adopting structured web mining frameworks enables organizations to build more adaptive user interfaces, enhance business intelligence, and automate knowledge discovery from online assets.

Decision-makers and engineering leaders should focus development on hybrid solutions that combine content and usage signals for personalization, invest in graph-capable machine learning tools, and explore automated wrappers to integrate heterogeneous data sources. Priority research should address the integration of dynamic web data, schema maintenance, wrapper robustness, and time-sensitive topic tracking.

The article notes several limitations, including the dynamic nature of web content, data quality challenges in user interaction logs due to proxy caching, and the nascent stage of multimedia data mining. Consequently, organizations should apply web mining models with appropriate safeguards, validating behavioral inferences and investing in robust data pre-processing pipelines.

arXiv: cs/0011033
Cover for Web mining research: a survey

Abstract

With the huge amount of information available online, the World Wide Web is a fertile area for data mining research. The Web mining research is at the cross road of research from several research communities, such as database, information retrieval, and within AI, especially the sub-areas of machine learning and natural language processing. However, there is a lot of confusions when comparing research efforts from different point of views. In this paper, we survey the research in the area of Web mining, point out some confusions regarded the usage of the term Web mining and suggest three Web mining categories. Then we situate some of the research with respect to these three categories. We also explore the connection between the Web mining categories and the related agent paradigm. For the survey, we focus on representation issues, on the process, on the learning algorithm, and on the application of the recent works as the criteria. We conclude the paper with some research issues.

Table of Contents

  • 1 Introduction
  • 2 Web Mining
  • 2.1 Overview
  • 2.1.1 Web Mining and Information Retrieval
  • 2.1.2 Web Mining and Information Extraction
  • 2.1.3 Web Mining and Machine Learning Applied on the Web
  • 2.2 Web Mining Categories
  • 2.3 Web Mining and the Agent Paradigm
  • 3 Web Content Mining
  • 3.1 Information Retrieval View
  • 3.1.1 Information Retrieval View for Unstructured Documents
  • 3.1.2 Information Retrieval View for Semi-Structured Documents
  • 3.2 Database View
  • 3.3 About Mining Multimedia Data
  • 4 Web Structure Mining
  • 5 Web Usage Mining
  • 6 Related Works
  • 7 Conclusions
  • 8 Acknowledgements
  • References

Knowls

  1. Knowl 1 — Definition and Subtask Decomposition of Web Mining

    definition

    Web mining is defined as the application of data mining and machine learning techniques to automatically discover, extract, and generalize information and patterns from World Wide Web documents, structures, and services. The knowledge discovery process on the Web can be structured into four sequential subtasks:

    1. Resource finding: The retrieval of target text or hypertext sources on the Web, including online web pages, electronic newsletters, newsgroups, text databases, and legacy data migrated to Web interfaces.
    2. Information selection and pre-processing: The automatic extraction and transformation of retrieved raw data into structured feature representations. This includes linguistic cleaning (stop-word removal, stemming), feature extraction (nn-grams, keyphrases, named entities), relational first-order logic transformations, or graph-based structural modeling.
    3. Generalization: The automated discovery of general patterns, predictive models, extraction rules, or structural schemas across individual Web sites or multiple sites using machine learning or data mining algorithms.
    4. Analysis: The validation, interpretation, and visualization of discovered patterns, which often incorporates human-in-the-loop validation due to the interactive nature of the Web.
  2. Knowl 2 — Three-Way Taxonomy of Web Mining

    definition

    Web mining is categorized into three primary domains based on the specific dimension of the Web being mined:

    1. Web Content Mining: The discovery of useful information and patterns from the actual content, documents, and data published on the Web. Content types include unstructured text, semi-structured documents (such as HTML pages), structured tables/databases, and multimedia data.
    2. Web Structure Mining: The discovery of models and topological patterns underlying the hyperlink graph of the Web. It applies social network analysis to evaluate inter-document link structures, identifying authoritative sources, hub pages, and emergent communities.
    3. Web Usage Mining: The discovery of behavioral patterns from secondary data generated during user interactions with Web systems. It analyzes clickstreams and access logs from client browsers, proxy servers, and Web servers to model navigation behaviors and user profiles.
  3. Knowl 3 — Comparative Dimensions of Web Mining Categories

    data/table

    The primary categories of Web mining differ systematically across data views, primary input sources, underlying representations, computational methods, and target application domains:

    Dimension Content Mining (IR View) Content Mining (DB View) Web Structure Mining Web Usage Mining
    View of Data Unstructured / Semi-structured Semi-structured / Web site as DB Hyperlink topology User interactivity
    Main Data Text documents, Hypertext documents Hypertext documents Directed hyperlink graph Web server logs, Browser logs, Proxy logs
    Representation Bag of words, nn-grams, phrases, concepts/ontologies, relational (first-order logic) Edge-labeled graph (OEM), relational Directed graph Relational tables, directed graphs
    Method TF-IDF, Machine learning, Statistical NLP, Inductive Logic Programming (ILP) Proprietary algorithms, ILP, modified association rules, attribute-oriented induction Link analysis algorithms (e.g., HITS, PageRank, Clever) Machine learning, Statistical modeling, sequence discovery, association rules
    Applications Document categorization, clustering, extraction rule learning, user modeling Schema extraction, DataGuide construction, frequent substructure discovery, multilevel databases Page categorization, authority/hub discovery, cyber-community trawling Site restructuring, adaptive personalization, e-commerce marketing, user modeling

    This comparison highlights how the Information Retrieval view operates globally across documents, the Database view focuses locally on schema modeling and database transformation, Web structure mining exploits hyperlink graph topologies, and Web usage mining analyzes behavioral interaction trails.

  4. Knowl 4 — Dual Perspectives on Web Content Mining: IR View vs. DB View

    model/method

    Web content mining is bifurcated into two primary paradigms reflecting distinct community perspectives:

    • Information Retrieval (IR) View: Adopts a global scope aimed at improving document retrieval, filtering, and organization based on explicit or inferred user profiles. Unstructured text is commonly modeled via vector space bags-of-words, terms, or relational logic. Semi-structured hypertext incorporates HTML tag hierarchies and anchor text. Tasks include text classification, clustering, event detection and tracking (TDT), and automated rule learning for information extraction.
    • Database (DB) View: Adopts a local scope aimed at integrating and transforming semi-structured Web site data into queryable database models to support structured queries beyond keyword search. Web pages are modeled as semi-structured data graphs, predominantly using the Object Exchange Model (OEM), where vertices represent objects and labeled edges represent relations. Primary tasks include building structural summaries (schemas or DataGuides), discovering frequent substructures, constructing Multi-Layered Databases (MLDB), and executing declarative query languages over semi-structured Web repositories.
  5. Knowl 5 — Mapping Web Mining Categories to Intelligent Agent Paradigms

    model/method

    Intelligent software agents that filter and retrieve Web information correspond directly to specific Web mining categories based on their underlying filtering mechanisms:

    Agent Information Filtering Mechanism Associated Web Mining Category
    Content-based filters Web Content Mining
    Reputation-based filters Web Structure Mining (and Content Mining)
    Collaborative / social-based filters Web Usage Mining
    Event-based filters Web Usage Mining
    Hybrid filters Combination of Web Mining Categories
    • Content-based agents: Analyze textual or structural page attributes to match document features with explicit user preferences (Web content mining).
    • Reputation-based agents: Determine document importance based on network link prestige, citations, and authority topology (Web structure mining).
    • Collaborative/social agents: Aggregate ratings, profiles, and browsing histories across similar users to recommend relevant content (Web usage mining).
    • Event-based agents: Track discrete user actions—such as bookmarking URLs, mouse clicks, scrolling, and link traversal paths—to adaptively personalize interfaces (Web usage mining).
    • Hybrid agents: Integrate content characteristics, hyperlink structures, and historical user interaction logs.
  6. Knowl 6 — Distinctions Between Web Mining, Information Retrieval, and Information Extraction

    definition

    Web mining has specific boundaries and functional overlaps with Information Retrieval (IR) and Information Extraction (IE):

    • Web Mining vs. Information Retrieval: IR focuses on retrieving relevant documents while excluding non-relevant documents from a large repository, treating documents typically as bags of unordered words. Web mining serves as a component within IR when machine learning techniques are applied to automate document categorization, clustering, and index term construction. However, IR encompasses non-mining tasks such as indexing structures, query parsing, visualization, and user interface design.
    • Web Mining vs. Information Extraction: IE operates at a finer granularity than IR, focusing on extracting specific semantic facts, entities, and relations from within documents to populate structured databases. Web mining intersects with IE through wrapper induction and rule learning, where machine learning algorithms automatically infer extraction patterns from annotated Web documents, bypassing manual wrapper engineering.
  7. Knowl 7 — Web Structure Mining Principles and Applications

    model/method

    Web structure mining models the topology of the World Wide Web using graph-theoretic algorithms and social network analysis. It analyzes inter-document relationships (hyperlinks connecting distinct pages) as opposed to intra-document markup.

    • Authority and Hub Modeling: Algorithms such as HITS and PageRank determine quality ranking and relevance via directed graph structure. Authorities are central Web pages referenced by many relevant sources. Hubs are catalog or directory pages containing multiple outgoing links to authoritative pages. Extensions to HITS integrate anchor-text content and outlier filtering to mitigate topic drift.
    • Cyber-Community Discovery: Analysis of dense bipartite subgraphs in the Web graph enables the automatic identification of emerging micro-communities sharing specialized interests.
    • Web Warehouse Link Metrics: Hyperlink topology is used to measure site completeness (ratio of internal vs. external links), detect mirror sites across distributed servers through structural replication, and evaluate hierarchical site designs to optimize user navigation flow.
  8. Knowl 8 — Web Usage Mining Architecture, Preprocessing, and Application Domains

    model/method

    Web usage mining discovers patterns in the secondary behavioral logs generated as users interact with the Web.

    • Data Sources Across System Layers: Interaction data is captured across three tiers: clients (browser histories, cookies, local click/scroll events), proxy servers (intermediate routing and caching logs), and Web servers (access logs, referrer logs, session identifiers, user queries, registration databases).
    • Preprocessing Challenges: Critical data preparation tasks include user identification, session boundary identification, episode delineation, and compensating for incomplete logs caused by browser caching and proxy masking.
    • Data Mining Techniques: Approaches include transforming access logs into relational tables for association rule and sequence discovery (e.g., composite association rules, MIDAS algorithm), and modeling navigation paths using directed graphs and hypertext probabilistic grammars.
    • Application Domains: Applications divide into: (1) personalized user modeling, which learns individual user interest profiles to adapt user interfaces dynamically; and (2) impersonalized navigation pattern mining, which extracts aggregate usage trends to optimize site organization, caching strategies, and business marketing decisions.

Coverage note — Specific individual third-party algorithms and systems cited throughout the survey tables (such as individual Inductive Logic Programming algorithms, Naive Bayes variants, Support Vector Machine baselines, and specific wrapper systems) were omitted as standalone knowls because they represent surveyed external literature rather than original algorithmic or theoretical contributions of this survey.

References

  1. 1.S. Abiteboul. Querying semi-structured data. In F. N. Afrati and P. Kolaitis, editors, Database Theory ICDT '97, 6th International Conference, Delphi, Greece, January 8-10, 1997, Proceedings, volume 1186 of Lecture Notes in Computer Science, pages 1-18. Springer, 1997.
  2. 2.S. Abiteboul, D. Quass, J. McHugh, J. Widom, and J. L. Wiener. The lorel query language for semistructured data. Int. J. on Digital Libraries, 1(1):68-88, 1997.
  3. 3.H. Ahonen, O. Heinonen, M. Klemettinen, and A. Verkamo. Applying data mining techniques for descriptive phrase extraction in digital document collections. In Advances in Digital Libraries (ADL'98), Santa Barbara, California, USA, April 1998, 1998.
  4. 4.H. Ahonen, O. Heinonen, M. Klemettinen, and A. Verkamo. Finding co-occurring text phrases by combining sequence and frequent set discovery. In R. Feldman, editor, Proceedings of 16th International Joint Conference on Artificial Intelligence IJCAI-99 Workshop on Te~ Mining: Foundations, Techniques and Applications, pages 1-9, 1999.
  5. 5.J. Allan, J. Carbonell, G. Doddington, J. Yamron, and Y. Yang. Topic detection and tracking pilot study: Final report. In Proceedings of the DARPA Broadcast News Transcription and Understanding Workshop, 1998, 1998.
  6. 6.J. Allan, R. Papka, and V. Lavrenko. On-line new event detection and tracking. In Proceedings of the glst annual international ACM SIGIR conference on Research and development in information retrieval August 2~ - gS, 1998, pages 37-45, Melbourne Australia, 1998.
  7. 7.D. E. Appelt and D. Israel. Introduction to information extraction technology. In Proceedings of 16th International Joint Conference on Artificial Intelligence IJCAI-99, Tutorial, 1999.
  8. 8.G. O. Arocena and A. O. Mendelzon. Weboqh Restructuring documents, databases, and webs. Theory and Practice of Object Systems, 5(3):127-141, 1999.
  9. 9.P. Atzeni and G. Mecca. Cut & paste. In Proceedings of the Sixteenth A CM SIGA CT-SIGMOD-SIGART Symposium on Principles of Database Systems, May 1P-1~, 1997, Tucson, Arizona, pages 144-153. ACM Press, 1997.
  10. 10.R. Baeza-Yates and e. Berthier Ribeiro-Neto. Modern Information Retrieval. Addison-Wesley Longman Publishing Company, 1999.
  11. 11.M. Balabanovi'c and Y. Shoham. Fab: Contentbased, collaborative recommendation. Communications of the ACM, 40(3):66-70, 1997.
  12. 12.A. Biichner, M. Baumgarten, S. Anand, M. Mulvenna, and J. Hughes. Navigation pattern discovery from internet data. In Proceedings of the WEBKDD '99 Workshop on Web Usage Analysis and User Profiling, August 15, 1999, San Diego, CA, USA, 1999.
  13. 13.K. Bharat and M. 1~. Henzinger. Improved algorithms for topic distillation in a hyperlinked environment. In Proceedings of the 21st annual international ACM SIGIR conference on Research and development in information retrieval August PJ - ~8, 1998, pages 104-111, Melbourne Australia, 1998.
  14. 14.D. Billsus and M. Pazzani. A hybrid user model for news story classification. In Proceedings of the Seventh International Conference on User Modeling (UM '99), Banff, Canada, 1999.
  15. 15.J. Borges and M. Levene. Mining association rules in hypertext databases. In Proceedings of the Fourth International Conference on Knowledge Discovery and Data Mining (KDD-98), August 27-31, 1998, New York City, New York, USA, 1998.
  16. 16.
    1. Borges and M. Levene. Data mining of user navigation patterns. In Proceedings of the WBBKDD'99 Workshop on Web Usage Analysis and User Profiling, August 15, 1999, San Diego, CA, USA, pages 31-36, 1999.
  17. 17.S. Brin and L. Page. The anatomy of a large-scale hypertextual Web search engine. In Seventh International World Wide Web Conference, Brisbane, Australia, 1998.
  18. 18.P. Buneman. Semistructured data. In Proceedings of the Sixteenth A CM SIGA CT-SIGMOD-SIGART Symposium on Principles of Database Systems, May 1~-14, 1997, Tucson, Arizona, pages 117-121. ACM Press, 1997.
  19. 19.P. Buneman, S. B. Davidson, G. G. Hillebrand, and D. Suciu. A query language and optimization techniques for unstructured data. In H. V. Jagadish and I. S. Mumick, editors, Proceedings of the 1996 ACM SIGMOD International Conference on Management of Data, Montreal, Quebec, Canada, June J-6, 1996, pages 505-516. ACM Press, 1996.
  20. 20.J. Carbonell, M. Craven, S. Fienberg, T. Mitchell, and Y. Yang. Report on the conald workshop on learning from text and the web. In CONALD Workshop on Learning from Text and the Web, June, 1998, 1998.
  21. 21.J. Caxbonell, Y. Yang, and W. Cohen. Special issue of machine learning on information retrieval introduction. Machine Learning, 39:99-101, 2000.
  22. 22.C. Cardie. Empirical methods in information extraction. AI Magazine, 18(4):65-79, 1997.
  23. 23.S. Chakrabarti. Data mining for hypertext: A tutorial survey. ACM SIGKDD Explorations, 1(2):1-11, 2000.
  24. 24.S. Chakrabarti, B. Dora, D. Gibson, J. Kleinberg, S. Kumar, P. Pghavan, S. tLjagopalan, and A. Tomkins. Mining the link structure of the world wide web. IEEE Computer, 32(8):60-67, 1999.
  25. 25.S. Chakrabarti, B. Dora, and P. Indyk. Enhanced hypertext categorization using hyperlinks. In L. M. Haas and A. Tiwary, editors, SIGMOD 1998, Proceedings ACM SIGMOD International Conference on Management of Data, June 2-~, 1998, Seattle, Washington, USA, pages 307-318. ACM Press, 1998.
  26. 26.S. Chawathe, H. Garcia-Molina, J. Hammer, K. Ireland, Y. Papakonstantinou, J. Ullman, and J. Widom. The tsimmis project: Integration of heterogeneous information sources. In Proceedings of the lOth Meeting of the Information Processing Society of Japan, pages 7-18, 1994.
  27. 27.W.W. Cohen. Learning to classify english text with ilp methods. In Advances in Inductive Logic Programming (Ed. L. De Raedt), IOS Press, 1995.
  28. 28.W.W. Cohen. Some practical observations on integration of web information. In A CM SIGMOD Workshop on The Web and Databases (WebDB'99), pages 55-60, Philadelphia, Pennsylvania, USA, 1999.
  29. 29.W. W. Cohen. What can we learn from the web? In Proceedings of the Sixteenth International Conference on Machine Learning (ICML '99), pages 515-521, 1999.
  30. 30.R. Cooley, B. Mobasher, and J. Srivastava. Web mining: Information and pattern discovery on the world wide web. In Proceedings of the 9th IEEE International Conference on Tools with Artificial Intelligence (ICTAI'97), 1997.
  31. 31.R. Cooley, B. Mobasher, and J. Srivastava. Data preparation for mining world wide web browsing patterns. Knowledge and Information Systems, 1(1), 1999.
  32. 32.R. W. Cooley. Web Usage Mining: Discovery and Application of Interesting Patterns from Web data. PhD thesis, Dept. of Computer Science, University of Minnesota, May 2000.
  33. 33.J. Cowie and W. Lehnert. Information extraction. Communications of the ACM, 39(1):80-91, 1996.
  34. 34.M. Craven, D. DiPasquo, D. Freitag, A. McCallum, T. Mitchell, K. Nigam, mad S. Slattery. Learning to extract symbolic knowledge from the world wide web. In Proceedings of the Fifteenth National Conference on Artificial Intellligence (AAAI98), pages 509-516, 1998.
  35. 35.F. Crimmins, A. Smeaton, T. Dkaki, and J. Mothe. T~trafusion: Information discovery on the internet. IEEE Intelligent Systems, 14(4):55-62, 1999.
  36. 36.S. Deerwester, S. Dumais, G. Furnas, T. Landauer, and R. Harshman. Indexing by latent semantic analysis. Journal of the American Society for Information Science, 41(6):391-407, 1990.
  37. 37.L. Dehaspe and L. de Raedt. Mining association rules in multiple relations. In Proceedings of the 7th International Workshop on Inductive Logic Programming, volume 1297 of Lecture Notes in Computer Science, pages 125-132, Prague, Czech Republic, 1997. Springer.
  38. 38.J. A. Delgado. Agent-Based Information Filtering and Recommender System On the Internet. PhD thesis, Dept. of Intelligence Computer Science, Nagoya Institute of Technology, March 2000.
  39. 39.S. Dumais, J. Platt, D. Heckerman, and M. Sahami. Inductive learning algorithms and representations for text categorization. In Proceedings of the 1998 A CM 7th international conference on Information and knowledge management, pages 148-155, Washington United States, 1998.
  40. 40.J. S.. T. Eliassi-Rad. Intelligent agents for web-based tasks: An advice-taking approach. In Working Notes of the AAAI/ICML-98 Workshop on Learning for Text Categorization, Madison, WI, pages 588-589, 1999.
  41. 41.O. Etzioni. The world wide web: Quagmire or gold mine. Communications of the ACM, 39(i1):65-68, 1996.
  42. 42.U. Fayyad, S. Djorgovski, and N. Weir. Automating the analysis and cataloging of sky surveys. In Advances in Knowledge Discovery and Data Mining, pages 471- 493. AAAI Press, 1996.
  43. 43.U. Fayyad, G. Piatetsky-Shapiro, and P. Smyth. From data mining to knowledge discovery: An overview. In Advances in Knowledge Discovery and Data Mining, pages 1-34. AAAI Press, 1996.
  44. 44.U. Fayyad, G. Piatetsky-Shapiro, and P. Smyth. Knowledge discovery and data mining: toward a unifying framework. In Proceeding of The Second Int. Conference on Knowledge Discovery and Data Mining, pages 82-88, 1996.
  45. 45.R. Feldman and I. Dagan. Knowledge discovery in textual databases (kdt). In Proceedings of the First International Conference on Knowledge Discovery and Data Mining (KDD-95), pages 112-117, Montreal, Canada, 1995.
  46. 46.R. Feldman, M. Fresko, Y. Kinar, Y. Lindell, O. Liphstar, M. Rajman, Y. Schler, and O. Zamir. Text mining at the term level. In Principles of Data Mining and Knowledge Discovery, Second European Symposium, PKDD '98, volume 1510 of Lecture Notes in Computer Science, pages 56-64. Springer, 1998.
  47. 47.D. Fensel, C. Knoblock, N. Kushmerick, and M.-C. Rousset. Workshop on intelligent information integration (iii'99). AI Magazine, 21(1):91-94, 2000.
  48. 48.M. F. Fernandez, D. Floreseu, A. Y. Levy, and D. Suciu. A query language for a web-site management system. SIGMOD Record, 26(3):4-11, 1997.
  49. 49.R. E. Filman and S. Pant. Searching the internet - guest editors' introduction. IEEE Internet Computing, 2(4):21-23, 1998.
  50. 50.D. Florescu, A. Y. Levy, and A. O. Mendelzon. Database techniques for the world-wide web: A survey. SIGMOD Record, 27(3):59-74, 1998.
  51. 51.E. Frank, G. W. Paynter, I. H. Witten, C. Gutwin, and C. G. Nevill-Manning. Domain-specific keyphrase extraction. In Proceedings of 16th International Joint Conference on Artificial Intelligence IJCAI-99, pages 668-673, 1999.
  52. 52.D. Freitag. Information extraction from htmh Application of a general learning approach. In Proceedings of the Fifteenth Conference on Artificial Intelligence AAAI-98 (1998), pages 517-523, 1998.
  53. 53.D. Freitag and A. McCallum. Information extraction with hmms and shrinkage. In Proceedings of the AAAL 99 Workshop on Machine Learning for Information Extraction, 1999.
  54. 54.J. Fiirnkranz. Exploiting structural information for text classification on the www. In Advances in Intelligent Data Analysis, Third International Symposium, IDA-99, pages 487-498, 1999.
  55. 55.M. N. Garofalakis, R. Rastogi, S. Seshadri, and K. Shim. Data mining and the web: Past, present and future. In Workshop on Web Information and Data Management, 1999, pages 43-47, 1999.
  56. 56.R. Goldman and J. Widom. Dataguides: Enabling query formulation and optimization in semistructured databases. In M. Jarke, M. J. Carey, K. R. Dittrich, F. H. Lochovsky, P. Loucopoulos, and M. A. Jeusfeld, editors, VLDB'97, Proceedings of P3rd International Conference on Very Large Data Bases, August 25-29, 1997, Athens, Greece, pages 436-445. Morgan Kaufmann, 1997.
  57. 57.R. Goldman and J. Widom. Approximate dataguides. In Proceedings of the Workshop on Query Processing for Semistructured Data and Non-Standard Data Formats, 1999.
  58. 58.S. Green, L. Hurst, B. Nangle, P. Cunningham, F. Somers, and R. Evans. Software agents: A review. Technical Report TCD-CS-1997-06, Technical Report of Trinity College, University of Dublin, 1997.
  59. 59.S. Grumbach and G. Mecca. In search of the lost schema. In Database Theory - ICDT '99, 7th International Conference, pages 314-331, 1999.
  60. 60.J. Hammer, H. Garcia-Molina, J. Cho, A. Crespo, and R. Aranha. Extracting semistructured information from the web. In Proceedings of the Workshop on Management of Semistructured Data, pages 18-25, 1997.
  61. 61.A. Hauptmann. Integrating and using large databases of text, image, video and audio. IEEGE Intelligent Systems, 14(5):34-35, 1999.
  62. 62.M. A. Hearst. Untangling text data mining. In Proceedings of ACL'99: the 37th Annual Meeting of the Association for Computational Linguistics, 1999.
  63. 63.T. Hofmann. The cluster-abstraction model: Unsupervised learning of topic hierarchies from text data. In Proceedings of 16th International Joint Conference on Artificial Intelligence IJCAI-99, pages 682-687, 1999.
  64. 64.S. J. Hong and S. M. Weiss. Advances in predictive model generation for data mining. Technical Report Report RC-21570, IBM Research Report, 1999.
  65. 65.T. Honkela, S. Kaski, K. Lagus, and T. Kohonen. Websom - self-organizing maps of document collections. In Proc. of Workshop on Self-Organizing Maps 1997 (WSOM'97), pages 310-315, 1997.
  66. 66.A. Houston, H. Chen, S. M. Hubbard, B. R. Schatz, T. D. Ng, R. R. Sewell, and K. M. Tolle. Medical data mining on the internet: Research on a cancer information system. Artificial Intelligence Review, 13:437-446, 1999.
  67. 67.C.-N. Hsu and M.-T. Dung. Generating finite-state transducers for semi-structured data extraction from the web. Information Systems, 23(8):521-538, 1998.
  68. 68.T. Joachims, D. Freitag, and T. Mitchell. Webwatcher: A tour guide for the world wide web. In Proceedings of the International Joint Conference on Artificial Intelligence IJCAI-97, pages 770-777, 1997.
  69. 69.M. Junker, M. Sintek, and M. Rinck. Learning for text categorization and information extraction with ilp. In Proceedings of the Workshop on Learning Language in Logic, Bled, Slovenia, 1999, 1999.
  70. 70.H. L. K. Wang. Discovering association of structure from semistructured objects. To appear in IEgEgE Transactions on Knowledge and Data Engineering, 1999.
  71. 71.H. Kargupta, I. Hamzaogiu, and B. Stafford. Distributed data mining using an agent based architecture. In Proceedings of Knowledge Discovery And Data Mining, pages 211-214. AAAI Press, 1997.
  72. 72.H. Kautz, B. Selman, and M. Shah. The hidden web. A1 magazine, 18(2):27-36, 1997.
  73. 73.S. Khoshafian and A. B. Baker. Multimedia and Imaging Databases. Morgan Kaufmann Publishers, 1996.
  74. 74.J. M. Kleinberg. Authoritative sources in a hyperlinked environment. In Proc. of A CM-SIAM Symposium on Discrete Algorithms, 1998, pages 668-677, 1998.
  75. 75.Y. Kodratoff. About knowledge discovery in texts: A definition and an example. In Proc. of Advanced Course on Artificial Intelligence 1999 (ACAI-99) on Machine Learning Applications (Invited talk), 1999.
  76. 76.S. R. Kumar, P. Raghavan, S. Kajagopalan, and A. Tomkins. Trawling the web for emerging cybercommunities. In Proceedings of the Eighth World Wide Web Conference (WWWS), 1999.
  77. 77.N. Kushmerick. Gleaning the web. IBBE Intelligent Systems, 14(2):20-22, 1999.
  78. 78.N. Kushmerick, D. Weld, and R. Doorenbos. Wrapper induction for information extraction. In Proceedings of the International Joint Conference on Artificial Intelligence IJCAI-97, pages 729--737, 1997.
  79. 79.L. Lakshmanem, F. Sadri, and I. Subramanian. A declarative language for querying and restructuring the web. In Proceedings of 6th. International Workshop on Research Issues in Data Engineering, RIDGE '96, pages 12-21, 1996.
  80. 80.P. Langley. User modeling in adaptive interfaces. In Proceedings of the Seventh International Conference on User Modeling, pages 357-370, 1999.
  81. 81.S. Lawrence and C. L. Giles. Accessibility of information on the web. Nature, 400:107-109, 1999.
  82. 82.lberto O. Mendelzon, G. A. Mihalla, and T. Milo. Querying the world wide web. In Proceedings of the Fourth International Conference on Parallel and Distributed Information Systems, pages 80-91, 1996.
  83. 83.B. Lent, R. Agrawal, and R. Srikant. Discovering trends in text databases. In Proc. 3 rd Int Conf. On Knowledge Discovery and Data Mining (KDD 1997), pages 227-230, 1997.
  84. 84.A. Y. Levy and D. S. Weld. Intelligent internet systems. Artificial Intelligence, 118(1-2), 2000.
  85. 85.S. K. Madria, S. S. Bhowmick, W. K. Ng, and E.-P. Lira. Research issues in web data mining. In Proceedings of Data Warehousing and Knowledge Discovery, First International Conference, DaWaK '99, pages 303-312, 1999.
  86. 86.P. Maes. Agents that reduce work and information overload. Communications of the ACM, 37(7):30-40, 1994.
  87. 87.B. Masand and M. Spiliopoulou. Webkdd-99: Workshop on web usage analysis and user profiling. SIGKDD Explorations, 1(2), 2000.
  88. 88.A. McCallum, K. Nigam, J. Rennie, and K. Seymore. A machine learning approach to building domainspecific search engines. In Proceedings of the International Joint Conference on Artificial Intelligence IJCAI-99, pages 662-667, 1999.
  89. 89.T. Mitchell. Machine Learning. McGraw Hill, 1997.
  90. 90.T. M. Mitchell. Machine learning and data mining. Communications of the ACM, 42(11):30-36, 1999.
  91. 91.D. Mladenic. Text-learning and related intelligent agents. 1EgEE Intelligent Systems, 14(4):44-54, 1999.
  92. 92.D. Mladenic and M. Grobelnik. Feature selection for unbalanced class distribution and naive bayes. In Proceedings of the 16th International Conference on Machine Learning ICML-99, pages 258-267, 1999.
  93. 93.I. Muslea. Extraction patterns for information extraction tasks: A survey. In AAAI-99 Workshop on Machine Learning for Information Extraction, 1999.
  94. 94.I. Muslea, S. Minton, and C. Knoblock. Wrapper induction for semistructured, web-based information sources. In Proceedings of the Conference on Automatic Learning and Discovery CONALD-98, 1998.
  95. 95.U. Y. Nahm and R. J. Mooney. Ua mutually beneficial integration of data mining and information extraction. In Proceedings of the Seventeenth National Conference on Artificial Intelligence (AAAI-O0), 2000.
  96. 96.S. Nestorov, S. Abiteboul, and R. Motwani. Infering structure in semistructured data. SIGMOD Record, 26(4), 1997.
  97. 97.S. Nestorov, S. Abiteboul, and R. Motwani. Extracting schema from semistructured data. In L. M. Haas and A. Tiwary, editors, SIGMOD 1998, Proceedings ACM SIGMOD International Conference on Management of Data, June P-G, 1998, Seattle, Washington, USA, pages 295-306. ACM Press, 1998.
  98. 98.K. Nigam, J. Lafferty, and A. McCallum. Using maximum entropy for text classification. In Proceedings of the International Joint Conference on Artificial Intelligence IJCAI-99 Workshop on Machine Learning for Information Filtering, pages 61-67, 1999.
  99. 99.G. Paliouras, C. Papatheodorou, V. Karkaletsis, P. Tzitziras, and C. D. Spyropoulos. Large-scale mining of usage data on web sites. In AAAI PO00 Sprin 9 Symposium on Adaptive User Interfaces, 2000.
  100. 100.M. T. Pazienza, editor. Information Extraction: A multidisciplinary Approach to an Emerging Information Technology, volume 1299 of Lecture Notes in Computer Science. International Summer School, SCIE-97, Frascati (Rome), Springer, 1997.
  101. 101.M. T. Pazienza, editor. Information Extraction, Frascati (Rome), 1999. International Summer" School, SCIE-99 , Frascati (Rome).
  102. 102.G. Piatetsky-Shapiro, It. Braachman, T. Khabaza, W. Kloesgen, and E. Simoudis. An overview of issues in developing industrial data mining and knowledge discovery applications. In Proceeding of The Second Int. Conference on Knowledge Discovery and Data Mining, 1996, pages 89-95, 1996.
  103. 103.M. Rajman and R. Besan~on. Text mining - knowledge extraction from unstructured textual data. In Proc. of 6th Conference of International Federation of Classification Societies (IFCS-98)? Roma (Italy), pages 473- 480, 1998.
  104. 104.A. Rauber and D. Merkl. Automatic labeling of selforganizing maps: Making a treasure-map reveal its secrets. In Proc of the Pacific Asia Conf on Knowledge Discovery and Data Mining (PAKDD'99), Beijing, China, 1999.
  105. 105.J. ttennie and A. McCallum. Using reinforcement learning to spider the web efficiently. In Proceedings of the 16th International Conference on Machine Learning ICML-99, 1999.
  106. 106.E. Riloff. Little words can make a big difference for text classification. In E. A. Fox, P. Ingwersen, and R. Fidel, editors, SIGIR'95, Proceedings off the 18th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval. Seattle, Washington, USA, July 9-13, 1995 (Special Issue of the SIGIR Forum), pages 130-136. ACM Press, 1995.
  107. 107.G. Salton and M. McGill. Introduction to Modern Information Retrieval. McGraw Hill, 1983.
  108. 108.S. Scott and S. Matwin. Feature engineering for text classification. In Proceedings of the 16th International Conference on Machine Learning ICML-99, 1999.
  109. 109.L. Singh, B. Chen, It. Halght, P. Scheuermann, and K. Aoki. A robust system architecture for mining semistructured data. In Proceeding of The Second Int. Conference on Knowledge Discovery and Data Mining, 1998, pages 329-333, 1998.
  110. 110.P. Smyth, U. M. Fayyad, M. C. Burl, and P. Perona. Modeling subjective uncertainty in image annotation. Advances in Knowledge Discovery and Data Mining, pages 517-539, 1996.
  111. 111.S. Soderland. Learning information extraction rules for semi-structured and free text. Machine Learning, 34(1-3):233-272, 1996.
  112. 112.M. Spiliopoulou. Data mining for the web. In Principles of Data Mining and Knowledge Discovery, Second European Symposium, PKDD '99, pages 588-589, 1999.
  113. 113.J. Srivastava, R. Cooley, M. Deshpande, and P.-N. Tan. Web usage mining: Discovery and applications of usage patterns from web data. SIGKDD Explorations, 1(2), 2000.
  114. 114.V. S. Subrahmanian. Principles of Multimedia Database Systems. Morgan Kaufmann Publishers, 1998.
  115. 115.A.-H. Tan. Text mining: The state of the art and the challenges. In Proc of the Pacific Asia Conf on Knowledge Discovery and Data Mining PAKDD'99 workshop on Knowledge Discovery from Advanced Databases, pages 65-70, 1999.
  116. 116.H. Toivonen. On knowledge discovery in graphstructured data. In Workshop on Knowledge Discovery from Advanced Databases (KDAD'99), pages 26-31, 1999.
  117. 117.R. Uthurusamy. From data mining to knowledge discovery: Current challenges and future directions. In Advances in Knowledge Discovery and Data Mining, pages 561-569, 1996.
  118. 118.S. Vaithyanathan. Introduction: Data mining on the internet. Artificial Intelligence Review, 13(5/6):343- 344, 1999.
  119. 119.C. 3. van Rijsbergen. Information Retrieval. Butterworths, 1979.
  120. 120.K. Wang and H. Liu. Schema discovery for semistructured data. In Proceedings of the Third International Conference on Knowledge Discovery and Data Mining (KDD'97), pages 271-274, 1997.
  121. 121.S. M. Weiss, C. Apt6, F. Damerau, D. E. Johnson, F. J. Oles, T. Goetz, and T. Hampp. Maximizing text-mining performance. IEEE Intelligent Systems, 14(4):63-69, 1999.
  122. 122.W. Wiener, 3. Pedersen, and A. Weigend. A neural network approach to topic spotting. In Proceedings of the ~th Symposium on Document Analysis and Information Retrieval (SDAIR 95), pages 317-332, 1995.
  123. 123.Y. Wilks. Information Extraction as a core language technology, volume 1299 of Lecture Notes in Computer Science, chapter In M-T. Pazienza (ed.), Information Extraction, pages 1-9. Springer, 1997.
  124. 124.I. H. Witten, Z. Bray, M. Mahoui, and W. J. Teahan. Text mining: A new frontier for lossless compression. In Data Compression Conference 1999, pages 198-207, 1999.
  125. 125.Y. Yang, J. Carbonell, It. Brown, T. Pierce, B. T. Archibald, and X. Liu. Learning approaches for detecting and tracking news events. Iggg Intelligent Systems, 14(4):32-43, 1999.
  126. 126.Y. Yang and J. Pedersen. Guest editors' introduction: Intelligent information retrieval. IEEE Intelligent Systems, 14(4):30-31, 1999.
  127. 127.O. Zg/ane and J. Han. Webmh Querying the worldwide web for resources and knowledge. In Proc. A CM CIKM'98 Workshop on Web Information and Data Management (WIDM'98), pages 9-12, 1998.
  128. 128.O. R. Zaiane, J. Han, Z.-N. Li, S. H. Chee, and J. Chiang. Multimediaminer: a system prototype for multimedia data mining. In Proc. A CM SIGMOD Intl. Conf. on Management of Data, pages 581-583, 1998.

Citation

MLA
Kosala, R., and H. Blockeel. “Web Mining Research: A Survey”. ACM SIGKDD Explorations, 2(1):1-15, 2000, 2000, http://arxiv.org/abs/cs/0011033v1.
APA
Kosala, R., & Blockeel, H. (2000). Web Mining Research: A Survey. ACM SIGKDD Explorations, 2(1):1-15, 2000. http://arxiv.org/abs/cs/0011033v1
Chicago
Kosala, R., and H. Blockeel. 2000. “Web Mining Research: A Survey”. ACM SIGKDD Explorations, 2(1):1-15, 2000. http://arxiv.org/abs/cs/0011033v1.
Harvard
Kosala, R. and Blockeel, H. (2000) “Web Mining Research: A Survey”, ACM SIGKDD Explorations, 2(1):1-15, 2000 [Preprint]. Available at: http://arxiv.org/abs/cs/0011033v1.
Vancouver
1. Kosala R, Blockeel H (2000) Web Mining Research: A Survey. ACM SIGKDD Explorations, 2(1):1-15, 2000

BibTeX

@article{kosala2000web,
  title = {Web Mining Research: A Survey},
  author = {Kosala, Raymond and Blockeel, Hendrik},
  year = {2000},
  journal = {ACM SIGKDD Explorations, 2(1):1-15, 2000},
  url = {http://arxiv.org/abs/cs/0011033v1},
  eprint = {cs/0011033}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF