Web document clustering: a feasibility demonstration

Oren ZamirOren Etzioni

article1998SIGIR1,317 citations

Introduces Suffix Tree Clustering, a linear-time algorithm that groups web search results into overlapping, phrase-labeled clusters using only search engine snippets rather than full document texts.

Listen

Web search engines typically present results as long, ranked lists of document summaries. Because search engines suffer from low precision and users face extensive result sets, finding relevant information by sifting through these linear lists is inefficient and time-consuming. Grouping search results into topical clusters offers an alternative browsing method, but traditional clustering techniques are generally too slow for real-time web use and fail to address the specific constraints of web search results.

The article evaluates whether post-retrieval document clustering is feasible for web search results and introduces a novel algorithm called Suffix Tree Clustering to meet the performance and usability demands of the web.

To evaluate this approach, the authors developed a prototype meta-search system and gathered 10 distinct web document collections, each containing 200 search snippets and their full web pages, across defined search queries with manually assigned relevance ratings. The analysis benchmarked Suffix Tree Clustering against standard algorithms—including hierarchical and iterative techniques—across retrieval precision, processing speed, and the impact of clustering short snippets versus full document texts.

The findings demonstrate four major outcomes. First, Suffix Tree Clustering outperformed traditional clustering algorithms and the default ranked list, achieving the highest average precision. Second, clustering short snippets (averaging about 20 informative words) resulted in only a slight reduction in cluster quality compared to clustering full web pages (averaging 220 informative words). Third, the algorithm operates in linear time relative to collection size and processes documents incrementally as they arrive over the network, returning clustered results in roughly 0.01 seconds after the final document download. Finally, allowing documents to belong to multiple clusters and identifying multi-word phrases were both critical to performance; 72% of documents were placed in multiple clusters, and 55% of the base clusters relied on multi-word phrases.

These results indicate that search engines and meta-search services do not need to download entire web pages or invest heavy server resources to group results meaningfully. By operating incrementally on snippet text alone, clustering can be deployed on separate intermediary machines without creating perceptible latency for users. Furthermore, extracting shared multi-word phrases naturally generates informative, browsable labels for each cluster.

Organizations developing search and information retrieval interfaces should consider deploying incremental phrase-based clustering on snippets rather than relying solely on ranked lists or computationally expensive full-text clustering. Future implementation should include controlled user studies to measure how human searchers interact with cluster labels and navigate overlapping categories in practice.

The conclusions are subject to certain limitations. The evaluation relied on a relatively small test set of 10 queries and assumed an idealized user who consistently identifies and selects the most relevant cluster. While technical performance and speed metrics are robust, direct validation of user productivity and satisfaction in live environments remains necessary.

Cover for Web document clustering: a feasibility demonstration

Abstract

Users of Web search engines are often forced to sift through the long ordered list of document “snippets” returned by the engines. The IR community has explored document clustering as an alternative method of organizing retrieval results, but clustering has yet to be deployed on the major search engines.

The paper articulates the unique requirements of Web document clustering and reports on the first evaluation of clustering methods in this domain. A key requirement is that the methods create their clusters based on the short snippets returned by Web search engines. Surprisingly, we find that clusters based on snippets are almost as good as clusters created using the full text of Web documents.

To satisfy the stringent requirements of the Web domain, we introduce an incremental, linear time (in the document collection size) algorithm called Suffix Tree Clustering (STC), which creates clusters based on phrases shared between documents. We show that STC is faster than standard clustering methods in this domain, and argue that Web document clustering via STC is both feasible and potentially beneficial.

Table of Contents

  • 1 Introduction
  • 2 Previous Work on Document Clustering
  • 3 Suffix Tree Clustering
  • 3.1 Step 1 - Document 'Cleaning'
  • 3.2 Step 2 - Identifying Base Clusters
  • 3.3 Step 3 - Combining Base Clusters
  • 4 Experiments
  • 4.1 Effectiveness for Information Retrieval
  • 4.2 Snippets versus Whole Document
  • 4.3 Execution Time
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — Suffix Tree Clustering Algorithm

    algorithm

    Suffix Tree Clustering (STC) is a linear-time, incremental post-retrieval clustering algorithm that operates on ordered sequences of words (phrases) rather than bag-of-words document vectors. STC groups documents sharing common phrases and consists of three stages:

    1. Document Cleaning: Strip non-word tokens (HTML tags, numbers, punctuation), apply light stemming (prefix and suffix removal, plural reduction), mark sentence boundaries, and record pointers from transformed tokens to the original string for cluster presentation.
    2. Base Cluster Identification: Insert all word-level suffixes of all sentences from all documents into a generalized suffix tree. Each internal node vv with at least two distinct document descendants defines a base cluster BvB_v whose shared phrase PvP_v is the concatenated edge label from the root to vv. Each base cluster receives a score s(Bv)=∣Bv∣⋅f(∣Pv∣)s(B_v) = |B_v| \cdot f(|P_v|).
    3. Base Cluster Merging: Construct a base cluster graph over the top kk highest-scoring base clusters. An undirected edge connects two base clusters if their document overlap exceeds a similarity threshold τ=0.5\tau = 0.5 in both directions. The connected components of this graph form the final clusters.
    Input: Document collection D={d1,d2,…,dn}D = \{d_1, d_2, \dots, d_n\}, max base clusters k=500k = 500, overlap threshold τ=0.5\tau = 0.5
    Output: Ranked list of document clusters CC with phrase labels
    1. Initialize generalized suffix tree T←EmptyTree()T \leftarrow \text{EmptyTree}()
    2. For each document di∈Dd_i \in D:
         di′←CleanAndStem(di)d_i' \leftarrow \text{CleanAndStem}(d_i)
         For each sentence S∈di′S \in d_i':
           Insert all word-level suffixes of SS tagged with document ID ii into TT
    3. BaseClusters ←∅\leftarrow \emptyset
    4. For each node v∈Tv \in T:
         Dv←Set of document IDs in subtree beneath vD_v \leftarrow \text{Set of document IDs in subtree beneath } v
         If ∣Dv∣≥2|D_v| \ge 2:
           Pv←Phrase from root to vP_v \leftarrow \text{Phrase from root to } v
           score←∣Dv∣⋅f(∣Pv∣)score \leftarrow |D_v| \cdot f(|P_v|)
           If score>0score > 0:
             Add base cluster Bv=(Dv,Pv,score)B_v = (D_v, P_v, score) to BaseClusters
    5. TopBaseClusters ←Select top k base clusters with highest score from BaseClusters\leftarrow \text{Select top } k \text{ base clusters with highest } score \text{ from } BaseClusters
    6. Initialize graph G=(V,E)G = (V, E) with V=TopBaseClustersV = TopBaseClusters, E=∅E = \emptyset
    7. For each distinct pair (Bm,Bn)∈TopBaseClusters×TopBaseClusters(B_m, B_n) \in TopBaseClusters \times TopBaseClusters:
         If ∣Bm∩Bn∣/∣Bm∣>τ|B_m \cap B_n| / |B_m| > \tau and ∣Bm∩Bn∣/∣Bn∣>τ|B_m \cap B_n| / |B_n| > \tau:
           Add edge (Bm,Bn)(B_m, B_n) to EE
    8. Clusters C←∅C \leftarrow \emptyset
    9. For each connected component KK in GG:
         DK←⋃B∈KB.documentsD_K \leftarrow \bigcup_{B \in K} B.\text{documents}
         PhrasesK←⋃B∈KB.phrasePhrases_K \leftarrow \bigcup_{B \in K} B.\text{phrase}
         ScoreK←∑B∈KB.scoreScore_K \leftarrow \sum_{B \in K} B.score
         Add (DK,PhrasesK,ScoreK)(D_K, Phrases_K, Score_K) to CC
    10. Return CC sorted descending by ScoreKScore_K
  2. Knowl 2 — Base Cluster Scoring Function in STC

    equation

    In Suffix Tree Clustering (STC), each base cluster BB defined by a shared phrase PP across a set of documents is assigned a quality score s(B)s(B):

    s(B)=∣B∣⋅f(∣P∣)s(B) = |B| \cdot f(|P|)

    where:

    • ∣B∣∈N|B| \in \mathbb{N} is the number of documents in base cluster BB (∣B∣≥2|B| \ge 2).
    • ∣P∣∈N|P| \in \mathbb{N} is the effective length of phrase PP, defined as the count of words in PP that have a non-zero score. A word receives a score of zero if it belongs to a stoplist (containing standard stop words as well as Web-specific terms such as "previous", "java", "frames", and "mail"), appears in 3 or fewer documents in the collection, or appears in more than 40% of the documents in the collection.
    • f(∣P∣)f(|P|) is a piece-wise length-weighting function that penalizes single-word phrases (∣P∣=1|P| = 1), increases linearly for phrases containing between 2 and 6 effective words (2≤∣P∣≤62 \le |P| \le 6), and remains constant for phrases of 7 or more effective words (∣P∣≥7|P| \ge 7).
  3. Knowl 3 — Base Cluster Merging Criterion and Cluster Graph

    model/method

    In Suffix Tree Clustering (STC), documents often share multiple distinct phrases, yielding base clusters with overlapping document sets. To merge near-duplicate base clusters without conflating distinct topics, STC defines a binary similarity measure between two base clusters BmB_m and BnB_n based on document overlap:

    Similarity(Bm,Bn)={1if ∣Bm∩Bn∣∣Bm∣>0.5 and ∣Bm∩Bn∣∣Bn∣>0.50otherwise\text{Similarity}(B_m, B_n) = \begin{cases} 1 & \text{if } \frac{|B_m \cap B_n|}{|B_m|} > 0.5 \text{ and } \frac{|B_m \cap B_n|}{|B_n|} > 0.5 \\ 0 & \text{otherwise} \end{cases}

    where ∣Bm∣|B_m| and ∣Bn∣|B_n| denote the document counts of base clusters BmB_m and BnB_n, and ∣Bm∩Bn∣|B_m \cap B_n| is the number of documents common to both.

    A base cluster graph G=(V,E)G = (V, E) is constructed where the vertex set VV consists of the kk highest-scoring base clusters (with k=500k = 500). An undirected edge connects BmB_m and BnB_n if and only if Similarity(Bm,Bn)=1\text{Similarity}(B_m, B_n) = 1. The final document clusters correspond to the connected components of GG. The document set of each final cluster is the union of the document sets of its constituent base clusters, naturally permitting documents to belong to multiple non-identical clusters.

  4. Knowl 4 — Linear Time Complexity and Incrementality of Suffix Tree Clustering

    theoretical result

    Suffix Tree Clustering (STC) achieves an asymptotic time complexity of O(n)O(n), where nn is the number of documents in the collection, under the assumption that document length in words is bounded by a constant:

    1. Document cleaning and word-level stemming require O(n)O(n) operations.
    2. Suffix tree construction using Ukkonen's online algorithm requires O(n)O(n) time and allows documents to be inserted incrementally as they stream over the network.
    3. The number of suffix tree nodes modified or created per incoming document is O(1)O(1).
    4. Base cluster similarity recalculation is bounded to O(1)O(1) per document arrival by checking modified base clusters only against the top kk highest-scoring base clusters (k=500k = 500).

    Because STC processes snippets incrementally during network download latency, clustering finishes almost instantaneously (0.010.01 seconds after receiving the last document in prototype tests).

  5. Knowl 5 — Comparative Retrieval Precision of STC and Baseline Clustering Algorithms

    empirical result

    In an empirical evaluation across 10 Web search queries (each comprising 200 documents and approximately 40 relevant documents), clustering algorithms were evaluated on their ability to reorder search results when a user browses the highest-density cluster until inspecting the top 10% of the collection.

    The resulting average precisions across the 10 collections were:

    • Suffix Tree Clustering (STC): ≈0.38\approx 0.38
    • Group-Average Agglomerative Hierarchical Clustering (GAHC): ≈0.36\approx 0.36
    • Fractionation: ≈0.31\approx 0.31
    • Buckshot: ≈0.28\approx 0.28
    • K-Means: ≈0.25\approx 0.25
    • Single-Pass: ≈0.22\approx 0.22
    • Original search engine ranked list: ≈0.17\approx 0.17

    STC achieved the highest average precision among all evaluated algorithms, more than doubling the precision of the unclustered search engine ranked list.

  6. Knowl 6 — Ablation Study on Multi-Word Phrases and Document Overlap in STC

    empirical result

    An ablation study on 10 Web document collections assessed the individual contributions of multi-word phrases and overlapping document assignments to Suffix Tree Clustering (STC):

    • Full STC (multi-word phrases and overlapping clusters): average precision ≈0.38\approx 0.38.
    • STC-no-overlap (multi-cluster documents forced into a single partition by assignment to the nearest cluster centroid): average precision dropped to ≈0.33\approx 0.33.
    • STC-no-phrases (base clusters restricted strictly to single-word unigrams): average precision dropped to ≈0.22\approx 0.22.

    In contrast, adding suffix-tree-extracted multi-word phrases as additional dimensions in vector-space clustering algorithms did not consistently improve performance (GAHC average precision changed from ≈0.36\approx 0.36 to ≈0.34\approx 0.34, while K-Means changed from ≈0.25\approx 0.25 to ≈0.26\approx 0.26). This demonstrates that phrase extraction and cluster overlap are effective specifically because of the base cluster generation and merging mechanism in STC.

  7. Knowl 7 — Cluster Overlap Rates for Relevant versus Irrelevant Documents

    data/table

    Allowing documents to appear in multiple clusters improves precision only if relevant documents receive higher overlap than irrelevant documents. The average number of cluster memberships per document for relevant and irrelevant documents across 10 Web document collections is summarized below:

    Metric K-Means Buckshot STC
    Avg. clusters per relevant document 1.40 1.40 2.60
    Avg. clusters per irrelevant document 1.55 1.35 1.90
    Ratio (Relevant / Irrelevant) 0.90 1.04 1.37

    STC places relevant documents into substantially more clusters than irrelevant documents (ratio of 1.371.37), whereas K-Means places irrelevant documents into more clusters than relevant ones (ratio of 0.900.90) and Buckshot exhibits almost equal overlap (ratio of 1.041.04). In STC, 72% of documents were placed in more than one cluster, with an overall mean of 2.1 clusters per document.

  8. Knowl 8 — Snippet Clustering versus Full Document Clustering

    empirical result

    Clustering was evaluated on short search engine snippets (averaging 50 words per snippet, 20 words after stop-word and frequency pruning) versus full downloaded HTML documents (averaging 760 words per document, 220 words after pruning). Across all evaluated algorithms, snippet-only clustering resulted in only minor performance drops compared to full-document clustering:

    • STC: average precision decreased from ≈0.38\approx 0.38 (full text) to ≈0.36\approx 0.36 (snippets).
    • GAHC: average precision decreased from ≈0.36\approx 0.36 (full text) to ≈0.34\approx 0.34 (snippets).
    • Fractionation: average precision decreased from ≈0.31\approx 0.31 (full text) to ≈0.30\approx 0.30 (snippets).
    • Buckshot: average precision decreased from ≈0.28\approx 0.28 (full text) to ≈0.27\approx 0.27 (snippets).
    • K-Means: average precision decreased from ≈0.25\approx 0.25 (full text) to ≈0.24\approx 0.24 (snippets).
    • Single-Pass: average precision decreased from ≈0.22\approx 0.22 (full text) to ≈0.21\approx 0.21 (snippets).

    Because snippets extracted by search engines concentrate salient query-related phrases while omitting peripheral off-topic content present in full Web pages, snippet clustering achieves comparable cluster quality without the network overhead of downloading full Web documents.

  9. Knowl 9 — Execution Time Scaling of Document Clustering Algorithms

    empirical result

    Execution times were benchmarked on a Linux machine with a Pentium 200 MHz processor across snippet collections ranging from 100 to 1000 snippets (averaged over 10 collections per size):

    • At 100 snippets: all algorithms finished in under 2 seconds (STC ≈0.3\approx 0.3s, K-Means ≈0.5\approx 0.5s, Single-Pass ≈0.5\approx 0.5s, Buckshot ≈1.2\approx 1.2s, Fractionation ≈1.5\approx 1.5s, GAHC ≈1.8\approx 1.8s).
    • At 500 snippets: STC required ≈1.0\approx 1.0s, K-Means ≈3.5\approx 3.5s, Buckshot ≈4.0\approx 4.0s, Fractionation ≈5.5\approx 5.5s, Single-Pass ≈7.0\approx 7.0s, GAHC ≈11.5\approx 11.5s.
    • At 1000 snippets: STC required ≈2.0\approx 2.0s, K-Means ≈7.0\approx 7.0s, Buckshot ≈7.5\approx 7.5s, Fractionation ≈10.5\approx 10.5s, Single-Pass ≈14.0\approx 14.0s, while GAHC exceeded 18.018.0s due to its O(n2)O(n^2) complexity.

    STC scaled linearly and executed faster than all tested baseline algorithms across all collection sizes.

  10. Knowl 10 — Evaluation Methodology for Post-Retrieval Web Document Clustering

    experimental setup

    To evaluate clustering as an interactive post-retrieval browsing mechanism on the Web:

    1. Ten query topics and descriptions are defined (e.g., topic "black bear attacks").
    2. The MetaCrawler meta-search engine retrieves the top 200 snippets for each query, and the corresponding full HTML documents are downloaded.
    3. Binary relevance judgments (relevant or non-relevant) are manually assigned to all 200 documents per query based on the query description (averaging ≈40\approx 40 relevant documents per collection).
    4. All clustering algorithms are configured to generate 10 clusters to enable controlled comparison.
    5. User browsing behavior is modeled by ranking clusters by their density of relevant documents and accumulating documents starting from the highest-density cluster until exactly 10% of the collection (20 documents) has been selected.
    6. If an overlapping clustering algorithm includes an already-seen document in a subsequent cluster, that re-viewed instance is counted as irrelevant. All remaining 90% uninspected documents in the collection are considered irrelevant when calculating average precision.

Coverage note — None was omitted; all contributed models, algorithms, formal equations, complexity analyses, experimental setups, and empirical findings are represented.

References

  1. 1.R. B. Allen, P. Obry and M. Littman. An interface for navigating clustered document sets returned by queries. In Proceedings of the ACM Conference on Organizational Computing Systems, pages 166-7 1, 1993.
  2. 2.C. Buckley, G. Salton, J. Allen and A. Singhal. Automatic query expansion using SMART: TREC-3. In: D. K. Harman (ed.), The Third Text Retrieval Conference (TREC-3). U.S. Department of Commerce, 1995.
  3. 3.D. R. Cutting, D. R. Karger, J. O. Pedersen and J. W. Tukey. Scatter/Gather: a cluster-based approach to browsing large document collections. In Proceedings of the 15th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 318-29, 1992.
  4. 4.D. R. Cutting, D. R. Karger and J. O. Pedersen. Constant interaction-time Scatter/Gather browsing of large document collections. In Proceedings of the 16th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 126-35, 1993.
  5. 5.W. B. Croft. Organizing and searching large files of documents. Ph.D. Thesis. University of Cambridge, October 1978.
  6. 6.J. L. Fagan. Experiments in automatic phrase indexing for document retrieval: a comparison of syntactic and non-syntactic methods. Ph.D. Thesis, Cornell University, 1987.
  7. 7.D. Gusfield. Algorithms on strings, trees and sequences: computer science and computational biology, chapter 6. Cambridge University Press, 1997.
  8. 8.M. A. Hearst and J. O. Pedersen. Reexamining the cluster hypothesis: Scatter/Gather on retrieval results. In Proceedings of the 19th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 76-84, 1996.
  9. 9.M. A. Hearst. The use of categories and clusters in information access interfaces. In T. Strzalkowski (ed.), Natural Language Information Retrieval, Kluwer Academic Publishers, to appear.
  10. 10.D. R. Hill. A vector clustering technique. In Samuelson (ed.), Mechanised Information Storage, Retrieval and Dissemination, North-Holland, Amsterdam, 1968.
  11. 11.D. A. Hull, G. Grefenstette, B. M. Schulze, E. Gaussier, H. Schütze and L. O. Pedersen. Xerox TREC-5 site report: routing, filtering, NLP, and Spanish tracks. In: D. K. Harman (ed.), The Fifth Text Retrieval Conference (TREC-5). NIST Special Publication, 1997.
  12. 12.A. V. Leouski and W. B. Croft. An evaluation of techniques for clustering search results. Technical Report IR-76, Department of Computer Science, University of Massachusetts, Amherst, 1996.
  13. 13.Y. S. Maarek and A. J. Wecker. The Librarian's Assistant: automatically organizing on-line books into dynamic bookshelves. In Proceedings of RIAO '94, 1994.
  14. 14.G. W. Milligan and M. C. Cooper. An examination of procedures for detecting the number of clusters in a data set. Psychometrika, 50: 159-79, 1985.
  15. 15.E. Rasmussen. Clustering Algorithms. In W. B. Frakes and R. Baeza-Yates (eds.), Information Retrieval, pages 419-42. Prentice Hall, Eaglewood Cliffs, N. J., 1992.
  16. 16.J. J. Rocchio, Document retrieval systems - optimization and evaluation. Ph.D. Thesis, Harvard University, 1966.
  17. 17.G. Salton, C. S. Yang and C. T. Yu. A theory of term importance in automatic text analysis. JASIS, 26( 1):33-44, 1975.
  18. 18.H. Schütze and C. Silverstein. Projections for efficient document clustering. In Proceedings of the 20th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 74-8 1, 1997.
  19. 19.E. Selberg and O. Etzioni. Multi-service search and comparison using the MetaCrawler. In Proceedings of the 4th World Wide Web Conference, 1995.
  20. 20.J. Shakes, M. Langheinrich and O. Etzioni. Ahoy! the home page finder. In Proceedings of the 6th World Wide Web Conference, 1997.
  21. 21.C. Silverstein and J. O. Pedersen. Almost-constant time clustering of arbitrary corpus subsets. In Proceedings of the 20th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 60-66, 1997.
  22. 22.Proceedings of the TDT Workshop, University of Maryland, College Park, MD, October 1997.
  23. 23.E. Ukkonen. On-line construction of suffix trees. Algorithmica, 14:249-60, 1995.
  24. 24.C. J. van Rijsbergen, Information Retrieval, Butterworths, London, 2nd ed., 1979.
  25. 25.E. M. Voorhees. Implementing agglomerative hierarchical clustering algorithms for use in document retrieval. Information Processing and Management, 22:465-76, 1986.
  26. 26.P. Weiner, Linear pattern matching algorithms. In Proceedings of the 14th Annual Symposium on Foundations of Computer Science (FOCS), pages 1-1 1, 1973.
  27. 27.P. Willet. Recent trends in hierarchical document clustering: a critical review. Information Processing and Management, 24:577-97. 1988.
  28. 28.O. Zamir, O. Etzioni, O. Madani and R. M. Karp. Fast and intuitive clustering of Web documents. In Proceedings of the 3rd International Conference on Knowledge Discovery and Data Mining, pages 287-290, 1997.
  29. 29.C. Zhai, X. Tong, N. Milic-Frayling and D. A. Evans. Evaluation of syntactic phrase indexing - CLARIT NLP track report. In: D. K. Harman (ed.), The Fifth Text Retrieval Conference (TREC-5). NIST Special Publication, 1997.

Citation

MLA
Zamir, O., and O. Etzioni. “Web Document Clustering”. Proceedings of the 21st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, 1998, pp. 46–54, https://doi.org/10.1145/290941.290956.
APA
Zamir, O., & Etzioni, O. (1998). Web document clustering. Proceedings of the 21st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, 46–54. https://doi.org/10.1145/290941.290956
Chicago
Zamir, O., and O. Etzioni. 1998. “Web Document Clustering”. Proceedings of the 21st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, 46–54. https://doi.org/10.1145/290941.290956.
Harvard
Zamir, O. and Etzioni, O. (1998) “Web document clustering”, Proceedings of the 21st annual international ACM SIGIR conference on Research and development in information retrieval. ACM, pp. 46–54. Available at: https://doi.org/10.1145/290941.290956.
Vancouver
1. Zamir O, Etzioni O (1998) Web document clustering. In: Proceedings of the 21st annual international ACM SIGIR conference on Research and development in information retrieval. ACM, pp 46–54

BibTeX

@inproceedings{Zamir_1998, series={SIGIR98}, title={Web document clustering: a feasibility demonstration}, url={http://dx.doi.org/10.1145/290941.290956}, DOI={10.1145/290941.290956}, booktitle={Proceedings of the 21st annual international ACM SIGIR conference on Research and development in information retrieval}, publisher={ACM}, author={Zamir, Oren and Etzioni, Oren}, year={1998}, month=Aug, pages={46–54}, collection={SIGIR98} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF