Scatter/Gather: a cluster-based approach to browsing large document collections
Douglass R. CuttingJan O. PedersenDavid R KargerJohn W. Tukey
Introduces Scatter/Gather, an interactive information access paradigm paired with linear-time clustering algorithms that allows users to explore large document collections through iterative topic-based summarization rather than traditional keyword queries.
Traditional information retrieval systems rely heavily on keyword queries, which assume users know precisely what terms to search for. When users have vague information goals or lack familiarity with the vocabulary of a large text collection, standard keyword searches often fail. Historical attempts to apply document grouping to search were hindered by slow processing speeds and minimal retrieval improvements when forced into traditional search frameworks.
The article demonstrates an interactive browsing technique called Scatter/Gather and evaluates fast grouping procedures that make dynamic exploration feasible. Rather than using automated grouping merely to accelerate standard keyword searches, the authors propose using it as an intuitive, dynamic table-of-contents metaphor for navigating large text collections.
To evaluate this framework, the authors developed two linear-time grouping procedures—one tailored for fast interactive updates and another for more accurate initial corpus organization—and tested them on an experimental dataset of approximately 5,000 news articles. The interactive interface automatically generates concise summaries for each cluster using representative titles and high-frequency terms, allowing users to select subsets of interest and dynamically re-cluster them.
The investigation produced several key findings. First, localizing grouping operations to small sample sets or bounded partitions reduces computational processing time from quadratic to linear time, allowing interactive responses in seconds. Second, the faster sampling-based grouping procedure reliably identifies natural topic groups with high probability when document collections contain distinct subject areas. Third, dynamic re-clustering enables users to quickly isolate broad themes and drill down into obscure, localized stories that would otherwise remain hidden under broader subjects without requiring specific search keywords.
These findings indicate that content grouping can serve as an effective standalone discovery tool, reducing the cognitive effort and time needed to explore unfamiliar document archives. Instead of replacing standard keyword search, this browsing model complements focused lookup tools by helping users understand the thematic structure of a corpus and formulate better targeted queries.
For future implementation, the authors recommend combining offline structural preprocessing with fast sampling-based re-clustering during live user sessions. Because the interactive sampling procedure is non-deterministic and the browsing experience is exploratory, rigorous quantitative evaluation remains challenging until standardized metrics for open-ended browsing tasks are developed.
No sufficiently relevant recommendations were found.
- Paper: Web document clustering: a feasibility demonstration, Oren Zamir et al. (1998). See how the cluster-browsing idea moves from document archives to live web search, where Suffix Tree Clustering organizes retrieved results for exploration.
