Querying Heterogeneous Information Sources Using Source Descriptions
Alon Y. LevyAnand RajaramanJoann J. Ordille
Presents the Information Manifold system, establishing scalable algorithms that use declarative descriptions of web source contents and query capabilities to prune irrelevant databases and generate executable query plans across hundreds of heterogeneous online sources.
Online platforms host an expanding volume of structured databases covering products, stock markets, entertainment, and enterprise directories. Standard search engines rely on keyword indexing over unstructured text and cannot execute complex, relational queries across structured web forms. Consequently, users must manually locate individual databases, submit separate queries, and stitch the results together by hand. The article evaluates and demonstrates the Information Manifold, a deployed system designed to provide a unified query interface across more than 100 heterogeneous, structured web sources without requiring users to navigate individual databases.
The evaluated approach uses a global relational and object-oriented world view against which users formulate queries. Crucially, external sources are described declaratively as queries over this world view rather than as rigid schemas. The system also introduces capability records that explicitly define the specific input-output parameter requirements and query restrictions of external systems. Using these descriptions, the architecture executes a two-stage query generation process: it first groups relevant databases into target buckets to prune irrelevant sources, and then applies a polynomial-time algorithm to order subgoals into valid, executable execution plans.
The findings show that this approach prevents exponential computational bottlenecks during query planning. Pruning based on declarative descriptions reduced candidate plan evaluations by several orders of magnitude; for instance, in a 100-source scenario where unpruned generation would evaluate over 1,000,000 combinations, the bucket algorithm evaluated only 26. Across empirical tests scaling from 20 to 100 information sources, the average generation time per plan remained below one second. The authors prove that finding an executable order for a plan operates in polynomial time when restricted to single capability records, whereas permitting multiple capability records per source causes the ordering problem to become NP-complete.
These results indicate that enterprises can scale federated query systems across hundreds of online data sources without suffering steep planning latency or rebuilding integration pipelines whenever databases are added or modified. Pipelining plans to stream initial results to users substantially lowers perceived waiting times compared to standard batch execution. For organizations managing distributed structured assets, the article demonstrates that declarative source modeling provides a practical and cost-effective alternative to hand-coded data integration wrappers.
Organizations pursuing large-scale data integration should implement declarative content modeling and explicit input-output capability constraints to streamline multi-source querying. However, decision-makers should note key operational boundaries: the system provides read-only query capabilities and intentionally omits transaction processing or data update mechanisms. Furthermore, source descriptions in the evaluated system were generated manually, and the framework relies on single capability records per source. Future operational initiatives will require automated tooling to generate source descriptions and probabilistic modeling to handle partially relevant sources.
- Paper: KQML as an agent communication language, Tim Finin et al. (1994). Introduces foundational agent communication protocols and mediator architectures for runtime information exchange across heterogeneous distributed sources.
- Paper: Web mining research: a survey, Raymond Kosala et al. (2000). Surveys the evolution of web content and structure mining, providing broader taxonomy for database and information retrieval perspectives on web data integration.
- Paper: Knowledge Graphs, Aidan Hogan et al. (2020). Extends the principles of integrating and querying heterogeneous structured web sources into the modern framework of knowledge graphs and federated graph querying.
