Built independently by an author, for readers. Read the story and support ChapterPal

keyword

capability descriptions

Capability descriptions are declarative specifications that define the query-processing functionalities, supported operations, and access constraints of individual data sources within a heterogeneous data integration system. Because distributed and autonomous sources vary widely in their technical capabilities, ranging from fully featured relational databases to web forms and file repositories with strict input requirements or limited search filters, capability descriptions formally detail what queries or parameter bindings each source can natively handle. Mediators and query planners rely on these descriptions to construct valid execution plans, decomposing complex global queries into subqueries that match the operational limits of each underlying source while scheduling any unsupported filtering, joining, or post-processing tasks at the integration layer.

1 item

Querying Heterogeneous Information Sources Using Source Descriptions

Querying Heterogeneous Information Sources Using Source Descriptions

Alon Y. Levy, Anand Rajaraman, Joann J. Ordille

OrganizationsAT&T Labs—ResearchBell LaboratoriesStanford University

Why you should read this

Presents the Information Manifold system, establishing scalable algorithms that use declarative descriptions of web source contents and query capabilities to prune irrelevant databases and generate executable query plans across hundreds of heterogeneous online sources.

We witness a rapid increase in the number of structured information sources that are available online, especially on the World-Wide Web. These sources store interrelated data on topics such as product information, stock market information, entertainment, etc. We would like to use the data stored in these databases to answer complex queries that go beyond keyword searches. We describe the Information Manifold, an implemented system that provides uniform access to a heterogeneous collection of more than 100 information sources on the WWW. IM contains declarative descriptions of the contents and capabilities of the information sources. We describe algorithms that use the source descriptions to prune efficiently the set of information sources for a given query and practical algorithms to generate executable query plans. We also present experimental studies indicating that the architecture and algorithms used in the Information Manifold scale up well to several hundred information sources.

Added

2026-09-25