COMA - A System for Flexible Combination of Schema Matching Approaches
Hai-Do HongE. Rahm
Introduces a schema matching platform that flexibly combines multiple match algorithms and reuses previous match results to improve matching accuracy across diverse real-world database and XML schemas.
Modern data integration, data warehouse loading, and web service communication depend heavily on schema matching—the process of identifying semantic correspondences between distinct database structures or XML formats. Traditionally, this mapping has been conducted manually by domain experts, creating a severe operational bottleneck that increases development time and costs as organizations manage an ever-growing number of external data sources and interfaces.
The article evaluates whether an automated, flexible platform combining multiple matching algorithms can systematically improve match accuracy and reduce manual integration effort. It also introduces and tests a strategy to reuse previous schema matching results to solve new matching tasks.
To address this, the article develops COMA, a prototype system equipped with an extensible library of simple, hybrid, and reuse-oriented matching algorithms. COMA standardizes schemas into directed graphs, stores intermediate similarity values in a dedicated repository, and executes configurable aggregation and candidate selection strategies. The approach was tested across 12,312 experimental series on five real-world XML purchase order schemas from BizTalk, comprising 10 distinct pairwise matching tasks. Performance was measured using standard precision, recall, and an overall quality metric that reflects post-match manual correction effort by penalizing both false positives and missed correspondences.
The investigation produced four primary findings. First, combining multiple matchers decisively outperformed individual algorithms; combining all hybrid matchers yielded an overall score of 0.73, compared to just 0.45 for the best single non-reuse matcher. Second, reusing prior match results was highly effective, with the manual reuse combination achieving the highest overall score of 0.82 along with over 90% precision. Third, for aggregation and selection strategies, taking the average similarity value across matchers and enforcing mutual agreement between both match directions substantially outperformed optimistic maximum aggregation or one-way selection. Finally, matching difficulty scaled directly with schema size and structural differences; while smaller tasks achieved near-perfect matching scores, larger and more heterogeneous schemas experienced performance degradation, dropping overall scores to around 0.6 to 0.7.
These results demonstrate that organizations can significantly cut manual data integration effort and reduce human error by adopting composite, multi-matcher tools. Relying on single match techniques (such as simple element name comparison) introduces high error rates, whereas composite platforms provide stable, high-precision mappings. Furthermore, building a repository to preserve and reuse verified historical mappings creates a compounding efficiency advantage for ongoing data integration workflows.
Organizations handling complex schema integration should implement composite matching architectures that default to average-based aggregation, bidirectional candidate validation, and historical reuse. Future technical work should expand match capabilities by integrating instance-level data analysis, connecting large-scale standardized ontologies, and testing more sophisticated candidate selection models.
The findings are supported by comprehensive testing on real-world schemas, though the scope is currently limited to XML purchase order structures evaluated in automatic mode. While user feedback loops are supported by the architecture to resolve residual ambiguities in production environments, confidence remains high that composite and reuse-based strategies consistently outperform single-algorithm methods.
- Paper: Generic Schema Matching with Cupid, Jayant Madhavan et al. (2001). This paper introduces Cupid, establishing the foundational hybrid approach that combines linguistic normalization and structural tree matching upon which COMA builds and extends.
- Paper: Decision Combination in Multiple Classifier Systems, Tin Kam Ho et al. (1994). It provides foundational principles and aggregation strategies for combining multiple distinct classification and matching algorithms into a unified decision framework.
- Paper: Automatic Retrieval and Clustering of Similar Words, Dekang Lin (1998). It details information-theoretic word similarity measures that serve as key building blocks for lexical schema matching components.
- Paper: Duplicate Record Detection: A Survey, Ahmed K. Elmagarmid et al. (2007). This comprehensive survey expands on data integration concepts, evaluating how matching and entity resolution techniques are applied to detect duplicate records across heterogeneous databases.
- Paper: WordNet::Similarity - Measuring the Relatedness of Concepts, Ted Pedersen et al. (2004). It formalizes and expands software tools for computing taxonomic and lexical relatedness in WordNet, which can serve as pluggable similarity modules in composite matchers like COMA.
