COMA - A System for Flexible Combination of Schema Matching Approaches

Hai-Do HongE. Rahm

article2002VLDB1,348 citations

Introduces a schema matching platform that flexibly combines multiple match algorithms and reuses previous match results to improve matching accuracy across diverse real-world database and XML schemas.

Listen

Modern data integration, data warehouse loading, and web service communication depend heavily on schema matching—the process of identifying semantic correspondences between distinct database structures or XML formats. Traditionally, this mapping has been conducted manually by domain experts, creating a severe operational bottleneck that increases development time and costs as organizations manage an ever-growing number of external data sources and interfaces.

The article evaluates whether an automated, flexible platform combining multiple matching algorithms can systematically improve match accuracy and reduce manual integration effort. It also introduces and tests a strategy to reuse previous schema matching results to solve new matching tasks.

To address this, the article develops COMA, a prototype system equipped with an extensible library of simple, hybrid, and reuse-oriented matching algorithms. COMA standardizes schemas into directed graphs, stores intermediate similarity values in a dedicated repository, and executes configurable aggregation and candidate selection strategies. The approach was tested across 12,312 experimental series on five real-world XML purchase order schemas from BizTalk, comprising 10 distinct pairwise matching tasks. Performance was measured using standard precision, recall, and an overall quality metric that reflects post-match manual correction effort by penalizing both false positives and missed correspondences.

The investigation produced four primary findings. First, combining multiple matchers decisively outperformed individual algorithms; combining all hybrid matchers yielded an overall score of 0.73, compared to just 0.45 for the best single non-reuse matcher. Second, reusing prior match results was highly effective, with the manual reuse combination achieving the highest overall score of 0.82 along with over 90% precision. Third, for aggregation and selection strategies, taking the average similarity value across matchers and enforcing mutual agreement between both match directions substantially outperformed optimistic maximum aggregation or one-way selection. Finally, matching difficulty scaled directly with schema size and structural differences; while smaller tasks achieved near-perfect matching scores, larger and more heterogeneous schemas experienced performance degradation, dropping overall scores to around 0.6 to 0.7.

These results demonstrate that organizations can significantly cut manual data integration effort and reduce human error by adopting composite, multi-matcher tools. Relying on single match techniques (such as simple element name comparison) introduces high error rates, whereas composite platforms provide stable, high-precision mappings. Furthermore, building a repository to preserve and reuse verified historical mappings creates a compounding efficiency advantage for ongoing data integration workflows.

Organizations handling complex schema integration should implement composite matching architectures that default to average-based aggregation, bidirectional candidate validation, and historical reuse. Future technical work should expand match capabilities by integrating instance-level data analysis, connecting large-scale standardized ontologies, and testing more sophisticated candidate selection models.

The findings are supported by comprehensive testing on real-world schemas, though the scope is currently limited to XML purchase order structures evaluated in automatic mode. While user feedback loops are supported by the architecture to resolve residual ambiguities in production environments, confidence remains high that composite and reuse-based strategies consistently outperform single-algorithm methods.

  • Paper: Generic Schema Matching with Cupid, Jayant Madhavan et al. (2001). This paper introduces Cupid, establishing the foundational hybrid approach that combines linguistic normalization and structural tree matching upon which COMA builds and extends.
  • Paper: Decision Combination in Multiple Classifier Systems, Tin Kam Ho et al. (1994). It provides foundational principles and aggregation strategies for combining multiple distinct classification and matching algorithms into a unified decision framework.
  • Paper: Automatic Retrieval and Clustering of Similar Words, Dekang Lin (1998). It details information-theoretic word similarity measures that serve as key building blocks for lexical schema matching components.
  • Paper: Duplicate Record Detection: A Survey, Ahmed K. Elmagarmid et al. (2007). This comprehensive survey expands on data integration concepts, evaluating how matching and entity resolution techniques are applied to detect duplicate records across heterogeneous databases.
  • Paper: WordNet::Similarity - Measuring the Relatedness of Concepts, Ted Pedersen et al. (2004). It formalizes and expands software tools for computing taxonomic and lexical relatedness in WordNet, which can serve as pluggable similarity modules in composite matchers like COMA.
Cover for COMA - A System for Flexible Combination of Schema Matching Approaches

Abstract

Schema matching is the task of finding semantic correspondences between elements of two schemas. It is needed in many database applications, such as integration of web data sources, data warehouse loading and XML message mapping. To reduce the amount of user effort as much as possible, automatic approaches combining several match techniques are required. While such match approaches have found considerable interest recently, the problem of how to best combine different match algorithms still requires further work. We have thus developed the COMA schema matching system as a platform to combine multiple matchers in a flexible way. We provide a large spectrum of individual matchers, in particular a novel approach aiming at reusing results from previous match operations, and several mechanisms to combine the results of matcher executions. We use COMA as a framework to comprehensively evaluate the effectiveness of different matchers and their combinations for real-world schemas. The results obtained so far show the superiority of combined match approaches and indicate the high value of reuse-oriented strategies.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Overview of COMA
  • 4 Matcher library
  • 4.1 Simple matchers
  • 4.2 Hybrid matchers
  • 5 Reuse of previous match results
  • 5.1 The MatchCompose operation
  • 5.2 The Schema reuse matcher
  • 6 Combination of similarity values
  • 6.1 Aggregation of matcher-specific results
  • 6.2 Direction and selection of match candidates
  • 6.3 Computation of combined similarity
  • 6.4 Construction of hybrid matchers
  • 7 Evaluation on real world schemas
  • 7.1 Experimental design
  • 7.2 Combination strategies
  • 7.3 Single matchers and matcher combinations
  • 7.4 Match sensitivity
  • 7.5 Evaluation conclusions
  • 8 Summary and future work
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — COMA Schema Representation and Similarity Cube Architecture

    model/method

    In the COMA (COmbination of MAtching algorithms) schema matching architecture, schemas are parsed from external representations (such as XML DTDs/schemas or relational DBMS catalogs) into rooted directed acyclic graphs (DAGs). Graph nodes represent schema elements (tables, columns, complex XML elements, attributes), and directed edges represent structural containment or referential relationships. Schema elements are uniquely identified by their paths, which are sequences of nodes from the root to the target element, enabling distinct contextual representation for shared schema fragments.

    Matching two schemas S1S_1 (with mm elements) and S2S_2 (with nn elements) using an ensemble of kk distinct matchers produces an intermediate data structure termed the similarity cube, C∈[0,1]k×m×nC \in [0, 1]^{k \times m \times n}. Each entry C(i,u,v)C(i, u, v) stores the similarity score between element u∈S1u \in S_1 and element v∈S2v \in S_2 as calculated by matcher i∈{1,…,k}i \in \{1, \dots, k\}.

    COMA supports both single-pass automatic execution and iterative interactive execution. In interactive mode, a special UserFeedback matcher captures user-specified matches (assigned a similarity of 1.01.0) and mismatches (assigned 0.00.0). These user-fixed values are preserved throughout downstream matcher executions and propagate constraint information to structural matcher computations in the neighborhoods of those nodes.

  2. Knowl 2 — Three-Step Matcher Combination Framework

    algorithm

    COMA combines independent matcher predictions stored in a similarity cube C∈[0,1]k×m×nC \in [0, 1]^{k \times m \times n} into a final schema mapping via a three-step aggregation and selection pipeline:

    Input: Similarity cube C∈[0,1]k×m×nC \in [0, 1]^{k \times m \times n} for schemas S1S_1 (mm elements) and S2S_2 (nn elements), aggregation operator Agg∈{Max,Min,Average,Weighted}Agg \in \{\text{Max}, \text{Min}, \text{Average}, \text{Weighted}\}, match direction Dir∈{LargeSmall,SmallLarge,Both}Dir \in \{\text{LargeSmall}, \text{SmallLarge}, \text{Both}\}, selection criteria SelSel, combined similarity metric CombSim∈{Average,Dice}∪{None}CombSim \in \{\text{Average}, \text{Dice}\}\cup\{\text{None}\}
    Output: Final mapping Mfinal⊆S1×S2×[0,1]M_{final} \subseteq S_1 \times S_2 \times [0, 1] and optional schema similarity sschema∈[0,1]s_{schema} \in [0, 1]
    // Step 1: Aggregation of matcher-specific results into matrix Smat∈[0,1]m×nS_{mat} \in [0, 1]^{m \times n}
    for each element u∈S1u \in S_1 do
        for each element v∈S2v \in S_2 do
            if Agg==MaxAgg == \text{Max} then
                Smat(u,v)←max⁡1≤i≤kC(i,u,v)S_{mat}(u, v) \leftarrow \max_{1 \le i \le k} C(i, u, v)
            else if Agg==MinAgg == \text{Min} then
                Smat(u,v)←min⁡1≤i≤kC(i,u,v)S_{mat}(u, v) \leftarrow \min_{1 \le i \le k} C(i, u, v)
            else if Agg==AverageAgg == \text{Average} then
                Smat(u,v)←1k∑i=1kC(i,u,v)S_{mat}(u, v) \leftarrow \frac{1}{k} \sum_{i=1}^k C(i, u, v)
            else if Agg==WeightedAgg == \text{Weighted} with weights w1,…,wkw_1, \dots, w_k where ∑wi=1\sum w_i = 1 then
                Smat(u,v)←∑i=1kwi⋅C(i,u,v)S_{mat}(u, v) \leftarrow \sum_{i=1}^k w_i \cdot C(i, u, v)
            end if
        end for
    end for
    // Step 2: Directional filtering and candidate selection
    Cand12←∅Cand_{12} \leftarrow \emptyset
    for each u∈S1u \in S_1 do
        Rank elements v∈S2v \in S_2 descending by Smat(u,v)S_{mat}(u, v)
        Cand12←Cand12∪{(u,v,Smat(u,v))∣v satisfies Sel relative to u}Cand_{12} \leftarrow Cand_{12} \cup \{ (u, v, S_{mat}(u, v)) \mid v \text{ satisfies } Sel \text{ relative to } u \}
    end for
    Cand21←∅Cand_{21} \leftarrow \emptyset
    for each v∈S2v \in S_2 do
        Rank elements u∈S1u \in S_1 descending by Smat(u,v)S_{mat}(u, v)
        Cand21←Cand21∪{(u,v,Smat(u,v))∣u satisfies Sel relative to v}Cand_{21} \leftarrow Cand_{21} \cup \{ (u, v, S_{mat}(u, v)) \mid u \text{ satisfies } Sel \text{ relative to } v \}
    end for
    if Dir==LargeSmallDir == \text{LargeSmall} and ∣S1∣≥∣S2∣|S_1| \ge |S_2| then
        Mfinal←Cand21M_{final} \leftarrow Cand_{21}
    else if Dir==SmallLargeDir == \text{SmallLarge} and ∣S1∣≥∣S2∣|S_1| \ge |S_2| then
        Mfinal←Cand12M_{final} \leftarrow Cand_{12}
    else if Dir==BothDir == \text{Both} then
        Mfinal←{(u,v,s)∈Cand12∣∃s′:(u,v,s′)∈Cand21}M_{final} \leftarrow \{ (u, v, s) \in Cand_{12} \mid \exists s' : (u, v, s') \in Cand_{21} \}
    end if
    // Step 3: Computation of combined similarity (optional for schema-level mapping, required for hybrid matchers)
    if CombSim≠NoneCombSim \ne \text{None} then
        sschema←CombSim(S1,S2,Mfinal)s_{schema} \leftarrow CombSim(S_1, S_2, M_{final})
    end if
    return Mfinal,sschemaM_{final}, s_{schema}

    The candidate selection filter SelSel supports:

    1. MaxN(n)\text{MaxN}(n): retains the top nn candidates per element.
    2. MaxDelta(d)\text{MaxDelta}(d): retains the top candidate with maximum similarity smaxs_{max} plus any candidate with similarity s≥smax−ds \ge s_{max} - d (or relative ratio (smax−s)/smax≤d(s_{max} - s)/s_{max} \le d).
    3. Threshold(t)\text{Threshold}(t): retains all candidates with similarity s≥ts \ge t.
    4. Hybrid filter Threshold(t)+MaxDelta(d)\text{Threshold}(t) + \text{MaxDelta}(d): applies MaxDelta(d)\text{MaxDelta}(d) restricted to candidates satisfying s≥ts \ge t.
  3. Knowl 3 — MatchCompose Operation and Schema Reuse Matcher

    model/method

    The MatchCompose operation derives a new mapping between two schemas S1S_1 and S3S_3 by chaining previously computed, confirmed mappings that share an intermediate schema S2S_2. Given mapping tables match1:S1↔S2match_1: S_1 \leftrightarrow S_2 with similarities sim12(u,v)sim_{12}(u, v) and match2:S2↔S3match_2: S_2 \leftrightarrow S_3 with similarities sim23(v,w)sim_{23}(v, w), MatchCompose computes the relational natural join on the shared elements of S2S_2.

    To prevent aggressive score degradation associated with multiplicative chaining (sim12⋅sim23sim_{12} \cdot sim_{23}), MatchCompose computes the transitive correspondence similarity using the arithmetic mean: sim13(u,w)=sim12(u,v)+sim23(v,w)2\text{sim}_{13}(u, w) = \frac{\text{sim}_{12}(u, v) + \text{sim}_{23}(v, w)}{2}

    The Schema reuse matcher operationalizes this composition over a repository of historical match results:

    1. Given a target match task between schemas S1S_1 and S2S_2, it queries the repository for all schemas SiS_i for which both mappings S1↔SiS_1 \leftrightarrow S_i and S2↔SiS_2 \leftrightarrow S_i exist.
    2. For each intermediate schema SiS_i, MatchCompose is executed to generate candidate mapping S1↔S2S_1 \leftrightarrow S_2.
    3. If multiple intermediate schemas yield candidate mappings, the resulting similarity matrices are aggregated (e.g., using Average) and filtered to produce a single similarity matrix for the Schema matcher.
  4. Knowl 4 — Construction of Hybrid Matchers

    model/method

    COMA provides five hybrid matchers constructed by composing simple element and structural matchers through the combination pipeline:

    1. Name: Tokenizes element names into constituent word tokens and expands common abbreviations/acronyms. It executes simple string matchers (e.g., Trigram) and semantic dictionary matchers (Synonym) over all pairs of tokens from the two element names, aggregates matcher scores using Max, applies bidirectional candidate selection (Both, Max1), and computes the final name similarity as the Average of the selected token similarity values.
    2. TypeName: Combines data type compatibility (DataType, derived from a predefined data type compatibility lookup table) and element name similarity (Name) using weighted aggregation: sim=0.7⋅simName+0.3⋅simDataType\text{sim} = 0.7 \cdot \text{sim}_{\text{Name}} + 0.3 \cdot \text{sim}_{\text{DataType}}.
    3. NamePath: Concatenates all element names along the path from the root node to the target node into a single hierarchical name string and evaluates similarity between path strings using the Name matcher.
    4. Children: Evaluates the structural similarity between two inner schema nodes by comparing the sets of their immediate child elements using a leaf-level matcher (defaulting to TypeName), filtering candidates via Both, Max1, and computing set similarity using Average.
    5. Leaves: Evaluates the structural similarity between two inner schema nodes by comparing the sets of all leaf elements in their respective subtrees using TypeName as the leaf matcher, filtering candidates via Both, Max1, and computing set similarity using Average.
  5. Knowl 5 — Overall Schema Match Quality Metric

    equation

    To evaluate the quality of an automatic schema matching operation against a gold standard of manually determined real matches RR, the returned match set PP is partitioned into true positives I=P∩RI = P \cap R, false positives F=P∖IF = P \setminus I, and false negatives (missed matches) M=R∖IM = R \setminus I. The performance is measured using Precision, Recall, and Overall:

    Precision=∣I∣∣P∣=∣I∣∣I∣+∣F∣\text{Precision} = \frac{|I|}{|P|} = \frac{|I|}{|I| + |F|}

    Recall=∣I∣∣R∣\text{Recall} = \frac{|I|}{|R|}

    Overall=1−∣F∣+∣M∣∣R∣=∣I∣−∣F∣∣R∣=Recall⋅(2−1Precision)\text{Overall} = 1 - \frac{|F| + |M|}{|R|} = \frac{|I| - |F|}{|R|} = \text{Recall} \cdot \left(2 - \frac{1}{\text{Precision}}\right)

    where Precision,Recall∈[0,1]\text{Precision}, \text{Recall} \in [0, 1] and Overall∈(−∞,1]\text{Overall} \in (-\infty, 1].

    Overall\text{Overall} quantifies the post-match manual effort required to correct the result by adding missed matches and removing false matches. If Precision<0.5\text{Precision} < 0.5, then ∣F∣>∣I∣|F| > |I|, yielding Overall<0\text{Overall} < 0, which signifies that the manual effort required to clean up incorrect matches outweighs the utility of the automated match prediction. Overall=1\text{Overall} = 1 if and only if P=RP = R.

  6. Knowl 6 — Component Set Combined Similarity Metrics

    equation

    When evaluating the similarity between two sets of schema element components S1S_1 and S2S_2 (such as name tokens, child nodes, or leaf nodes), COMA computes a combined similarity score from the 1:1 match candidates identified between elements of S1S_1 and S2S_2. Assuming at most one match candidate per element in each set, the combined similarity is computed using either Average or Dice:

    Average(S1,S2)=∑(u,v)∈M12s(u,v)+∑(v,u)∈M21s(v,u)∣S1∣+∣S2∣\text{Average}(S_1, S_2) = \frac{\sum_{(u, v) \in M_{12}} s(u, v) + \sum_{(v, u) \in M_{21}} s(v, u)}{|S_1| + |S_2|}

    Dice(S1,S2)=∣M12∣+∣M21∣∣S1∣+∣S2∣\text{Dice}(S_1, S_2) = \frac{|M_{12}| + |M_{21}|}{|S_1| + |S_2|}

    where M12⊆S1×S2M_{12} \subseteq S_1 \times S_2 is the set of selected directional matches from S1S_1 to S2S_2, M21⊆S2×S1M_{21} \subseteq S_2 \times S_1 is the set of selected directional matches from S2S_2 to S1S_1, and s(u,v)∈[0,1]s(u, v) \in [0, 1] is the similarity value assigned to correspondence (u,v)(u, v).

    Dice considers only the proportion of matched elements without weighting by individual correspondence similarities, making it more optimistic than Average. When all element correspondence similarities equal 1.01.0, Average(S1,S2)=Dice(S1,S2)\text{Average}(S_1, S_2) = \text{Dice}(S_1, S_2).

  7. Knowl 7 — Empirical Effectiveness of Matcher Combination Strategies

    empirical result

    An exhaustive evaluation over 12,312 configuration series across 10 XML purchase order schema matching tasks established the comparative effectiveness of different combination strategies:

    1. Aggregation: Average consistently outperformed Max and Min. Max is overly optimistic and severely affected by false positives generated by constituent matchers; all series employing Max aggregation yielded average Overall<0.1\text{Overall} < 0.1. Min and Average were resilient to inaccurate constituent matchers, with Average being the only aggregation strategy producing series with average Overall>0.6\text{Overall} > 0.6.
    2. Match Direction: The bidirectional Both strategy significantly outperformed unidirectional matching (LargeSmall and SmallLarge). SmallLarge (matching smaller source schemas against larger target schemas) produced poor results (average Overall≤0.3\text{Overall} \le 0.3) due to abundant false positives, while LargeSmall reached Overall≤0.6\text{Overall} \le 0.6. Both enforces mutual candidate agreement and achieved the highest match quality across all schema sizes.
    3. Candidate Selection: Standard Threshold filtering performed worst among selection methods; even the best variant, Threshold(0.8)\text{Threshold}(0.8), had average Overall<0.3\text{Overall} < 0.3. MaxN(1)\text{MaxN}(1) and Threshold(0.5)+MaxN(1)\text{Threshold}(0.5) + \text{MaxN}(1) achieved average Overall\text{Overall} up to 0.700.70. The most effective selection strategies were Delta(0.02)\text{Delta}(0.02) and Threshold(0.5)+Delta(0.02)\text{Threshold}(0.5) + \text{Delta}(0.02), which exceeded 0.700.70 Overall\text{Overall}.
    4. Set Similarity: Average outperformed Dice for computing component set similarity in hybrid matchers, achieving a peak average Overall\text{Overall} of 0.730.73 compared to 0.670.67 for Dice.

    The optimal default combination strategy identified is (Average,Both,Threshold(0.5)+Delta(0.02))(\text{Average}, \text{Both}, \text{Threshold}(0.5) + \text{Delta}(0.02)) with Average set similarity.

  8. Knowl 8 — Empirical Performance of Standalone Matchers versus Matcher Combinations

    empirical result

    Evaluation on real-world XML purchase order schemas demonstrated the superiority of composite matchers and reuse strategies over standalone matchers:

    1. Standalone No-Reuse Matchers: Single hybrid matchers that ignore path context (Name, TypeName, Children, Leaves) suffered from schema redundancy (shared sub-elements), yielding high false positive rates and negative average Overall\text{Overall} scores across several tasks. NamePath was the most effective no-reuse single matcher (average Overall=0.45\text{Overall} = 0.45), as hierarchical path names effectively disambiguated shared element contexts at the cost of slightly reduced Recall.
    2. Standalone Reuse Matchers: The schema reuse matchers substantially outperformed all no-reuse single matchers. Standalone reuse of automatically generated match results (SchemaA) achieved an average Overall\text{Overall} of 0.620.62, while reuse of manually confirmed match results (SchemaM) reached an average Overall\text{Overall} of 0.730.73, achieving Precision of 0.850.85 and 0.880.88, respectively.
    3. Matcher Combinations: Combining all five hybrid matchers without reuse (All) achieved an average Overall\text{Overall} of 0.730.73. Combining NamePath and Leaves achieved an average Overall\text{Overall} improvement of 20%20\% over NamePath alone and 80%80\% over Leaves alone. The reuse combination All + SchemaM achieved the highest performance across all experiments, with an average Overall\text{Overall} of 0.820.82, average Precision>0.90\text{Precision} > 0.90, and average Recall=0.89\text{Recall} = 0.89, outperforming the best no-reuse approach by more than 10%10\%.
  9. Knowl 9 — Characteristics of the Purchase Order XML Schema Matching Benchmark

    data/table

    The empirical evaluation of COMA utilized 5 real-world purchase order XML schemas from BizTalk.org: CIDX (Schema 1), Excel (Schema 2), Noris (Schema 3), Paragon (Schema 4), and Apertum (Schema 5). The 10 pairwise matching tasks formed a benchmark characterized by shared structures (indicated where path count exceeds node count) and moderate schema overlap (Dice similarity of matched paths ≈0.5\approx 0.5 across tasks).

    Characteristic Schema 1 Schema 2 Schema 3 Schema 4 Schema 5
    Max depth 4 4 4 6 5
    #Nodes / paths 40 / 40 35 / 54 46 / 65 74 / 80 80 / 145
    #Inner nodes / paths 7 / 7 9 / 12 8 / 11 11 / 12 23 / 29
    #Leaf nodes / paths 33 / 33 26 / 42 38 / 54 63 / 68 57 / 116

    The benchmark demonstrates varying structural scale, ranging from 40 elements up to 145 paths, allowing assessment of match candidate search spaces under realistic schema heterogeneity.

  10. Knowl 10 — Impact of Schema Scale and Heterogeneity on Match Quality and Stability

    empirical result

    Across 10 pairwise matching tasks, automatic schema matching effectiveness correlated inversely with schema size and positively with schema similarity:

    1. Schema Scale and Similarity Degradation: For small schema pairs with high structural correspondence (e.g., matching Schema 1 against Schema 2, with 40 to 54 paths), both no-reuse and manual reuse approaches achieved near-optimal Overall\text{Overall} scores approaching 1.01.0. However, as schema size expanded to 80–145 paths (e.g., Schema 4 against Schema 5) and pairwise schema similarity dropped towards 0.50.5, maximum attainable Overall\text{Overall} degraded to 0.6−0.70.6 - 0.7.
    2. Matcher Stability: The composite strategies All (no-reuse) and All + SchemaM (reuse) demonstrated superior stability across varying schema difficulties. Out of the 10 matching tasks, All and All + SchemaM yielded the maximal Overall\text{Overall} in 5 and 6 tasks, respectively, while exhibiting minor degradation of at most 10%10\% relative to the task-specific best matcher in the remaining tasks.

Coverage note — None was omitted; all key contributions—including the COMA architecture, matcher library, reuse mechanism (MatchCompose and Schema), combination pipeline, evaluation metrics, and comprehensive experimental results—are fully represented.

References

  1. 1.Berlin, J., A. Motro: Autoplex: Automated Discovery of Content for Virtual Databases. CoopIS 2001, 108-122
  2. 2.Bergamaschi, S., S. Castano, M. Vincini, D. Beneventano: Semantic Integration of Heterogeneous Information Sources. Data & Knowledge Engineering 36: 3, 215–249, 2001
  3. 3.Bernstein, P.A., A. Halevy, R. A. Pottinger: A Vision for Management of Complex Models. SIGMOD Record 29: 4, 55–63, 2000
  4. 4.Bright, M.W. et al: Automated Resolution of Semantic Heterogeneity in Multidatabase. ACM Trans. Database Systems 19: 2, 1994
  5. 5.Castano, S., V. De Antonellis: A Schema Analysis and Reconciliation Tool Environment. IDEAS 1999, 53-62
  6. 6.Castano, S, V. De Antonellis, M.G. Fugini, B. Pernici: Conceptual Schema Analysis: Techniques and Applications. ACM Trans. Database Systems 23: 3, 286-333, 1998
  7. 7.Doan, A.H., P. Domingos, A. Halevy: Reconciling Schemas of Disparate Data Sources: A Machine-Learning Approach. SIGMOD 2001
  8. 8.Doan, A.H., J. Madhavan, P. Domingos, A. Halevy: Learning to Map between Ontologies on the Semantic Web. WWW 2002
  9. 9.Embley, D.W. et al.: Multifaceted Exploitation of Metadata for Attribute Match Discovery in Information Integration. WIIW 2001
  10. 10.Hall, P., G. Dowling: Approximate String Matching. Computing Survey 12: 4, 381-402, 1980
  11. 11.Li, W., C. Clifton: SemInt: A Tool for Identifying Attribute Correspondences in Heterogeneous Databases Using Neural Network. Data and Knowledge Engineering 33: 1, 49-84, 2000
  12. 12.Madhavan, J., P.A. Bernstein, E. Rahm: Generic Schema Matching with Cupid. VLDB 2001
  13. 13.Melnik, S., H. Garcia-Molina, E. Rahm: Similarity Flooding: A Versatile Graph Matching Algorithm. ICDE 2002
  14. 14.Miller, R.J. et al.: The Clio Project: Managing Heterogeneity. SIGMOD Record 30:1, 78-83, 2001
  15. 15.Milo, T., S. Zohar: Using Schema Matching to Simplify Heterogeneous Data Translation. VLDB 1998, 122-133
  16. 16.Palopoli, L., G. Terracina, D. Ursino: The System DIKE: Towards the Semi-Automatic Synthesis of Cooperative Information Systems and Data Warehouses. ADBIS-DASFAA 2000, 108–117
  17. 17.Rada, R. et al.: Development and Application of a Metric on Semantic Nets. IEEE Trans. Systems, Man, and Cybernetics 19: 1, 1989
  18. 18.Rahm, E., P.A. Bernstein: A Survey of Approaches to Automatic Schema Matching. VLDB Journal 10: 4, 2001
  19. 19.Rahm, E., Do, H.H.: Data Cleaning: Problems and Current Approaches. IEEE Bulletin on Data Engineering 23:4, 2000
  20. 20.Winkler, W.E.: Advanced Methods for Record Linking. Section on Survey Research Methods (American Statistical Association), 1994

Citation

MLA
Do, H.-H., and E. Rahm. “COMA — A System for Flexible Combination of Schema Matching Approaches”. VLDB '02: Proceedings of the 28th International Conference on Very Large Databases, Elsevier, 2002, pp. 610–21, https://doi.org/10.1016/B978-155860869-6/50060-3.
APA
Do, H.-H., & Rahm, E. (2002). COMA — A system for flexible combination of schema matching approaches. In VLDB '02: Proceedings of the 28th International Conference on Very Large Databases (pp. 610–621). Elsevier. https://doi.org/10.1016/B978-155860869-6/50060-3
Chicago
Do, H.-H., and E. Rahm. 2002. “COMA — A System for Flexible Combination of Schema Matching Approaches”. In VLDB '02: Proceedings of the 28th International Conference on Very Large Databases. Elsevier. https://doi.org/10.1016/B978-155860869-6/50060-3.
Harvard
Do, H.-H. and Rahm, E. (2002) “COMA — A system for flexible combination of schema matching approaches”, VLDB '02: Proceedings of the 28th International Conference on Very Large Databases. Elsevier, pp. 610–621. Available at: https://doi.org/10.1016/B978-155860869-6/50060-3.
Vancouver
1. Do H-H, Rahm E (2002) COMA — A system for flexible combination of schema matching approaches. In: VLDB '02: Proceedings of the 28th International Conference on Very Large Databases. Elsevier, pp 610–621

BibTeX

@inbook{Do_2002, title={COMA — A system for flexible combination of schema matching approaches}, ISBN={9781558608696}, url={http://dx.doi.org/10.1016/B978-155860869-6/50060-3}, DOI={10.1016/b978-155860869-6/50060-3}, booktitle={VLDB ’02: Proceedings of the 28th International Conference on Very Large Databases}, publisher={Elsevier}, author={Do, Hong-Hai and Rahm, Erhard}, year={2002}, pages={610–621} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF