Generic Schema Matching with Cupid

Jayant MadhavanP. BernsteinE. Rahm

article2001VLDB1,657 citations10-Year Best Paper Award

Proposes Cupid, a generic schema matching algorithm that combines linguistic analysis with tree-based structural matching to accurately discover correspondences across diverse data models and applications.

Listen

Modern data management frequently requires integrating disparate data sources, translating messages between systems, and consolidating databases into centralized data warehouses. A central bottleneck in these operations is schema matching, the process of identifying corresponding elements across different data structures. Because variations in naming conventions, nested formats, and structural designs are common across independently designed systems, schema matching is largely performed through labor-intensive and error-prone manual effort. Automatic tools often fail when encountering subtle structural or vocabulary differences, driving the need for a generic, standalone matching framework.

The main objective of the article is to introduce and evaluate Cupid, a generic schema-matching algorithm that discovers correspondences across diverse data models independently of specific applications. The article demonstrates that combining automated linguistic normalization with bottom-up structural analysis significantly improves match accuracy across disparate relational and semi-structured schemas.

To evaluate this approach, the authors developed Cupid and compared its performance against two established schema-integration systems, DIKE and MOMIS, using six isolated canonical test cases and two complex real-world scenarios. The real-world evaluations involved aligning disparate electronic business purchase orders in XML format and mapping a normalized relational database schema to a dimensional data warehouse star schema. The algorithm executes across three primary phases: normalizing and categorizing element names using thesauri to compute linguistic similarity; calculating structural similarity via a bottom-up tree traversal that strongly weights atomic leaf elements and contextual constraints; and generating weighted mappings based on these combined scores.

The findings show that Cupid successfully identifies correct element mappings across varied and complex structures where existing tools struggle. First, Cupid proved to be the only evaluated system capable of automatically generating context-dependent mappings, correctly distinguishing identical sub-structures (such as shared address formats) based on their specific parent contexts. Second, the bottom-up structural approach biased toward leaf elements enabled Cupid to accurately map schemas with substantial differences in nesting and hierarchy, whereas baseline systems required manual interventions or produced fractured clusters. Third, automated linguistic normalization—including tokenization and expansion—successfully mapped elements despite significant naming variations without requiring manual dictionary entries for each variation. In the real-world XML evaluation, Cupid detected 100% of the correct attribute matches, outperforming baseline systems that required significant schema remodeling or produced false cluster groupings. However, Cupid’s reliance on simple path-based or unconstrained leaf-matching heuristics produced several false positives, such as detecting seven erroneous mappings when structural context was removed and generating duplicate target associations in the purchase order test.

These results imply that generic, automated schema matching can substantially reduce the human overhead, operational costs, and project timelines associated with large-scale enterprise data integration. By demonstrating that structural matching is most effective when anchored to leaf-level data content rather than top-level hierarchies, the article establishes a practical design pattern for data translation components. Nevertheless, the presence of false positives underscores that fully automated matching cannot entirely replace human oversight, and generated mappings must still be validated by domain experts.

For enterprise practitioners and developers, the article recommends deploying matching tools that combine both linguistic normalization and deep structural context rather than relying on isolated name or structure algorithms. Future development should focus on integrating automated parameter tuning to remove the need for manual threshold adjustments, dynamically incorporating user feedback to refine iterative match results, and integrating off-the-shelf thesauri with continuous machine learning. Decision-makers should note key limitations: the algorithm struggles with cyclic schema definitions arising from recursive types and currently relies on manual threshold calibrations. Overall, confidence in the algorithm's core matching capabilities is high for hierarchical and relational structures, but cautious validation remains essential for complex non-tree graph schemas and ambiguous naming environments.

Madhavan et al (2001).pdf
Cover for Generic Schema Matching with Cupid

Abstract

Schema matching is a critical step in many applications, such as XML message mapping, data warehouse loading, and schema integration. In this paper, we investigate algorithms for generic schema matching, outside of any particular data model or application. We first present a taxonomy for past solutions, showing that a rich range of techniques is available. We then propose a new algorithm, Cupid, that discovers mappings between schema elements based on their names, data types, constraints, and schema structure, using a broader set of techniques than past approaches. Some of our innovations are the integrated use of linguistic and structural matching, context-dependent matching of shared types, and a bias toward leaf structure where much of the schema content resides. After describing our algorithm, we present experimental results that compare Cupid to two other schema matching systems.

This is an extended version of a paper published at the 27th VLDB Conference [7].

Table of Contents

  • 1 Introduction
  • 2 The Schema Matching Problem
  • 3 A Taxonomy of Matching Techniques
  • 4 The Cupid Approach
  • 5 Linguistic Matching
  • 5.1 Normalization
  • 5.2 Categorization
  • Name Similarity
  • 5.3 Comparison
  • 6 Structure Matching
  • 7 Mapping Generation
  • 8 Extending to General Schemas
  • 8.1 Schema Graphs
  • 8.2 Matching Shared Types
  • 8.3 Matching Referential Constraints
  • 8.4 Other Features
  • 9 Comparative Study
  • 9.1 Canonical Examples
  • 9.2 Real world example
  • 9.3 Experimental Conclusions
  • 10 Summary and Future Work
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Cupid Schema Matching Framework

    model/method

    Cupid is a generic, schema-based, hybrid matching algorithm designed to discover correspondences across heterogeneous schemas independently of specific data models (e.g., XML Schema, relational databases). The algorithm operates in three successive phases:

    1. Linguistic Matching: Computes a linguistic similarity coefficient lsim(m1,m2)∈[0,1]lsim(m_1, m_2) \in [0, 1] for pairs of schema elements m1m_1 and m2m_2 using name normalization (tokenization, abbreviation expansion, noise elimination, concept tagging), thesaurus lookups, categorization into broad clusters (concept, datatype, container), and token-weighted name similarity.

    2. Structural Matching: Computes a structural similarity coefficient ssim(s,t)∈[0,1]ssim(s, t) \in [0, 1] between schema tree nodes ss and tt using a bottom-up post-order tree traversal. Leaf similarity is initialized via datatype compatibility, and non-leaf similarity is computed as the proportion of subtree leaves that form strongly mapped pairs. Structural similarities of descendant leaves are recursively rewarded or penalized depending on whether their ancestor nodes exhibit high or low overall similarity.

    3. Combined Weighted Similarity and Mapping Generation: For each element pair (s,t)(s, t), the overall weighted similarity wsim(s,t)wsim(s, t) is calculated as:

    wsim(s,t)=wstruct⋅ssim(s,t)+(1−wstruct)⋅lsim(s,t)wsim(s, t) = w_{struct} \cdot ssim(s, t) + (1 - w_{struct}) \cdot lsim(s, t)

    where wstruct∈[0,1]w_{struct} \in [0, 1] is a tunable weight factor (typically set between 0.50.5 and 0.60.6). Mapping elements are generated by selecting source elements ss that maximize wsim(s,t)wsim(s, t) above an acceptance threshold thacceptth_{accept} for target element tt.

  2. Knowl 2 — TreeMatch Structural Matching Algorithm

    algorithm

    The TreeMatch algorithm evaluates structural similarity between schema elements organized as trees SS and TT. The algorithm biases matching toward schema leaves (where atomic data values reside) and propagates similarity context bottom-up.

    Leaf-level structural similarities are initialized using a data type compatibility table (assigning scores in [0,0.5][0, 0.5]). The algorithm performs a post-order traversal of both trees. For non-leaf nodes, structural similarity is computed as the fraction of leaves in the two subtrees that have strong links (wsim≥thacceptwsim \ge th_{accept}). If the combined weighted similarity wsim(s,t)wsim(s, t) exceeds a high threshold thhighth_{high}, the structural similarities of all leaf pairs in leaves(s)×leaves(t)leaves(s) \times leaves(t) are multiplied by an increase factor cincc_{inc}. If wsim(s,t)wsim(s, t) falls below a low threshold thlowth_{low}, leaf similarities are multiplied by a decrease factor cdecc_{dec}.

    Input: SourceTree SS, TargetTree TT, linguistic similarity matrix lsimlsim, thresholds thhigh,thlow,thacceptth_{high}, th_{low}, th_{accept}, structural weight wstructw_{struct}, factors cinc,cdecc_{inc}, c_{dec}
    Output: Structural similarity matrix ssimssim and weighted similarity matrix wsimwsim
    for each s∈S,t∈Ts \in S, t \in T where s,ts, t are leaf nodes:
        ssim(s,t)=datatype-compatibility(s,t)ssim(s, t) = \text{datatype-compatibility}(s, t)
    S′=post-order(S)S' = \text{post-order}(S)
    T′=post-order(T)T' = \text{post-order}(T)
    for each s∈S′s \in S':
        for each t∈T′t \in T':
            ssim(s,t)=structural-similarity(s,t)ssim(s, t) = \text{structural-similarity}(s, t)
            wsim(s,t)=wstruct⋅ssim(s,t)+(1−wstruct)⋅lsim(s,t)wsim(s, t) = w_{struct} \cdot ssim(s, t) + (1 - w_{struct}) \cdot lsim(s, t)
            if wsim(s,t)>thhighwsim(s, t) > th_{high}:
                increase-struct-similarity(leaves(s),leaves(t),cinc)\text{increase-struct-similarity}(leaves(s), leaves(t), c_{inc})
            if wsim(s,t)<thlowwsim(s, t) < th_{low}:
                decrease-struct-similarity(leaves(s),leaves(t),cdec)\text{decrease-struct-similarity}(leaves(s), leaves(t), c_{dec})
  3. Knowl 3 — Schema Tree Construction for Context-Dependent Matching

    algorithm

    Generic schemas are represented as rooted graphs where nodes are schema elements connected by three relationship types:

    • Containment: Strict physical nesting where each element (except the root) has a single parent.
    • Aggregation: Weaker grouping permitting multiple parents without delete propagation (e.g., compound keys).
    • IsDerivedFrom: Generalization/specialization or shared complex type usage.

    When a shared type is referenced by multiple parent elements, its constituent fields must be matched in a context-dependent manner (e.g., an Address type used by both ShippingAddress and BillingAddress). To materialize distinct contexts, Cupid converts the schema graph into an expanded schema tree via a pre-order traversal that duplicates the subschema rooted at the target of each IsDerivedFrom relation for each parent context.

    Input: Schema element current_securrent\_se, Schema tree node current_stncurrent\_stn
    Output: Expanded schema tree node current_stncurrent\_stn
    function construct_schema_tree(current_se, current_stn):
        if current_securrent\_se is root or current_securrent\_se was reached through a containment relationship:
            if current_securrent\_se is tagged not-instantiated:
                return current_stncurrent\_stn
            new_stn=new schema tree node corresponding to current_senew\_stn = \text{new schema tree node corresponding to } current\_se
            set new_stnnew\_stn as a child of current_stncurrent\_stn
            current_stn=new_stncurrent\_stn = new\_stn
            for each outgoing containment or IsDerivedFrom relationship from current_securrent\_se:
                new_se=target schema element of the relationshipnew\_se = \text{target schema element of the relationship}
                construct_schema_tree(new_senew\_se, current_stncurrent\_stn)
            return current_stncurrent\_stn
    schema_tree=construct_schema_tree(schema.root,NULL)schema\_tree = \text{construct\_schema\_tree}(schema.root, \text{NULL})
  4. Knowl 4 — Referential Constraint Reification as Join View Nodes

    model/method

    In Cupid, referential integrity constraints (such as relational foreign keys, XML DTD ID/IDREF pairs, and XML Schema key/keyref pairs) are represented as RefIntRefInt elements that aggregate source referencing columns and reference target primary key columns.

    To exploit referential constraints during structural matching, each foreign key constraint is reified as a synthetic join view node added to the schema tree:

    1. A new join view node representing the join between the two participating tables is created.
    2. The join view node is placed as a child under the lowest common ancestor of the two tables in the schema tree.
    3. The columns of both tables are attached as children of the join view node.

    This transformation provides two structural matching benefits:

    • When two schemas contain analogous referential relationships, matching their join view nodes reinforces the structural similarity of the constituent columns.
    • It enables the discovery of mappings between a normalized schema with joins and a denormalized schema containing single flat tables or pre-aggregated views.
  5. Knowl 5 — Linguistic Matching in Cupid

    model/method

    Linguistic matching in Cupid computes element-level similarities through three stages:

    1. Normalization: Schema element names are parsed into token sets by splitting on punctuation, casing transitions, symbols, and digits. Abbreviations and acronyms are expanded via a thesaurus (e.g., Qty →\rightarrow Quantity), noise words (articles, conjunctions, prepositions) are eliminated, and domain concepts are tagged (e.g., Price and Cost tagged as Money). Each token is classified into one of five token types: number, special symbol, common word, concept, or content.

    2. Categorization: Elements from both schemas are clustered into categories based on concept tags, broad data types (e.g., Number), and container elements (e.g., elements nested within Address). Two categories c1c_1 and c2c_2 are deemed compatible if the name similarity between their keyword sets satisfies ns(c1,c2)≥thnsns(c_1, c_2) \ge th_{ns}. Detailed element-to-element comparisons are pruned to elements that belong to compatible categories.

    3. Comparison: For elements m1∈C1m_1 \in C_1 and m2∈C2m_2 \in C_2 in compatible categories, per-token-type weighted name similarities are calculated and scaled by the maximum compatibility of their categories:

    lsim(m1,m2)=ns(m1,m2)×max⁡c1∈C1,c2∈C2ns(c1,c2)lsim(m_1, m_2) = ns(m_1, m_2) \times \max_{c_1 \in C_1, c_2 \in C_2} ns(c_1, c_2)

    If two elements share no compatible categories, lsim(m1,m2)lsim(m_1, m_2) is set to 0.

  6. Knowl 6 — Token-Level and Name Similarity Formulations in Cupid

    equation

    Let T1T_1 and T2T_2 be two sets of name tokens. The name similarity ns(T1,T2)ns(T_1, T_2) is defined as the average maximum similarity of each token in one set to any token in the opposite set:

    ns(T1,T2)=∑t1∈T1max⁡t2∈T2sim(t1,t2)+∑t2∈T2max⁡t1∈T1sim(t1,t2)∣T1∣+∣T2∣ns(T_1, T_2) = \frac{\sum_{t_1 \in T_1} \max_{t_2 \in T_2} sim(t_1, t_2) + \sum_{t_2 \in T_2} \max_{t_1 \in T_1} sim(t_1, t_2)}{|T_1| + |T_2|}

    where sim(t1,t2)∈[0,1]sim(t_1, t_2) \in [0, 1] is looked up in an annotated synonym and hypernym thesaurus, or estimated via common prefix/suffix substring matching in the absence of a thesaurus entry.

    To compute the overall name similarity between two schema elements m1m_1 and m2m_2, tokens are partitioned into five disjoint token types i∈TokenType={number,special symbol,common word,concept,content}i \in \text{TokenType} = \{\text{number}, \text{special symbol}, \text{common word}, \text{concept}, \text{content}\}. Let T1iT_{1i} and T2iT_{2i} denote the tokens of type ii for m1m_1 and m2m_2, respectively. The overall element name similarity ns(m1,m2)ns(m_1, m_2) is a weighted average across token types:

    ns(m1,m2)=∑i∈TokenTypewi⋅ns(T1i,T2i)⋅(∣T1i∣+∣T2i∣)∑i∈TokenTypewi⋅(∣T1i∣+∣T2i∣)ns(m_1, m_2) = \frac{\sum_{i \in \text{TokenType}} w_i \cdot ns(T_{1i}, T_{2i}) \cdot (|T_{1i}| + |T_{2i}|)}{\sum_{i \in \text{TokenType}} w_i \cdot (|T_{1i}| + |T_{2i}|)}

    where weights satisfy ∑i∈TokenTypewi=1\sum_{i \in \text{TokenType}} w_i = 1, with higher weights assigned to concept and content tokens than to numbers and common grammatical particles.

  7. Knowl 7 — Structural Similarity Measure Based on Leaf Strong Links

    equation

    In Cupid's TreeMatch algorithm, the structural similarity ssim(s,t)ssim(s, t) between two schema tree nodes ss and tt (where at least one is a non-leaf node) is defined as the fraction of leaves in their respective subtrees that possess at least one strong link to a leaf in the opposing subtree:

    ssim(s,t)=∣{x∈leaves(s)∣∃y∈leaves(t),stronglink(x,y)}∪{x∈leaves(t)∣∃y∈leaves(s),stronglink(y,x)}∣∣leaves(s)∪leaves(t)∣ssim(s, t) = \frac{|\{x \in leaves(s) \mid \exists y \in leaves(t), stronglink(x, y)\} \cup \{x \in leaves(t) \mid \exists y \in leaves(s), stronglink(y, x)\}|}{|leaves(s) \cup leaves(t)|}

    where leaves(u)leaves(u) denotes the set of leaf nodes in the subtree rooted at uu. A leaf pair (x,y)(x, y) forms a strong link, denoted stronglink(x,y)stronglink(x, y), if and only if their combined weighted similarity satisfies:

    wsim(x,y)≥thacceptwsim(x, y) \ge th_{accept}

    where thaccept∈[0,1]th_{accept} \in [0, 1] is the acceptance threshold (typically 0.50.5).

    For two leaf nodes ss and tt, ssim(s,t)ssim(s, t) is initialized to their datatype compatibility score in [0,0.5][0, 0.5] (with identical types assigned 0.50.5 to allow structural reinforcement during bottom-up traversal).

  8. Knowl 8 — Computational Optimizations in Cupid Schema Matching

    model/method

    Cupid employs several heuristics and pruning strategies to maintain tractability on complex schemas:

    1. Lazy Schema Tree Expansion: Expanding shared types upfront can cause duplicate structural evaluations across identical instantiated subtrees. In lazy expansion, nodes in the schema graph are traversed in inverse topological order of containment and IsDerivedFrom edges. After an element tt with multiple incoming edges is matched, copies of its evaluated subtree and similarity state are propagated to its parent contexts, preventing redundant recomputations.

    2. Leaf Count Pruning: Nodes ss and tt whose subtrees differ drastically in size (e.g., leaf count ratio exceeding 2:1) are excluded from comparison, as their structural compatibility is reliably poor.

    3. Depth-Bounded Matching: High-level nodes with large subtrees can restrict leaf evaluations to subtrees within depth kk of the node.

    4. Immediate Child Shortcut: If immediate children of two nodes match with very high confidence, expensive leaf-level structural evaluations are bypassed.

    5. Optionality Handling: Leaves reachable exclusively through optional schema nodes (e.g., optional XML attributes) that lack strong links are removed from both the numerator and denominator of ssim(s,t)ssim(s, t), mitigating false structural mismatch penalties.

  9. Knowl 9 — Hyperparameter Configuration and Typical Values in Cupid

    data/table

    The behavior of Cupid is governed by structural similarity thresholds, reinforcement factors, and weighting parameters.

    Parameter Description Typical Value
    thnsth_{ns} Name similarity threshold for determining compatible categories to prune pairwise comparisons. 0.5
    thhighth_{high} High weighted similarity threshold. If wsim(s,t)≥thhighwsim(s,t) \ge th_{high}, structural similarity between leaf pairs in leaves(s)×leaves(t)leaves(s) \times leaves(t) is increased (thhigh>thacceptth_{high} > th_{accept}). 0.6
    thlowth_{low} Low weighted similarity threshold. If wsim(s,t)≤thlowwsim(s,t) \le th_{low}, structural similarity between leaf pairs in leaves(s)×leaves(t)leaves(s) \times leaves(t) is decreased (thlow<thacceptth_{low} < th_{accept}). 0.35
    cincc_{inc} Multiplicative increase factor applied to leaf structural similarities (ssim←min⁡(1.0,ssim⋅cinc)ssim \leftarrow \min(1.0, ssim \cdot c_{inc})). 1.2
    cdecc_{dec} Multiplicative decrease factor applied to leaf structural similarities (ssim←ssim⋅cdecssim \leftarrow ssim \cdot c_{dec}, typically ≈cinc−1\approx c_{inc}^{-1}). 0.9
    thacceptth_{accept} Threshold for identifying strong links (wsim(s,t)≥thacceptwsim(s,t) \ge th_{accept}) and acceptable mapping elements. 0.5
    wstructw_{struct} Weight of structural similarity in computing weighted similarity (wsim=wstruct⋅ssim+(1−wstruct)⋅lsimwsim = w_{struct} \cdot ssim + (1 - w_{struct}) \cdot lsim). 0.5–0.6

    The choice of thnsth_{ns} is not sensitive as it serves solely to prune non-candidate pairs. Leaf structural update factor cincc_{inc} is selected as a function of maximum schema depth, while wstructw_{struct} is typically configured lower for leaf-leaf pairs than for non-leaf pairs.

  10. Knowl 10 — Comparative Evaluation of Cupid, DIKE, and MOMIS on Canonical and Real-World Schemas

    empirical result

    Cupid was empirically compared against two hybrid schema-matching prototypes, DIKE and MOMIS (using ARTEMIS), across six canonical synthetic cases and two real-world integration benchmarks:

    Canonical Cases:

    1. Identical schemas: Cupid correctly mapped all elements without auxiliary thesauri; DIKE duplicated key attributes in the abstracted schema; MOMIS required manual selection of WordNet synsets.
    2. Identical names, different data types: All three handled data type conversions via compatibility tables.
    3. Similar data types, name prefixes/suffixes: Cupid mapped correctly using token normalization; DIKE required explicit Lexical Synonymy Property Dictionary (LSPD) entries; MOMIS required manual user-defined synonym links.
    4. Different class names, identical attributes: Cupid mapped correctly due to leaf-level matching; DIKE merged entities; MOMIS succeeded after manual hypernym selection.
    5. Nesting variations (nested vs. flat customer schemas): Cupid and DIKE correctly matched elements; MOMIS failed to cluster the non-Customer classes together.
    6. Type substitution / context-dependent mapping: Cupid successfully disambiguated shared types (e.g., mapping Address under ShippingAddress vs. BillingAddress to distinct target elements ShipTo and BillTo). DIKE required user intervention, and MOMIS clustered shared classes into independent global clusters.

    Real-World Benchmarks:

    • CIDX vs. Excel XML Purchase Orders: Cupid correctly identified all leaf-level XML attribute matches, including CIDX.line to Excel.itemNumber (based purely on data types and structure without thesaurus support). MOMIS clustered classes without mapping internal attributes 1:1, and DIKE's results depended on the chosen ER modeling abstraction.
    • Relational RDB to Star Data Warehouse: Cupid's join view nodes correctly matched the join of Orders and OrderDetails to the Sales fact table, and Geography to the join of Territories and Region, whereas DIKE and MOMIS failed to cluster several corresponding multi-table combinations.

Coverage note — The high-level taxonomy of past schema matching approaches from Section 3 was omitted as standalone knowls because it summarizes existing literature, focusing extraction instead on the novel algorithms, formal equations, architectural components, optimizations, and empirical findings of Cupid.

References

  1. 1.S. Bergamashchi, S. Castano, M. Vincini: Semantic Integration of Semistructured and Structured Data Sources. SIGMOD Record 28(1), 1999, pp. 54-59.
  2. 2.P.A. Bernstein, A. Halevy, R.A. Pottinger: A Vision for Management of Complex Models. SIGMOD Record 29(4), 2000, pp. 55-63.
  3. 3.S. Castano, V. De Antonellis: A Schema Analysis and Reconciliation Tool Environment. IDEAS’99, pp. 53-62.
  4. 4.C. Clifton, E. Hausman, A. Rosenthal: Experience with a Combined Approach to Attribute-Matching Across Heterogeneous Databases. Proc. 7th IFIP Conf. On DB Semantics, 1997.
  5. 5.A. Doan, P. Domingos, A. Halevy: Reconciling Schemas of Disparate Data Sources: A Machine-Learning Approach. SIGMOD 2001, pp. 509-520.
  6. 6.W. Li, C. Clifton: SEMINT: A tool for identifying attribute correspondences in heterogeneous databases using neural networks. Data & Knowledge Engineering, 33(1), 2000, pp. 49-84.
  7. 7.J. Madhavan, P.A. Bernstein, E. Rahm: Generic Schema Matching using Cupid. VLDB 2001.
  8. 8.Microsoft Corp., BizTalk Mapper: http://www.microsoft.com/technet/biztalk/btsdocs
  9. 9.R. Miller, L. Haas, M.A. Hernandez: Schema Mapping as Query Discovery. VLDB 2000, pp. 77-88.
  10. 10.T. Milo, S. Zohar: Using Schema Matching to Simplify Heterogeneous Data Translation. VLDB 1998.
  11. 11.P. Mitra, G. Weiderhold, J. Jannink: Semi-automatic Integration of Knowledge Sources, FUSION 99.
  12. 12.L. Palopoli, G. Terracina, D. Ursino: The System DIKE: Towards the Semi-Automatic Synthesis of Cooperative Information Systems and Data Warehouses. ADBIS-DASFAA 2000, Matfyzpress, 108-117.
  13. 13.E. Rahm, P.A. Bernstein: On Matching Schemas Automatically. MSR Tech. Report MSR-TR-2001-17, 2001.
  14. 14.J.A. Wald, P.G. Sorenson: Explaining Ambiguity in a Formal Query Language. ACM TODS 15(2), 1990, 125-161.
  15. 15.Q. Wang, J. Yu, K. Wong: Approximate Graph Schema Extraction for Semi-Structured Data. EDBT 2000, pp. 302-316.
  16. 16.WordNet – a lexical database for English: http://www.cogsci.princeton.edu/~wn/.
  17. 17.XML Schema: http://www.w3.org/XML/Schema.

Citation

MLA
Madhavan, J., et al. “Generic Schema Matching with Cupid”. Qucosa (Saxon State and University Library Dresden), 2001, pp. 49–58, http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.216.5548.
APA
Madhavan, J., Bernstein, P. A., & Rahm, E. (2001). Generic Schema Matching with Cupid. Qucosa (Saxon State and University Library Dresden), 49–58. http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.216.5548
Chicago
Madhavan, J., P. A. Bernstein, and E. Rahm. 2001. “Generic Schema Matching with Cupid”. Qucosa (Saxon State and University Library Dresden), 49–58. http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.216.5548.
Harvard
Madhavan, J., Bernstein, P.A. and Rahm, E. (2001) “Generic Schema Matching with Cupid”, Qucosa (Saxon State and University Library Dresden), pp. 49–58. Available at: http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.216.5548.
Vancouver
1. Madhavan J, Bernstein PA, Rahm E (2001) Generic Schema Matching with Cupid. Qucosa (Saxon State and University Library Dresden) 49–58

BibTeX

@article{madhavan2001generic,
  title = {Generic Schema Matching with Cupid},
  author = {Madhavan, Jayant and Bernstein, Philip A. and Rahm, Erhard},
  year = {2001},
  journal = {Qucosa (Saxon State and University Library Dresden)},
  pages = {49-58},
  url = {http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.216.5548}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF