BERTMap: A BERT-Based Ontology Alignment System

Yuan HeJiaoyan ChenDenvar AntonyrajahIan Horrocks

article2022AAAI120 citations

Proposes an ontology alignment system that combines fine-tuned BERT classifiers on ontology-derived corpora with sub-word candidate selection and logic-based mapping refinement to outperform traditional matchers on complex biomedical benchmarks.

Listen

Organizations increasingly rely on structured knowledge bases, known as ontologies, to manage complex domain information. However, independently created ontologies often use different naming conventions and hierarchical structures for identical concepts, creating severe integration bottlenecks. While traditional matching systems rely on simple surface-text matching and logic rules, emerging machine learning techniques often require costly manual data annotation or use static word representations that miss nuanced contextual meanings. The article introduces and evaluates BERTMap, an automated alignment system that combines contextual language models with graph structure and logical reasoning to match equivalent concepts across ontologies without requiring heavy manual supervision.

The system addresses the matching challenge through a four-step pipeline. First, it extracts domain synonyms and non-synonyms directly from the ontologies to construct training corpora. Second, it fine-tunes a contextual language model to score the semantic similarity between class labels. Third, it reduces computational complexity by filtering candidate matches using sub-word indexing before scoring. Finally, it refines predictions by extending matches to neighboring parent and child concepts and executing logic-based repairs to eliminate contradictory pairings. The authors evaluated the system on large-scale biomedical benchmark tasks—including alignments between the Foundational Model of Anatomy, SNOMED Clinical Terms, and the National Cancer Institute Thesaurus—under both unsupervised and semi-supervised configurations.

The evaluation yielded several key findings regarding system performance. First, BERTMap surpassed leading rule-based systems on two out of three large-scale tasks, outperforming top tools like AML and LogMap by approximately 1.4% to 5.4% in overall accuracy balance (F1 score). Second, incorporating a small set of known mappings in a semi-supervised setup consistently improved alignment accuracy over purely unsupervised runs. Third, utilizing complementary auxiliary label sources proved highly impactful when ontologies lacked rich naming metadata, boosting overall matching accuracy by roughly 50% compared to systems restricted to sparse internal text. Fourth, while the system slightly trailed leading baselines by about 2.3% to 2.6% on the FMA-NCI task, it consistently outperformed existing machine learning alternatives across all benchmarks because it effectively learned domain semantics rather than relying on brittle heuristic training samples.

These findings demonstrate that contextual artificial intelligence models can successfully replace traditional surface-level text matching in automated data integration. By capturing deeper contextual synonyms—such as linking spinal abbreviations to anatomical terms—the system reduces the manual effort and operational cost required to harmonize enterprise knowledge representations. Furthermore, combining machine learning predictions with automated logical repairs mitigates data quality risks by ensuring that merged knowledge bases remain coherent and usable for downstream analytics.

Organizations handling complex knowledge integration should consider adopting contextual language model pipelines to automate entity matching, particularly when dealing with extensive synonym variation. When implementing such pipelines, teams should prioritize supplementing target knowledge bases with auxiliary domain dictionaries and applying post-prediction logical repairs to maximize precision. For next steps, the authors recommend expanding evaluations to broader industrial environments and exploring deeper integration between language representations and structural graph embeddings to refine performance on complex structural tasks.

Confidence in these findings is high for biomedical domains with established expert reference standards. However, decision-makers should note that the system's performance varies depending on the structural characteristics of the input data and may slightly lag specialized rule-based systems when source ontologies have distinct structural constraints. Consequently, pilot testing on representative internal datasets is advised before full deployment.

arXiv: 2112.02682
Cover for BERTMap: A BERT-Based Ontology Alignment System

Abstract

Ontology alignment (a.k.a ontology matching (OM)) plays a critical role in knowledge integration. Owing to the success of machine learning in many domains, it has been applied in OM. However, the existing methods, which often adopt ad-hoc feature engineering or non-contextual word embeddings, have not yet outperformed rule-based systems especially in an unsupervised setting. In this paper, we propose a novel OM system named BERTMap which can support both unsupervised and semi-supervised settings. It first predicts mappings using a classifier based on fine-tuning the contextual embedding model BERT on text semantics corpora extracted from ontologies, and then refines the mappings through extension and repair by utilizing the ontology structure and logic. Our evaluation with three alignment tasks on biomedical ontologies demonstrates that BERTMap can often perform better than the leading OM systems LogMap and AML.

Table of Contents

  • Introduction
  • Preliminaries
  • Problem Formulation
  • BERT: Pre-Training and Fine-Tuning
  • BERTMap
  • Corpus Construction and BERT Fine-Tuning
  • Mapping Prediction
  • Mapping Refinement
  • Evaluation
  • Experiment Settings
  • Results
  • Related Work
  • Conclusion and Future Work
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — BERTMap System Architecture

    model/method

    BERTMap is an ontology alignment framework designed to discover equivalence relationships between classes of two ontologies O\mathcal{O} and O′\mathcal{O}' by integrating contextual language representation with ontology structure and logical reasoning. The pipeline consists of four major steps:

    1. Corpus Construction: Extracts synonym and non-synonym label pairs from the input ontologies (intra-ontology), optional known mappings (cross-ontology), and optional auxiliary domain ontologies (complementary corpus).
    2. Fine-Tuning: Fine-tunes a pre-trained contextual language representation model (e.g., Bio-Clinical BERT) paired with a binary classification head on the extracted synonym/non-synonym corpora to predict whether any pair of entity labels are synonymous.
    3. Mapping Prediction: Filters candidate class pairs using a sub-word inverted index based on WordPiece tokenization and IDF scoring. Candidate pairs are then scored using exact string matching or by averaging BERT synonym classification scores across all pairwise label combinations.
    4. Mapping Refinement: Applies an iterative mapping extension step based on ontology hierarchy (locality principle) to improve recall, followed by a propositional logic-based mapping repair step to eliminate mappings causing logical inconsistency.
  2. Knowl 2 — Text Semantics Corpora Construction for Ontology Label Fine-Tuning

    model/method

    To fine-tune contextual language representation models without manual annotations, BERTMap extracts positive (synonym) and negative (non-synonym) label pairs from three sources:

    • Intra-ontology Corpus (ioio): For each named class cc with preprocessed label set Ω(c)\Omega(c), all distinct pairs (ω1,ω2)∈Ω(c)×Ω(c)(\omega_1, \omega_2) \in \Omega(c) \times \Omega(c) are extracted as synonyms. Label identity pairs (ω,ω)(\omega, \omega) can optionally be included as identity synonyms (idsids). Negative samples consist of:
      1. Soft non-synonyms: label pairs (ω1,ω2)(\omega_1, \omega_2) from two randomly selected distinct classes.
      2. Hard non-synonyms: label pairs from structurally disjoint classes, where sibling classes sharing an immediate common superclass are assumed disjoint.
    • Cross-ontology Corpus (coco): In a semi-supervised setting with a seed set of known equivalence class mappings (c,c′)(c, c'), synonyms are extracted from the Cartesian product Ω(c)×Ω(c′)\Omega(c) \times \Omega(c'), and non-synonyms are extracted from randomly aligned class pairs across the two ontologies.
    • Complementary Corpus (cpcp): Synonyms and non-synonyms are extracted following the intra-ontology method from an external auxiliary ontology in the same domain, restricted to classes that share at least one label with the input ontologies.

    To ensure symmetry, for every synonym pair (ω1,ω2)(\omega_1, \omega_2), the reversed pair (ω2,ω1)(\omega_2, \omega_1) is added. If any randomly generated non-synonym pair appears in the synonym set, it is removed.

  3. Knowl 3 — Sub-Word Inverted Index-Based Candidate Selection

    model/method

    To eliminate the O(∣C∣⋅∣C′∣)O(|C| \cdot |C'|) cost of evaluating all class pairs between source class set CC and target class set C′C', BERTMap prunes the candidate search space using sub-word inverted indices built from BERT's inherent WordPiece tokenizer.

    Let T(c)T(c) denote the set of WordPiece sub-word tokens extracted from all preprocessed labels Ω(c)\Omega(c) of class c∈Cc \in C. An inverted index I′I' maps each sub-word token tt to the set of target classes I′[t]⊆C′I'[t] \subseteq C' whose labels contain tt.

    For each source class cc, candidate target classes in ⋃t∈T(c)I′[t]\bigcup_{t \in T(c)} I'[t] are ranked using an inverted document frequency (IDF) scoring metric:

    Sselect(c,c′)=∑t∈T(c)∩T(c′)idf(t)=∑t∈T(c)∩T(c′)log⁡10∣C′∣∣I′[t]∣S_{select}(c, c') = \sum_{t \in T(c) \cap T(c')} \text{idf}(t) = \sum_{t \in T(c) \cap T(c')} \log_{10} \frac{|C'|}{|I'[t]|}

    where ∣⋅∣|\cdot| denotes set cardinality. Only the top-kk scored target classes (with k≪∣C′∣k \ll |C'|; e.g., k=200k = 200) are retained for subsequent scoring, reducing overall candidate selection complexity to O(k∣C∣)O(k |C|).

  4. Knowl 4 — Mapping Score Computation via String Match and BERT Classifier

    equation

    For a source class c∈Cc \in C with label set Ω(c)\Omega(c) and a candidate target class c′∈C′c' \in C' with label set Ω(c′)\Omega(c'), the mapping score Smap(c,c′)∈[0,1]S_{map}(c, c') \in [0, 1] is defined as:

    Smap(c,c′)={1.0if Ω(c)∩Ω(c′)≠∅Sbert(Ω(c),Ω(c′))otherwiseS_{map}(c, c') = \begin{cases} 1.0 & \text{if } \Omega(c) \cap \Omega(c') \neq \emptyset \\ S_{bert}(\Omega(c), \Omega(c')) & \text{otherwise} \end{cases}

    where Ω(c)∩Ω(c′)≠∅\Omega(c) \cap \Omega(c') \neq \emptyset indicates that cc and c′c' share at least one identical preprocessed label (serving as a string-matching fast path).

    When no identical label exists, Sbert(Ω(c),Ω(c′))S_{bert}(\Omega(c), \Omega(c')) computes the arithmetic mean of synonym classification scores across all label pairs in the Cartesian product:

    Sbert(Ω(c),Ω(c′))=1∣Ω(c)∣⋅∣Ω(c′)∣∑ω∈Ω(c)∑ω′∈Ω(c′)s(ω,ω′)S_{bert}(\Omega(c), \Omega(c')) = \frac{1}{|\Omega(c)| \cdot |\Omega(c')|} \sum_{\omega \in \Omega(c)} \sum_{\omega' \in \Omega(c')} s(\omega, \omega')

    where s(ω,ω′)∈[0,1]s(\omega, \omega') \in [0, 1] is the positive (synonymous) softmax probability generated by the fine-tuned BERT classifier for the tokenized label pair (ω,ω′)(\omega, \omega'). The top-scoring target class c′=arg⁡max⁡c′Smap(c,c′)c' = \arg\max_{c'} S_{map}(c, c') is selected as the candidate mapping.

    Mappings can be constructed in three directions:

    1. src2tgtsrc2tgt: searching the best target class c′∈C′c' \in C' for each source class c∈Cc \in C.
    2. tgt2srctgt2src: searching the best source class c∈Cc \in C for each target class c′∈C′c' \in C'.
    3. combinedcombined: the deduplicated union of src2tgtsrc2tgt and tgt2srctgt2src mappings.
  5. Knowl 5 — Iterative Mapping Extension Algorithm

    algorithm

    Under the ontological locality principle, if two classes cc and c′c' are equivalent, their respective superclasses and subclasses are also likely to correspond. BERTMap expands high-confidence mapping predictions using an iterative extension procedure:

    Input: High-confidence mapping set MM, extension threshold κ\kappa
    Output: Extended mapping set MexM_{ex}
    Mfr←MM_{fr} \leftarrow M
    Mex←∅M_{ex} \leftarrow \emptyset
    Let Sup(c)\text{Sup}(c) return the direct superclasses of cc
    Let Sub(c)\text{Sub}(c) return the direct subclasses of cc
    while Mfr≠∅M_{fr} \neq \emptyset do
        Mnew←∅M_{new} \leftarrow \emptyset
        for each mapping (c,c′,Smap(c,c′))∈Mfr(c, c', S_{map}(c, c')) \in M_{fr} do
            for each (x,x′)∈(Sup(c)×Sup(c′))∪(Sub(c)×Sub(c′))(x, x') \in (\text{Sup}(c) \times \text{Sup}(c')) \cup (\text{Sub}(c) \times \text{Sub}(c')) do
                m←(x,x′,Smap(x,x′))m \leftarrow (x, x', S_{map}(x, x'))
                if Smap(x,x′)≥κS_{map}(x, x') \ge \kappa and m∉Mm \notin M and m∉Mexm \notin M_{ex} then
                    Mnew←Mnew∪{m}M_{new} \leftarrow M_{new} \cup \{m\}
                end if
            end for
        end for
        Mex←Mex∪MnewM_{ex} \leftarrow M_{ex} \cup M_{new}
        Mfr←MnewM_{fr} \leftarrow M_{new}
    end while
    return MexM_{ex}

    The algorithm iteratively evaluates candidate pairs across superclass and subclass Cartesian products. Any newly identified mapping with score Smap(x,x′)≥κS_{map}(x, x') \ge \kappa (where κ=0.9\kappa = 0.9) is added to the frontier set MfrM_{fr} and extension set MexM_{ex} until no further qualifying mappings are found.

  6. Knowl 6 — Propositional Logic-Based Mapping Repair

    model/method

    Integrating two ontologies via equivalence mappings can induce logical inconsistencies, such as unsatisfiable classes. A diagnosis (perfect repair) is a minimal set of mappings whose deletion eliminates all logical incoherence, but computing exact diagnoses on large ontologies is computationally prohibitive.

    BERTMap applies an approximate propositional logic-based repair algorithm (Jiménez-Ruiz et al. 2013) to prune inconsistent mappings. This repair module computes an approximate repair set RR of mappings to remove such that:

    1. RR is guaranteed to be a subset of a valid diagnosis, preventing the unnecessary removal of correct mappings;
    2. Only a minimal number of unsatisfiable classes remain after deletion of RR.

    Because mapping extension is restricted to high-confidence mappings and repair is performed via propositional approximations, refinement consistently enhances mapping precision without prohibitive computational overhead.

  7. Knowl 7 — Biomedical Ontology Alignment Evaluation Setup and Tasks

    experimental setup

    BERTMap is evaluated on three biomedical ontology alignment benchmark tasks from the OAEI LargeBio track:

    1. FMA-SNOMED: Matching the Foundational Model of Anatomy (10,157 classes) with SNOMED CT (13,412 classes), containing 6,026 target reference mappings (M=M_=) and 2,982 ignored mappings (M?M_?) that cause logical inconsistencies.
    2. FMA-NCI: Matching FMA (3,696 classes) with the National Cancer Institute Thesaurus (6,488 classes), containing 2,686 reference mappings (M=M_=) and 338 ignored mappings (M?M_?).
    3. FMA-SNOMED+: An extended FMA-SNOMED task where the SNOMED ontology is augmented with labels and synonyms searched from the 2021 release of SNOMED CT.

    Performance is evaluated using Precision (PP), Recall (RR), and Macro-F1F_1 (F1F_1):

    P=∣Mout∩M=∖M?∣∣Mout∖M?∣,R=∣Mout∩M=∖M?∣∣M=∖M?∣,F1=2PRP+RP = \frac{|M_{out} \cap M_= \setminus M_?|}{|M_{out} \setminus M_?|}, \quad R = \frac{|M_{out} \cap M_= \setminus M_?|}{|M_= \setminus M_?|}, \quad F_1 = \frac{2PR}{P + R}

    In the unsupervised setting, M=M_= is split into validation (MvalM_{val}, 10%) and test (MtestM_{test}, 90%). In the semi-supervised setting, M=M_= is split into training (MtrainM_{train}, 20%), validation (MvalM_{val}, 10%), and test (MtestM_{test}, 70%). Mappings not in the active evaluation set are treated as ignored mappings (M?M_?).

    Fine-tuning utilizes Bio-Clinical BERT with sequence length 128, batch size 32, Adam optimizer, a positive-to-negative sample ratio of 1:4 (2 soft and 2 hard non-synonyms per synonym in ioio and cpcp, 4 non-synonyms in coco), trained for 3 epochs with checkpoint evaluations every 0.1 epoch.

  8. Knowl 8 — Comparative Performance of BERTMap and Baselines on Biomedical Ontologies

    empirical result

    BERTMap was benchmarked against leading rule-based systems (LogMap, AML, LogMapLt) and machine learning-based approaches (LogMap-ML*, string-matching, edit-similarity) across the FMA-SNOMED, FMA-SNOMED+, and FMA-NCI tasks.

    FMA-SNOMED (90% Test) FMA-SNOMED+ (90% Test)
    System / Setting Precision Recall Macro-F1 Precision Recall Macro-F1
    String-match 0.987 0.194 0.324 0.978 0.672 0.797
    Edit-similarity 0.971 0.209 0.343 0.978 0.728 0.834
    LogMapLt 0.965 0.206 0.339 0.953 0.717 0.819
    LogMap-ML* 0.944 0.205 0.337 0.955 0.684 0.797
    LogMap 0.935 0.685 0.791 0.869 0.867 0.868
    AML 0.892 0.757 0.819 0.895 0.829 0.861
    BERTMap (unsupervised best) 0.905 0.771 0.833 0.924 0.851 0.886
    BERTMap (semi-supervised best)* 0.892 0.786 0.836 0.908 0.852 0.879

    *(Semi-supervised results measured on the 70% test set, where AML achieves F1=0.806F_1 = 0.806 on FMA-SNOMED and F1=0.846F_1 = 0.846 on FMA-SNOMED+; LogMap achieves F1=0.782F_1 = 0.782 on FMA-SNOMED and F1=0.852F_1 = 0.852 on FMA-SNOMED+).

    On FMA-NCI (90% test set), LogMap achieves F1=0.919F_1 = 0.919 (P=0.938,R=0.900P=0.938, R=0.900) and AML achieves F1=0.918F_1 = 0.918 (P=0.936,R=0.900P=0.936, R=0.900), whereas unsupervised BERTMap achieves F1=0.893F_1 = 0.893 (P=0.938,R=0.852P=0.938, R=0.852) and semi-supervised BERTMap achieves F1=0.880F_1 = 0.880 (P=0.959,R=0.813P=0.959, R=0.813).

    Key takeaways:

    1. BERTMap outperforms all baselines on FMA-SNOMED (surpassing AML by 1.4% unsupervised and 3.0% semi-supervised) and FMA-SNOMED+ (surpassing AML by 2.5% unsupervised and 3.3% semi-supervised).
    2. On ontologies with sparse class labels (FMA-SNOMED), BERTMap augmented with a complementary corpus (+cp+cp) outperforms lexical and word-embedding baselines by approximately 50% in Macro-F1F_1.
    3. The machine learning baseline LogMap-ML* suffers from noisy heuristic anchor selection (F1≈0.337–0.822F_1 \approx 0.337\text{--}0.822), whereas BERTMap successfully fine-tunes BERT using self-supervised intra-ontology label pairs.
  9. Knowl 9 — Threshold Sensitivity and Mapping Validation Dynamics in BERTMap

    empirical result

    Validation across mapping score threshold λ∈[0,1]\lambda \in [0, 1] reveals a consistent optimization profile across all BERTMap experimental settings:

    • As threshold λ\lambda increases from 0 toward 1.0, precision increases steeply while recall drops only slightly, causing Macro-F1F_1 to increase monotonically and peak at λ=0.999\lambda = 0.999.
    • In a two-step validation procedure on validation set MvalM_{val}—where step 1 optimizes {τ,λ}\{\tau, \lambda\} on initial predictions and step 2 optimizes λ\lambda after iterative mapping extension—the best threshold λ\lambda determined in step 1 consistently coincides with the best threshold in step 2 (λ=0.999\lambda = 0.999).
    • The mapping selection direction τ=src2tgt\tau = \text{src2tgt} combined with λ=0.999\lambda = 0.999 consistently provides optimal validation performance across both unsupervised and semi-supervised configurations.
  10. Knowl 10 — Contextual Semantic Generalization of BERTMap Over Lexical Matchers

    empirical result

    BERTMap successfully retrieves semantically equivalent class mappings where lexical and string-matching OM systems fail due to divergent surface forms. Examples from the FMA-SNOMED benchmark include:

    • Third_cervical_spinal_ganglion (FMA) ≡\equiv C3_spinal_ganglion (SNOMED): BERTMap captures the clinical equivalence between the abbreviation "C3" and the phrase "third cervical".
    • Deep_posterior_sacrococcygeal_ligament (FMA) ≡\equiv Structure_of_deep_dorsal_sacrococcygeal_ligament (SNOMED): BERTMap identifies anatomical synonymy between "posterior" and "dorsal".
    • Wall_of_smooth_endoplasmic_reticulum (FMA) ≡\equiv Agranular_endoplasmic_reticulum_membrane (SNOMED): BERTMap identifies semantic equivalence between "smooth" and "agranular", as well as "wall" and "membrane" in cellular anatomy.

    These examples confirm that fine-tuned contextual transformers effectively capture domain-specific semantic synonyms beyond surface sub-string overlap.

Coverage note — None was omitted; all key contributions—including system architecture, corpus construction, candidate selection, mapping scoring, iterative extension, logic repair, benchmark setups, empirical results, threshold dynamics, and qualitative case analyses—are represented.

References

  1. 1.Alsentzer, E.; Murphy, J.; Boag, W.; Weng, W.-H.; Jindi, D.; Naumann, T.; and McDermott, M. 2019. Publicly Available Clinical BERT Embeddings. In Proceedings of the 2nd Clinical Natural Language Processing Workshop, 72–78.
  2. 2.Bojanowski, P.; Grave, E.; Joulin, A.; and Mikolov, T. 2017. Enriching Word Vectors with Subword Information. Transactions of the Association for Computational Linguistics, 5: 135–146.
  3. 3.Chen, J.; Hu, P.; Jimenez-Ruiz, E.; Holter, O. M.; Antonyrajah, D.; and Horrocks, I. 2021a. OWL2Vec*: Embedding of OWL ontologies. Machine Learning, 1–33.
  4. 4.Chen, J.; Jiménez-Ruiz, E.; Horrocks, I.; Antonyrajah, D.; Hadian, A.; and Lee, J. 2021b. Augmenting ontology alignment by semantic embedding and distant supervision. In European Semantic Web Conference, 392–408. Springer.
  5. 5.Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of NAACL-HLT, 4171–4186.
  6. 6.Faria, D.; Pesquita, C.; Santos, E.; Palmonari, M.; Cruz, I. F.; and Couto, F. M. 2013. The AgreementMakerLight Ontology Matching System. In Meersman, R.; Panetto, H.; Dillon, T.; Eder, J.; Bellahsene, Z.; Ritter, N.; De Leenheer, P.; and Dou, D., eds., On the Move to Meaningful Internet Systems: OTM 2013 Conferences, 527–541. Berlin, Heidelberg: Springer Berlin Heidelberg. ISBN 978-3-642-41030-7.
  7. 7.Grau, B. C.; Horrocks, I.; Kazakov, Y.; and Sattler, U. 2007. A Logical Framework for Modularity of Ontologies. In IJCAI.
  8. 8.Iyer, V.; Agarwal, A.; and Kumar, H. 2020. VeeAlign: a supervised deep learning approach to ontology alignment. In OM@ISWC.
  9. 9.Jiménez-Ruiz, E.; Agibetov, A.; Chen, J.; Samwald, M.; and Cross, V. V. 2020. Dividing the Ontology Alignment Task with Semantic Embeddings and Logic-based Modules. ArXiv, abs/2003.05370.
  10. 10.Jiménez-Ruiz, E.; and Cuenca Grau, B. 2011. LogMap: Logic-Based and Scalable Ontology Matching. In Aroyo, L.; Welty, C.; Alani, H.; Taylor, J.; Bernstein, A.; Kagal, L.; Noy, N.; and Blomqvist, E., eds., The Semantic Web – ISWC 2011, 273–288. Berlin, Heidelberg: Springer Berlin Heidelberg. ISBN 978-3-642-25073-6.
  11. 11.Jiménez-Ruiz, E.; Meilicke, C.; Grau, B. C.; and Horrocks, I. 2013. Evaluating Mapping Repair Systems with Large Biomedical Ontologies. In Description Logics.
  12. 12.Kolyvakis, P.; Kalousis, A.; and Kiritsis, D. 2018. DeepAlignment: Unsupervised Ontology Matching with Refined Word Vectors. In Proceedings of NAACL-HLT, 787–798.
  13. 13.Loshchilov, I.; and Hutter, F. 2017. Fixing Weight Decay Regularization in Adam. ArXiv, abs/1711.05101.
  14. 14.Mikolov, T.; Chen, K.; Corrado, G. S.; and Dean, J. 2013. Efficient Estimation of Word Representations in Vector Space. In ICLR.
  15. 15.Neutel, S.; and Boer, M. D. 2021. Towards Automatic Ontology Alignment using BERT. In AAAI Spring Symposium: Combining Machine Learning with Knowledge Engineering.
  16. 16.Nkisi-Orji, I.; Wiratunga, N.; Massie, S.; Hui, K.-Y.; and Heaven, R. 2018. Ontology alignment based on word embedding and random forest classification. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases, 557–572. Springer.
  17. 17.Otero-Cerdeira, L.; Rodríguez-Martínez, F. J.; and Gómez-Rodríguez, A. 2015. Ontology matching: A literature review. Expert Systems with Applications, 42(2): 949–971.
  18. 18.Portisch, J.; Hladik, M.; and Paulheim, H. 2019. Wiktionary Matcher. In OM@ISWC.
  19. 19.Portisch, J.; and Paulheim, H. 2018. ALOD2Vec matcher. In OM@ISWC.
  20. 20.Reimers, N.; and Gurevych, I. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. ArXiv, abs/1908.10084.
  21. 21.Shvaiko, P.; and Euzenat, J. 2013. Ontology Matching: State of the Art and Future Challenges. IEEE Transactions on Knowledge and Data Engineering, 25(1): 158–176.
  22. 22.Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L. u.; and Polosukhin, I. 2017. Attention is All you Need. In Guyon, I.; Luxburg, U. V.; Bengio, S.; Wallach, H.; Fergus, R.; Vishwanathan, S.; and Garnett, R., eds., Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc.
  23. 23.Wang, L.; Bhagavatula, C.; Neumann, M.; Lo, K.; Wilhelm, C.; and Ammar, W. 2018. Ontology alignment in the biomedical domain using entity definitions and context. In Proceedings of the BioNLP 2018 workshop, 47–55.
  24. 24.Wu, Y.; Schuster, M.; Chen, Z.; Le, Q. V.; Norouzi, M.; Macherey, W.; Krikun, M.; Cao, Y.; Gao, Q.; Macherey, K.; Klingner, J.; Shah, A.; Johnson, M.; Liu, X.; Kaiser, L.; Gouws, S.; Kato, Y.; Kudo, T.; Kazawa, H.; Stevens, K.; Kurian, G.; Patil, N.; Wang, W.; Young, C.; Smith, J.; Riesa, J.; Rudnick, A.; Vinyals, O.; Corrado, G.; Hughes, M.; and Dean, J. 2016. Google’s Neural Machine Translation System: Bridging the Gap between Human and Machine Translation. CoRR.
  25. 25.Xiang, C.; Jiang, T.; Chang, B.; and Sui, Z. 2015. Ersom: A structural ontology matching approach using automatically learned entity representation. In Proceedings of the 2015 conference on empirical methods in natural language processing, 2419–2429.

Citation

MLA
He, Y., et al. “BERTMap: A BERT-based Ontology Alignment System”. arXiv, 2021, http://arxiv.org/abs/2112.02682v4.
APA
He, Y., Chen, J., Antonyrajah, D., & Horrocks, I. (2021). BERTMap: A BERT-based Ontology Alignment System. arXiv. http://arxiv.org/abs/2112.02682v4
Chicago
He, Y., J. Chen, D. Antonyrajah, and I. Horrocks. 2021. “BERTMap: A BERT-based Ontology Alignment System”. arXiv. http://arxiv.org/abs/2112.02682v4.
Harvard
He, Y. et al. (2021) “BERTMap: A BERT-based Ontology Alignment System”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2112.02682v4.
Vancouver
1. He Y, Chen J, Antonyrajah D, Horrocks I (2021) BERTMap: A BERT-based Ontology Alignment System. arXiv

BibTeX

@article{he2021bertmap,
  title = {BERTMap: A BERT-based Ontology Alignment System},
  author = {He, Yuan and Chen, Jiaoyan and Antonyrajah, Denvar and Horrocks, Ian},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2112.02682v4},
  eprint = {2112.02682}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF