XL-LEXEME: WiC Pretrained Model for Cross-Lingual LEXical sEMantic changE

Pierluigi CassottiLucia SicilianiMarco de GemmisGiovanni SemeraroPierpaolo Basile

article2023ACL77 citations

Presents XL-LEXEME, a bi-encoder model that adapts Sentence-BERT to the Word-in-Context task with target-word highlighting to generate scalable, comparable lexical representations that achieve state-of-the-art semantic change detection across multiple languages.

Listen

Tracking how word meanings evolve over time—known as lexical semantic change detection—is essential for accurate historical text analysis, information retrieval, and language technology. Traditional methods often depend on rigid dictionary sense inventories that fail to capture emerging or obsolete meanings. While modern deep learning models can determine whether a word shares the same meaning across different contexts, existing architectures typically rely on joint-sentence cross-encoders. These cross-encoders are computationally expensive and cannot produce standalone, directly comparable word representations across vast historical corpora.

The article demonstrates an efficient, multilingual approach called XL-LEXEME, which adapts Sentence-BERT architectures using target word delimiters to perform lexical semantic change detection. The model evaluates whether knowledge learned from synchronic Word-in-Context datasets can successfully transfer to historical, diachronic language shifts across five languages: English, German, Swedish, Latin, and Russian.

To achieve this, the authors built a Siamese bi-encoder network utilizing XLM-RoBERTa Large and trained it with contrastive loss across merged multilingual Word-in-Context benchmark datasets. Target words within input sentences are marked with special delimiters, allowing the network to encode sentences independently into separate, comparable vector representations. The model was evaluated against established historical benchmarks, specifically the SemEval-2020 Task 1 ranking subtask across four languages and the RuShiftEval benchmark across three historical periods in Russian.

The evaluation yielded several key findings. First, XL-LEXEME outperformed previous state-of-the-art systems and baselines in English (0.757 correlation), German (0.877 correlation), and Swedish (0.754 correlation). Second, on the Russian RuShiftEval benchmark, XL-LEXEME achieved state-of-the-art performance with an average correlation of 0.802, which further increased to 0.825 when fine-tuned on target historical data. Third, the Siamese bi-encoder architecture reduced theoretical computational complexity compared to standard cross-encoders while maintaining superior accuracy. Finally, the model failed on Latin (-0.056 correlation), demonstrating no significant alignment with human annotations.

These findings indicate that general-purpose contextual training can effectively identify semantic evolution over centuries without needing explicit temporal annotations or costly cross-encoding. This reduction in computational requirements lowers infrastructure costs and processing timelines when analyzing large-scale text archives. However, the contrast between strong performance on modern languages and failure on Latin shows that the model relies heavily on language representation within the underlying pre-trained multilingual model.

Stakeholders and practitioners analyzing evolving terminology across large archives can deploy XL-LEXEME as an efficient, high-performing alternative to heavy cross-encoder pipelines. When applying the model to new languages, teams should prioritize languages well-represented in underlying foundation models or related language families. Future work should focus on developing dedicated cross-lingual semantic change evaluation benchmarks and expanding training coverage for low-resource and ancient languages.

Confidence in the reported improvements is high for modern European languages with sufficient pre-training representation. However, users should exercise caution regarding small evaluation sample sizes (ranging from 31 to 48 target words per language in SemEval-2020) and acknowledge performance limitations on ancient or severely underrepresented languages.

No sufficiently relevant recommendations were found.

Cover for XL-LEXEME: WiC Pretrained Model for Cross-Lingual LEXical sEMantic changE

Abstract

The recent introduction of large-scale datasets for the WiC (Word in Context) task enables the creation of more reliable and meaningful contextualized word embeddings. However, most of the approaches to the WiC task use cross-encoders, which prevent the possibility of deriving comparable word embeddings. In this work, we introduce XL-LEXEME, a Lexical Semantic Change Detection model. XL-LEXEME extends SBERT, highlighting the target word in the sentence. We evaluate XL-LEXEME on the multilingual benchmarks for SemEval-2020 Task 1 - Lexical Semantic Change (LSC) Detection and the RuShiftEval shared task involving five languages: English, German, Swedish, Latin, and Russian. XL-LEXEME outperforms the state-of-the-art in English, German and Swedish with statistically significant differences from the baseline results and obtains state-of-the-art performance in the RuShiftEval shared task.

Table of Contents

  • 1 Introduction and Motivation
  • 2 Related Work
  • 3 XL-LEXEME
  • 4 Experimental setting
  • 4.1 Lexical Semantic Change Detection
  • 4.2 Training details
  • 5 Results
  • 6 Conclusion
  • 7 Limitations
  • Acknowledgements
  • References
  • A Hyper-parameters

Knowls

  1. Knowl 1 — XL-LEXEME produces target-focused, comparable context representations

    model/method

    XL-LEXEME adapts a Siamese sentence encoder to Word-in-Context (WiC) training for lexical semantic change detection. Given two sentences containing the same target word, it marks each target occurrence with <t> and </t>, tokenizes the sentences into subwords, and encodes them separately using shared XLM-RoBERTa-large weights. The target markers identify the word whose sense is being compared, while each representation is formed by summing the encoded vectors for all subwords in that sentence. The two resulting vectors are distinct sentence representations in a shared embedding space, rather than a single joint representation produced by a cross-encoder.

  2. Knowl 2 — Contrastive WiC training objective

    equation

    XL-LEXEME is trained to bring sentence representations for same-meaning target-word uses together and separate different-meaning uses, using the contrastive loss

    ℓ=12[yδ2+(1−y)max⁡(0,m−δ)2].\ell=\frac{1}{2}\left[y\delta^2+(1-y)\max(0,m-\delta)^2\right].

    Here, y∈{0,1}y\in\{0,1\} is the WiC pair label, with y=1y=1 for a same-meaning pair; δ\delta is the cosine distance between the two encoded sentence representations; and the margin is fixed at m=0.5m=0.5. The loss is applied to pairs of sentences that contain the target word in context.

  3. Knowl 3 — Lexical semantic change score from cross-period context pairs

    equation

    For a target word, XL-LEXEME estimates lexical semantic change by averaging distances between context representations sampled from two time periods. Let n0n_0 and n1n_1 be the numbers of sampled contexts from periods t0t_0 and t1t_1, respectively; let sit0\mathbf{s}^{t_0}_i and sjt1\mathbf{s}^{t_1}_j be their XL-LEXEME sentence representations; and let δ\delta denote cosine distance. The score is

    LSC⁡(t0,t1)=1n0n1∑i=1n0∑j=1n1δ ⁣(sit0,sjt1).\operatorname{LSC}(t_0,t_1)=\frac{1}{n_0n_1}\sum_{i=1}^{n_0}\sum_{j=1}^{n_1}\delta\!\left(\mathbf{s}^{t_0}_i,\mathbf{s}^{t_1}_j\right).

    A larger average cross-period distance indicates greater semantic change. The same cross-period averaging procedure is used with the cross-encoder baseline, using its model-specific pair distance instead of XL-LEXEME's cosine distance.

  4. Knowl 4 — Separate encoding reduces pairwise inference cost

    theoretical result

    The paper compares self-attention costs for a cross-encoder and XL-LEXEME under a fixed total input length. If NN is the total number of tokens processed by the cross-encoder and dd is the hidden-vector dimension, the cross-encoder cost is stated as O(N2d)O(N^2d). XL-LEXEME encodes two sequences separately, each of length approximately N/2N/2, for a stated cost of O ⁣(2(N/2)2d)O\!\left(2(N/2)^2d\right). The reduction follows from avoiding joint attention across both sentences; XL-LEXEME also yields an individual reusable representation for each sentence.

  5. Knowl 5 — Training and evaluation configuration

    experimental setup

    XL-LEXEME and the cross-encoder baseline use fine-tuned XLM-RoBERTa-large. WiC training data combine MCL-WiC, AM2iCo, and XL-WiC with a randomly selected 75% of each dataset's development data; the remaining 25% of development data is used for hyperparameter tuning. The cross-encoder training set is augmented by swapping sentence order. Both systems use AdamW and linear learning-rate warm-up over 10% of the training data. The learning-rate search is {10−6,2×10−6,5×10−6,10−5,2×10−5}\{10^{-6},2\times10^{-6},5\times10^{-6},10^{-5},2\times10^{-5}\} and the weight-decay search is {0,0.01}\{0,0.01\}; the selected learning rate is 10−510^{-5} for both, with weight decay 0.00 for XL-LEXEME and 0.01 for the cross-encoder. Maximum sequence lengths are 128 per XL-LEXEME sentence and 256 for the cross-encoder pair. The XLM-R-large configuration has 24 layers, hidden size 1024, and 16 attention heads. Experiments ran on an NVIDIA GeForce RTX 3090.

    Evaluation covers SemEval-2020 Task 1 Subtask 2, which ranks target words by semantic-change degree in English, German, Swedish, and Latin, and RuShiftEval, which evaluates Russian pre-Soviet-to-Soviet, Soviet-to-post-Soviet, and pre-Soviet-to-post-Soviet period pairs. For each language and period, 200 target-word contexts are sampled; sampling is repeated ten times and reported results are averaged across repetitions.

  6. Knowl 6 — SemEval results: strongest correlations in English, German, and Swedish

    data/table

    The table reports Spearman correlations for SemEval-2020 Task 1 Subtask 2. XL-LEXEME has the highest reported correlation among the listed systems for English, German, and Swedish, but not Latin; the Latin score is slightly negative. A dagger marks a score for which the paper reports no statistically significant difference from XL-LEXEME. The paper reports a significant correlation for German (p<0.001p<0.001) and no significant correlation for Latin. Its Fisher z-transformation comparisons report significant differences from leaderboard systems for English and German (p<0.05p<0.05), except for TempoBERT and Temporal Attention; in Swedish, XL-LEXEME is reported as significantly different from the Count baseline.

    Language UG-Student-Intern Jiaxin-Jinan cs2020 UWB Count Freq. TempoBERT Temporal Attn. Cross-encoder XL-LEXEME
    EN 0.422 0.325 0.375 0.367 0.022 -0.217 0.467 †\dagger0.520 †\dagger0.752 0.757
    DE 0.725 0.717 0.702 0.697 0.216 0.014 – †\dagger0.763 †\dagger0.837 0.877
    SV †\dagger0.547 †\dagger0.588 †\dagger0.536 †\dagger0.604 -0.022 -0.150 – – †\dagger0.680 0.754
    LA 0.412 0.440 0.399 0.254 0.359 †\dagger0.020 0.512 0.565 †\dagger0.016 -0.056
    Average 0.527 0.518 0.503 0.481 0.144 -0.083 – – 0.571 0.583

    The cross-encoder scores are 0.752, 0.837, 0.680, and -0.056 for English, German, Swedish, and Latin, respectively; XL-LEXEME improves on the first three and matches Latin. XL-LEXEME's average is 0.583, compared with 0.571 for the cross-encoder.

  7. Knowl 7 — RuShiftEval results: fine-tuning improves all three period-pair scores

    data/table

    The table gives Spearman correlations on the three Russian RuShiftEval test sets and their average. Untuned XL-LEXEME is competitive with the leaderboard systems; XL-LEXEME fine-tuned on RuSemShift obtains the highest reported correlation on each test set and the highest average. A dagger indicates that the leaderboard score is not statistically different from the untuned XL-LEXEME result. The paper states that untuned XL-LEXEME's differences from the two leading leaderboard systems are not significant, and that the fine-tuned model's differences from DeepMistake and GlossReader are also not significant.

    Dataset GlossReader DeepMistake UWB Baseline Cross-encoder XL-LEXEME XL-LEXEME fine-tuned
    RuShiftEval1 †\dagger0.781 †\dagger0.798 0.362 0.314 †\dagger0.727 0.775 0.799
    RuShiftEval2 †\dagger0.803 †\dagger0.773 0.354 0.302 †\dagger0.753 0.822 0.833
    RuShiftEval3 †\dagger0.822 †\dagger0.803 0.533 0.381 †\dagger0.748 0.809 0.842
    Average 0.802 0.791 0.417 0.332 0.743 0.802 0.825
  8. Knowl 8 — Cross-lingual transfer is promising but not established for semantic change

    limitation

    Because XL-LEXEME uses a shared multilingual encoder, its representations across languages may be comparable in a common geometric space. The paper does not evaluate cross-lingual semantic change, however, because suitable cross-lingual lexical-semantic-change resources are unavailable. The SemEval evaluation sets are also small: 37 English, 48 German, 31 Swedish, and 40 Latin target words. Latin results are poor, and the authors identify limited Latin representation in XLM-R training, lack of related-language coverage in the WiC training data, and the historical distance of the ancient Latin corpus as possible factors. They state that the effects of language representation in XLM-R and WiC training data cannot yet be quantified precisely.

Coverage note — No other substantial contributed material was omitted; target-word marking, representation construction, training configuration, evaluation procedure, benchmark results, efficiency comparison, and stated limitations are represented.

References

  1. 1.Nikolay Arefyev, Daniil Homskiy, Maksim Fedoseev, Adis Davletov, Vitaly Protasov, and Alexander Panchenko. 2021. DeepMistake: Which Senses are Hard to Distinguish for a WordinContext Model. In Computational Linguistics and Intellectual Technologies - Papers from the Annual International Conference "Dialogue" 2021, volume 2021-June. Section: 20.
  2. 2.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020a. Unsupervised Cross-lingual Representation Learning at Scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 8440–8451. Association for Computational Linguistics.
  3. 3.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020b. Unsupervised Cross-lingual Representation Learning at Scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 8440–8451. Association for Computational Linguistics.
  4. 4.Raia Hadsell, Sumit Chopra, and Yann LeCun. 2006. Dimensionality reduction by learning an invariant mapping. In 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR 2006), 17-22 June 2006, New York, NY, USA, pages 1735–1742. IEEE Computer Society.
  5. 5.William L. Hamilton, Jure Leskovec, and Dan Jurafsky. 2016. Diachronic Word Embeddings Reveal Statistical Laws of Semantic Change. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1489–1501, Berlin, Germany. Association for Computational Linguistics.
  6. 6.David R. Hardoon, Sandor Szedmak, and John Shawe-Taylor. 2004. Canonical Correlation Analysis: An Overview with Application to Learning Methods. Neural Computation, 16(12):2639–2664.
  7. 7.Taku Kudo and John Richardson. 2018. SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 66–71, Brussels, Belgium. Association for Computational Linguistics.
  8. 8.Severin Laicher, Sinan Kurtyigit, Dominik Schlechtweg, Jonas Kuhn, and Sabine Schulte im Walde. 2021. Explaining and improving BERT performance on lexical semantic change detection. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Student Research Workshop, pages 192–202, Online. Association for Computational Linguistics.
  9. 9.Qianchu Liu, Edoardo Maria Ponti, Diana McCarthy, Ivan Vulic, and Anna Korhonen. 2021. AM2iCo: Evaluating Word Meaning in Context across Low-Resource Languages with Adversarial Examples. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, pages 7151–7162. Association for Computational Linguistics.
  10. 10.Federico Martelli, Najla Kalach, Gabriele Tola, and Roberto Navigli. 2021. SemEval-2021 Task 2: Multilingual and Cross-lingual Word-in-Context Disambiguation (MCL-WiC). In Proceedings of the 15th International Workshop on Semantic Evaluation (SemEval-2021), pages 24–36, Online. Association for Computational Linguistics.
  11. 11.Tomás Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient estimation of word representations in vector space. In 1st International Conference on Learning Representations, ICLR 2013, Workshop Track Proceedings.
  12. 12.George A. Miller. 1992. WordNet: A Lexical Database for English. In Speech and Natural Language: Proceedings of a Workshop Held at Harriman, New York, February 23-26, 1992.
  13. 13.George A. Miller, Martin Chodorow, Shari Landes, Claudia Leacock, and Robert G. Thomas. 1994. Using a Semantic Concordance for Sense Identification. In Human Language Technology, Proceedings of a Workshop held at Plainsboro, New Jerey, USA, March 8-11, 1994. Morgan Kaufmann.
  14. 14.Stefano Montanelli and Francesco Periti. 2023. A survey on contextualised semantic shift detection. arXiv preprint arXiv:2304.01666.
  15. 15.Roberto Navigli. 2009. Word Sense Disambiguation: A Survey. ACM Comput. Surv., 41(2).
  16. 16.Mohammad Taher Pilehvar and José Camacho-Collados. 2019. WiC: the Word-in-Context Dataset for Evaluating Context-Sensitive Meaning Representations. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers), pages 1267–1273. Association for Computational Linguistics.
  17. 17.Ondrej Prazák, Pavel Pribán, and Stephen Taylor. 2021. UWB@ RuShiftEval Measuring Semantic Difference as per-word Variation in Aligned Semantic Spaces. In Computational Linguistics and Intellectual Technologies - Papers from the Annual International Conference "Dialogue" 2021, volume 2021-June. Section: 20.
  18. 18.Ondrej Prazák, Pavel Pribán, Stephen Taylor, and Jakub Sido. 2020. UWB at SemEval-2020 Task 1: Lexical Semantic Change Detection. In Proceedings of the Fourteenth Workshop on Semantic Evaluation, SemEval@COLING2020, pages 246–254. International Committee for Computational Linguistics.
  19. 19.William H. Press. 2002. Numerical recipes in C++: the art of scientific computing, 2nd Edition (C++ ed., print. is corrected to software version 2.10). Cambridge University Press.
  20. 20.Maxim Rachinskiy and Nikolay Arefyev. 2021. Zeroshot Crosslingual Transfer of a Gloss Language Model for Semantic Change Detection. In Computational Linguistics and Intellectual Technologies - Papers from the Annual International Conference "Dialogue" 2021, volume 2021-June. Section: 20.
  21. 21.Alessandro Raganato, Tommaso Pasini, José Camacho-Collados, and Mohammad Taher Pilehvar. 2020. XL-WiC: A Multilingual Benchmark for Evaluating Semantic Contextualization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 7193–7206. Association for Computational Linguistics.
  22. 22.Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3982–3992, Hong Kong, China. Association for Computational Linguistics.
  23. 23.Julia Rodina and Andrey Kutuzov. 2020. RuSemShift: a dataset of historical lexical semantic change in Russian. In Proceedings of the 28th International Conference on Computational Linguistics, pages 1037–1047, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  24. 24.Guy D. Rosin, Ido Guy, and Kira Radinsky. 2022. Time Masking for Temporal Language Models. In WSDM ’22: The Fifteenth ACM International Conference on Web Search and Data Mining, Virtual Event / Tempe, AZ, USA, February 21 - 25, 2022, pages 833–841. ACM.
  25. 25.Guy D. Rosin and Kira Radinsky. 2022. Temporal Attention for Language Models. CoRR, abs/2202.02093.
  26. 26.Dominik Schlechtweg, Barbara McGillivray, Simon Hengchen, Haim Dubossarsky, and Nina Tahmasebi. 2020. SemEval-2020 Task 1: Unsupervised Lexical Semantic Change Detection. In Proceedings of the Fourteenth Workshop on Semantic Evaluation, SemEval@COLING2020, pages 1–23. International Committee for Computational Linguistics.
  27. 27.Dominik Schlechtweg, Sabine Schulte im Walde, and Stefanie Eckmann. 2018. Diachronic Usage Relatedness (DURel): A Framework for the Annotation of Lexical Semantic Change. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 169–174, New Orleans, Louisiana. Association for Computational Linguistics.
  28. 28.Shuyi Xie, Jian Ma, Haiqin Yang, Lianxin Jiang, Yang Mo, and Jianping Shen. 2021. PALI at SemEval-2021 Task 2: Fine-Tune XLM-RoBERTa for Word in Context Disambiguation. In Proceedings of the 15th International Workshop on Semantic Evaluation (SemEval-2021), pages 713–718, Online. Association for Computational Linguistics.

Citation

MLA
Cassotti, P., et al. “XL-LEXEME: WiC Pretrained Model for Cross-Lingual LEXical sEMantic changE”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 2023, pp. 1577–85, https://doi.org/10.18653/V1/2023.ACL-SHORT.135.
APA
Cassotti, P., Siciliani, L., DeGemmis, M., Semeraro, G., & Basile, P. (2023). XL-LEXEME: WiC Pretrained Model for Cross-Lingual LEXical sEMantic changE. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 1577–1585. https://doi.org/10.18653/V1/2023.ACL-SHORT.135
Chicago
Cassotti, P., L. Siciliani, M. DeGemmis, G. Semeraro, and P. Basile. 2023. “XL-LEXEME: WiC Pretrained Model for Cross-Lingual LEXical sEMantic changE”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 1577–85. https://doi.org/10.18653/V1/2023.ACL-SHORT.135.
Harvard
Cassotti, P. et al. (2023) “XL-LEXEME: WiC Pretrained Model for Cross-Lingual LEXical sEMantic changE”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). Association for Computational Linguistics, pp. 1577–1585. Available at: https://doi.org/10.18653/V1/2023.ACL-SHORT.135.
Vancouver
1. Cassotti P, Siciliani L, DeGemmis M, Semeraro G, Basile P (2023) XL-LEXEME: WiC Pretrained Model for Cross-Lingual LEXical sEMantic changE. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). Association for Computational Linguistics, pp 1577–1585

BibTeX

@inproceedings{Cassotti_2023, title={XL-LEXEME: WiC Pretrained Model for Cross-Lingual LEXical sEMantic changE}, url={http://dx.doi.org/10.18653/V1/2023.ACL-SHORT.135}, DOI={10.18653/v1/2023.acl-short.135}, booktitle={Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)}, publisher={Association for Computational Linguistics}, author={Cassotti, Pierluigi and Siciliani, Lucia and DeGemmis, Marco and Semeraro, Giovanni and Basile, Pierpaolo}, year={2023}, pages={1577–1585} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/