XL-LEXEME: WiC Pretrained Model for Cross-Lingual LEXical sEMantic changE
Pierluigi CassottiLucia SicilianiMarco de GemmisGiovanni SemeraroPierpaolo Basile
Presents XL-LEXEME, a bi-encoder model that adapts Sentence-BERT to the Word-in-Context task with target-word highlighting to generate scalable, comparable lexical representations that achieve state-of-the-art semantic change detection across multiple languages.
Tracking how word meanings evolve over time—known as lexical semantic change detection—is essential for accurate historical text analysis, information retrieval, and language technology. Traditional methods often depend on rigid dictionary sense inventories that fail to capture emerging or obsolete meanings. While modern deep learning models can determine whether a word shares the same meaning across different contexts, existing architectures typically rely on joint-sentence cross-encoders. These cross-encoders are computationally expensive and cannot produce standalone, directly comparable word representations across vast historical corpora.
The article demonstrates an efficient, multilingual approach called XL-LEXEME, which adapts Sentence-BERT architectures using target word delimiters to perform lexical semantic change detection. The model evaluates whether knowledge learned from synchronic Word-in-Context datasets can successfully transfer to historical, diachronic language shifts across five languages: English, German, Swedish, Latin, and Russian.
To achieve this, the authors built a Siamese bi-encoder network utilizing XLM-RoBERTa Large and trained it with contrastive loss across merged multilingual Word-in-Context benchmark datasets. Target words within input sentences are marked with special delimiters, allowing the network to encode sentences independently into separate, comparable vector representations. The model was evaluated against established historical benchmarks, specifically the SemEval-2020 Task 1 ranking subtask across four languages and the RuShiftEval benchmark across three historical periods in Russian.
The evaluation yielded several key findings. First, XL-LEXEME outperformed previous state-of-the-art systems and baselines in English (0.757 correlation), German (0.877 correlation), and Swedish (0.754 correlation). Second, on the Russian RuShiftEval benchmark, XL-LEXEME achieved state-of-the-art performance with an average correlation of 0.802, which further increased to 0.825 when fine-tuned on target historical data. Third, the Siamese bi-encoder architecture reduced theoretical computational complexity compared to standard cross-encoders while maintaining superior accuracy. Finally, the model failed on Latin (-0.056 correlation), demonstrating no significant alignment with human annotations.
These findings indicate that general-purpose contextual training can effectively identify semantic evolution over centuries without needing explicit temporal annotations or costly cross-encoding. This reduction in computational requirements lowers infrastructure costs and processing timelines when analyzing large-scale text archives. However, the contrast between strong performance on modern languages and failure on Latin shows that the model relies heavily on language representation within the underlying pre-trained multilingual model.
Stakeholders and practitioners analyzing evolving terminology across large archives can deploy XL-LEXEME as an efficient, high-performing alternative to heavy cross-encoder pipelines. When applying the model to new languages, teams should prioritize languages well-represented in underlying foundation models or related language families. Future work should focus on developing dedicated cross-lingual semantic change evaluation benchmarks and expanding training coverage for low-resource and ancient languages.
Confidence in the reported improvements is high for modern European languages with sufficient pre-training representation. However, users should exercise caution regarding small evaluation sample sizes (ranging from 31 to 48 target words per language in SemEval-2020) and acknowledge performance limitations on ancient or severely underrepresented languages.
- Paper: Cross-lingual Language Model Pretraining, Guillaume Lample et al. (2019). Its multilingual pretraining methods provide context for the cross-lingual Transformer representations that XL-LEXEME adapts.
- Paper: Making Monolingual Sentence Embeddings Multilingual Using Knowledge Distillation, Nils Reimers et al. (2020). Its XLM-RoBERTa student encoder shows how multilingual sentence representations can be distilled into a bi-encoder, a key design choice behind XL-LEXEME.
- Paper: Language-agnostic BERT Sentence Embedding, Fangxiaoyu Feng et al. (2020). LaBSE’s multilingual dual-encoder approach prepares readers to understand XL-LEXEME’s independently encoded, comparable sentence vectors.
- Paper: SimCSE: Simple Contrastive Learning of Sentence Embeddings, Tianyu Gao et al. (2021). SimCSE introduces contrastive training for sentence embeddings, clarifying the learning objective XL-LEXEME adapts to word-in-context comparisons.
No sufficiently relevant recommendations were found.
