Built independently by an author, for readers. Read the story and support ChapterPal

keyword

vocabulary overlap

Vocabulary overlap refers to the degree or proportion of shared words, subwords, or tokens present between two or more distinct text corpora, vocabularies, or languages. In natural language processing and machine learning, it serves as a measure of lexical similarity to determine how much the terminology of one text distribution matches another, such as comparing a general pretraining dataset to specialized domain data, downstream task prompts, or different languages within multilingual systems. A higher vocabulary overlap indicates a substantial common lexical foundation, which generally facilitates effective knowledge transfer, cross-lingual consistency, and model adaptation, whereas low overlap signals domain or linguistic divergence that may require specialized tokenization, vocabulary expansion, or targeted domain-adaptive training.

3 items

On the Effect of Pretraining Corpora on In-context Learning by a Large-scale Language Model

On the Effect of Pretraining Corpora on In-context Learning by a Large-scale Language Model

Seongjin Shin, Sang-Woo Lee, Hwijeen Ahn, Sungdong Kim, HyoungSeok Kim, Boseop Kim, Kyunghyun Cho, Gichang Lee, Woo-Myoung Park, Jung-Woo Ha, Nako Sung

OrganizationsClova AI ResearchNaver AI LabNew York University

Why you should read this

Reveals that large language model in-context learning capabilities depend heavily on pretraining data sources and combinations rather than corpus size alone, showing that low validation perplexity does not reliably predict few-shot performance.

Many recent studies on large-scale language models have reported successful in-context zero- and few-shot learning ability. However, the in-depth analysis of when in-context learning occurs is still lacking. For example, it is unknown how in-context learning performance changes as the training corpus varies. Here, we investigate the effects of the source and size of the pretraining corpus on in-context learning in HyperCLOVA, a Korean-centric GPT-3 model. From our in-depth investigation, we introduce the following observations: (1) in-context learning performance heavily depends on the corpus domain source, and the size of the pretraining corpus does not necessarily determine the emergence of in-context learning, (2) in-context learning ability can emerge when a language model is trained on a combination of multiple corpora, even when each corpus does not result in in-context learning on its own, (3) pretraining with a corpus related to a downstream task does not always guarantee the competitive in-context learning performance of the downstream task, especially in the few-shot setting, and (4) the relationship between language modeling (measured in perplexity) and in-context learning does not always correlate: e.g., low perplexity does not always imply high in-context few-shot learning performance.

Added

2026-10-01

Cross-Lingual Consistency of Factual Knowledge in Multilingual Language Models

Cross-Lingual Consistency of Factual Knowledge in Multilingual Language Models

Jirui Qi, Raquel Fernández, Arianna Bisazza

OrganizationsUniversity of AmsterdamUniversity of Groningen

Why you should read this

Introduces a ranking-based consistency metric and a multi-parallel benchmark to evaluate factual knowledge alignment across languages in multilingual language models, showing that scaling model size fails to improve cross-lingual consistency and that consistency scores predict whether edited facts transfer across languages.

Multilingual large-scale Pretrained Language Models (PLMs) have been shown to store considerable amounts of factual knowledge, but large variations are observed across languages. With the ultimate goal of ensuring that users with different language backgrounds obtain consistent feedback from the same model, we study the cross-lingual consistency (CLC) of factual knowledge in various multilingual PLMs. To this end, we propose a Ranking-based Consistency (RankC) metric to evaluate knowledge consistency across languages independently from accuracy. Using this metric, we conduct an in-depth analysis of the determining factors for CLC, both at model level and at language-pair level. Among other results, we find that increasing model size leads to higher factual probing accuracy in most languages, but does not improve cross-lingual consistency. Finally, we conduct a case study on CLC when new factual associations are inserted in the PLMs via model editing. Results on a small sample of facts inserted in English reveal a clear pattern whereby the new piece of knowledge transfers only to languages with which English has a high RankC score.

Added

2026-10-01