keyword
high-resource languages
High-resource languages are languages that possess an abundance of digital data, annotated corpora, and technological tools necessary to train, evaluate, and deploy machine learning and natural language processing models effectively. In computational contexts, this classification encompasses widely documented human spoken languages as well as prevalent programming languages that are heavily represented in large-scale datasets and standardized benchmarks. Because modern artificial intelligence systems rely heavily on massive volumes of text and code for pre-training, high-resource languages consistently achieve higher performance, accuracy, and linguistic fluency in automated tasks compared to languages with limited digital records.
3 items

Measuring the Impact of Programming Language Distribution
Gabriel Orlanski, Kefan Xiao, Xavier Garcia, Jeffrey Hui, Joshua Howland, Jonathan Malmaud, Jacob Austin, Rishabh Singh, Michele Catasta
Why you should read this
Presents the BabelCode execution-based evaluation framework and demonstrates that balancing multilingual training distributions dramatically improves code language model performance on low-resource programming languages with minimal degradation to high-resource ones.
Current benchmarks for evaluating neural code models focus on only a small subset of programming languages, excluding many popular languages such as Go or Rust. To ameliorate this issue, we present the BabelCode framework for execution-based evaluation of any benchmark in any language. BabelCode enables new investigations into the qualitative performance of models’ memory, runtime, and individual test case results. Additionally, we present a new code translation dataset called Translating Python Programming Puzzles (TP3) from the Python Programming Puzzles (Schuster et al., 2021) benchmark that involves translating expert-level python functions to any language. With both BabelCode and the TP3 benchmark, we investigate if balancing the distributions of 14 languages in a training dataset improves a large language model’s performance on low-resource languages. Training a model on a balanced corpus results in, on average, 12.34% higher pass@k across all tasks and languages compared to the baseline. We find that this strategy achieves 66.48% better pass@k on low-resource languages at the cost of only a 12.94% decrease to high-resource languages. In our three translation tasks, this strategy yields, on average, 30.77% better low-resource pass@k while having 19.58% worse high-resource pass@k.
Added
2026-10-03

AI Diffusion in Low Resource Language Countries
Amit Misra, Syed Waqas Zamir, Wassim Hamidouche, Inbal Becker-Reshef, Juan M. Lavista Ferres
Why you should read this
Demonstrates that low-resource language countries suffer a twenty percent reduction in artificial intelligence adoption rates, isolating linguistic accessibility as an independent barrier to global technology diffusion.
Artificial intelligence (AI) is diffusing globally at unprecedented speed, but adoption remains uneven. Frontier Large Language Models (LLMs) are known to perform poorly on low-resource languages due to data scarcity. We hypothesize that this performance deficit reduces the utility of AI, thereby slowing adoption in Low-Resource Language Countries (LRLCs). To test this, we use a weighted regression model to isolate the language effect from socioeconomic and demographic factors, finding that LRLCs have a share of AI users that is approximately 20% lower relative to their baseline. These results indicate that linguistic accessibility is a significant, independent barrier to equitable AI diffusion.
Added
2026-09-29

Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages
Zihao Li, Yucheng Shi, Zirui Liu, Fan Yang, Ali Payani, Ninghao Liu, Mengnan Du
Why you should read this
Proposes Language Ranker, a metric that quantifies and benchmarks multilingual large language model capabilities across low- and high-resource languages by measuring cosine similarity between internal hidden representations and an English baseline.
The development of Large Language Models (LLMs) relies on extensive text corpora, which are often unevenly distributed across languages. This imbalance results in LLMs performing significantly better on high-resource languages like English, German, and French, while their capabilities in low-resource languages remain inadequate. Currently, there is a lack of quantitative methods to evaluate the performance of LLMs in these low-resource languages. To address this gap, we propose the Language Ranker, an intrinsic metric designed to benchmark and rank languages based on LLM performance using internal representations. By comparing the LLM’s internal representation of various languages against a baseline derived from English, we can assess the model’s multilingual capabilities in a robust and language-agnostic manner. Our analysis reveals that high-resource languages exhibit higher similarity scores with English, demonstrating superior performance, while low-resource languages show lower similarity scores, underscoring the effectiveness of our metric in assessing language-specific capabilities. Besides, the experiments show that there is a strong correlation between the LLM’s performance in different languages and the proportion of those languages in its pre-training corpus. These insights underscore the efficacy of the Language Ranker as a tool for evaluating LLM performance across different languages, particularly those with limited resources.
Added
2026-09-26
