Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages
Zihao LiYucheng ShiZirui LiuFan YangAli PayaniNinghao LiuMengnan Du
Proposes Language Ranker, a metric that quantifies and benchmarks multilingual large language model capabilities across low- and high-resource languages by measuring cosine similarity between internal hidden representations and an English baseline.
Modern large language models rely on vast amounts of training data that are heavily skewed toward English and a few high-resource languages, often allocating less than ten percent of their training tokens to all other languages combined. This uneven foundation creates sharp performance disparities, causing models to struggle with context, idioms, and reasoning in low-resource languages commonly spoken in developing regions. Despite the clear operational and equity risks of this divide, developers and decision-makers have lacked a standardized, quantitative method to evaluate model capabilities across diverse languages.
The article introduces and evaluates the Language Ranker, a language-agnostic metric designed to quantify and benchmark multilingual capabilities using a model's internal representations. By establishing English as a reference baseline and measuring how closely other languages map to it within the intermediate layers of a model, the framework systematically grades language-specific proficiency.
To validate this approach, the authors evaluated five prominent open-source model families—LlaMa2, LlaMa3, Qwen, Mistral, and Gemma—across 94 languages using parallel translation pairs from the OPUS-100 corpus and cross-lingual translation data from the Tatoeba Challenge. The methodology computes average cosine similarity between non-English sentences and their English counterparts across multiple Transformer layers, verifies these metrics against multi-language reasoning benchmarks such as ARC and MMLU, and examines geometric embedding space distributions across high- and low-resource languages.
The analysis produced three primary findings. First, representation similarity scores clearly separate high-resource and low-resource languages: high-resource languages such as German, Spanish, and French consistently achieved cosine similarities above 0.60, whereas low-resource languages like Igbo, Kazakh, and Kannada remained below 0.40. Second, model performance directly mirrors training corpus volume; languages representing greater shares of the pre-training data exhibited significantly higher similarity scores and superior reasoning accuracy. Third, geometric analysis revealed that high-resource languages occupy a balanced, multi-dimensional embedding space, whereas low-resource representations collapse into a narrow, constrained subspace, confirming that internal representation alignment directly dictates functional capability.
These findings indicate that internal representation similarity serves as a dependable, computationally efficient proxy for actual multilingual performance without requiring extensive task-specific testing for every language. For organizations deploying AI globally, these results highlight the substantial performance and reliability risks of using baseline models in low-resource regions, while showing that targeted fine-tuning on specific languages—as demonstrated by Qwen in Chinese—can successfully bridge representation gaps.
Leaders and practitioners should adopt representation-based benchmarking to audit multilingual capabilities before deploying systems across diverse language markets. Furthermore, development teams should use these metrics to optimize the composition of pre-training corpora and guide targeted fine-tuning investments to achieve broader global performance.
While the current evaluation relies primarily on English as a baseline and centers on 7-billion and 13-billion parameter open-source models, cross-checks using other high-resource baselines and larger architectures support high confidence in the overall framework. However, readers should note that the metric assesses semantic and structural alignment rather than subtle, dialect-specific nuances, indicating that full production deployments in critical low-resource settings should pair this tool with qualitative local evaluations.
- Paper: MEGA: Multilingual Evaluation of Generative AI, Kabir Ahuja et al. (2023). Reading MEGA provides essential background on the cross-lingual performance disparities across high- and low-resource languages in modern LLMs that Language Ranker seeks to efficiently quantify.
- Paper: XNLI: Evaluating Cross-lingual Sentence Representations, Alexis Conneau et al. (2018). Understanding XNLI's foundational cross-lingual evaluation framework clarifies why sentence-level multilingual representations and English-anchored transfer benchmarks are central to evaluating model capabilities.
- Paper: The State and Fate of Linguistic Diversity and Inclusion in the NLP World, Pratik Joshi et al. (2020). This paper establishes the taxonomy and empirical disparities between high- and low-resource languages across NLP, motivating the resource divide that Language Ranker measures.
- Paper: Language-agnostic BERT Sentence Embedding, Fangxiaoyu Feng et al. (2020). LaBSE introduces key mechanisms for mapping multilingual sentences into a shared, language-agnostic embedding space that directly precedes representation-similarity evaluation methods.
- Paper: Unsupervised Cross-lingual Representation Learning at Scale, Alexis Conneau et al. (2019). XLM-R demonstrates the scaling properties and capacity trade-offs of multilingual representations across 100 languages, providing key foundational context for embedding alignment.
- Paper: No Language Left Behind: Scaling Human-Centered Machine Translation, NLLB Team et al. (2022). NLLB details large-scale multilingual evaluation and translation data across 200+ languages, illustrating the low-resource representation challenges tackled by the Language Ranker metric.
- Paper: Making Monolingual Sentence Embeddings Multilingual Using Knowledge Distillation, Nils Reimers et al. (2020). This work explains how multilingual sentence embeddings are aligned to English teacher models, laying the conceptual groundwork for using English-anchored cosine similarities as a capability proxy.
- Paper: Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages, Ayyoob Imani et al. (2023). Glot500 provides crucial empirical evidence on scaling LLM pretraining to tail languages, illustrating the representation collapse in low-resource settings.
- Paper: TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI Agents, Bofei Zhang et al. (2026). TongUI extends the study of internal LLM representations by probing neuron-level activations and shared 'Lingua Franca' mechanisms across languages, offering a mechanistic continuation of Language Ranker's geometric findings.
- Paper: MMTEB: Massive Multilingual Text Embedding Benchmark, Kenneth C. Enevoldsen et al. (2025). MMTEB builds on multilingual embedding assessment by providing a massive, 500-task downstream benchmark across 250+ languages to test whether internal alignment correlates with diverse task performance.
- Paper: Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models, Yanzhao Zhang et al. (2025). Qwen3 Embedding applies large-scale foundation model representations to downstream multilingual embedding and reranking tasks, demonstrating how model families evaluated in Language Ranker can be adapted in practice.
- Paper: Omnilingual MT: Machine Translation for 1,600 Languages, Omnilingual MT Team et al. (2026). Omnilingual MT extends cross-lingual scaling to over 1,600 languages, providing an extreme-scale testbed for the low-resource representation and evaluation principles highlighted in Language Ranker.
