Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages

Zihao LiYucheng ShiZirui LiuFan YangAli PayaniNinghao LiuMengnan Du

article2025AAAI74 citations

Proposes Language Ranker, a metric that quantifies and benchmarks multilingual large language model capabilities across low- and high-resource languages by measuring cosine similarity between internal hidden representations and an English baseline.

Listen

Modern large language models rely on vast amounts of training data that are heavily skewed toward English and a few high-resource languages, often allocating less than ten percent of their training tokens to all other languages combined. This uneven foundation creates sharp performance disparities, causing models to struggle with context, idioms, and reasoning in low-resource languages commonly spoken in developing regions. Despite the clear operational and equity risks of this divide, developers and decision-makers have lacked a standardized, quantitative method to evaluate model capabilities across diverse languages.

The article introduces and evaluates the Language Ranker, a language-agnostic metric designed to quantify and benchmark multilingual capabilities using a model's internal representations. By establishing English as a reference baseline and measuring how closely other languages map to it within the intermediate layers of a model, the framework systematically grades language-specific proficiency.

To validate this approach, the authors evaluated five prominent open-source model families—LlaMa2, LlaMa3, Qwen, Mistral, and Gemma—across 94 languages using parallel translation pairs from the OPUS-100 corpus and cross-lingual translation data from the Tatoeba Challenge. The methodology computes average cosine similarity between non-English sentences and their English counterparts across multiple Transformer layers, verifies these metrics against multi-language reasoning benchmarks such as ARC and MMLU, and examines geometric embedding space distributions across high- and low-resource languages.

The analysis produced three primary findings. First, representation similarity scores clearly separate high-resource and low-resource languages: high-resource languages such as German, Spanish, and French consistently achieved cosine similarities above 0.60, whereas low-resource languages like Igbo, Kazakh, and Kannada remained below 0.40. Second, model performance directly mirrors training corpus volume; languages representing greater shares of the pre-training data exhibited significantly higher similarity scores and superior reasoning accuracy. Third, geometric analysis revealed that high-resource languages occupy a balanced, multi-dimensional embedding space, whereas low-resource representations collapse into a narrow, constrained subspace, confirming that internal representation alignment directly dictates functional capability.

These findings indicate that internal representation similarity serves as a dependable, computationally efficient proxy for actual multilingual performance without requiring extensive task-specific testing for every language. For organizations deploying AI globally, these results highlight the substantial performance and reliability risks of using baseline models in low-resource regions, while showing that targeted fine-tuning on specific languages—as demonstrated by Qwen in Chinese—can successfully bridge representation gaps.

Leaders and practitioners should adopt representation-based benchmarking to audit multilingual capabilities before deploying systems across diverse language markets. Furthermore, development teams should use these metrics to optimize the composition of pre-training corpora and guide targeted fine-tuning investments to achieve broader global performance.

While the current evaluation relies primarily on English as a baseline and centers on 7-billion and 13-billion parameter open-source models, cross-checks using other high-resource baselines and larger architectures support high confidence in the overall framework. However, readers should note that the metric assesses semantic and structural alignment rather than subtle, dialect-specific nuances, indicating that full production deployments in critical low-resource settings should pair this tool with qualitative local evaluations.

Cover for Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages

Abstract

The development of Large Language Models (LLMs) relies on extensive text corpora, which are often unevenly distributed across languages. This imbalance results in LLMs performing significantly better on high-resource languages like English, German, and French, while their capabilities in low-resource languages remain inadequate. Currently, there is a lack of quantitative methods to evaluate the performance of LLMs in these low-resource languages. To address this gap, we propose the Language Ranker, an intrinsic metric designed to benchmark and rank languages based on LLM performance using internal representations. By comparing the LLM’s internal representation of various languages against a baseline derived from English, we can assess the model’s multilingual capabilities in a robust and language-agnostic manner. Our analysis reveals that high-resource languages exhibit higher similarity scores with English, demonstrating superior performance, while low-resource languages show lower similarity scores, underscoring the effectiveness of our metric in assessing language-specific capabilities. Besides, the experiments show that there is a strong correlation between the LLM’s performance in different languages and the proportion of those languages in its pre-training corpus. These insights underscore the efficacy of the Language Ranker as a tool for evaluating LLM performance across different languages, particularly those with limited resources.

Table of Contents

  • Introduction
  • The Proposed Method
  • Probing Datasets
  • Similarity Measurement
  • Rank Correlation Measurement
  • Experiments
  • Can Language Ranker Quantify LLM Performance Across Languages? (RQ1)
  • Comparison Across Different LLMs (RQ2)
  • Relationship to Ratio of Training Corpus? (RQ3)
  • Correlation with Other Inference Tasks? (RQ4)
  • Further Analysis of Proposed Metric
  • Why Using English as Baseline? (RQ5)
  • Deeper Analysis of the Embedding Space (RQ6)
  • Why Using Cosine Similarity? (RQ7)
  • Related Work
  • Conclusions and Future Work
  • References

Knowls

  1. Knowl 1 — Language Ranker Metric for Multilingual Representation Similarity

    model/method

    The Language Ranker evaluates an autoregressive large language model's (LLM's) multilingual capabilities by computing the layer-averaged cosine similarity between the final token internal representations of parallel sentences in English and a non-English target language.

    Given an English sentence X=(x1,x2,…,xn)X = (x_1, x_2, \dots, x_n) and its corresponding target language sentence Y=(y1,y2,…,ym)Y = (y_1, y_2, \dots, y_m), the representation vectors of the last tokens at transformer layer l∈{1,…,H}l \in \{1, \dots, H\} (where HH is the total number of layers) are denoted as xnlx_n^l and yml∈Rdy_m^l \in \mathbb{R}^d. The layer-wise intermediate hidden representation is given by: xl+1=MLP(xl+MHA(xl))x^{l+1} = \text{MLP}(x^l + \text{MHA}(x^l)) where MHA\text{MHA} denotes multi-head or multi-group attention, and MLP\text{MLP} denotes a multilayer perceptron layer.

    The cosine similarity between the two text representations at layer ii is calculated as: Simi=xni⋅ymi∥xni∥2∥ymi∥2\text{Sim}_i = \frac{x_n^i \cdot y_m^i}{\|x_n^i\|_2 \|y_m^i\|_2}

    To achieve a robust similarity measurement across network depths, the Language Ranker computes the arithmetic mean over a selected subset of intermediate layers lsub={5,10,15,20,25}l_{sub} = \{5, 10, 15, 20, 25\}: Sim=1∣lsub∣∑i∈lsubSimi\text{Sim} = \frac{1}{|l_{sub}|} \sum_{i \in l_{sub}} \text{Sim}_i

    The language-level score is computed by averaging Sim\text{Sim} across 2,000 parallel test sentences from the OPUS-100 dataset.

  2. Knowl 2 — Longest Common Partial Order Sublist Metric for Ranking Consistency

    definition

    To evaluate the consistency of multilingual capability rankings produced by two distinct large language models, the ranking similarity is computed as the normalized length of their longest common partial order sublist.

    Let A=(a1,a2,…,aN)A = (a_1, a_2, \dots, a_N) and B=(b1,b2,…,bN)B = (b_1, b_2, \dots, b_N) denote two ordered lists containing the same set of NN languages, sorted in descending order of their Language Ranker similarity scores. A sublist C=(c1,c2,…,ck)C = (c_1, c_2, \dots, c_k) is a common partial order sublist of AA and BB if C⊆AC \subseteq A, C⊆BC \subseteq B, and for every sequence of indices 1≤i1≤i2≤⋯≤ik≤k1 \le i_1 \le i_2 \le \dots \le i_k \le k, the relative ordering condition: IndexA(ci1)≤IndexA(ci2)≤⋯≤IndexA(cik)andIndexB(ci1)≤IndexB(ci2)≤⋯≤IndexB(cik)\text{Index}_A(c_{i_1}) \le \text{Index}_A(c_{i_2}) \le \dots \le \text{Index}_A(c_{i_k}) \quad \text{and} \quad \text{Index}_B(c_{i_1}) \le \text{Index}_B(c_{i_2}) \le \dots \le \text{Index}_B(c_{i_k}) holds simultaneously.

    The rank correlation between models AA and BB is defined as the ratio of the length of the longest common partial order sublist C∗C^* to the total number of ranked languages NN: Correlation(A,B)=∣C∗∣N\text{Correlation}(A, B) = \frac{|C^*|}{N}

  3. Knowl 3 — Double Variance Metric for Quantifying Subspace Distribution Quality

    model/method

    The double variance metric quantifies the geometric uniformity and isotropy of a language's embedding subspace in a large language model using principal component analysis (PCA), without referencing a baseline language.

    Let {Xi}i=1n⊂Rd\{X_i\}_{i=1}^n \subset \mathbb{R}^d be the centralized sentence embedding vectors for a given language. The projection variance along a unit direction vector ω∈Rd\omega \in \mathbb{R}^d (with ωTω=1\omega^T \omega = 1) is: Var(X,ω)=1n∑i=1n(XiTω)2=ωTCov(X)ω\text{Var}(X, \omega) = \frac{1}{n} \sum_{i=1}^n (X_i^T \omega)^2 = \omega^T \text{Cov}(X) \omega where Cov(X)=1n∑i=1nXiXiT\text{Cov}(X) = \frac{1}{n} \sum_{i=1}^n X_i X_i^T is the sample covariance matrix. Under PCA, the variances Var(X,ωi)\text{Var}(X, \omega_i) along the principal projection directions correspond to the eigenvalues λi\lambda_i of Cov(X)\text{Cov}(X).

    The double variance metric is defined as the sample variance of the first KK eigenvalues {λ1,λ2,…,λK}\{\lambda_1, \lambda_2, \dots, \lambda_K\}: DoubleVar(X)=1K∑i=1K(λi−λˉ)2,where λˉ=1K∑i=1Kλi\text{DoubleVar}(X) = \frac{1}{K} \sum_{i=1}^K (\lambda_i - \bar{\lambda})^2, \quad \text{where } \bar{\lambda} = \frac{1}{K} \sum_{i=1}^K \lambda_i

    A low double variance indicates balanced projection variance across directions and a well-distributed, isotropic embedding space (characteristic of high-resource languages). A high double variance indicates an unbalanced distribution where embeddings collapse into a narrow line (characteristic of low-resource languages).

  4. Knowl 4 — Disparities in LLM Internal Representation Similarities between High- and Low-Resource Languages

    empirical result

    Evaluating 7B-parameter models (LlaMa2 7B, Gemma 7B, Mistral-v0.1 7B, and Qwen 7B) and 13B-parameter models (LlaMa2 13B) using the Language Ranker across transformer layers reveals a clear gap between language resource levels:

    • High-resource languages (German, Spanish, French, Malay, Indonesian, Chinese) consistently maintain cosine similarity scores above 0.600.60 relative to English across intermediate layers, with Spanish and French typically displaying the highest scores.
    • Low-resource languages (Igbo, Kazakh, Kannada, Oriya, Turkmen) display substantially lower similarity scores, frequently remaining below 0.400.40.
    • In LlaMa2, Gemma, and Mistral, Chinese has slightly lower similarity scores than European high-resource languages, but in Qwen (which includes dedicated Chinese pre-training and fine-tuning), Chinese similarity reaches parity with other high-resource languages and improves in deeper layers.
    • In LlaMa2 13B, the separation between high- and low-resource languages matches that in LlaMa2 7B, while exhibiting larger representation fluctuations in earlier layers.
  5. Knowl 5 — Correlation between Training Corpus Proportion and Language Ranker Similarity in LlaMa2

    data/table

    In LlaMa2 (where English comprises 89.70% of the pre-training data), the Language Ranker similarity score of non-English languages relative to English strongly correlates with their proportion in the pre-training corpus. High-resource languages with pre-training proportions above 0.10% achieve similarity scores exceeding 0.55, whereas low-resource languages with pre-training proportions ≤0.01%\le 0.01\% achieve similarity scores below 0.40:

    Language Proportion (%) Similarity Language Proportion (%) Similarity
    German 0.17% 0.723 Welsh ≤0.01%\le 0.01\% 0.396
    French 0.16% 0.737 Persian ≤0.01%\le 0.01\% 0.300
    Swedish 0.15% 0.662 Urdu ≤0.01%\le 0.01\% 0.275
    Chinese 0.13% 0.552 Kannada ≤0.01%\le 0.01\% 0.236
  6. Knowl 6 — Reasoning Benchmark Accuracy Across High- and Low-Resource Languages on ARC and MMLU

    data/table

    Evaluation on 4-shot multiple-choice reasoning tasks from the MLMM-evaluation benchmark (100 randomly sampled questions per language from ARC and MMLU) reveals significant accuracy deficits in low-resource languages compared to high-resource languages across LlaMa2 7B, Gemma 7B, Mistral 7B, and Qwen 7B:

    Language ARC (4-shot Accuracy) MMLU (4-shot Accuracy)
    LlaMa2 7B Gemma 7B Mistral 7B Qwen 7B LlaMa2 7B Gemma 7B Mistral 7B Qwen 7B
    Chinese 27% 71% 57% 66% 32% 54% 37% 44%
    German 27% 68% 63% 32% 25% 57% 47% 27%
    French 31% 76% 59% 42% 24% 58% 48% 28%
    Spanish 31% 77% 60% 46% 29% 56% 52% 33%
    Italian 29% 77% 67% 44% 23% 56% 44% 32%
    Kannada 24% 48% 27% 21% 21% 40% 22% 19%
    Hindi 28% 60% 42% 22% 25% 45% 32% 23%
    Armenian 19% 40% 36% 20% 20% 36% 30% 25%
    Marathi 28% 46% 26% 25% 27% 42% 30% 26%
    Telugu 30% 42% 30% 30% 24% 33% 34% 23%

    High-resource languages achieve substantially higher accuracy on both benchmarks (e.g., Gemma reaches 68%--77% on ARC for high-resource vs. 40%--60% for low-resource; 54%--58% on MMLU for high-resource vs. 33%--45% for low-resource). LlaMa2 7B exhibits lower overall performance across all languages.

  7. Knowl 7 — Internal Representation Similarities Across Language Pair Resource Categories

    data/table

    Using parallel sentences from the Tatoeba-Challenge dataset across Gemma 7B and LlaMa2 7B, language pairs grouped into High Resource--High Resource (H-H), High Resource--Low Resource (H-L), and Low Resource--Low Resource (L-L) demonstrate that representations of high-resource languages cluster closely together, whereas low-resource languages diverge significantly from high-resource languages and each other:

    Pair Category Language Pair Gemma 7B Pair Category Language Pair Gemma 7B
    High–High English–German 0.72 High–Low German–Silesian 0.48
    High–High Italian–French 0.68 High–Low French–Erzya 0.35
    High–High German–French 0.67 High–Low Italian–Romany 0.32
    High–High French–Chinese 0.59 High–Low Italian–Uighur 0.27
    Low–Low Azerbaijani–Turkmen 0.51 Low–Low Hungarian–Yiddish 0.24
    Low–Low Kabyle–SMT 0.36 Low–Low Mari–Tatar 0.48

    For LlaMa2 7B, the corresponding similarity scores are:

    • High--High: English--German (0.72), Italian--French (0.69), German--French (0.68), French--Chinese (0.56).
    • High--Low: German--Silesian (0.44), French--Erzya (0.31), Italian--Romany (0.15), Italian--Uighur (0.20).
    • Low--Low: Azerbaijani--Turkmen (0.51), Hungarian--Yiddish (0.19), Kabyle--SMT (0.40), Mari--Tatar (0.42).

    Because all H-H pairs maintain high mutual similarity (0.56--0.72), choosing English as the primary baseline language is representative of high-resource baselines generally.

  8. Knowl 8 — Inverse Relationship Between Subspace Double Variance and Cosine Similarity

    data/table

    Comparing the PCA double variance of sentence embeddings against the Language Ranker cosine similarity for Gemma 7B and Mistral 7B reveals a strong negative relationship: high-resource languages exhibit low double variance and high similarity, whereas low-resource languages exhibit high double variance and low similarity:

    Language Gemma 7B Mistral 7B
    Double Variance Similarity Double Variance Similarity
    High-Resource Languages
    Italian 0.04 0.66 0.12 0.59
    French 0.09 0.69 0.17 0.65
    Spanish 0.06 0.68 0.17 0.64
    German 0.10 0.72 0.19 0.66
    Low-Resource Languages
    Nepali 0.75 0.42 0.60 0.28
    Kazakh 0.85 0.38 0.32 0.32
    Burmese 0.36 0.30 0.40 0.20
    Pashto 0.72 0.36 0.73 0.29

    This shows that lower similarity to English corresponds directly to structural degradation (narrow clustering and anisotropy) in the model's internal representation space.

  9. Knowl 9 — Multilingual Performance Ranking Consistency Across LLM Architectures

    empirical result

    Evaluating the rank correlation between language performance lists generated by four 7B models (LlaMa2 7B, Gemma 7B, Mistral 7B, and Qwen 7B) using the longest common partial order sublist metric shows consistent relative performance ordering across distinct architectures:

    • LlaMa2 vs. Gemma: correlation ratio = 0.33
    • LlaMa2 vs. Mistral: correlation ratio = 0.32
    • LlaMa2 vs. Qwen: correlation ratio = 0.33
    • Gemma vs. Mistral: correlation ratio = 0.41
    • Gemma vs. Qwen: correlation ratio = 0.38
    • Mistral vs. Qwen: correlation ratio = 0.40

    Despite variations in training corpora, tokenizer designs, and network architectures, the relative proficiency hierarchy across languages remains largely consistent across models.

Coverage note — Visual 2D PCA scatter plots (Figure 4) and box plots (Figure 5) were summarized conceptually within the double variance and embedding distribution knowls rather than transcribed as raw coordinates.

References

  1. 1.Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774.
  2. 2.Ahuja, S.; Aggarwal, D.; Gumma, V.; Watts, I.; Sathe, A.; Ochieng, M.; Hada, R.; Jain, P.; Axmed, M.; Bali, K.; and Sitaram, S. 2024. MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks. arXiv:2311.07463.
  3. 3.Bai, J.; Bai, S.; Chu, Y.; Cui, Z.; Dang, K.; Deng, X.; Fan, Y.; Ge, W.; Han, Y.; Huang, F.; et al. 2023. Qwen technical report. arXiv preprint arXiv:2309.16609.
  4. 4.Blasi, D.; Anastasopoulos, A.; and Neubig, G. 2021. Systematic Inequalities in Language Technology Performance across the World’s Languages. arXiv preprint arXiv:2110.06733.
  5. 5.Clark, P.; Cowhey, I.; Etzioni, O.; Khot, T.; Sabharwal, A.; Schoenick, C.; and Tafjord, O. 2018. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457.
  6. 6.Conneau, A.; Khandelwal, K.; Goyal, N.; Chaudhary, V.; Wenzek, G.; Guzman, F.; Grave, E.; Ott, M.; Zettlemoyer, L.; and Stoyanov, V. 2019. Unsupervised cross-lingual representation learning at scale. arXiv preprint arXiv:1911.02116.
  7. 7.Gurnee, W.; and Tegmark, M. 2023. Language models represent space and time. arXiv preprint arXiv:2310.02207.
  8. 8.Hendrycks, D.; Burns, C.; Basart, S.; Zou, A.; Mazeika, M.; Song, D.; and Steinhardt, J. 2020. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300.
  9. 9.Huang, H.; Tang, T.; Zhang, D.; Zhao, W. X.; Song, T.; Xia, Y.; and Wei, F. 2023. Not all languages are created equal in llms: Improving multilingual capability by cross-lingual-thought prompting. arXiv preprint arXiv:2305.07004.
  10. 10.Jiang, A. Q.; Sablayrolles, A.; Mensch, A.; Bamford, C.; Chaplot, D. S.; Casas, D. d. l.; Bressand, F.; Lengyel, G.; Lample, G.; Saulnier, L.; et al. 2023. Mistral 7B. arXiv preprint arXiv:2310.06825.
  11. 11.Kalyan, K. S.; Rajasekharan, A.; and Sangeetha, S. 2021. Ammus: A survey of transformer-based pretrained models in natural language processing. arXiv preprint arXiv:2108.05542.
  12. 12.Lankford, S.; Afli, H.; and Way, A. 2024. Transformers for Low-Resource Languages: Is F'eidir Linn! arXiv preprint arXiv:2403.01985.
  13. 13.Li, K.; Patel, O.; Viégas, F.; Pfister, H.; and Wattenberg, M. 2024. Inference-time intervention: Eliciting truthful answers from a language model. Advances in Neural Information Processing Systems, 36.
  14. 14.Liu, C.; Zhang, W.; Zhao, Y.; Luu, A. T.; and Bing, L. 2024. Is Translation All You Need? A Study on Solving Multilingual Tasks with Large Language Models. arXiv preprint arXiv:2403.10258.
  15. 15.Marks, S.; and Tegmark, M. 2023. The geometry of truth: Emergent linear structure in large language model representations of true/false datasets. arXiv preprint arXiv:2310.06824.
  16. 16.Meta-AI. 2024. LlaMa3. https://github.com/meta-llama/llama3. Accessed: 2024-06-14.
  17. 17.OpenAI. 2023. GPT-3 Dataset Statistics. https://github.com/openai/gpt-3/tree/master/dataset_statistics.
  18. 18.Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. 2022. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35: 27730–27744.
  19. 19.Qin, L.; Chen, Q.; Wei, F.; Huang, S.; and Che, W. 2023. Cross-lingual prompting: Improving zero-shot chain-of-thought reasoning across languages. arXiv preprint arXiv:2310.14799.
  20. 20.Qin, L.; Chen, Q.; Zhou, Y.; Chen, Z.; Li, Y.; Liao, L.; Li, M.; Che, W.; and Yu, P. S. 2024. Multilingual large language model: A survey of resources, taxonomy and frontiers. arXiv preprint arXiv:2404.04925.
  21. 21.Schäfer, A.; Ravfogel, S.; Hofmann, T.; Pimentel, T.; and Schlag, I. 2024. Language Imbalance Can Boost Cross-lingual Generalisation. arXiv preprint arXiv:2404.07982.
  22. 22.Shen, L.; Tan, W.; Chen, S.; Chen, Y.; Zhang, J.; Xu, H.; Zheng, B.; Koehn, P.; and Khashabi, D. 2024. The language barrier: Dissecting safety challenges of llms in multilingual contexts. arXiv preprint arXiv:2401.13136.
  23. 23.Steck, H.; Ekanadham, C.; and Kallus, N. 2024. Is Cosine-Similarity of Embeddings Really About Similarity? In Companion Proceedings of the ACM on Web Conference 2024, WWW ’24. ACM.
  24. 24.Team, G.; Mesnard, T.; Hardin, C.; Dadashi, R.; Bhupatiraju, S.; Pathak, S.; Sifre, L.; Rivière, M.; Kale, M. S.; Love, J.; et al. 2024. Gemma: Open models based on gemini research and technology. arXiv preprint arXiv:2403.08295.
  25. 25.Tiedemann, J. 2020. The Tatoeba Translation Challenge – Realistic Data Sets for Low Resource and Multilingual MT. In Proceedings of the Fifth Conference on Machine Translation, 1174–1182. Online: Association for Computational Linguistics.
  26. 26.Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  27. 27.Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017. Attention is all you need. Advances in neural information processing systems, 30.
  28. 28.Wendler, C.; Veselovsky, V.; Monea, G.; and West, R. 2024. Do Llamas Work in English? On the Latent Language of Multilingual Transformers. arXiv:2402.10588.
  29. 29.Xie, Y.; Chen, X.; Zhan, H.; Shivakumara, P.; Yin, B.; Liu, C.; and Lu, Y. 2024. Weakly supervised scene text generation for low-resource languages. Expert Systems with Applications, 237: 121622.
  30. 30.Zellers, R.; Holtzman, A.; Bisk, Y.; Farhadi, A.; and Choi, Y. 2019. Hellaswag: Can a machine really finish your sentence? arXiv preprint arXiv:1905.07830.
  31. 31.Zhang, B.; Williams, P.; Titov, I.; and Sennrich, R. 2020. Improving Massively Multilingual Neural Machine Translation and Zero-Shot Translation. In Jurafsky, D.; Chai, J.; Schluter, N.; and Tetreault, J., eds., Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 1628–1639. Online: Association for Computational Linguistics.
  32. 32.Zhang, X.; Li, S.; Hauer, B.; Shi, N.; and Kondrak, G. 2023. Don’t trust ChatGPT when your question is not in English: A study of multilingual abilities and types of LLMs. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 7915–7927.
  33. 33.Zou, A.; Phan, L.; Chen, S.; Campbell, J.; Guo, P.; Ren, R.; Pan, A.; Yin, X.; Mazeika, M.; Dombrowski, A.-K.; et al. 2023. Representation engineering: A top-down approach to ai transparency. arXiv preprint arXiv:2310.01405.

Citation

MLA
Li, Z., et al. “Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages”. Proceedings of the AAAI Conference on Artificial Intelligence, vol. 39, no. 27, 2025, pp. 28186–94, https://doi.org/10.1609/aaai.v39i27.35038.
APA
Li, Z., Shi, Y., Liu, Z., Yang, F., Payani, A., Liu, N., & Du, M. (2025). Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages. Proceedings of the AAAI Conference on Artificial Intelligence, 39(27), 28186–28194. https://doi.org/10.1609/aaai.v39i27.35038
Chicago
Li, Z., Y. Shi, Z. Liu, et al. 2025. “Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages”. Proceedings of the AAAI Conference on Artificial Intelligence 39 (27): 28186–94. https://doi.org/10.1609/aaai.v39i27.35038.
Harvard
Li, Z. et al. (2025) “Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages”, Proceedings of the AAAI Conference on Artificial Intelligence, 39(27), pp. 28186–28194. Available at: https://doi.org/10.1609/aaai.v39i27.35038.
Vancouver
1. Li Z, Shi Y, Liu Z, Yang F, Payani A, Liu N, Du M (2025) Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages. Proceedings of the AAAI Conference on Artificial Intelligence 39:28186–28194

BibTeX

@article{Li_2025, title={Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages}, volume={39}, ISSN={2159-5399}, url={http://dx.doi.org/10.1609/aaai.v39i27.35038}, DOI={10.1609/aaai.v39i27.35038}, number={27}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, publisher={Association for the Advancement of Artificial Intelligence (AAAI)}, author={Li, Zihao and Shi, Yucheng and Liu, Zirui and Yang, Fan and Payani, Ali and Liu, Ninghao and Du, Mengnan}, year={2025}, month=Apr, pages={28186–28194} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF