AI Diffusion in Low Resource Language Countries
Amit MisraSyed Waqas ZamirWassim HamidoucheInbal Becker-ReshefJuan M. Lavista Ferres
Demonstrates that low-resource language countries suffer a twenty percent reduction in artificial intelligence adoption rates, isolating linguistic accessibility as an independent barrier to global technology diffusion.
Artificial intelligence is expanding globally at a rapid pace, yet its adoption varies widely across nations. While past digital divides have been driven by economic and infrastructural gaps, modern generative AI relies heavily on large language models that depend on vast quantities of online text. Because digital text is overwhelmingly concentrated in a few high-resource languages, models perform significantly worse in low-resource languages, creating a potential barrier to AI adoption for non-English and underrepresented language communities.
The article aims to evaluate whether language resourcing independently drives country-level AI adoption and to quantify the adoption gap between Low-Resource Language Countries and countries with higher-resource languages.
To conduct this evaluation, the authors built a taxonomy classifying languages into high, mid, and low resources based on digital presence and model performance, assigning readiness labels to countries using dominant national language data. They merged this taxonomy with AI tool usage telemetry across 147 countries, controlling for key socioeconomic and demographic factors, including gross domestic product per capita, electricity access, internet penetration, and population age structure. The analysis used weighted regression models and difference-in-differences estimators to isolate the specific impact of language availability.
The findings show that raw AI user share in Low-Resource Language Countries is less than half that of other nations, recording 9.9% in 2025 compared to 21.3% for non-low-resource countries. After controlling for income and infrastructure, Low-Resource Language Countries still face an estimated 2.1 percentage point adoption deficit, representing an approximate 20% shortfall relative to their baseline adoption rate. Furthermore, longitudinal analysis shows no statistically significant widening or narrowing of this gap between 2024 and 2025, indicating that the disparity is persistent rather than closing naturally.
These results demonstrate that economic growth and digital infrastructure alone are insufficient to guarantee inclusive technological diffusion; linguistic accessibility operates as an independent obstacle. Without intervention, populations speaking low-resource languages risk exclusion from AI productivity benefits, compounding existing global economic disparities.
The article concludes that closing this adoption gap requires direct action to build high-quality digital training datasets for low-resource languages, as technological workarounds cannot replace sufficient data. Research and policy initiatives must prioritize multilingual data collection to make AI accessible across linguistic boundaries.
Confidence in these findings is supported by consistent results across multiple statistical weighting and regression models. However, readers should consider certain limitations: country-level language classification is complicated by widespread multilingualism and subnational literacy or urban-rural disparities, and the one-year dataset offers a limited time horizon that requires longer-term monitoring.
- Paper: The State and Fate of Linguistic Diversity and Inclusion in the NLP World, Pratik Joshi et al. (2020). This paper establishes the foundational taxonomy and global digital resource disparities across languages, providing essential context for the data scarcity underlying low-resource language deficits.
- Paper: MEGA: Multilingual Evaluation of Generative AI, Kabir Ahuja et al. (2023). This study demonstrates the generative performance deficits of frontier LLMs on low-resource and non-Latin script languages, supplying the direct technical premise for reduced user utility.
- Paper: The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants, Lucas Bandarkar et al. (2024). This benchmark provides critical empirical evidence of how English-centric large language models exhibit severe performance drop-offs on medium- and low-resource languages.
- Paper: MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks, Sanchit Ahuja et al. (2024). This work documents the wide performance divide across global language families in modern foundation models, corroborating the technological barriers facing low-resource language regions.
- Paper: No Language Left Behind: Scaling Human-Centered Machine Translation, NLLB Team et al. (2022). This paper analyzes the systemic technological exclusion of underserved linguistic communities and the technical hurdles of expanding language support beyond high-resource languages.
- Paper: Unintended Impacts of LLM Alignment on Global Representation, Michael J. Ryan et al. (2024). This paper explores how post-training alignment widens disparities across global non-Western dialects and regions, offering key insight into why frontier models underperform globally.
- Paper: Large Language Models are Geographically Biased, Rohin Manvi et al. (2024). This study details systemic geographic and socioeconomic disparities in foundation models, contextualizing the non-linguistic factors that interact with language accessibility in developing regions.
- Paper: Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages, Zihao Li et al. (2025). This work introduces a quantitative, representation-level metric to systematically grade and rank LLM proficiency across high- and low-resource languages, operationalizing the linguistic performance gaps identified in the source.
- Paper: Designing Culturally Aligned AI Systems For Social Good in Non-Western Contexts, Deepak Varuvel Dennison et al. (2025). This paper investigates real-world deployments and practical design requirements for contextualizing and deploying AI systems across non-Western, low-resource language environments to bridge adoption barriers.
- Paper: Omnilingual MT: Machine Translation for 1,600 Languages, Omnilingual MT Team et al. (2026). This work presents an architectural approach to expand translation capabilities across more than 1,600 languages, directly tackling the generation bottleneck in low-resource communities.
