One Country, 700+ Languages: NLP Challenges for Underrepresented Languages and Dialects in Indonesia
Alham Fikri AjiGenta Indra WinataFajri KotoSamuel CahyawijayaAde RomadhonyRahmad MahendraKemal KurniawanDavid MoeljadiRadityo Eko PrasojoTimothy Baldwin
Examines the severe resource scarcity across Indonesia’s 700+ languages to expose why state-of-the-art NLP models fail on dialect-rich, understudied languages and provides actionable strategies to build inclusive language technologies.
Natural language processing (NLP) technology—the computational handling of human language—has focused overwhelmingly on English and a small set of data-rich languages. This creates a severe digital divide for multilingual nations such as Indonesia, which is the fourth most populous country in the world and home to more than 700 spoken languages. Many of these indigenous languages are endangered, lack written standardization, and are largely absent from modern digital systems.
The article provides an empirical overview of NLP research across Indonesia's diverse linguistic landscape. Its main objective is to evaluate the availability of linguistic resources, identify key linguistic and infrastructural bottlenecks, and demonstrate how dialectal variations degrade the performance of current language technologies.
The authors conducted a multi-part analysis combining a historical literature review of Indonesian language research, corpus-scale audits comparing data availability against speaker populations, and targeted empirical experiments. Specifically, the authors evaluated popular off-the-shelf language identification systems—including langid.py, FastText, and CLD3—across 29 sentence sets translated into multiple regional dialects and formality styles of Javanese.
The findings reveal substantial disparities and system weaknesses. First, Indonesian languages suffer from severe digital underrepresentation: while Italian and Javanese have comparable speaker populations, Wikipedia contains over 3,000 megabytes of Italian text compared to less than 50 megabytes for Javanese. Second, language technologies exhibit severe dialectal bias. Evaluated language identification tools showed substantial accuracy swings depending on the dialect; for instance, FastText's top-1 accuracy ranged from 37.9% on Central Ngoko down to 6.9% on Western Ngoko. Third, complex linguistic phenomena such as code-mixing (blending languages at word and morpheme levels) and non-standard orthography significantly expand vocabulary complexity and degrade baseline model reliability. Finally, severe compute and infrastructure constraints persist locally, where top university computer science faculties frequently operate with fewer than ten graphical processing units (GPUs).
These findings mean that applying standard NLP pipelines directly to Indonesian language contexts risks high operational failure rates, poor user uptake, and the systematic exclusion of regional populations. Deploying tools that only recognize dominant dialects creates a biased feedback loop in automated data collection, further marginalizing underrepresented dialects. Conversely, developing robust, localized NLP systems offers significant societal value by bridging ethnic divides and facilitating public service delivery.
To address these gaps, practitioners and researchers should pursue several actionable next steps: (1) document detailed metadata—such as regional dialect, formality register, and style—within all newly created datasets; (2) prioritize data- and compute-efficient model designs, including parameter distillation, model pruning, and subword or token-free architectures; (3) build parallel translation corpora pivoting on standard Indonesian to generate synthetic training data; (4) expand spoken language processing for predominantly unwritten languages; and (5) partner directly with local linguistic communities to align development with genuine user needs.
The primary limitations of the article involve the constrained sample sizes used in the Javanese dialect experiments (29 sentences evaluated across three dialect regions) and the scarcity of standardized benchmarks for the remaining hundreds of long-tail languages. Nevertheless, the broad findings provide strong confidence that current mainstream language models require substantial adaptation before they can perform reliably across diverse Indonesian language environments.
- Paper: The State and Fate of Linguistic Diversity and Inclusion in the NLP World, Pratik Joshi et al. (2020). Its global analysis of NLP’s linguistic resource disparities provides the context for understanding why Indonesia’s many underrepresented languages pose a distinctive challenge.
- Paper: Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages, Ayyoob Imani et al. (2023). It turns the source’s call to address underrepresented languages into a large-scale effort to collect data and extend multilingual models to hundreds of languages.
- Paper: BYOL: Bring Your Own Language into LLMs, Waqas Zamir et al. (2026). It carries the challenge of unequal language resources into a practical framework for choosing adaptation strategies for languages with different levels of available text.
- Paper: Omnilingual MT: Machine Translation for 1,600 Languages, Omnilingual MT Team et al. (2026). It extends multilingual machine translation toward broad coverage, addressing the data and evaluation barriers the source identifies for underrepresented languages.
