One Country, 700+ Languages: NLP Challenges for Underrepresented Languages and Dialects in Indonesia

Alham Fikri AjiGenta Indra WinataFajri KotoSamuel CahyawijayaAde RomadhonyRahmad MahendraKemal KurniawanDavid MoeljadiRadityo Eko PrasojoTimothy Baldwin

article2022ACL141 citations

Examines the severe resource scarcity across Indonesia’s 700+ languages to expose why state-of-the-art NLP models fail on dialect-rich, understudied languages and provides actionable strategies to build inclusive language technologies.

Listen

Natural language processing (NLP) technology—the computational handling of human language—has focused overwhelmingly on English and a small set of data-rich languages. This creates a severe digital divide for multilingual nations such as Indonesia, which is the fourth most populous country in the world and home to more than 700 spoken languages. Many of these indigenous languages are endangered, lack written standardization, and are largely absent from modern digital systems.

The article provides an empirical overview of NLP research across Indonesia's diverse linguistic landscape. Its main objective is to evaluate the availability of linguistic resources, identify key linguistic and infrastructural bottlenecks, and demonstrate how dialectal variations degrade the performance of current language technologies.

The authors conducted a multi-part analysis combining a historical literature review of Indonesian language research, corpus-scale audits comparing data availability against speaker populations, and targeted empirical experiments. Specifically, the authors evaluated popular off-the-shelf language identification systems—including langid.py, FastText, and CLD3—across 29 sentence sets translated into multiple regional dialects and formality styles of Javanese.

The findings reveal substantial disparities and system weaknesses. First, Indonesian languages suffer from severe digital underrepresentation: while Italian and Javanese have comparable speaker populations, Wikipedia contains over 3,000 megabytes of Italian text compared to less than 50 megabytes for Javanese. Second, language technologies exhibit severe dialectal bias. Evaluated language identification tools showed substantial accuracy swings depending on the dialect; for instance, FastText's top-1 accuracy ranged from 37.9% on Central Ngoko down to 6.9% on Western Ngoko. Third, complex linguistic phenomena such as code-mixing (blending languages at word and morpheme levels) and non-standard orthography significantly expand vocabulary complexity and degrade baseline model reliability. Finally, severe compute and infrastructure constraints persist locally, where top university computer science faculties frequently operate with fewer than ten graphical processing units (GPUs).

These findings mean that applying standard NLP pipelines directly to Indonesian language contexts risks high operational failure rates, poor user uptake, and the systematic exclusion of regional populations. Deploying tools that only recognize dominant dialects creates a biased feedback loop in automated data collection, further marginalizing underrepresented dialects. Conversely, developing robust, localized NLP systems offers significant societal value by bridging ethnic divides and facilitating public service delivery.

To address these gaps, practitioners and researchers should pursue several actionable next steps: (1) document detailed metadata—such as regional dialect, formality register, and style—within all newly created datasets; (2) prioritize data- and compute-efficient model designs, including parameter distillation, model pruning, and subword or token-free architectures; (3) build parallel translation corpora pivoting on standard Indonesian to generate synthetic training data; (4) expand spoken language processing for predominantly unwritten languages; and (5) partner directly with local linguistic communities to align development with genuine user needs.

The primary limitations of the article involve the constrained sample sizes used in the Javanese dialect experiments (29 sentences evaluated across three dialect regions) and the scarcity of standardized benchmarks for the remaining hundreds of long-tail languages. Nevertheless, the broad findings provide strong confidence that current mainstream language models require substantial adaptation before they can perform reliably across diverse Indonesian language environments.

Aji et al (2022).pdf
Cover for One Country, 700+ Languages: NLP Challenges for Underrepresented Languages and Dialects in Indonesia

Abstract

NLP research is impeded by a lack of resources and awareness of the challenges presented by underrepresented languages and dialects. Focusing on the languages spoken in Indonesia, the second most linguistically diverse and the fourth most populous nation of the world, we provide an overview of the current state of NLP research for Indonesia’s 700+ languages. We highlight challenges in Indonesian NLP and how these affect the performance of current NLP systems. Finally, we provide general recommendations to help develop NLP technology not only for languages of Indonesia but also other underrepresented languages.

Table of Contents

  • 1 Introduction
  • 2 Background and Related Work
  • 2.1 History and Taxonomy
  • 2.2 Efforts in Multilingual Research
  • 2.3 Progress in Indonesian NLP
  • 3 Challenges for Indonesian NLP
  • 3.1 Limited Resources
  • 3.2 Language Diversity
  • 3.2.1 Regional Dialects and Style Differences
  • Case Study in Javanese
  • 3.2.2 Code-Mixing
  • 3.3 Orthography Variation
  • 3.4 Societal Challenges
  • 4 Opportunities
  • 4.1 Better Documentation
  • 4.2 Potential Research Directions
  • 4.3 Engagement with Communities
  • 5 Conclusion
  • Acknowledgments
  • References
  • A Language Statistics
  • B Wikipedia Availability
  • C Wikipedia Vocabulary Overlap
  • D Dialect Differences
  • E Local Language Classification
  • F Indonesian NLP Resources

Knowls

  1. Knowl 1 — NLP research is highly uneven across Indonesian languages

    empirical result

    An audit of ACL Anthology papers from 2000 to 2020 found that publications mentioning Indonesian local languages increased little compared with publications on Indonesian, and that Indonesian languages received far less research attention per million speakers than the European and Asian languages included in the comparison. This imbalance includes languages with large speaker populations: the paper reports approximately 198 million Indonesian speakers, 84 million Javanese speakers, and 34 million Sundanese speakers. Indonesia has more than 700 languages; Ethnologue classifies 440 as endangered and 12 as extinct. The audit therefore identifies a substantial mismatch between the linguistic scale and vitality of Indonesia’s languages and their visibility in NLP research.

  2. Knowl 2 — Available text data is far below what speaker counts would suggest

    empirical result

    The paper’s corpus and Wikipedia comparisons show a severe shortage of text for Indonesian local languages, including languages with many speakers. In CC-100, Javanese and Sundanese account for 0.001% and 0.002% of the corpus, respectively; in mC4, the paper reports 0.6 million Javanese and 0.3 million Sundanese tokens out of 6.3 trillion total tokens. Indo4B-Plus contains approximately 10% Javanese and Sundanese data combined. Wikipedia comparisons show that Italian has more than 3 GB of articles while Javanese has less than 50 MB, despite comparable speaker counts; Sundanese has less than 25 MB, versus more than 1.5 GB for languages with similar numbers of speakers. Many other Indonesian local languages have no Wikipedia edition, and the paper notes that high-quality written sources are difficult to find because local-language use is often primarily spoken and many relevant websites publish in Indonesian.

  3. Knowl 3 — Regional dialects and styles produce substantial lexical differences

    empirical result

    Lexical variation within a language can be large enough to complicate NLP even when speakers and resources are grouped under one language label. The paper reports prior findings of more than 50% lexical variation between Javanese varieties in different cities in central and eastern Java, and up to 13% between Javanese districts in Lamongan. A study of Saluan reported up to 23.5% variation across 31 observation points for 200 basic vocabulary items. The paper also elicited Javanese translations across Western, Central, and Eastern regional varieties and distinguished Ngoko, a daily-use style, from Krama, a polite style used with elders and people of higher social status. For example, elicited equivalents for ‘I/me’ include ‘inyong’ in Western Ngoko, ‘aku’ in Central and Eastern Ngoko, and ‘kulo’ in Eastern Krama. These regional and social-style distinctions can change the words an NLP system encounters.

  4. Knowl 4 — Language identification accuracy changes across Javanese varieties

    empirical result

    To test whether language identification systems handle Javanese varieties uniformly, the authors had native speakers translate 29 simple sentences under specified regional-dialect and style conditions, then evaluated langid.py, FastText, and CLD3. Accuracy values below are fractions from 0 to 1; each pair gives Top-1 and Top-3 accuracy, except CLD3, for which only Top-1 is reported. For Western Ngoko, langid.py scored 0.241/0.621, FastText 0.069/0.379, and CLD3 0.759. For Central Ngoko, the scores were 0.345/0.690, 0.379/0.724, and 0.828. For Eastern Ngoko, they were 0.276/0.552, 0.103/0.379, and 0.552. For Eastern Krama, they were 0.345/0.759, 0.379/0.586, and 0.897. The authors attribute stronger identification of Central Ngoko and Krama, in general, to the systems’ training data, which includes Javanese Wikipedia written in those varieties. The results demonstrate that language-identification performance depends on the dialect and style tested; consequently, automatic collection based on language identification can underrepresent varieties that the system detects poorly.

  5. Knowl 5 — Code-mixing occurs within words as well as between languages

    empirical result

    In Indonesian conversation and social-media text, speakers mix Indonesian with local languages and other languages, producing colloquial forms that can be difficult for NLP systems to process. The paper highlights mixing at the morpheme level: ‘quotenya’ combines English ‘quote’ with the Indonesian possessive suffix ‘-nya’, while ‘ngetag’ combines the Betawinese prefix ‘nge-’ with the English word ‘tag’. Code-mixing can also involve complete words or utterances, and may arise in multilingual border communities; the paper gives Jember, where speakers combine Javanese and Madurese in daily conversation, as an example. Thus, language variation is not limited to switching between intact language-specific words: mixed morphological forms also occur.

  6. Knowl 6 — Non-standard orthography creates multiple written forms for local-language words

    empirical result

    Many Indonesian local languages lack a widely adopted standard orthography, and Romanized spellings can vary even when forms are pronounced the same and are mutually intelligible. Native-speaker-confirmed examples include Javanese ‘apa’ and ‘opo’ for ‘what’, Balinese ‘inggih’ and ‘nggih’ for ‘yes’, and Sundanese ‘punten’ and ‘punteun’ for ‘please/sorry’. Some languages have older writing systems or transliteration standards, but those systems are not necessarily widely practiced. The resulting spelling variation enlarges the vocabulary seen by NLP systems, particularly systems based on word-level representations, and makes it harder to ensure that different spellings of the same word receive similar representations.

  7. Knowl 7 — Publicly available labeled datasets cover few local-language tasks

    empirical result

    The paper’s resource survey finds that many earlier Indonesian NLP studies did not release their data or models, while standardized and publicly available resources remain concentrated on Indonesian. Its inventory of publicly available local-language datasets lists six examples: Minangkabau sentiment analysis from Twitter and reviews (5,000 sentences); Sundanese emotion classification from Twitter (2,518); Javanese hate-speech detection from Twitter (3,478); Sundanese hate-speech detection from Twitter (2,209); Javanese dependency parsing from Wikipedia (125); and Javanese sentiment analysis on translated IMDb reviews (100,000). The paper defines these dataset sizes as sentence counts. The small number of listed language-task combinations illustrates how limited labeled coverage is beyond Indonesian, even for better-resourced local languages.

  8. Knowl 8 — Record dialect, region, style, and register in dataset and model metadata

    model/method

    The authors recommend documenting regional dialect and style/register in NLP datasets and models, in addition to identifying the language. Regional metadata can disclose which varieties a system can handle, helping users and stakeholders set realistic expectations; for crawled data, it can also provide information about the topics represented. Style and register metadata should capture distinctions such as politeness and formality, which are relevant both for describing data and for research on style modeling or style transfer. The recommendation is especially pertinent when linguistic differences across regions or social styles are substantial.

  9. Knowl 9 — Prioritize data-efficient, compute-efficient, robust, and speech-capable NLP

    model/method

    The paper proposes a research agenda for underrepresented Indonesian languages and other low-resource languages. To reduce dependence on large monolingual corpora, it recommends data-efficient adaptation, few-shot learning, and transfer from related languages. It also recommends collecting parallel data between Indonesian and individual local languages: many Indonesians are bilingual, and vocabulary overlap may help with translation from relatively little parallel data; such data could additionally support synthetic-data generation. To address limited computing resources, the authors favor lightweight and fast models, using approaches such as distillation, factorization, pruning, or more efficient training, and recommend studying the efficiency–quality trade-off, including for non-neural methods. To handle code-mixing and spelling variation, they suggest exploring subword and token-free approaches. Because some local languages are used mainly in speech, they also call for work on spoken-language understanding, speech recognition, and multimodal NLP.

  10. Knowl 10 — Language-technology design must account for local use and infrastructure

    limitation

    The paper cautions that there is no single language-technology solution suited to every Indonesian community. Some residents use a local language daily and are less proficient in Indonesian, while other speakers increasingly use Indonesian and may pass their local language on less often. Access to technology is also uneven: the paper reports 73.7% internet penetration in Indonesia in 2020, concentrated mainly on Java; among people without internet access, 39% said they did not understand the technology and 15% said they lacked a device. It also describes limited funding and scarce GPU resources outside Java. The authors therefore recommend assessing whether a proposed technology serves local needs and working with native speakers, local communities, and linguists to develop appropriate use cases and resources.

Coverage note — The historical account of language families and the exhaustive catalogs of Indonesian-only corpora, benchmarks, and tools were omitted because they provide background or resource listings rather than additional central findings; representative local-language dataset coverage is included.

References

  1. 1.Husen Abas. 1987. Indonesian as a unifying language of wider communication: a historical and sociolinguistic perspective. Number 73 in Pacific Linguistics Series D. The Australian National University, Canberra.
  2. 2.Zaenal Abidin, Permata, Imam Ahmad, and Rusliyawati. 2021. Effect of mono corpus quantity on statistical machine translation Indonesian–Lampung dialect of nyo. In Journal of Physics: Conference Series, volume 1751. IOP Publishing.
  3. 3.Alexander Adelaar. 2005. The Austronesian languages of Asia and Madagascar: A historical perspective. In Alexander Adelaar and Nikolaus P. Himmelmann, editors, The Austronesian Languages of Asia and Madagascar, chapter 1, pages 1–42. Routledge, Oxon.
  4. 4.Mirna Adriani, Jelita Asian, Bobby Nazief, Seyed MM Tahaghoghi, and Hugh E Williams. 2007. Stemming Indonesian: A confix-stripping approach. ACM Transactions on Asian Language Information Processing (TALIP), 6(4):1–33.
  5. 5.Željko Agic and Ivan Vuli ´ c. 2019. ´ JW300: A wide-coverage parallel corpus for low-resource languages. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3204–3210, Florence, Italy. Association for Computational Linguistics.
  6. 6.Alham Fikri Aji, Nikolay Bogoychev, Kenneth Heafield, and Rico Sennrich. 2020. In neural machine translation, what does transfer learning transfer? In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7701–7710.
  7. 7.Alham Fikri Aji, Radityo Eko Prasojo Tirana Noor Fatyanosa, Philip Arthur, Suci Fitriany, Salma Qonitah, Nadhifa Zulfa, Tomi Santoso, and Mahendra Data. 2021. Paracotta: Synthetic multilingual paraphrase corpora from the most diverse translation sample pair. In Proceedings of the 35th Pacific Asia Conference on Language, Information and Computation, pages 666–675.
  8. 8.Alham Fikri Aji and Kenneth Heafield. 2017. Sparse communication for distributed gradient descent. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 440–445.
  9. 9.Ika Alfina, Ruli Manurung, and Mohamad Ivan Fanany. 2016. DBpedia entities expansion in automatically building dataset for Indonesian NER. In 2016 International Conference on Advanced Computer Science and Information Systems (ICACSIS), pages 335–340. IEEE.
  10. 10.Ika Alfina, Rio Mulia, Mohamad Ivan Fanany, and Yudo Ekanata. 2017. Hate speech detection in the Indonesian language: A dataset and preliminary study. In 2017 International Conference on Advanced Computer Science and Information Systems (ICACSIS), pages 233–238. IEEE.
  11. 11.Zaid Alyafeai, Maraim Masoud, Mustafa Ghaleb, and Maged S Al-shaibani. 2021. Masader: Metadata sourcing for arabic text and speech data resources. arXiv preprint arXiv:2110.06744.
  12. 12.Antonios Anastasopoulos, Christopher Cox, Graham Neubig, and Hilaria Cruz. 2020. Endangered languages meet Modern NLP. In Proceedings of the 28th International Conference on Computational Linguistics: Tutorial Abstracts, pages 39–45, Barcelona, Spain (Online). International Committee for Computational Linguistics.
  13. 13.Karl Ronald Anderbeck. 2008. Malay dialects of the Batanghari river basin (Jambi, Sumatra). SIL International.
  14. 14.Anisya O. Anindyatri and Imarotul Mufidah. 2020. Gambaran Kondisi Vitalitas Bahasa Daerah di Indonesia. Kementrian Pendidikan dan Kebudayaan Pusat Data dan Teknologi Informasi, Tangerang Selatan, Indonesia.
  15. 15.Tri Apriani, Herry Sujaini, and Novi Safriadi. 2016. Pengaruh kuantitas korpus terhadap akurasi mesin penerjemah statistik bahasa bugis wajo ke Bahasa Indonesia. JUSTIN (Jurnal Sistem Dan Teknologi Informasi), 4(1):168–173.
  16. 16.Valentina Kania Prameswara Artari, Rahmad Mahendra, Meganingrum Arista Jiwanggi, Adityo Anggraito, and Indra Budi. 2021. A multi-pass sieve coreference resolution for Indonesian. In Proceedings of the International Conference on Recent Advances in Natural Language Processing (RANLP 2021), pages 79–85.
  17. 17.Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. 2020. On the cross-lingual transferability of monolingual representations. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4623–4637.
  18. 18.Jessica Naraiswari Arwidarasti, Ika Alfina, and Adila Alfa Krisnadhi. 2019. Converting an Indonesian constituency treebank to the Penn treebank format. In 2019 International Conference on Asian Language Processing (IALP), pages 331–336.
  19. 19.Annisa Nurul Azhar, Masayu Leylia Khodra, and Arie Pratama Sutiono. 2019. Multi-label aspect categorization with convolutional neural networks and extreme gradient boosting. In 2019 International Conference on Electrical Engineering and Informatics (ICEEI), pages 35–40. IEEE.
  20. 20.Kurniawati Azizah, Mirna Adriani, and Wisnu Jatmiko. 2020. Hierarchical transfer learning for multilingual, multi-speaker, and style transfer dnn-based tts on low-resource languages. IEEE Access, 8:179798–179812.
  21. 21.Anab Maulana Barik, Rahmad Mahendra, and Mirna Adriani. 2019. Normalization of Indonesian-English code-mixed Twitter data. In Proceedings of the 5th Workshop on Noisy User-generated Text (W-NUT 2019), pages 417–424, Hong Kong, China. Association for Computational Linguistics.
  22. 22.Sadar Baskoro and Mirna Adriani. 2008. Developing an Indonesian speech recognition system. In Second MALINDO Workshop. Selangor, Malaysia.
  23. 23.Peter Bellwood. 1997. Prehistory of the Indo-Malaysian Archipelago. University of Hawaii Press, Honolulu.
  24. 24.Peter Bellwood, Geoffrey Chambers, Malcolm Ross, and Hsiao-chun Hung. 2011. Are ‘cultures’ inherited? Multidisciplinary perspectives on the origins and migrations of Austronesian-speaking peoples prior to 1000 BC. In Investigating archaeological cultures, pages 321–354. Springer.
  25. 25.Emily M Bender and Batya Friedman. 2018. Data statements for natural language processing: Toward mitigating system bias and enabling better science. Transactions of the Association for Computational Linguistics, 6:587–604.
  26. 26.Jacques Bertrand. 2004. Nationalism and ethnic conflict in Indonesia. Cambridge University Press.
  27. 27.Laurent Besacier, Etienne Barnard, Alexey Karpov, and Tanja Schultz. 2014. Automatic speech recognition for under-resourced languages: A survey. Speech communication, 56:85–100.
  28. 28.Steven Bird. 2020. Decolonising speech and language technology. In Proceedings of the 28th International Conference on Computational Linguistics, pages 3504–3519, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  29. 29.Robert Blust. 1980. Austronesian etymologies. Oceanic Linguistics, 19:1–181.
  30. 30.Francis Bond and Timothy Baldwin. 2016. Introduction to Japanese Computational Linguistics. CSLI Publications, Stanford, USA.
  31. 31.Francis Bond, Lian Tze Lim, Enya Kong Tang, and Hammam Riza. 2014. The combined Wordnet Bahasa. NUSA: Linguistic studies of languages in and around Indonesia, 57:83–100.
  32. 32.Luiz Henrique Bonifacio, Israel Campiotti, Vitor Jeronymo, Roberto Lotufo, and Rodrigo Nogueira. 2021. mmarco: A multilingual version of the ms marco passage ranking dataset. arXiv preprint arXiv:2108.13897.
  33. 33.Indra Budi, Stéphane Bressan, Gatot Wahyudi, Zainal A. Hasibuan, and Bobby A. A. Nazief. 2005. Named entity recognition for the Indonesian language: Combining contextual, morphological and part-of-speech features into a knowledge engineering approach. In Discovery Science, pages 57–69, Berlin, Heidelberg. Springer Berlin Heidelberg.
  34. 34.Budiono, Hammam Riza, and Chairil Hakim. 2009. Resource report: building parallel text corpora for multi-domain translation system. In Proceedings of the 7th Workshop on Asian Language Resources (ALR7), pages 92–95.
  35. 35.Samuel Cahyawijaya. 2021. Greenformers: Improving computation and memory efficiency in transformer models via low-rank approximation.
  36. 36.Samuel Cahyawijaya, Genta Indra Winata, Bryan Wilie, Karissa Vincentio, Xiaohong Li, Adhiguna Kuncoro, Sebastian Ruder, Zhi Yuan Lim, Syafri Bahar, Masayu Khodra, Ayu Purwarianti, and Pascale Fung. 2021. IndoNLG: Benchmark and resources for evaluating Indonesian natural language generation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 8875–8898, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  37. 37.Isaac Caswell, Julia Kreutzer, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Auguste Tapo, Nishant Subramani, Artem Sokolov, Claytone Sikasote, Monang Setyawan, Supheakmungkol Sarin, Sokhar Samb, Benoît Sagot, Clara Rivera, Annette Rios Gonzales, Isabel Papadimitriou, Salomey Osei, Pedro Ortiz Suarez, Iroro Orife, Kelechi Ogueji, Rubungo Andre Niyongabo, Toan Q. Nguyen, Mathias Muller, Andr’e Muller, Shamsuddeen Hassan Muhammad, Nanda Firdausi Muhammad, Ayanda Mnyakeni, Jamshidbek Mirzakhalov, Tapiwanashe Matangira, Colin Leong, Nze Lawson, Sneha Kudugunta, Yacine Jernite, M. Jenny, Orhan Firat, Bonaventure F. P. Dossou, Sakhile Dlamini, Nisansa de Silva, Sakine cCabuk Balli, Stella Rose Biderman, Alessia Battisti, Ahmed Baruwa, Ankur Bapna, Pallavi N. Baljekar, Israel Abebe Azime, Ayodele Awokoya, Duygu Ataman, Orevaoghene Ahia, Oghenefego Ahia, Sweta Agrawal, and Mofetoluwa Adeyemi. 2022. Quality at a glance: An audit of web-crawled multilingual datasets. Transactions of the Association for Computational Linguistics, 10:50–72.
  38. 38.Yu-An Chung, Chenguang Zhu, and Michael Zeng. 2021. SPLAT: Speech-language joint pre-training for spoken language understanding. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1897–1907, Online. Association for Computational Linguistics.
  39. 39.Jonathan H. Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki. 2020. TyDi QA: A benchmark for information-seeking question answering in typologically diverse languages. Transactions of the Association for Computational Linguistics, 8:454–470.
  40. 40.Abigail C Cohn and Maya Ravindranath. 2014. Local languages in indonesia: Language maintenance or language shift. Linguistik Indonesia, 32(2):131–148.
  41. 41.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8440–8451, Online. Association for Computational Linguistics.
  42. 42.Alexander R. Coupe and František Kratochvíl. 2020. Asia before English. In Kingsley Bolton, Werner Botha, and Andy Kirkpatrick, editors, The Handbook of Asian Englishes, pages 15–48. John Wiley & Sons, Inc.
  43. 43.Wenliang Dai, Samuel Cahyawijaya, Zihan Liu, and Pascale Fung. 2021. Multimodal end-to-end sparse model for emotion recognition. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5305–5316, Online. Association for Computational Linguistics.
  44. 44.Wenliang Dai, Zihan Liu, Tiezheng Yu, and Pascale Fung. 2020. Modality-transferable emotion embeddings for low-resource multimodal emotion recognition. In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing, pages 269–280, Suzhou, China. Association for Computational Linguistics.
  45. 45.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  46. 46.Nindian Puspa Dewi, Joan Santoso, Ubaidi Ubaidi, and Eka Rahayu Setyaningsih. 2020. Combination of genetic algorithm and Brill tagger algorithm for part of speech tagging Bahasa Madura. Proceeding of the Electrical Engineering Computer Science and Informatics, 7(2):38–42.
  47. 47.Arawinda Dinakaramani, Fam Rashel, Andry Luthfi, and Ruli Manurung. 2014. Designing an Indonesian part of speech tagset and manually tagged Indonesian corpus. In 2014 International Conference on Asian Language Processing (IALP), pages 66–69.
  48. 48.Michael Diskin, Alexey Bukhtiyarov, Max Ryabinin, Lucile Saulnier, Anton Sinitsin, Dmitry Popov, Dmitry Pyrkin, Maxim Kashirin, Alexander Borzunov, Albert Villanova del Moral, et al. 2021. Distributed deep learning in open collaborations. Advances in Neural Information Processing Systems, 34.
  49. 49.A. Seza Dogruöz, Sunayana Sitaram, Barbara E. Bullock, and Almeida Jacqueline Toribio. 2021. A survey of code-switching: Linguistic and social perspectives for language technologies. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1654–1666, Online. Association for Computational Linguistics.
  50. 50.David M. Eberhard, Gary F. Simons, and Charles D. Fennig. 2021. Ethnologue: Languages of the World. Twenty-fourth edition. Dallas, Texas: SIL International.
  51. 51.European Language Resources Association. 2019. BLT4All: Language Technologies for All. https://lt4all.elra.info/en/. [Online; accessed Dec. 2019.].
  52. 52.Muhammad Fachri. 2014. Named entity recognition for Indonesian text using hidden markov model. Undergraduate Thesis.
  53. 53.Andri Imam Fauzi and Dwi Puspitorini. 2018. Dialect and identity: A case study of Javanese use in WhatsApp and Line. In IOP Conference Series: Earth and Environmental Science, volume 175, page 012111. IOP Publishing.
  54. 54.Abdurrisyad Fikri and Ayu Purwarianti. 2012. Case based Indonesian closed domain question answering system with real world questions. In 2012 7th International Conference on Telecommunication Systems, Services, and Applications (TSSA), pages 181–186.
  55. 55.Philip Gage. 1994. A new algorithm for data compression. C Users J., 12(2):23–38.
  56. 56.Dan Gillick, Cliff Brunk, Oriol Vinyals, and Amarnag Subramanya. 2016. Multilingual language processing from bytes. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1296–1306, San Diego, California. Association for Computational Linguistics.
  57. 57.Dirk Goldhahn, Thomas Eckart, and Uwe Quasthoff. 2012. Building large monolingual dictionaries at the Leipzig corpora collection: From 100 to 200 languages. In Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC’12), pages 759–765, Istanbul, Turkey. European Language Resources Association (ELRA).
  58. 58.Russell D. Gray and Fiona M. Jordan. 2000. Language trees support the express-train sequence of Austronesian expansion. Nature, 405:1052–1055.
  59. 59.William Gunawan, Derwin Suhartono, Fredy Purnomo, and Andrew Ongko. 2018. Named-entity recognition for Indonesian language using bidirectional lstm-cnns. Procedia Computer Science, 135:425–432.
  60. 60.Tri Wahyu Guntara, Alham Fikri Aji, and Radityo Eko Prasojo. 2020. Benchmarking multidomain English-Indonesian machine translation. In Proceedings of the 13th Workshop on Building and Using Comparable Corpora, pages 35–43, Marseille, France. European Language Resources Association.
  61. 61.Suchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A Smith. 2020. Don’t stop pretraining: Adapt language models to domains and tasks. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8342–8360.
  62. 62.Francisco Guzmán, Peng-Jen Chen, Myle Ott, Juan Pino, Guillaume Lample, Philipp Koehn, Vishrav Chaudhary, and Marc’Aurelio Ranzato. 2019. The FLORES evaluation datasets for low-resource machine translation: Nepali–English and Sinhala–English. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 6098–6111, Hong Kong, China. Association for Computational Linguistics.
  63. 63.Muhammad Yudistira Hanifmuti and Ika Alfina. 2020. Aksara: An Indonesian morphological analyzer that conforms to the UD v2 annotation guidelines. In 2020 International Conference on Asian Language Processing (IALP), pages 86–91. IEEE.
  64. 64.Akhmad Haryono. 2012. Perubahan dan perkembangan bahasa: Tinjauan historis dan sosiolinguistik. Ph.D. thesis, Udayana University.
  65. 65.Tahmid Hasan, Abhik Bhattacharjee, Md. Saiful Islam, Kazi Mubasshir, Yuan-Fang Li, Yong-Bin Kang, M. Sohel Rahman, and Rifat Shahriyar. 2021. XL-sum: Large-scale multilingual abstractive summarization for 44 languages. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 4693–4703, Online. Association for Computational Linguistics.
  66. 66.Muhammad Hasbiansyah, Herry Sujaini, and Novi Safriadi. 2016. Tuning for quality untuk uji akurasi mesin penerjemah statistik (MPS) Bahasa Indonesia-bahasa dayak kanayatn. JUSTIN (Jurnal Sistem dan Teknologi Informasi), 4(1):209–213.
  67. 67.Andi Hermanto, Teguh Bharata Adji, and Noor Akhmad Setiawan. 2015. Recurrent neural network language model for English-Indonesian machine translation: Experimental study. In 2015 International conference on science in information technology (ICSITech), pages 132–136. IEEE.
  68. 68.Devin Hoesen and Ayu Purwarianti. 2018. Investigating bi-LSTM and CRF with POS tag embedding for Indonesian named entity tagger. In 2018 International Conference on Asian Language Processing (IALP), pages 35–38. IEEE.
  69. 69.Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020. XTREME: A Massively Multilingual Multitask Benchmark for Evaluating Cross-lingual Generalization. In Proceedings of ICML 2020.
  70. 70.Baden Hughes, Timothy Baldwin, Steven Bird, Jeremy Nicholson, and Andrew MacKinlay. 2006. Reconsidering language identification for written language resources. In Proceedings of LREC 2006, pages 485–488, Genoa, Italy.
  71. 71.Muhammad Okky Ibrohim and Indra Budi. 2018. A dataset and preliminaries study for abusive language detection in Indonesian social media. Procedia Computer Science, 135:222–229. The 3rd International Conference on Computer Science and Computational Intelligence (ICCSCI 2018) : Empowering Smart Technology in Digital Era for a Better Life.
  72. 72.Muhammad Okky Ibrohim and Indra Budi. 2019. Multi-label hate speech and abusive language detection in Indonesian Twitter. In Proceedings of the Third Workshop on Abusive Language Online, pages 46–57, Florence, Italy. Association for Computational Linguistics.
  73. 73.Arfinda Ilmania, Abdurrahman, Samuel Cahyawijaya, and Ayu Purwarianti. 2018. Aspect detection and sentiment classification using deep neural network for Indonesian aspect-based sentiment analysis. In 2018 International Conference on Asian Language Processing (IALP), pages 62–67.
  74. 74.Benediktus Sridin Sulu Jahang and Zita Meirina. 2021. 1,3 juta anak di ntt belum bisa berbahasa indonesia. Last accessed on 05/10/2021.
  75. 75.Rini Jannati, Rahmad Mahendra, Cakra Wishnu Wardhana, and Mirna Adriani. 2018. Stance classification towards political figures on blog writing. In 2018 International Conference on Asian Language Processing (IALP), pages 96–101.
  76. 76.Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2020. TinyBERT: Distilling BERT for natural language understanding. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 4163–4174.
  77. 77.Ricky Chandra Johanes, Rahmad Mahendra, and Brahmastro Kresnaraman. 2020. Structuring code-switched product titles in Indonesian e-commerce platform. In Computational Data and Social Networks, pages 217–227, Cham. Springer International Publishing.
  78. 78.Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. 2020. The state and fate of linguistic diversity and inclusion in the NLP world. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6282–6293, Online. Association for Computational Linguistics.
  79. 79.Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov. 2017. Bag of tricks for efficient text classification. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers, pages 427–431. Association for Computational Linguistics.
  80. 80.Erlin Kartikasari, Kisyani Laksono, Dian Savitri Agusniar, and Diah Yovita Suryarini. 2018. A study of dialectology on Javanese "Ngoko" in Banyuwangi, Surabaya, Magetan, and Solo. Humaniora, 30(2):128.
  81. 81.Siti Oryza Khairunnisa, Aizhan Imankulova, and Mamoru Komachi. 2020. Towards a standardized dataset on Indonesian named entity recognition. In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing: Student Research Workshop, pages 64–71, Suzhou, China. Association for Computational Linguistics.
  82. 82.Simran Khanuja, Diksha Bansal, Sarvesh Mehtani, Savya Khosla, Atreyee Dey, Balaji Gopalan, Dilip Kumar Margam, Pooja Aggarwal, Rajiv Teja Nagipogu, Shachi Dave, Shruti Gupta, Subhash Chandra Bose Gali, Vish Subramanian, and Partha Talukdar. 2021. MuRIL: Multilingual Representations for Indian Languages. arXiv preprint arXiv:2103.10730.
  83. 83.Yash Khemchandani, Sarvesh Mehtani, Vaidehi Patil, Abhijeet Awasthi, Partha Talukdar, and Sunita Sarawagi. 2021. Exploiting language relatedness for low web-resource language model adaptation: An Indic languages study. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1312–1323, Online. Association for Computational Linguistics.
  84. 84.Young Jin Kim, Marcin Junczys-Dowmunt, Hany Hassan, Alham Fikri Aji, Kenneth Heafield, Roman Grundkiewicz, and Nikolay Bogoychev. 2019. From research to production and back: Ludicrously fast neural machine translation. In Proceedings of the 3rd Workshop on Neural Generation and Translation, pages 280–288.
  85. 85.Marian Klamer. 2018. Documenting the linguistic diversity of Indonesia: Time is running out. In Proceedings of ‘Revitalization of local languages as the pillar of pluralism’, pages 1–10. APBL (Asosiasi Peneliti Bahasa-bahasa Lokal) and Nusa Cendana University, Kupang, Satya Wacana Press.
  86. 86.Fajri Koto and Ikhwan Koto. 2020. Towards computational linguistics in Minangkabau language: Studies on sentiment analysis and machine translation. In Proceedings of the 34th Pacific Asia Conference on Language, Information and Computation, pages 138–148, Hanoi, Vietnam. Association for Computational Linguistics.
  87. 87.Fajri Koto, Jey Han Lau, and Timothy Baldwin. 2020a. Liputan6: A large-scale Indonesian dataset for text summarization. In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing, pages 598–608, Suzhou, China. Association for Computational Linguistics.
  88. 88.Fajri Koto, Jey Han Lau, and Timothy Baldwin. 2021. IndoBERTweet: A pretrained language model for Indonesian Twitter with effective domain-specific vocabulary initialization. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 10660–10668, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  89. 89.Fajri Koto, Afshin Rahimi, Jey Han Lau, and Timothy Baldwin. 2020b. IndoLEM and IndoBERT: A benchmark dataset and pre-trained language model for Indonesian NLP. In Proceedings of the 28th International Conference on Computational Linguistics, pages 757–770, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  90. 90.Fajri Koto and Gemala Y Rahmaningtyas. 2017. Inset lexicon: Evaluation of a word list for Indonesian sentiment analysis in microblogs. In 2017 International Conference on Asian Language Processing (IALP), pages 391–394. IEEE.
  91. 91.Taku Kudo. 2018. Subword regularization: Improving neural network translation models with multiple subword candidates. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 66–75, Melbourne, Australia. Association for Computational Linguistics.
  92. 92.Kemal Kurniawan and Alham Fikri Aji. 2018. Toward a standardized and more accurate Indonesian part-of-speech tagging. In 2018 International Conference on Asian Language Processing (IALP), pages 303–307. IEEE.
  93. 93.Kemal Kurniawan, Lea Frermann, Philip Schulz, and Trevor Cohn. 2021. PPT: Parsimonious parser transfer for unsupervised cross-lingual adaptation. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 2907–2918, Online. Association for Computational Linguistics.
  94. 94.Kemal Kurniawan and Samuel Louvan. 2018. Indosum: A new benchmark dataset for Indonesian text summarization. In 2018 International Conference on Asian Language Processing (IALP), pages 215–220. IEEE.
  95. 95.Retno Kusumaningrum, M. Ihsan Aji Wiedjayanto, Satriyo Adhy, and Suryono. 2016. Classification of Indonesian news articles based on latent Dirichlet allocation. In 2016 International Conference on Data and Software Engineering (ICoDSE), pages 1–5.
  96. 96.Septina Dian Larasati. 2012. IDENTIC corpus: Morphologically enriched Indonesian-English parallel corpus. In Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC’12), pages 902–906, Istanbul, Turkey. European Language Resources Association (ELRA).
  97. 97.Septina Dian Larasati, Vladislav Kubon, and Daniel Zeman. 2011. Indonesian morphology tool (morphind): Towards an Indonesian corpus. In International Workshop on Systems and Frameworks for Computational Morphology, pages 119–129. Springer.
  98. 98.Teven Le Scao and Alexander M Rush. 2021. How many data points is a prompt worth? In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2627–2636.
  99. 99.Dessi Puji Lestari, Koji Iwano, and Sadaoki Furui. 2006. A large vocabulary continuous speech recognition system for Indonesian language. In 15th Indonesian Scientific Conference in Japan Proceedings, pages 17–22.
  100. 100.Zhaojiang Lin, Zihan Liu, Genta Indra Winata, Samuel Cahyawijaya, Andrea Madotto, Yejin Bang, Etsuko Ishii, and Pascale Fung. 2021a. XPersona: Evaluating multilingual personalized chatbot. In Proceedings of the 3rd Workshop on Natural Language Processing for Conversational AI, pages 102–112, Online. Association for Computational Linguistics.
  101. 101.Zhaojiang Lin, Andrea Madotto, Genta Indra Winata, Peng Xu, Feijun Jiang, Yuxiang Hu, Chen Shi, and Pascale Fung. 2021b. BiToD: A bilingual multi-domain dataset for task-oriented dialogue modeling. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1).
  102. 102.Bing Liu, Minqing Hu, and Junsheng Cheng. 2005. Opinion observer: analyzing and comparing opinions on the web. In Proceedings of the 14th international conference on World Wide Web, pages 342–351.
  103. 103.Fangyu Liu, Emanuele Bugliarello, Edoardo Maria Ponti, Siva Reddy, Nigel Collier, and Desmond Elliott. 2021. Visually grounded reasoning across languages and cultures. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 10467–10485, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  104. 104.Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020. Multilingual denoising pre-training for neural machine translation. Transactions of the Association for Computational Linguistics, 8:726–742.
  105. 105.Marco Lui and Timothy Baldwin. 2012. langid. py: An off-the-shelf language identification tool. In Proceedings of the ACL 2012 system demonstrations, pages 25–30.
  106. 106.Edwin Lunando and Ayu Purwarianti. 2013. Indonesian social media sentiment analysis with sarcasm detection. In 2013 International Conference on Advanced Computer Science and Information Systems (ICACSIS), pages 195–198.
  107. 107.Andrea Madotto, Zhaojiang Lin, Genta Indra Winata, and Pascale Fung. 2021. Few-shot bot: Prompt-based learning for dialogue systems. arXiv preprint arXiv:2110.08118.
  108. 108.Putu Devi Maharani and Komang Dian Puspita Candra. 2018. Variasi leksikal bahasa Bali dialek kuta selatan. Mudra Jurnal Seni Budaya, 33(1):76–84.
  109. 109.Rahmad Mahendra, Alham Fikri Aji, Samuel Louvan, Fahrurrozi Rahman, and Clara Vania. 2021. IndoNLI: A natural language inference dataset for Indonesian. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 10511–10527, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  110. 110.Rahmad Mahendra, Septina Dian Larasati, and Ruli Manurung. 2008. Extending an Indonesian semantic analysis-based question answering system with linguistic and world knowledge axioms. In Proceedings of the 22nd Pacific Asia Conference on Language, Information and Computation, pages 262–271, The University of the Philippines Visayas Cebu College, Cebu City, Philippines. De La Salle University, Manila, Philippines.
  111. 111.Rahmad Mahendra, Heninggar Septiantri, Haryo Akbarianto Wibowo, Ruli Manurung, and Mirna Adriani. 2018. Cross-lingual and supervised learning approach for Indonesian word sense disambiguation task. In Proceedings of the 9th Global Wordnet Conference, pages 245–250, Nanyang Technological University (NTU), Singapore. Global Wordnet Association.
  112. 112.Miftahul Mahfuzh, Sidik Soleman, and Ayu Purwarianti. 2019. Improving joint layer RNN based keyphrase extraction by using syntactical features. In 2019 International Conference of Advanced Informatics: Concepts, Theory and Applications (ICAICTA), pages 1–6. IEEE.
  113. 113.Ryan McDonald, Joakim Nivre, Yvonne Quirmbach-Brundage, Yoav Goldberg, Dipanjan Das, Kuzman Ganchev, Keith Hall, Slav Petrov, Hao Zhang, Oscar Täckström, Claudia Bedini, Núria Bertomeu Castelló, and Jungmee Lee. 2013. Universal Dependency annotation for multilingual parsing. In Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 92–97, Sofia, Bulgaria. Association for Computational Linguistics.
  114. 114.Angelina McMillan-Major, Zaid Alyafeai, Stella Biderman, Kimbo Chen, Francesco De Toni, Gérard Dupont, Hady Elsahar, Chris Emezue, Alham Fikri Aji, Suzana Ilic, et al. 2022. Documenting geographically and contextually diverse data sources: The bigscience catalogue of language data and resources. arXiv preprint arXiv:2201.10066.
  115. 115.David Moeljadi. 2017. Building JATI: A Treebank for Indonesian. In Proceedings of The 4th Atma Jaya Conference on Corpus Studies (ConCorps 4), pages 1–9, Jakarta.
  116. 116.David Moeljadi, Francis Bond, and Sanghoun Song. 2015. Building an HPSG-based Indonesian resource grammar (INDRA). In Proceedings of the Grammar Engineering Across Frameworks (GEAF) 2015 Workshop, pages 9–16, Beijing, China. Association for Computational Linguistics.
  117. 117.David Moeljadi, Aditya Kurniawan, and Debaditya Goswami. 2019. Building Cendana: a treebank for informal Indonesian. In Proceedings of the 33rd Pacific Asia Conference on Language, Information and Computation, pages 156–164. Waseda Institute for the Study of Language and Information.
  118. 118.Aqsath Rasyid Naradhipa and Ayu Purwarianti. 2011. Sentiment classification for Indonesian message in social media. In Proceedings of the 2011 International Conference on Electrical Engineering and Informatics, pages 1–4.
  119. 119.Arbi Haza Nasution, Yohei Murakami, and Toru Ishida. 2017. A generalized constraint approach to bilingual dictionary induction for low-resource language families. ACM Transactions on Asian and Low-Resource Language Information Processing (TALLIP), 17(2):1–29.
  120. 120.Arbi Haza Nasution, Yohei Murakami, and Toru Ishida. 2021. Plan optimization to bilingual dictionary induction for low-resource language families. ACM Transactions Asian Low-Resource Language Information Processing, 20(2).
  121. 121.Toan Q Nguyen and David Chiang. 2017. Transfer learning across low-resource, related languages for neural machine translation. In Proceedings of the Eighth International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pages 296–301.
  122. 122.Della Widya Ningtyas, Herry Sujaini, and Novi Safriadi. 2018. Penggunaan pivot language pada mesin penerjemah statistik bahasa inggris ke bahasa melayu sambas. JEPIN (Jurnal Edukasi dan Penelitian Informatika), 4(2):173–178.
  123. 123.Made Nindyatama Nityasya, Haryo Akbarianto Wibowo, Radityo Eko Prasojo, and Alham Fikri Aji. 2021. Costs to consider in adopting NLP for your business.
  124. 124.Hiroki Nomoto, Hannah Choi, David Moeljadi, and Francis Bond. 2018. MALINDO Morph: Morphological dictionary and analyser for Malay/Indonesian. In Proceedings of the LREC 2018 Workshop “The 13th Workshop on Asian Language Resources”, pages 36–43.
  125. 125.Sashi Novitasari, Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura. 2020. Cross-lingual machine speech chain for Javanese, Sundanese, Balinese, and Bataks speech recognition and synthesis. In Proceedings of the 1st Joint Workshop on Spoken Language Technologies for Under-resourced languages (SLTU) and Collaboration and Computing for Under-Resourced Languages (CCURL), pages 131–138, Marseille, France. European Language Resources association.
  126. 126.Fajrin Nurjanah et al. 2018. Pengembangan kemampuan berbahasa indonesia siswa sekolah dasar desa terpencil melalui metode karyawisata berbasis potensi lokal. FKIP e-PROCEEDING, pages 167–176.
  127. 127.Bill Palmer. 2018. The languages and linguistics of the New Guinea area: A comprehensive guide, volume 4. Walter de Gruyter GmbH.
  128. 128.Valantino Ateng Pamolango. 2012. Geografi dialek bahasa Saluan. PARAFRASE: Jurnal Kajian Kebahasaan & Kesastraan, 12(02).
  129. 129.Xiaoman Pan, Boliang Zhang, Jonathan May, Joel Nothman, Kevin Knight, and Heng Ji. 2017. Cross-lingual name tagging and linking for 282 languages. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1946–1958, Vancouver, Canada. Association for Computational Linguistics.
  130. 130.Jonas Pfeiffer, Gregor Geigle, Aishwarya Kamath, Jan-Martin O. Steitz, Stefan Roth, Ivan Vulic, and Iryna Gurevych. 2021. xGQA: Cross-lingual visual question answering. ArXiv, abs/2109.06082.
  131. 131.Tiago Pimentel, Maria Ryskina, Sabrina J. Mielke, Shijie Wu, Eleanor Chodroff, Brian Leonard, Garrett Nicolai, Yustinus Ghanggo Ate, Salam Khalifa, Nizar Habash, Charbel El-Khaissi, Omer Goldman, Michael Gasser, William Lane, Matt Coler, Arturo Oncevay, Jaime Rafael Montoya Samame, Gema Celeste Silva Villegas, Adam Ek, Jean-Philippe Bernardy, Andrey Shcherbakov, Aziyana Bayyr-ool, Karina Sheifer, Sofya Ganieva, Matvey Plugaryov, Elena Klyachko, Ali Salehi, Andrew Krizhanovsky, Natalia Krizhanovsky, Clara Vania, Sardana Ivanova, Aelita Salchak, Christopher Straughn, Zoey Liu, Jonathan North Washington, Duygu Ataman, Witold Kieras, Marcin Wolinski, Totok Suhardijanto, Niklas Stoehr, Zahroh Nuriah, Shyam Ratan, Francis M. Tyers, Edoardo M. Ponti, Grant Aiton, Richard J. Hatcher, Emily Prud’hommeaux, Ritesh Kumar, Mans Hulden, Botond Barta, Dorina Lakatos, Gábor Szolnok, Judit Ács, Mohit Raj, David Yarowsky, Ryan Cotterell, Ben Ambridge, and Ekaterina Vylomova. 2021. Sigmorphon 2021 shared task on morphological reinflection: Generalization across languages. In Proceedings of the 18th SIGMORPHON Workshop on Computational Research in Phonetics, Phonology, and Morphology, pages 229–259, Online. Association for Computational Linguistics.
  132. 132.Femphy Pisceldo, Rahmad Mahendra, Ruli Manurung, and I Wayan Arka. 2008. A two-level morphological analyser for the Indonesian language. In Proceedings of the Australasian Language Technology Association Workshop 2008, pages 142–150.
  133. 133.Edoardo Maria Ponti, Goran Glavaš, Olga Majewska, Qianchu Liu, Ivan Vulic, and Anna Korhonen. 2020. XCOPA: A multilingual dataset for causal commonsense reasoning. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2362–2376, Online. Association for Computational Linguistics.
  134. 134.Apriyani Purwaningsih. 2017. Geografi dialek bahasa jawa pesisiran di desa paciran kabupaten lamongan. In Proceeding of International Conference on Art, Language, and Culture, pages 594–605.
  135. 135.Ayu Purwarianti, Alvin Andhika, Alfan Farizki Wicaksono, Irfan Afif, and Filman Ferdian. 2016. Inanlp: Indonesia natural language processing toolkit, case study: Complaint tweet classification. In 2016 International Conference On Advanced Informatics: Concepts, Theory And Application (ICAICTA), pages 1–5. IEEE.
  136. 136.Ayu Purwarianti and Ida Ayu Putu Ari Crisdayanti. 2019. Improving bi-LSTM performance for Indonesian sentiment analysis using paragraph vector. In 2019 International Conference of Advanced Informatics: Concepts, Theory and Applications (ICAICTA), pages 1–5. IEEE.
  137. 137.Ayu Purwarianti, Masatoshi Tsuchiya, and Seiichi Nakagawa. 2007. A machine learning approach for Indonesian question answering system. In Artificial Intelligence and Applications, pages 573–578.
  138. 138.Desmond Darma Putra, Abdul Arfan, and Ruli Manurung. 2008. Building an Indonesian wordnet. In Proceedings of the 2nd International MALINDO Workshop.
  139. 139.Oddy Virgantara Putra, Fathin Muhammad Wasmanson, Triana Harmini, and Shoffin Nahwa Utama. 2020. Sundanese twitter dataset for emotion classification. In 2020 International Conference on Computer Engineering, Network, and Intelligent Multimedia (CENIM), pages 391–395. IEEE.
  140. 140.Shofianina Dwi Ananda Putri, Muhammad Okky Ibrohim, and Indra Budi. 2021. Abusive language and hate speech detection for Indonesian-local language in social media text. In Recent Advances in Information and Communication Technology 2021, pages 88–98, Cham. Springer International Publishing.
  141. 141.Valdi Rachman, Septiviana Savitri, Fithriannisa Augustianti, and Rahmad Mahendra. 2017. Named entity recognition on Indonesian Twitter posts using long short-term memory networks. In 2017 International Conference on Advanced Computer Science and Information Systems (ICACSIS), pages 228–232.
  142. 142.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  143. 143.Anna Rogers, Timothy Baldwin, and Kobi Leins. 2021. ‘just what do you think you’re doing, dave?’ a checklist for responsible data use in NLP. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 4821–4833, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  144. 144.Malcolm Ross. 2005. Pronouns as a preliminary diagnostic for grouping Papuan languages. Papuan pasts: Cultural, linguistic and biological histories of Papuan-speaking peoples, 572:15–65.
  145. 145.Nur Endah Safitri, Amalia Zahra, and Mirna Adriani. 2016. Spoken language identification with phonotactics methods on Minangkabau, Sundanese, and Javanese languages. Procedia Computer Science, 81:182–187.
  146. 146.Nikmatun Aliyah Salsabila, Yosef Ardhito Winatmoko, Ali Akbar Septiandri, and Ade Jamal. 2018. Colloquial indonesian lexicon. In 2018 International Conference on Asian Language Processing (IALP), pages 226–229.
  147. 147.Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019. DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108.
  148. 148.Ilham Fathy Saputra, Rahmad Mahendra, and Alfan Farizki Wicaksono. 2018. Keyphrases extraction from user-generated contents in healthcare domain using long short-term memory networks. In Proceedings of the BioNLP 2018 workshop, pages 28–34, Melbourne, Australia. Association for Computational Linguistics.
  149. 149.Mei Silviana Saputri, Rahmad Mahendra, and Mirna Adriani. 2018. Emotion classification on Indonesian Twitter dataset. In 2018 International Conference on Asian Language Processing (IALP), pages 90–95. IEEE.
  150. 150.Gita Sarwadi, Mahsun Mahsun, and Burhanuddin Burhanuddin. 2019. Lexical variation of Sasak Kuto-Kute dialect in North Lombok district. Jurnal Kata: Penelitian tentang Ilmu Bahasa dan Sastra, 3(1):155–169.
  151. 151.Roy Schwartz, Jesse Dodge, Noah A Smith, and Oren Etzioni. 2020. Green AI. Communications of the ACM, 63(12):54–63.
  152. 152.Holger Schwenk, Vishrav Chaudhary, Shuo Sun, Hongyu Gong, and Francisco Guzmán. 2021. WikiMatrix: Mining 135M parallel sentences in 1620 language pairs from Wikipedia. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 1351–1361, Online. Association for Computational Linguistics.
  153. 153.Dmitriy Serdyuk, Yongqiang Wang, Christian Fuegen, Anuj Kumar, Baiyang Liu, and Yoshua Bengio. 2018. Towards end-to-end spoken language understanding. In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP).
  154. 154.Ken Nabila Setya and Rahmad Mahendra. 2018. Semi-supervised textual entailment on Indonesian wikipedia data. In 2018 International Conference on Computational Linguistics and Intelligent Text Processing (CICLing).
  155. 155.HS Simon and Ayu Purwarianti. 2013. Experiments on Indonesian-Japanese statistical machine translation. In 2013 IEEE international conference on computational intelligence and cybernetics (CYBERNETICSCOM), pages 80–84. IEEE.
  156. 156.Masitowarni Siregar, Syamsul Bahri, Dedi Sanjaya, et al. 2014. Code switching and code mixing in indonesia: Study in sociolinguistics. English Language and Literature Studies, 4(1):77–92.
  157. 157.Sunayana Sitaram, Khyathi Raghavi Chandu, Sai Krishna Rallabandi, and Alan W Black. 2019. A survey of code-switched speech and language processing. arXiv preprint arXiv:1904.00784.
  158. 158.James Neil Sneddon. 2003. The Indonesian language: Its history and role in modern society. UNSW Press, Sydney.
  159. 159.Soeparno. 2015. Kerancuan fono-ortografis dan ortofonologis bahasa Indonesia ragam lisan dan tulis. Diksi, 12(2).
  160. 160.Hein Steinhauer. 2005. Colonial history and language policy in insular Southeast Asia and Madagascar. In Alexander Adelaar and Nikolaus P. Himmelmann, editors, The Austronesian Languages of Asia and Madagascar, chapter 3, pages 65–86. Routledge, Oxon.
  161. 161.Made Agus Putra Subali and Chastine Fatichah. 2019. Kombinasi metode rule-based dan n-gram stemming untuk mengenali stemmer bahasa bali. Jurnal Teknologi Informasi dan Ilmu Komputer, 6(2):219–228.
  162. 162.Arie Ardiyanti Suryani, Dwi Hendratmo Widyantoro, Ayu Purwarianti, and Yayat Sudaryat. 2015. Experiment on a phrase-based statistical machine translation using PoS tag information for Sundanese into Indonesian. In 2015 International Conference on Information Technology Systems and Innovation (ICITSI), pages 1–6.
  163. 163.Arie Ardiyanti Suryani, Dwi Hendratmo Widyantoro, Ayu Purwarianti, and Yayat Sudaryat. 2018. The rule-based sundanese stemmer. ACM Transactions on Asian and Low-Resource Language Information Processing (TALLIP), 17(4):1–28.
  164. 164.Taufic Leonardo Sutejo and Dessi Puji Lestari. 2018. Indonesia hate speech detection using deep learning. In 2018 International Conference on Asian Language Processing (IALP), pages 39–43. IEEE.
  165. 165.Bejo Sutrisno and Yessika Ariesta. 2019. Beyond the use of code mixing by social media influencers in Instagram. Advances in Language and Literary Studies, 10(6):143–151.
  166. 166.Samson Tan and Shafiq Joty. 2021. Code-mixing on sesame street: Dawn of the adversarial polyglots. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3596–3616.
  167. 167.Dipta Tanaya and Mirna Adriani. 2016. Dictionary-based word segmentation for Javanese. Procedia Computer Science, 81:208–213.
  168. 168.Dipta Tanaya and Mirna Adriani. 2018. Word segmentation for Javanese character using dictionary, SVM, and CRF. In 2018 International Conference on Asian Language Processing (IALP), pages 240–243.
  169. 169.Yi Tay, Vinh Q Tran, Sebastian Ruder, Jai Gupta, Hyung Won Chung, Dara Bahri, Zhen Qin, Simon Baumgartner, Cong Yu, and Donald Metzler. 2021. Charformer: Fast character transformers via gradient-based subword tokenization. arXiv preprint arXiv:2106.12672.
  170. 170.I Nyoman Prayana Trisna and Arif Nurwidyantoro. 2020. Single document keywords extraction in Bahasa Indonesia using phrase chunking. Telkomnika, 18(4):1917–1925.
  171. 171.Rob van der Goot, Alan Ramponi, Arkaitz Zubiaga, Barbara Plank, Benjamin Muller, Iñaki San Vicente Roncal, Nikola Ljubešic, Özlem Çetinoğlu, Rahmad Mahendra, Talha Çolakoğlu, Timothy Baldwin, Tommaso Caselli, and Wladimir Sidorenko. 2021. MultiLexNorm: A shared task on multilingual lexical normalization. In Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021), pages 493–509, Online. Association for Computational Linguistics.
  172. 172.Daan van Esch, Tamar Lucassen, Sebastian Ruder, Isaac Caswell, and Clara Rivera. 2022. Writing system and speaker metadata for 2,800+ language varieties.
  173. 173.Jozina Vander Klok. 2015. The dichotomy of auxiliaries in Javanese: Evidence from two dialects. Australian Journal of Linguistics, 35(2):142–167.
  174. 174.Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019. Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 5797–5808.
  175. 175.Denny Vrandeciˇ c and Markus Krötzsch. 2014. Wikidata: a free collaborative knowledgebase. Communications of the ACM, 57(10):78–85.
  176. 176.Devid Haryalesmana Wahid and SN Azhari. 2016. Peringkasan sentimen esktraktif di twitter menggunakan hybrid tf-idf dan cosine similarity. IJCCS (Indonesian Journal of Computing and Cybernetics Systems), 10(2):207–218.
  177. 177.Haryo Akbarianto Wibowo, Made Nindyatama Nityasya, Afra Feyza Akyürek, Suci Fitriany, Alham Fikri Aji, Radityo Eko Prasojo, and Derry Tanti Wijaya. 2021. IndoCollex: A testbed for morphological transformation of Indonesian word colloquialism. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 3170–3183, Online. Association for Computational Linguistics.
  178. 178.Haryo Akbarianto Wibowo, Tatag Aziz Prawiro, Muhammad Ihsan, Alham Fikri Aji, Radityo Eko Prasojo, Rahmad Mahendra, and Suci Fitriany. 2020. Semi-supervised low-resource style transfer of Indonesian informal to formal language with iterative forward-translation. In 2020 International Conference on Asian Language Processing (IALP), pages 310–315. IEEE.
  179. 179.Alfan Farizki Wicaksono and Ayu Purwarianti. 2010. HMM based part-of-speech tagger for Bahasa Indonesia. In 4th International MALINDO (Malaysian-Indonesian Language) Workshop.
  180. 180.Alfan Farizki Wicaksono, Clara Vania, Bayu Distiawan, and Mirna Adriani. 2014. Automatically building a corpus for sentiment analysis on Indonesian tweets. In Proceedings of the 28th Pacific Asia conference on language, information and computing, pages 185–194.
  181. 181.Bryan Wilie, Karissa Vincentio, Genta Indra Winata, Samuel Cahyawijaya, Xiaohong Li, Zhi Yuan Lim, Sidik Soleman, Rahmad Mahendra, Pascale Fung, Syafri Bahar, and Ayu Purwarianti. 2020. IndoNLU: Benchmark and resources for evaluating Indonesian natural language understanding. In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing, pages 843–857, Suzhou, China. Association for Computational Linguistics.
  182. 182.Andika William and Yunita Sari. 2020. CLICK-ID: A novel dataset for Indonesian clickbait headlines. Data in Brief, 32:106231.
  183. 183.Genta Indra Winata. 2021. Multilingual transfer learning for code-switched language and speech neural modeling. Hong Kong University of Science and Technology (Hong Kong).
  184. 184.Genta Indra Winata, Samuel Cahyawijaya, Zhaojiang Lin, Zihan Liu, and Pascale Fung. 2020a. Lightweight and efficient end-to-end speech recognition using low-rank transformer. ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 6144–6148.
  185. 185.Genta Indra Winata, Samuel Cahyawijaya, Zihan Liu, Zhaojiang Lin, Andrea Madotto, Peng Xu, and Pascale Fung. 2020b. Learning fast adaptation on cross-accented speech recognition. Proc. Interspeech 2020, pages 1276–1280.
  186. 186.Genta Indra Winata and Masayu Leylia Khodra. 2015. Handling imbalanced dataset in multi-label text categorization using bagging and adaptive boosting. In 2015 International Conference on Electrical Engineering and Informatics (ICEEI), pages 500–505. IEEE.
  187. 187.Genta Indra Winata, Andrea Madotto, Zhaojiang Lin, Rosanne Liu, Jason Yosinski, and Pascale Fung. 2021. Language models are few-shot multilingual learners. In Proceedings of the 1st Workshop on Multilingual Representation Learning, pages 1–15.
  188. 188.Genta Indra Winata, Andrea Madotto, Jamin Shin, Elham J Barezi, and Pascale Fung. 2019a. On the effectiveness of low-rank matrix factorization for lstm model compression. In Proceedings of the 33rd Pacific Asia Conference on Language, Information and Computation, pages 253–262. Waseda Institute for the Study of Language and Information.
  189. 189.Genta Indra Winata, Andrea Madotto, Chien-Sheng Wu, and Pascale Fung. 2018. Code-switching language modeling using syntax-aware multi-task learning. In Proceedings of the Third Workshop on Computational Approaches to Linguistic Code-Switching, pages 62–67.
  190. 190.Genta Indra Winata, Andrea Madotto, Chien-Sheng Wu, and Pascale Fung. 2019b. Code-switched language models using neural based synthetic data from parallel sentences. In Proceedings of the 23rd Conference on Computational Natural Language Learning (CoNLL), pages 271–280.
  191. 191.Wilson Wongso, David Samuel Setiawan, and Derwin Suhartono. 2021. Causal and masked language modeling of Javanese language using transformer-based architectures. In 2021 International Conference on Advanced Computer Science and Information Systems (ICACSIS), pages 1–7. IEEE.
  192. 192.Linting Xue, Aditya Barua, Noah Constant, Rami Al-Rfou, Sharan Narang, Mihir Kale, Adam Roberts, and Colin Raffel. 2021a. Byt5: Towards a token-free future with pre-trained byte-to-byte models.
  193. 193.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021b. mT5: A massively multilingual pre-trained text-to-text transformer. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 483–498, Online. Association for Computational Linguistics.
  194. 194.Evi Yulianti, Indra Budi, Achmad N. Hidayanto, Hisar M. Manurung, and Mirna Adriani. 2011. Developing Indonesian-English hybrid machine translation system. In 2011 International Conference on Advanced Computer Science and Information Systems, pages 265–270.
  195. 195.Amalia Zahra, Sadar Baskoro, and Mirna Adriani. 2009. Building a pronunciation dictionary for Indonesian speech recognition system. In The TCAST workshop, Singapore.
  196. 196.Daniel Zeman, Jan Hajic, Martin Popel, Martin Potthast, Milan Straka, Filip Ginter, Joakim Nivre, and Slav Petrov. 2018. CoNLL 2018 shared task: Multilingual parsing from raw text to Universal Dependencies. In Proceedings of the CoNLL 2018 Shared Task: Multilingual Parsing from Raw Text to Universal Dependencies, pages 1–21, Brussels, Belgium. Association for Computational Linguistics.

Citation

MLA
Aji, A. F., et al. “One Country, 700+ Languages: NLP Challenges for Underrepresented Languages and Dialects in Indonesia”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 7226–49, https://doi.org/10.18653/v1/2022.acl-long.500.
APA
Aji, A. F., Winata, G. I., Koto, F., Cahyawijaya, S., Romadhony, A., Mahendra, R., Kurniawan, K., Moeljadi, D., Prasojo, R. E., Baldwin, T., Lau, J. H., & Ruder, S. (2022). One Country, 700+ Languages: NLP Challenges for Underrepresented Languages and Dialects in Indonesia. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7226–7249. https://doi.org/10.18653/v1/2022.acl-long.500
Chicago
Aji, A. F., G. I. Winata, F. Koto, et al. 2022. “One Country, 700+ Languages: NLP Challenges for Underrepresented Languages and Dialects in Indonesia”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7226–49. https://doi.org/10.18653/v1/2022.acl-long.500.
Harvard
Aji, A.F. et al. (2022) “One Country, 700+ Languages: NLP Challenges for Underrepresented Languages and Dialects in Indonesia”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 7226–7249. Available at: https://doi.org/10.18653/v1/2022.acl-long.500.
Vancouver
1. Aji AF, Winata GI, Koto F, et al (2022) One Country, 700+ Languages: NLP Challenges for Underrepresented Languages and Dialects in Indonesia. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 7226–7249

BibTeX

@inproceedings{aji-etal-2022-one,
    title = "One Country, 700+ Languages: {NLP} Challenges for Underrepresented Languages and Dialects in {I}ndonesia",
    author = "Aji, Alham Fikri  and
      Winata, Genta Indra  and
      Koto, Fajri  and
      Cahyawijaya, Samuel  and
      Romadhony, Ade  and
      Mahendra, Rahmad  and
      Kurniawan, Kemal  and
      Moeljadi, David  and
      Prasojo, Radityo Eko  and
      Baldwin, Timothy  and
      Lau, Jey Han  and
      Ruder, Sebastian",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.500/",
    doi = "10.18653/v1/2022.acl-long.500",
    pages = "7226--7249"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/