Ethical Considerations for Machine Translation of Indigenous Languages: Giving a Voice to the Speakers
Manuel MagerElisabeth MagerKatharina KannNgoc Thang Vu
Examines the ethical challenges of developing machine translation for Indigenous languages through direct interviews with community leaders and language activists, establishing practical guidance for respectful data collection and collaborative system design.
Recent advancements in natural language processing have spurred growing interest in developing machine translation systems for low-resource and endangered Indigenous languages. However, these languages are inextricably tied to the cultural identity, sovereignty, and historical experiences of their speaking communities, raising complex ethical challenges around data collection, privacy, and intellectual property. Machine translation projects risk repeating historical patterns of colonial exploitation if outside researchers treat Indigenous communities purely as passive data sources rather than active partners.
The main objective of the article is to examine the ethical considerations of developing machine translation for Indigenous languages and to evaluate the perspectives, expectations, and concerns of native speakers regarding automated translation technologies.
To achieve this, the article surveys established ethical guidelines across linguistics and machine translation before conducting a mixed quantitative and qualitative case study. The study involved a 12-question survey administered to 22 Indigenous language activists, teachers, and community leaders across 12 language groups in the Americas—including Aymara, Maya, Nahua, Quechua, and Zapotec—supplemented by in-depth individual interviews with two participants.
The investigation produced several key findings. First, an overwhelming majority of participants view machine translation positively, expecting it to assist language revitalization, education, and access to public services. Second, 77.3% reported that their communities place no restrictions on sharing their language with outsiders, although researchers must not assume this applies universally. Third, participants emphasized significant ethical risks, specifically pointing to religious and sacred content as sensitive domains that should not be translated without explicit consent; this highlights the hazard of using religious texts like the Bible as standard default training data. Fourth, views on data governance were mixed: while most favored public access to resources, 29.4% believed communities should own the data, 20.6% favored ownership by contributing speakers, and 17% assigned ownership to external research groups. Finally, participants expressed mixed views on releasing imperfect experimental models, with 54.8% open to iterative refinement and 45.5% concerned that inaccurate translations could distort cultural meanings or harm language learning.
These findings demonstrate that ethical machine translation cannot rely on rigid, top-down technical standards. Instead, research teams must actively engage with Indigenous community structures, respect data sovereignty, and navigate traditional knowledge ownership alongside statutory copyright laws. Without direct community oversight, translation systems risk misrepresenting nuanced cultural concepts, threatening internal community privacy, or enabling unauthorized commercial exploitation.
Organizations and researchers working on Indigenous language technologies should establish collaborative partnerships and formal consultation agreements before initiating projects. Practical steps include agreeing on data licensing terms, avoiding the uncritical use of religious corpora, and training Indigenous researchers and engineers to build long-term technical sovereignty. When external permissions and governance structures are ambiguous, researchers can adopt emerging attribution frameworks such as Traditional Knowledge labels.
The conclusions of the article reflect the perspectives of a specific group of 22 participants from the Americas and cannot be generalized to all Indigenous groups globally. While confidence in the identified ethical themes and general community attitudes is high, practitioners must exercise caution and treat each language community as a distinct cultural and political entity requiring tailored local engagement.
- Paper: Language (Technology) is Power: A Critical Survey of “Bias” in NLP, Su Lin Blodgett et al. (2020). Its account of representational and allocational harms gives readers a grounding for the source’s critique of extractive language-technology practices.
- Paper: The State and Fate of Linguistic Diversity and Inclusion in the NLP World, Pratik Joshi et al. (2020). Its analysis of systemic linguistic exclusion establishes the digital-inequality context needed to understand why Indigenous-language MT raises distinct ethical stakes.
- Paper: Ethical and social risks of harm from Language Models, Laura Weidinger et al. (2022). Its taxonomy of language-model harms prepares readers to recognize the risks of exposing sensitive data and deploying unreliable language systems.
- Paper: No Language Left Behind: Scaling Human-Centered Machine Translation, NLLB Team et al. (2022). Its human-centered MT framework provides technical and community-engagement context for the source’s evaluation of Indigenous speakers’ concerns.
No sufficiently relevant recommendations were found.
