Ethical Considerations for Machine Translation of Indigenous Languages: Giving a Voice to the Speakers

Manuel MagerElisabeth MagerKatharina KannNgoc Thang Vu

article2023ACL50 citations

Examines the ethical challenges of developing machine translation for Indigenous languages through direct interviews with community leaders and language activists, establishing practical guidance for respectful data collection and collaborative system design.

Listen

Recent advancements in natural language processing have spurred growing interest in developing machine translation systems for low-resource and endangered Indigenous languages. However, these languages are inextricably tied to the cultural identity, sovereignty, and historical experiences of their speaking communities, raising complex ethical challenges around data collection, privacy, and intellectual property. Machine translation projects risk repeating historical patterns of colonial exploitation if outside researchers treat Indigenous communities purely as passive data sources rather than active partners.

The main objective of the article is to examine the ethical considerations of developing machine translation for Indigenous languages and to evaluate the perspectives, expectations, and concerns of native speakers regarding automated translation technologies.

To achieve this, the article surveys established ethical guidelines across linguistics and machine translation before conducting a mixed quantitative and qualitative case study. The study involved a 12-question survey administered to 22 Indigenous language activists, teachers, and community leaders across 12 language groups in the Americas—including Aymara, Maya, Nahua, Quechua, and Zapotec—supplemented by in-depth individual interviews with two participants.

The investigation produced several key findings. First, an overwhelming majority of participants view machine translation positively, expecting it to assist language revitalization, education, and access to public services. Second, 77.3% reported that their communities place no restrictions on sharing their language with outsiders, although researchers must not assume this applies universally. Third, participants emphasized significant ethical risks, specifically pointing to religious and sacred content as sensitive domains that should not be translated without explicit consent; this highlights the hazard of using religious texts like the Bible as standard default training data. Fourth, views on data governance were mixed: while most favored public access to resources, 29.4% believed communities should own the data, 20.6% favored ownership by contributing speakers, and 17% assigned ownership to external research groups. Finally, participants expressed mixed views on releasing imperfect experimental models, with 54.8% open to iterative refinement and 45.5% concerned that inaccurate translations could distort cultural meanings or harm language learning.

These findings demonstrate that ethical machine translation cannot rely on rigid, top-down technical standards. Instead, research teams must actively engage with Indigenous community structures, respect data sovereignty, and navigate traditional knowledge ownership alongside statutory copyright laws. Without direct community oversight, translation systems risk misrepresenting nuanced cultural concepts, threatening internal community privacy, or enabling unauthorized commercial exploitation.

Organizations and researchers working on Indigenous language technologies should establish collaborative partnerships and formal consultation agreements before initiating projects. Practical steps include agreeing on data licensing terms, avoiding the uncritical use of religious corpora, and training Indigenous researchers and engineers to build long-term technical sovereignty. When external permissions and governance structures are ambiguous, researchers can adopt emerging attribution frameworks such as Traditional Knowledge labels.

The conclusions of the article reflect the perspectives of a specific group of 22 participants from the Americas and cannot be generalized to all Indigenous groups globally. While confidence in the identified ethical themes and general community attitudes is high, practitioners must exercise caution and treat each language community as a distinct cultural and political entity requiring tailored local engagement.

arXiv: 2305.19474

No sufficiently relevant recommendations were found.

Cover for Ethical Considerations for Machine Translation of Indigenous Languages: Giving a Voice to the Speakers

Abstract

In recent years, machine translation has become very successful for high-resource language pairs. This has also sparked new interest in research on the automatic translation of low-resource languages, including Indigenous languages. However, the latter are deeply related to the ethnic and cultural groups that speak (or used to speak) them. The data collection, modeling and deploying machine translation systems thus result in new ethical questions that must be addressed. Motivated by this, we first survey the existing literature on ethical considerations for the documentation, translation, and general natural language processing for Indigenous languages. Afterward, we conduct and analyze an interview study to shed light on the positions of community leaders, teachers, and language activists regarding ethical concerns for the automatic translation of their languages. Our results show that the inclusion, at different degrees, of native speakers and community members is vital to performing better and more ethical research on Indigenous languages.

Table of Contents

  • 1 Introduction
  • 2 Defining “Endangered Language”
  • 3 Ethics and MT
  • 3.1 Ethics and Data
  • 3.2 Ethics and Human Translation
  • 3.3 Ethics and Machine Translation
  • 4 The Speakers’ Opinions
  • 4.1 Study Design
  • 4.2 Results
  • 5 Discussion
  • 6 Conclusion
  • References
  • A Questionnaire
  • B Complete answers of the open questions
  • C Spanish Version: Consideraciones éticas para la traducción automática de lenguas indígenas: dar voz a los hablantes
  • C.1 Introducción
  • C.2 Definiendo “lenguaje en peligro”
  • C.3 Ética y MT
  • C.4 Las opiniones de los oradores
  • C.5 Diseño del estudio
  • C.6 Resultados
  • C.7 Discusión
  • C.8 Conclusión

Knowls

  1. Knowl 1 — Community involvement is central, but permissions cannot be presumed

    empirical result

    In a case study of 22 Indigenous language activists, teachers, and community leaders from the Americas, 85% favored integrating community members into machine-translation (MT) research. At the same time, 77.3% reported that their community had no restrictions on sharing its language with outsiders, while some communities did have restrictions; only one participant selected a requirement for explicit permission from a community body. The findings support consulting and involving speakers, but do not justify assuming that any community consents to research or language sharing.

  2. Knowl 2 — Study design and scope

    experimental setup

    The study collected survey responses from 22 volunteer language activists, teachers, and community leaders associated with 12 Indigenous language communities in the Americas: Aymara, Chatino, Maya, Mazatec, Mixe, Nahua, Otomí, Quechua, Tenek, Tepehuano, Kichwa of Otavalo, and Zapotec. A 12-question questionnaire covered participants’ language background, sharing restrictions, MT benefits and risks, research inclusion, data ownership, and best practices; open responses supplied qualitative detail. The researchers also conducted two one-to-one interviews, with a Mixe activist and a Chatino linguist. Participants were informed about the study, volunteered, and were not asked for personal identifiers. The study was designed as a case study, not as a representative survey of Indigenous peoples.

  3. Knowl 3 — Participants saw potential benefits in MT, especially for everyday communication and public services

    empirical result

    Most survey participants viewed an MT system for their language as beneficial. In the study’s topic selections, everyday conversation was chosen by 15 participants, while science and education, culture and traditions, and medicine and health were each chosen by 14. Open responses described possible uses including language learning and teaching, communication with people who do not speak the language, support in hospitals and government offices, increased visibility and social use, and preservation or revitalization. These responses show interest in MT alongside the participants’ concerns about how systems are built and used.

  4. Knowl 4 — Participants’ assessments of MT danger were mixed, with translation quality and cultural fit as concerns

    empirical result

    On a five-point scale where 1 meant not dangerous and 5 meant dangerous, the survey chart records 13 participants selecting 1, two selecting 2, four selecting 3, none selecting 4, and three selecting 5. Open responses identify risks that include inaccurate translations of cultural concepts, distortion of appropriate language use, and loss of meaning where expressions are context-dependent or not directly translatable. Participants also worried that numerous language variants could lead developers to impose a standardized form, which they viewed as a threat to linguistic diversity.

  5. Knowl 5 — Sensitive domains require community-specific decisions; the paper advises caution with Bible corpora

    model/method

    Although most open answers did not name a topic that should be excluded from MT, religion was a recurring concern. Participants mentioned sacred songs, ritual or ceremonial speech, sacred knowledge, and the possible disclosure of ceremonial secrets; some also raised concerns about Western religion, political matters, laws, medicine, and private community issues. The authors recommend that researchers consult the relevant community before using religious material, including Bible translations. Where researchers lack a close relationship with a community—as in a large-scale multilingual MT experiment—the paper recommends avoiding the Bible as a default corpus for an Indigenous language.

  6. Knowl 6 — Data accessibility and ownership received different levels of support

    empirical result

    Participants selected among four data-policy options, and their ownership preferences were divided: 32.4% selected public availability without restrictions, 29.4% community ownership, 20.6% ownership by the speakers who produced the translations, and 17.6% ownership by the research group that compiled the data. The authors report that open access to research materials was valued in participants’ comments, but the distribution of ownership preferences shows that accessibility and ownership should not be treated as the same question. The paper recommends discussing data rights with participants or community representatives before work begins, taking account of local governance and the possible conflict between communal understandings of knowledge and national legal systems. It also calls for special agreements, including financial terms, if collected data are used commercially.

  7. Knowl 7 — Views on releasing a low-quality experimental system were nearly evenly split

    empirical result

    Participants disagreed about making an MT system available when its translation quality was poor. The survey graphic reports that 54.5% thought it would be useful to make the system available so progress could be seen and corrections provided, while 45.5% preferred that it not be used or released to the public. The accompanying prose gives the first figure as 54.8%, an inconsistency within the paper. Participants who opposed release connected poor quality to incorrect cultural translations and the risk of teaching users the language incorrectly; supporters commonly framed early release as acceptable if the system could be improved and corrected.

  8. Knowl 8 — Ethical practice combines consultation, cultural respect, and sharing research outputs

    model/method

    The paper synthesizes recurring principles in ethical guidance for research with Indigenous communities into three themes: consultation, negotiation, and mutual understanding about project aims and outcomes; respect for each community’s history, culture, and worldview, including involvement of local researchers, speakers, or governing bodies; and making research outputs and materials available for community use. The authors emphasize that applying these principles requires continued, case-specific discussion rather than a single universal rulebook. For MT, they recommend incorporating community members into development, feedback, and quality checking, and recognizing their contributions.

  9. Knowl 9 — Data and technology sovereignty imply community control and participation in development

    model/method

    The paper treats data sovereignty as community control over data, knowledge, and cultural expressions created by Indigenous communities, including decisions about ownership and licensing of resulting data products. It recommends reaching these arrangements through consultation and direct collaboration rather than assuming that standard legal or licensing frameworks fit every community; it notes that Creative Commons licenses may not be designed for Indigenous needs. The authors also report participants’ interest in having technology for their own languages and in taking part directly in developing it. They recommend greater inclusion of Indigenous researchers in NLP and valuing the training of Indigenous researchers and engineers.

  10. Knowl 10 — Findings are limited to a nonrepresentative case study in the Americas

    limitation

    The study involved 22 individuals from selected Indigenous language communities in the Americas and does not establish the views of all Indigenous communities, nations, or language speakers. The authors state that participants expressed personal opinions that should not be taken as their communities’ official positions. Because communities have different histories and circumstances, the paper presents its recommendations as broad, non-normative guidance to be adapted through consultation, not as universal rules for MT research.

Coverage note — The broad background review of endangered-language categories and historical examples of colonial translation is not extracted as separate knowls because it primarily contextualizes the study; the recurring ethical principles and the study’s own recommendations are included. Repetitive individual open-response quotations and the Spanish translation appendix are omitted.

References

  1. 1.Željko Agic and Ivan Vuli ´ c. 2019. Jw300: A wide- ´ coverage parallel corpus for low-resource languages. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3204–3210.
  2. 2.Yasnaya Elena Aguilar-Gil. 2020. A modest proposal to save the world. https://restofworld.org/2020/saving-the-world-through-tequiology/.
  3. 3.Tecla Alfredo and Garza Ramos Alberto. 1978. Teoría, métodos y técnicas en la investigación social. México. Ediciones de Cultura Popular, pages 11–74.
  4. 4.Mikel Artetxe, Gorka Labaka, Eneko Agirre, and Kyunghyun Cho. 2018. Unsupervised neural machine translation. In 6th International Conference on Learning Representations, ICLR 2018.
  5. 5.Peter K Austin and Julia Sallabank. 2013. Endangered languages: An introduction.
  6. 6.Heriberto Avelino. 2021. Nuevas perspectivas en documentación lingÜÍstica. Ciencias Antropologicas, pages 93–119.
  7. 7.Dzmitry Bahdanau, Kyung Hyun Cho, and Yoshua Bengio. 2015. Neural machine translation by jointly learning to align and translate. In 3rd International Conference on Learning Representations, ICLR 2015.
  8. 8.Carlos Barron-Romero, Mager Manuel, and Reyes Avilés Fernando. 2016. Richard feynman, los alfabetos y los lenguajes.
  9. 9.Nicolas Beauclair. 2010. Éticas andinas y discursos de reivindicaciones indígenas: asociando tradición y alter-mundialización. Tinkuy: Boletín de investigación y debate, (12):9–34.
  10. 10.Steven Bird. 2020. Decolonising speech and language technology. In Proceedings of the 28th International Conference on Computational Linguistics, pages 3504–3519, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  11. 11.Gina Bustamante, Arturo Oncevay, and Roberto Zariquiey. 2020. No data to crawl? monolingual corpus creation from PDF files of truly low-resource languages in Peru. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 2914–2923, Marseille, France. European Language Resources Association.
  12. 12.Hilary M Carey. 2010. Lancelot threlkeld, biraban, and the colonial bible in australia. Comparative Studies in Society and History, 52(2):447–478.
  13. 13.Andrew Chesterman. 2001. Proposal for a hieronymic oath. The translator, 7(2):139–154.
  14. 14.Eric Cheyfitz. 1997. The poetics of imperialism: Translation and colonization from The Tempest to Tarzan. University of Pennsylvania Press.
  15. 15.Christos Christodouloupoulos and Mark Steedman. 2015. A massively parallel corpus: the bible in 100 languages. Language resources and evaluation, 49(2):375–395.
  16. 16.Alain Couillault, Karën Fort, Gilles Adda, and Hugues de Mazancourt. 2014. Evaluating corpora documentation with regards to the ethics and big data charter. In Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC’14), pages 4225–4229, Reykjavik, Iceland. European Language Resources Association (ELRA).
  17. 17.Ewa Czaykowska-Higgins. 2009. Research models, community engagement, and linguistic fieldwork: Reflections on working within canadian indigenous communities.
  18. 18.Erica-Irene A Daes. 1993. Discrimination against Indigenous peoples: Study on the protection of the cultural and intellectual property of Indigenous peoples. United Nations, Economic and Social Council.
  19. 19.TA DelValls. 1978. El instituto lingüístico de verano, instrumento del imperialismo. Nueva Antropología, 3(9):117–142.
  20. 20.Lise M Dobrin. 2009. Sil international and the disciplinary culture of linguistics: Introduction. Language, 85(3):618–619.
  21. 21.Stephen Doherty. 2016. Translations| the impact of translation technologies on the process and product of translation. International journal of communication, 10:23.
  22. 22.Arienne M Dwyer. 2006. Ethics and practicalities of cooperative fieldwork and analysis. Essentials of language documentation, 178.
  23. 23.Paola Enríquez. 2019. El rol de la lengua kichwa en la construcción de la identidad en la población indígena de cañar. Lenguas en contacto: desafíos en la diversidad, page 267.
  24. 24.Joseph Errington. 2001. Colonial linguistics. Annual review of anthropology, 30(1):19–39.
  25. 25.José Luis Soberanes Fernández. 2022. La libertad religiosa en el pueblo huichol religious freedom in the huichol community. Revista de Garantismo y Derechos Humanos, page 43.
  26. 26.Karën Fort, Gilles Adda, Benoît Sagot, Joseph Mariani, and Alain Couillault. 2011. Crowdsourcing for language resource development: criticisms about amazon mechanical turk overpowering use. In Language and Technology Conference, pages 303–314. Springer.
  27. 27.Rachael Gilmour. 2007. Missionaries, colonialism and language in nineteenth-century south africa. History Compass, 5(6):1761–1777.
  28. 28.Ken Hale. 1992. Endangered languages: On endangered languages and the safeguarding of diversity. language, 68(1):1–42.
  29. 29.Mika Hämäläinen. 2021. Endangered languages are not low-resourced! arXiv preprint arXiv:2103.09567.
  30. 30.Daniel Hershcovich, Stella Frank, Heather Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, et al. 2022. Challenges and strategies in cross-cultural nlp. arXiv preprint arXiv:2203.10020.
  31. 31.Kenneth C Hill. 2002. On publishing the hopi dictionary. Making dictionaries: Preserving Indigenous languages of the Americas, pages 299–311.
  32. 32.Csilla Horváth, Norbert Szilágyi, Veronika Vincze, and Ágoston Nagy. 2017. Language technology resources and tools for mansi: an overview. In Proceedings of the Third Workshop on Computational Linguistics for Uralic Languages, pages 56–65.
  33. 33.ILO Ilo. 1989. C169-indigenous and tribal peoples convention. In Convention concerning Indigenous and Tribal Peoples in Independent Countries. Geneva: ILO.
  34. 34.Alfredo Tecla Jiménez and Alberto Garza Ramos. 1985. Teoría, métodos y técnicas en la investigación social. Eds. Taller Abierto.
  35. 35.Marcin Junczys-Dowmunt. 2019. Microsoft translator at wmt 2019: Towards large-scale document-level neural machine translation. arXiv preprint arXiv:1907.06170.
  36. 36.Judith L. Klavans. 2018. Computational challenges for polysynthetic languages. In Proceedings of the Workshop on Computational Modeling of Polysynthetic Languages, pages 1–11, Santa Fe, New Mexico, USA. Association for Computational Linguistics.
  37. 37.Guillaume Lample, Alexis Conneau, Ludovic Denoyer, and Marc’Aurelio Ranzato. 2018. Unsupervised machine translation using monolingual corpora only. In International Conference on Learning Representations.
  38. 38.Jochen L. Leidner and Vassilis Plachouras. 2017. Ethical by design: Ethics best practices for natural language processing. In Proceedings of the First ACL Workshop on Ethics in Natural Language Processing, pages 30–40, Valencia, Spain. Association for Computational Linguistics.
  39. 39.Wesley Leonard. 2020. Producing language reclamation by decolonising’language’.
  40. 40.Zoey Liu, Crystal Richardson, Richard Hatcher Jr, and Emily Prud’hommeaux. 2022. Not always about you: Prioritizing community needs when developing endangered language technology. arXiv preprint arXiv:2204.05541.
  41. 41.Monika Ludescher. 2001. Instituciones y prácticas coloniales en la amazonía peruana: pasado y presente. Indiana, pages 313–359.
  42. 42.Manuel Mager, Ximena Gutierrez-Vasques, Gerardo Sierra, and Ivan Meza-Ruiz. 2018. Challenges of language technologies for the indigenous languages of the Americas. In Proceedings of the 27th International Conference on Computational Linguistics, pages 55–69, Santa Fe, New Mexico, USA. Association for Computational Linguistics.
  43. 43.Arya D McCarthy, Rachel Wicks, Dylan Lewis, Aaron Mueller, Winston Wu, Oliver Adams, Garrett Nicolai, Matt Post, and David Yarowsky. 2020. The johns hopkins university bible corpus: 1600+ tongues for typological exploration. In Proceedings of the 12th language resources and evaluation conference, pages 2884–2892.
  44. 44.Guillermo Meza Salcedo. 2017. Ética de la investigación desde el pensamiento indígena: derechos colectivos y el principio de la comunalidad. Revista de Bioética y Derecho, (41):141–159.
  45. 45.Christopher Moseley. 2010. Atlas of the World’s Languages in Danger. Unesco.
  46. 46.Tejaswini Niranjana. 1990. Translation, colonialism and rise of english. Economic and Political Weekly, pages 773–779.
  47. 47.Tejaswini Niranjana. 1992. Siting translation: History, post-structuralism and the colonial context. URL: https://doi. org/10.1525/9780520911369 [2021-04-01].
  48. 48.Hanns-Gregor Nissing and Jörn Müller. 2009. Grundpositionen philosophischer Ethik: von Aristoteles bis Jürgen Habermas. WBG.
  49. 49.Azucena Palacios. 2008. La lengua como instrumento de identidad y diferenciación: más allá de la influencia de las lenguas amerindias. De moneda nunca usada.
  50. 50.Ellie Pavlick, Matt Post, Ann Irvine, Dmitry Kachaev, and Chris Callison-Burch. 2014. The language demographics of amazon mechanical turk. Transactions of the Association for Computational Linguistics, 2:79–92.
  51. 51.Suwilai Premsrirat and Dennis Malone. 2003. Language development and language revitalization in asia. SIL International, Mahidol University.
  52. 52.James Rachels and Stuart Rachels. 1986. The elements of moral philosophy. Temple University Press Philadelphia.
  53. 53.Guillermo-Meza Salcedo. 2016. El ‘vivir nosotros’ amerindio vs ‘decir nosotros’ de la globalización. Cuadernos de filosofía latinoamericana, 37(114):151–166.
  54. 54.Florian Alexander Schmidt. 2013. The good, the bad and the ugly: Why crowdsourcing needs ethics. In 2013 International Conference on Cloud and Green Computing, pages 531–535. IEEE.
  55. 55.Lane Schwartz. 2022. Primum non nocere: Before working with indigenous data, the acl must confront ongoing colonialism.
  56. 56.Linda Tuhiwai Smith. 2021. Decolonizing methodologies: Research and indigenous peoples. Bloomsbury Publishing.
  57. 57.Gabriel Stanovsky, Noah A Smith, and Luke Zettlemoyer. 2019. Evaluating gender bias in machine translation. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 1679–1684.
  58. 58.Maria Tymoczko. 2006. Translation: Ethics, ideology, action. The Massachusetts Review, 47(3):442–461.
  59. 59.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In NIPS.
  60. 60.Shiyue Zhang, Ben Frey, and Mohit Bansal. 2022. How can nlp help revitalize endangered languages? a case study and roadmap for the cherokee language. arXiv preprint arXiv:2204.11909.

Citation

MLA
Mager, M., et al. “Ethical Considerations for Machine Translation of Indigenous Languages: Giving a Voice to the Speakers”. arXiv, 2023, http://arxiv.org/abs/2305.19474v1.
APA
Mager, M., Mager, E., Kann, K., & Vu, N. T. (2023). Ethical Considerations for Machine Translation of Indigenous Languages: Giving a Voice to the Speakers. arXiv. http://arxiv.org/abs/2305.19474v1
Chicago
Mager, M., E. Mager, K. Kann, and N. T. Vu. 2023. “Ethical Considerations for Machine Translation of Indigenous Languages: Giving a Voice to the Speakers”. arXiv. http://arxiv.org/abs/2305.19474v1.
Harvard
Mager, M. et al. (2023) “Ethical Considerations for Machine Translation of Indigenous Languages: Giving a Voice to the Speakers”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2305.19474v1.
Vancouver
1. Mager M, Mager E, Kann K, Vu NT (2023) Ethical Considerations for Machine Translation of Indigenous Languages: Giving a Voice to the Speakers. arXiv

BibTeX

@article{mager2023ethical,
  title = {Ethical Considerations for Machine Translation of Indigenous Languages: Giving a Voice to the Speakers},
  author = {Mager, Manuel and Mager, Elisabeth and Kann, Katharina and Vu, Ngoc Thang},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2305.19474v1},
  eprint = {2305.19474}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/