Designing Culturally Aligned AI Systems For Social Good in Non-Western Contexts

Deepak Varuvel DennisonMohit JainTanuja GanuAditya Vashistha

article2025International Conference on Human Factors in Computing Systems5 citations

Derives twelve practical design guidelines for building culturally grounded AI systems in non-Western contexts by analyzing real-world deployments across seven countries and eighteen languages.

Listen

Artificial intelligence applications are expanding rapidly into high-stakes social sectors—including healthcare, education, agriculture, and law—across non-Western regions. While these technologies offer significant potential to expand access and improve service delivery, standard commercial models frequently fail in these environments due to Western-centric training data, language deficits, and a lack of cultural awareness. In high-stakes settings, such deficiencies risk generating harmful medical advice, inaccurate pedagogical materials, or misleading legal translations. The article evaluates how real-world artificial intelligence systems are contextualized, deployed, and sustained in non-Western environments, demonstrating the operational and human requirements necessary for safe, equitable, and effective implementation.

To examine these dynamics, the article analyzes eight deployments operating across seven countries (India, Bangladesh, Colombia, Kenya, Ethiopia, Nigeria, and Ghana) and supporting 18 languages, with user bases ranging from hundreds to approximately 350,000 individuals. The researchers conducted semi-structured interviews with 17 key informants—comprising 10 technical developers and 7 domain experts—and complemented these qualitative accounts with secondary project documentation, followed by member-checking to validate findings against practitioner experiences.

The analysis identifies six core factors that dictate implementation success: Language, Institution, Safety, Task, End-User Demography, and Domain. First, foundational language support requires substantial intervention because standard models perform poorly on low-resource languages and dialects, often demanding custom translation pipelines, local glossaries, or newly trained speech-recognition engines. Second, institutional alignment and trust represent the primary gatekeepers for scale; systems succeed only when they adhere strictly to local policies, mandated reporting formats, and existing organizational workflows. Third, human-in-the-loop oversight and curated knowledge bases—rather than automated safeguards alone—serve as the critical backbone for safety and risk mitigation. Fourth, technical configurations must adapt to task and environmental constraints, including high ambient field noise, low-connectivity offline needs, and latency tolerances. Fifth, demographic realities such as low literacy, device limitations, and age require flexible interaction channels, making voice modalities and accessible platforms like basic SMS or WhatsApp far more effective than standalone web applications.

These findings demonstrate that technology cannot substitute for institutional capacity; rather, artificial intelligence acts as an amplifier that requires strong existing human infrastructures to function safely. Successful deployments rely heavily on continuous, intensive human labor from domain experts and field workers who curate specialized data, validate outputs, and manage edge cases. This evidence challenges the common techno-solutionist assumption that advanced foundation models can deliver immediate social impact out of the box, showing instead that long-term efficacy depends on continuous contextual adaptation and local alignment.

For practitioners, funders, and policymakers, the article provides actionable guidance structured around twelve practical recommendations. Stakeholders should treat technical developers and local domain experts as equal partners across the entire lifecycle, favor modular system designs that can adapt to rapid model shifts, and conduct small-scale pilots to identify real-world failure modes before broad deployment. Furthermore, organizations should establish layered safety architectures that combine lightweight screening models, reliable frontier models, and dedicated human escalation paths. Before committing large budgets to resource-heavy model pretraining, teams should leverage accessible in-context methods, such as prompt engineering and domain-specific retrieval-augmented generation.

The conclusions are drawn from a qualitative study of eight non-profit and cross-organizational initiatives, which introduces certain contextual boundaries. The findings reflect specific operational realities within selected geographies and low-resource settings, and long-term longitudinal maintenance costs remain an ongoing operational uncertainty. Nevertheless, the empirical breadth across four diverse sectors provides high confidence in the central conclusion: deploying artificial intelligence for social good in non-Western contexts requires sustained investment in human collaboration, institutional integration, and localized safeguards.

arXiv: 2509.16158
Cover for Designing Culturally Aligned AI Systems For Social Good in Non-Western Contexts

Abstract

AI technologies are increasingly deployed in high-stakes domains such as education, healthcare, law, and agriculture to address complex challenges in non-Western contexts. This paper examines eight real-world deployments spanning seven countries and 18 languages, combining 17 interviews with AI developers and domain experts with secondary research. Our findings identify six cross-cutting factors - Language, Institution, Safety, Task, End-User Demography, and Domain - that structured how systems were designed and deployed. These factors were shaped by Sociocultural (diversity, practices), Institutional (resources, policies), and Technological (capabilities, limits) influences. We find that building effective AI systems required extensive collaboration between AI developers and domain experts, with human resources proving more critical to achieving safe and effective outcomes in high-stakes domains than technological expertise alone. Additionally, we present 12 guidelines synthesizing these dynamics for designing AI for social good systems that are culturally grounded, equitable, and responsive to the needs of non-Western contexts.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methods
  • 3.1 AI Project Selection
  • 3.2 Semi-structured Interviews
  • 3.3 Secondary Research
  • 3.4 Analysis
  • 3.5 Researcher Positionality
  • 4 Findings
  • 4.1 Language
  • 4.1.1 Language and Technical Considerations
  • 4.1.2 Human Infrastructure
  • 4.2 Institution
  • 4.2.1 Aligning with Institutional Requirements and Knowledge
  • 4.2.2 Institutional Support
  • 4.3 Safety
  • 4.3.1 Safety Measures Before Processing Inputs
  • 4.3.2 Measures During Workflow Processing
  • 4.4 Task
  • 4.4.1 Environmental Requirements
  • 4.4.2 Task Requirements
  • 4.4.3 Latency Needs
  • 4.5 End-User Demography
  • 4.5.1 Geography and Culture
  • 4.5.2 Other Demographic Factors
  • 4.6 Domain
  • 4.6.1 Curated Knowledge Bases over LLM World Knowledge
  • 4.6.2 Human-in-the-Loop Workflows
  • 5 Discussion
  • 5.1 Sociocultural Influences
  • 5.2 Institutional Influences
  • 5.3 Technological Influences
  • 6 Conclusion
  • References
  • A AI Adaptation Techniques
  • B AI Developer Interview Protocol
  • C Domain Expert Interview Protocol

Knowls

  1. Knowl 1 — LISTED Framework for Culturally Aligned AI in Non-Western Settings

    theoretical result

    The LISTED framework conceptualizes six cross-cutting sociotechnical factors that govern the design, implementation, and long-term viability of AI systems deployed for social good in non-Western high-stakes contexts (such as healthcare, agriculture, education, and law). These six factors are shaped by three higher-order forces—Sociocultural, Institutional, and Technological influences:

    • Language: Grounding systems in local linguistic realities by addressing foundational capability gaps in low-resource languages, managing regional disparities in foundation model support, and building localized vocabularies and voice resources.
    • Institution: Aligning system behavior and data flows with government regulations, statutory document formats, regional agricultural/health directives, privacy laws, and leveraging institutional authority to establish community trust.
    • Safety: Structuring multi-stage protection mechanisms that combine input-stage filtering (such as custom classifiers for code-mixed queries and prompt guardrails), automated LLM auditing agents, and human expert verification workflows.
    • Task: Tailoring AI pipelines to physical operating environments (such as high ambient background noise), specific functional goals (such as preserving user speech errors in assessment rather than smoothing them), domain translation constraints, and latency tolerance.
    • End-User Demography: Customizing modalities (voice-first vs. text), delivery mediums (SMS vs. messaging apps vs. web interfaces), and content generation to accommodate varying literacy levels, age, visual impairment, gendered vocabularies, and socio-economic realities.
    • Domain: Grounding generative outputs in expert-curated knowledge bases via Retrieval-Augmented Generation (RAG) rather than generic model world knowledge, resolving local lexical nuances that cause retrieval failures, and establishing domain-specific Human-in-the-Loop (HIL) workflows.
  2. Knowl 2 — Empirical Study of Eight Real-World AI Deployments in Non-Western High-Stakes Domains

    empirical result

    An empirical qualitative investigation of eight real-world, non-profit AI systems deployed in high-stakes social domains across seven non-Western countries (India, Bangladesh, Colombia, Kenya, Ethiopia, Nigeria, and Ghana) and supporting 18 languages yielded concrete operational characteristics of AI for Social Good deployments:

    • VetBot (India, 2025, Agriculture): Web-based veterinary assistant (<500 users) supporting Tamil, English, Kannada, Hindi, and Malayalam using LLMs, Speech-to-Text (STT), and Text-to-Speech (TTS) with RAG and Prompt Engineering.
    • FarmerChat (Ethiopia, India, Kenya, Nigeria, 2023, Agriculture): Mobile agricultural advisory (~15,000 farmers) supporting Swahili, Amharic, Hausa, Hindi, Odiya, Telugu, Kannada, and English using LLMs, STT, and TTS with RAG, Prompt Engineering, Glossaries, and Fine-Tuning.
    • PROMPTS (Ghana, Kenya, Nigeria, 2023, Healthcare): SMS-based maternal health system (~350,000 women; ~1.1 million SMS messages/month) supporting Swahili, Hausa, Yoruba, Xhosa, and Zulu using LLMs and BERT-based classifiers with Continual Pre-training, Instruction Tuning, and Human-in-the-Loop triage.
    • CataractBot (India, 2024, Healthcare): WhatsApp-based perioperative ophthalmic assistant (~2,000 patients) supporting English, Kannada, Hindi, Telugu, Tamil, and Urdu using LLMs, STT, and TTS with RAG, Human-in-the-Loop doctor verification, and Prompt Engineering.
    • ASHABot (India, 2024, Healthcare): WhatsApp-based assistant for community health workers (~5,000 workers) supporting Hindi, English, Telugu, and Marathi using LLMs, STT, and TTS with RAG, Glossaries, Prompt Engineering, and Human-in-the-Loop escalation.
    • ReadAI (Colombia, India, 2025, Education): Mobile child reading assessment tool (~5,000 children) supporting Hindi and Marathi using custom-trained STT and Human-in-the-Loop evaluator review.
    • Shiksha Copilot (India, 2024, Education): Web-based teacher lesson planner (~1,100 teachers) supporting Kannada, Telugu, Hindi, and English using LLMs with RAG, Instruction Tuning, Glossaries, Human-in-the-Loop teacher curation, and Prompt Engineering.
    • LegalTranslateAI (Bangladesh, India, 2019, Legal): Web-based court judgment translation system (~120,000 judgments) supporting Hindi, Kannada, Tamil, Telugu, Punjabi, Marathi, Gujarati, Malayalam, Bengali, and Urdu using Neural Machine Translation (NMT) with Fine-Tuning, Glossaries, and Human-in-the-Loop translator/judicial verification.
  3. Knowl 3 — Taxonomy of In-Weight and In-Context Adaptation Techniques for AI Localization

    definition

    Techniques for adapting and culturally aligning AI systems for non-Western, high-stakes deployments are categorized into two structural paradigms based on parameter modification:

    1. In-Weight Adaptations: Modifications that directly alter the learnable parameters or internal representations of the model:

      • Model Training: Constructing and training models from scratch on specialized, locally collected datasets (e.g., training custom acoustic STT models on domain- and child-specific speech recordings).
      • Continual Pre-training: Unsupervised training of an existing pre-trained model on domain-specific or low-resource language corpora to expand linguistic and cultural representation.
      • Fine-Tuning: Supervised training of a pre-trained model on labeled, domain-specific paired datasets (such as parallel legal corpora or agricultural query-response sets) to improve specialized task accuracy.
      • Instruction Tuning: Updating model weights using curated instruction–response pairs to ensure models follow specialized directives and conform to regional communication styles.
    2. In-Context Adaptations: Interventions that steer model behavior during orchestration and inference while leaving model parameters intact:

      • Prompt Engineering: Refining natural language instructions, persona definitions, and few-shot examples supplied in the inference prompt.
      • Retrieval-Augmented Generation (RAG): Querying external, authoritative domain knowledge bases to retrieve factual chunks and ground generative model responses.
      • Glossary Adaptation: Deterministically injecting curated domain- or dialect-specific vocabulary mappings into input pre-processing or output post-processing.
      • Human-in-the-Loop (HIL): System workflows that route generated outputs or out-of-distribution user queries to human experts for verification, editing, or manual response generation before final delivery.
  4. Knowl 4 — Language Adaptation Mechanisms: Hybrid Model Pipelining, Domain-Penalized ASR Scoring, and Glossary Pivoting

    empirical result

    Deploying language technologies in non-Western contexts requires specialized technical strategies to navigate low baseline capabilities and high evaluation error in non-English languages:

    • Hybrid Generative Pipelining: When an individual model cannot provide both pedagogical/domain reasoning and local language fluency, tasks are split across models. In the Shiksha Copilot educational system, GPT-4 generated high-quality pedagogical lesson plans in English, while Sarvam 2B (an Indic LLM) translated the generated structure into fluent Kannada. This decoupled pipeline resolved the trade-off where frontier models struggled with Kannada generation and smaller localized models lacked complex pedagogical planning depth.
    • Domain-Penalized ASR Evaluation with Agree Score: Standard Word Error Rate (WER) fails to capture critical domain errors in agricultural speech. In FarmerChat, ASR model selection was evaluated against gold-standard human transcripts using a fine-tuned Llama 3 model that computed an "agree score." This metric heavily penalized recognition errors on agricultural terms (crop varieties, pests, diseases, chemical units, numerals), which represented ~20% of user vocabulary. A mis-transcription of "masoor" (red lentil) as "mushroom" completely alters the agricultural recommendation, rendering general WER insufficient.
    • Dialect Glossary Pivoting: For unsupported regional dialects lacking base LLM support (such as Bhojpuri in Bihar, India), systems co-develop deterministic bidirectional glossaries mapping dialect terms to a supported regional language (Hindi) before model processing, avoiding the prohibitive cost of training custom dialect models from scratch.
  5. Knowl 5 — Multi-Tiered Safety Architectures in Multilingual and High-Stakes Deployments

    model/method

    Ensuring safety in high-stakes non-Western AI deployments requires a multi-tiered architecture that integrates lightweight triage models, in-weight data hardening, automated LLM auditing, and human escalation:

    1. Pre-Input Code-Mixed Triage: Standard commercial safety classifiers fail on code-mixed queries (vernacular languages mixed with English). Deployments utilize custom lightweight classifiers (e.g., a RoBERTa model fine-tuned on ~100,000 annotated user queries in PROMPTS) to classify message intent and assign emergency-risk scores. In maternal health, this automated triage routes ~80% of routine non-medical questions safely to LLMs while deflecting critical or sensitive maternal health queries to dedicated human medical desks.
    2. In-Weight Adversarial Hardening: Continual pre-training and instruction tuning incorporate multi-lingual safety datasets with adversarial prompts (such as BeaverTails) to prevent safety guardrail degradation when prompts are submitted in non-English languages.
    3. LLM Validation Agents: Independent LLM auditing agents evaluate generated responses against safety guidelines and factual references before output delivery, rejecting unaligned candidate responses.
    4. Domain-Specific Human Verification: In high-liability legal and medical contexts, fully automated output is restricted. In CataractBot, ophthalmic responses require doctor review before delivery and carry visual verification badges. In LegalTranslateAI, NMT-generated court drafts require manual editing by professional court translators, judicial sign-off, and explicit disclaimers stating that the translated document carries no legal standing.
  6. Knowl 6 — Strategic Regional Clustering and Demographic Localization of AI Systems

    empirical result

    Accommodating granular cultural, demographic, and linguistic variation within scalable AI systems is operationalized through specific localization techniques:

    • Strategic Regional Clustering: Because hyper-local contextualization across every village dialect and cultural preference is computationally and logistically impossible at scale, systems aggregate geographic zones into representative clusters. In Shiksha Copilot, domain experts clustered Karnataka into three distinct regions (North, South, and Coastal) based on dialect, cuisine, climate, flora, and fauna. Generative prompts produced options for these three clusters, delegating final selection to local classroom teachers.
    • Linguistic Diversity Index for Data Sampling: In training custom acoustic STT models for diverse children (ReadAI), geographic coverage was optimized across districts using the Linguistic Diversity Index—which quantifies the number and distribution of spoken mother tongues—to capture maximum linguistic variance across 10 representative districts in Uttar Pradesh rather than attempting exhaustive state-wide collection.
    • Demography-Driven Modality and Channel Selection: In rural agricultural and healthcare settings (VetBot, FarmerChat, CataractBot, ASHABot), low literacy, visual impairments, and manual labor require voice-first interfaces (STT/TTS) over text. At the infrastructure layer, services serving low-income mothers (PROMPTS) operate via SMS to eliminate mobile data costs, whereas Indian deployments leverage WhatsApp to minimize user onboarding overhead despite significant logistical complexity with commercial messaging APIs.
  7. Knowl 7 — Task-Specific Architectural Adaptations: Acoustic Preservation, Environmental Conditioning, and Offline Inference

    empirical result

    Domain task requirements in high-stakes non-Western settings necessitate fundamental departures from standard AI modeling assumptions:

    • Acoustic-Preserving ASR vs. Language Model Autocorrection: Standard ASR systems prioritize semantic reconstruction by using internal language models to smooth out and correct acoustic errors. In reading assessment applications (ReadAI), this smoothing behavior undermines the core task of identifying student pronunciation errors. The ASR architecture was modified to prioritize the raw acoustic signal over language model prediction, ensuring that student miscues were accurately captured and preserved for evaluation.
    • In-Situ Environmental Acoustic Conditioning: Speech models trained on studio-recorded or noise-canceled audio failed when deployed in real-world settings. ASR models had to be retrained on audio captured directly within operational environments, including noisy rural classrooms (ReadAI) and open agricultural fields (FarmerChat), to ensure robust recognition under ambient environmental noise.
    • Neural Machine Translation vs. LLM Generative Translation: In legal translation (LegalTranslateAI), developers intentionally retained fine-tuned Neural Machine Translation (NMT) models over Large Language Models despite the latter's superior handling of long-form text, because legal translations prioritized strict, deterministic vocabulary control and standardized statutory phrasing over generative fluency.
    • Offline On-Device Inference: In low-connectivity environments where assessors conduct hundreds of evaluations daily (ReadAI), speech models were deployed for local on-device inference on mobile hardware, eliminating network latency and dependency on cellular data.
  8. Knowl 8 — Institutional Alignment, Legitimacy, and Human-in-the-Loop Capacity Bottlenecks

    empirical result

    Institutional dynamics directly govern whether AI interventions transition from pilot projects to sustained real-world adoption:

    • Mandatory Administrative Alignment: System adoption is contingent on matching official administrative templates and regulatory policies. In Shiksha Copilot, teachers only adopted generated lesson plans once the formatting matched state-mandated administrative structures (which varied between states such as Karnataka and Telangana), as lesson plans function as formal supervisory records. In FarmerChat, knowledge retrieval was strictly isolated to local state agriculture department policies (e.g., restricting retrieval to natural farming documents in Andhra Pradesh) to ensure advice complied with binding regional agricultural mandates.
    • Operational Capacity in Human-in-the-Loop Systems: HIL workflows fail when they impose unmanageable operational burdens on existing human infrastructure. While CataractBot successfully routed out-of-domain queries to doctors who were already supporting patients via existing WhatsApp groups, ASHABot failed when routing unanswered questions to community health supervisors via WhatsApp crowdsourcing. Supervisors lacked the bandwidth to respond synchronously, leaving frontline health workers without timely answers during active patient visits.
    • Institutional Legitimacy and Co-Branding: System adoption scales according to institutional trust. Services deployed under formal Memorandums of Understanding (MOUs) and co-branded under local county or ministry identities (e.g., PROMPTS) achieved immediate community trust, whereas independent systems relying on social media recruitment (e.g., FarmerChat) or unsupported volunteer labor (e.g., VetBot) faced persistent community skepticism, slow adoption, and onboarding bottlenecks.
  9. Knowl 9 — Twelve Design Guidelines for Culturally Aligned AI for Social Good

    theoretical result

    A set of twelve design guidelines synthesizes lessons across Sociocultural, Institutional, and Technological influences for building and deploying culturally aligned AI systems in non-Western contexts:

    Sociocultural Guidelines:

    • G1 (Full-Lifecycle Equal Partnership): AI developers and local domain experts must collaborate as equal partners across all phases of design, implementation, evaluation, and scaling to resolve cultural nuances and long-tailed linguistic spaces.
    • G2 (Dedicated Low-Resource Linguistic Attention): Foundational language gaps in under-resourced languages require sustained, long-term human and financial investments in dialect vocabularies, translation auditing, and instruction tuning.
    • G3 (Dominant Language Workarounds): In multilingual contexts where specific local dialects are unsupported by base models, developers should build adaptations (e.g., dialect glossaries) on top of closely related or widely spoken regional languages.
    • G4 (Active Trust Calibration): System design must actively calibrate community trust, implementing institutional branding to overcome initial skepticism while using disclaimers and human verification badges to prevent unwarranted over-reliance.

    Institutional Guidelines:

    • G5 (Compliance with Mandates and Workflows): AI tools must integrate seamlessly into existing institutional submission formats, policies, and workplace routines rather than requiring new administrative workflows.
    • G6 (Amplifying Existing Capacity): AI systems should be designed to augment and streamline existing, functional institutional capacities rather than attempting to substitute for missing infrastructure or absent staff.
    • G7 (Planning for Long-Term Maintenance): System roadmaps must budget for continuous organizational commitments, including periodic expert knowledge curation, pipeline monitoring, data refreshes, and scaling inference expenses.
    • G8 (Capacity-Driven Adaptation Staging): Organizations should prioritize lightweight in-context adaptation methods (RAG, prompt engineering, glossaries) and adopt resource-intensive in-weight modifications (pre-training, fine-tuning) only when in-context approaches fail.

    Technological Guidelines:

    • G9 (Modular, Change-Tolerant Architectures): System pipelines must maintain modular decoupling to seamlessly integrate rapidly advancing foundation models without requiring full architectural overhauls.
    • G10 (Small-Scale Risk-Buffering Rollouts): Early, limited deployments must be conducted to expose probabilistic, acoustic, and long-tail contextual failures that cannot be detected in controlled lab settings.
    • G11 (Contextualized, Participatory Evaluation): Generic foundation benchmarks must be augmented with localized evaluation paradigms, such as domain-specific gold-standard transcripts and feedback-driven safety auditing.
    • G12 (Multi-Layered Safety Infrastructure): System reliability in high-stakes domains must combine lightweight pre-input triage models, frontier LLMs for general reasoning, validation agents, and dedicated human expert workflows.
  10. Knowl 10 — Resource-Driven Neglect of Subtle Cultural and Identity Biases in Production Deployments

    limitation

    In production deployments of AI for Social Good in non-Western settings, safety and harm mitigation efforts focus predominantly on acute, overt risks (such as clinical misinformation, self-harm, weapons, or statutory violations like prenatal sex determination) and factual accuracy. Subtle and systemic representational harms—including cultural stereotypes, caste discrimination, gender biases, and religious biases embedded in base LLMs—remain largely unaddressed in practice. Development teams and non-profit organizations face severe resource, time, and infrastructure constraints, forcing them to prioritize basic linguistic functionality, domain ground-truth verification, and latency management over comprehensive socio-cultural debiasing and intersectional fairness auditing.

Coverage note — No substantial contributed material was omitted. All key empirical findings regarding the LISTED framework, the 8 deployed systems, the technical adaptation taxonomy, domain-specific HIL mechanisms, and the 12 design guidelines have been comprehensively captured.

References

  1. 1.Dhruv Agarwal, Mor Naaman, and Aditya Vashistha. 2025. AI Suggestions Homogenize Writing Toward Western Styles and Diminish Cultural Nuances. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems. 1–21. doi:10.1145/3706598.3713564 arXiv:2409.11360 [cs].
  2. 2.Humaid Al Naqbi, Zied Bahroun, and Vian Ahmed. 2024. Enhancing Work Productivity through Generative Artificial Intelligence: A Comprehensive Literature Review. Sustainability 16, 3 (Jan. 2024), 1166. doi:10.3390/su16031166 Number: 3 Publisher: Multidisciplinary Digital Publishing Institute.
  3. 3.Morgan G. Ames. 2019. The Charisma Machine: The Life, Death, and Legacy of One Laptop per Child. MIT Press. Google-Books-ID: yYy5DwAAQBAJ.
  4. 4.Urvashi Aneja, Aarushi Gupta, and Aditya Vashistha. 2025. Beyond Semantics: Examining Gender Bias in LLMs Deployed within Low-resource Contexts in India. (2025).
  5. 5.Oghenemaro Anuyah, Ruyuan Wan, Cornelius Adejoro, Tom Yeh, Ronald Metoyer, and Karla Badillo-Urquiola. 2024. Cultural Considerations in AI Systems for the Global South: A Systematic Review. In Proceedings of the 4th African Human Computer Interaction Conference (AfriCHI ’23). Association for Computing Machinery, New York, NY, USA, 125–134. doi:10.1145/3628096.3629046
  6. 6.Tita Alissa Bach, Amna Khan, Harry Hallock, Gabriela Beltrão, and Sonia Sousa. 2024. A Systematic Literature Review of User Trust in AI-Enabled Systems: An HCI Perspective. International Journal of Human–Computer Interaction 40, 5 (March 2024), 1251–1266. doi:10.1080/10447318.2022.2138826 Publisher: Taylor & Francis _eprint: https://doi.org/10.1080/10447318.2022.2138826.
  7. 7.David Baidoo-anu and Leticia Owusu Ansah. 2023. Education in the Era of Generative Artificial Intelligence (AI): Understanding the Potential Benefits of ChatGPT in Promoting Teaching and Learning. Journal of AI 7, 1 (Dec. 2023), 52–62. doi:10.61969/jai.1337500 Number: 1 Publisher: İzmir Academy Association.
  8. 8.Aaron J. Barnes, Yuanyuan Zhang, and Ana Valenzuela. 2024. AI and culture: Culturally dependent responses to AI systems. Current Opinion in Psychology 58 (Aug. 2024), 101838. doi:10.1016/j.copsyc.2024.101838
  9. 9.Samuel Becher and Benjamin Alarie. 2025. LexOptima: The promise of AI-enabled legal systems. University of Toronto Law Journal 75, 1 (Jan. 2025), 73–121. doi:10.3138/utlj-2024-0002 Publisher: University of Toronto Press.
  10. 10.Emily M. Bender, View Profile, Timnit Gebru, View Profile, Angelina McMillan-Major, View Profile, Shmargaret Shmitchell, and View Profile. 2021. On the Dangers of Stochastic Parrots. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency. 610–623. doi:10.1145/3442188.3445922
  11. 11.Yoshua Bengio, Geoffrey Hinton, Andrew Yao, Dawn Song, Pieter Abbeel, Trevor Darrell, Yuval Noah Harari, Ya-Qin Zhang, Lan Xue, Shai Shalev-Shwartz, Gillian Hadfield, Jeff Clune, Tegan Maharaj, Frank Hutter, Atılım Güneş Baydin, Sheila McIlraith, Qiqi Gao, Ashwin Acharya, David Krueger, Anca Dragan, Philip Torr, Stuart Russell, Daniel Kahneman, Jan Brauner, and Sören Mindermann. 2024. Managing extreme AI risks amid rapid progress. Science 384, 6698 (May 2024), 842–845. doi:10.1126/science.adn0117 arXiv:2310.17688 [cs].
  12. 12.Kirti Bhagat, Kinshuk Vasisht, and Danish Pruthi. 2024. Richer Output for Richer Countries: Uncovering Geographical Disparities in Generated Stories and Travel Recommendations. https://arxiv.org/abs/2411.07320v2
  13. 13.Elizabeth Bondi, Lily Xu, Diana Acosta-Navas, and Jackson A. Killian. 2021. Envisioning Communities: A Participatory Approach Towards AI for Social Good. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society (AIES ’21). Association for Computing Machinery, New York, NY, USA, 425–436. doi:10.1145/3461702.3462612
  14. 14.Yu Ying Chiu, Liwei Jiang, Maria Antoniak, Chan Young Park, Shuyue Stella Li, Mehar Bhatia, Sahithya Ravi, Yulia Tsvetkov, Vered Shwartz, and Yejin Choi. 2024. CulturalTeaming: AI-Assisted Interactive Red-Teaming for Challenging LLMs’ (Lack of) Multicultural Knowledge. doi:10.48550/arXiv.2404.06664 arXiv:2404.06664 [cs].
  15. 15.Jun Ho Choi, Oliver Garrod, Paul Atherton, Andrew Joyce-Gibbons, Miriam Mason-Sesay, and Daniel Björkegren. 2024. Are LLMs Useful in the Poorest Schools? TheTeacher.AI in Sierra Leone. doi:10.48550/arXiv.2310.02982 arXiv:2310.02982.
  16. 16.Yunjae J. Choi, Minha Lee, and Sangsu Lee. 2023. Toward a Multilingual Conversational Agent: Challenges and Expectations of Code-mixing Multilingual Users. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (CHI ’23). Association for Computing Machinery, New York, NY, USA, 1–17. doi:10.1145/3544548.3581445
  17. 17.Eric Corbett, Remi Denton, and Sheena Erete. 2023. Power and Public Participation in AI. In Equity and Access in Algorithms Mechanisms and Optimization. ACM, Boston MA USA, 1–13. doi:10.1145/3617694.3623228
  18. 18.Josh Cowls, Thomas King, Mariarosaria Taddeo, and Luciano Floridi. 2019. Designing AI for Social Good: Seven Essential Factors. SSRN Electronic Journal (2019). doi:10.2139/ssrn.3388669
  19. 19.Diego Moreira Da Rosa, Leandro Soares Guedes, Monica Landoni, and Milene Silveira. 2024. Investigating technology users’ behavior and difficulties in two different multilingual contexts. In Proceedings of the XXIII Brazilian Symposium on Human Factors in Computing Systems. ACM, Brasília DF Brazil, 1–12. doi:10.1145/3702038.3702071
  20. 20.Diego Moreira Da Rosa, Maria Luisa Lamb Souto, and Milene Selbach Silveira. 2022. Designing interfaces for multilingual users: a pattern language. In Proceedings of the 27th European Conference on Pattern Languages of Programs. ACM, Irsee Germany, 1–7. doi:10.1145/3551902.3551964
  21. 21.Qinpu Dang and Guiquan Li. 2025. Unveiling trust in AI: the interplay of antecedents, consequences, and cultural dynamics. AI & SOCIETY (July 2025). doi:10.1007/s00146-025-02477-6
  22. 22.Roberts Dar‘gis, Guntis Barzdi ¯ n,š, Inguna Skadin, a, Normunds Gruz¯ ¯ıtis, and Baiba Saul¯ıte. 2024. Evaluating Open-Source LLMs in Low-Resource Languages: Insights from Latvian High School Exams. In Proceedings of the 4th International Conference on Natural Language Processing for Digital Humanities, Mika Hämäläinen, Emily Öhman, So Miyagawa, Khalid Alnajjar, and Yuri Bizzoni (Eds.). Association for Computational Linguistics, Miami, USA, 289–293. doi:10.18653/v1/2024.nlp4dh-1.28
  23. 23.Fred D. Davis and Andrina Granić. 2024. The Technology Acceptance Model: 30 Years of TAM. Springer International Publishing, Cham. doi:10.1007/978-3-030-45274-2
  24. 24.Fernando Delgado, Stephen Yang, Michael Madaio, and Qian Yang. 2023. The Participatory Turn in AI Design: Theoretical Foundations and the Current State of Practice. doi:10.48550/arXiv.2310.00907 arXiv:2310.00907 [cs].
  25. 25.Deepak Varuvel Dennison, Bakhtawar Ahtisham, Kavyansh Chourasia, Nirmit Arora, Rahul Singh, Rene F. Kizilcec, Akshay Nambi, Tanuja Ganu, and Aditya Vashistha. 2025. Teacher-AI Collaboration for Curating and Customizing Lesson Plans in Low-Resource Schools. doi:10.48550/arXiv.2507.00456 arXiv:2507.00456 [cs].
  26. 26.Lindsey Dewitt Prat, Olivia Nercy Ndlovu Lucas, Christopher Golias, and Mia Lewis. 2024. Decolonizing LLMs: An Ethnographic Framework for AI in African Contexts. Ethnographic Praxis in Industry Conference Proceedings 2024, 1 (2024), 46–85. doi:10.1111/epic.12196 _eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/epic.12196.
  27. 27.Anett Erdmann and Luis Toro-Dupouy. 2025. The influence of the institutional environment on AI adoption in universities: identifying value drivers and necessary conditions. European Journal of Innovation Management (Feb. 2025). doi:10.1108/EJIM-04-2024-0407
  28. 28.Maria Eriksson, Erasmo Purificato, Arman Noroozian, Joao Vinagre, Guillaume Chaslot, Emilia Gomez, and David Fernandez-Llorca. 2025. Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation. doi:10.48550/arXiv.2502.06559 arXiv:2502.06559 [cs].
  29. 29.Vanessa Evers, Agnes Kukulska-Hulme, and Ann Jones. 1999. Cross-Cultural Understanding of Interface Design: A Cross-Cultural Analysis of Icon Recognition, G. V. Prahbu and E. M. del Galdo (Eds.). Rochester, NY, 173–182. http://www.iwips.org/proceedings/1999/papers/cross_cultural_understanding_of_interface_design_a_cross_cultural_analysis_of_icon_recognition/
  30. 30.Ruchao Fan, Yunzheng Zhu, Jinhan Wang, and Abeer Alwan. 2022. Towards Better Domain Adaptation for Self-supervised Models: A Case Study of Child ASR. IEEE Journal of Selected Topics in Signal Processing 16, 6 (Oct. 2022), 1242–1252. doi:10.1109/JSTSP.2022.3200910 arXiv:2305.00115 [eess].
  31. 31.Jennifer Fereday and Eimear Muir-Cochrane. 2006. Demonstrating Rigor Using Thematic Analysis: A Hybrid Approach of Inductive and Deductive Coding and Theme Development. International Journal of Qualitative Methods 5, 1 (March 2006), 80–92. doi:10.1177/160940690600500107 Publisher: SAGE Publications Inc.
  32. 32.Luciano Floridi, Josh Cowls, Thomas C. King, and Mariarosaria Taddeo. 2020. How to Design AI for Social Good: Seven Essential Factors. Science and Engineering Ethics 26, 3 (June 2020), 1771–1796. doi:10.1007/s11948-020-00213-5
  33. 33.Carolina Fuentes, Iyubanit Rodríguez, Gabriela Cajamarca, Laura Cabrera-Quiros, Andrés Lucero, Valeria Herskovic, and Kenton O’Hara. 2024. Opportunities and Challenges of Emerging Human-AI Interactions to Support Healthcare in the Global South. In Companion Publication of the 2024 Conference on Computer-Supported Cooperative Work and Social Computing (CSCW Companion ’24). Association for Computing Machinery, New York, NY, USA, 724–727. doi:10.1145/3678884.3681834
  34. 34.Fiona Fui-Hoon Nah, Zheng , Ruilin, Cai , Jingyuan, Siau , Keng, , and Langtao Chen. 2023. Generative AI and ChatGPT: Applications, challenges, and AI-human collaboration. Journal of Information Technology Case and Application Research 25, 3 (July 2023), 277–304. doi:10.1080/15228053.2023.2233814 Publisher: Routledge _eprint: https://doi.org/10.1080/15228053.2023.2233814.
  35. 35.Isabel O. Gallegos, Ryan A. Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, and Nesreen K. Ahmed. 2024. Bias and Fairness in Large Language Models: A Survey. Computational Linguistics 50, 3 (Sept. 2024), 1097–1179. doi:10.1162/coli_a_00524 Place: Cambridge, MA Publisher: MIT Press.
  36. 36.Marjan Ghazvininejad, Hila Gonen, and Luke Zettlemoyer. 2023. Dictionary-based Phrase-level Prompting of Large Language Models for Machine Translation. doi:10.48550/arXiv.2302.07856 arXiv:2302.07856 [cs].
  37. 37.Amelia Hardy, Anka Reuel, Kiana Jafari Meimandi, Lisa Soder, Allie Griffith, Dylan M Asmar, Sanmi Koyejo, Michael S. Bernstein, and Mykel John Kochenderfer. 2025. More than Marketing? On the Information Value of AI Benchmarks for Practitioners. In Proceedings of the 30th International Conference on Intelligent User Interfaces (IUI ’25). Association for Computing Machinery, New York, NY, USA, 1032–1047. doi:10.1145/3708359.3712152
  38. 38.Gillian R. Hayes. 2020. Inclusive and engaged HCI. interactions 27, 2 (Feb. 2020), 26–31. doi:10.1145/3378561
  39. 39.Rüdiger Heimgärtner. 2018. Culturally-Aware HCI Systems. In Advances in Culturally-Aware Intelligent Systems and in Cross-Cultural Psychological Studies, Colette Faucher (Ed.). Springer International Publishing, Cham, 11–37. doi:10.1007/978-3-319-67024-9_2
  40. 40.Daniel Hershcovich, Stella Frank, Heather Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, Constanza Fierro, Katerina Margatina, Phillip Rust, and Anders Søgaard. 2022. Challenges and Strategies in Cross-Cultural NLP. doi:10.48550/arXiv.2203.10020 arXiv:2203.10020 [cs].
  41. 41.Delene Heukelman and Seraphin Eyono Obono. 2009. Exploring the African Village metaphor for computer user interface icons. In Proceedings of the 2009 Annual Research Conference of the South African Institute of Computer Scientists and Information Technologists (SAICSIT ’09). Association for Computing Machinery, New York, NY, USA, 132–140. doi:10.1145/1632149.1632167
  42. 42.Omar Mohammed Horani, Ahmad Samed Al-Adwan, Husam Yaseen, Hazar Hmoud, Waleed Mugahed Al-Rahmi, and Ali Alkhalifah. 2025. The critical determinants impacting artificial intelligence adoption at the organizational level. Information Development 41, 3 (Sept. 2025), 1055–1079. doi:10.1177/02666669231166889 Publisher: SAGE Publications Ltd.
  43. 43.Lilly Irani, Janet Vertesi, Paul Dourish, Kavita Philip, and Rebecca E. Grinter. 2010. Postcolonial computing: a lens on design and development. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems. ACM, Atlanta Georgia USA, 1311–1320. doi:10.1145/1753326.1753522
  44. 44.Jiaming Ji, Mickel Liu, Juntao Dai, Xuehai Pan, Chi Zhang, Ce Bian, Boyuan Chen, Ruiyang Sun, Yizhou Wang, and Yaodong Yang. 2023. BEAVERTAILS: towards improved safety alignment of llm via a human-preference dataset. In Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS ’23). Curran Associates Inc., Red Hook, NY, USA, 24678–24704.
  45. 45.Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. 2021. The State and Fate of Linguistic Diversity and Inclusion in the NLP World. doi:10.48550/arXiv.2004.09095 arXiv:2004.09095 [cs].
  46. 46.Antonia Karamolegkou, Angana Borah, Eunjung Cho, Sagnik Ray Choudhury, Martina Galletti, Rajarshi Ghosh, Pranav Gupta, Oana Ignat, Priyanka Kargupta, Neema Kotonya, Hemank Lamba, Sun-Joo Lee, Arushi Mangla, Ishani Mondal, Deniz Nazarova, Poli Nemkova, Dina Pisarevskaya, Naquee Rizwan, Nazanin Sabri, Dominik Stammbach, Anna Steinberg, David Tomás, Steven R. Wilson, Bowen Yi, Jessica H. Zhu, Arkaitz Zubiaga, Anders Søgaard, Alexander Fraser, Zhijing Jin, Rada Mihalcea, Joel R. Tetreault, and Daryna Dementieva. 2025. NLP for Social Good: A Survey of Challenges, Opportunities, and Responsible Deployment. doi:10.48550/arXiv.2505.22327 arXiv:2505.22327 [cs] version: 1.
  47. 47.Naveena Karusala, Aditya Vishwanath, Aditya Vashistha, Sunita Kumar, and Neha Kumar. 2018. "Only if you use English you will get to more things": Using Smartphones to Navigate Multilingualism. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems. ACM, Montreal QC Canada, 1–14. doi:10.1145/3173574.3174147
  48. 48.Ji Hye Kim and Kun Pyo Lee. 2005. Cultural difference and mobile phone interface design: icon recognition according to level of abstraction. In Proceedings of the 7th international conference on Human computer interaction with mobile devices & services (MobileHCI ’05). Association for Computing Machinery, New York, NY, USA, 307–310. doi:10.1145/1085777.1085841
  49. 49.Nir Kshetri. 2024. Linguistic Challenges in Generative Artificial Intelligence: Implications for Low-Resource Languages in the Developing World. Journal of Global Information Technology Management 27, 2 (April 2024), 95–99. doi:10.1080/1097198X.2024.2341496 Publisher: Routledge _eprint: https://doi.org/10.1080/1097198X.2024.2341496.
  50. 50.Sreejith Kurup and Vivek Gupta. 2022. Factors Influencing the AI Adoption in Organizations. Metamorphosis 21, 2 (Dec. 2022), 129–139. doi:10.1177/09726225221124035 Publisher: SAGE Publications India.
  51. 51.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-augmented generation for knowledge-intensive NLP tasks. In Advances in neural information processing systems, H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.), Vol. 33. Curran Associates, Inc., 9459–9474. https://proceedings.neurips.cc/paper_files/paper/2020/file/6b493230205f780e1bc26945df7481e5-Paper.pdf
  52. 52.Zihao Li, Yucheng Shi, Zirui Liu, Fan Yang, Ali Payani, Ninghao Liu, and Mengnan Du. 2025. Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages. Proceedings of the AAAI Conference on Artificial Intelligence 39, 27 (April 2025), 28186–28194. doi:10.1609/aaai.v39i27.35038
  53. 53.Weixin Liang, Yaohui Zhang, Mihai Codreanu, Jiayu Wang, Hancheng Cao, and James Zou. 2025. The Widespread Adoption of Large Language Model-Assisted Writing Across Society. doi:10.48550/arXiv.2502.09747 arXiv:2502.09747 [cs] version: 2.
  54. 54.Q. Vera Liao and Ziang Xiao. 2025. Rethinking Model Evaluation as Narrowing the Socio-Technical Gap. doi:10.48550/arXiv.2306.03100 arXiv:2306.03100 [cs].
  55. 55.Chen Cecilia Liu, Iryna Gurevych, and Anna Korhonen. 2025. Culturally Aware and Adapted NLP: A Taxonomy and a Survey of the State of the Art. Transactions of the Association for Computational Linguistics 13 (July 2025), 652–689. doi:10.1162/tacl_a_00760
  56. 56.Nicole Tsz Yeung Liu, Samuel N. Kirshner, and Eric T. K. Lim. 2023. Is algorithm aversion WEIRD? A cross-country comparison of individual-differences and algorithm aversion. Journal of Retailing and Consumer Services 72 (May 2023), 103259. doi:10.1016/j.jretconser.2023.103259
  57. 57.Peter John Loewen, Blake Lee-Whiting, Maggie Arai, Thomas Bergeron, Thomas Galipeau, Isaac Gazendam, Hugh Needham Lee Slinger, and Sofiya Yusypovych. [n. d.]. Global public opinion on artificial intelligence (GPO-AI). Technical Report. Tech. Rep., 2024.[Online]. Available: https://srinstitute. utoronto. ca . . . .
  58. 58.Bilal A. Mateen, Vaishnavi Menon, Ambrose Agweyu, Robert Korom, Elizabeth Omoluabi, David McAfee, Natnael Shimelash, Samuel Rutunda, Crystal Rugege, Gwydion Williams, Mira Emmanuel-Fabula, Alastair K. Denniston, Xiaoxuan Liu, and Melissa Miles. 2025. Trials for LLM-supported clinical decisions in African primary healthcare. Nature Medicine (July 2025), 1–3. doi:10.1038/s41591-025-03815-3 Publisher: Nature Publishing Group.
  59. 59.Josh McGiff and Nikola S. Nikolov. 2025. Overcoming Data Scarcity in Generative Language Modelling for Low-Resource Languages: A Systematic Review. doi:10.48550/arXiv.2505.04531 arXiv:2505.04531 [cs].
  60. 60.Timothy R McIntosh, Teo Susnjak, Nalin Arachchilage, Tong Liu, Dan Xu, Paul Watters, and Malka N Halgamuge. 2025. Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence. IEEE Transactions on Artificial Intelligence (2025), 1–18. doi:10.1109/TAI.2025.3569516
  61. 61.Courtney McKim. 2023. Meaningful Member-Checking: A Structured Approach to Member-Checking. American Journal of Qualitative Research 7, 2 (Feb. 2023), 41–52. https://www.ajqr.org/article/meaningful-member-checking-a-structured-approach-to-member-checking-12973 Publisher: HA Publication.
  62. 62.Indrani Medhi, Aman Sagar, and Kentaro Toyama. 2006. Text-Free User Interfaces for Illiterate and Semi-Literate Users. In 2006 International Conference on Information and Communication Technologies and Development. 72–82. doi:10.1109/ICTD.2006.301841
  63. 63.Shujaat Mirza, Bruno Coelho, Yuyuan Cui, Christina Pöpper, and Damon McCoy. 2024. Global-Liar: Factuality of LLMs over Time and Geographic Regions. doi:10.48550/arXiv.2401.17839 arXiv:2401.17839 [cs].
  64. 64.Eduardo Mosqueira-Rey, Elena Hernández-Pereira, David Alonso-Ríos, José Bobes-Bascarán, and Ángel Fernández-Leal. 2023. Human-in-the-loop machine learning: a state of the art. Artificial Intelligence Review 56, 4 (April 2023), 3005–3054. doi:10.1007/s10462-022-10246-w
  65. 65.Khadijeh Moulaei, Atiye Yadegari, Mahdi Baharestani, Shayan Farzanbakhsh, Babak Sabet, and Mohammad Reza Afrash. 2024. Generative artificial intelligence in healthcare: A scoping review on benefits, challenges and applications. International Journal of Medical Informatics 188 (Aug. 2024), 105474. doi:10.1016/j.ijmedinf.2024.105474
  66. 66.Arthur G. O. Mutambara. 2025. Artificial Intelligence: A Driver of Inclusive Development and Shared Prosperity for the Global South. CRC Press. Google-Books-ID: UElLEQAAQBAJ.
  67. 67.Jessica Ojo, Odunayo Ogundepo, Akintunde Oladipo, Kelechi Ogueji, Jimmy Lin, Pontus Stenetorp, and David Ifeoluwa Adelani. 2025. AfroBench: How Good are Large Language Models on African Languages?. In Findings of the Association for Computational Linguistics: ACL 2025, Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar (Eds.). Association for Computational Linguistics, Vienna, Austria, 19048–19095. doi:10.18653/v1/2025.findings-acl.976
  68. 68.Priyadarsini Patnaik and Mahmoud Bakkar. 2024. Exploring determinants influencing artificial intelligence adoption, reference to diffusion of innovation theory. Technology in Society 79 (Dec. 2024), 102750. doi:10.1016/j.techsoc.2024.102750
  69. 69.Roberto Pereira and Maria Cecília Calani Baranauskas. 2015. A value-oriented and culturally informed approach to the design of interactive systems. International Journal of Human-Computer Studies 80 (Aug. 2015), 66–82. doi:10.1016/j.ijhcs.2015.04.001
  70. 70.Michael A. Peters and Shivali Tukdeo. 2025. Beyond the Utopic/Dyspotic Frames: Towards a Research Agenda for AI in Education and Development (AI4ED) in the Global South. Contemporary Education Dialogue 22, 1 (Jan. 2025), 176–184. doi:10.1177/09731849241298118 Publisher: SAGE Publications India.
  71. 71.Kavita Philip, Lilly Irani, and Paul Dourish. 2012. Postcolonial Computing: A Tactical Survey. Science, Technology, & Human Values 37, 1 (2012), 3–29. https://www.jstor.org/stable/41511154 Publisher: Sage Publications, Inc..
  72. 72.Mahika Phutane, Ananya Seelam, and Aditya Vashistha. 2025. "Cold, Calculated, and Condescending": How AI Identifies and Explains Ableism Compared to Disabled People. doi:10.1145/3715275.3732128 arXiv:2410.03448 [cs].
  73. 73.Vinodkumar Prabhakaran, Rida Qadri, and Ben Hutchinson. 2022. Cultural Incongruencies in Artificial Intelligence. doi:10.48550/arXiv.2211.13069 arXiv:2211.13069 [cs].
  74. 74.Sonel Pyram. 2024. Future directions for context in ICT4D: A systematic literature review. Information Development (May 2024), 02666669241248149. doi:10.1177/02666669241248149 Publisher: SAGE Publications Ltd.
  75. 75.Libo Qin, Qiguang Chen, Yuhang Zhou, Zhi Chen, Yinghui Li, Lizi Liao, Min Li, Wanxiang Che, and Philip S. Yu. 2025. A survey of multilingual large language models. Patterns 6, 1 (Jan. 2025), 101118. doi:10.1016/j.patter.2024.101118
  76. 76.Pragnya Ramjee, Mehak Chhokar, Bhuvan Sachdeva, Mahendra Meena, Hamid Abdullah, Aditya Vashistha, Ruchit Nagar, and Mohit Jain. 2025. ASHABot: An LLM-Powered Chatbot to Support the Informational Needs of Community Health Workers. doi:10.48550/arXiv.2409.10913 arXiv:2409.10913 [cs].
  77. 77.Pragnya Ramjee, Bhuvan Sachdeva, Satvik Golechha, Shreyas Kulkarni, Geeta Fulari, Kaushik Murali, and Mohit Jain. 2025. CataractBot: An LLM-Powered Expert-in-the-Loop Chatbot for Cataract Patients. doi:10.48550/arXiv.2402.04620 arXiv:2402.04620 [cs].
  78. 78.Surangika Ranathunga and Nisansa de Silva. 2022. Some Languages are More Equal than Others: Probing Deeper into the Linguistic Disparity in the NLP World. doi:10.48550/arXiv.2210.08523 arXiv:2210.08523 [cs].
  79. 79.Anka Reuel, Amelia Hardy, Chandler Smith, Max Lamparth, Malcolm Hardy, and Mykel J. Kochenderfer. 2024. BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices. Advances in Neural Information Processing Systems 37 (Dec. 2024), 21763–21813. doi:10.52202/079017-0685
  80. 80.EVERETT M. ROGERS, ARVIND SINGHAL, and MARGARET M. QUINLAN. 2008. Diffusion of Innovations. In An Integrated Approach to Communication Theory and Research (2 ed.). Routledge. Num Pages: 17.
  81. 81.Pranab Sahoo, Ayush Kumar Singh, Sriparna Saha, Vinija Jain, Samrat Mondal, and Aman Chadha. 2025. A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications. doi:10.48550/arXiv.2402.07927 arXiv:2402.07927 [cs].
  82. 82.Luciana Salgado, Roberto Pereira, and Isabela Gasparini. 2015. Cultural Issues in HCI: Challenges and Opportunities. In Human-Computer Interaction: Design and Evaluation, Masaaki Kurosu (Ed.). Springer International Publishing, Cham, 60–70. doi:10.1007/978-3-319-20901-2_6
  83. 83.Ronny Scherer, Fazilat Siddiq, and Jo Tondeur. 2019. The technology acceptance model (TAM): A meta-analytic structural equation modeling approach to explaining teachers’ adoption of digital technology in education. Computers & Education 128 (Jan. 2019), 13–35. doi:10.1016/j.compedu.2018.09.009
  84. 84.Agrima Seth, Monojit Choudhary, Sunayana Sitaram, Kentaro Toyama, Aditya Vashistha, and Kalika Bali. 2018. How Deep Is Representational Bias in LLMs? The Cases of Caste and Religion. (2018).
  85. 85.Haroon Sheikh, Corien Prins, and Erik Schrijvers. 2023. Contextualization. In Mission AI: The New System Technology, Haroon Sheikh, Corien Prins, and Erik Schrijvers (Eds.). Springer International Publishing, Cham, 179–209. doi:10.1007/978-3-031-21448-6_6
  86. 86.Keith Shepherd. [n. d.]. Virtual Agronomist - an AI-assisted chatbot for guiding crop management decisions of smallholder farmers in Africa. ([n. d.]).
  87. 87.Zheyuan Ryan Shi, Claire Wang, and Fei Fang. 2020. Artificial Intelligence for Social Good: A Survey. doi:10.48550/arXiv.2001.01818 arXiv:2001.01818 [cs].
  88. 88.Namita Singh, Jacqueline Wang’ombe, Nereah Okanga, Tetyana Zelenska, Jona Repishti, Jayasankar G. K, Sanjeev Mishra, Rajsekar Manokaran, Vineet Singh, Mohammed Irfan Rafiq, Rikin Gandhi, and Akshay Nambi. 2024. Farmer.Chat: Scaling AI-Powered Agricultural Services for Smallholder Farmers. doi:10.48550/arXiv.2409.08916 arXiv:2409.08916 [cs].
  89. 89.Shivalika Singh, Angelika Romanou, Clémentine Fourrier, David Ifeoluwa Adelani, Jian Gang Ngui, Daniel Vila-Suero, Peerat Limkonchotiwat, Kelly Marchisio, Wei Qi Leong, Yosephine Susanto, Raymond Ng, Shayne Longpre, Sebastian Ruder, Wei-Yin Ko, Antoine Bosselut, Alice Oh, Andre Martins, Leshem Choshen, Daphne Ippolito, Enzo Ferrante, Marzieh Fadaee, Beyza Ermis, and Sara Hooker. 2025. Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar (Eds.). Association for Computational Linguistics, Vienna, Austria, 18761–18799. doi:10.18653/v1/2025.acl-long.919
  90. 90.Katarzyna Stawarz, Dmitri Katz, Amid Ayobi, Paul Marshall, Taku Yamagata, Raul Santos-Rodriguez, Peter Flach, and Aisling Ann O’Kane. 2023. Co-designing opportunities for Human-Centred Machine Learning in supporting Type 1 diabetes decision-making. International Journal of Human-Computer Studies 173 (May 2023), 103003. doi:10.1016/j.ijhcs.2023.103003
  91. 91.David W Stewart and Michael A Kamins. 1993. Secondary research: Information sources and methods. Vol. 4. Sage.
  92. 92.Xuli Tang, Xin Li, Ying Ding, Min Song, and Yi Bu. 2020. The Pace of Artificial Intelligence Innovations: Speed, Talent, and Trial-and-Error. doi:10.48550/arXiv.2009.01812 arXiv:2009.01812 [cs].
  93. 93.Yan Tao, Olga Viberg, Ryan S Baker, and René F Kizilcec. 2024. Cultural bias and cultural alignment of large language models. PNAS Nexus 3, 9 (Sept. 2024), pgae346. doi:10.1093/pnasnexus/pgae346
  94. 94.Kentaro Toyama. 2015. Geek Heresy: Rescuing Social Change from the Cult of Technology. PublicAffairs. Google-Books-ID: 2w6CBgAAQBAJ.
  95. 95.Azmine Toushik Wasi. 2025. AI for Social Good and Public Impact in Marginalized Communities: A Cross-Cultural Framework. In Social Impact of AI: Research, Diversity and Inclusion Frameworks, Yetunde Folajimi, Leonidas Deligiannidis, Salem Othman, and Shawren Singh (Eds.). Springer Nature Switzerland, Cham, 1–12. doi:10.1007/978-3-031-98949-0_1
  96. 96.Robin Williams and David Edge. 1996. The social shaping of technology. Research Policy 25, 6 (Sept. 1996), 865–899. doi:10.1016/0048-7333(96)00885-2
  97. 97.Susan Wyche. 2020. Using Cultural Probes in HCI4D/ICTD: A Design Case Study from Bungoma, Kenya. Proc. ACM Hum.-Comput. Interact. 4, CSCW1 (May 2020), 63:1–63:23. doi:10.1145/3392873
  98. 98.Jinying Xu and Weisheng Lu. 2022. Developing a human-organization-technology fit model for information technology adoption in organizations. Technology in Society 70 (Aug. 2022), 102010. doi:10.1016/j.techsoc.2022.102010
  99. 99.Eric Zhou and Dokyun Lee. 2024. Generative artificial intelligence, human creativity, and art. PNAS Nexus 3, 3 (March 2024), pgae052. doi:10.1093/pnasnexus/pgae052
  100. 100.Mi Zhou, Vibhanshu Abhishek, Timothy Derdenger, Jaymo Kim, and Kannan Srinivasan. 2024. Bias in Generative AI. doi:10.48550/arXiv.2403.02726 arXiv:2403.02726 [econ].

Citation

MLA
Dennison, D. V., et al. “Designing Culturally Aligned AI Systems For Social Good in Non-Western Contexts”. arXiv, 2025, https://doi.org/10.48550/arxiv.2509.16158.
APA
Dennison, D. V., Jain, M., Ganu, T., & Vashistha, A. (2025). Designing Culturally Aligned AI Systems For Social Good in Non-Western Contexts. arXiv. https://doi.org/10.48550/arxiv.2509.16158
Chicago
Dennison, D. V., M. Jain, T. Ganu, and A. Vashistha. 2025. “Designing Culturally Aligned AI Systems For Social Good in Non-Western Contexts”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2509.16158.
Harvard
Dennison, D.V. et al. (2025) “Designing Culturally Aligned AI Systems For Social Good in Non-Western Contexts”. arXiv. Available at: https://doi.org/10.48550/arxiv.2509.16158.
Vancouver
1. Dennison DV, Jain M, Ganu T, Vashistha A (2025) Designing Culturally Aligned AI Systems For Social Good in Non-Western Contexts. https://doi.org/10.48550/arxiv.2509.16158

BibTeX

@misc{https://doi.org/10.48550/arxiv.2509.16158,
  doi = {10.48550/ARXIV.2509.16158},
  url = {https://arxiv.org/abs/2509.16158},
  author = {Dennison, Deepak Varuvel and Jain, Mohit and Ganu, Tanuja and Vashistha, Aditya},
  keywords = {Human-Computer Interaction (cs.HC), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {Designing Culturally Aligned AI Systems For Social Good in Non-Western Contexts},
  publisher = {arXiv},
  year = {2025},
  copyright = {Creative Commons Attribution Share Alike 4.0 International}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by-sa/4.0/