“I Don’t Think RAI Applies to My Model” – Engaging Non-champions with Sticky Stories for Responsible AI Work

Nadia NaharChenyang YangYanxin ChenWesley DengKenneth HolsteinMotahhare EslamiChristian Kastner

article2025International Conference on Human Factors in Computing Systems2 citations

Introduces sticky stories—concrete narratives of unexpected machine learning failures—to overcome practitioner apathy and significantly increase the breadth and depth of ethical risks identified during development.

Listen

Organizations increasingly mandate Responsible AI (RAI) processes, checklists, and templates to prevent algorithmic harms, yet these mechanisms frequently fail in practice. While formal RAI champions and ethics advocates readily adopt these tools, the vast majority of machine learning practitioners—non-champions who lack prior intrinsic motivation or formal ethics roles—routinely treat RAI assessments as superficial, bureaucratic compliance tasks or dismiss them as completely irrelevant to their specific models.

The article aims to evaluate the nature of this practitioner disengagement and demonstrate an effective, theory-informed intervention to foster meaningful, critical engagement with RAI among non-champion practitioners. Specifically, it introduces "sticky stories"—tailored narratives of unexpected, severe, and concrete machine learning harms designed to disrupt practitioners' default assumptions and provoke genuine deliberation during risk assessment workflows.

To address this challenge, the authors followed a multi-stage approach. First, they conducted an ethnographic formative study at a partner technology organization, collecting roughly 22 hours of observation and interviews across governance teams and data scientists. Drawing on psychological theories of transformative learning and cognitive dissonance, they then developed an eight-step compound generative AI pipeline that produces stories tailored to a user's system across five core dimensions: concreteness, severity, surprisingness, diversity, and relevance. They validated the pipeline offline across 240 stories (120 sticky and 120 baseline prompts) using both human annotators and automated model evaluation. Finally, they conducted a controlled user study with 29 active practitioners analyzing their own machine learning projects under three conditions: no stories, baseline generic stories, and sticky stories, accompanied by a two-month post-study follow-up.

The investigation produced several key findings. First, formative observations confirmed that data scientists routinely bypassed RAI evaluations because mainstream media narratives had desensitized them into believing fairness concerns only applied to obvious demographic categories like race and gender. Second, the automated pipeline successfully embodied the intended qualities, outperforming baseline zero-shot prompts by 98% in concreteness, 43% in surprisingness, and 31% in perceived severity, though requiring 5.5 times more processing time and 46 times more computational tokens. Third, in the user study, practitioners presented with sticky stories increased their time spent on harm assessment by roughly 200% (from 5.4 to 16.6 minutes), whereas baseline stories produced only an 11% increase. Fourth, sticky stories prompted participants to identify 4.5 times more distinct harm categories and 3.5 times more subcategories compared to baseline stories. Finally, sticky stories triggered qualitative behavioral markers of critical reflection—such as questioning baseline assumptions, exploring non-obvious stakeholder perspectives, and forming concrete mitigation plans—even among skeptical or previously indifferent practitioners.

These findings indicate that providing structured governance templates alone cannot ensure AI safety if technical practitioners remain unmotivated. When practitioners view RAI as irrelevant, high-stakes operational, reputational, legal, and financial risks remain unnoticed before deployment. Rather than relying on short-term behavioral nudges or generic compliance exercises, organizations can successfully engage reluctant teams by confronting them with highly tailored, surprising scenarios that expose realistic failure modes within their own technical systems.

Based on these results, organizational leaders and AI engineering teams should integrate automated, context-specific narrative prompts directly into standard risk assessment pipelines, model cards, and pre-deployment auditing workflows. AI tools should emphasize diverse edge cases and severe, overlooked impacts rather than exhaustive, generic checklists. Furthermore, leadership should recognize different practitioner profiles—such as active resistors, compliant followers, and indifferent developers—and tailor organizational interventions accordingly to systematically build institutional competence.

Decision-makers should consider several limitations of the findings. The empirical evaluation relied on a modest sample size of 29 practitioners, and the qualitative formative findings were drawn from a single enterprise environment. In addition, the study captured immediate, single-session reflections and early intent; long-term behavioral transformation across daily development cycles over multi-year horizons remains to be demonstrated through broader, longitudinal field deployments.

arXiv: 2509.22858

No sufficiently relevant recommendations were found.

No sufficiently relevant recommendations were found.

Cover for “I Don’t Think RAI Applies to My Model” – Engaging Non-champions with Sticky Stories for Responsible AI Work

Abstract

Responsible AI (RAI) tools -- checklists, templates, and governance processes -- often engage RAI champions, individuals intrinsically motivated to advocate ethical practices, but fail to reach non-champions, who frequently dismiss them as bureaucratic tasks. To explore this gap, we shadowed meetings and interviewed data scientists at an organization, finding that practitioners perceived RAI as irrelevant to their work. Building on these insights and theoretical foundations, we derived design principles for engaging non-champions, and introduced sticky stories -- narratives of unexpected ML harms designed to be concrete, severe, surprising, diverse, and relevant, unlike widely circulated media to which practitioners are desensitized. Using a compound AI system, we generated and evaluated sticky stories through human and LLM assessments at scale, confirming they embodied the intended qualities. In a study with 29 practitioners, we found that, compared to regular stories, sticky stories significantly increased time spent on harm identification, broadened the range of harms recognized, and fostered deeper reflection.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 State of Responsible AI in Industry
  • 2.2 Supports for RAI Harm Identification
  • 2.3 Engaging Non-Champions
  • 3 Phase I: Formative Study
  • 3.1 Research Method
  • 3.2 Key Finding: Despite Substantial Governance Efforts, RAI Had Minimal Impact, with Practitioners Largely Ignoring or Dismissing RAI Concerns
  • 4 Phase II: Sticky Story Intervention Design
  • 4.1 Theoretical Background and Design Principles
  • 4.2 Generating Sticky Stories
  • 4.3 Sticky Story Integration in a Tool
  • 5 Evaluation I: Evaluating Stickiness of Harm Stories
  • 5.1 Experiment Setups
  • 5.1.1 Data
  • 5.1.2 Methods
  • 5.1.3 Metrics
  • 5.1.4 Limitations (Threats to Validity)
  • 5.2 Results
  • 6 Evaluation II: User Study
  • 6.1 Study Design
  • 6.1.1 Participants and Recruitment
  • 6.1.2 Experimental Conditions
  • 6.1.3 Tasks
  • 6.1.4 Treatment Groups
  • 6.1.5 Study Protocol
  • 6.2 Data Analysis
  • 6.3 Limitations (Threats to Validity/Credibility)
  • 6.4 Findings
  • 6.4.1 Finding 1: Practitioners spent significantly more time on harm identification tasks when sticky stories were shown
  • 6.4.2 Finding 2: Practitioners identified significantly more diverse harms when sticky stories were shown
  • 6.4.3 Finding 3: Practitioners critically reflected on the sticky stories more, but only skimmed the baseline stories
  • 6.4.4 Finding 4: Follow-up survey revealed more post-study actions than the sticky story group
  • 6.4.5 Finding 5: Practitioners exhibit distinct trajectories in shifting from initial indifference or resistance toward a more engaged stance on RAI
  • 7 Discussion
  • References

Knowls

  1. Knowl 1 — Sticky stories are designed to make RAI harms personally consequential

    model/method

    Sticky stories are narratives of unexpected harms that could arise from a practitioner’s own machine-learning system. They are intended to engage non-champions—practitioners who are indifferent or skeptical about Responsible AI (RAI)—by disrupting assumptions that RAI risks are irrelevant to their work and prompting critical reflection. The design draws on transformative learning and cognitive dissonance: an unfamiliar, consequential scenario may challenge existing beliefs, while reflection gives practitioners an opportunity to reconsider them. The stories are designed around five qualities: concreteness (tangible details and observable consequences), severity (substantial magnitude or scope of harm), surprisingness (non-obvious outcomes that challenge expectations), relevance (connection to the practitioner’s system and work context), and diversity (varied stakeholders, harms, and system behaviors). The intended near-term outcome is deeper engagement and reflection; a single exposure is not claimed to establish lasting transformation.

  2. Knowl 2 — A compound AI pipeline generates and selects sticky stories

    algorithm

    The pipeline takes an ML system description, its intended purpose, and a representative stakeholder use case as input, and produces two high-severity harm stories intended to satisfy the five sticky-story qualities. It uses GPT-4o for most generation, GPT-4o-mini to evaluate refinements, and mxbai-embed-large-v1 for sentence embeddings.

    Input: ML system description, intended purpose, representative stakeholder use case
    Output: Two selected harm stories
    1. Predefine harm types, including cultural misrepresentation, reinforcement of biases,
       unequal access to opportunities, and erasure of minorities.
    2. Prompt an LLM to identify direct, indirect, direct-surprising, and
       indirect-surprising stakeholders for the system.
    3. Generate potentially relevant demographic attributes for each stakeholder.
    4. For each harm type–stakeholder combination, generate initial stories.
    5. Regenerate stories for surprisingness, using earlier outputs as counterexamples
       to discourage typical or generic narratives.
    6. Embed the stories, cluster them with K-means using k = 10, and sample from the
       five least-populated clusters to favor less redundant stories.
    7. Refine sampled stories for concreteness and severity; use a second model to
       evaluate them and revise insufficiently specific or clear stories, up to three times.
    8. Ask an LLM to select the two stories with the greatest severity while requiring
       that the selected stories satisfy all five intended qualities.
    9. Return the selected stories.

    The predefined harm categories are grounded in existing harm and RAI assessment taxonomies. Identifying surprising stakeholder groups is intended to reveal people who might otherwise be overlooked; the combination matrix broadens the harm scenarios considered. The clustering step favors diversity, while the refinement and final selection steps prioritize concrete, high-impact stories.

  3. Knowl 3 — Sticky stories substantially increased time spent identifying harms

    empirical result

    In a study of 29 practitioners, each person first completed a harm-identification task without stories and then completed a second task with either baseline stories or sticky stories. Participants in the baseline-story group spent a mean of 7.5 minutes on the no-story task and 8.3 minutes on the baseline-story task, an increase of 11%. Participants in the sticky-story group spent a mean of 5.4 minutes on the no-story task and 16.6 minutes on the sticky-story task, an increase of 207% relative to their own first task. The time values are reported as mean ± standard deviation: 7.5 ± 4.3 and 8.3 ± 3.6 minutes for the baseline group; 5.4 ± 2.7 and 16.6 ± 7.4 minutes for the sticky-story group.

    The story condition predicted the relative increase in time, controlling for task/fairness-goal order, RAI awareness, championship, and prior AI/ML experience (F=31.37F=31.37, p<0.001p<0.001). For second-task time, an ANCOVA controlling for first-task time and the same participant and task factors found a story-condition effect of F=28.47F=28.47, p<0.001p<0.001. Because every participant completed the no-story task first, the comparison remains subject to possible order and carryover effects.

  4. Knowl 4 — Sticky stories broadened the categories of harms practitioners identified

    empirical result

    Practitioners exposed to sticky stories identified more distinct harm categories and subcategories than practitioners exposed to baseline stories; the paper reports approximately 4.5 times as many new categories and 3.5 times as many new subcategories. These figures concern additional distinct harms on the second task relative to the first, not simply the total number of harms listed. In the baseline-story group, the second-task additional counts averaged 0.31 ± 0.48 categories and 0.54 ± 0.66 subcategories; in the sticky-story group, they averaged 1.38 ± 0.62 categories and 1.88 ± 0.62 subcategories. The story-condition effects on second-task category and subcategory counts were statistically significant after adjustment: categories, F=19.43F=19.43, p=0.0002p=0.0002; subcategories, F=22.39F=22.39, p=0.0001p=0.0001. The analyses accounted for task order, RAI awareness, championship, and prior AI/ML experience. The authors emphasize category diversity rather than raw harm counts because multiple listed harms can be similar.

  5. Knowl 5 — Sticky stories elicited more critical reflection than baseline stories

    empirical result

    Think-aloud transcripts showed more signs of critical reflection among practitioners who saw sticky stories than among those who saw baseline stories. Surprise or enthusiasm was the most frequently observed reflection indicator: it appeared in 17 participants overall, 15 of whom were in the sticky-story condition. Connecting a story to wider systems or past incidents was observed in nine sticky-story participants and four baseline-story participants. Challenging assumptions, exploring multiple perspectives, and iterative reconsideration appeared only in the sticky-story condition. By contrast, the coded shallow-engagement behaviors—such as skimming, dismissing a story, or copying it as a harm without further reasoning—appeared only in the baseline condition. These indicators operationalized critical reflection as examining assumptions, considering perspectives and wider implications, revising one’s reasoning, and contemplating intentional change.

    Two months later, nine of the 25 participants contacted for follow-up had responded: six from the sticky-story group and three from the baseline group. Sticky-story respondents described concrete changes or discussions, including more diverse data recruitment, expert review and stricter annotation checks, safety checks, and discussions with colleagues about RAI. Baseline-group responses were generally less action-oriented. The small, self-selected follow-up sample is descriptive evidence only and does not establish a comparative or lasting effect.

  6. Knowl 6 — The generation pipeline produced stories with stronger measured stickiness, at higher cost

    empirical result

    An offline evaluation compared pipeline-generated stories with zero-shot GPT-4o baseline stories across 15 AI application scenarios. For each scenario and each of two fairness goals—quality of service and allocation of resources and opportunities—the researchers generated two stories per method, yielding 120 stories per method. Four qualities were rated as binary judgments, scaled here as proportions: severity was 0.992 for pipeline stories versus 0.683 for baseline stories; surprisingness, 0.783 versus 0.354; concreteness, 1.000 versus 0.017; and relevance, 0.892 versus 0.979. Diversity was measured as cosine distance between embeddings of story titles and was 0.156 for pipeline stories versus 0.098 for baseline stories. Thus pipeline stories scored higher on severity, surprisingness, concreteness, and measured diversity, while baseline stories scored higher on relevance. The authors note that some pipeline stories were overly dramatic and consequently harder to relate to.

    The pipeline used 56,665 tokens and took 50 seconds on average, compared with 1,232 tokens and 9 seconds for the baseline. Human–LLM agreement was high for concreteness (Cohen’s κ=1.000\kappa=1.000), surprisingness (0.93560.9356), and relevance (0.83870.8387), but lower for severity (0.47830.4783); the severity disagreements arose because the LLM judge rated baseline stories as severe too generously. The small scenario set, binary quality judgments, and potential bias in LLM ratings constrain interpretation.

  7. Knowl 7 — Practitioners showed five distinct RAI engagement profiles

    empirical result

    The authors describe five qualitative profiles based on participants’ prior orientation and observed responses: resistors explicitly treated RAI as hype or irrelevant; indifferents acknowledged RAI’s importance but showed little prior follow-through; followers participated mainly because their organizations required RAI processes; learners had limited RAI knowledge; and champions were already intrinsically motivated advocates. The observed counts were two resistors, six indifferents, seven followers, nine learners, and five champions.

    The authors report these as descriptive trajectories, not statistically validated participant types. In the sticky-story condition, resistors moved from dismissal toward skepticism, recognition of overlooked risks, and reframing some issues as RAI-relevant. Indifferents showed a path from curiosity to reflection and concrete plans; followers sometimes moved from compliance toward greater ownership; and learners sometimes recognized knowledge gaps or broadened their perspectives. Champions were already engaged, and the study did not show substantial differences in their harm identification across story conditions. The authors conjecture that different qualities may especially resonate with different non-champion profiles—for example, surprisingness with indifferents and diversity with learners—but do not establish these as causal, profile-specific effects.

  8. Knowl 8 — The user study compared no-story, baseline-story, and sticky-story conditions on participants’ own projects

    experimental setup

    Twenty-nine practitioners completed two harm-identification and mitigation tasks based on their own ML projects. The tasks addressed two fairness goals from an RAI assessment guide: allocation of resources and opportunities, and minimization of stereotyping, demeaning, and erasing outputs. Every participant completed one task without stories first. For the second task, participants were assigned either two zero-shot baseline harm stories or two sticky stories. The fairness-goal order was counterbalanced, allowing a within-participant comparison of the first, no-story task with the second task and a between-group comparison of baseline versus sticky stories. The stories were presented through an interactive tool adapting an RAI assessment template; stories could be generated in advance because generation typically took 3–4 minutes. The study measured task time, harm counts and taxonomic diversity, think-aloud evidence of reflection, and self-reported follow-up actions. The design included non-champions with varied RAI motivation, rather than recruiting only participants already motivated to engage with RAI.

  9. Knowl 9 — A formative study identified an engagement gap despite organizational RAI support

    empirical result

    In a partner technology organization seeking organization-wide RAI adoption, the researchers observed three governance-team meetings, eight project-team meetings, and four one-on-one working sessions, and conducted informal conversations and five semi-structured interviews—two with governance champions and three with data scientists. The work totalled roughly 22 hours of observation. The researchers found that governance champions actively developed and supported RAI processes, but many data scientists viewed RAI as abstract or peripheral and were unsure why it mattered to their projects. RAI topics were largely absent from project meetings or subordinated to client deadlines; practitioners skipped assessment-template sections unrelated to their immediate tasks, used templates only when mandated, or completed them retrospectively to demonstrate compliance. Even active coaching and organizational support often resulted in minimal, check-the-box participation. Some practitioners also dismissed concerns when they did not match familiar public narratives about race or gender; the authors present the role of such narratives as a conjecture. The findings motivated the focus on interventions that make harms feel relevant rather than simply supplying more guidance. This interpretive qualitative study concerns a single organizational setting and relied on contemporaneous notes rather than recordings.

  10. Knowl 10 — Evidence supports short-term engagement, not durable transformation

    limitation

    The user study had 29 participants and measured a single researcher-facilitated exposure, so its results do not establish sustained changes in workplace behavior or generalize confidently across organizations and domains. All participants completed the no-story task first, confounding story exposure with appearing second and leaving learning, priming, fatigue, and carryover as possible influences. Think-aloud procedures may also alter participants’ reasoning, and time on task can reflect confusion as well as engagement. Although the authors sent a two-month follow-up, only nine participants responded, which is insufficient to infer durable or comparative effects. The offline story evaluation used only 15 curated application scenarios and binary ratings; LLM-as-judge biases were especially evident for severity, where the judge was overly generous to baseline stories. The authors therefore treat the observed reflection and reported actions as early signs, not proof that sticky stories produce long-term transformation.

Coverage note — No substantial contributed material was omitted. Screen-level details of the assessment tool and individual story vignettes were left out because they illustrate the implementation rather than add independent findings.

References

  1. 1.Baker-Brunnbauer, J. 2021. Management perspective of ethics in artificial intelligence. AI and ethics. 1, 2 (2021), 173–181.
  2. 2.Balayn, A., Yurrita, M., Yang, J. and Gadiraju, U. 2023. “Fairness toolkits, A checkbox culture?” on the factors that fragment developer practices in handling algorithmic harms. Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society (2023), 482–495.
  3. 3.Ballard, S., Chappell, K.M. and Kennedy, K. 2019. Judgment call the game: Using value sensitive design and design fiction to surface ethical concerns related to technology. Proceedings of the 2019 on Designing Interactive Systems Conference (2019), 421–433.
  4. 4.Basili, V., Heidrich, J., Lindvall, M., Münch, J., Regardie, M., Rombach, D., Seaman, C. and Trendowicz, A. 2014. GQM+strategies: A comprehensive methodology for aligning business strategies with software measurement. arXiv [cs.SE].
  5. 5.Bessen, J., Impink, S.M. and Seamans, R. 2022. The cost of ethical AI development for AI startups. Proceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society (2022), 92–106.
  6. 6.Bhat, A., Coursey, A., Hu, G., Li, S., Nahar, N., Zhou, S., Kästner, C. and Guo, J.L.C. 2023. Aspirations and Practice of ML Model Documentation: Moving the Needle with Nudging and Traceability. Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (2023), 1–17.
  7. 7.Bogucka, E., Constantinides, M., Šćepanović, S. and Quercia, D. 2024. Co-designing an AI impact assessment report template with AI practitioners and AI compliance experts. arXiv [cs.HC].
  8. 8.Boren, T. and Ramey, J. 2000. Thinking aloud: reconciling theory and practice. IEEE transactions on professional communication. 43, 3 (2000), 261–278.
  9. 9.Boyd, K. 2022. Designing up with value-sensitive design: Building a field guide for ethical ML development. Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency (2022), 2069–2082.
  10. 10.Brookfield, S.D. 2017. Becoming a critically reflective teacher. Jossey-Bass.
  11. 11.Buçinca, Z., Pham, C.M., Jakesch, M., Ribeiro, M.T., Olteanu, A. and Amershi, S. 2023. AHA!: Facilitating AI Impact Assessment by Generating Examples of Harms. arXiv [cs.HC].
  12. 12.Chang, J. and Custis, C. 2022. Understanding implementation challenges in machine learning documentation. Proceedings of the 2nd ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization (2022), 1–8.
  13. 13.Collision Between Vehicle Controlled by Developmental Automated Driving System and Pedestrian: https://www.ntsb.gov/investigations/ AccidentReports/ Reports/HAR1903.pdf .
  14. 14.Costanza-Chock, S. 2020. Design Justice: Community-Led Practices to Build the Worlds We Need. MIT Press.
  15. 15.Dastin, J. 2022. Amazon Scraps Secret AI Recruiting Tool that Showed Bias against Women. Ethics of Data and Analytics. Auerbach Publications. 296–299.
  16. 16.Deng, W.H., Barocas, S. and Wortman Vaughan, J. 2025. Supporting industry computing researchers in assessing, articulating, and addressing the potential negative societal impact of their work. Proceedings of the ACM on human-computer interaction. 9, 2 (2025), 1–37.
  17. 17.Deng, W.H., Guo, B.B., Devos, A., Shen, H., Eslami, M. and Holstein, K. 2023. Understanding Practices, Challenges, and Opportunities for User-Engaged Algorithm Auditing in Industry Practice. Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (2023), 1–18.
  18. 18.Deng, W.H., Yildirim, N., Chang, M., Eslami, M., Holstein, K. and Madaio, M. 2023. Investigating practices and opportunities for cross-functional collaboration around AI fairness in industry practice. Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency. (2023), 705–716.
  19. 19.Dominique, B., Maghraoui, K.E., Piorkowski, D. and Herger, L. 2023. FactSheets for hardware-aware AI models: A case study of analog in memory computing AI models. Proceedings of the 2023 IEEE International Conference on Software Services Engineering (SSE) (2023), 148–158.
  20. 20.Dual-coding theory: 2025. https:// en.wikipedia.org/wiki/Dual-coding_theory.
  21. 21.Ehsan, U., Liao, Q.V., Passi, S., Riedl, M.O. and Daumé, H., III 2024. Seamful XAI: Operationalizing seamful design in Explainable AI. Proceedings of the ACM on human-computer interaction. 8, CSCW1 (2024), 1–29.
  22. 22.Elsayed-Ali, S., Berger, S. E., Santana, V. F. D., & Becerra Sandoval, J. C. 2023. Responsible & inclusive cards: An online card tool to promote critical reflection in technology industry work practices. Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (2023), 1–14.
  23. 23.Ericsson, K.A. and Simon, H.A. 1993. Protocol Analysis. MIT Press.
  24. 24.Festinger, L. 1957. A Theory of Cognitive Dissonance. Stanford University Press.
  25. 25.Four stages of competence: 2025. https:// en.wikipedia.org/wiki/Four_stages_of_competence.
  26. 26.Fredricks, J.A., Blumenfeld, P.C. and Paris, A.H. 2004. School engagement: Potential of the concept, state of the evidence. Review of educational research. 74, 1 (2004), 59–109.
  27. 27.Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J.W., Wallach, H., Iii, H.D. and Crawford, K. 2021. Datasheets for datasets. Communications of the ACM. 64, 12 (2021), 86–92.
  28. 28.Gravett, S. 2002. Transformative Learning through Action Research: A Case Study from South Africa. Adult Education Research Conference (2002).
  29. 29.Green, M.C. and Appel, M. 2024. Narrative transportation: How stories shape how we see ourselves and the world. Advances in Experimental Social Psychology. Elsevier. 1–82.
  30. 30.Hadi Mogavi, R., Guo, B., Zhang, Y., Haq, E.-U., Hui, P. and Ma, X. 2022. When gamification spoils your learning: A qualitative case study of gamification misuse in a language-learning app. Proceedings of the Ninth ACM Conference on Learning @ Scale (2022).
  31. 31.Heath, D. and Heath, C. 2009. Made to Stick. Random House Trade.
  32. 32.Holstein, K., Vaughan, J.W., Daumé, H., III, Dudík, M. and Wallach, H. 2019. Improving fairness in machine learning systems: What do industry practitioners need? Proceedings of the 2019 CHI conference on human factors in computing systems (2019), 1–16.
  33. 33.Hsieh, H.-F. and Shannon, S.E. 2005. Three approaches to qualitative content analysis. Qualitative health research. 15, 9 (2005), 1277–1288.
  34. 34.Hummel, D. and Maedche, A. 2019. How effective is nudging? A quantitative review on the effect sizes and limits of empirical nudging studies. Journal of behavioral and experimental economics. 80, (Jun. 2019), 47–58.
  35. 35.Itti, L. and Baldi, P. 2009. Bayesian surprise attracts human attention. Vision research. 49, 10 (2009), 1295–1306.
  36. 36.Kallina, E., Bohné, T. and Singh, J. 2025. Stakeholder participation for responsible AI development: Disconnects between guidance and current practice. Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency (2025), 1060–1079.
  37. 37.Kaur, H., Conrad, M.R., Rule, D., Lampe, C. and Gilbert, E. 2024. Interpretability gone bad: The role of bounded rationality in how practitioners understand machine learning. Proceedings of the ACM on human-computer interaction. 8, CSCW1 (2024), 1–34.
  38. 38.Kihlstrom, J.F. 2021. Ecological validity and “ecological validity.” Perspectives on psychological science: a journal of the Association for Psychological Science. 16, 2 (2021), 466–471.
  39. 39.Kim, S.-E., Kim, K., Lee, J., Ko, Y., Wang, Y. and So, H.-J. 2025. Dilemmas in AI ethics: A digital game for moral reasoning and collective decision-making. Proceedings of the International Conference on Artificial Intelligence in Education (Cham, 2025), 434–447.
  40. 40.Lanne, M., Nieminen, M. and Leikas, J. 2025. Organisational tensions in introducing socially sustainable AI. AI & society. (2025). DOI:https://doi.org/10.1007/s00146-025-02293-y.
  41. 41.Lazar, J., Feng, J.H. and Hochheiser, H. 2017. Research Methods in Human-Computer Interaction. Morgan Kaufmann.
  42. 42.Leveson, N.G. 2016. Engineering a safer world: Systems thinking applied to safety. The MIT Press.
  43. 43.Liao, Q.V., Subramonyam, H., Wang, J. and Wortman Vaughan, J. 2023. Designerly understanding: Information needs for model transparency to support design ideation for AI-powered user experience. Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (2023), 1–21.
  44. 44.Louviere, J.J., Hensher, D.A. and Swait, J.D. 2014. Stated choice methods. Cambridge University Press.
  45. 45.Madaio, M.A., Stark, L., Wortman Vaughan, J. and Wallach, H. 2020. Co-Designing Checklists to Understand Organizational Challenges and Opportunities around Fairness in AI. Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems (2020), 1–14.
  46. 46.Madaio, M., Egede, L., Subramonyam, H., Wortman Vaughan, J. and Wallach, H. 2022. Assessing the fairness of AI systems: AI practitioners’ processes, challenges, and needs for support. Proceedings of the ACM on human-computer interaction. 6, CSCW1 (2022), 1–26.
  47. 47.Madaio, M., Kapania, S., Qadri, R., Wang, D., Zaldivar, A., Denton, R. and Wilcox, L. 2024. Learning about responsible AI on-the-job: Learning pathways, orientations, and aspirations. Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency (2024), 1544–1558.
  48. 48.Mann, K., Gordon, J. and MacLeod, A. 2009. Reflection and reflective practice in health professions education: a systematic review. Advances in health sciences education: theory and practice. 14, 4 (2009), 595–621.
  49. 49.Martelaro, N. and Ju, W. 2020. What could go wrong? Exploring the downsides of autonomous vehicles. Proceedings of the 12th International Conference on Automotive User Interfaces and Interactive Vehicular Applications (2020).
  50. 50.Massaro, D.W., Petty, R.E. and Cacioppo, J.T. 1988. Communication and Persuasion: Central and Peripheral Routes to Attitude Change. The American journal of psychology. 101, 1 (1988), 155.
  51. 51.McAdams, D.P. 2011. Narrative Identity. Handbook of Identity Theory and Research. Springer New York. 99–115.
  52. 52.Mezirow, J. 2000. Learning as transformation: Critical perspectives on a theory in progress. Jossey-Bass.
  53. 53.Mezirow, J. 1991. Transformative dimensions of adult learning. Jossey-Bass.
  54. 54.Mezirow, J. 2018. Transformative learning theory. Contemporary Theories of Learning. Routledge. 114–128.
  55. 55.Microsoft RAI Impact Assessment Template: https://blogs.microsoft.com/wp-content/ uploads/prod/ sites/ 5/ 2022/ 06/Microsoft-RAI-Impact-Assessment-Template.pdf .
  56. 56.Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., Vasserman, L., Hutchinson, B., Spitzer, E., Raji, I.D. and Gebru, T. 2019. Model cards for model reporting. Proceedings of the Conference on Fairness, Accountability, and Transparency (2019), 220–229.
  57. 57.ml-practical-usecases: A database of 450 Machine Learning (ML) system design case studies from 100+ companies: https://github.com/mallahyari/ml-practical-usecases. Accessed: 2025-09-09.
  58. 58.Nahar, N., Zhang, H., Lewis, G., Zhou, S. and Kästner, C. 2023. A Meta-Summary of Challenges in Building Products with ML Components – Collecting Experiences from 4758+ Practitioners. Proceedings of the IEEE/ACM 2nd International Conference on AI Engineering – Software Engineering for AI (CAIN) (2023), 171–183.
  59. 59.Nahar, N., Zhou, S., Lewis, G. and Kästner, C. 2022. Collaboration challenges in building ML-enabled systems. Proceedings of the 44th International Conference on Software Engineering (2022).
  60. 60.Nathan, L.P., Friedman, B., Klasnja, P., Kane, S.K. and Miller, J.K. 2008. Envisioning systemic effects on persons and society throughout interactive system design. Proceedings of the 7th ACM conference on Designing interactive systems (2008), 1–10.
  61. 61.Omar, Z.A., Nahar, N., Tjaden, J., Gilles, I.M., Mekonnen, F., Hsieh, J., Kästner, C. and Menon, A. 2025. Beyond Accuracy, SHAP, and Anchors – On the difficulty of designing effective end-user explanations. arXiv [cs.HC].
  62. 62.O’Neil, C. 2016. Weapons of Math Destruction: How Big Data Increases Inequality and Threatens Democracy. Crown.
  63. 63.OpenAI et al. 2023. GPT-4 Technical Report. arXiv [cs.CL].
  64. 64.Pang, R.Y., Santy, S., Just, R. and Reinecke, K. 2024. BLIP: Facilitating the exploration of undesirable consequences of digital technologies. Proceedings of the CHI Conference on Human Factors in Computing Systems (2024), 1–18.
  65. 65.Passi, S. and Barocas, S. 2019. Problem Formulation and Fairness. Proceedings of the Conference on Fairness, Accountability, and Transparency (New York, NY, USA, Jan. 2019), 39–48.
  66. 66.Paunesku, D., Walton, G.M., Romero, C., Smith, E.N., Yeager, D.S. and Dweck, C.S. 2015. Mind-set interventions are a scalable treatment for academic underachievement. Psychological science. 26, 6 (2015), 784–793.
  67. 67.Polman, E. and Maglio, S.J. 2024. The Problem With Behavioral Nudges. The Wall Street Journal. The Wall Street Journal.
  68. 68.Popular Machine Learning Applications and Use Cases in our Daily Life: 2019. https://www.analyticsvidhya.com/blog/ 2019/ 07/ ultimate-list-popular-machine-learning-use-cases/. Accessed: 2025-09-09.
  69. 69.ProjectPro, B.Y. 2021. 15 Machine Learning Use Cases and Applications in 2025. ProjectPro.
  70. 70.Raji, I.D., Smart, A., White, R.N., Mitchell, M., Gebru, T., Hutchinson, B., Smith-Loud, J., Theron, D. and Barnes, P. 2020. Closing the AI accountability gap: Defining an end-to-end framework for internal algorithmic auditing. Proceedings of the 2020 conference on fairness, accountability, and transparency (2020), 33–44.
  71. 71.Rakova, B., Yang, J., Cramer, H. and Chowdhury, R. 2021. Where Responsible AI meets Reality: Practitioner Perspectives on Enablers for shifting Organizational Practices. Proceedings of the ACM on Human-Computer Interaction. 5, CSCW1 (2021), 1–13.
  72. 72.Reimers, N. and Gurevych, I. 2019. Sentence-BERT: Sentence embeddings using Siamese BERT-networks. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) (2019).
  73. 73.Responsible AI: https://www.ibm.com/ trust/ responsible-ai. Accessed: 2025-09-11.
  74. 74.Responsible AI: Ethical policies and practices: https://www.microsoft.com/ en-us/ai/ responsible-ai. Accessed: 2025-09-11.
  75. 75.Responsible AI Maturity Model: 2023. https://www.microsoft.com/ en-us/ research/publication/ responsible-ai-maturity-model/. Accessed: 2025-09-11.
  76. 76.Responsible AI Transparency Report: https:// cdn-dynmedia-1.microsoft.com/is/ content/microsoftcorp/microsoft/msc/documents/presentations/CSR/ Responsible-AI-Transparency-Report-2024.pdf . Accessed: 2025-09-11.
  77. 77.Richardson, B., Garcia-Gathright, J., Way, S.F., Thom, J. and Cramer, H. 2021. Towards fairness in practice: A practitioner-oriented rubric for evaluating fair ML toolkits. Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (2021), 1–13.
  78. 78.Rismani, S. and Moon, A. 2023. What does it mean to be a responsible AI practitioner: An ontology of roles and skills. Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society (2023), 584–595.
  79. 79.Rismani, S., Shelby, R., Smart, A., Jatho, E., Kroll, J., Moon, A. and Rostamzadeh, N. 2023. From Plane Crashes to Algorithmic Harm: Applicability of Safety Engineering Frameworks for Responsible ML. Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (2023), 1–18.
  80. 80.Rodrigues, L., Pereira, F.D., Toda, A.M., Palomino, P.T., Pessoa, M., Carvalho, L.S.G., Fernandes, D., Oliveira, E.H.T., Cristea, A.I. and Isotani, S. 2022. Gamification suffers from the novelty effect but benefits from the familiarization effect: Findings from a longitudinal study. International journal of educational technology in higher education. 19, 1 (2022), 13.
  81. 81.Ryan, C.L., Cant, R., McAllister, M.M., Vanderburg, R. and Batty, C. 2022. Transformative learning theory applications in health professional and nursing education: An umbrella review. Nurse education today. 119, 105604 (Dec. 2022), 105604.
  82. 82.Sambasivan, N., Kapania, S., Highfill, H., Akrong, D., Paritosh, P. and Aroyo, L.M. 2021. “Everyone wants to do the model work, not the data work”: Data Cascades in High-Stakes AI. Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (2021), 1–15.
  83. 83.Schoen, D.A. 2017. The reflective practitioner: How professionals think in action. Routledge.
  84. 84.Shelby, R., Rismani, S., Henne, K., Moon, A., Rostamzadeh, N., Nicholas, P., Yilla-Akbari, N. ’mah, Gallegos, J., Smart, A., Garcia, E. and Virk, G. 2023. Sociotechnical harms of algorithmic systems: Scoping a taxonomy for harm reduction. Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society (2023), 723–741.
  85. 85.Slovic, P., Finucane, M.L., Peters, E. and MacGregor, D.G. 2007. The affect heuristic. European journal of operational research. 177, 3 (2007), 1333–1352.
  86. 86.Smith, J.J., Madaio, M., Burke, R. and Fiesler, C. 2025. Pragmatic fairness: Evaluating ML fairness within the constraints of industry. Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency (2025), 628–638.
  87. 87.The board game: 2025. https:// tethics.eu/ the-board-game/. Accessed: 2025-09-09.
  88. 88.The building blocks of Microsoft’s responsible AI program: 2021. https://blogs.microsoft.com/on-the-issues/2021/01/19/microsoft-responsible-ai-program/. Accessed: 2025-08-04.
  89. 89.The Ethical Dilemmas Board Game: https:// cfrr.worldbank.org/publications/ ethical-dilemmas-board-game. Accessed: 2025-09-09.
  90. 90.The Shift from Models to Compound AI Systems: http://bair.berkeley.edu/blog/ 2024/ 02/ 18/ compound-ai-systems/. Accessed: 2025-08-27.
  91. 91.Thorne, S. 2016. Interpretive description: Qualitative research for applied practice. Routledge.
  92. 92.Thorne, S., Kirkham, S.R. and MacDonald-Emes, J. 1997. Interpretive description: a noncategorical qualitative alternative for developing nursing knowledge. Research in nursing & health. 20, 2 (1997), 169–177.
  93. 93.Vaast, E. 2025. Experiencing and addressing the moral ambivalence of developing digital technology: Insights from artificial intelligence developers. Proceedings of the Annual Hawaii International Conference on System Sciences (2025).
  94. 94.Vakkuri, V., Kemell, K.-K., Tolvanen, J., Jantunen, M., Halme, E. and Abrahamsson, P. 2022. How do software companies deal with artificial intelligence ethics? A gap analysis. Proceedings of the 26th International Conference on Evaluation and Assessment in Software Engineering (2022), 100–109.
  95. 95.Wang, A., Datta, T. and Dickerson, J.P. 2024. Strategies for increasing corporate responsible AI prioritization. Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society (2024), 1514–1526.
  96. 96.Wang, Q., Madaio, M., Kane, S., Kapania, S., Terry, M. and Wilcox, L. 2023. Designing responsible AI: Adaptations of UX practice to meet responsible AI challenges. Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (2023), 1–16.
  97. 97.Wang, Z.J., Kulkarni, C., Wilcox, L., Terry, M. and Madaio, M. 2024. Farsight: Fostering responsible AI awareness during AI application prototyping. Proceedings of the CHI Conference on Human Factors in Computing Systems (2024), 1–40.
  98. 98.Watkins-Hayes, C. 2019. Remaking a life. University of California Press.
  99. 99.Widder, D.G., Dabbish, L., Herbsleb, J.D. and Martelaro, N. 2024. Power and play: Investigating “license to critique” in teams’ AI ethics discussions. Proceedings of the ACM on human-computer interaction. 8, CSCW2 (2024), 1–23.
  100. 100.Widder, D.G. and Nafus, D. 2023. Dislocated accountabilities in the “AI supply chain”: Modularity and developers’ notions of responsibility. Big data & society. 10, 1 (2023). DOI:https://doi.org/10.1177/20539517231177620.
  101. 101.Winecoff, A.A. and Watkins, E.A. 2022. Artificial concepts of artificial intelligence. Proceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society (2022), 788–799.
  102. 102.Winecoff, A. and Bogen, M. 2025. Improving governance outcomes through AI documentation: Bridging theory and practice. Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (2025), 1–18.
  103. 103.Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E.P., Zhang, H., Gonzalez, J.E. and Stoica, I. 2023. Judging LLM-as-a-judge with MT-bench and Chatbot Arena. Advances in neural information processing systems. 36, (2023), 46595–46623.
  104. 104.Ziosi, M. and Pruss, D. 2024. Evidence of what, for whom? The socially contested role of algorithmic bias in a predictive policing tool. Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency (2024), 1596–1608.
  105. 105.
    1. The problem with the nudge effect: it can make you buy more carrots – but it can’t make you eat them. The Guardian. The Guardian.

Citation

MLA
Nahar, N., et al. “"I Don't Think RAI Applies to My Model'' -- Engaging Non-champions with Sticky Stories for Responsible AI Work”. arXiv, 2025, http://arxiv.org/abs/2509.22858v1.
APA
Nahar, N., Yang, C., Chen, Y., Deng, W. H., Holstein, K., Eslami, M., & Kästner, C. (2025). "I Don't Think RAI Applies to My Model'' -- Engaging Non-champions with Sticky Stories for Responsible AI Work. arXiv. http://arxiv.org/abs/2509.22858v1
Chicago
Nahar, N., C. Yang, Y. Chen, et al. 2025. “"I Don't Think RAI Applies to My Model'' -- Engaging Non-champions with Sticky Stories for Responsible AI Work”. arXiv. http://arxiv.org/abs/2509.22858v1.
Harvard
Nahar, N. et al. (2025) “"I Don't Think RAI Applies to My Model'' -- Engaging Non-champions with Sticky Stories for Responsible AI Work”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2509.22858v1.
Vancouver
1. Nahar N, Yang C, Chen Y, Deng WH, Holstein K, Eslami M, Kästner C (2025) "I Don't Think RAI Applies to My Model'' -- Engaging Non-champions with Sticky Stories for Responsible AI Work. arXiv

BibTeX

@article{nahar2025don,
  title = {"I Don't Think RAI Applies to My Model'' -- Engaging Non-champions with Sticky Stories for Responsible AI Work},
  author = {Nahar, Nadia and Yang, Chenyang and Chen, Yanxin and Deng, Wesley Hanwen and Holstein, Ken and Eslami, Motahhare and Kästner, Christian},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2509.22858v1},
  eprint = {2509.22858}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/