“I Don’t Think RAI Applies to My Model” – Engaging Non-champions with Sticky Stories for Responsible AI Work
Nadia NaharChenyang YangYanxin ChenWesley DengKenneth HolsteinMotahhare EslamiChristian Kastner
Introduces sticky stories—concrete narratives of unexpected machine learning failures—to overcome practitioner apathy and significantly increase the breadth and depth of ethical risks identified during development.
Organizations increasingly mandate Responsible AI (RAI) processes, checklists, and templates to prevent algorithmic harms, yet these mechanisms frequently fail in practice. While formal RAI champions and ethics advocates readily adopt these tools, the vast majority of machine learning practitioners—non-champions who lack prior intrinsic motivation or formal ethics roles—routinely treat RAI assessments as superficial, bureaucratic compliance tasks or dismiss them as completely irrelevant to their specific models.
The article aims to evaluate the nature of this practitioner disengagement and demonstrate an effective, theory-informed intervention to foster meaningful, critical engagement with RAI among non-champion practitioners. Specifically, it introduces "sticky stories"—tailored narratives of unexpected, severe, and concrete machine learning harms designed to disrupt practitioners' default assumptions and provoke genuine deliberation during risk assessment workflows.
To address this challenge, the authors followed a multi-stage approach. First, they conducted an ethnographic formative study at a partner technology organization, collecting roughly 22 hours of observation and interviews across governance teams and data scientists. Drawing on psychological theories of transformative learning and cognitive dissonance, they then developed an eight-step compound generative AI pipeline that produces stories tailored to a user's system across five core dimensions: concreteness, severity, surprisingness, diversity, and relevance. They validated the pipeline offline across 240 stories (120 sticky and 120 baseline prompts) using both human annotators and automated model evaluation. Finally, they conducted a controlled user study with 29 active practitioners analyzing their own machine learning projects under three conditions: no stories, baseline generic stories, and sticky stories, accompanied by a two-month post-study follow-up.
The investigation produced several key findings. First, formative observations confirmed that data scientists routinely bypassed RAI evaluations because mainstream media narratives had desensitized them into believing fairness concerns only applied to obvious demographic categories like race and gender. Second, the automated pipeline successfully embodied the intended qualities, outperforming baseline zero-shot prompts by 98% in concreteness, 43% in surprisingness, and 31% in perceived severity, though requiring 5.5 times more processing time and 46 times more computational tokens. Third, in the user study, practitioners presented with sticky stories increased their time spent on harm assessment by roughly 200% (from 5.4 to 16.6 minutes), whereas baseline stories produced only an 11% increase. Fourth, sticky stories prompted participants to identify 4.5 times more distinct harm categories and 3.5 times more subcategories compared to baseline stories. Finally, sticky stories triggered qualitative behavioral markers of critical reflection—such as questioning baseline assumptions, exploring non-obvious stakeholder perspectives, and forming concrete mitigation plans—even among skeptical or previously indifferent practitioners.
These findings indicate that providing structured governance templates alone cannot ensure AI safety if technical practitioners remain unmotivated. When practitioners view RAI as irrelevant, high-stakes operational, reputational, legal, and financial risks remain unnoticed before deployment. Rather than relying on short-term behavioral nudges or generic compliance exercises, organizations can successfully engage reluctant teams by confronting them with highly tailored, surprising scenarios that expose realistic failure modes within their own technical systems.
Based on these results, organizational leaders and AI engineering teams should integrate automated, context-specific narrative prompts directly into standard risk assessment pipelines, model cards, and pre-deployment auditing workflows. AI tools should emphasize diverse edge cases and severe, overlooked impacts rather than exhaustive, generic checklists. Furthermore, leadership should recognize different practitioner profiles—such as active resistors, compliant followers, and indifferent developers—and tailor organizational interventions accordingly to systematically build institutional competence.
Decision-makers should consider several limitations of the findings. The empirical evaluation relied on a modest sample size of 29 practitioners, and the qualitative formative findings were drawn from a single enterprise environment. In addition, the study captured immediate, single-session reflections and early intent; long-term behavioral transformation across daily development cycles over multi-year horizons remains to be demonstrated through broader, longitudinal field deployments.
No sufficiently relevant recommendations were found.
No sufficiently relevant recommendations were found.
