DuoDrama: Supporting Screenplay Refinement Through LLM-Assisted Human Reflection
Yuying TangXinyi ChenHaotian LiXing XieXiaojuan MaHuamin Qu
Presents DuoDrama, an LLM-assisted system that coordinates first-person character experience with structural story evaluation to provide actionable dual-perspective feedback for screenplay refinement.
Screenplay refinement is a decisive phase in film and television production where writers must balance two complementary viewpoints: an internal perspective that examines character emotions and intentions, and an external perspective that assesses overarching story structure, pacing, and theme. Existing artificial intelligence (AI) tools struggle to support this reflective process because they either generate critique from a single vantage point or deploy multi-agent personas independently without coordinating them, forcing writers to spend significant time filtering fragmented or misaligned suggestions.
The article demonstrates and evaluates DuoDrama, an AI-powered system designed to assist screenwriters during script refinement by coordinating internal character immersion with external evaluative critique. The system implements a novel workflow called ExReflect, inspired by performance theory: an AI agent first adopts a character’s internal experience role to simulate personal thoughts and then shifts to an actor’s external evaluation role to pose grounded feedback questions across five core narrative dimensions.
To develop and evaluate the system, the authors conducted a formative study with nine professional screenwriters to establish key design goals and feedback requirements. They then implemented DuoDrama using a multi-agent architecture where individual character agents maintain short- and long-term memory pipelines to generate line-level instant feedback and scene-level post-hoc feedback. Finally, the authors conducted a two-session mixed-methods user study with 14 professional screenwriters, gathering usability metrics, Likert-scale ratings, and qualitative interviews, while also performing statistical comparative testing against three baseline conditions.
The evaluation revealed several key findings regarding the effectiveness of experiential grounding and perspective shifting. First, DuoDrama achieved an average System Usability Scale score of 84.46, demonstrating strong usability, explanatory clarity, and logical coherence. Second, comparative analysis showed that grounding evaluation in simulated personal experience produced statistically significant improvements in feedback quality—specifically in content richness, detail specificity, and narrative relevance—compared to standard review approaches without experiential grounding. Third, the system significantly outperformed single-perspective and ungrounded baselines in fostering deep user reflection and motivation to revise, with 13 of 14 writers confirming that the tool effectively surfaced blind spots and hidden gaps.
These findings indicate that effective creative feedback requires a careful balance between situated context and critical distance. Rather than offering superficial text edits or overly broad critiques, AI can stimulate meaningful human reflection by demonstrating why a narrative issue emerges from within a scene and framing constructive inquiries. This approach reduces the cognitive burden on writers and helps bridge the gap between creative intention and execution, showing strong potential to generalize to other domains that require reflective evaluation, such as education and user experience design.
For future development, the article recommends implementing adaptive regulation mechanisms that allow users to customize feedback density and timing, as screenwriters showed divergent preferences between immediate line-by-line interruptions and post-hoc scene reviews. Future systems should also support interactive, conversational follow-ups to explore critiques in greater depth. Decision-makers should note that the study evaluated short-term interactions focused on text-based screenplay refinement; further longitudinal pilots and explorations into multimodal sensory cues (e.g., visual and auditory pacing) are necessary to fully assess long-term creative collaboration.
- Paper: CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society, Guohao Li et al. (2023). It introduces foundational inception prompting and role-playing mechanisms between collaborative LLM agents, which underpins DuoDrama's agent-based perspective-taking.
- Paper: InCharacter: Evaluating Personality Fidelity in Role-Playing Agents through Psychological Interviews, Xintao Wang et al. (2024). It provides the methodology for evaluating persona and psychological fidelity in role-playing LLM agents, which is crucial for DuoDrama's internal character experience modeling.
- Paper: Reflexion: language agents with verbal reinforcement learning, Noah Shinn et al. (2023). It formalizes verbal reinforcement learning and self-reflective feedback loops in LLM agents, establishing key concepts for reflection-driven refinement workflows.
- Paper: Re3: Generating Longer Stories With Recursive Reprompting and Revision, Kevin Yang et al. (2022). It demonstrates how decomposing narrative writing into structured planning, drafting, and iterative revision modules enhances long-form creative consistency.
- Paper: Dungeons and Dragons as a Dialog Challenge for Artificial Intelligence, Chris Callison-Burch et al. (2022). It examines how conditioning language models on explicit character roles versus narrator perspectives supports collaborative, interactive story generation.
- Paper: Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate, Tian Liang et al. (2024). It shows how simulating multi-agent debate and distinct viewpoints prevents degeneration-of-thought and enhances divergent thinking during reflection and evaluation.
- Paper: Personalizing Dialogue Agents: I have a dog, do you have pets too?, Saizheng Zhang et al. (2018). It establishes techniques for maintaining consistent agent personas and viewpoints across dialogue, which DuoDrama adapts for character-grounded feedback.
- Paper: Reasoning Models Generate Societies of Thought, Junsol Kim et al. (2026). It analyzes the internal mechanics of how reasoning models simulate multi-perspective dialogues and societies of thought, offering a deeper theoretical foundation for role-shifting workflows like ExReflect.
- Paper: PaperOrchestra: A Multi-Agent Framework for Automated AI Research Paper Writing, Yiwen Song et al. (2026). It extends role-based multi-agent collaboration and feedback-driven refinement to the automated authoring and review of complete scientific papers.
- Paper: Idea2Story: An Automated Pipeline for Transforming Research Concepts into Complete Scientific Narratives, Tengyue Xu et al. (2026). It applies structured retrieval and simulated review workflows to turn high-level research concepts into full scientific narratives.
- Paper: LEGOMem: Modular Procedural Memory for Multi-agent LLM Systems for Workflow Automation, Dongge Han et al. (2026). It broadens the reuse of operational experience by developing modular procedural memory architectures across multi-agent workflow systems.
