DuoDrama: Supporting Screenplay Refinement Through LLM-Assisted Human Reflection

Yuying TangXinyi ChenHaotian LiXing XieXiaojuan MaHuamin Qu

article2026CHI4 citations

Presents DuoDrama, an LLM-assisted system that coordinates first-person character experience with structural story evaluation to provide actionable dual-perspective feedback for screenplay refinement.

Listen

Screenplay refinement is a decisive phase in film and television production where writers must balance two complementary viewpoints: an internal perspective that examines character emotions and intentions, and an external perspective that assesses overarching story structure, pacing, and theme. Existing artificial intelligence (AI) tools struggle to support this reflective process because they either generate critique from a single vantage point or deploy multi-agent personas independently without coordinating them, forcing writers to spend significant time filtering fragmented or misaligned suggestions.

The article demonstrates and evaluates DuoDrama, an AI-powered system designed to assist screenwriters during script refinement by coordinating internal character immersion with external evaluative critique. The system implements a novel workflow called ExReflect, inspired by performance theory: an AI agent first adopts a character’s internal experience role to simulate personal thoughts and then shifts to an actor’s external evaluation role to pose grounded feedback questions across five core narrative dimensions.

To develop and evaluate the system, the authors conducted a formative study with nine professional screenwriters to establish key design goals and feedback requirements. They then implemented DuoDrama using a multi-agent architecture where individual character agents maintain short- and long-term memory pipelines to generate line-level instant feedback and scene-level post-hoc feedback. Finally, the authors conducted a two-session mixed-methods user study with 14 professional screenwriters, gathering usability metrics, Likert-scale ratings, and qualitative interviews, while also performing statistical comparative testing against three baseline conditions.

The evaluation revealed several key findings regarding the effectiveness of experiential grounding and perspective shifting. First, DuoDrama achieved an average System Usability Scale score of 84.46, demonstrating strong usability, explanatory clarity, and logical coherence. Second, comparative analysis showed that grounding evaluation in simulated personal experience produced statistically significant improvements in feedback quality—specifically in content richness, detail specificity, and narrative relevance—compared to standard review approaches without experiential grounding. Third, the system significantly outperformed single-perspective and ungrounded baselines in fostering deep user reflection and motivation to revise, with 13 of 14 writers confirming that the tool effectively surfaced blind spots and hidden gaps.

These findings indicate that effective creative feedback requires a careful balance between situated context and critical distance. Rather than offering superficial text edits or overly broad critiques, AI can stimulate meaningful human reflection by demonstrating why a narrative issue emerges from within a scene and framing constructive inquiries. This approach reduces the cognitive burden on writers and helps bridge the gap between creative intention and execution, showing strong potential to generalize to other domains that require reflective evaluation, such as education and user experience design.

For future development, the article recommends implementing adaptive regulation mechanisms that allow users to customize feedback density and timing, as screenwriters showed divergent preferences between immediate line-by-line interruptions and post-hoc scene reviews. Future systems should also support interactive, conversational follow-ups to explore critiques in greater depth. Decision-makers should note that the study evaluated short-term interactions focused on text-based screenplay refinement; further longitudinal pilots and explorations into multimodal sensory cues (e.g., visual and auditory pacing) are necessary to fully assess long-term creative collaboration.

arXiv: 2602.05854
Cover for DuoDrama: Supporting Screenplay Refinement Through LLM-Assisted Human Reflection

Abstract

AI has been increasingly integrated into screenwriting practice. In refinement, screenwriters expect AI to provide feedback that supports reflection across the internal perspective of characters and the external perspective of the overall story. However, existing AI tools cannot sufficiently coordinate the two perspectives to meet screenwriters' needs. To address this gap, we present DuoDrama, an AI system that generates feedback to assist screenwriters' reflection in refinement. To enable DuoDrama, based on performance theories and a formative study with nine professional screenwriters, we design the Experience-Grounded Feedback Generation Workflow for Human Reflection (ExReflect). In ExReflect, an AI agent adopts an experience role to generate experience and then shifts to an evaluation role to generate feedback based on the experience. A study with fourteen professional screenwriters shows that DuoDrama improves feedback quality and alignment and enhances the effectiveness, depth, and richness of reflection. We conclude by discussing broader implications and future directions.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Reflection in Screenwriting Refinement
  • 2.2 AI for Screenwriting Reflection
  • 2.3 AI-Assisted Human Reflection Strategy
  • 3 The DuoDrama System
  • 3.1 ExReflect: An Experience-Grounded Feedback Generation Workflow for Human Reflection Informed by Performance Theory
  • 3.1.1 Background
  • 3.1.2 ExReflect Workflow
  • 3.2 Formative Study
  • 3.2.1 Participants and Procedure
  • 3.2.2 Data Analysis and Results
  • 3.3 System Design
  • 3.3.1 Screenplay Pre-processing and Multi-Agent Orchestration
  • 3.3.2 ExReflect-powered Agent
  • 3.3.3 User Interface Design
  • 3.4 System Walk-Through
  • 3.5 DuoDrama System Implementation
  • 4 User Study
  • 4.1 Study Design
  • 4.1.1 Session 1: Experience and Effect Test
  • 4.1.2 Session 2: Offline Comparative Evaluation
  • 4.2 Study Participants and Procedure
  • 4.3 Data Collection and Analysis
  • 4.4 Study Session 1 Results: Experience and Effect Test
  • 4.4.1 DuoDrama Provides High-Quality Logic and Clarity Feedback for Reflection
  • 4.4.2 DuoDrama Supports Smooth Interaction and Contextually Aligned Feedback for Reflection
  • 4.4.3 DuoDrama Leads to Insight and Refinement Willingness for Reflection
  • 4.4.4 DuoDrama Provides Acceptable Timing and Quantity of Feedback
  • 4.5 Study Session 2 Results: Offline Comparative Evaluation
  • 4.5.1 DuoDrama Provides Higher Quality Feedback for Reflection
  • 4.5.2 DuoDrama Shows Stronger Aligned Feedback for Reflection
  • 4.5.3 DuoDrama Enhances the Perceived Effectiveness, Depth, and Richness of Human Reflection
  • 4.6 Summary of DuoDrama Advantages
  • 5 Discussion
  • 5.1 Experience-Grounded Feedback in Supporting Human Reflection
  • 5.2 Adaptive Regulation: Timing, Quantity, and Personalization
  • 5.3 Limitations and Future Work
  • 6 Conclusion
  • References

Knowls

  1. Knowl 1 — ExReflect: Experience-Grounded Feedback Generation Workflow for Human Reflection

    model/method

    The Experience-Grounded Feedback Generation Workflow for Human Reflection (ExReflect) is a workflow designed to balance internal immersion with external critical evaluation within a single AI agent. Drawing on two traditions from performance theory—Konstantin Stanislavski's system of psychological embodiment and living through given circumstances, and Bertolt Brecht's epic theatre technique of critical distancing via the alienation effect (Verfremdungseffekt)—ExReflect operates via sequential stakeholder role-switching:

    1. Experience Role (Immersive Enactment): The agent first adopts an internal stakeholder identity (e.g., a screenplay character). Given the character's persona (background, personality traits, motivations) and the immediate interaction context (scene, environment, co-characters), the agent generates private, situated inner thoughts that represent its firsthand personal experience (PEPE).
    2. Evaluation Role (Critical Analysis): The agent then shifts to an external stakeholder identity (e.g., the actor portraying that character). It evaluates the scene from a critical distance, using the generated personal experience (PEPE) as supplementary short-term context. The evaluation role generates constructive, question-based feedback targeting narrative, emotional, and structural coherence.

    This two-step process ensures that critical evaluation is deeply grounded in internally situated character logic rather than functioning as generic or detached critique.

  2. Knowl 2 — Multi-Agent Orchestration and ExReflect Pipeline in DuoDrama

    algorithm

    DuoDrama operationalizes ExReflect across multi-character screenplays using a coordinated multi-agent architecture implemented via LangChain Expression Language (LCEL), vector retrieval (FAISS), and large language models (GPT-4.1 for reasoning and GPT-4o for visual generation).

    Input: Screenplay text SS, Character profiles CC, Long-term vector store MLTM_{LT}
    Output: Inner thoughts TT, Instant feedback FinstF_{inst}, Post-hoc feedback FpostF_{post}
    Parse SS into scene blocks S={s1,s2,…,sK}S = \{s_1, s_2, \dots, s_K\} via LLM prompt with regex fallback
    for each scene s∈{s1,…,sK}s \in \{s_1, \dots, s_K\} do
        Segment ss into line-by-line dialogue and action units
        Initialize dynamic short-term memory MST[c]←∅M_{ST}[c] \leftarrow \emptyset for each active character cc
        for each line l∈sl \in s do
            Identify character cc performing or speaking ll
            Activate agent AcA_c in Experience Role
            Retrieve relevant past scene memories mpast←FAISS_Query(MLT,l)m_{past} \leftarrow \text{FAISS\_Query}(M_{LT}, l)
            Execute Chain-of-Thought enactment:
                rinit←InitialInterpretation(l,MST[c])r_{init} \leftarrow \text{InitialInterpretation}(l, M_{ST}[c])
                rmem←MemoryRecall(rinit,C[c],mpast)r_{mem} \leftarrow \text{MemoryRecall}(r_{init}, C[c], m_{past})
                robj←ObjectiveDefinition(rmem,l)r_{obj} \leftarrow \text{ObjectiveDefinition}(r_{mem}, l)
                t←SynthesizeInnerThoughts(rinit,rmem,robj)t \leftarrow \text{SynthesizeInnerThoughts}(r_{init}, r_{mem}, r_{obj})
            Append tt to MST[c]M_{ST}[c] and record tt into TT
            
            Shift agent AcA_c to Evaluation Role (actor portraying cc)
            ctxinst←Combine(MST[c],mpast,l)ctx_{inst} \leftarrow \text{Combine}(M_{ST}[c], m_{past}, l)
            qinst←GenerateActorFeedback(ctxinst)q_{inst} \leftarrow \text{GenerateActorFeedback}(ctx_{inst})
            f←VerifyAndFilterFeedback(qinst,ctxinst,mode=instant)f \leftarrow \text{VerifyAndFilterFeedback}(q_{inst}, ctx_{inst}, \text{mode}=\text{instant})
            if f≠nullf \neq \text{null} then
                Append ff to FinstF_{inst}
            end if
        end for
        
        for each active agent AcA_c in scene ss do
            ctxpost←CombineFullScene(s,MST[c],MLT)ctx_{post} \leftarrow \text{CombineFullScene}(s, M_{ST}[c], M_{LT})
            qpost←GenerateSceneFeedback(ctxpost)q_{post} \leftarrow \text{GenerateSceneFeedback}(ctx_{post})
            fscene←VerifyAndFilterFeedback(qpost,ctxpost,mode=post_hoc)f_{scene} \leftarrow \text{VerifyAndFilterFeedback}(q_{post}, ctx_{post}, \text{mode}=\text{post\_hoc})
            if fscene≠nullf_{scene} \neq \text{null} then
                Append fscenef_{scene} to FpostF_{post}
            end if
        end for
    end for
    return T,Finst,FpostT, F_{inst}, F_{post}
  3. Knowl 3 — Self-Verification and Feedback Timing Assessment Algorithm

    algorithm

    To ensure that AI feedback is accurate, stylistically diverse, and well-timed without breaking the screenwriter's creative flow, DuoDrama employs a neutral self-verification module that evaluates candidate feedback prior to displaying it.

    Input: Candidate feedback question qq, Interaction context CctxC_{ctx}, Timing mode τ∈{instant,post_hoc}\tau \in \{\text{instant}, \text{post\_hoc}\}
    Output: Decision accept∈{true,false}accept \in \{\text{true}, \text{false}\}
    1. Evidence Verification:
    is_grounded←CheckFactualGrounding(q,Cctx)is\_grounded \leftarrow \text{CheckFactualGrounding}(q, C_{ctx})
    if is_grounded=falseis\_grounded = \text{false} then
        return false
    end if
    2. Target Dimension Check:
    has_dimension←CoversAnyDimension(q,{emotion,motivation,relationship,pacing,theme})has\_dimension \leftarrow \text{CoversAnyDimension}(q, \{\text{emotion}, \text{motivation}, \text{relationship}, \text{pacing}, \text{theme}\})
    if has_dimension=falsehas\_dimension = \text{false} then
        return false
    end if
    3. Impact and Timing Evaluation:
    if τ=instant\tau = \text{instant} then
        is_critical←JudgeImmediateSeverity(q)is\_critical \leftarrow \text{JudgeImmediateSeverity}(q)
        if is_critical=falseis\_critical = \text{false} then
            return false
        end if
    else
        is_holistic←JudgeSceneLevelImpact(q)is\_holistic \leftarrow \text{JudgeSceneLevelImpact}(q)
        if is_holistic=falseis\_holistic = \text{false} then
            return false
        end if
    end if
    4. Expression Diversity Check (Soft Constraint):
    has_cliche←DetectFormulaicJargon(q)has\_cliche \leftarrow \text{DetectFormulaicJargon}(q)
    if has_cliche=truehas\_cliche = \text{true} then
        has_high_value←EvaluateInsightTradeoff(q)has\_high\_value \leftarrow \text{EvaluateInsightTradeoff}(q)
        if has_high_value=falsehas\_high\_value = \text{false} then
            return false
        end if
    end if
    return true
  4. Knowl 4 — Comparative Experimental Setup for AI-Assisted Screenwriting Refinement

    experimental setup

    The effectiveness of DuoDrama's feedback mechanism was evaluated through an offline comparative study involving N=14N = 14 professional screenwriters (self-reported screenwriting experience ranging from 2 to 15 years; all had prior AI screenwriting tool experience). Each participant submitted an original screenplay draft containing an outline, character descriptions, and full scene dialogue/actions.

    Participants evaluated four systematically varied feedback conditions on a new screenplay excerpt (Scene B) in a randomized order:

    1. Eval-PE (DuoDrama): Evaluation role perspective (actor) grounded in the experience role's generated personal experience (PEPE). Delivers line-level instant feedback and scene-level post-hoc feedback across five reflection dimensions.
    2. Exp-PE: Experience role perspective (character) with PEPE. Feedback is posed as first-person character inquiries constrained strictly to what is knowable to that character.
    3. Eval-NoPE: Evaluation role perspective without personal experience simulation. Generates post-hoc feedback along the five dimensions based solely on script text without line-anchored inner thoughts.
    4. Rev-NoPE (Industry Baseline): Screenplay reviewer perspective without personal experience. Relies solely on screenplay text to generate post-hoc critiques from an external reviewer standpoint.

    Evaluation utilized an 18-item 7-point Likert scale measuring four primary constructs: Feedback Alignment (DG2), Feedback Quality (DG1), Perceived Effectiveness (DG3), and User Reflection (DG3). Statistical comparisons were conducted using two-tailed Wilcoxon signed-rank tests for matched pairs, reporting the test statistic (WW), pp-value, and effect size r=Z/nr = Z / \sqrt{n}.

  5. Knowl 5 — Comparative Evaluation of DuoDrama Against Alternative Feedback Conditions

    data/table

    The quantitative results from the offline comparative evaluation (N=14N = 14 professional screenwriters) demonstrate the superiority of DuoDrama over conditions lacking either evaluative distance (Exp-PE) or experiential grounding (Eval-NoPE, Rev-NoPE).

    Dimension Subdimension DuoDrama DuoDrama vs. Exp-PE DuoDrama vs. Eval-NoPE DuoDrama vs. Rev-NoPE
    MM MM WW (pp) rr MM WW (pp) rr MM WW (pp) rr
    Alignment Character emotion 5.0 5.0 43.5 (.634) 0.13 4.5 26.5 (.231) 0.32 3.5 16.0 (.044*) 0.54
    Behavioral motivation 6.0 5.0 26.5 (.180) 0.36 5.0 20.0 (.146) 0.39 3.0 7.0 (.014*) 0.66
    Character relationship 5.0 5.0 33.0 (.437) 0.21 4.0 13.5 (.041*) 0.55 4.0 9.0 (.018*) 0.63
    Plot pacing 5.0 3.5 29.0 (.238) 0.32 3.0 7.0 (.006**) 0.73 4.0 10.0 (.053) 0.52
    Thematic consistency 6.0 5.0 33.0 (.295) 0.28 4.5 20.0 (.078) 0.47 5.0 19.0 (.129) 0.41
    Quality Content richness 6.0 5.0 32.0 (.199) 0.34 3.0 11.0 (.008**) 0.70 4.0 10.0 (.016*) 0.64
    Clarity 5.0 5.0 22.0 (.187) 0.35 4.5 9.5 (.031*) 0.58 6.0 41.5 (.610) 0.14
    Relevance 6.0 6.0 31.0 (.611) 0.14 5.5 10.0 (.085) 0.46 4.5 0.0 (.009**) 0.70
    Comprehensibility 6.5 6.0 30.0 (.563) 0.15 5.0 19.0 (.186) 0.35 6.0 28.0 (.356) 0.25
    Specificity of detail 6.0 5.0 30.0 (.215) 0.33 3.5 5.0 (.005**) 0.76 3.0 3.0 (.002**) 0.84
    Perceived Emotional insight 6.0 4.0 25.5 (.160) 0.38 4.0 7.5 (.023*) 0.61 4.0 7.0 (.009**) 0.69
    Effectiveness Motivational insight 6.0 5.0 19.0 (.130) 0.40 4.5 14.0 (.032*) 0.57 5.0 25.5 (.126) 0.41
    Relationship insight 6.0 4.0 27.0 (.191) 0.35 4.0 7.5 (.010**) 0.69 3.5 0.0 (.004**) 0.78
    Plot pacing insight 5.0 3.5 22.5 (.087) 0.46 3.0 18.0 (.080) 0.47 5.0 36.0 (.559) 0.16
    Thematic insight 5.0 2.0 16.5 (.065) 0.49 3.0 21.5 (.250) 0.31 4.0 29.0 (.519) 0.17
    Revision motivation 6.0 2.5 0.0 (.004**) 0.78 4.0 0.0 (.004**) 0.78 6.0 10.0 (.085) 0.46
    User Depth of reflection 6.0 4.5 7.0 (.009**) 0.69 5.0 8.5 (.017*) 0.64 5.5 20.0 (.308) 0.27
    Reflection Richness of reflection 5.5 3.5 23.0 (.076) 0.47 2.0 0.0 (.002**) 0.84 3.5 13.5 (.015*) 0.65

    *Note: MM represents median score on a 7-point Likert scale. Significance levels: * p<.05p < .05, ** p<.01p < .01. Key conclusions:

    1. DuoDrama vs. Exp-PE: Shifting from an experience role to an external actor evaluation role significantly increased revision motivation (p=.004,r=0.78p = .004, r = 0.78) and depth of reflection (p=.009,r=0.69p = .009, r = 0.69), overcoming the passive immersion of purely character-based agents.
    2. DuoDrama vs. Eval-NoPE: Grounding the evaluation role in simulated personal experience (PEPE) yielded statistically significant improvements in content richness (p=.008p = .008), specificity (p=.005p = .005), pacing alignment (p=.006p = .006), emotional insight (p=.023p = .023), relationship insight (p=.010p = .010), revision motivation (p=.004p = .004), reflection depth (p=.017p = .017), and reflection richness (p=.002p = .002).
    3. DuoDrama vs. Rev-NoPE: Compared to standard script reviewing, DuoDrama produced significantly more relevant (p=.009p = .009), detailed (p=.002p = .002), and emotionally/motivationally aligned feedback (p<.05p < .05 across emotion, motivation, and relationships).
  6. Knowl 6 — Usability and User Experience of DuoDrama

    empirical result

    In interactive user testing with 14 professional screenwriters evaluating Scene A of their screenplays:

    • System Usability: DuoDrama achieved an average System Usability Scale (SUS) score of 84.4684.46 (standard benchmark for good usability is ≥68\ge 68).
    • Simulation Credibility and Clarity (DG1): 10 of 14 participants reported positive ratings (scores 5–7 on a 7-point scale) for logical coherence (Q5), 12 of 14 for explanatory clarity (Q8), and 9 of 14 for character inner-thought credibility (Q2). Character immersion (Q1) had 7 of 14 positive ratings, with participants noting that while logical consistency was strong, complex layered psychological subtexts were occasionally simplified.
    • Interaction and Alignment (DG2): 13 of 14 participants rated the transition from character inner thoughts to actor feedback as smooth (Q6), and 10 of 14 rated feedback as relevant to plot and character state (Q7). Character imagination alignment (Q3) had 6 of 14 positive ratings; minor deviations were viewed by some writers as constructive creative variation.
    • Insight and Refinement Willingness (DG3): 11 of 14 participants rated the feedback as constructive for deeper exploration (Q9), and 9 of 14 confirmed it prompted them to rethink character or plot directions (Q10).
    • Feedback Timing Preferences (DG4): 10 of 14 screenwriters expressed a preference for scene-level post-hoc feedback over sentence-level instant feedback, citing that post-hoc feedback provided holistic structural guidance without interrupting line-by-line creative reading.
  7. Knowl 7 — Design Goals and Reflection Dimensions for Screenplay Refinement AI

    definition

    Synthesized from a 3W1H (who, what, when, how) formative study with nine professional screenwriters, screenplay refinement requires supporting two complementary reflection perspectives (internal character logic vs. external structural overview) across five narrative dimensions and four design goals (DGs):

    Five Reflection Dimensions

    1. Character Emotion: Psychological credibility and moment-to-moment emotional transitions.
    2. Behavioral Motivation: Consistency and justification of character goals and actions.
    3. Character Relationships: Dynamics, power shifts, and interpersonal subtexts across scenes.
    4. Plot Pacing: Narrative momentum, causality, and scene transitions.
    5. Thematic Consistency: Alignment between local dialogue/actions and the script's core thematic premise.

    Four Design Goals (DGs)

    • DG1 (Simulation & Critique): High-quality experiential simulation through visible character inner thoughts paired with understandable, transparent critical feedback from an actor persona.
    • DG2 (Alignment): Feedback must maintain tight alignment with full narrative context, backstories, and the five reflection dimensions rather than analyzing isolated lines.
    • DG3 (Creative Insight): Feedback must provoke reflection by surfacing blind spots, hidden gaps, and actionable refinement directions rather than prescribing direct textual fixes.
    • DG4 (Timing Balance): Feedback timing must balance creative flow and intervention urgency by offering granular in-action instant feedback for critical disruptions and holistic on-action post-hoc feedback at scene boundaries.
  8. Knowl 8 — User Interface Architecture of DuoDrama

    model/method

    DuoDrama's frontend interface is structured into four coordinated panels corresponding to the design goals:

    1. Control Panel (DG1, DG2): Enables screenwriters to upload screenplay drafts, trigger automatic segmentation, view the clickable line-by-line script navigation overview, and toggle active ExReflect character agents.
    2. Enactment Panel (DG1, DG2, DG3): Presents the screenplay line-by-line interleaved with the active characters' simulated inner thoughts (personal experience). Displays clickable red notification icons beside specific sentences to reveal collapsible instant feedback boxes nested under target lines.
    3. Critique Panel (DG1, DG2, DG3): Displays a scrollable bubble-style list of post-hoc feedback questions generated at the scene level, guiding reflection on macro-narrative structure, character relationship arcs, and thematic progression.
    4. Value Marking Function (DG3, DG4): Allows users to click a checkmark on useful inner thoughts, instant feedback, or post-hoc critique items. Marked items are highlighted and aggregated into a background JSON object containing contextual metadata (character, scene ID, feedback type) for later revision review.
  9. Knowl 9 — Limitations of DuoDrama and AI-Assisted Reflection

    limitation

    The authors identify several limitations of DuoDrama and its evaluation:

    1. Single Domain and Creative Stage: The workflow was evaluated exclusively in screenwriting refinement following a completed draft. Its generalizability to other creative stages (e.g., initial idea drafting or late-stage polishing) and other dual-perspective domains (e.g., student-teacher educational feedback or user-designer interface evaluation) has not been empirically verified.
    2. Short-Term Evaluation Horizon: The user study comprised two lab sessions per participant; longitudinal deployments are required to assess how continuous interaction with ExReflect agents influences writing practices and authorial voice over time.
    3. Text-Only Experiential Grounding: Personal experience generation relies strictly on textual dialogue and action descriptions, omitting multi-modal narrative cues such as visual framing, lighting, vocal inflection, tone, and pacing.
    4. Static Feedback Density: The rule- and prompt-based filtering mechanism applies fixed thresholds, which led to divided preferences among screenwriters regarding feedback volume and interruption frequency.
    5. Lack of Inter-Agent Negotiation: In multi-character scenes, agents generate personal experiences and actor critiques independently without communicating or resolving conflicting perspectives prior to presenting feedback to the writer.

Coverage note — Specific qualitative screenplay excerpts and illustrative participant case scripts from Table 3 were omitted as standalone knowls because they represent anecdotal evidence rather than generalizable mechanisms, though their qualitative findings are fully integrated into the comparative evaluation knowls.

References

  1. 1.Adnan Abbas, Caleb Wohn, Donghan Hu, Eugenia H Rho, and Sang Won Lee. 2025. PITCH: Designing Agentic Conversational Support for Planning and Self-reflection. In Proceedings of the 7th ACM Conference on Conversational User Interfaces (CUI ’25). Association for Computing Machinery, New York, NY, USA, Article 62, 22 pages. doi:10.1145/3719160.3736634
  2. 2.Piotr D. Adamczyk and Brian P. Bailey. 2004. If Not Now, When? The Effects of Interruption at Different Moments Within Task Execution. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Vienna, Austria) (CHI ’04). Association for Computing Machinery, New York, NY, USA, 271–278. doi:10.1145/985692.985727
  3. 3.J.E. Allen, C.I. Guinn, and E. Horvtz. 1999. Mixed-Initiative Interaction. IEEE Intelligent Systems and their Applications 14, 5 (1999), 14–23. doi:10.1109/5254.796083
  4. 4.Craig Batty, Radha O’Meara, Stayci Taylor, Hester Joyce, Philippa Burne, Noel Maloney, Mark Poole, and Marilyn Tofler. 2018. Script Development as a ’Wicked Problem’. Journal of Screenwriting 9, 2 (2018), 153–174. doi:10.1386/josc.9.2.153_1
  5. 5.Craig Batty and Stayci Taylor. 2025. Interrogating Writing Practices: Perspectives from the Screenwriting Industry. Sat (2025). doi:10.62959/WIP-01-2015-08
  6. 6.Eric P.S. Baumer, Vera Khovanskaya, Mark Matthews, Lindsay Reynolds, Victoria Schwanda Sosik, and Geri Gay. 2014. Reviewing Reflection: On the Use of Reflection in Interactive System Design. In Proceedings of the 2014 Conference on Designing Interactive Systems (Vancouver, BC, Canada) (DIS ’14). Association for Computing Machinery, New York, NY, USA, 93–102. doi:10.1145/2598510.2598598
  7. 7.Karim Benharrak, Tim Zindulka, Florian Lehmann, Hendrik Heuer, and Daniel Buschek. 2024. Writer-Defined AI Personas for On-Demand Feedback Generation. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24). Association for Computing Machinery, New York, NY, USA, Article 1049, 18 pages. doi:10.1145/3613904.3642406
  8. 8.Eric Bentley. 1964. Are Stanislavski and Brecht Commensurable? Tulane Drama Review 9, 1 (1964), 69–76. doi:10.2307/1124779
  9. 9.Marit Bentvelzen, Paweł W. Woźniak, Pia S.F. Herbes, Evropi Stefanidi, and Jasmin Niess. 2022. Revisiting Reflection in HCI: Four Design Resources for Technologies that Support Reflection. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 6, 1, Article 2 (March 2022), 27 pages. doi:10.1145/3517233
  10. 10.David Bordwell. 2006. The Way Hollywood Tells It: Story and Style in Modern Movies. Univ of California Press. doi:10.1525/9780520932326
  11. 11.Virginia Braun and Victoria Clarke. 2006. Using Thematic Analysis in Psychology. Qualitative Research in Psychology 3, 2 (2006), 77–101. doi:10.1191/1478088706qp063oa
  12. 12.Bertolt Brecht. 2013. Short Description of a New Technique of Acting Which Produces an Alienation Effect. In The Twentieth Century Performance Reader. Routledge, 101–112. doi:10.4324/9780203125236-14
  13. 13.Kevin Michael Brooks. 1999. Metalinear Cinematic Narrative: Theory, Process, and Tool. Ph. D. Dissertation. Massachusetts Institute of Technology. doi:1721.1/9544
  14. 14.Sharon Marie Carnicke. 2020. Stanislavsky in Focus: An Acting Master for the Twenty-First Century. Routledge. doi:10.4324/9781003134381
  15. 15.Jing Chen, Xinyu Zhu, Cheng Yang, Chufan Shi, Yadong Xi, Yuxiang Zhang, Junjie Wang, Jiashu Pu, Tian Feng, Yujiu Yang, and Rongsheng Zhang. 2024. HoLLMwood: Unleashing the Creativity of Large Language Models in Screenwriting via Role Playing. In Findings of the Association for Computational Linguistics: EMNLP 2024. Association for Computational Linguistics, Miami, Florida, USA, 8075–8121. doi:10.18653/v1/2024.findings-emnlp.474
  16. 16.Xinyue Chen, Lev Tankelevitch, Rishi Vanukuru, Ava Elizabeth Scott, Payod Panda, and Sean Rintel. 2025. Are We On Track? AI-Assisted Active and Passive Goal Reflection During Meetings. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25). Association for Computing Machinery, New York, NY, USA, Article 705, 22 pages. doi:10.1145/3706598.3714052
  17. 17.Amy Cook, Jessica Hammer, Salma Elsayed-Ali, and Steven Dow. 2019. How Guiding Questions Facilitate Feedback Exchange in Project-Based Learning. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (Glasgow, Scotland Uk) (CHI ’19). Association for Computing Machinery, New York, NY, USA, 1–12. doi:10.1145/3290605.3300368
  18. 18.Jane L. E, Yu-Chun Grace Yen, Isabelle Yan Pan, Grace Lin, Mingyi Li, Hyoungwook Jin, Mengyi Chen, Haijun Xia, and Steven P. Dow. 2024. When to Give Feedback: Exploring Tradeoffs in the Timing of Design Feedback. In Proceedings of the 16th Conference on Creativity & Cognition (Chicago, IL, USA) (C&C ’24). Association for Computing Machinery, New York, NY, USA, 292–310. doi:10.1145/3635636.3656183
  19. 19.Jack Epps Jr. 2016. Screenwriting Is Rewriting: The Art and Craft of Professional Revision. Bloomsbury Publishing USA.
  20. 20.Gerhard Fischer, Kumiyo Nakakoji, Jonathan Ostwald, Gerry Stahl, and Tamara Sumner. 1993. Embedding Critics in Design Environments. The Knowledge Engineering Review 8, 4 (1993), 285–307. doi:10.1017/S026988890000031X
  21. 21.Nicholas Flanagan, Christopher Gist, Sue Joseph, and Craig Batty. 2025. The Autoethnographic Screenwriter: Deepening Insights and Producing Knowledge in the Creative PhD. New Writing 22, 1 (2025), 37–54. doi:10.1080/14790726.2024.2424593
  22. 22.Katy Ilonka Gero and Lydia B. Chilton. 2019. Metaphoria: An Algorithmic Companion for Metaphor Creation. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (Glasgow, Scotland Uk) (CHI ’19). Association for Computing Machinery, New York, NY, USA, 1–12. doi:10.1145/3290605.3300526
  23. 23.Leo A. Goodman. 1961. Snowball Sampling. The Annals of Mathematical Statistics 32 (1961), 148–170. doi:10.1214/aoms/1177705148
  24. 24.Paolo Grigis and Antonella De Angeli. 2024. Playwriting with Large Language Models: Perceived Features, Interaction Strategies and Outcomes. In Proceedings of the 2024 International Conference on Advanced Visual Interfaces (Arenzano, Genoa, Italy) (AVI ’24). Association for Computing Machinery, New York, NY, USA, Article 38, 9 pages. doi:10.1145/3656650.3656688
  25. 25.Paul Joseph Gulino. 2024. Screenwriting: The Sequence Approach. Bloomsbury Publishing.
  26. 26.Paul Joseph Gulino and Connie Shears. 2018. The Science of Screenwriting: The Neuroscience Behind Storytelling Strategies. Bloomsbury Publishing USA.
  27. 27.Alicia Guo, Shreya Sathyanarayanan, Leijie Wang, Jeffrey Heer, and Amy X. Zhang. 2025. From Pen to Prompt: How Creative Writers Integrate AI into Their Writing Practice. In Proceedings of the 2025 Conference on Creativity and Cognition (C&C ’25). Association for Computing Machinery, New York, NY, USA, 527–545. doi:10.1145/3698061.3726910
  28. 28.Brett A. Halperin and Daniela K. Rosner. 2025. “AI is Soulless”: Hollywood Film Workers’ Strike and Emerging Perceptions of Generative Cinema. ACM Trans. Comput.-Hum. Interact. 32, 2, Article 19 (April 2025), 27 pages. doi:10.1145/3716135
  29. 29.Christina Kallas. 2017. Creative Screenwriting: Understanding Emotional Structure. Bloomsbury Publishing.
  30. 30.Max Kreminski, John Joon Young Chung, and Melanie Dickinson. 2024. Intent Elicitation in Mixed-Initiative Co-Creativity. In IUI Workshops. https://ceur-ws.org/Vol-3660/paper6.pdf
  31. 31.Tom Lazarus. 2007. Rewriting Secrets for Screenwriters: Seven Strategies to Improve and Sell Your Work. Macmillan+ ORM.
  32. 32.Hyunseung Lim, Dasom Choi, DaEun Choi, Sooyohn Nam, and Hwajung Hong. 2026. Feed-O-Meter: Investigating AI-Generated Mentee Personas as Interactive Agents for Scaffolding Design Feedback Practice. International Journal of Human-Computer Studies 208 (2026), 103687. doi:10.1016/j.ijhcs.2025.103687
  33. 33.Xingyu Bruce Liu, Shitao Fang, Weiyan Shi, Chien-Sheng Wu, Takeo Igarashi, and Xiang ’Anthony’ Chen. 2025. Proactive Conversational Agents with Inner Thoughts. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25). Association for Computing Machinery, New York, NY, USA, Article 184, 19 pages. doi:10.1145/3706598.3713760
  34. 34.Jonas Lowgren and Erik Stolterman. 2007. Thoughtful Interaction Design: A Design Perspective on Information Technology. Mit Press.
  35. 35.Robert McKee. 1997. Substance, Structure, Style, and the Principles of Screenwriting.
  36. 36.Robert McKee. 2005. Story. Vol. 3. Dixit.
  37. 37.Jack Mezirow et al. 1990. Fostering Critical Reflection in Adulthood. Vol. 366. San Francisco: Jossey-Bass.
  38. 38.Piotr Mirowski, Kory W. Mathewson, Jaylen Pittman, and Richard Evans. 2023. Co-Writing Screenplays and Theatre Scripts with Language Models: Evaluation by Industry Professionals. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23). Association for Computing Machinery, New York, NY, USA, Article 355, 34 pages. doi:10.1145/3544548.3581225
  39. 39.Sonia Moore. 1984. The Stanislavski System: The Professional Training of an Actor. Penguin.
  40. 40.Charlie Moritz. 2013. Scriptwriting for the Screen. Routledge.
  41. 41.Meg Mumford. 1995. Brecht Studies Stanislavski: Just a Tactical Move? New Theatre Quarterly 11, 43 (1995), 241–258. doi:10.1017/S0266464X0000912X
  42. 42.Hugh Munby. 1989. Reflection-in-Action and Reflection-on-Action. Current Issues in Education 9, 1 (1989), 31–42. doi:10.1353/eac.1989.a592219
  43. 43.Jill Nelmes. 2008. Developing the Screenplay Wingwalking: An Analysis of the Writing and Rewriting Process. Journal of British Cinema and Television 5, 2 (2008), 335–352. doi:10.3366/E174345210800040X
  44. 44.Changhoon Oh, Jungwoo Song, Jinhan Choi, Seonghyeon Kim, Sungwoo Lee, and Bongwon Suh. 2018. I Lead, You Help but Only with Enough Details: Understanding User Experience of Co-Creation with Artificial Intelligence. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems (Montreal QC, Canada) (CHI ’18). Association for Computing Machinery, New York, NY, USA, 1–13. doi:10.1145/3173574.3174223
  45. 45.Adiba Orzikulova, Han Xiao, Zhipeng Li, Yukang Yan, Yuntao Wang, Yuanchun Shi, Marzyeh Ghassemi, Sung-Ju Lee, Anind K Dey, and Xuhai Xu. 2024. Time2Stop: Adaptive and Explainable Human-AI Loop for Smartphone Overuse Intervention. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24). Association for Computing Machinery, New York, NY, USA, Article 250, 20 pages. doi:10.1145/3613904.3642747
  46. 46.Qiyu Pan, Jianqiao Zeng, Jie Wang, Junyu Liu, Yihan Qiu, Kangyu Yuan, and Zhenhui Peng. 2025. AMQuestioner: Training Critical Thinking with Question-Driven Interactive Argument Maps in Online Discussion. Proc. ACM Hum.-Comput. Interact. 9, 7, Article CSCW370 (Oct. 2025), 48 pages. doi:10.1145/3757551
  47. 47.Carl Plantinga. 2018. Screen Stories: Emotion and the Ethics of Engagement. Oxford University Press. doi:10.1093/oso/9780190867133.001.0001
  48. 48.Hua Xuan Qin, Shan Jin, Ze Gao, Mingming Fan, and Pan Hui. 2024. CharacterMeet: Supporting Creative Writers’ Entire Story Character Construction Processes Through Conversation with LLM-Powered Chatbot Avatars. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24). Association for Computing Machinery, New York, NY, USA, Article 1051, 19 pages. doi:10.1145/3613904.3642105
  49. 49.Hannah Rashkin, Elizabeth Clark, Fantine Huot, and Mirella Lapata. 2025. Help Me Write a Story: Evaluating LLMs’ Ability to Generate Writing Feedback. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (ACL ’25). Association for Computational Linguistics, Vienna, Austria, 25827–25847. doi:0.18653/v1/2025.acl-long.1254
  50. 50.Mohi Reza, Jeb Thomas-Mitchell, Peter Dushniku, Nathan Laundry, Joseph Jay Williams, and Anastasia Kuzminykh. 2025. Co-Writing with AI, on Human Terms: Aligning Research with User Demands Across the Writing Process. Proc. ACM Hum.-Comput. Interact. 9, 7, Article CSCW385 (Oct. 2025), 37 pages. doi:10.1145/3757566
  51. 51.Peter Robertson, Judyand Wiemer-Hastings. 2002. Feedback on Children’s Stories via Multiple Interface Agents. In Intelligent Tutoring Systems (ITS 2002). Springer, Berlin, Heidelberg, 923–932. doi:10.1007/3-540-47987-2_92
  52. 52.Jeff Sauro and James R Lewis. 2016. Quantifying the User Experience: Practical Statistics for User Research. Morgan Kaufmann.
  53. 53.Jule Britt Selbo. 2011. The Constructive Use of Film Genre for the Screenwriter: Creating Film Genre’s Mental Space. University of Exeter (United Kingdom).
  54. 54.Siri Senje. 2017. Formatting the Imagination: A Reflection on Screenwriting as a Creative Practice. Journal of Screenwriting 8, 3 (2017), 267–285. doi:10.1386/josc.8.3.267_1
  55. 55.Inhwa Song, SoHyun Park, Sachin R Pendse, Jessica Lee Schleider, Munmun De Choudhury, and Young-Ho Kim. 2025. ExploreSelf: Fostering User-driven Exploration and Reflection on Personal Challenges with Adaptive Guidance by Large Language Models. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25). Association for Computing Machinery, New York, NY, USA, Article 306, 22 pages. doi:10.1145/3706598.3713883
  56. 56.WAJ Steer. 1968. Brecht’s Epic Theatre: Theory and Practice. The Modern Language Review 63, 3 (1968), 636–649. doi:10.2307/3722205
  57. 57.Debbie Stone, Caroline Jarrett, Mark Woodroffe, and Shailey Minocha. 2005. User Interface Design and Evaluation. Elsevier.
  58. 58.Yuqian Sun, Yuying Tang, Ze Gao, Zhijun Pan, Chuyan Xu, Yurou Chen, Kejiang Qian, Zhigang Wang, Tristan Braud, Chang Hee Lee, and Ali Asadipour. 2023. AI Nüshu: An Exploration of Language Emergence in Sisterhood Through the Lens of Computational Linguistics. In SIGGRAPH Asia 2023 Art Papers (Sydney, NSW, Australia) (SA ’23). Association for Computing Machinery, New York, NY, USA, Article 4, 7 pages. doi:10.1145/3610591.3616427
  59. 59.Shayan Talaei, Meijin Li, Kanu Grover, James Kent Hippler, Diyi Yang, and Amin Saberi. 2025. StorySage: Conversational Autobiography Writing Powered by a Multi-Agent Framework. In Proceedings of the 38th Annual ACM Symposium on User Interface Software and Technology (UIST ’25). Association for Computing Machinery, New York, NY, USA, Article 209, 26 pages. doi:10.1145/3746059.3747681
  60. 60.Yuying Tang, Haotian Li, Minghe Lan, Xiaojuan Ma, and Huamin Qu. 2025. Understanding Screenwriters’ Practices, Attitudes, and Future Expectations in Human-AI Co-Creation. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25). Association for Computing Machinery, New York, NY, USA, Article 26, 18 pages. doi:10.1145/3706598.3714120
  61. 61.Yuying Tang, Yuqian Sun, Ze Gao, Zhijun Pan, Zhigang Wang, Tristan Braud, Chang Hee Lee, and Ali Asadipour. 2023. AI Nüshu (Women’s scripts) - An Exploration of Language Emergence in Sisterhood. In SIGGRAPH Asia 2023 Art Gallery (Sydney, NSW, Australia) (SA ’23). Association for Computing Machinery, New York, NY, USA, Article 4, 2 pages. doi:10.1145/3610537.3622957
  62. 62.Yuying Tang, Ningning Zhang, Mariana Ciancia, and Zhigang Wang. 2024. Exploring the Impact of AI-generated Image Tools on Professional and Non-professional Users in the Art and Design Fields. In Companion Publication of the 2024 Conference on Computer-Supported Cooperative Work and Social Computing (San Jose, Costa Rica) (CSCW Companion ’24). Association for Computing Machinery, New York, NY, USA, 451–458. doi:10.1145/3678884.3681890
  63. 63.James Thomas. 2013. Script Analysis for Actors, Directors, and Designers. Routledge. doi:10.4324/9780203797020
  64. 64.Kristin Thompson. 1999. Storytelling in the New Hollywood: Understanding Classical Narrative Technique. Harvard University Press. doi:10.2307/j.ctv1nzfgkr
  65. 65.Nadine Wagener, Leon Reicherts, Nima Zargham, Natalia Bartłomiejczyk, Ava Elizabeth Scott, Katherine Wang, Marit Bentvelzen, Evropi Stefanidi, Thomas Mildner, Yvonne Rogers, and Jasmin Niess. 2023. SelVReflect: A Guided VR Experience Fostering Reflection on Personal Challenges. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23). Association for Computing Machinery, New York, NY, USA, Article 323, 17 pages. doi:10.1145/3544548.3580763
  66. 66.ShunYi Yeo, Gionnieve Lim, Jie Gao, Weiyu Zhang, and Simon Tangi Perrault. 2024. Help Me Reflect: Leveraging Self-Reflection Interface Nudges to Enhance Deliberativeness on Online Deliberation Platforms. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24). Association for Computing Machinery, New York, NY, USA, Article 806, 32 pages. doi:10.1145/3613904.3642530
  67. 67.Yonghai Yu and Yun Bi. 2010. A Study on “5W1H” User Analysis on Interaction Design of Interface. In 2010 IEEE 11th International Conference on Computer-Aided Industrial Design & Conceptual Design 1 (CAIDCD 2010, Vol. 1). IEEE, Yiwu, China, 329–332. doi:10.1109/CAIDCD.2010.5681344
  68. 68.Chao Zhang, Kexin Ju, Peter Bidoshi, Yu-Chun Grace Yen, and Jeffrey M. Rzeszotarski. 2025. Friction: Deciphering Writing Feedback into Writing Revisions through LLM-Assisted Reflection. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25). Association for Computing Machinery, New York, NY, USA, Article 935, 27 pages. doi:10.1145/3706598.3714316
  69. 69.Yu Zhang, Kexue Fu, and Zhicong Lu. 2025. RevTogether: Supporting Science Story Revision with Multiple AI Agents. In Proceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems (CHI EA ’25). Association for Computing Machinery, New York, NY, USA, Article 462, 7 pages. doi:10.1145/3706599.3719888

Citation

MLA
Tang, Y., et al. “DuoDrama: Supporting Screenplay Refinement Through LLM-Assisted Human Reflection”. arXiv, 2026, http://arxiv.org/abs/2602.05854v1.
APA
Tang, Y., Chen, X., Li, H., Xie, X., Ma, X., & Qu, H. (2026). DuoDrama: Supporting Screenplay Refinement Through LLM-Assisted Human Reflection. arXiv. http://arxiv.org/abs/2602.05854v1
Chicago
Tang, Y., X. Chen, H. Li, X. Xie, X. Ma, and H. Qu. 2026. “DuoDrama: Supporting Screenplay Refinement Through LLM-Assisted Human Reflection”. arXiv. http://arxiv.org/abs/2602.05854v1.
Harvard
Tang, Y. et al. (2026) “DuoDrama: Supporting Screenplay Refinement Through LLM-Assisted Human Reflection”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2602.05854v1.
Vancouver
1. Tang Y, Chen X, Li H, Xie X, Ma X, Qu H (2026) DuoDrama: Supporting Screenplay Refinement Through LLM-Assisted Human Reflection. arXiv

BibTeX

@article{tang2026duodrama,
  title = {DuoDrama: Supporting Screenplay Refinement Through LLM-Assisted Human Reflection},
  author = {Tang, Yuying and Chen, Xinyi and Li, Haotian and Xie, Xing and Ma, Xiaojuan and Qu, Huamin},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2602.05854v1},
  eprint = {2602.05854}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF