InCharacter: Evaluating Personality Fidelity in Role-Playing Agents through Psychological Interviews
Xintao WangYunze XiaoJen-tse HuangSiyu YuanRui XuHaoran GuoQuan TuYaying FeiZiang LengWei Wang
Proposes an interview-based psychological evaluation framework, InCharacter, that assesses the personality fidelity of large language model role-playing agents more accurately than traditional self-report methods across 14 standard personality scales.
Role-playing agents powered by large language models are increasingly used in digital assistants, gaming non-player characters, and interactive simulations. However, validating whether these agents authentically reproduce target characters has historically relied on evaluating factual knowledge and linguistic style. These existing approaches require labor-intensive, character-specific datasets and overlook the underlying behavioral, emotional, and cognitive patterns that define a character's true persona.
The article establishes a standardized framework to evaluate personality fidelity in role-playing agents using established psychological scales. It introduces an interview-based methodology named INCHARACTER to measure whether artificial agents accurately reflect the human-perceived personalities of intended personas across diverse psychological dimensions.
To overcome the flaws of traditional multiple-choice self-reporting—where artificial agents often break character or produce responses skewed by base model training data—the article implements a two-stage clinical interview approach. First, standardized personality test items are converted into open-ended questions administered in isolated conversational sessions. Second, an advanced language model acts as an expert clinician, analyzing the open-ended dialogue to derive dimensional ratings. The evaluation covered 32 diverse characters across 14 psychological scales, including the Big Five Inventory and 16Personalities, comparing agent results against a benchmark constructed from crowdsourced character profiles and 93 human expert annotations.
The investigation produced four central findings. First, top-tier role-playing agents achieve high personality fidelity, reaching an average dimensional alignment accuracy of 78.9% across all 14 psychological scales and up to 80.7% on standard personality assessments. Second, the interview-based evaluation significantly outperformed self-report baselines, yielding higher alignment accuracy and producing more distinct, character-consistent behavior across repeated tests. Third, foundational character descriptions serve as the primary driver of personality replication, with memory modules providing supplementary behavioral nuance. Finally, commercial agents engineered to please users, such as Character.ai, exhibited poor personality fidelity (roughly 52% dimensional accuracy), agreeing with user prompts 65.8% of the time rather than portraying authentic personas.
These findings demonstrate that psychological interviewing provides a scalable, character-agnostic evaluation framework that avoids the cost of creating bespoke test datasets for every new persona. For practitioners, the results clarify that high-capability general models with explicit character descriptions deliver stronger persona fidelity than heavily constrained conversational agents that prioritize agreeable, flattering user interactions.
Organizations developing interactive artificial agents should replace standard self-report benchmarks with open-ended psychological interviews to rigorously audit persona fidelity and monitor conversational safety risks, such as dark personality traits. Developers should also adopt character-specific question adaptation to avoid knowledge hallucinations when testing fictional roles. Future efforts should evaluate how agent personalities evolve across prolonged multi-turn interactions and account for temporal developments in character storylines.
- Paper: Generative Agents: Interactive Simulacra of Human Behavior, Joon Sung Park et al. (2023). Introduces generative agents simulating believable human personas and social behaviors via interview-based evaluations, providing the foundational role-playing agent paradigm evaluated in InCharacter.
- Paper: Personalizing Dialogue Agents: I have a dog, do you have pets too?, Saizheng Zhang et al. (2018). Establishes foundational techniques for conditioning conversational agents on explicit persona profiles to maintain character consistency across dialogue turns.
- Paper: CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society, Guohao Li et al. (2023). Demonstrates communicative inception prompting for multi-character role-playing agents, setting up the prompt-based agent architectures tested in the source.
- Paper: Out of One, Many: Using Language Models to Simulate Human Samples, Lisa P. Argyle et al. (2022). Explores algorithmic fidelity and the capacity of large language models to mirror specific human sub-populations and psychological/attitudinal profiles.
- Paper: Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies, Gati V. Aher et al. (2023). Pioneers the simulation of human subjects across classic psychological and behavioral experiments using large language models.
- Paper: Whose Opinions Do Language Models Reflect?, Shibani Santurkar et al. (2023). Analyzes the default opinion distributions and demographic alignments of language models when subjected to psychometric survey questions.
- Paper: A survey on large language model based autonomous agents, Lei Wang et al. (2023). Provides a comprehensive architectural taxonomy of agent profiling, memory, and evaluation strategies in LLM-based autonomous agents.
- Paper: CogBench: a large language model walks into a psychology lab, Julian Coda-Forno et al. (2024). Extends the psychological evaluation of language models from character personality scales to standardized cognitive psychology laboratory tasks measuring decision-making traits.
- Paper: Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models, Paul Röttger et al. (2024). Critically examines the methodological validity of administering forced-choice psychometric surveys to language models, directly addressing measurement nuances relevant to InCharacter's interview design.
- Paper: Is this the real life? Is this just fantasy? The Misleading Success of Simulating Social Interactions With LLMs, Xuhui Zhou et al. (2024). Investigates the validity of simulating social interactions under realistic information asymmetry, probing beyond individual persona fidelity to multi-agent interactive dynamics.
- Paper: Learning Personalized Agents from Human Feedback, Kaiqu Liang et al. (2026). Applies real-time interactive feedback loops to adapt and align personalized agent personas continuously over live interactions.
- Paper: Evaluating Very Long-Term Conversational Memory of LLM Agents, Adyasha Maharana et al. (2024). Evaluates the temporal persistence and memory of persona-grounded conversational agents over extended multi-session interactions.
