InCharacter: Evaluating Personality Fidelity in Role-Playing Agents through Psychological Interviews

Xintao WangYunze XiaoJen-tse HuangSiyu YuanRui XuHaoran GuoQuan TuYaying FeiZiang LengWei Wang

article2024ACL185 citations

Proposes an interview-based psychological evaluation framework, InCharacter, that assesses the personality fidelity of large language model role-playing agents more accurately than traditional self-report methods across 14 standard personality scales.

Listen

Role-playing agents powered by large language models are increasingly used in digital assistants, gaming non-player characters, and interactive simulations. However, validating whether these agents authentically reproduce target characters has historically relied on evaluating factual knowledge and linguistic style. These existing approaches require labor-intensive, character-specific datasets and overlook the underlying behavioral, emotional, and cognitive patterns that define a character's true persona.

The article establishes a standardized framework to evaluate personality fidelity in role-playing agents using established psychological scales. It introduces an interview-based methodology named INCHARACTER to measure whether artificial agents accurately reflect the human-perceived personalities of intended personas across diverse psychological dimensions.

To overcome the flaws of traditional multiple-choice self-reporting—where artificial agents often break character or produce responses skewed by base model training data—the article implements a two-stage clinical interview approach. First, standardized personality test items are converted into open-ended questions administered in isolated conversational sessions. Second, an advanced language model acts as an expert clinician, analyzing the open-ended dialogue to derive dimensional ratings. The evaluation covered 32 diverse characters across 14 psychological scales, including the Big Five Inventory and 16Personalities, comparing agent results against a benchmark constructed from crowdsourced character profiles and 93 human expert annotations.

The investigation produced four central findings. First, top-tier role-playing agents achieve high personality fidelity, reaching an average dimensional alignment accuracy of 78.9% across all 14 psychological scales and up to 80.7% on standard personality assessments. Second, the interview-based evaluation significantly outperformed self-report baselines, yielding higher alignment accuracy and producing more distinct, character-consistent behavior across repeated tests. Third, foundational character descriptions serve as the primary driver of personality replication, with memory modules providing supplementary behavioral nuance. Finally, commercial agents engineered to please users, such as Character.ai, exhibited poor personality fidelity (roughly 52% dimensional accuracy), agreeing with user prompts 65.8% of the time rather than portraying authentic personas.

These findings demonstrate that psychological interviewing provides a scalable, character-agnostic evaluation framework that avoids the cost of creating bespoke test datasets for every new persona. For practitioners, the results clarify that high-capability general models with explicit character descriptions deliver stronger persona fidelity than heavily constrained conversational agents that prioritize agreeable, flattering user interactions.

Organizations developing interactive artificial agents should replace standard self-report benchmarks with open-ended psychological interviews to rigorously audit persona fidelity and monitor conversational safety risks, such as dark personality traits. Developers should also adopt character-specific question adaptation to avoid knowledge hallucinations when testing fictional roles. Future efforts should evaluate how agent personalities evolve across prolonged multi-turn interactions and account for temporal developments in character storylines.

arXiv: 2310.17976
Cover for InCharacter: Evaluating Personality Fidelity in Role-Playing Agents through Psychological Interviews

Abstract

Role-playing agents (RPAs), powered by large language models, have emerged as a flourishing field of applications. However, a key challenge lies in assessing whether RPAs accurately reproduce the personas of target characters, namely their character fidelity. Existing methods mainly focus on the knowledge and linguistic patterns of characters. This paper, instead, introduces a novel perspective to evaluate the personality fidelity of RPAs with psychological scales. Overcoming drawbacks of previous self-report assessments on RPAs, we propose InCharacter, namely Interviewing Character agents for personality tests. Experiments include various types of RPAs and LLMs, covering 32 distinct characters on 14 widely used psychological scales. The results validate the effectiveness of InCharacter in measuring RPA personalities. Then, with InCharacter, we show that state-of-the-art RPAs exhibit personalities highly aligned with the human-perceived personalities of the characters, achieving an accuracy up to 80.7%.

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 2.1 Role-Playing Agents
  • 2.2 Psychological Scales
  • 3 INCHARACTER
  • 3.1 Interview
  • 3.2 Assessment
  • 4 Experimental Setup
  • 4.1 Preliminary Study
  • 4.2 Experimental Settings
  • 5 Experimental Results
  • 5.1 Personality Tests on RPAs
  • Alignment between RPAs' Measured Personalities and Characters' Labeled Personalities
  • 5.2 Personality Fidelity of Different RPAs
  • 6 Related Work
  • 7 Conclusion
  • Limitations
  • Ethical Statement
  • Acknowledgment
  • References
  • A Notation Table
  • Definition
  • Task Formulation
  • Methods
  • Metrics
  • B Psychological Scales
  • Eysenck Personality Questionnaire (Revised)
  • C Character Selection
  • D Human Annotations
  • E Implementation Details
  • E.1 RPA Inference and Post-processing
  • E.2 Fine-Tuning
  • F Additional Results
  • F.1 Interviewer LLMs
  • F.2 Self-report v.s. Interview-based Methods
  • F.3 RPAs from Different Works
  • F.4 Foundation Models for RPAs
  • F.5 Comprehensive Results on 14 Scales
  • F.6 Importance of Using Personality Scales
  • F.7 Character-specific Question Adaptation
  • F.8 Enhancing Self-report Methods with In-context Learning
  • G Prompts List
  • H Case Study
  • H.1 Visualization
  • H.2 Compliant Responses from character.ai
  • H.3 Example Responses
  • I Other Statements

Knowls

  1. Knowl 1 — The INCHARACTER Interview and Assessment Framework

    model/method

    INCHARACTER is a two-stage evaluation framework designed to assess the personality fidelity of Role-Playing Agents (RPAs) using structured psychological interviews rather than direct self-report questionnaires.

    Given a psychological scale L=(P,D,O,f)\mathcal{L} = (P, D, O, f), where PP is a set of scale items, DD is the set of personality dimensions, OO is the set of response options (e.g., Likert levels 1 to 5), and ff is the scoring scheme:

    1. Interview Phase: Each scale item p∈Pp \in P is transformed into an open-ended conversational question q∈Qq \in Q (for example, rephrasing the Big Five Inventory item "Values artistic, aesthetic experiences" into "Do you value artistic, aesthetic experiences?"). An RPA C\mathcal{C} embodying character cc is interviewed by presenting each question qq in an isolated context to prevent conversational context bias, recording its natural persona-driven response rr.

    2. Assessment Phase: Evaluates the quantitative score sds_d across each dimension d∈Dd \in D from interview transcripts using an interviewer Large Language Model (LLM) via two alternative methodologies:

    • Option Conversion (OC and d-OC): The interviewer LLM converts each individual response rr for question qq into an option a∈Oa \in O. In dimension-specific option conversion (d-OC), standard Likert options (e.g., "Agree") are replaced with dimension-descriptive options (e.g., "Extroverted"). The collected response array A=(a1,…,a∣P∣)A = (a_1, \dots, a_{|P|}) is then evaluated using the scale's standard scoring function S=f(A)S = f(A).
    • Expert Rating (ER): Drawing on structured clinical psychiatric interviews, the interviewer LLM receives dimension definitions and scoring ranges, directly predicting the dimension score sds_d from the set of (q,r)(q, r) pairs without equal-weight aggregation. ER can be executed over all dimension questions simultaneously (ERall\text{ER}_{\text{all}}) or in batches of 3–4 items (ERbatch\text{ER}_{\text{batch}}).
  2. Knowl 2 — Evaluation Metrics for RPA Personality Alignment and Consistency

    experimental setup

    RPA personality fidelity and stability are evaluated across two sets of metrics:

    1. Measured Alignment (MA): Quantifies agreement between the RPA's measured personality and human ground-truth labels for character cc. Personality scores are binarized as positive or negative on each dimension relative to the midpoint of the scoring scale, excluding ambiguous marginal dimensions:
    • Mean Absolute Error (MAE): The average absolute difference between measured and annotated scores on dimension dd, normalized by the dimension score range length to lie in [0,1][0, 1] (or reported as a percentage).
    • Dimensional Accuracy (AccDim\text{Acc}_{\text{Dim}}): The percentage of individual dimensions where the RPA's binary trait classification matches the character ground truth.
    • Full Accuracy (AccFull\text{Acc}_{\text{Full}}): The percentage of complete scales where the RPA's classifications match the character ground truth across all dimensions simultaneously.
    1. Personality Consistency (PC): Evaluates variance across repeated runs, normalized by the scale range:
    • Item-level Standard Variance (StdItem\text{Std}_{\text{Item}}): The standard deviation of converted item scores on the exact same item across independent runs.
    • Dimension-level Standard Variance (StdDim\text{Std}_{\text{Dim}}): The standard deviation of converted item scores across different items within the same dimension in a single run.
    • Score-level Standard Variance (StdScore\text{Std}_{\text{Score}}): The standard deviation of final dimension scores across independent evaluation runs.
  3. Knowl 3 — Empirical Comparison of INCHARACTER against Self-Report Baselines

    empirical result

    Evaluating 32 RPAs on the Big Five Inventory (BFI) and 16Personalities (16P) demonstrates that the interview-based INCHARACTER framework achieves significantly higher alignment with human-annotated character personalities than direct Self-Report (SR) and Chain-of-Thought Self-Report (SR-CoT).

    Method Interviewer Model The Big Five Inventory The 16 Personalities
    AccDim\text{Acc}_{\text{Dim}} (%) AccFull\text{Acc}_{\text{Full}} (%) MAE (%) AccDim\text{Acc}_{\text{Dim}} (%) AccFull\text{Acc}_{\text{Full}} (%) MAE (%)
    SR GPT-4 63.3 7.3 23.2 65.6 21.9 26.5
    SR-CoT GPT-4 67.1 9.4 22.3 66.9 24.0 25.6
    INCHARACTER (OC) GPT-4 64.3 6.2 21.6 75.5 34.4 23.1
    INCHARACTER (d-OC) GPT-4 72.2 14.6 18.6 80.2 45.8 21.2
    INCHARACTER (ERall\text{ER}_{\text{all}}) GPT-4 76.6 30.2 18.9 79.6 43.8 20.1
    INCHARACTER (ERbatch\text{ER}_{\text{batch}}) GPT-4 76.6 31.2 18.2 80.7 44.8 20.5

    Direct self-report prompts cause RPAs to default to base LLM behavioral patterns or decline answering due to non-compliant character traits, whereas open-ended interviews allow natural expression of underlying mindsets. Utilizing Expert Rating with batching (ERbatch\text{ER}_{\text{batch}}) and GPT-4 achieves the highest dimensional accuracy (76.6%76.6\% on BFI and 80.7%80.7\% on 16P) and superior personality distinctiveness across characters (average standard variance across the 5 BFI dimensions of 1.031.03 for ERbatch\text{ER}_{\text{batch}} vs. 0.710.71 for SR).

  4. Knowl 4 — Validation of LLMs as Psychological Interview Evaluators

    empirical result

    To evaluate whether LLMs can accurately replicate human psychological examiners, LLM predictions on 100 sample question-response cases from the Big Five Inventory (BFI) were compared against human expert annotations under Option Conversion (OC), Dimension-specific Option Conversion (d-OC), and Expert Rating (ER).

    Assessment Task LLM Acc. (%) Pearson's rr Spearman's ρ\rho Kendall's τ\tau
    Option Conversion (OC) Gemini 69.5 54.5 55.9 53.2
    GPT-3.5 57.5 34.6 36.2 32.4
    GPT-4 71.0 60.0 64.3 59.5
    Dimension-specific OC (d-OC) Gemini 79.0 79.6 80.6 75.9
    GPT-3.5 76.5 79.2 81.7 74.5
    GPT-4 82.0 84.7 85.3 80.6
    Expert Rating (ER, batch) Gemini 84.0 83.9 85.7 76.6
    GPT-3.5 84.0 90.6 89.9 80.4
    GPT-4 89.0 92.5 92.7 83.7

    Accuracy treats score differences from human labels <1<1 point as correct, exactly 11 point as half-correct, and >1>1 point as incorrect. GPT-4 in the Expert Rating (batch) setting achieves 89.0%89.0\% accuracy and a Spearman rank correlation of ρ=92.7%\rho = 92.7\% with human evaluators, making only 4%4\% notable errors (primarily on contradictory RPA statements). Replacing generic Likert options with dimension-descriptive options in d-OC raises GPT-4 accuracy from 71.0%71.0\% to 82.0%82.0\%.

  5. Knowl 5 — Comprehensive Benchmark across 14 Psychological Scales

    empirical result

    Evaluating state-of-the-art RPAs (GPT-3.5 foundation model with ChatHaruhi and RoleLLM persona data) using INCHARACTER (ERbatch\text{ER}_{\text{batch}}, GPT-3.5 interviewer) across 14 psychological scales (73 total dimensions) achieves an overall average dimensional accuracy (AccDim\text{Acc}_{\text{Dim}}) of 78.9%78.9\% and an average MAE of 8.1%8.1\%.

    Scale Target Construct AccDim\text{Acc}_{\text{Dim}} (%) AccFull\text{Acc}_{\text{Full}} (%) MAE (%)
    BFI Big Five Inventory (OCEAN traits) 72.0 21.9 13.4
    16P NERIS Type Explorer (MBTI dimensions) 79.3 43.8 10.9
    BSRI Bem's Sex Role Inventory (Masculine/Feminine) 85.2 74.2 3.8
    DTDD Dark Triad Dirty Dozen (Narcissism, Mach., Psych.) 75.6 51.6 9.9
    ECR-R Experiences in Close Relationships (Attachment) 68.2 53.6 9.7
    EIS Emotional Intelligence Scale 79.2 79.2 6.1
    Empathy Cognitive and Emotional Empathy 84.6 84.6 5.6
    EPQ-R Revised Eysenck Personality Questionnaire 72.4 25.0 11.6
    GSE General Self-Efficacy Scale 93.1 93.1 2.9
    ICB Implicit Culture Belief Scale 83.3 83.3 4.7
    LMS Love of Money Scale 85.9 71.4 7.1
    LOT-R Life Orientation Test Revised (Optimism) 76.2 76.2 5.1
    WLEIS Wong and Law Emotional Intelligence Scale 73.4 54.8 10.8
    CABIN Comprehensive Assessment of Basic Interests 75.5 6.3 11.8
    Average Multi-Scale Mean 78.9 58.5 8.1

    RPAs align with character traits across diverse domains, including self-efficacy (93.1%93.1\%), money motivation (85.9%85.9\%), sex roles (85.2%85.2\%), and dark triad traits (75.6%75.6\%).

  6. Knowl 6 — Impact of Persona Descriptions versus Dialogue Memories on Personality Fidelity

    empirical result

    RPAs typically combine system prompt character descriptions (DD) and retrieved dialogue/experience memories (MM). Evaluating GPT-3.5-based RPAs under different data configurations with INCHARACTER (ERbatch\text{ER}_{\text{batch}}, GPT-3.5) on 32 characters isolates their contributions:

    Persona Data Setup The Big Five Inventory The 16 Personalities
    AccDim\text{Acc}_{\text{Dim}} (%) AccFull\text{Acc}_{\text{Full}} (%) MAE (%) AccDim\text{Acc}_{\text{Dim}} (%) AccFull\text{Acc}_{\text{Full}} (%) MAE (%)
    Description only (DD) 71.3 21.9 21.1 78.5 43.8 22.0
    Memories only (MM) 71.3 18.8 21.8 71.9 31.2 26.0
    Full Setup (D+MD + M) 72.0 21.9 18.8 79.3 43.8 22.6

    Character descriptions (DD) alone establish the core personality structure, achieving alignment near the full D+MD+M configuration. Character dialogue memories (MM) alone allow RPAs to implicitly manifest traits (e.g., extraversion or openness) even when the stored dialogues do not directly discuss the interview questions, though MM-only configurations exhibit lower full-profile accuracy (AccFull\text{Acc}_{\text{Full}} of 31.2%31.2\% on 16P vs. 43.8%43.8\% for DD).

  7. Knowl 7 — Personality Fidelity across Foundation Models and Commercial RPAs

    empirical result

    Testing RPAs configured with different foundation LLMs and comparing against commercial agents (Character.ai) using INCHARACTER (ERbatch\text{ER}_{\text{batch}}, GPT-3.5) highlights substantial differences in personality reproduction:

    Foundation LLM Model Category Big Five Inventory 16 Personalities
    AccDim\text{Acc}_{\text{Dim}} (%) MAE (%) AccDim\text{Acc}_{\text{Dim}} (%) MAE (%)
    Qwen 7B General Open-Source 60.5 24.3 67.8 27.9
    OpenChat-3.5 7B General Open-Source 63.1 23.1 76.9 24.6
    Mistral-2 7B General Open-Source 66.2 21.3 68.6 26.0
    LLaMA-2-Chat 13B General Open-Source 66.9 26.8 66.9 27.7
    Mixtral 8x7B General Open-Source 68.2 20.8 71.9 25.3
    CharacterGLM 6B Role-Play Fine-Tuned 54.1 25.8 52.1 29.7
    RP-Qwen 7B Role-Play Fine-Tuned 60.5 23.8 64.5 28.6
    RP-Mistral-2 7B Role-Play Fine-Tuned 70.1 21.7 69.4 26.1
    character.ai (D∗D^*) Commercial Closed-Source 52.2 31.2 52.9 31.6
    GPT-3.5 Closed-Source API 72.0 18.8 79.3 22.6
    GPT-4 Closed-Source API 73.9 19.8 76.0 23.2

    Key findings:

    1. Proprietary agents from Character.ai perform poorly (52.2%52.2\% BFI AccDim\text{Acc}_{\text{Dim}}, 31.2%31.2\% MAE). Converting Character.ai responses into Likert options shows that they agree or strongly agree with user questions 65.8%65.8\% of the time (and disagree only 17.1%17.1\%), demonstrating a compliance/sycophancy bias that overrides authentic character personas.
    2. GPT-3.5 and GPT-4 achieve the highest alignment (72.0–73.9%72.0\text{--}73.9\% on BFI, 76.0–79.3%76.0\text{--}79.3\% on 16P).
    3. Fine-tuning open models on role-playing dialogue datasets (RP-Mistral-2 7B) marginally improves personality accuracy while substantially reducing generation anomalies (lowering un-immersed responses from 0.8%0.8\% to 0.1%0.1\% and repetition from 3.5%3.5\% to 0.4%0.4\%).
  8. Knowl 8 — Character-Specific Question Adaptation in Psychological Scales

    model/method

    Standard psychometric questions frequently incorporate concepts outside the knowledge scope of specific fictional or historical characters (e.g., modern terms like "movies", "phones", or "deadlines" presented to fantasy characters from Harry Potter or Genshin Impact), which can trigger character hallucinations.

    The Character-Specific Question Adaptation strategy uses an LLM (e.g., GPT-4) to detect out-of-universe concepts in scale items and minimally adapt them into character-consistent equivalents (e.g., converting "books and movies" into "tales and legends" for wizarding world characters).

    Applying this adaptation to the 16Personalities scale (where 7 of 60 items contained domain mismatches, leading to adaptation of 60.7%60.7\% of eligible item-character pairs) improves measurement accuracy with INCHARACTER (ERbatch\text{ER}_{\text{batch}}, GPT-4):

    • Without Adaptation: AccDim=80.7%\text{Acc}_{\text{Dim}} = 80.7\%, AccFull=44.8%\text{Acc}_{\text{Full}} = 44.8\%, MAE=20.5%\text{MAE} = 20.5\%
    • With Adaptation: AccDim=81.8%\text{Acc}_{\text{Dim}} = 81.8\%, AccFull=46.9%\text{Acc}_{\text{Full}} = 46.9\%, MAE=20.1%\text{MAE} = 20.1\%
  9. Knowl 9 — Performance Degradation of Self-Report under In-Context Learning

    empirical result

    Augmenting direct Self-Report (SR) or Chain-of-Thought Self-Report (SR-CoT) questionnaires with 3-shot in-context learning (ICL) demonstrations degrades personality fidelity across RPAs.

    Method The Big Five Inventory The 16 Personalities
    AccDim\text{Acc}_{\text{Dim}} (%) AccFull\text{Acc}_{\text{Full}} (%) MAE (%) AccDim\text{Acc}_{\text{Dim}} (%) AccFull\text{Acc}_{\text{Full}} (%) MAE (%)
    Without ICL
    SR 63.3 7.3 23.2 65.6 21.9 26.5
    SR-CoT 67.1 9.4 22.3 66.9 24.0 25.6
    With ICL
    SR-ICL 56.1 3.1 27.5 59.5 12.5 30.4
    SR-CoT-ICL 60.5 3.1 24.6 64.9 25.0 27.3

    Because standardized few-shot demonstration examples cannot be customized for every fictional character, introducing static few-shot examples presents external persona behaviors that interfere with the RPA's target persona prompt, increasing score distortion and lowering AccDim\text{Acc}_{\text{Dim}} by 6.6–7.26.6\text{--}7.2 percentage points on BFI and 2.0–6.12.0\text{--}6.1 percentage points on 16P.

  10. Knowl 10 — Limitations of INCHARACTER and RPA Personality Evaluation

    limitation

    The methodology for assessing personality fidelity in role-playing agents is constrained by two primary factors:

    1. Interviewer Model Error and Bias: INCHARACTER relies on interviewer LLMs to convert open responses or predict expert ratings. Errors or internal biases of the interviewer LLM (which generates incorrect ratings in approximately 4%4\% of cases for GPT-4 during expert rating evaluation) can propagate into test scores, potentially underestimating true RPA personality fidelity.
    2. Static Personality Assumption for Dynamic Personas: Ground-truth annotations represent fixed personality snapshots. In practice, literary and cinematic characters (e.g., James Bond across multiple decades) undergo substantial psychological evolution throughout narrative arcs. A static single-vector annotation fails to capture the progressive personality dynamics of evolving characters.

Coverage note — None was omitted; all primary frameworks, empirical benchmarks, comparative experiments, ablation studies, and stated limitations are included.

References

  1. 1.Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. 2023. Qwen technical report. ArXiv preprint, abs/2309.16609.
  2. 2.Murray R Barrick and Michael K Mount. 1991. The big five personality dimensions and job performance: a meta-analysis. Personnel psychology, 44(1):1–26.
  3. 3.Sandra L Bem. 1981. Bem sex role inventory. Journal of personality and social psychology.
  4. 4.Bojana Bodroza, Bojana M Dinic, and Ljubisa Bojic. 2023. Personality testing of gpt-3: Limited temporal reliability, but highlighted social desirability of gpt-3’s personality instruments results. ArXiv preprint, abs/2306.04308.
  5. 5.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  6. 6.Patrick Butlin, Robert Long, Eric Elmoznino, Yoshua Bengio, Jonathan Birch, Axel Constant, George Deane, Stephen M Fleming, Chris Frith, Xu Ji, et al. 2023. Consciousness in artificial intelligence: Insights from the science of consciousness. ArXiv preprint, abs/2308.08708.
  7. 7.Julian Coda-Forno, Kristin Witte, Akshay K Jagadish, Marcel Binz, Zeynep Akata, and Eric Schulz. 2023. Inducing anxiety in large language models increases exploration and bias. ArXiv preprint, abs/2304.11111.
  8. 8.Jacob. Cohen. 1968. Weighed kappa: Nominal scale agreement with provision for scaled disagreement or partial credit. Psychological Bulletin, 70(4):213–220.
  9. 9.Pim Cuijpers, Juan Li, Stefan G Hofmann, and Gerhard Andersson. 2010. Self-reported versus clinician-rated symptoms of depression as outcome measures in psychotherapy research on depression: a meta-analysis. Clinical psychology review, 30(6):768–778.
  10. 10.Haodong Duan, Jueqi Wei, Chonghua Wang, Hongwei Liu, Yixiao Fang, Songyang Zhang, Dahua Lin, and Kai Chen. 2023. Botchat: Evaluating llms’ capabilities of having multi-turn dialogues. ArXiv preprint, abs/2310.13650.
  11. 11.Michael B First. 2014. Structured clinical interview for the dsm (scid). The encyclopedia of clinical psychology, pages 1–6.
  12. 12.Jingsheng Gao, Yixin Lian, Ziyi Zhou, Yuzhuo Fu, and Baoyuan Wang. 2023. LiveChat: A large-scale personalized dialogue dataset automatically constructed from live streaming. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 15387–15405, Toronto, Canada. Association for Computational Linguistics.
  13. 13.Jochen Hartmann, Jasper Schwenzow, and Maximilian Witte. 2023. The political ideology of conversational ai: Converging evidence on chatgpt’s pro-environmental, left-libertarian orientation. Available at SSRN 4316084.
  14. 14.Jen-tse Huang, Man Ho Lam, Eric John Li, Shujie Ren, Wenxuan Wang, Wenxiang Jiao, Zhaopeng Tu, and Michael R Lyu. 2023a. Emotionally numb or empathetic? evaluating how llms feel using emotionbench. ArXiv preprint, abs/2308.03656.
  15. 15.Jen-tse Huang, Wenxuan Wang, Man Ho Lam, Eric John Li, Wenxiang Jiao, and Michael R Lyu. 2023b. Chatgpt an enfj, bard an istj: Empirical study on personalities of large language models. ArXiv preprint, abs/2305.19926.
  16. 16.Jen-tse Huang, Wenxuan Wang, Man Ho Lam, Eric John Li, Wenxiang Jiao, and Michael R Lyu. 2023c. Revisiting the reliability of psychological scales on large language models. ArXiv preprint, abs/2305.19926.
  17. 17.Jen-tse Huang, Wenxuan Wang, Eric John Li, Man Ho Lam, Shujie Ren, Youliang Yuan, Wenxiang Jiao, Zhaopeng Tu, and Michael R Lyu. 2024. Who is chatgpt? benchmarking llms’ psychological portrayal using psychobench. In Proceedings of the Twelfth International Conference on Learning Representations.
  18. 18.Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023. Mistral 7b. ArXiv preprint, abs/2310.06825.
  19. 19.Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al. 2024. Mixtral of experts. ArXiv preprint, abs/2401.04088.
  20. 20.Guangyuan Jiang, Manjie Xu, Song-Chun Zhu, Wenjuan Han, Chi Zhang, and Yixin Zhu. 2022. Mpi: Evaluating and inducing personality in pre-trained language models. ArXiv preprint, abs/2206.07550.
  21. 21.Saketh Reddy Karra, Son The Nguyen, and Theja Tulabandhula. 2022. Estimating the personality of white-box language models. ArXiv preprint, abs/2204.12000.
  22. 22.Maurice G Kendall. 1938. A new measure of rank correlation. Biometrika, 30(1/2):81–93.
  23. 23.Cheng Li, Ziang Leng, Chenxi Yan, Junyi Shen, Hao Wang, Weishi MI, Yaying Fei, Xiaoyang Feng, Song Yan, HaoSheng Wang, et al. 2023. Chatharuhi: Reviving anime character in reality via large language model. ArXiv preprint, abs/2308.09597.
  24. 24.Xingxuan Li, Yutong Li, Shafiq Joty, Linlin Liu, Fei Huang, Lin Qiu, and Lidong Bing. 2022. Does gpt-3 demonstrate psychopathy? evaluating large language models from a psychological perspective. ArXiv preprint, abs/2212.10529.
  25. 25.Marilù Miotto, Nicola Rossberg, and Bennett Kleinberg. 2022. Who is GPT-3? an exploration of personality, values and demographics. In Proceedings of the Fifth Workshop on Natural Language Processing and Computational Social Science (NLP+CSS), pages 218–227, Abu Dhabi, UAE. Association for Computational Linguistics.
  26. 26.Danko Nikolić. 2010. The brain is a context machine. Review of psychology, 17(1):33–38.
  27. 27.OpenAI. 2022. Openai: Introducing chatgpt.
  28. 28.OpenAI. 2023. Gpt-4 technical report.
  29. 29.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems, volume 35, pages 27730–27744. Curran Associates, Inc.
  30. 30.Keyu Pan and Yawen Zeng. 2023. Do llms possess a personality? making the mbti test an amazing evaluation for large language models. ArXiv preprint, abs/2307.16180.
  31. 31.Karl Pearson. 1920. Notes on the history of correlation. Biometrika, 13(1):25–45.
  32. 32.Peter Romero, Stephen Fitz, and Teruo Nakatsuma. 2023. Do gpt language models suffer from split personality disorder? the advent of substrate-free psychometrics. ResearchSquare preprint.
  33. 33.A John Rush, William Hiser, and Donna E Giles. 1987. A comparison of self-reported versus clinician-related symptoms in depression. The Journal of clinical psychiatry, 48(6):246–248.
  34. 34.Jérôme Rutinowski, Sven Franke, Jan Endendyk, Ina Dormuth, Moritz Roidl, and Markus Pauly. 2024. The self-perception and political biases of chatgpt. Human Behavior and Emerging Technologies, 2024(1):7115633.
  35. 35.Mustafa Safdari, Greg Serapio-García, Clément Crepy, Stephen Fitz, Peter Romero, Luning Sun, Marwa Abdulhai, Aleksandra Faust, and Maja Matarić. 2023. Personality traits in large language models. ArXiv preprint, abs/2307.00184.
  36. 36.Yunfan Shao, Linyang Li, Junqi Dai, and Xipeng Qiu. 2023a. Character-LLM: A trainable agent for role-playing. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 13153–13187, Singapore. Association for Computational Linguistics.
  37. 37.Yunfan Shao, Linyang Li, Junqi Dai, and Xipeng Qiu. 2023b. Character-LLM: A trainable agent for role-playing. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 13153–13187, Singapore. Association for Computational Linguistics.
  38. 38.Vera Sorin, Danna Brin, Yiftach Barash, Eli Konen, Alexander Charney, Girish Nadkarni, and Eyal Klang. 2023. Large language models (llms) and empathy-a systematic review. medRxiv, pages 2023–08.
  39. 39.Charles Spearman. 1961. The proof and measurement of association between two things.
  40. 40.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023. Llama: Open and efficient foundation language models.
  41. 41.Timothy J Trull, Thomas A Widiger, J David Useda, Jay Holcomb, Bao-Tran Doan, Seth R Axelrod, Barry L Stern, and Beth S Gershuny. 1998. A structured interview for the assessment of the five-factor model of personality. Psychological assessment, 10(3):229.
  42. 42.Quan Tu, Chuanqi Chen, Jinpeng Li, Yanran Li, Shuo Shang, Dongyan Zhao, Ran Wang, and Rui Yan. 2023. Characterchat: Learning towards conversational ai with personalized social support. ArXiv preprint, abs/2308.10278.
  43. 43.Quan Tu, Shilong Fan, Zihang Tian, and Rui Yan. 2024. Charactereval: A chinese benchmark for role-playing conversational agent evaluation. ArXiv preprint, abs/2401.01275.
  44. 44.Rudolf Uher, Roy H Perlis, Anna Placentino, Mojca Zvezdana Dernovšek, Neven Henigsberg, Ole Mors, Wolfgang Maier, Peter McGuffin, and Anne Farmer. 2012. Self-report and clinician-rated measures of depression severity: can one replace the other? Depression and anxiety, 29(12):1043–1049.
  45. 45.Guan Wang, Sijie Cheng, Xianyuan Zhan, Xiangang Li, Sen Song, and Yang Liu. 2024. Openchat: Advancing open-source language models with mixed-quality data. In The Twelfth International Conference on Learning Representations.
  46. 46.Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar. 2023a. Voyager: An open-ended embodied agent with large language models. arXiv preprint arXiv: Arxiv-2305.16291.
  47. 47.Zekun Moore Wang, Zhongyuan Peng, Haoran Que, Jiaheng Liu, Wangchunshu Zhou, Yuhan Wu, Hongcheng Guo, Ruitong Gan, Zehao Ni, Man Zhang, et al. 2023b. Rolellm: Benchmarking, eliciting, and enhancing role-playing abilities of large language models. ArXiv preprint, abs/2310.00746.
  48. 48.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou. 2022. Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems, volume 35, pages 24824–24837. Curran Associates, Inc.
  49. 49.Jinfeng Zhou, Zhuang Chen, Dazhen Wan, Bosi Wen, Yi Song, Jifan Yu, Yongkang Huang, Libiao Peng, Jiaming Yang, Xiyao Xiao, et al. 2023. Characterglm: Customizing chinese conversational ai characters with large language models. ArXiv preprint, abs/2311.16832.

Citation

MLA
Wang, X., et al. “InCharacter: Evaluating Personality Fidelity in Role-Playing Agents Through Psychological Interviews”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 1840–73, https://doi.org/10.18653/v1/2024.acl-long.102.
APA
Wang, X., Xiao, Y., Huang, J.-. tse ., Yuan, S., Xu, R., Guo, H., Tu, Q., Fei, Y., Leng, Z., Wang, W., Chen, J., Li, C., & Xiao, Y. (2024). InCharacter: Evaluating Personality Fidelity in Role-Playing Agents through Psychological Interviews. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1840–1873. https://doi.org/10.18653/v1/2024.acl-long.102
Chicago
Wang, X., Y. Xiao, J.-. tse . Huang, et al. 2024. “InCharacter: Evaluating Personality Fidelity in Role-Playing Agents Through Psychological Interviews”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1840–73. https://doi.org/10.18653/v1/2024.acl-long.102.
Harvard
Wang, X. et al. (2024) “InCharacter: Evaluating Personality Fidelity in Role-Playing Agents through Psychological Interviews”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 1840–1873. Available at: https://doi.org/10.18653/v1/2024.acl-long.102.
Vancouver
1. Wang X, Xiao Y, Huang J-tse, et al (2024) InCharacter: Evaluating Personality Fidelity in Role-Playing Agents through Psychological Interviews. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 1840–1873

BibTeX

@inproceedings{wang-etal-2024-incharacter,
    title = "{I}n{C}haracter: Evaluating Personality Fidelity in Role-Playing Agents through Psychological Interviews",
    author = "Wang, Xintao  and
      Xiao, Yunze  and
      Huang, Jen-tse  and
      Yuan, Siyu  and
      Xu, Rui  and
      Guo, Haoran  and
      Tu, Quan  and
      Fei, Yaying  and
      Leng, Ziang  and
      Wang, Wei  and
      Chen, Jiangjie  and
      Li, Cheng  and
      Xiao, Yanghua",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.102/",
    doi = "10.18653/v1/2024.acl-long.102",
    pages = "1840--1873"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/