Mimetic Alignment with ASPECT: Evaluation of AI-inferred Personal Profiles

Ruoxi ShangDan MarshallEdward CutrellDenae Ford

article2026arXiv0 citations

Presents ASPECT, a pipeline that infers validated communication profiles from workplace behavioral data without per-person training, allowing language models to accurately mirror individual communication styles while giving users transparent evidence to review and correct their representations.

Listen

As artificial intelligence agents increasingly communicate on behalf of individuals in workplace settings, capturing a person's authentic communication style remains a major challenge. Existing personalization methods typically depend on costly per-person fine-tuning, produce generic outputs from shallow persona prompts, or optimize outputs for general human preference rather than individual communication style. The article introduces and evaluates ASPECT (Automated Social Psychometric Evaluation of Communication Traits), a prompt-based pipeline that directs large language models to construct structured, evidence-grounded communication profiles from observed workplace interaction data without requiring individual model training.

The article evaluates how accurately an automated system can infer personal communication styles from workplace interaction data and examines whether these inferred profiles produce appropriate, socially aligned communication in workplace scenarios. The approach builds upon the Communication Styles Inventory, a validated 92-item psychometric instrument spanning six communication dimensions. The system compresses 90 days of local chat and meeting transcripts by extracting conversational evidence facet by facet and generating item-level scores with traceable rationales. The evaluation involved a case study of 20 professionals across diverse roles within a single organization, encompassing 1,840 paired item ratings and 600 blinded scenario evaluations comparing profiled outputs against self-report and generic baselines.

The findings demonstrate three core outcomes. First, the automated profiling pipeline achieved moderate overall alignment with participant self-assessments (mean absolute error of 1.39 on a 5-point scale and a within-person rank correlation of 0.39), successfully preserving the relative profile shape and distinguishing overt interpersonal tones, though it exhibited systematic positive biases by over-rating structural traits like Preciseness. Second, in downstream workplace scenarios, responses generated using the inferred profiles were preferred on aggregate, capturing 42.5% of first-place rankings compared to 32.5% for generic responses and 25.0% for self-report baselines, with significantly higher mean alignment ratings. Third, interactive profile reviews shifted profiling into a collaborative negotiation: participants revised their self-ratings in approximately 17.7% of facet evaluations after reviewing linked behavioral evidence, reconciling gaps between their aspirational self-image and observed workplace behavior.

These results indicate that grounding personalization in concrete behavioral evidence and psychometric frameworks outperforms self-reported profiles, while avoiding the compute costs and opacity of fine-tuning. However, strong individual differences in preference emerged, and participants reported an uncanny valley effect where flawed personalization felt more unsettling than a standard, neutral generic response. Furthermore, participants actively drew boundaries between their natural styles and context-specific professional personas, demonstrating that effective digital representation requires context-aware boundaries rather than simple behavioral averages.

Organizations developing or deploying conversational proxies should implement human-in-the-loop review mechanisms that allow users to inspect behavioral evidence, calibrate scores, and establish context-specific persona boundaries before deployment. Automated shrinkage or calibration corrections should also be applied to counter predictable model biases on structural communication traits. Decision-makers should note key limitations: the findings reflect a sample of 20 technical and professional participants within a single enterprise, relying on 90 days of workplace text without multimodal cues. Further longitudinal pilot studies across broader industries and communication channels are recommended before deploying autonomous representative agents in high-stakes environments.

arXiv: 2603.26922
Cover for Mimetic Alignment with ASPECT: Evaluation of AI-inferred Personal Profiles

Abstract

AI agents that communicate on behalf of individuals need to capture how each person actually communicates, yet current approaches either require costly per-person fine-tuning, produce generic outputs from shallow persona descriptions, or optimize preferences without modeling communication style. We present ASPECT (Automated Social Psychometric Evaluation of Communication Traits), a pipeline that directs LLMs to assess constructs from a validated communication scale against behavioral evidence from workplace data, without per-person training. In a case study with 20 participants (1,840 paired item ratings, 600 scenario evaluations), ASPECT-generated profiles achieved moderate alignment with self-assessments, and ASPECT-generated responses were preferred over generic and self-report baselines on aggregate, with substantial variation across individuals and scenarios. During the profile review phase, linked evidence helped participants identify mischaracterizations, recalibrate their own self-ratings, and negotiate context-appropriate representations. We discuss implications for building inspectable, individually scoped communication profiles that let individuals control how agents represent them at work.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 From communication support to AI representation
  • 2.2 Building structured profiles of individuals
  • 3 System Design
  • 3.1 Conceptual Model
  • 3.2 Profile Pipeline Development
  • 3.3 User Profile Evaluation
  • 4 Study Design
  • 4.1 Evaluation in a Workplace Setting
  • 4.2 Participants
  • 4.3 Procedure
  • 5 Findings
  • 5.1 RQ1: Inference Alignment
  • 5.1.1 Statistical Results
  • 5.1.2 Auditing as Bidirectional Alignment
  • 5.1.3 Where ASPECT matched participants’ self-assessments
  • 5.1.4 Correct inference and self-rating bias emerge upon reflection
  • 5.1.5 Calibrations are needed to find middle ground
  • 5.1.6 Sources of mischaracterizations and errors
  • 5.2 RQ2: From Profiles to Social Performance
  • 5.2.1 Statistical Results
  • 5.2.2 Individual Variation Dominates Aggregate Patterns
  • 5.2.3 Signals of Condition-based Preferences
  • 6 Discussion
  • 6.1 What we learned about building data-grounded profiles
  • 6.2 Representation boundaries and sources of mischaracterization
  • 6.3 Using ASPECT framework in practice
  • 7 Limitations
  • 8 Conclusion
  • References
  • A Evidence Extraction Schema
  • B Communication Styles Inventory (CSI) - Complete Scale
  • C Scenario templates used in study
  • D Data Collected
  • D.1 Raw data from participants.
  • D.2 Derived data (model-produced).
  • E Figures

Knowls

  1. Knowl 1 — ASPECT Automated Communication Profiling Pipeline

    model/method

    The Automated Social Psychometric Evaluation of Communication Traits (ASPECT) pipeline constructs structured, evidence-grounded communication style profiles of individuals from unstructured workplace communication logs (e.g., meeting transcripts, direct messages, group chats) without requiring per-person model fine-tuning.

    The pipeline operates in two sequential phases using an LLM reasoning model (e.g., OpenAI o1):

    1. Phase 1: Construct-Guided Evidence Extraction: Communication data is partitioned into token-budgeted batches. The LLM scans the data construct-by-construct across the facets of a validated psychometric instrument (such as the Communication Styles Inventory). For each facet, the model receives the construct definition, associated scale items, and a communication batch, acting as an observer to extract 2 to 5 concrete behavioral instances. Each evidence instance contains a structured context summary (situational background, social dynamics, communication setting, and behavioral analysis) paired with a 2-to-5-turn conversational excerpt.

    2. Phase 2: Inferred Assessment from Item-Level Scoring: The LLM evaluates each individual item in the psychometric instrument independently. It receives the item text alongside the aggregated evidence pool extracted for its parent facet in Phase 1. The model outputs an integer rating on a 1--5 Likert scale and a natural language rationale grounding the score in the extracted evidence. If no behavioral evidence was identified for a facet, the pipeline assigns a default baseline rating reflecting the absence of the behavior.

  2. Knowl 2 — Evidence Extraction Data Schema for Psychometric Profiling

    definition

    In the evidence extraction phase of the ASPECT pipeline, each behavioral instance extracted from communication logs is represented as a structured JSON object containing a contextual summary and a conversational dialogue snippet:

    {
      "situational_background": "Meeting purpose, topic, and timing",
      "social_dynamics": "Participants involved and their roles relative to the target user",
      "communication_setting": "1-on-1 vs. group; formal vs. informal; planned vs. spontaneous; stakes level",
      "behavioral_analysis": "How the setting shaped the target user's manifestation of the facet, and potential differences elsewhere",
      "contextual_significance": "Why this instance demonstrates the facet given the situation and dynamics",
      "conversational_excerpt": [
        {"speaker": "TargetUser", "message": "Dialogue turn text..."},
        {"speaker": "OtherParty", "message": "Response turn text..."},
        {"speaker": "TargetUser", "message": "Follow-up turn text..."}
      ]
    }
    

    This schema ensures that downstream psychometric trait scoring remains traceable to concrete observational data.

  3. Knowl 3 — Psychometric Agreement Between AI-Inferred Communication Profiles and Self-Assessments

    empirical result

    In a study of N=20N = 20 workplace participants evaluating 92 items of the Communication Styles Inventory (CSI) across 23 facets and 6 dimensions (1,840 paired item ratings), initial profiles inferred by ASPECT showed moderate alignment with participants' self-assessments:

    • Overall Item-Level Metrics: Exact numeric matches occurred on 23.8%23.8\% of items. Mean Absolute Error (MAE) was 1.391.39 on a 1--5 scale (95%95\% CI [1.34,1.45][1.34, 1.45]). Agreement beyond chance was fair, with weighted κ=0.34\kappa = 0.34. Mean within-person rank correlation across items was ρ=0.39\rho = 0.39 (95%95\% CI [0.31,0.44][0.31, 0.44]), increasing to a median ρ=0.55\rho = 0.55 at the facet level and ρ=0.72\rho = 0.72 at the dimension level.
    • Inter-Rater Reliability: Across items, two-rater reliability was fair (ICC(A,1)=0.345ICC(A,1) = 0.345 for absolute agreement, ICC(C,1)=0.349ICC(C,1) = 0.349 for consistency).
    • Dimension-Level Variation:
    CSI Dimension MAE Bias (Signed Diff) Rank Correlation (ρ\rho)
    Verbal Aggressiveness 0.59 -0.40 0.39
    Emotionality 0.74 -0.05 0.20
    Questioningness 0.74 +0.23 0.26
    Impression Manipulativeness 0.82 -0.47 -0.03
    Expressiveness 1.02 +0.99 0.48
    Preciseness 1.69 +1.69 0.18

    ASPECT accurately recovered overt interpersonal tone (Verbal Aggressiveness, Emotionality) while exhibiting large positive bias on structural/organizational traits (Preciseness, Expressiveness).

  4. Knowl 4 — Downstream Social Performance and Preference Across Agent Response Conditions

    empirical result

    In a blinded, within-subject triad evaluation across 10 workplace scenario templates (N=20N = 20 participants, 600 evaluations total), responses generated from ASPECT-inferred profiles (Profiled) outperformed unpersonalized baselines (Generic) and self-reported trait baselines (Self-Report):

    • Win Rates (First-Place Ranks): Profiled responses were ranked first in 42.5%42.5\% of scenarios (85/200, 95%95\% CI [35.9%,49.4%][35.9\%, 49.4\%]), compared to 32.5%32.5\% for Generic (65/200) and 25.0%25.0\% for Self-Report (50/200, 95%95\% CI [19.5%,31.4%][19.5\%, 31.4\%]).
    • Mean Rank Order: Profiled (M=1.84,SD=0.82M = 1.84, SD = 0.82) was ranked best, followed by Generic (M=2.00,SD=0.81M = 2.00, SD = 0.81) and Self-Report (M=2.15,SD=0.79M = 2.15, SD = 0.79). A Friedman test revealed a significant main effect of condition (χ2=9.31,p=.0095,Kendall’s W=0.023\chi^2 = 9.31, p = .0095, \text{Kendall's } W = 0.023). Post-hoc Wilcoxon signed-rank tests with Holm--Bonferroni correction confirmed a significant advantage for Profiled over Self-Report (p=.0067,r=.22p = .0067, r = .22).
    • Alignment Ratings (1--5 Likert Scale): Profiled achieved a mean rating of 3.333.33 (SD=1.29SD = 1.29), Generic achieved 3.093.09 (SD=1.28SD = 1.28), and Self-Report achieved 2.952.95 (SD=1.19SD = 1.19). Linear mixed-effects modeling showed Profiled was rated significantly higher than Generic (β=0.24,p=.045\beta = 0.24, p = .045), with pairwise Cohen's d=0.30d = 0.30 versus Self-Report and d=0.19d = 0.19 versus Generic.

    Providing psychometric self-ratings alone without behavioral grounding (Self-Report) yielded the worst overall alignment, below unpersonalized generic generation.

  5. Knowl 5 — Typology of Root Causes for AI-Self Communication Profile Misalignments

    definition

    Divergences between LLM-inferred communication profiles and individual self-assessments stem from five distinct root causes categorized across data limits, methodological issues, and foundational LLM constraints:

    1. T1: Coverage and Observability Gaps (Data limitation): Ratings are derived from a narrow digital sample (e.g., recorded meetings, enterprise chat), omitting unrecorded pre-meeting small talk, offline/in-person conversations, internal emotional states, and rare-but-salient interpersonal events.
    2. T2: Situational Norms Misread as Traits (Method issue): Meeting-, role-, or task-mandated communication behaviors (such as a project manager structuring a meeting, leading presentations, or conducting risk triage) are incorrectly modeled as fixed, dispositional personality traits.
    3. T3: Tone and Valence Misinterpretation (LLM limitation): The language model takes pragmatic nuances literally, misinterpreting sarcasm, playful banter among familiar peers, emojis, or routine politeness/praise as genuine aggression, anxiety, or manipulation.
    4. T4: Construct and Item Misalignment (Method issue): Discrepancies between participants' intuitive definitions of terms and the psychometric instrument's operational scope (e.g., interpreting constructive intellectual challenges as interpersonal argumentativeness, or equating general emotion with sentimentality).
    5. T5: Evidence Use and Scoring Integrity Problems (Method/Tooling issue): Model scoring errors including score inflation from a single isolated incident, misattributing words spoken by other meeting attendees to the target user due to transcript speaker diarization errors, or inconsistent handling of missing evidence.
  6. Knowl 6 — Bidirectional Calibration and Reflection Dynamics in Profile Auditing

    empirical result

    When participants inspected side-by-side comparisons of their self-ratings against ASPECT's inferred ratings along with linked behavioral evidence and rationales across 411 facet-level audits, the review functioned as a bidirectional calibration tool:

    • Audit Decision Breakdown:
      • Totally Aligned: n=141n = 141 (34.3%34.3\%)
      • Misalign--Disapprove: n=142n = 142 (34.5%34.5\%)
      • Misalign--Middle Ground: n=42n = 42 (10.2%10.2\%)
      • Misalign--Approve AI: n=31n = 31 (7.5%7.5\%)
      • Unsure / Not Mentioned: n=55n = 55 (13.4%13.4\%)

    In 17.7%17.7\% (73/41173/411) of facet evaluations, participants revised their own assessment: 7.5%7.5\% fully adopted the AI score after being confronted with concrete behavioral instances they had forgotten or minimized, and 10.2%10.2\% negotiated an explicit middle ground.

    Discrepancies often revealed a gap between participants' internal self-concept (aspirational or comfort level) and their enacted workplace persona (effortful, audience-adapted professional habits), demonstrating that pre-audit statistical alignment metrics represent a conservative lower bound.

  7. Knowl 7 — Heterogeneity and Clustering of User Personalization Preferences

    data/table

    Participants exhibited strong individual differences in preferred agent response conditions, with overall inter-participant agreement on response quality being near zero (Kendall's W=0.077W = 0.077). Random slopes analysis of condition effects confirmed substantial variance (SD=0.87SD = 0.87 for Self-Report, SD=0.94SD = 0.94 for Profiled).

    Cluster NN Mean Win Rate Mean Rating Margin Participants
    Prefers Profiled 10 0.60 +0.73 P2, P4, P6, P10, P12, P13, P15, P16, P17, P18
    Prefers Generic 4 0.57 +0.75 P1, P5, P8, P9
    Prefers Self-Report 2 0.70 +0.85 P14, P19
    Mixed / No Clear Preference 4 – +0.05 P3, P7, P11, P20

    Participants were assigned to clusters if a condition won at least two of three summary metrics (win rate margin ≥0.20\ge 0.20, rating margin ≥0.25\ge 0.25, or rank margin ≥0.20\ge 0.20) or won one metric by a strong margin (win rate ≥0.30\ge 0.30, rating ≥0.40\ge 0.40, rank ≥0.30\ge 0.30). Half of the participants (50%50\%) strongly favored data-profiled responses, while 20%20\% preferred unpersonalized generic responses, highlighting the necessity for user-configurable personalization depth.

  8. Knowl 8 — Scenario-Specific Contextual Drivers of Agent Personalization Preference

    empirical result

    Cross-analysis of 10 APRACE-parameterized workplace communication scenarios (varying hierarchy, familiarity, purpose, stakes, formality, and mode) revealed systematic task-dependent preferences for personalized versus generic agent responses:

    • Profiled Preference Scenarios: Profiled responses won the majority in 5 of 10 scenarios: explaining work to a distant peer (Scenario 1), responding to peer challenge in chat (Scenario 3), planning an initiative with a manager (Scenario 4), sharing a standup update (Scenario 7), and discussing credit attribution (Scenario 9). These tasks have defined objectives where personal communication style (structuring habits, pacing, balance of firmness and support) is critical to success.
    • Generic Preference Scenarios: Generic responses won in 4 of 10 scenarios: team check-in openings (Scenario 1), handling last-minute schedule changes (Scenario 5), acknowledging unexpected process shifts (Scenario 8), and planning team celebrations (Scenario 10). In these settings, communication goals involve light coordination or delicate emotional tone where neutral, unpersonalized phrasing is considered safer.
    • Self-Report Preference Scenario: Won in only 1 scenario (Scenario 6: informal catch-up with an influential colleague) with a minimal margin.
  9. Knowl 9 — The Uncanny Valley of Mimetic Agent Representation

    theoretical result

    User tolerance for AI persona imitation exhibits an 'uncanny valley' threshold effect in social representation: imperfect personalization is perceived as significantly worse than no personalization.

    When an AI agent communicates generically, users perceive it as an unpersonalized, neutral tool. However, when an agent attempts to mimic a specific individual's voice but incorporates subtle mischaracterizations (e.g., expressing out-of-character enthusiasm, improper valence, or misinterpreting professional norms), users experience the output as a distorted, uncomfortable identity representation. To be acceptable to the user being represented, mimetic AI agents must either achieve very high fidelity or clearly default to neutral, unpersonalized communication.

  10. Knowl 10 — Systematic Over-Rating Bias in Text-Based Psychometric Profiling

    limitation

    LLM-based psychometric profiling from workplace interaction data introduces predictable systematic positive bias on specific traits, prominently Preciseness (extbias=+1.69,MAE=1.69,ICC≈0 ext{bias} = +1.69, \text{MAE} = 1.69, ICC \approx 0) and Expressiveness (extbias=+0.99,MAE=1.02 ext{bias} = +0.99, \text{MAE} = 1.02).

    This inflation is driven by two underlying mechanisms:

    1. Asymmetry in Default Absence Scoring: When no behavioral evidence is detected, assigning a baseline low score works well for traits whose absence in workplace logs is meaningful (e.g., absence of aggressive text correctly maps to low Angriness), but fails for traits that are inherently unobservable in textual records (e.g., internal deliberation before speaking under Thoughtfulness).
    2. Professional Norm Filtering: Workplace communication channels enforce organizational norms of purposeful, structured, and edited messaging. The model mistakes these ubiquitous situational workplace norms for enduring, dispositional traits of the individual.

Coverage note — Omitted the full 92-item list of the Communication Styles Inventory and the detailed cell-by-cell breakdown of the 11 APRACE parameters for all 10 scenario templates, as these represent standardized instruments and structural templates rather than core scientific findings.

References

  1. 1.Character.AI 2025. Character.AI. Character.AI. https://character.ai/ Founded 2021; public beta launched September 16, 2022; founders Noam Shazeer and Daniel de Freitas.
  2. 2.Lisa P Argyle, Christopher A Bail, Ethan C Busby, Joshua R Gubler, Thomas Howe, Christopher Rytting, Taylor Sorensen, and David Wingate. 2023. Leveraging AI for democratic discourse: Chat interventions can improve online political conversations at scale. Proceedings of the National Academy of Sciences 120, 41 (2023), e2311627120.
  3. 3.Rebekah Lee Baik, Stephanie Lee, Serena Jinchen Xie, Wang Liao, Elina H Hwang, and Weichao Yuwen. 2025. Adapting Communication Styles in Health Chatbot using Large Language Models to Support Family Caregivers from Multicultural Backgrounds. In Proceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems. 1–8.
  4. 4.Angelique Bakker-Pieper and Reinout E de Vries. 2013. The incremental validity of communication styles over personality traits for leader outcomes. Human Performance 26, 1 (2013), 1–19.
  5. 5.Allan Bell. 1984. Language style as audience design. Language in society 13, 2 (1984), 145–204.
  6. 6.Penelope Brown. 1987. Politeness: Some universals in language usage.
  7. 7.Souradip Chakraborty, Jiahao Qin, Evrard Garcelon, Alessandro Lazaric, Matteo Pirotta, and Andrea Zanette. 2024. MaxMin-RLHF: Alignment with diverse human preferences. In Proceedings of the 41st International Conference on Machine Learning.
  8. 8.Ti-Chung Cheng, Carmen Badea, Christian Bird, Thomas Zimmermann, Robert DeLine, Nicole Forsgren, and Denae Ford. 2024. GEMS: Generative Expert Metric System through Iterative Prompt Priming. arXiv:2410.00880 [cs.SE] https://arxiv.org/abs/2410.00880
  9. 9.Yi Fei Cheng, Hirokazu Shirado, and Shunichi Kasahara. 2025. Conversational Agents on Your Behalf: Opportunities and Challenges of Shared Autonomy in Voice Communication for Multitasking. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems. 1–18.
  10. 10.Michelene TH Chi, Nicholas De Leeuw, Mei-Hung Chiu, and Christian LaVancher. 1994. Eliciting self-explanations improves understanding. Cognitive science 18, 3 (1994), 439–477.
  11. 11.Brian S Connelly and Deniz S Ones. 2010. An other perspective on personality: meta-analytic integration of observers’ accuracy and predictive validity. Psychological bulletin 136, 6 (2010), 1092.
  12. 12.Paul T Costa and Robert R McCrae. 2008. The revised neo personality inventory (neo-pi-r). The SAGE handbook of personality theory and assessment 2, 2 (2008), 179–198.
  13. 13.Lee J Cronbach and Paul E Meehl. 1955. Construct validity in psychological tests. Psychological bulletin 52, 4 (1955), 281.
  14. 14.Reinout E De Vries, Angelique Bakker-Pieper, Femke E Konings, and Barbara Schouten. 2013. The communication styles inventory (CSI) a six-dimensional behavioral model of communication styles and its relation with personality. Communication Research 40, 4 (2013), 506–532.
  15. 15.Robert F DeVellis and Carolyn T Thorpe. 2021. Scale development: Theory and applications. Sage publications.
  16. 16.Pierluigi Diotaiuti, Giuseppe Valente, Stefania Mancone, and Angela Grambone. 2020. Psychometric properties and a preliminary validation study of the Italian brief version of the communication styles inventory (CSI-B/I). Frontiers in Psychology 11 (2020), 1421.
  17. 17.Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daume Iii, and Kate Crawford. 2021. Datasheets for datasets. Commun. ACM 64, 12 (2021), 86–92.
  18. 18.Howard Giles, Nikolas Coupland, and Justine Coupland. 1991. Accommodation theory: Communication, context, and consequence. Contexts of accommodation: Developments in applied sociolinguistics 1 (1991), 1–68.
  19. 19.Erving Goffman. 1959. The presentation of self in everyday life, Double Day Anchor. Garden City, NY (1959).
  20. 20.Lewis R Goldberg. 1993. The structure of phenotypic personality traits. American psychologist 48, 1 (1993), 26.
  21. 21.Jeffrey T Hancock, Mor Naaman, and Karen Levy. 2020. AI-mediated communication: Definition, research agenda, and ethical considerations. Journal of Computer-Mediated Communication 25, 1 (2020), 89–100.
  22. 22.Jess Hohenstein and Malte Jung. 2018. AI-supported messaging: An investigation of human-human text conversation with AI support. In Extended abstracts of the 2018 CHI conference on human factors in computing systems. 1–6.
  23. 23.Sarah Susanna Hoppler, Robin Segerer, and Jana Nikitin. 2022. The Six Components of Social Interactions: Actor, Partner, Relation, Activities, Context, and Evaluation. Frontiers in Psychology Volume 12 - 2021 (2022). doi:10.3389/fpsyg.2021.743074
  24. 24.Jessica Huang, Ig-Jae Kim, and Dongwook Yoon. 2025. Mirror to Companion: Exploring Roles, Values, and Risks of AI Self-Clones through Story Completion. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems. 1–15.
  25. 25.Angel Hsing-Chi Hwang, John Oliver Siy, Renee Shelby, and Alison Lentz. 2024. In whose voice?: examining AI agent representation of people in social interaction through generative speech. In Proceedings of the 2024 ACM Designing Interactive Systems Conference. 224–245.
  26. 26.Maurice Jakesch, Megan French, Xiao Ma, Jeffrey T Hancock, and Mor Naaman. 2019. AI-mediated communication: How the perception that profile text was written by AI affects trustworthiness. In Proceedings of the 2019 CHI conference on human factors in computing systems. 1–13.
  27. 27.Taewook Kim, Jung Soo Lee, Zhenhui Peng, and Xiaojuan Ma. 2019. Love in lyrics: An exploration of supporting textual manifestation of affection in social messaging. Proceedings of the ACM on Human-Computer Interaction 3, CSCW (2019), 1–27.
  28. 28.Avraham N Kluger and Angelo DeNisi. 1996. The effects of feedback interventions on performance: a historical review, a meta-analysis, and a preliminary feedback intervention theory. Psychological bulletin 119, 2 (1996), 254.
  29. 29.Michal Kosinski, Sandra C Matz, Samuel D Gosling, Vesselin Popov, and David Stillwell. 2015. Facebook as a research tool for the social sciences: Opportunities, challenges, ethical considerations, and practical guidelines. American psychologist 70, 6 (2015), 543.
  30. 30.Michal Kosinski, David Stillwell, and Thore Graepel. 2013. Private traits and attributes are predictable from digital records of human behavior. Proceedings of the national academy of sciences 110, 15 (2013), 5802–5805.
  31. 31.Justin Kruger and David Dunning. 1999. Unskilled and unaware of it: how difficulties in recognizing one’s own incompetence lead to inflated self-assessments. Journal of personality and social psychology 77, 6 (1999), 1121.
  32. 32.John D Lee and Katrina A See. 2004. Trust in automation: Designing for appropriate reliance. Human factors 46, 1 (2004), 50–80.
  33. 33.Patrick Yung Kang Lee, Ning F Ma, Ig-Jae Kim, and Dongwook Yoon. 2023. Speculating on risks of AI clones to selfhood and relationships: Doppelgangerphobia, identity fragmentation, and living memories. Proceedings of the ACM on Human-computer Interaction 7, CSCW1 (2023), 1–28.
  34. 34.Joanne Leong, John Tang, Edward Cutrell, Sasa Junuzovic, Gregory Paul Baribault, and Kori Inkpen. 2024. Dittos: Personalized, embodied agents that participate in meetings when you are unavailable. Proceedings of the ACM on Human-Computer Interaction 8, CSCW2 (2024), 1–28.
  35. 35.Xiao Ma, Jeffrey T Hancock, Kenneth Lim Mingjie, and Mor Naaman. 2017. Self-disclosure and perceived trustworthiness of Airbnb host profiles. In Proceedings of the 2017 ACM conference on computer supported cooperative work and social computing. 2397–2409.
  36. 36.Paul A Mabe and Stephen G West. 1982. Validity of self-evaluation of ability: A review and meta-analysis. Journal of applied Psychology 67, 3 (1982), 280.
  37. 37.Reid McIlroy-Young, Jon Kleinberg, Siddhartha Sen, Solon Barocas, and Ashton Anderson. 2022. Mimetic models: Ethical implications of ai that acts like you. In Proceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society. 479–490.
  38. 38.Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. 2019. Model cards for model reporting. In Proceedings of the conference on fairness, accountability, and transparency. 220–229.
  39. 39.Lene Nielsen and Kira Storgaard Hansen. 2014. Personas is applicable: a study on the use of personas in Denmark. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems. 1665–1674.
  40. 40.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in neural information processing systems 35 (2022), 27730–27744.
  41. 41.Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2023. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th annual acm symposium on user interface software and technology. 1–22.
  42. 42.Joon Sung Park, Carolyn Q Zou, Aaron Shaw, Benjamin Mako Hill, Carrie Cai, Meredith Ringel Morris, Robb Willer, Percy Liang, and Michael S Bernstein. 2024. Generative agent simulations of 1,000 people. arXiv preprint arXiv:2411.10109 (2024).
  43. 43.Nilay Patel. 2024. The CEO of Zoom wants AI clones in meetings. https://www.theverge.com/2024/6/3/24168733/zoom-ceo-ai-clones-digital-twins-videoconferencing-decoder-interview Accessed: 2024-09-04.
  44. 44.Ellie Pavlick and Joel Tetreault. 2016. An empirical analysis of formality in online communication. Transactions of the association for computational linguistics 4 (2016), 61–74.
  45. 45.Sriyash Poddar, Yanming Wan, Hamish Ivison, Siddharth Choudhury, and Natasha Jaques. 2024. Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning. In NeurIPS 2024 Workshop on Pluralistic Alignment.
  46. 46.Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2023. Direct preference optimization: Your language model is secretly a reward model. Advances in neural information processing systems 36 (2023).
  47. 47.Inioluwa Deborah Raji, Andrew Smart, Rebecca N White, Margaret Mitchell, Timnit Gebru, Ben Hutchinson, Jamila Smith-Loud, Daniel Theron, and Parker Barnes. 2020. Closing the AI accountability gap: Defining an end-to-end framework for internal algorithmic auditing. In Proceedings of the 2020 conference on fairness, accountability, and transparency. 33–44.
  48. 48.Omar Shaikh, Shardul Sapkota, Shan Rizvi, Eric Horvitz, Joon Sung Park, Diyi Yang, and Michael S Bernstein. 2025. Creating General User Models from Computer Use. arXiv preprint arXiv:2505.10831 (2025).
  49. 49.Ashish Sharma, Inna W Lin, Adam S Miner, David C Atkins, and Tim Althoff. 2021. Towards facilitating empathic conversations in online mental health support: A reinforcement learning approach. In Proceedings of the web conference 2021. 194–205.
  50. 50.Simine Vazire. 2010. Who knows what about a person? The self–other knowledge asymmetry (SOKA) model. Journal of personality and social psychology 98, 2 (2010), 281.
  51. 51.Wu Youyou, Michal Kosinski, and David Stillwell. 2015. Computer-based personality judgments are more accurate than those made by humans. Proceedings of the National Academy of Sciences 112, 4 (2015), 1036–1040.
  52. 52.Ivan Zakazov, Mikolaj Boronski, Lorenzo Drudi, and Robert West. 2024. Assessing Social Alignment: Do Personality-Prompted Large Language Models Behave Like Humans? arXiv preprint arXiv:2412.16772 (2024).
  53. 53.Shuo Zhou, Zhe Zhang, and Timothy Bickmore. 2017. Adapting a persuasive conversational agent for the Chinese culture. In 2017 international conference on culture and computing (culture and computing). IEEE, 89–96.
  54. 54.Caleb Ziems, Minzhi Li, Anthony Zhang, and Diyi Yang. 2022. Inducing positive perspectives with text reframing. arXiv preprint arXiv:2204.02952 (2022).

Citation

MLA
Shang, R., et al. “Mimetic Alignment with ASPECT: Evaluation of AI-inferred Personal Profiles”. arXiv, 2026, https://doi.org/10.48550/arxiv.2603.26922.
APA
Shang, R., Marshall, D., Cutrell, E., & Ford, D. (2026). Mimetic Alignment with ASPECT: Evaluation of AI-inferred Personal Profiles. arXiv. https://doi.org/10.48550/arxiv.2603.26922
Chicago
Shang, R., D. Marshall, E. Cutrell, and D. Ford. 2026. “Mimetic Alignment with ASPECT: Evaluation of AI-inferred Personal Profiles”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2603.26922.
Harvard
Shang, R. et al. (2026) “Mimetic Alignment with ASPECT: Evaluation of AI-inferred Personal Profiles”. arXiv. Available at: https://doi.org/10.48550/arxiv.2603.26922.
Vancouver
1. Shang R, Marshall D, Cutrell E, Ford D (2026) Mimetic Alignment with ASPECT: Evaluation of AI-inferred Personal Profiles. https://doi.org/10.48550/arxiv.2603.26922

BibTeX

@misc{https://doi.org/10.48550/arxiv.2603.26922,
  doi = {10.48550/ARXIV.2603.26922},
  url = {https://arxiv.org/abs/2603.26922},
  author = {Shang, Ruoxi and Marshall, Dan and Cutrell, Edward and Ford, Denae},
  keywords = {Human-Computer Interaction (cs.HC), Artificial Intelligence (cs.AI), FOS: Computer and information sciences, FOS: Computer and information sciences, H.5.3; H.1.2},
  title = {Mimetic Alignment with ASPECT: Evaluation of AI-inferred Personal Profiles},
  publisher = {arXiv},
  year = {2026},
  copyright = {Creative Commons Attribution 4.0 International}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/