Mimetic Alignment with ASPECT: Evaluation of AI-inferred Personal Profiles
Ruoxi ShangDan MarshallEdward CutrellDenae Ford
Presents ASPECT, a pipeline that infers validated communication profiles from workplace behavioral data without per-person training, allowing language models to accurately mirror individual communication styles while giving users transparent evidence to review and correct their representations.
As artificial intelligence agents increasingly communicate on behalf of individuals in workplace settings, capturing a person's authentic communication style remains a major challenge. Existing personalization methods typically depend on costly per-person fine-tuning, produce generic outputs from shallow persona prompts, or optimize outputs for general human preference rather than individual communication style. The article introduces and evaluates ASPECT (Automated Social Psychometric Evaluation of Communication Traits), a prompt-based pipeline that directs large language models to construct structured, evidence-grounded communication profiles from observed workplace interaction data without requiring individual model training.
The article evaluates how accurately an automated system can infer personal communication styles from workplace interaction data and examines whether these inferred profiles produce appropriate, socially aligned communication in workplace scenarios. The approach builds upon the Communication Styles Inventory, a validated 92-item psychometric instrument spanning six communication dimensions. The system compresses 90 days of local chat and meeting transcripts by extracting conversational evidence facet by facet and generating item-level scores with traceable rationales. The evaluation involved a case study of 20 professionals across diverse roles within a single organization, encompassing 1,840 paired item ratings and 600 blinded scenario evaluations comparing profiled outputs against self-report and generic baselines.
The findings demonstrate three core outcomes. First, the automated profiling pipeline achieved moderate overall alignment with participant self-assessments (mean absolute error of 1.39 on a 5-point scale and a within-person rank correlation of 0.39), successfully preserving the relative profile shape and distinguishing overt interpersonal tones, though it exhibited systematic positive biases by over-rating structural traits like Preciseness. Second, in downstream workplace scenarios, responses generated using the inferred profiles were preferred on aggregate, capturing 42.5% of first-place rankings compared to 32.5% for generic responses and 25.0% for self-report baselines, with significantly higher mean alignment ratings. Third, interactive profile reviews shifted profiling into a collaborative negotiation: participants revised their self-ratings in approximately 17.7% of facet evaluations after reviewing linked behavioral evidence, reconciling gaps between their aspirational self-image and observed workplace behavior.
These results indicate that grounding personalization in concrete behavioral evidence and psychometric frameworks outperforms self-reported profiles, while avoiding the compute costs and opacity of fine-tuning. However, strong individual differences in preference emerged, and participants reported an uncanny valley effect where flawed personalization felt more unsettling than a standard, neutral generic response. Furthermore, participants actively drew boundaries between their natural styles and context-specific professional personas, demonstrating that effective digital representation requires context-aware boundaries rather than simple behavioral averages.
Organizations developing or deploying conversational proxies should implement human-in-the-loop review mechanisms that allow users to inspect behavioral evidence, calibrate scores, and establish context-specific persona boundaries before deployment. Automated shrinkage or calibration corrections should also be applied to counter predictable model biases on structural communication traits. Decision-makers should note key limitations: the findings reflect a sample of 20 technical and professional participants within a single enterprise, relying on 90 days of workplace text without multimodal cues. Further longitudinal pilot studies across broader industries and communication channels are recommended before deploying autonomous representative agents in high-stakes environments.
- Paper: LaMP: When Large Language Models Meet Personalization, Alireza Salemi et al. (2024). Establishes foundational benchmarks and retrieval-augmented methods for grounding LLM outputs in historical user profile data without full model retraining.
- Paper: Personalizing Dialogue Agents: I have a dog, do you have pets too?, Saizheng Zhang et al. (2018). Introduces the classic paradigm of conditioning conversational dialogue agents on explicit persona profiles to maintain conversational consistency.
- Paper: InCharacter: Evaluating Personality Fidelity in Role-Playing Agents through Psychological Interviews, Xintao Wang et al. (2024). Provides the psychometric framework for evaluating LLM agent fidelity and communication traits against validated psychological scales via structured interview and evaluation pipelines.
- Paper: Out of One, Many: Using Language Models to Simulate Human Samples, Lisa P. Argyle et al. (2022). Demonstrates how conditioning language models on personal and demographic sub-profiles enables realistic simulation of individual human attitudes and communication patterns.
- Paper: CogBench: a large language model walks into a psychology lab, Julian Coda-Forno et al. (2024). Adapts canonical cognitive and psychological behavioral metrics to assess the underlying traits, decision-making, and alignment of LLM agents.
- Paper: VizCopilot: Fostering Appropriate Reliance on Enterprise Chatbots with Context Visualization, Sam Yu-Te Lee et al. (2025). Examines workplace user inspection and interactive steering of underlying data contexts to prevent model mischaracterization and foster appropriate reliance.
- Paper: Learning Personalized Agents from Human Feedback, Kaiqu Liang et al. (2026). Extends personalized agent representation by introducing a dynamic human-in-the-loop framework where agents refine explicit user memory and clarify preferences during live interactions.
- Paper: SENSE-7: Taxonomy and Dataset for Measuring User Perceptions of Empathy in Sustained Human-AI Conversations, Jina Suh et al. (2026). Applies psychometric profiling and personalized agent interactions to the sustained evaluation of perceived digital empathy and relational dynamics in workplace conversations.
- Paper: LLMs Get Lost in Evolving User Intent, Jihoon Tack et al. (2026). Investigates the breakdown and adaptation challenges LLM agents face when tracking dynamically shifting user intents over ongoing multi-turn interactions.
