LaMP: When Large Language Models Meet Personalization
Alireza SalemiSheshera MysoreMichael BenderskyHamed Zamani
Introduces the LaMP benchmark and effective retrieval-augmented methods to evaluate and improve how large language models adapt their text generation and classification to individual user profiles across seven diverse tasks.
Modern natural language processing systems rely heavily on large language models that generate standard, non-personalized responses. While personalization is well established in search and recommendation platforms, current language model benchmarks predominantly use a one-size-fits-all paradigm that fails to account for individual user styles, preferences, and histories. As language models are increasingly integrated into customer-facing products, enterprise workflows, and content generation tools, adapting outputs to specific users is essential for delivering higher-quality experiences.
The article introduces the Language Model Personalization benchmark, designed to train and evaluate language models on producing personalized text. It systematically evaluates how retrieval-augmented techniques—which select relevant historical user profile items to ground model outputs—can enhance both classification and generation tasks across different user contexts.
The researchers developed a standardized benchmark comprising seven distinct tasks: three personalized classification tasks (citation identification, movie tagging, and product rating) and four personalized generation tasks (news headline generation, scholarly title generation, email subject generation, and tweet paraphrasing). The evaluation covers two operational environments: generalizing to new users and predicting future interactions of existing users over time. The authors assessed two primary retrieval augmentation frameworks—in-prompt augmentation and fusion-in-decoder—using dense semantic matching, keyword search, recency, and random baseline retrievers across fine-tuned and zero-shot open and commercial models.
The primary findings demonstrate that integrating personalized profile data substantially improves model performance across nearly all settings. Fine-tuning a language model with retrieval-augmented personalization achieved an average relative improvement of 23.5% across the benchmark compared to non-personalized baselines. In zero-shot settings without model fine-tuning, retrieval augmentation yielded a 12.2% average relative improvement. Crucially, retrieving semantically relevant or recent profile items consistently outperformed random profile sampling, showing that selective context injection is vital. In architectural comparisons, the fusion-in-decoder method delivered superior results for classification tasks, whereas in-prompt augmentation proved most effective for text generation.
These findings indicate that organizations do not need to incur the high computational and storage costs of maintaining separate fine-tuned models for each user. Instead, shared language models combined with effective user-profile retrieval provide a cost-effective, scalable architecture for personalized AI applications. Moreover, smaller fine-tuned models augmented with user profiles frequently outperformed significantly larger zero-shot models, offering immediate opportunities to optimize compute costs, latency, and operational efficiency.
Organizations developing user-facing language model applications should implement retrieval pipelines over user interaction histories rather than relying on generic prompting. When selecting deployment strategies, teams should weigh architectural trade-offs: use in-prompt augmentation for flexible, generation-heavy workflows across arbitrary models, and explore fusion-in-decoder architectures for specialized classification tasks. Future technical initiatives should focus on developing hybrid retrieval models that combine temporal recency and semantic relevance, as well as optimizing prompt compression to handle extensive user profiles within context window limits.
These conclusions are supported by empirical results across diverse benchmark tasks, though decision-makers should consider specific limitations. Most datasets rely on public web sources where pre-training data overlap cannot be fully ruled out, and evaluations focus primarily on short text tasks. Furthermore, fine-tuning on personal user data introduces potential privacy risks, requiring robust governance, privacy-preserving techniques, or secure in-house deployment when handling sensitive personal profiles.
- Paper: Personalizing Dialogue Agents: I have a dog, do you have pets too?, Saizheng Zhang et al. (2018). This foundational paper establishes the paradigm of conditioning conversational and generative models on explicit user persona profiles to achieve personalized text generation.
- Paper: In-Context Retrieval-Augmented Language Models, Ori Ram et al. (2023). This work introduces in-context retrieval-augmented language modeling by prepending retrieved external context into prompts without retraining, which forms the direct basis for LaMP's in-prompt retrieval architecture.
- Paper: Time Waits for No One! Analysis and Challenges of Temporal Misalignment, Kelvin Luu et al. (2022). This paper analyzes temporal degradation and misalignment across historical text corpora, motivating LaMP's evaluation setup for predicting user interactions over time.
- Paper: Self-Adaptive In-Context Learning: An Information Compression Perspective for In-Context Example Selection and Ordering, Zhiyong Wu et al. (2023). This study demonstrates dynamic, semantically aligned example selection and ranking for in-context learning, underpinning LaMP's selective retrieval over user interaction histories.
- Paper: When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories, Alex Troy Mallen et al. (2022). This work investigates when non-parametric retrieval memory is necessary versus parametric memory, providing the core rationale for augmenting language models with external user histories.
- Paper: Improving language models by retrieving from trillions of tokens, Sebastian Borgeaud et al. (2022). This paper establishes retrieval-augmented language modeling architectures (such as RETRO) that ground generation in external data indexes rather than sole parameter memory.
- Paper: Out of One, Many: Using Language Models to Simulate Human Samples, Lisa P. Argyle et al. (2022). This research pioneers conditioning large language models on individual demographic and attitudinal profiles to simulate human-specific behavior, preceding LaMP's systematic personalization benchmark.
- Paper: Evaluating Large Language Models as Generative User Simulators for Conversational Recommendation, Se-eun Yoon et al. (2024). This study extends the evaluation of personalized language models conditioned on user history to interactive, multi-turn conversational recommendation and user simulation.
- Paper: Evaluating Very Long-Term Conversational Memory of LLM Agents, Adyasha Maharana et al. (2024). This paper continues the exploration of user-grounded context by evaluating very long-term conversational memory, event dynamics, and personalization over extended time horizons.
- Paper: Enhancing Noise Robustness of Retrieval-Augmented Language Models with Adaptive Adversarial Training, Feiteng Fang et al. (2024). This work addresses a key practical vulnerability in retrieval-augmented language models like those tested in LaMP by developing adaptive adversarial training against retrieval noise.
- Paper: InstructRetro: Instruction Tuning post Retrieval-Augmented Pretraining, Boxin Wang et al. (2024). This paper advances retrieval-augmented architectures by demonstrating large-scale continued pretraining and instruction tuning over dynamic retrieval indexes.
