LaMP: When Large Language Models Meet Personalization

Alireza SalemiSheshera MysoreMichael BenderskyHamed Zamani

article2024ACL586 citations

Introduces the LaMP benchmark and effective retrieval-augmented methods to evaluate and improve how large language models adapt their text generation and classification to individual user profiles across seven diverse tasks.

Listen

Modern natural language processing systems rely heavily on large language models that generate standard, non-personalized responses. While personalization is well established in search and recommendation platforms, current language model benchmarks predominantly use a one-size-fits-all paradigm that fails to account for individual user styles, preferences, and histories. As language models are increasingly integrated into customer-facing products, enterprise workflows, and content generation tools, adapting outputs to specific users is essential for delivering higher-quality experiences.

The article introduces the Language Model Personalization benchmark, designed to train and evaluate language models on producing personalized text. It systematically evaluates how retrieval-augmented techniques—which select relevant historical user profile items to ground model outputs—can enhance both classification and generation tasks across different user contexts.

The researchers developed a standardized benchmark comprising seven distinct tasks: three personalized classification tasks (citation identification, movie tagging, and product rating) and four personalized generation tasks (news headline generation, scholarly title generation, email subject generation, and tweet paraphrasing). The evaluation covers two operational environments: generalizing to new users and predicting future interactions of existing users over time. The authors assessed two primary retrieval augmentation frameworks—in-prompt augmentation and fusion-in-decoder—using dense semantic matching, keyword search, recency, and random baseline retrievers across fine-tuned and zero-shot open and commercial models.

The primary findings demonstrate that integrating personalized profile data substantially improves model performance across nearly all settings. Fine-tuning a language model with retrieval-augmented personalization achieved an average relative improvement of 23.5% across the benchmark compared to non-personalized baselines. In zero-shot settings without model fine-tuning, retrieval augmentation yielded a 12.2% average relative improvement. Crucially, retrieving semantically relevant or recent profile items consistently outperformed random profile sampling, showing that selective context injection is vital. In architectural comparisons, the fusion-in-decoder method delivered superior results for classification tasks, whereas in-prompt augmentation proved most effective for text generation.

These findings indicate that organizations do not need to incur the high computational and storage costs of maintaining separate fine-tuned models for each user. Instead, shared language models combined with effective user-profile retrieval provide a cost-effective, scalable architecture for personalized AI applications. Moreover, smaller fine-tuned models augmented with user profiles frequently outperformed significantly larger zero-shot models, offering immediate opportunities to optimize compute costs, latency, and operational efficiency.

Organizations developing user-facing language model applications should implement retrieval pipelines over user interaction histories rather than relying on generic prompting. When selecting deployment strategies, teams should weigh architectural trade-offs: use in-prompt augmentation for flexible, generation-heavy workflows across arbitrary models, and explore fusion-in-decoder architectures for specialized classification tasks. Future technical initiatives should focus on developing hybrid retrieval models that combine temporal recency and semantic relevance, as well as optimizing prompt compression to handle extensive user profiles within context window limits.

These conclusions are supported by empirical results across diverse benchmark tasks, though decision-makers should consider specific limitations. Most datasets rely on public web sources where pre-training data overlap cannot be fully ruled out, and evaluations focus primarily on short text tasks. Furthermore, fine-tuning on personal user data introduces potential privacy risks, requiring robust governance, privacy-preserving techniques, or secure in-house deployment when handling sensitive personal profiles.

arXiv: 2304.11406
Cover for LaMP: When Large Language Models Meet Personalization

Abstract

This paper highlights the importance of personalization in large language models and introduces the LaMP benchmark — a novel benchmark for training and evaluating language models for producing personalized outputs. LaMP offers a comprehensive evaluation framework with diverse language tasks and multiple entries for each user profile. It consists of seven personalized tasks, spanning three text classification and four text generation tasks. We additionally propose two retrieval augmentation approaches that retrieve personal items from each user profile for personalizing language model outputs. To this aim, we study various retrieval models, including term matching, semantic matching, and time-aware methods. Extensive experiments on LaMP for zero-shot and fine-tuned language models demonstrate the efficacy of the proposed retrieval augmentation approach and highlight the impact of personalization in various natural language tasks.

Table of Contents

  • 1 Introduction
  • 2 The LaMP Benchmark
  • 2.1 Tasks Definitions
  • 2.2 Data Splits
  • 2.3 Evaluation
  • 3 Retrieval Augmentation for Personalizing LLMs
  • 4 Experiments
  • 4.1 Experimental Setup
  • 4.2 Fine-Tuning Retrieval Augmented LMs for Personalization
  • 4.3 Zero-Shot Personalized Results for LLMs
  • 5 Research Problems Enabled by LaMP
  • 6 Related Work
  • 7 Conclusion
  • Acknowledgments
  • Limitations
  • References
  • A Data Creation Details for Tasks in the LaMP benchmark
  • A.1 User-based Separation Setting
  • A.2 Time-based Separation Setting
  • B Samples of the Tasks Introduced in the LaMP Benchmark
  • C Prompts Used for Adding User Profile to the Language Model's Input
  • D Performance of the Models on the Validation Set
  • E Performance of Some Other Non-Personalized Baselines on the LaMP Benchmark
  • F Dataset Licenses
  • G AI Assistance Usage

Knowls

  1. Knowl 1 — Problem Formulation of Language Model Personalization and Evaluation Splits in LaMP

    definition

    In the Language Model Personalization (LaMP) framework, the goal is to generate a personalized output sequence yy given an input sequence xx and a user profile PuP_u corresponding to user uu. The user profile is defined as a collection of historical input-output pairs created or endorsed by that user:

    Pu={(xu1,yu1),(xu2,yu2),…,(xumu,yumu)}P_u = \{(x_{u1}, y_{u1}), (x_{u2}, y_{u2}), \dots, (x_{u m_u}, y_{u m_u})\}

    where mum_u is the number of historical interactions for user uu. Each data entry across tasks is structured as a triple (x,y,Pu)(x, y, P_u).

    The benchmark defines two distinct data splitting regimes to evaluate different personalization environments:

    1. User-based separation (LaMP-iiU for task ii): Partitions the dataset across disjoint sets of users such that the training, validation, and test splits have no overlapping users. This setting evaluates the ability of models to personalize for newly observed users.
    2. Time-based separation (LaMP-iiT for task ii): Partitions each user's chronological activity history. The earliest interactions form the user profile PuP_u, intermediate interactions form training and validation instances, and the most recent interactions form the test instances. This setting evaluates personalization for future interactions of known users.
  2. Knowl 2 — The LaMP Benchmark Task Suite and Dataset Statistics

    data/table

    The LaMP benchmark consists of seven diverse language processing tasks spanning text classification and text generation:

    1. LaMP-1: Personalized Citation Identification: Binary classification determining which of two candidate references an author will cite in an input paper abstract (Source: Citation Network Dataset V14; Metric: Accuracy).
    2. LaMP-2: Personalized Movie Tagging: 15-class categorical classification predicting which user-specific tag applies to a movie description (Sources: MovieLens, MovieDB; Metrics: Accuracy, macro-F1).
    3. LaMP-3: Personalized Product Rating: 5-class ordinal classification (1 to 5 stars) predicting the star rating associated with an input product review (Source: Amazon Reviews; Metrics: MAE, RMSE).
    4. LaMP-4: Personalized News Headline Generation: Text generation producing an author-stylized news headline given an article body (Source: HuffPost News Category Dataset; Metrics: ROUGE-1, ROUGE-L).
    5. LaMP-5: Personalized Scholarly Title Generation: Text generation predicting the title of a research paper given its abstract (Source: Citation Network Dataset V14; Metrics: ROUGE-1, ROUGE-L).
    6. LaMP-6: Personalized Email Subject Generation: Text generation producing an email subject line given the email body (Source: Avocado Research Email Collection; Metrics: ROUGE-1, ROUGE-L).
    7. LaMP-7: Personalized Tweet Paraphrasing: Text generation rewriting a paraphrased tweet back into the specific author's linguistic style (Source: Sentiment140; Metrics: ROUGE-1, ROUGE-L).
    Task Split Type #Train #Dev #Test Input Len. Profile Size
    Citation Ident. User 9682 2500 2500 51.40±5.7251.40 \pm 5.72 90.61±53.8790.61 \pm 53.87
    Citation Ident. Time 6542 1500 1500 51.43±5.7051.43 \pm 5.70 84.15±47.5484.15 \pm 47.54
    Movie Tag. User 3820 692 870 92.27±20.8392.27 \pm 20.83 159.29±330.81159.29 \pm 330.81
    Movie Tag. Time 5073 1410 1557 92.39±21.9592.39 \pm 21.95 86.76±189.5286.76 \pm 189.52
    Product Rat. User 20000 2500 2500 145.14±157.96145.14 \pm 157.96 188.10±129.42188.10 \pm 129.42
    Product Rat. Time 20000 2500 2500 128.18±146.25128.18 \pm 146.25 185.40±129.30185.40 \pm 129.30
    News Headline User 12527 1925 2376 30.53±12.6730.53 \pm 12.67 287.16±360.62287.16 \pm 360.62
    News Headline Time 12500 1500 1800 29.97±12.0929.97 \pm 12.09 204.59±250.75204.59 \pm 250.75
    Scholarly Title User 9682 2500 2500 152.81±86.60152.81 \pm 86.60 89.61±53.8789.61 \pm 53.87
    Scholarly Title Time 14682 1500 1500 162.34±65.63162.34 \pm 65.63 87.88±53.6387.88 \pm 53.63
    Email Subject User 4840 1353 1246 436.15±805.54436.15 \pm 805.54 80.72±51.7380.72 \pm 51.73
    Email Subject Time 4821 1250 1250 454.87±889.41454.87 \pm 889.41 55.67±36.3255.67 \pm 36.32
    Tweet Para. User 10437 1500 1496 29.76±6.9429.76 \pm 6.94 17.74±15.1017.74 \pm 15.10
    Tweet Para. Time 13437 1498 1500 29.72±7.0129.72 \pm 7.01 15.71±14.8615.71 \pm 14.86
  3. Knowl 3 — Retrieval-Augmentation Architectures for Personalization: IPA and FiD

    model/method

    Because raw user profiles PuP_u are large and contain varying relevance to a given input xix_i, retrieval augmentation is used to select kk relevant profile entries. The framework uses a query construction function ϕq\phi_q, a retriever R(q,Pu,k)R(q, P_u, k), and a prompt construction function ϕp\phi_p.

    Two architectures integrate the retrieved entries:

    1. In-Prompt Augmentation (IPA): The retrieved items are concatenated with the task instructions and test input into a single augmented input prompt xˉi\bar{x}_i:

    xˉi=ϕp(xi,R(ϕq(xi),Pu,k))\bar{x}_i = \phi_p(x_i, R(\phi_q(x_i), P_u, k))

    IPA can be used directly with both fine-tuned and off-the-shelf zero-shot language models across any model architecture (encoder-decoder or decoder-only).

    1. Fusion-in-Decoder (FiD): For an encoder-decoder architecture, each retrieved profile item dij∈R(ϕq(xi),Pu,k)d_{ij} \in R(\phi_q(x_i), P_u, k) is combined with the input xix_i into an individual sequence:

    xˉij=ϕp(xi,dij)\bar{x}_{ij} = \phi_p(x_i, d_{ij})

    Each xˉij\bar{x}_{ij} is encoded independently by the transformer encoder. The decoder then applies cross-attention across the concatenated representations of all kk encoder outputs to produce the target tokens yy. FiD requires model training and encoder-decoder backbones, but allows processing a larger number of retrieved items without exhausting standard context windows.

  4. Knowl 4 — Two-Stage Prompt Construction and Context Budget Trimming

    algorithm

    To combine kk retrieved profile items into a prompt without exceeding the maximum context window of a language model, a two-stage formatting and token-budget trimming procedure is applied.

    Let LL denote the maximum context window of the language model, Lˉ\bar{L} the reserved token capacity for the task input and instruction, and kk the number of profile items retrieved by R(ϕq(xi),Pu,k)R(\phi_q(x_i), P_u, k). Each individual profile entry prompt is constrained to at most L−Lˉk\frac{L - \bar{L}}{k} tokens.

    Input: Task input xix_i, retrieved profile entries {di1,di2,…,dik}⊆Pu\{d_{i1}, d_{i2}, \dots, d_{ik}\} \subseteq P_u, context limit LL, input budget Lˉ\bar{L}
    Output: Aggregated personalized input prompt xˉi\bar{x}_i
    b←⌊(L−Lˉ)/k⌋b \leftarrow \lfloor (L - \bar{L}) / k \rfloor
    Pprompts←empty listP_{prompts} \leftarrow \text{empty list}
    for each retrieved item dijd_{ij} in {di1,…,dik}\{d_{i1}, \dots, d_{ik}\} do
        pj←PPEP(dij)p_{j} \leftarrow \text{PPEP}(d_{ij}) # Apply task-specific Per Profile Entry Prompt template
        if token_length(pjp_j) > bb then
            pj←trim(pj,b)p_{j} \leftarrow \text{trim}(p_j, b) # Truncate non-template content (e.g., text/abstract body) preserving titles, scores, or tags
        end if
        Pprompts.append(pj)P_{prompts}.\text{append}(p_j)
    end for
    xˉi←AIP(Pprompts,xi)\bar{x}_i \leftarrow \text{AIP}(P_{prompts}, x_i) # Combine PPEPs with task instruction and xix_i via Aggregated Input Prompt template
    return xˉi\bar{x}_i
  5. Knowl 5 — Performance of Fine-Tuned FlanT5-base on the LaMP Benchmark

    empirical result

    Fine-tuning FlanT5-base (250M parameters) with retrieval-augmented user profiles achieves a relative average performance improvement of 23.5% over non-personalized baselines across the LaMP benchmark. Dense semantic retrieval (Contriever) and BM25 substantially outperform random profile entry selection and non-personalized models across both user-based and time-based splits.

    Non-Personalized Personalized (Tuned Profile)
    Dataset Metric No-Retr. Global Rand. Profile Rand. Tuned IPA Tuned FiD (k=16k=16)
    LaMP-1U Accuracy ↑\uparrow 0.518 0.539 0.598 0.734 0.754
    LaMP-2U Accuracy ↑\uparrow 0.468 0.442 0.497 0.556 0.642
    LaMP-2U F1 ↑\uparrow 0.435 0.403 0.459 0.519 0.607
    LaMP-3U MAE ↓\downarrow 0.275 0.286 0.284 0.246 0.236
    LaMP-3U RMSE ↓\downarrow 0.581 0.607 0.602 0.565 0.539
    LaMP-4U ROUGE-1 ↑\uparrow 0.153 0.159 0.162 0.186 0.180
    LaMP-4U ROUGE-L ↑\uparrow 0.140 0.147 0.148 0.171 0.166
    LaMP-5U ROUGE-1 ↑\uparrow 0.418 0.408 0.409 0.450 0.431
    LaMP-5U ROUGE-L ↑\uparrow 0.378 0.370 0.371 0.409 0.392
    LaMP-6U ROUGE-1 ↑\uparrow 0.379 0.473 0.486 0.587 0.567
    LaMP-6U ROUGE-L ↑\uparrow 0.358 0.457 0.470 0.575 0.555
    LaMP-7U ROUGE-1 ↑\uparrow 0.509 0.510 0.514 0.528 0.517
    LaMP-7U ROUGE-L ↑\uparrow 0.455 0.457 0.460 0.475 0.464
    LaMP-1T Accuracy ↑\uparrow 0.628 0.625 0.657 0.714 0.698
    LaMP-2T Accuracy ↑\uparrow 0.506 0.513 0.518 0.564 0.661
    LaMP-2T F1 ↑\uparrow 0.443 0.449 0.456 0.519 0.624
    LaMP-3T MAE ↓\downarrow 0.280 0.280 0.279 0.266 0.250
    LaMP-3T RMSE ↓\downarrow 0.615 0.616 0.612 0.598 0.598
    LaMP-4T ROUGE-1 ↑\uparrow 0.159 0.160 0.169 0.177 0.170
    LaMP-4T ROUGE-L ↑\uparrow 0.145 0.147 0.155 0.162 0.157
    LaMP-5T ROUGE-1 ↑\uparrow 0.462 0.459 0.460 0.479 0.456
    LaMP-5T ROUGE-L ↑\uparrow 0.416 0.412 0.414 0.431 0.414
    LaMP-6T ROUGE-1 ↑\uparrow 0.479 0.500 0.525 0.547 0.540
    LaMP-6T ROUGE-L ↑\uparrow 0.463 0.452 0.507 0.533 0.525
    LaMP-7T ROUGE-1 ↑\uparrow 0.462 0.474 0.505 0.516 0.502
    LaMP-7T ROUGE-L ↑\uparrow 0.416 0.457 0.456 0.465 0.450
  6. Knowl 6 — Zero-Shot Personalization Performance with FlanT5-XXL and GPT-3.5

    empirical result

    In zero-shot evaluation without task-specific fine-tuning, incorporating retrieved user profile entries via In-Prompt Augmentation (IPA) achieves a relative average performance improvement of 12.2% across tasks compared to non-personalized prompting. Personalization improves performance across all tasks except LaMP-7 (Personalized Tweet Paraphrasing).

    User-based Separation Time-based Separation
    Non-Personalized Personalized Non-Personalized Personalized
    Dataset Metric FlanT5 GPT-3.5 FlanT5 GPT-3.5 FlanT5 GPT-3.5 FlanT5 GPT-3.5
    LaMP-1 Acc ↑\uparrow 0.520 0.541 0.699 0.695 0.502 0.508 0.636 0.634
    LaMP-2 Acc ↑\uparrow 0.365 0.408 0.414 0.508 0.360 0.382 0.396 0.466
    F1 ↑\uparrow 0.308 0.314 0.364 0.457 0.276 0.299 0.304 0.418
    LaMP-3 MAE ↓\downarrow 0.344 0.706 0.267 0.620 0.333 0.677 0.299 0.603
    RMSE ↓\downarrow 0.650 0.972 0.552 1.049 0.650 0.948 0.616 1.002
    LaMP-4 R-1 ↑\uparrow 0.163 0.136 0.182 0.150 0.176 0.146 0.188 0.158
    R-L ↑\uparrow 0.147 0.119 0.167 0.133 0.160 0.128 0.172 0.140
    LaMP-5 R-1 ↑\uparrow 0.442 0.387 0.450 0.390 0.471 0.424 0.483 0.425
    R-L ↑\uparrow 0.400 0.329 0.411 0.329 0.422 0.355 0.433 0.351
    LaMP-6 R-1 ↑\uparrow 0.362 – 0.482 – 0.335 – 0.401 –
    R-L ↑\uparrow 0.343 – 0.471 – 0.319 – 0.387 –
    LaMP-7 R-1 ↑\uparrow 0.453 0.399 0.448 0.390 0.448 0.390 0.440 0.382
    R-L ↑\uparrow 0.395 0.336 0.394 0.322 0.396 0.330 0.389 0.318

    Note: GPT-3.5 was not evaluated on LaMP-6 due to privacy and licensing terms of the Avocado dataset. For classification outputs that do not match a valid label, predictions are mapped to the most semantically similar class using BERTScore (GPT-3.5 produced invalid class strings 2% to 8% of the time, while FlanT5-XXL strictly adhered to the candidate labels).

  7. Knowl 7 — Architectural Task Specialization: FiD vs. IPA for Personalization

    empirical result

    When fine-tuning language models with user profile augmentation, Fusion-in-Decoder (FiD) and In-Prompt Augmentation (IPA) exhibit distinct task-type advantages:

    1. Text Classification: FiD (k=16k=16) systematically outperforms IPA across classification tasks where trained. On LaMP-2U (Movie Tagging), FiD reaches 0.642 Accuracy / 0.607 F1 versus IPA's 0.556 Accuracy / 0.519 F1. On LaMP-2T, FiD achieves 0.661 Accuracy / 0.624 F1 versus IPA's 0.564 Accuracy / 0.519 F1. On LaMP-3 (Product Rating), FiD achieves lower MAE (0.236 vs 0.246 on User; 0.250 vs 0.266 on Time).
    2. Text Generation: IPA consistently outperforms FiD across all four generation tasks (LaMP-4 News Headline, LaMP-5 Scholarly Title, LaMP-6 Email Subject, and LaMP-7 Tweet Paraphrasing) on both user-based and time-based splits. For instance, on LaMP-6U, IPA achieves 0.587 ROUGE-1 / 0.575 ROUGE-L compared to FiD's 0.567 ROUGE-1 / 0.555 ROUGE-L.
  8. Knowl 8 — Fine-Tuned Smaller Language Models Outperform Zero-Shot LLMs on Personalized Tasks

    empirical result

    A smaller model (FlanT5-base with 250M parameters) fine-tuned on personalized tasks consistently outperforms substantially larger models (FlanT5-XXL with 11B parameters and GPT-3.5) evaluated in a zero-shot setting.

    • On LaMP-1U (Citation Identification), fine-tuned FlanT5-base achieves an accuracy of 0.754 (FiD) and 0.734 (IPA), whereas zero-shot FlanT5-XXL achieves 0.699 and zero-shot GPT-3.5 achieves 0.695.
    • On LaMP-2U (Movie Tagging), fine-tuned FlanT5-base achieves 0.642 accuracy / 0.607 F1, while zero-shot FlanT5-XXL reaches only 0.414 accuracy / 0.364 F1 and zero-shot GPT-3.5 reaches 0.508 accuracy / 0.457 F1.
    • On LaMP-6U (Email Subject Generation), fine-tuned FlanT5-base achieves 0.587 ROUGE-1 / 0.575 ROUGE-L, compared to zero-shot FlanT5-XXL at 0.482 ROUGE-1 / 0.471 ROUGE-L.
  9. Knowl 9 — Profile Size Scaling and Retrieval Model Selection in Personalized Prompting

    empirical result

    The effectiveness of retrieval-augmented personalization depends on the retrieval mechanism and the number of retrieved items kk:

    1. Retriever Selection: No single retrieval strategy is universally best across all tasks. Dense semantic retrieval (Contriever) performs best for classification tasks (LaMP-1, LaMP-2, LaMP-3) and semantic generation tasks (LaMP-4U, LaMP-7U). Term-matching retrieval (BM25) performs best for text generation tasks requiring precise lexical matching (LaMP-5U Scholarly Titles and LaMP-6U Email Subjects). Recency retrieval outperforms Contriever in time-based rating and headline tasks (LaMP-3T and LaMP-4T).
    2. Item Count (kk): In In-Prompt Augmentation (IPA), increasing kk from 1 to 4 generally monotonically improves downstream accuracy and ROUGE scores across tasks before plateauing. However, further increases can lead to performance degradation due to fixed context length constraints requiring aggressive prompt trimming.
  10. Knowl 10 — Methodological, Evaluative, and Privacy Limitations in Language Model Personalization

    limitation

    The LaMP benchmark and retrieval-based personalization approach have four core limitations:

    1. Task Simplifications: Real-world citation recommendation involves large-scale open corpus candidate ranking rather than binary classification, and product review rating prediction is framed without access to accompanying numerical scores.
    2. Pre-training Data Contamination: Six of the seven datasets are derived from public web sources (HuffPost, Amazon, Twitter, Citation Network, MovieLens) and may have been seen during LLM pretraining, potentially inflating zero-shot benchmark scores. Only the private Avocado dataset (LaMP-6) provides guarantees against pre-training leakage.
    3. Metric Limitations for Personalization: Standard generation metrics (ROUGE-1, ROUGE-L, BLEU) assess n-gram overlap against a single reference but do not capture subjective alignment, idiosyncratic user writing preferences, or stylistic satisfaction.
    4. Privacy and Membership Inference Risks: Personalizing language models via fine-tuning on user profile data creates vulnerability to membership inference and privacy extraction attacks; zero-shot prompting avoids gradient updates on private profiles but requires trusted hosting.

Coverage note — Non-personalized auxiliary baseline experiments from the appendix using SVM, BERT, and BART were omitted as they serve only as conventional non-personalized reference points rather than core contributions.

References

  1. 1.Xiang Ao, Xiting Wang, Ling Luo, Ying Qiao, Qing He, and Xing Xie. 2021. PENS: A dataset and generic framework for personalized news headline generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 82–92, Online. Association for Computational Linguistics.
  2. 2.Satanjeev Banerjee and Alon Lavie. 2005. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, pages 65–72, Ann Arbor, Michigan. Association for Computational Linguistics.
  3. 3.Paul N. Bennett, Ryen W. White, Wei Chu, Susan T. Dumais, Peter Bailey, Fedor Borisyuk, and Xiaoyuan Cui. 2012. Modeling the impact of short- and long-term behavior on search personalization. In Proceedings of the 35th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’12, page 185–194, New York, NY, USA. Association for Computing Machinery.
  4. 4.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pellat, Kevin Robinson, Dasha Valter, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Zhao, Yanping Huang, Andrew Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei. 2022. Scaling instruction-finetuned language models.
  5. 5.Bruce W. Croft, Stephen Cronen-Townsend, and Victor Lavrenko. 2001. Relevance feedback and personalization: A language modeling perspective. In DELOS Workshop: Personalisation and Recommender Systems in Digital Libraries.
  6. 6.Abhinandan S. Das, Mayur Datar, Ashutosh Garg, and Shyam Rajaram. 2007. Google news personalization: Scalable online collaborative filtering. In Proceedings of the 16th International Conference on World Wide Web, WWW ’07, page 271–280, New York, NY, USA. Association for Computing Machinery.
  7. 7.James Davidson, Benjamin Liebald, Junning Liu, Palash Nandy, Taylor Van Vleet, Ullas Gargi, Sujoy Gupta, Yu He, Mike Lambert, Blake Livingston, and Dasarathi Sampath. 2010. The youtube video recommendation system. In Proceedings of the Fourth ACM Conference on Recommender Systems, RecSys ’10, page 293–296, New York, NY, USA. Association for Computing Machinery.
  8. 8.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In North American Chapter of the Association for Computational Linguistics.
  9. 9.Shiran Dudy, Steven Bedrick, and Bonnie Webber. 2021. Refocusing on relevance: Personalization in NLG. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 5190–5202, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  10. 10.Susan T. Dumais. 2016. Personalized search: Potential and pitfalls. In Proceedings of the 25th ACM International on Conference on Information and Knowledge Management, CIKM ’16, page 689, New York, NY, USA. Association for Computing Machinery.
  11. 11.Peter S. Fader, Bruce G.S. Hardie, and Ka Lok Lee. 2005. Rfm and clv: Using iso-value curves for customer base analysis. Journal of Marketing Research, 42(4):415–430.
  12. 12.Michael Färber and Adam Jatowt. 2020. Citation recommendation: Approaches and datasets. Int. J. Digit. Libr.
  13. 13.Lucie Flek. 2020. Returning the N to NLP: Towards contextually personalized classification models. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7828–7838, Online. Association for Computational Linguistics.
  14. 14.Andrew Fowler, Kurt Partridge, Ciprian Chelba, Xiaojun Bi, Tom Ouyang, and Shumin Zhai. 2015. Effects of language modeling and its personalization on touchscreen typing performance. In Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems, CHI ’15, page 649–658, New York, NY, USA. Association for Computing Machinery.
  15. 15.Markus Freitag and Yaser Al-Onaizan. 2017. Beam search strategies for neural machine translation. In Proceedings of the First Workshop on Neural Machine Translation, pages 56–60, Vancouver. Association for Computational Linguistics.
  16. 16.Sebastian Gehrmann, Tosin Adewumi, Karmanya Aggarwal, Pawan Sasanka Ammanamanchi, Anuoluwapo Aremu, Antoine Bosselut, Khyathi Raghavi Chandu, Miruna-Adriana Clinciu, Dipanjan Das, Kaustubh Dhole, Wanyu Du, Esin Durmus, Ondřej Dušek, Chris Chinenye Emezue, Varun Gangal, Cristina Garbacea, Tatsunori Hashimoto, Yufang Hou, Yacine Jernite, Harsh Jhamtani, Yangfeng Ji, Shailza Jolly, Mihir Kale, Dhruv Kumar, Faisal Ladhak, Aman Madaan, Mounica Maddela, Khyati Mahajan, Saad Mahamood, Bodhisattwa Prasad Majumder, Pedro Henrique Martins, Angelina McMillan-Major, Simon Mille, Emiel van Miltenburg, Moin Nadeem, Shashi Narayan, Vitaly Nikolaev, Andre Niyongabo Rubungo, Salomey Osei, Ankur Parikh, Laura Perez-Beltrachini, Niranjan Ramesh Rao, Vikas Raunak, Juan Diego Rodriguez, Sashank Santhanam, João Sedoc, Thibault Sellam, Samira Shaikh, Anastasia Shimorina, Marco Antonio Sobrevilla Cabezudo, Hendrik Strobelt, Nishant Subramani, Wei Xu, Diyi Yang, Akhila Yerukola, and Jiawei Zhou. 2021. The GEM benchmark: Natural language generation, its evaluation and metrics. In Proceedings of the 1st Workshop on Natural Language Generation, Evaluation, and Metrics (GEM 2021), pages 96–120, Online. Association for Computational Linguistics.
  17. 17.Alec Go, Richa Bhayani, and Lei Huang. 2009. Twitter sentiment classification using distant supervision.
  18. 18.Mitchell L. Gordon, Michelle S. Lam, Joon Sung Park, Kayur Patel, Jeff Hancock, Tatsunori Hashimoto, and Michael S. Bernstein. 2022. Jury learning: Integrating dissenting voices into machine learning models. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems, CHI ’22, New York, NY, USA. Association for Computing Machinery.
  19. 19.Manish Gupta, Rui Li, Zhijun Yin, and Jiawei Han. 2010. Survey on social tagging techniques. SIGKDD Explor. Newsl.
  20. 20.F. Maxwell Harper and Joseph A. Konstan. 2015. The movielens datasets: History and context. ACM Trans. Interact. Intell. Syst., 5(4).
  21. 21.M.A. Hearst, S.T. Dumais, E. Osuna, J. Platt, and B. Scholkopf. 1998. Support vector machines. IEEE Intelligent Systems and their Applications, 13(4):18–28.
  22. 22.Xiaolei Huang, Lucie Flek, Franck Dernoncourt, Charles Welch, Silvio Amir, Ramit Sawhney, and Diyi Yang. 2022. Usernlp’22: 2022 international workshop on user-centered natural language processing. In Companion Proceedings of the Web Conference 2022, WWW ’22, page 1176–1177, New York, NY, USA. Association for Computing Machinery.
  23. 23.Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bojanowski, Armand Joulin, and Edouard Grave. 2022. Unsupervised dense information retrieval with contrastive learning. Transactions on Machine Learning Research.
  24. 24.Gautier Izacard and Edouard Grave. 2021. Leveraging passage retrieval with generative models for open domain question answering. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 874–880, Online. Association for Computational Linguistics.
  25. 25.Aaron Jaech and Mari Ostendorf. 2018. Personalized language model for query auto-completion. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 700–705, Melbourne, Australia. Association for Computational Linguistics.
  26. 26.Milton King and Paul Cook. 2020. Evaluating approaches to personalizing language models. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 2461–2469, Marseille, France. European Language Resources Association.
  27. 27.Hannah Rose Kirk, Bertie Vidgen, Paul Röttger, and Scott A Hale. 2023. Personalisation within bounds: A risk taxonomy and policy framework for the alignment of large language models with personalised feedback. arXiv preprint arXiv:2303.05453.
  28. 28.Joseph Konstan and Loren Terveen. 2021. Human-centered recommender systems: Origins, advances, challenges, and opportunities. AI Magazine, 42(3):31–42.
  29. 29.Brian Lester, Rami Al-Rfou, and Noah Constant. 2021. The power of scale for parameter-efficient prompt tuning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 3045–3059, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  30. 30.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online. Association for Computational Linguistics.
  31. 31.Pan Li and Alexander Tuzhilin. 2019. Towards controllable and personalized review generation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3237–3245, Hong Kong, China. Association for Computational Linguistics.
  32. 32.Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74–81, Barcelona, Spain. Association for Computational Linguistics.
  33. 33.Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In International Conference on Learning Representations.
  34. 34.Bodhisattwa Prasad Majumder, Shuyang Li, Jianmo Ni, and Julian McAuley. 2019. Generating personalized recipes from historical user preferences. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 5976–5982, Hong Kong, China. Association for Computational Linguistics.
  35. 35.Justus Mattern, Fatemehsadat Mireshghallah, Zhijing Jin, Bernhard Schölkopf, Mrinmaya Sachan, and Taylor Berg-Kirkpatrick. 2023. Membership inference attacks against language models via neighbourhood comparison.
  36. 36.Pierre-Emmanuel Mazaré, Samuel Humeau, Martin Raison, and Antoine Bordes. 2018. Training millions of personalized dialogue agents. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2775–2779, Brussels, Belgium. Association for Computational Linguistics.
  37. 37.Fatemehsadat Mireshghallah, Vaishnavi Shrivastava, Milad Shokouhi, Taylor Berg-Kirkpatrick, Robert Sim, and Dimitrios Dimitriadis. 2022. UserIdentifier: Implicit user representations for simple and effective personalized sentiment analysis. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3449–3456, Seattle, United States. Association for Computational Linguistics.
  38. 38.Rishabh Misra. 2022. News category dataset. arXiv preprint arXiv:2209.11429.
  39. 39.Rishabh Misra and Jigyasa Grover. 2021. Sculpting Data for ML: The first act of Machine Learning.
  40. 40.Maxim Naumov, Dheevatsa Mudigere, Hao-Jun Michael Shi, Jianyu Huang, Narayanan Sundaraman, Jongsoo Park, Xiaodong Wang, Udit Gupta, Carole-Jean Wu, Alisson G. Azzolini, Dmytro Dzhulgakov, Andrey Mallevich, Ilia Cherniavskii, Yinghai Lu, Raghuraman Krishnamoorthi, Ansha Yu, Volodymyr Kondratenko, Stephanie Pereira, Xianjie Chen, Wenlin Chen, Vijay Rao, Bill Jia, Liang Xiong, and Misha Smelyanskiy. 2019. Deep learning recommendation model for personalization and recommendation systems.
  41. 41.Jianmo Ni, Jiacheng Li, and Julian McAuley. 2019. Justifying recommendations using distantly-labeled reviews and fine-grained aspects. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 188–197, Hong Kong, China. Association for Computational Linguistics.
  42. 42.Oard, Douglas, Webber, William, Kirsch, David A., and Golitsynskiy, Sergey. 2015. Avocado research email collection.
  43. 43.OpenAI. 2023. Gpt-4 technical report.
  44. 44.Sheena Panthaplackel, Adrian Benton, and Mark Dredze. 2022. Updated headline generation: Creating updated summaries for evolving news stories. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6438–6461, Dublin, Ireland. Association for Computational Linguistics.
  45. 45.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311–318, Philadelphia, Pennsylvania, USA. Association for Computational Linguistics.
  46. 46.Fabio Petroni, Aleksandra Piktus, Angela Fan, Patrick Lewis, Majid Yazdani, Nicola De Cao, James Thorne, Yacine Jernite, Vladimir Karpukhin, Jean Maillard, Vassilis Plachouras, Tim Rocktäschel, and Sebastian Riedel. 2021. KILT: a benchmark for knowledge intensive language tasks. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2523–2544, Online. Association for Computational Linguistics.
  47. 47.Barbara Plank. 2022. The “problem” of human label variation: On ground truth in data, modeling and evaluation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 10671–10682, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  48. 48.Joan Plepi, Béla Neuendorf, Lucie Flek, and Charles Welch. 2022. Unifying data perspectivism and personalization: An application to social norms. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 7391–7402, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  49. 49.Hongjin Qian, Xiaohe Li, Hanxun Zhong, Yu Guo, Yueyuan Ma, Yutao Zhu, Zhanliang Liu, Zhicheng Dou, and Ji-Rong Wen. 2021. Pchatbot: A large-scale dataset for personalized chatbot. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’21, page 2470–2477, New York, NY, USA. Association for Computing Machinery.
  50. 50.Omid Rafieian and Hema Yoganarasimhan. 2023. Ai and personalization. Artificial Intelligence in Marketing, pages 77–102.
  51. 51.Werner J. Reinartz and V. Kumar. 2000. On the profitability of long-life customers in a noncontractual setting: An empirical investigation and implications for marketing. Journal of Marketing, 64(4):17–35.
  52. 52.Werner J. Reinartz and V. Kumar. 2003. The impact of customer relationship characteristics on profitable lifetime duration. Journal of Marketing, 67(1):77–99.
  53. 53.Stephen Robertson, S. Walker, S. Jones, M. M. Hancock-Beaulieu, and M. Gatford. 1995. Okapi at trec-3. In Proceedings of the Third Text REtrieval Conference, TREC-3, pages 109–126. Gaithersburg, MD: NIST.
  54. 54.Paul Rottger, Bertie Vidgen, Dirk Hovy, and Janet Pierrehumbert. 2022. Two contrasting data annotation paradigms for subjective NLP tasks. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 175–190, Seattle, United States. Association for Computational Linguistics.
  55. 55.Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov. 2017. Membership inference attacks against machine learning models. In 2017 IEEE Symposium on Security and Privacy (SP), pages 3–18.
  56. 56.Nikita Soni, Matthew Matero, Niranjan Balasubramanian, and H. Andrew Schwartz. 2022. Human language modeling. In Findings of the Association for Computational Linguistics: ACL 2022, pages 622–636, Dublin, Ireland. Association for Computational Linguistics.
  57. 57.Shayan A. Tabrizi, Azadeh Shakery, Hamed Zamani, and Mohammad Ali Tavallaei. 2018. Person: Personalized information retrieval evaluation based on citation networks. Information Processing & Management, 54(4):630–656.
  58. 58.Jie Tang, Jing Zhang, Limin Yao, Juanzi Li, Li Zhang, and Zhong Su. 2008. Arnetminer: Extraction and mining of academic social networks. In KDD’08, pages 990–998.
  59. 59.Stojan Trajanovski, Chad Atalla, Kunho Kim, Vipul Agarwal, Milad Shokouhi, and Chris Quirk. 2021. When does text prediction benefit from additional context? an exploration of contextual signals for chat and email messages. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: Industry Papers, pages 1–9, Online. Association for Computational Linguistics.
  60. 60.Sebastian Vincent, Rowanne Sumner, Alice Dowek, Charlotte Blundell, Emily Preston, Chris Bayliss, Chris Oakley, and Carolina Scarton. 2023. Personalised language modelling of screen characters using rich metadata annotations. arXiv preprint arXiv:2303.16618.
  61. 61.Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2019. Superglue: A stickier benchmark for general-purpose language understanding systems. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc.
  62. 62.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018. GLUE: A multi-task benchmark and analysis platform for natural language understanding. In Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 353–355, Brussels, Belgium. Association for Computational Linguistics.
  63. 63.Charles Welch, Chenxi Gu, Jonathan K. Kummerfeld, Veronica Perez-Rosas, and Rada Mihalcea. 2022. Leveraging similar users for personalized language modeling with limited data. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1742–1752, Dublin, Ireland. Association for Computational Linguistics.
  64. 64.Yuwei Wu, Xuezhe Ma, and Diyi Yang. 2021. Personalized response generation via generative split memory network. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1956–1970, Online. Association for Computational Linguistics.
  65. 65.Joern Wuebker, Patrick Simianer, and John DeNero. 2018. Compact personalized models for neural machine translation. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 881–886, Brussels, Belgium. Association for Computational Linguistics.
  66. 66.Jiajing Xu, Andrew Zhai, and Charles Rosenberg. 2022. Rethinking personalized ranking at pinterest: An end-to-end approach. In Proceedings of the 16th ACM Conference on Recommender Systems, RecSys ’22, page 502–505, New York, NY, USA. Association for Computing Machinery.
  67. 67.Gui-Rong Xue, Jie Han, Yong Yu, and Qiang Yang. 2009. User language model for collaborative personalized search. ACM Trans. Inf. Syst., 27(2).
  68. 68.Hansi Zeng, Surya Kallumadi, Zaid Alibadi, Rodrigo Nogueira, and Hamed Zamani. 2023. A personalized dense retrieval framework for unified information access. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’23, New York, NY, USA. Association for Computing Machinery.
  69. 69.Saizheng Zhang, Emily Dinan, Jack Urbanek, Arthur Szlam, Douwe Kiela, and Jason Weston. 2018. Personalizing dialogue agents: I have a dog, do you have pets too? In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2204–2213, Melbourne, Australia. Association for Computational Linguistics.
  70. 70.Tianyi Zhang*, Varsha Kishore*, Felix Wu*, Kilian Q. Weinberger, and Yoav Artzi. 2020. Bertscore: Evaluating text generation with bert. In International Conference on Learning Representations.
  71. 71.Hanxun Zhong, Zhicheng Dou, Yutao Zhu, Hongjin Qian, and Ji-Rong Wen. 2022. Less is more: Learning to refine dialogue history for personalized dialogue generation. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5808–5820, Seattle, United States. Association for Computational Linguistics.
  72. 72.Wanjun Zhong, Duyu Tang, Jiahai Wang, Jian Yin, and Nan Duan. 2021. UserAdapter: Few-shot user learning in sentiment analysis. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 1484–1488, Online. Association for Computational Linguistics.
  73. 73.Jianing Zhou and Suma Bhat. 2021. Paraphrase generation: A survey of the state of the art. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 5075–5086, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  74. 74.Jian Zhu and David Jurgens. 2021. Idiosyncratic but not arbitrary: Learning idiolects in online registers reveals distinctive yet consistent individual styles. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing.

Citation

MLA
Salemi, A., et al. “LaMP: When Large Language Models Meet Personalization”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 7370–92, https://doi.org/10.18653/v1/2024.acl-long.399.
APA
Salemi, A., Mysore, S., Bendersky, M., & Zamani, H. (2024). LaMP: When Large Language Models Meet Personalization. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7370–7392. https://doi.org/10.18653/v1/2024.acl-long.399
Chicago
Salemi, A., S. Mysore, M. Bendersky, and H. Zamani. 2024. “LaMP: When Large Language Models Meet Personalization”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7370–92. https://doi.org/10.18653/v1/2024.acl-long.399.
Harvard
Salemi, A. et al. (2024) “LaMP: When Large Language Models Meet Personalization”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 7370–7392. Available at: https://doi.org/10.18653/v1/2024.acl-long.399.
Vancouver
1. Salemi A, Mysore S, Bendersky M, Zamani H (2024) LaMP: When Large Language Models Meet Personalization. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 7370–7392

BibTeX

@inproceedings{salemi-etal-2024-lamp,
    title = "{L}a{MP}: When Large Language Models Meet Personalization",
    author = "Salemi, Alireza  and
      Mysore, Sheshera  and
      Bendersky, Michael  and
      Zamani, Hamed",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.399/",
    doi = "10.18653/v1/2024.acl-long.399",
    pages = "7370--7392"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/