Bridging the Preference Gap between Retrievers and LLMs

Zixuan KeWeize KongCheng LiMingyang ZhangQiaozhu MeiMichael Bendersky

article2024ACL93 citations

Proposes a sequence-to-sequence bridge framework that connects frozen retrievers and large language models, training on supervised and reinforcement learning to select and reorder passages according to the model's actual context preferences rather than human ranking assumptions.

Listen

Retrieval-augmented generation enhances large language models by fetching external documents to answer queries accurately. However, existing systems treat the search component and the language model independently. Search engines are traditionally designed for human reading habits where ranking order is primary, whereas language models exhibit fundamentally different preferences: they can attend to information anywhere in their context window but are easily confused by irrelevant material. Modifying massive commercial models or enterprise search engines to fix this mismatch is often technically infeasible and cost-prohibitive.

The article demonstrates the existence of this "preference gap" between retrievers and language models and evaluates a novel framework called Bridging the Gap (BGM). The objective is to optimize the interface between search retrievers and language models without modifying either core component.

To achieve this, the authors placed a lightweight, sequence-to-sequence "bridge model" between a frozen retriever and a frozen language model. This intermediate model selects, reorders, and can even omit retrieved passages before passing them to the main generator. The bridge model is trained in two stages: first through supervised learning using a greedy search algorithm to identify high-performing passage combinations, and second through reinforcement learning using downstream task accuracy as a reward signal. The framework was evaluated across four benchmarks spanning open-domain question answering and personalized text generation tasks.

The investigation revealed four key findings. First, information selection impacts language model accuracy far more than ordering; randomizing the order of top retrieved passages altered performance by only about 1%, whereas changing which single passage was selected caused performance swings exceeding 5%. Second, the bridge framework consistently outperformed all standard retrieval baselines across all datasets, boosting exact-match accuracy on complex multi-hop questions from 25.80% to 35.64% (an absolute improvement of nearly 10 percentage points). Third, combining supervised learning with reinforcement learning proved essential, as supervised learning alone yielded inconsistent results. Finally, the bridge model successfully filtered out unhelpful or noisy context, choosing to provide zero retrieved passages when external data was irrelevant, thereby allowing the language model to rely correctly on its internal knowledge.

These findings indicate that traditional rerankers are insufficient for language model workflows because they score documents independently and fail to perform dynamic, sample-level selection. By selectively filtering out unneeded passages, this bridge approach not only boosts output quality but also lowers operating costs and processing latency by reducing prompt token volume.

Organizations developing generation systems should consider deploying dedicated bridge models rather than pursuing expensive fine-tuning of large models or retrievers. However, the article notes limitations: bridge models trained on a specific dataset or language model currently show reduced effectiveness when transferred to different domains or architectures. Additional research and pilot testing are recommended to improve domain generalization before applying the approach broadly across heterogeneous systems.

Ke et al (2024).pdf
Cover for Bridging the Preference Gap between Retrievers and LLMs

Abstract

Large Language Models (LLMs) have demonstrated superior results across a wide range of tasks, and Retrieval-augmented Generation (RAG) is an effective way to enhance the performance by locating relevant information and placing it into the context window of the LLM. However, the relationship between retrievers and LLM in a RAG is still under-investigated. Most existing work treats the retriever and the LLM as independent components and leaves a gap between retrieving human-“friendly” information and assembling a LLM-“friendly” context. In this work, we examine a novel bridge mechanism. We validate the ranking and selection assumptions of retrievers in the context of RAG and propose a framework that chains together supervised and reinforcement learning to train a bridge model that optimizes the connection between the retriever and the LLM. Empirical results demonstrate the effectiveness of our method in both question-answering and personalized generation tasks.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Problem Formulation
  • 4 Training the Bridge Model
  • 4.1 Supervised Learning
  • 4.2 Reinforcement Learning
  • 5 Experiments
  • 5.1 Datasets and Baselines
  • 5.2 LLMs and Hyperparameters
  • 5.3 Evaluation Results and Analysis
  • 6 Conclusion
  • 7 Limitations
  • 8 Ethics Statement
  • References
  • A Case Studies

Knowls

  1. Knowl 1 — RAG ranking order matters less than passage selection in the tested setting

    empirical result

    The paper identifies a preference gap between retrievers, which rank passages for human-oriented use, and LLMs, which may respond differently to the passages selected for their context. In experiments with GTR retrieving passages and a frozen PaLM2-S answering from them, randomly reordering the top five passages changed performance by around 1%, whereas changing which single passage was supplied (the top-1 selection under different orders) produced performance variation exceeding 5% across the evaluated datasets. The top-1 effect was small on Natural Questions and larger, with either positive or negative direction, on other datasets. In this setting, passage selection had a larger effect than the ordering of multiple selected passages.

  2. Knowl 2 — BGM uses a trainable sequence model to adapt frozen retrieval results for a frozen LLM

    model/method

    Bridging the Gap between retrievers and LLMs (BGM) leaves both the retriever and final LLM fixed and trains a bridge model between them. Given a query xx and a corpus D={di}i=1mD=\{d_i\}_{i=1}^{m} of passages, the frozen dense retriever encodes queries and passages with EE and scores passage dd by s(d,x)=cos⁡(E(d),E(x))s(d,x)=\cos(E(d),E(x)). It returns the KK highest-scoring passages, denoted (djretr)j=1K(d_j^{\mathrm{retr}})_{j=1}^{K}. A trainable sequence-to-sequence bridge BB takes the query and retrieved passages and produces an adapted sequence, (djbdr)j=1n=B(x,(djretr)j=1K) (d_j^{\mathrm{bdr}})_{j=1}^{n}=B\bigl(x,(d_j^{\mathrm{retr}})_{j=1}^{K}\bigr). Here mm is corpus size, KK is the number retrieved, and nn is the output sequence length; nn can be smaller or larger than KK.

    Each retrieved passage is prefixed with a unique passage-ID token. The bridge generates IDs rather than passage text, and the IDs are mapped back to the original passage content before that content is passed to the LLM. The sequence-generation formulation jointly supports reordering, selecting a subset by stopping generation, and repeating passages by emitting an ID more than once. In the experiments, the retriever and LLM are frozen while the bridge learns which retrieved information and arrangement best serve the LLM.

  3. Knowl 3 — Greedy search synthesizes silver passage sequences for supervised training

    algorithm

    BGM creates supervised targets called silver passage sequences (SPS) by greedily building a passage sequence that improves downstream task performance. Given the retrieved candidate passages and an evaluator R(S)R(S) that scores RAG output using the downstream task metric, the procedure starts with the empty sequence; R(∅)R(\varnothing) is the score without retrieval augmentation. At each step it evaluates appending each candidate not already in the sequence, appends the candidate giving the highest score only if that score strictly exceeds the current sequence's score, and stops when no candidate improves the score. The resulting sequence is used as the target for bridge-model cross-entropy training. Because candidates are not added twice during this search, the SPS procedure itself does not synthesize repetitions.

    Input: query x, retrieved candidate passages P, downstream evaluator R
    Output: silver passage sequence S
    S <- empty sequence
    currentScore <- R(S)
    repeat:
        bestPassage <- none
        bestScore <- currentScore
        for each passage p in P that is not in S:
            candidate <- S with p appended
            score <- R(candidate)
            if score > bestScore:
                bestPassage <- p
                bestScore <- score
        if bestPassage is none:
            return S
        append bestPassage to S
        currentScore <- bestScore

    For KK unique retrieved candidates, the procedure adds at most KK passages and, in the worst case, evaluates at most K(K+1)/2K(K+1)/2 candidate extensions.

  4. Knowl 4 — BGM chains supervised learning and downstream-reward reinforcement learning

    model/method

    BGM trains its bridge in two stages. First, supervised learning uses the greedy-search silver passage sequence for each query as the target and trains the sequence-to-sequence bridge with cross-entropy loss. This provides an initialization that already performs passage ranking and selection. Second, reinforcement learning treats the supervised bridge as a policy that generates passage-ID sequences. The allowed actions are passage IDs, so the bridge organizes retrieved passages rather than rewriting their content. The downstream task score is the reward—Exact Match for question answering or BLEU for personalized generation in the reported experiments—and the reward is evaluated on the final LLM output. Unlike the SPS supervision, policy generation can explore sequences that repeat a passage. The paper describes optimizing this policy with an off-the-shelf RL method such as PPO; it does not establish that PPO is the only permissible optimizer.

  5. Knowl 5 — Evaluation covers open-domain QA and personalized generation with fixed retrieval and LLM components

    experimental setup

    Experiments use four datasets: Natural Questions (NQ) and multi-hop HotpotQA for open-domain question answering, plus Avocado Email and Amazon Book for personalized document completion using a user's previous writing as candidate passages. NQ and HotpotQA are evaluated with Exact Match (EM); Email and Book are evaluated with BLEU. The retriever returns K=5K=5 passages. The main LLM is frozen PaLM2-S with temperature 0 for deterministic generation. The bridge is FLAN-T5-XXL (11B) in the main experiments; its supervised training uses learning rate 0.001, linear warmup for 1,000 steps, square-root-normalized learning-rate decay, validation-based convergence, and beam search of size 4. The downstream metric is also used as the RL reward.

    The comparisons include Naive (no retrieval), GTR (the retrieved ranking), Random (the GTR passages in randomized order), and PSR (perplexity-distillation passage reranking without dynamic selection). The dataset splits and average prompt lengths, where prompt length includes the query and retrieved passages, are:

    Dataset Training Validation Test Average prompt tokens
    NQ 79,168 8,757 3,610 517.82
    HotpotQA 68,659 5,600 5,600 564.83
    Email 13,305 764 1,227 173.85
    Book 20,789 41,331 41,331 124.52
  6. Knowl 6 — BGM outperforms retrieval and reranking baselines, while manual top-k selection from PSR remains weaker

    empirical result

    Across the four datasets, BGM obtains the best score among the compared systems. Its advantage over GTR is especially large on HotpotQA, where selecting useful passages from multiple-hop evidence matters; the relatively similar GTR and Random scores on NQ also indicate that changing order alone has limited impact in that comparison. On Book, Naive already scores 11.50 BLEU, and BGM's gain is smaller than on the other datasets; nevertheless, BGM reaches 12.07 BLEU and selects no passages for roughly half the test examples (about 25,000), showing that its selection can omit retrieval when it is not useful.

    The PSR ablation tests whether a fixed choice of the top kk passages after reranking can substitute for BGM's learned, per-example selection. Each metric below is listed in the order NQ EM, HotpotQA EM, Email BLEU, Book BLEU. No fixed top-kk PSR setting matches BGM across the four datasets.

    Model NQ EM HotpotQA EM Email BLEU Book BLEU
    Naive 33.07 28.01 5.57 11.5
    Random 43.71 26.10 8.55 8.61
    GTR 43.79 25.80 9.76 8.75
    PSR 43.60 25.51 9.08 9.14
    PSR (Top1) 42.02 32.69 7.28 11.53
    PSR (Top2) 42.54 31.05 7.77 10.11
    PSR (Top3) 42.85 32.71 8.21 9.70
    PSR (Top4) 43.71 32.37 8.26 9.11
    BGM 45.37 35.64 10.42 12.07
  7. Knowl 7 — Greedy task-scored SPS gives better bridge training targets than GTR or PSR sequences

    empirical result

    The paper compares alternative passage sequences used as supervised targets for training the bridge: the GTR retrieval ranking, the PSR reranked sequence, and the task-scored greedy SPS. The bridge trained from greedy SPS achieves the best reported result on all four datasets. PSR improves over GTR on HotpotQA, Email, and Book, but is slightly lower on NQ; therefore, target-sequence quality affects performance, and the greedy target is strongest in these experiments.

    Silver sequence used for training NQ EM HotpotQA EM Email BLEU Book BLEU
    GTR 43.79 25.80 9.76 8.75
    PSR 43.68 29.73 10.1 10.35
    Greedy (BGM) 45.37 35.64 10.42 12.07
  8. Knowl 8 — Reinforcement learning adds performance beyond supervised bridge training

    empirical result

    An ablation compares the complete BGM training pipeline (supervised learning followed by reinforcement learning) with BGM trained using supervised learning alone and with GTR. Adding RL improves the supervised-only result on all four datasets, although the Book gain is small: 12.07 versus 12.05 BLEU. The gains are larger on NQ, HotpotQA, and Email, supporting RL's role in optimizing for downstream task output rather than relying only on the silver sequences.

    Model NQ EM HotpotQA EM Email BLEU Book BLEU
    GTR 43.79 25.8 9.76 8.75
    BGM (SL only) 39.44 34.26 8.62 12.05
    BGM 45.37 35.64 10.42 12.07
  9. Knowl 9 — BGM benefits from bridge capacity and also improves results with a smaller LLM

    empirical result

    The bridge-size study compares FLAN-T5-Large, FLAN-T5-XL, and FLAN-T5-XXL (11B) against GTR. Each bridge size scores above GTR on every dataset, although the largest bridge is not best on every individual metric: FLAN-T5-Large has the highest HotpotQA EM (35.87), while FLAN-T5-XXL leads on NQ, Email, and Book. The LLM-size study compares PaLM2-XXS and PaLM2-S on NQ and HotpotQA. BGM beats all listed non-BGM baselines for both LLM sizes; these results indicate that the bridge is useful with the smaller LLM as well.

    Bridge model NQ EM HotpotQA EM Email BLEU Book BLEU
    GTR 43.79 25.80 9.76 8.75
    FLAN-T5-Large 44.15 35.87 10.18 10.19
    FLAN-T5-XL 44.87 35.41 9.64 10.7
    FLAN-T5-XXL 45.37 35.64 10.42 12.07
    LLM Model NQ EM HotpotQA EM
    PaLM2-XXS Naive 12.13 14.57
    PaLM2-XXS Random 31.19 24.41
    PaLM2-XXS GTR 31.91 23.17
    PaLM2-XXS PSR 32.04 22.53
    PaLM2-XXS BGM 39.88 28.69
    PaLM2-S Naive 33.07 28.01
    PaLM2-S Random 43.71 26.10
    PaLM2-S GTR 43.79 25.80
    PaLM2-S PSR 43.60 25.51
    PaLM2-S BGM 45.37 35.64
  10. Knowl 10 — BGM's cross-dataset and cross-LLM generalization is limited

    limitation

    BGM performs worse when transferred to an unseen dataset than when trained and evaluated in-domain, and transferring a bridge trained with PaLM2-S to PaLM2-XXS also reduces performance. In the results below, in-domain training uses PaLM2-S; the first block reports transfer across datasets, and the second compares bridges trained and tested with different LLM sizes. Dashes indicate results not reported. The authors identify dataset/domain generalization as limited and leave improved generalization methods for future work. They also note that the supervised silver sequences are generated greedily and that alternative synthesis approaches remain to be explored.

    Training condition and test NQ EM HotpotQA EM Email BLEU Book BLEU
    Test on PaLM2-S
    BGM (in-domain) 45.37 35.64 10.42 12.07
    BGM (trained on NQ) – 33.42 5.66 11.22
    BGM (trained on Email) 35.59 27.98 – 11.38
    Test on PaLM2-XXS
    BGM (trained with PaLM2-XXS) 39.88 28.69 – –
    BGM (trained with PaLM2-S) 30.63 24.55 – –

Coverage note — The paper's case studies are omitted because they illustrate individual retrieval decisions rather than establishing an additional general result; private-email examples are also unavailable for display because of data-license restrictions.

References

  1. 1.Leonard Adolphs, Benjamin Boerschinger, Christian Buck, Michelle Chen Huebscher, Massimiliano Ciaramita, Lasse Espeholt, Thomas Hofmann, Yannic Kilcher, Sascha Rothe, Pier Giuseppe Sessa, et al. 2021. Boosting search engines with interactive agents. arXiv preprint arXiv:2109.00527.
  2. 2.Rohan Anil, Andrew M. Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, Eric Chu, Jonathan H. Clark, Laurent El Shafey, Yanping Huang, Kathy Meier-Hellstern, Gaurav Mishra, Erica Moreira, Mark Omernick, Kevin Robinson, Sebastian Ruder, Yi Tay, Kefan Xiao, Yuanzhong Xu, Yujing Zhang, Gustavo Hernández Ábrego, Junwhan Ahn, Jacob Austin, Paul Barham, Jan A. Botha, James Bradbury, Siddhartha Brahma, Kevin Brooks, Michele Catasta, Yong Cheng, Colin Cherry, Christopher A. Choquette-Choo, Aakanksha Chowdhery, Clément Crepy, Shachi Dave, Mostafa Dehghani, Sunipa Dev, Jacob Devlin, Mark Díaz, Nan Du, Ethan Dyer, Vladimir Feinberg, Fangxiaoyu Feng, Vlad Fienber, Markus Freitag, Xavier Garcia, Sebastian Gehrmann, Lucas Gonzalez, and et al. 2023. Palm 2 technical report. CoRR, abs/2305.10403.
  3. 3.Andrea Bacciu, Florin Cocunasu, Federico Siciliano, Fabrizio Silvestri, Nicola Tonellotto, and Giovanni Trappolini. 2023. Rraml: Reinforced retrieval augmented machine learning. arXiv preprint arXiv:2307.12798.
  4. 4.Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, et al. 2022. Improving language models by retrieving from trillions of tokens. In International conference on machine learning, pages 2206–2240. PMLR.
  5. 5.Michiel De Jong, Yury Zemlyanskiy, Nicholas FitzGerald, Joshua Ainslie, Sumit Sanghai, Fei Sha, and William W Cohen. 2023. Pre-computed memory or on-the-fly encoding? a hybrid approach to retrieval augmentation makes the most of your compute. In International Conference on Machine Learning, pages 7329–7342. PMLR.
  6. 6.Michiel de Jong, Yury Zemlyanskiy, Nicholas FitzGerald, Sumit Sanghai, William W Cohen, and Joshua Ainslie. 2023. Glimmer: generalized late-interaction memory reranker. arXiv preprint arXiv:2306.10231.
  7. 7.Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. 2020. Retrieval augmented language model pre-training. In International conference on machine learning, pages 3929–3938. PMLR.
  8. 8.Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bojanowski, Armand Joulin, and Edouard Grave. 2021. Unsupervised dense information retrieval with contrastive learning. arXiv preprint arXiv:2112.09118.
  9. 9.Gautier Izacard and Edouard Grave. 2020. Leveraging passage retrieval with generative models for open domain question answering. arXiv preprint arXiv:2007.01282.
  10. 10.Gautier Izacard, Patrick Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel, and Edouard Grave. 2022. Few-shot learning with retrieval augmented language models. arXiv preprint arXiv:2208.03299.
  11. 11.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. arXiv preprint arXiv:2004.04906.
  12. 12.Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. 2020. Generalization through memorization: Nearest neighbor language models. In International Conference on Learning Representations.
  13. 13.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019. Natural questions: A benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:452–466.
  14. 14.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing Systems, 33:9459–9474.
  15. 15.Cheng Li, Mingyang Zhang, Qiaozhu Mei, Yaqing Wang, Spurthi Amba Hombaiah, Yi Liang, and Michael Bendersky. 2023. Teach llms to personalize–an approach inspired by writing education. arXiv preprint arXiv:2308.07968.
  16. 16.Nelson F Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2023. Lost in the middle: How language models use long contexts. arXiv preprint arXiv:2307.03172.
  17. 17.Jianmo Ni, Jiacheng Li, and Julian McAuley. 2019. Justifying recommendations using distantly-labeled reviews and fine-grained aspects. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 188–197, Hong Kong, China. Association for Computational Linguistics.
  18. 18.Jianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai, Gustavo Hernández Ábrego, Ji Ma, Vincent Y Zhao, Yi Luan, Keith B Hall, Ming-Wei Chang, et al. 2021. Large dual encoders are generalizable retrievers. arXiv preprint arXiv:2112.07899.
  19. 19.Rodrigo Nogueira and Kyunghyun Cho. 2017. Task-oriented query reformulation with reinforcement learning. arXiv preprint arXiv:1704.04572.
  20. 20.Douglas Oard, William Webber, David Kirsch, and Sergey Golitsynskiy. 2015. Avocado research email collection. Philadelphia: Linguistic Data Consortium.
  21. 21.OpenAI. 2023. Gpt-4 technical report.
  22. 22.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020a. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21(1):5485–5551.
  23. 23.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020b. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res.
  24. 24.Stephen E Robertson. 1977. The probability ranking principle in ir. Journal of documentation.
  25. 25.Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed H Chi, Nathanael Schärli, and Denny Zhou. 2023a. Large language models can be easily distracted by irrelevant context. In International Conference on Machine Learning, pages 31210–31227. PMLR.
  26. 26.Weijia Shi, Sewon Min, Michihiro Yasunaga, Minjoon Seo, Rich James, Mike Lewis, Luke Zettlemoyer, and Wen-tau Yih. 2023b. Replug: Retrieval-augmented black-box language models. arXiv preprint arXiv:2301.12652.
  27. 27.Zeng Wei, Jun Xu, Yanyan Lan, Jiafeng Guo, and Xueqi Cheng. 2017. Reinforcement learning to rank with markov decision process. In Proceedings of the 40th international ACM SIGIR conference on research and development in information retrieval.
  28. 28.Yuhuai Wu, Markus N Rabe, DeLesley Hutchins, and Christian Szegedy. 2022. Memorizing transformers. arXiv preprint arXiv:2203.08913.
  29. 29.Zeqiu Wu, Yi Luan, Hannah Rashkin, David Reitter, Hannaneh Hajishirzi, Mari Ostendorf, and Gaurav Singh Tomar. 2021. Conqrr: Conversational query rewriting for retrieval with reinforcement learning. arXiv preprint arXiv:2112.08558.
  30. 30.Long Xia, Jun Xu, Yanyan Lan, Jiafeng Guo, Wei Zeng, and Xueqi Cheng. 2017. Adapting markov decision process for search result diversification. In Proceedings of the 40th international ACM SIGIR conference on research and development in information retrieval.
  31. 31.Fangyuan Xu, Weijia Shi, and Eunsol Choi. 2023. Recomp: Improving retrieval-augmented lms with compression and selective augmentation. arXiv preprint arXiv:2310.04408.
  32. 32.Jun Xu, Zeng Wei, Long Xia, Yanyan Lan, Dawei Yin, Xueqi Cheng, and Ji-Rong Wen. 2020. Reinforcement learning to rank with pairwise policy gradient. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 509–518.
  33. 33.Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. Hotpotqa: A dataset for diverse, explainable multi-hop question answering. In EMNLP.
  34. 34.Michihiro Yasunaga, Armen Aghajanyan, Weijia Shi, Richard James, Jure Leskovec, Percy Liang, Mike Lewis, Luke Zettlemoyer, and Wen-tau Yih. 2023. Retrieval-augmented multimodal language modeling.
  35. 35.Wei Zeng, Jun Xu, Yanyan Lan, Jiafeng Guo, and Xueqi Cheng. 2018. Multi page search with reinforcement learning to rank. In Proceedings of the 2018 ACM SIGIR international conference on theory of information retrieval, pages 175–178.

Citation

MLA
Ke, Z., et al. “Bridging the Preference Gap Between Retrievers and LLMs”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 10438–51, https://doi.org/10.18653/v1/2024.acl-long.562.
APA
Ke, Z., Kong, W., Li, C., Zhang, M., Mei, Q., & Bendersky, M. (2024). Bridging the Preference Gap between Retrievers and LLMs. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 10438–10451. https://doi.org/10.18653/v1/2024.acl-long.562
Chicago
Ke, Z., W. Kong, C. Li, M. Zhang, Q. Mei, and M. Bendersky. 2024. “Bridging the Preference Gap Between Retrievers and LLMs”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 10438–51. https://doi.org/10.18653/v1/2024.acl-long.562.
Harvard
Ke, Z. et al. (2024) “Bridging the Preference Gap between Retrievers and LLMs”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 10438–10451. Available at: https://doi.org/10.18653/v1/2024.acl-long.562.
Vancouver
1. Ke Z, Kong W, Li C, Zhang M, Mei Q, Bendersky M (2024) Bridging the Preference Gap between Retrievers and LLMs. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 10438–10451

BibTeX

@inproceedings{ke-etal-2024-bridging,
    title = "Bridging the Preference Gap between Retrievers and {LLM}s",
    author = "Ke, Zixuan  and
      Kong, Weize  and
      Li, Cheng  and
      Zhang, Mingyang  and
      Mei, Qiaozhu  and
      Bendersky, Michael",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.562/",
    doi = "10.18653/v1/2024.acl-long.562",
    pages = "10438--10451"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/