Bridging the Preference Gap between Retrievers and LLMs
Zixuan KeWeize KongCheng LiMingyang ZhangQiaozhu MeiMichael Bendersky
Proposes a sequence-to-sequence bridge framework that connects frozen retrievers and large language models, training on supervised and reinforcement learning to select and reorder passages according to the model's actual context preferences rather than human ranking assumptions.
Retrieval-augmented generation enhances large language models by fetching external documents to answer queries accurately. However, existing systems treat the search component and the language model independently. Search engines are traditionally designed for human reading habits where ranking order is primary, whereas language models exhibit fundamentally different preferences: they can attend to information anywhere in their context window but are easily confused by irrelevant material. Modifying massive commercial models or enterprise search engines to fix this mismatch is often technically infeasible and cost-prohibitive.
The article demonstrates the existence of this "preference gap" between retrievers and language models and evaluates a novel framework called Bridging the Gap (BGM). The objective is to optimize the interface between search retrievers and language models without modifying either core component.
To achieve this, the authors placed a lightweight, sequence-to-sequence "bridge model" between a frozen retriever and a frozen language model. This intermediate model selects, reorders, and can even omit retrieved passages before passing them to the main generator. The bridge model is trained in two stages: first through supervised learning using a greedy search algorithm to identify high-performing passage combinations, and second through reinforcement learning using downstream task accuracy as a reward signal. The framework was evaluated across four benchmarks spanning open-domain question answering and personalized text generation tasks.
The investigation revealed four key findings. First, information selection impacts language model accuracy far more than ordering; randomizing the order of top retrieved passages altered performance by only about 1%, whereas changing which single passage was selected caused performance swings exceeding 5%. Second, the bridge framework consistently outperformed all standard retrieval baselines across all datasets, boosting exact-match accuracy on complex multi-hop questions from 25.80% to 35.64% (an absolute improvement of nearly 10 percentage points). Third, combining supervised learning with reinforcement learning proved essential, as supervised learning alone yielded inconsistent results. Finally, the bridge model successfully filtered out unhelpful or noisy context, choosing to provide zero retrieved passages when external data was irrelevant, thereby allowing the language model to rely correctly on its internal knowledge.
These findings indicate that traditional rerankers are insufficient for language model workflows because they score documents independently and fail to perform dynamic, sample-level selection. By selectively filtering out unneeded passages, this bridge approach not only boosts output quality but also lowers operating costs and processing latency by reducing prompt token volume.
Organizations developing generation systems should consider deploying dedicated bridge models rather than pursuing expensive fine-tuning of large models or retrievers. However, the article notes limitations: bridge models trained on a specific dataset or language model currently show reduced effectiveness when transferred to different domains or architectures. Additional research and pilot testing are recommended to improve domain generalization before applying the approach broadly across heterogeneous systems.
- Paper: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, Patrick Lewis et al. (2020). This foundational RAG paper establishes how retrievers and generators are combined, the setup this work adapts to optimize the retriever-to-LLM connection.
- Paper: Retrieval-Augmented Generation for Large Language Models: A Survey, Yunfan Gao et al. (2023). This survey maps RAG components and optimization approaches, providing the framework needed to situate the paper’s retriever–LLM bridge.
- Paper: Re2G: Retrieve, Rerank, Generate, Michael R. Glass et al. (2022). Re2G’s retrieve-rerank-generate pipeline supplies a concrete precedent for training and coordinating retrieval with generation, which the source further develops.
- Paper: RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs, Yue Yu et al. (2024). RankRAG extends the effort to align retrieval with generation by training one LLM to rerank retrieved contexts and answer from the selected evidence.
- Paper: A Multi-Task Embedder For Retrieval Augmented LLMs, Peitian Zhang et al. (2024). LLM-Embedder carries the idea of using downstream model utility to improve retrieval into a multi-task retriever spanning knowledge, memory, examples, and tools.
- Paper: DRAGIN: Dynamic Retrieval Augmented Generation based on the Real-time Information Needs of Large Language Models, Weihang Su et al. (2024). DRAGIN extends retriever–generator coordination by letting an LLM’s evolving information needs determine when to retrieve and what to search.
