M-RAG: Reinforcing Large Language Model Performance through Retrieval-Augmented Generation with Multiple Partitions

Zheng WangShu Xian TeoJieer OuyangYongjun XuWei Shi

article2024ACL67 citations

Proposes a multi-agent reinforcement learning framework that partitions external retrieval databases to isolate relevant memories and eliminate noise, boosting language model performance by up to 12% across summarization, translation, and dialogue generation tasks.

Listen

Retrieval-Augmented Generation (RAG) is widely used to ground Large Language Models (LLMs) in factual external information, reducing hallucinations and improving response accuracy. However, conventional systems query an entire monolithic database at once, which frequently introduces irrelevant noise, increases computational latency, and dilutes the model's focus on essential information. As enterprise databases expand, performing coarse-grained searches over massive datasets creates critical performance bottlenecks and quality degradations across automated generative workflows.

The article introduces and evaluates a multiple partition paradigm for RAG, known as M-RAG. The objective is to demonstrate how dividing an external database into specialized sub-partitions and utilizing multi-agent reinforcement learning can optimize fine-grained memory retrieval and enhance downstream language generation performance without requiring fine-tuning of the underlying language models.

To evaluate this framework, the authors conducted comprehensive experiments across seven benchmark datasets spanning three language generation tasks: text summarization, machine translation, and multi-turn dialogue. They benchmarked the approach against leading retrieval methods across five diverse language model architectures, including Mixtral 8x7B, Llama 2 13B, Gemma 7B, Mistral 7B, and Phi-2 2.7B. The approach structures the retrieval pipeline using two lightweight, collaboratively trained reinforcement learning agents: Agent-S, which dynamically routes incoming queries to the most suitable partition, and Agent-R, which iteratively refines and evaluates retrieved demonstration memories before final text generation.

The findings confirm substantial performance improvements across all evaluated domains. When compared against the strongest baseline methods, M-RAG improved text summarization metrics by up to 11%, machine translation quality by approximately 8%, and dialogue generation relevance by 12%. The experiments demonstrated that querying partitioned subsets consistently yields better generation outcomes than querying a single monolithic database. Furthermore, index construction for smaller partitions proved significantly faster than indexing an entire database, while retrieval overhead remained minimal compared to the overall text generation latency.

These results demonstrate that database architecture and fine-grained retrieval routing are critical levers for improving generative AI accuracy. By keeping the core language model frozen and utilizing lightweight external agents to manage memory selection, organizations can achieve higher quality outputs without incurring the massive computational and financial costs of model re-training. Additionally, maintaining partitioned database structures naturally aligns with enterprise data privacy controls, access rights management, and distributed cloud computing systems.

Organizations implementing retrieval-augmented workflows should consider shifting from single-index vector stores to modular, partitioned architectures matched to specific domains, such as category-based or graph-indexed partitioning. Implementation teams should adopt intelligent routing and memory verification mechanisms before feeding retrieved context into generation models. Decision-makers should note that while M-RAG maintains fast runtime during live inference, initial training requires multiple queries to the language model to optimize agent policies. The reported evaluations were conducted using 4-bit quantized model weights, providing high confidence in the architectural trends while suggesting that deployment teams conduct task-specific indexing pilots to establish optimal partition counts.

arXiv: 2405.16420
Cover for M-RAG: Reinforcing Large Language Model Performance through Retrieval-Augmented Generation with Multiple Partitions

Abstract

Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by retrieving relevant memories from an external database. However, existing RAG methods typically organize all memories in a whole database, potentially limiting focus on crucial memories and introducing noise. In this paper, we introduce a multiple partition paradigm for RAG (called M-RAG), where each database partition serves as a basic unit for RAG execution. Based on this paradigm, we propose a novel framework that leverages LLMs with Multi-Agent Reinforcement Learning to optimize different language generation tasks explicitly. Through comprehensive experiments conducted on seven datasets, spanning three language generation tasks and involving three distinct language model architectures, we confirm that M-RAG consistently outperforms various baseline methods, achieving improvements of 11%, 8%, and 12% for text summarization, machine translation, and dialogue generation, respectively.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methodology
  • 3.1 Discussion on Partitioning a Database
  • 3.2 Agent-S: Selecting a Database Partition
  • 3.3 Agent-R: Refining Memories in the Selected Partition
  • 3.4 The M-RAG Framework
  • 4 Experiments
  • 4.1 Experimental Setup
  • 4.2 Experimental Results
  • 5 Conclusion and Limitations
  • References
  • A Appendix
  • A.1 Other Evaluation Metrics for Machine Translation
  • A.2 Further Discussion

Knowls

  1. Knowl 1 — Database Partitioning Strategies in Multi-Partition RAG

    model/method

    In the multiple partition retrieval-augmented generation (M-RAG) paradigm, an external database (training corpus) D={(xi,yi)}i=1∣D∣\mathbb{D} = \{(x_i, y_i)\}_{i=1}^{|\mathbb{D}|} containing input-output pairs (x,y)(x, y) (such as document-summary pairs for summarization or context-response pairs for dialogue) is divided into MM partitions D={Dm}m=1M\mathbb{D} = \{\mathbb{D}_m\}_{m=1}^M, where each partition Dm\mathbb{D}_m serves as an independent basic unit for retrieval and memory storage. Four database partitioning strategies are evaluated:

    1. Randomization: Employs probability amplification techniques such as Locality-Sensitive Hashing (LSH) to hash similar data vectors into the same partition bucket with high probability.
    2. Clustering: Clusters data vector representations using algorithms such as KK-means, which aligns with Inverted File Index (IVF) structures for approximate nearest neighbor search.
    3. Indexing: Uses graph partitioning with spectral clustering on navigable graph structures (such as Hierarchical Navigable Small World, HNSW, or Navigable Small World, NSW) to group topologically proximate regions into discrete partitions.
    4. Category: Assigns data vectors to partitions based on categorical metadata or labels inherent in the data (e.g., emotion or topic categories). A single item may belong to multiple partitions if its categories overlap.

    Empirically, indexing with M=4M=4 is optimal for text summarization, randomization with M=3M=3 for machine translation, and category-based partitioning with M=10M=10 for dialogue generation.

  2. Knowl 2 — Agent-S Formulation for Partition Selection

    model/method

    Agent-S (denoted Ag-S\text{Ag-S}) formulates database partition selection as a Markov Decision Process (MDP) / multi-armed bandit problem over MM database partitions D={Dm}m=1M\mathbb{D} = \{\mathbb{D}_m\}_{m=1}^M.

    • State Space: For an input pair (x,y)(x, y) and partitions Dm\mathbb{D}_m, the continuous state vector s(S)∈RMs^{(S)} \in \mathbb{R}^M measures the maximum semantic relevance between the query-target concatenation and the top retrieved memory in each partition:

    s(S)={max⁡(x~,y~)∈Dmsim(σ(x~⊕y~),σ(x⊕y))}m=1Ms^{(S)} = \left\{ \max_{(\tilde{x}, \tilde{y}) \in \mathbb{D}_m} \text{sim}(\sigma(\tilde{x} \oplus \tilde{y}), \sigma(x \oplus y)) \right\}_{m=1}^M

    where ⊕\oplus denotes concatenation, σ(⋅)\sigma(\cdot) is a text embedding model (such as CPT-Text), and sim(⋅,⋅)\text{sim}(\cdot, \cdot) is cosine similarity. During inference, when target yy and memory targets y~\tilde{y} are unavailable, the state is constructed using query embeddings alone: sim(σ(x~),σ(x))\text{sim}(\sigma(\tilde{x}), \sigma(x)).

    • Action Space: The action a(S)∈{1,…,M}a^{(S)} \in \{1, \dots, M\} represents selecting partition Dm\mathbb{D}_m for the retrieval-augmented generation pipeline:

    a(S)=m(1≤m≤M)a^{(S)} = m \quad (1 \le m \le M)

    • Policy and Rewards: The selection policy πθ(a(S)∣s(S))\pi_\theta(a^{(S)} \mid s^{(S)}) is parameterized by a Deep Q-Network (DQN) with parameters θ\theta. Because partition exploration does not immediately yield downstream generation feedback, the reward r(S)r^{(S)} is delayed and set to the cumulative improvement achieved by the memory refinement process:

    r(S)=Δ(hN,y)−Δ(h1,y)r^{(S)} = \Delta(h_N, y) - \Delta(h_1, y)

    where Δ(⋅,y)\Delta(\cdot, y) evaluates generation performance against reference yy, h1h_1 is the initial hypothesis generated from the retrieved partition memory, and hNh_N is the final refined hypothesis.

  3. Knowl 3 — Agent-R Formulation for Memory Candidate Refinement

    model/method

    Agent-R (denoted Ag-R\text{Ag-R}) formulates demonstration memory refinement within a selected partition Dm\mathbb{D}_m as a Markov Decision Process (MDP) optimized with Deep Q-Networks (DQN).

    Given an input query xx and the top retrieved memory (x~,y~)∈Dm(\tilde{x}, \tilde{y}) \in \mathbb{D}_m, Agent-R generates a pool of KK candidate responses using a frozen Large Language Model (LLM):

    C={y^k←LLM(x~)}k=1K\mathcal{C} = \left\{\hat{y}_k \leftarrow \text{LLM}(\tilde{x})\right\}_{k=1}^K

    • State Space: At step tt, given current generation hypothesis h←LLM(x⊕(x~,y~))h \leftarrow \text{LLM}(x \oplus (\tilde{x}, \tilde{y})), the state s(R)∈RKs^{(R)} \in \mathbb{R}^K computes the semantic similarity between the embedding of hypothesis hh and each candidate in C\mathcal{C}:

    s(R)={sim(σ(h),σ(y^k))}k=1Ks^{(R)} = \left\{ \text{sim}(\sigma(h), \sigma(\hat{y}_k)) \right\}_{k=1}^K

    • Action Space: The action a(R)∈{1,…,K}a^{(R)} \in \{1, \dots, K\} selects candidate memory y^k\hat{y}_k:

    a(R)=k(1≤k≤K)a^{(R)} = k \quad (1 \le k \le K)

    • Reward and Memory Update: When action at(R)a^{(R)}_t is taken, the demonstration memory becomes (x~,y^k)(\tilde{x}, \hat{y}_k), yielding a new hypothesis h′←LLM(x⊕(x~,y^k))h' \leftarrow \text{LLM}(x \oplus (\tilde{x}, \hat{y}_k)). Given evaluation metric Δ(⋅,⋅)\Delta(\cdot, \cdot) (such as ROUGE for summarization, BLEU for translation, or Distinct for dialogue):

    {r(R)=Δ(h′,y)−Δ(h,y),Dm.y~←y^k,h←h′if Δ(h′,y)>Δ(h,y)r(R)=0otherwise\begin{cases} r^{(R)} = \Delta(h', y) - \Delta(h, y), \quad \mathbb{D}_m.\tilde{y} \leftarrow \hat{y}_k, \quad h \leftarrow h' & \text{if } \Delta(h', y) > \Delta(h, y) \\ r^{(R)} = 0 & \text{otherwise} \end{cases}

    Over NN steps, the undiscounted cumulative reward sums telescopically:

    ∑t=2Nrt−1(R)=∑t=2N(Δ(ht,y)−Δ(ht−1,y))=Δ(hN,y)−Δ(h1,y)\sum_{t=2}^N r^{(R)}_{t-1} = \sum_{t=2}^N (\Delta(h_t, y) - \Delta(h_{t-1}, y)) = \Delta(h_N, y) - \Delta(h_1, y)

    which directly aligns cumulative reward maximization with maximizing the final hypothesis quality Δ(hN,y)\Delta(h_N, y).

  4. Knowl 4 — M-RAG Multi-Agent Reinforcement Learning Algorithm

    algorithm

    The M-RAG algorithm optimizes partition selection (Agent-S, policy πθ\pi_\theta) and memory refinement (Agent-R, policy πϕ\pi_\phi) via multi-agent DQN while keeping the LLM parameters frozen.

    Input: External database D\mathbb{D}, frozen LLM(⋅)\text{LLM}(\cdot), evaluation metric Δ(⋅,⋅)\Delta(\cdot, \cdot), embedding model σ(⋅)\sigma(\cdot), partition count MM, candidate pool size KK
    Output: Trained policies πθ(a(S)∣s(S))\pi_\theta(a^{(S)} \mid s^{(S)}) and πϕ(a(R)∣s(R))\pi_\phi(a^{(R)} \mid s^{(R)}), refined database D\mathbb{D}
    obtain D={Dm}m=1M\mathbb{D} = \{\mathbb{D}_m\}_{m=1}^M via partitioning strategy
    initialize policies πθ(a(S)∣s(S))\pi_\theta(a^{(S)} \mid s^{(S)}) and πϕ(a(R)∣s(R))\pi_\phi(a^{(R)} \mid s^{(R)})
    while not converged on validation set do
        sample text pair (x,y)(x, y) from training set
        construct s1(S)s^{(S)}_1 using (x,y)(x, y) on D\mathbb{D}
        for episode step i=1,2,…i = 1, 2, \dots do
            sample partition action m=ai(S)∼πθ(a∣si(S))m = a^{(S)}_i \sim \pi_\theta(a \mid s^{(S)}_i)
            ri(S)←0r^{(S)}_i \leftarrow 0
            retrieve top memory (x~,y~)∈Dm(\tilde{x}, \tilde{y}) \in \mathbb{D}_m
            generate initial hypothesis h←LLM(x⊕(x~,y~))h \leftarrow \text{LLM}(x \oplus (\tilde{x}, \tilde{y}))
            generate candidate pool C={y^k←LLM(x~)}k=1K\mathcal{C} = \{\hat{y}_k \leftarrow \text{LLM}(\tilde{x})\}_{k=1}^K
            construct s1(R)s^{(R)}_1 with hh on C\mathcal{C}
            for refinement step j=1,2,…j = 1, 2, \dots do
                sample candidate action k=aj(R)∼πϕ(a∣sj(R))k = a^{(R)}_j \sim \pi_\phi(a \mid s^{(R)}_j)
                generate candidate hypothesis h′←LLM(x⊕(x~,y^k))h' \leftarrow \text{LLM}(x \oplus (\tilde{x}, \hat{y}_k))
                if Δ(h′,y)>Δ(h,y)\Delta(h', y) > \Delta(h, y) then
                    rj(R)←Δ(h′,y)−Δ(h,y)r^{(R)}_j \leftarrow \Delta(h', y) - \Delta(h, y)
                    Dm.y~←y^k\mathbb{D}_m.\tilde{y} \leftarrow \hat{y}_k
                    h←h′h \leftarrow h'
                else
                    rj(R)←0r^{(R)}_j \leftarrow 0
                generate new candidate pool C\mathcal{C} and construct sj+1(R)s^{(R)}_{j+1} with hh
                ri(S)←ri(S)+rj(R)r^{(S)}_i \leftarrow r^{(S)}_i + r^{(R)}_j
                store transition (sj(R),aj(R),rj(R),sj+1(R))(s^{(R)}_j, a^{(R)}_j, r^{(R)}_j, s^{(R)}_{j+1}) in replay memory
            construct si+1(S)s^{(S)}_{i+1} by updating (x~,y~)(\tilde{x}, \tilde{y}) and (x,y)(x, y)
            store transition (si(S),ai(S),ri(S),si+1(S))(s^{(S)}_i, a^{(S)}_i, r^{(S)}_i, s^{(S)}_{i+1}) in replay memory
            sample minibatches from replay memories and optimize πθ,πϕ\pi_\theta, \pi_\phi via DQN
    Inference:
    Given input query xx:
    construct s(S)s^{(S)} using xx and retrieve candidates across {Dm}m=1M\{\mathbb{D}_m\}_{m=1}^M
    select partition m∗=arg⁡max⁡aπθ(a∣s(S))m^* = \arg\max_a \pi_\theta(a \mid s^{(S)})
    retrieve top memory (x~,y~)∈Dm∗(\tilde{x}, \tilde{y}) \in \mathbb{D}_{m^*}
    generate output y∗←LLM(x⊕(x~,y~))y^* \leftarrow \text{LLM}(x \oplus (\tilde{x}, \tilde{y}))

    Agent-S and Agent-R networks are two-layer feedforward neural networks: the hidden layer comprises 25 neurons with tanh activation, and the output layer comprises MM (or KK) neurons corresponding to the action space with a linear activation function.

  5. Knowl 5 — Computational Complexity of M-RAG vs Naive RAG

    theoretical result

    The computational complexity of M-RAG across indexing, retrieval, and generation stages compares to Naive RAG as follows:

    1. Indexing Complexity:

      • M-RAG: Constructing MM separate index partitions (e.g., using Hierarchical Navigable Small World graphs, HNSW) requires O(M⋅Nlog⁡N)\mathcal{O}(M \cdot N \log N) time, where MM is the number of partitions and NN is the maximum number of memories in any single partition.
      • Naive RAG: Organizing all N′=M⋅NN' = M \cdot N memories in a single global index requires O(N′log⁡N′)\mathcal{O}(N' \log N') time. Because N<N′N < N', M-RAG achieves faster index construction due to smaller graph sizes per partition.
    2. Retrieval Complexity:

      • M-RAG: Agent-S performs an Approximate kk-Nearest Neighbor (AKNN) search within each of the MM partitions, taking O(Mlog⁡N)\mathcal{O}(M \log N) time. Action sampling through the lightweight Agent-S feedforward neural network requires O(1)\mathcal{O}(1) time.
      • Naive RAG: Retrieval over the single full database takes O(log⁡N′)\mathcal{O}(\log N') time, which is marginally faster than M-RAG's O(Mlog⁡N)\mathcal{O}(M \log N) retrieval.
    3. Generation Complexity:

      • During Inference: Both M-RAG and Naive RAG perform single-pass generation using the frozen LLM, requiring O(E2)\mathcal{O}(E^2) time where EE is the number of generated tokens under Transformer attention.
      • During Training: Agent-R runs CC refinement iterations per episode, where each step queries the LLM to generate candidate pool items and hypotheses, incurring an additional training-time complexity of O(C⋅E2)\mathcal{O}(C \cdot E^2).
  6. Knowl 6 — Summarization Benchmark Performance of M-RAG

    data/table

    Evaluation of M-RAG against no-retrieval (None), Naive RAG, Selfmem, and Self-RAG on text summarization datasets XSum (BBC news articles) and BigPatent (US patent documents). Metrics reported are ROUGE-1 (R-1), ROUGE-2 (R-2), and ROUGE-L (R-L).

    LLM RAG XSum BigPatent
    R-1 R-2 R-L R-1 R-2 R-L
    Mixtral 8×\times7B None 25.40 6.39 18.30 47.41 16.63 25.14
    Mixtral 8×\times7B Naive 43.82 22.07 37.44 60.11 38.33 43.44
    Mixtral 8×\times7B Selfmem 44.67 22.38 37.86 64.12 39.21 46.21
    Mixtral 8×\times7B Self-RAG 44.01 22.26 37.51 63.59 38.65 45.25
    Mixtral 8×\times7B M-RAG 48.13 24.66 39.43 71.34 42.24 47.22
    Llama 2 13B M-RAG 37.18 18.02 26.44 60.31 37.33 33.47
    Phi-2 2.7B M-RAG 30.70 11.57 26.20 31.25 14.72 18.98

    With Mixtral 8×\times7B as the base generator, M-RAG outperforms the strongest baseline (Selfmem) by 7.7% on XSum R-1 (48.13 vs 44.67) and 11.3% on BigPatent R-1 (71.34 vs 64.12), with all improvements confirmed statistically significant (p<0.05p < 0.05).

  7. Knowl 7 — Machine Translation Benchmark Performance of M-RAG

    data/table

    Evaluation of M-RAG on the JRC-Acquis benchmark for four translation directions (Es→En\text{Es}\rightarrow\text{En}, En→Es\text{En}\rightarrow\text{Es}, De→En\text{De}\rightarrow\text{En}, En→De\text{En}\rightarrow\text{De}) across validation (Dev) and test sets using BLEU, BLEURT (BLEURT-20), and COMET (wmt22-comet-da).

    LLM RAG Es→\rightarrowEn En→\rightarrowEs De→\rightarrowEn En→\rightarrowDe
    Dev Test Dev Test Dev Test Dev Test
    Mixtral 8×\times7B None 34.34 34.81 32.60 28.32 43.75 44.09 43.78 42.24
    Mixtral 8×\times7B Naive 36.64 36.22 33.18 30.70 47.84 46.77 45.83 44.23
    Mixtral 8×\times7B Selfmem 37.65 37.11 34.12 31.86 48.08 47.31 51.38 49.81
    Mixtral 8×\times7B Self-RAG 37.17 36.82 33.80 31.61 47.99 47.27 50.10 48.75
    Mixtral 8×\times7B M-RAG 39.11 39.98 35.18 32.70 49.16 48.15 53.76 50.75
    Llama 2 13B M-RAG 30.41 30.03 26.40 22.03 41.10 42.22 45.98 42.58
    Phi-2 2.7B M-RAG 22.83 24.22 17.64 16.60 34.21 34.71 40.01 37.08

    On Mixtral 8×\times7B, M-RAG also achieves higher neural metrics than Selfmem across all pairs:

    • BLEURT: Es→En\text{Es}\rightarrow\text{En} 71.74 vs 63.63; En→Es\text{En}\rightarrow\text{Es} 63.66 vs 53.26; De→En\text{De}\rightarrow\text{En} 66.77 vs 59.93; En→De\text{En}\rightarrow\text{De} 70.99 vs 59.91.
    • COMET: Es→En\text{Es}\rightarrow\text{En} 82.66 vs 75.65; En→Es\text{En}\rightarrow\text{Es} 80.29 vs 55.28; De→En\text{De}\rightarrow\text{En} 67.33 vs 60.41; En→De\text{En}\rightarrow\text{De} 85.14 vs 52.13.
  8. Knowl 8 — Dialogue Generation Benchmark Performance of M-RAG

    data/table

    Evaluation on the DailyDialog dataset comparing M-RAG with baseline methods using BLEU-1 (B-1), BLEU-2 (B-2), Distinct-1 (D-1), and Distinct-2 (D-2). M-RAG(D) denotes the variant where the Distinct score is used as the reinforcement learning optimization metric Δ(⋅,⋅)\Delta(\cdot, \cdot) instead of BLEU.

    LLM RAG DailyDialog
    B-1 B-2 D-1 D-2
    Mixtral 8×\times7B None 15.52 7.05 61.49 89.51
    Mixtral 8×\times7B Naive 37.44 29.16 89.42 92.55
    Mixtral 8×\times7B Selfmem 38.16 29.92 89.23 95.23
    Mixtral 8×\times7B Self-RAG 37.76 29.79 88.24 95.34
    Mixtral 8×\times7B M-RAG 42.61 32.97 88.82 95.74
    Llama 2 13B M-RAG 31.29 17.63 63.19 88.20
    Phi-2 2.7B M-RAG 7.71 3.93 44.21 82.86
    Mixtral 8×\times7B M-RAG(D) 39.14 30.98 93.14 98.34

    M-RAG optimized for BLEU achieves 42.61 B-1 (an 11.7% relative improvement over Selfmem's 38.16). When optimized with Distinct metrics, M-RAG(D) achieves peak lexical diversity (93.14 D-1, 98.34 D-2) while maintaining high generation quality (39.14 B-1, 30.98 B-2).

  9. Knowl 9 — Performance of M-RAG Across 7B Open-Source LLMs

    data/table

    Comparison between Selfmem and M-RAG on Gemma 7B and Mistral 7B across text summarization (XSum), machine translation (JRC-Acquis Es→En\text{Es}\rightarrow\text{En}), and dialogue generation (DailyDialog).

    LLM RAG Summarization Translation Dialogue
    R-1 R-2 R-L BLEU B-1 B-2
    Gemma 7B Selfmem 31.38 9.97 25.07 24.61 15.56 7.91
    Gemma 7B M-RAG 33.81 12.93 27.82 26.92 18.15 9.95
    Mistral 7B Selfmem 35.40 12.68 27.06 26.26 18.28 10.05
    Mistral 7B M-RAG 37.47 13.24 30.49 32.65 24.52 11.53

    M-RAG consistently outperforms Selfmem on both 7B model architectures across all three generation tasks, demonstrating that the multi-partition reinforcement learning approach generalizes across model scales and families.

  10. Knowl 10 — Ablation Analysis and Hyperparameter Sensitivity of M-RAG

    empirical result

    An ablation study and hyperparameter analysis on the XSum summarization dataset with Mixtral 8×\times7B reveals the individual contributions of Agent-S and Agent-R, as well as the impact of partition count MM and candidate pool size KK.

    1. Ablation of Framework Components:

      • Full M-RAG: R-1 = 48.13, R-2 = 24.66, R-L = 39.43.
      • Without Agent-S (single database, M=1M=1): R-1 drops to 44.20 (-3.93), demonstrating that partitioning and selective routing prevent retrieval degradation caused by searching the entire vector space.
      • Without Agent-R (fixed greedy similarity rule instead of RL): R-1 drops to 45.75 (-2.38), confirming the importance of reinforcement-learned memory candidate refinement.
      • Without both agents (degrading to Naive RAG): R-1 drops to 43.82 (-4.31).
    2. Impact of Partition Count MM (Agent-S Action Space):

      • M=1M=1: R-1 = 44.20, index construction time = 299s, retrieval time = 0.61s, generation time = 83.59s.
      • M=2M=2: R-1 = 44.53, index construction time = 278s, retrieval time = 1.09s, generation time = 84.88s.
      • M=3M=3: R-1 = 46.27, index construction time = 257s, retrieval time = 1.54s, generation time = 82.81s.
      • M=4M=4: R-1 = 48.13, index construction time = 246s, retrieval time = 2.19s, generation time = 82.89s.
      • M=5M=5: R-1 = 47.21, index construction time = 227s, retrieval time = 2.59s, generation time = 86.64s. Increasing MM accelerates index construction across smaller sub-graphs while slightly increasing retrieval latency. M=4M=4 provides the optimal trade-off.
    3. Impact of Candidate Pool Size KK (Agent-R Action Space):

      • K=1K=1: R-1 = 45.81, pool generation time = 76s.
      • K=2K=2: R-1 = 46.54, pool generation time = 191s.
      • K=3K=3: R-1 = 48.13, pool generation time = 267s.
      • K=4K=4: R-1 = 48.18, pool generation time = 290s.
      • K=5K=5: R-1 = 48.25, pool generation time = 359s. Performance peaks and saturates at K=3K=3; larger KK provides marginal improvement at significantly higher prompt generation overhead.
  11. Knowl 11 — Stated Limitations of M-RAG

    limitation

    Two primary limitations of the M-RAG framework are identified:

    1. Evaluation on Quantized Models: Due to computational resource constraints, experiments are conducted with 4-bit quantized versions of base large language models (Mixtral 8×\times7B, Llama 2 13B, Phi-2 2.7B, Gemma 7B, Mistral 7B), although relative gains across RAG architectures are expected to remain consistent for full-precision models.
    2. Training Time Overhead: Even though LLM weights remain completely frozen and only lightweight policy networks for Agent-S and Agent-R are updated, training efficiency is bottlenecked by the requirement to query the LLM repeatedly to generate candidate pools and evaluate intermediate hypotheses (yielding O(C⋅E2)\mathcal{O}(C \cdot E^2) training-time complexity per episode).

Coverage note — No substantial contributed material was omitted; all partitioning methods, agent formulations, training/inference algorithms, complexity bounds, benchmark task results, ablation studies, and stated limitations are included.

References

  1. 1.Marah Abdin, Jyoti Aneja, ebastien Bubeck, and Caio Cesar Teodoro Mendes et al. 2023. Phi-2: The surprising power of small language models. https://www.microsoft.com/en-us/research/blog/phi-2-the-surprising-power-of-small-language-models.
  2. 2.Nathan Anderson, Caleb Wilson, and Stephen D. Richardson. 2022. Lingua: Addressing scenarios for live interpretation and automatic dubbing. In AMTA, pages 202–209.
  3. 3.Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi. 2023. Self-rag: Learning to retrieve, generate, and critique through self-reflection. CoRR, abs/2310.11511.
  4. 4.V. Blagojevi. 2023. Enhancing rag pipelines in haystack: Introducing diversityranker and lostinthemiddleranker. https://towardsdatascience.com/enhancing-rag-pipelines-in-haystack-45f14e2bc9f5.
  5. 5.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. NeurIPS, 33:1877–1901.
  6. 6.Howard Chen, Ramakanth Pasunuru, Jason Weston, and Asli Celikyilmaz. 2023. Walking down the memory maze: Beyond context limit through interactive reading. CoRR, abs/2310.05029.
  7. 7.Daixuan Cheng, Shaohan Huang, Junyu Bi, Yuefeng Zhan, Jianfeng Liu, Yujing Wang, Hao Sun, Furu Wei, Weiwei Deng, and Qi Zhang. 2023a. UPRISE: universal prompt retrieval for improving zero-shot evaluation. In EMNLP, pages 12318–12337.
  8. 8.Xin Cheng, Di Luo, Xiuying Chen, Lemao Liu, Dongyan Zhao, and Rui Yan. 2023b. Lift yourself up: Retrieval-augmented text generation with self memory. NeurIPS.
  9. 9.Woon Sang Cho, Pengchuan Zhang, Yizhe Zhang, Xiujun Li, Michel Galley, Chris Brockett, Mengdi Wang, and Jianfeng Gao. 2018. Towards coherent and cohesive long-form text generation. CoRR.
  10. 10.Zhuyun Dai, Vincent Y. Zhao, Ji Ma, Yi Luan, Jianmo Ni, Jing Lu, Anton Bakalov, Kelvin Guu, Keith B. Hall, and Ming-Wei Chang. 2023. Promptagator: Few-shot dense retrieval from 8 examples. In ICLR.
  11. 11.Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Qianyu Guo, Meng Wang, and Haofen Wang. 2023. Retrieval-augmented generation for large language models: A survey. CoRR, abs/2312.10997.
  12. 12.Siddharth Gollapudi, Neel Karia, Varun Sivashankar, Ravishankar Krishnaswamy, Nikit Begwani, Swapnil Raz, Yiyong Lin, Yin Zhang, Neelam Mahapatro, Premkumar Srinivasan, et al. 2023. Filtered-diskann: Graph algorithms for approximate nearest neighbor search with filters. In WWW, pages 3406–3416.
  13. 13.Jiatao Gu, Yong Wang, Kyunghyun Cho, and Victor O. K. Li. 2018. Search engine guided neural machine translation. In AAAI, pages 5133–5140. AAAI Press.
  14. 14.Rentong Guo, Xiaofan Luan, Long Xiang, Xiao Yan, Xiaomeng Yi, Jigao Luo, Qianya Cheng, Weizhi Xu, Jiarui Luo, Frank Liu, et al. 2022. Manu: a cloud native vector database management system. PVLDB, 15(12):3548–3561.
  15. 15.Yikun Han, Chunjiang Liu, and Pengfei Wang. 2023. A comprehensive survey on vector database: Storage and retrieval technique, challenge. CoRR.
  16. 16.Nabil Hossain, Marjan Ghazvininejad, and Luke Zettlemoyer. 2020. Simple and effective retrieve-edit-rerank text generation. In ACL, pages 2532–2538.
  17. 17.Piotr Indyk and Rajeev Motwani. 1998. Approximate nearest neighbors: towards removing the curse of dimensionality. In STOC, pages 604–613.
  18. 18.Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen, Craig Ferguson, Agata Lapedriza, Noah Jones, Shixiang Gu, and Rosalind Picard. 2019. Way off-policy batch deep reinforcement learning of implicit human preferences in dialog. CoRR.
  19. 19.Herve Jegou, Matthijs Douze, and Cordelia Schmid. 2010. Product quantization for nearest neighbor search. TPAMI, 33(1):117–128.
  20. 20.Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, and et al. 2023a. Mistral 7b. CoRR, abs/2310.06825.
  21. 21.Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, and Arthur Mensch et al. 2024. Mixtral of experts. CoRR, abs/2401.04088.
  22. 22.Huiqiang Jiang, Qianhui Wu, Chin-Yew Lin, Yuqing Yang, and Lili Qiu. 2023b. Llmlingua: Compressing prompts for accelerated inference of large language models. In EMNLP, pages 13358–13376.
  23. 23.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick S. H. Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. In EMNLP (1), pages 6769–6781.
  24. 24.Julia Kreutzer, Shahram Khadivi, Evgeny Matusov, and Stefan Riezler. 2018. Can neural machine translation be improved with user feedback? CoRR.
  25. 25.Carolin Lawrence and Stefan Riezler. 2018. Improving a neural semantic parser by counterfactual learning from human bandit feedback. CoRR.
  26. 26.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. NeurIPS, 33:9459–9474.
  27. 27.Jinpeng Li, Yingce Xia, Rui Yan, Hongda Sun, Dongyan Zhao, and Tie-Yan Liu. 2021. Stylized dialogue generation with multi-pass dual learning. In NeurIPS, pages 28470–28481.
  28. 28.Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan. 2016. A diversity-promoting objective function for neural conversation models. In HLT-NAACL, pages 110–119.
  29. 29.Xinze Li, Zhenghao Liu, Chenyan Xiong, Shi Yu, Yu Gu, Zhiyuan Liu, and Ge Yu. 2023. Structure-aware language model pretraining improves dense retrieval on structured data. In ACL (Findings), pages 11560–11574.
  30. 30.Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu. 2017. Dailydialog: A manually labelled multi-turn dialogue dataset. In IJCNLP(1), pages 986–995.
  31. 31.Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74–81.
  32. 32.Xi Victoria Lin, Xilun Chen, Mingda Chen, Weijia Shi, Maria Lomeli, Rich James, Pedro Rodriguez, Jacob Kahn, Gergely Szilvasy, Mike Lewis, Luke Zettlemoyer, and Scott Yih. 2023. RA-DIT: retrieval-augmented dual instruction tuning. CoRR, abs/2310.01352.
  33. 33.Ron Litman, Oron Anschel, Shahar Tsiper, Roee Litman, Shai Mazor, and R. Manmatha. 2020. SCATTER: selective context attentional scene text recognizer. In CVPR, pages 11959–11969.
  34. 34.Xinbei Ma, Yeyun Gong, Pengcheng He, Hai Zhao, and Nan Duan. 2023. Query rewriting for retrieval-augmented large language models. EMNLP, pages 5303–5315.
  35. 35.Yu A Malkov and Dmitry A Yashunin. 2018. Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs. TPAMI, 42(4):824–836.
  36. 36.Yury Malkov, Alexander Ponomarenko, Andrey Logvinov, and Vladimir Krylov. 2014. Approximate nearest neighbor algorithm based on navigable small world graphs. Information Systems, 45:61–68.
  37. 37.Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, and et al. 2024. Gemma: Open models based on gemini research and technology. CoRR, abs/2403.08295.
  38. 38.Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller. 2013. Playing atari with deep reinforcement learning. CoRR.
  39. 39.Reiichiro Nakano, Jacob Hilton, Suchir Balaji, and Jeff Wu et al. 2021. Webgpt: Browser-assisted question-answering with human feedback. CoRR, abs/2112.09332.
  40. 40.Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018. Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization. In EMNLP, pages 1797–1807.
  41. 41.Arvind Neelakantan, Tao Xu, Raul Puri, Alec Radford, Jesse Michael Han, Jerry Tworek, Qiming Yuan, Nikolas Tezak, Jong Wook Kim, Chris Hallacy, et al. 2022. Text and code embeddings by contrastive pretraining. CoRR.
  42. 42.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. NeurIPS, 35:27730–27744.
  43. 43.James Jie Pan, Jianguo Wang, and Guoliang Li. 2023. Survey of vector database management systems. CoRR.
  44. 44.Matt Post. 2018. A call for clarity in reporting BLEU scores. In WMT, pages 186–191.
  45. 45.Eva Sharma, Chen Li, and Lu Wang. 2019. BIGPATENT: A large-scale dataset for abstractive and coherent summarization. In ACL (1), pages 2204–2213.
  46. 46.Aleksandrs Slivkins et al. 2019. Introduction to multi-armed bandits. Foundations and Trends® in Machine Learning, 12(1-2):1–286.
  47. 47.Ralf Steinberger, Bruno Pouliquen, Anna Widiger, Camelia Ignat, Tomaz Erjavec, Dan Tufis, and Dániel Varga. 2006. The jrc-acquis: A multilingual aligned parallel corpus with 20+ languages. In LREC, pages 2142–2147.
  48. 48.Hugo Touvron, Louis Martin, Kevin Stone, and Peter Albert et al. 2023. Llama 2: Open foundation and fine-tuned chat models. CoRR, abs/2307.09288.
  49. 49.Shuohang Wang, Yichong Xu, Yuwei Fang, Yang Liu, Siqi Sun, Ruochen Xu, Chenguang Zhu, and Michael Zeng. 2022. Training data is more valuable than you think: A simple and effective method by retrieving from training data. In ACL, pages 3170–3179.
  50. 50.Xintao Wang, Qianwen Yang, Yongting Qiu, Jiaqing Liang, Qianyu He, Zhouhong Gu, Yanghua Xiao, and Wei Wang. 2023a. Knowledgpt: Enhancing large language models with retrieval and storage access on knowledge bases. CoRR, abs/2308.11761.
  51. 51.Zheng Wang, Bingzheng Gan, and Wei Shi. 2024. Multimodal query suggestion with multi-agent reinforcement learning from human feedback. In WWW, pages 1374–1385.
  52. 52.Zheng Wang, Cheng Long, Gao Cong, and Christian S. Jensen. 2023b. Collectively simplifying trajectories in a database: A query accuracy driven approach. CoRR, abs/2311.11204.
  53. 53.Zheng Wang, Cheng Long, Gao Cong, and Qianru Zhang. 2021. Error-bounded online trajectory simplification with multi-agent reinforcement learning. In KDD, pages 1758–1768.
  54. 54.Jeff Wu, Long Ouyang, Daniel M Ziegler, Nisan Stiennon, Ryan Lowe, Jan Leike, and Paul Christiano. 2021a. Recursively summarizing books with human feedback. CoRR.
  55. 55.Sixing Wu, Ying Li, Minghui Wang, Dawei Zhang, Yang Zhou, and Zhonghai Wu. 2021b. More is better: Enhancing open-domain dialogue generation via multi-source heterogeneous knowledge. In EMNLP, pages 2286–2300.
  56. 56.Sixing Wu, Ying Li, Dawei Zhang, and Zhonghai Wu. 2022. KSAM: infusing multi-source knowledge into dialogue generation via knowledge source aware multi-head decoding. In ACL (Findings), pages 353–363.
  57. 57.Shitao Xiao, Zheng Liu, Peitian Zhang, and Xingrun Xing. 2023. Lm-cocktail: Resilient tuning of language models via model merging. CoRR, abs/2311.13534.
  58. 58.Fangyuan Xu, Weijia Shi, and Eunsol Choi. 2023. RECOMP: improving retrieval-augmented lms with compression and selective augmentation. CoRR, abs/2310.04408.
  59. 59.Wenzhuo Xue, Hui Li, Yanguo Peng, Jiangtao Cui, and Yu Shi. 2017. Secure k nearest neighbors query for high-dimensional vectors in outsourced environments. IEEE TBD, 4(4):586–599.
  60. 60.Sanghyun Yi, Rahul Goel, Chandra Khatri, Alessandra Cervone, Tagyoung Chung, Behnam Hedayatnia, Anu Venkatesh, Raefer Gabriel, and Dilek Hakkani-Tur. 2019. Towards coherent and engaging spoken dialog response generation using automatic conversation evaluators. CoRR.
  61. 61.Shengyao Zhuang, Bing Liu, Bevan Koopman, and Guido Zuccon. 2023. Open-source large language models are strong zero-shot query likelihood models for document ranking. In EMNLP (Findings), pages 8807–8817.

Citation

MLA
Wang, Z., et al. “M-RAG: Reinforcing Large Language Model Performance Through Retrieval-Augmented Generation with Multiple Partitions”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 1966–78, https://doi.org/10.18653/v1/2024.acl-long.108.
APA
Wang, Z., Teo, S., Ouyang, J., Xu, Y., & Shi, W. (2024). M-RAG: Reinforcing Large Language Model Performance through Retrieval-Augmented Generation with Multiple Partitions. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1966–1978. https://doi.org/10.18653/v1/2024.acl-long.108
Chicago
Wang, Z., S. Teo, J. Ouyang, Y. Xu, and W. Shi. 2024. “M-RAG: Reinforcing Large Language Model Performance Through Retrieval-Augmented Generation with Multiple Partitions”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1966–78. https://doi.org/10.18653/v1/2024.acl-long.108.
Harvard
Wang, Z. et al. (2024) “M-RAG: Reinforcing Large Language Model Performance through Retrieval-Augmented Generation with Multiple Partitions”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 1966–1978. Available at: https://doi.org/10.18653/v1/2024.acl-long.108.
Vancouver
1. Wang Z, Teo S, Ouyang J, Xu Y, Shi W (2024) M-RAG: Reinforcing Large Language Model Performance through Retrieval-Augmented Generation with Multiple Partitions. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 1966–1978

BibTeX

@inproceedings{wang-etal-2024-rag,
    title = "{M}-{RAG}: Reinforcing Large Language Model Performance through Retrieval-Augmented Generation with Multiple Partitions",
    author = "Wang, Zheng  and
      Teo, Shu  and
      Ouyang, Jieer  and
      Xu, Yongjun  and
      Shi, Wei",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.108/",
    doi = "10.18653/v1/2024.acl-long.108",
    pages = "1966--1978"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/