M-RAG: Reinforcing Large Language Model Performance through Retrieval-Augmented Generation with Multiple Partitions
Zheng WangShu Xian TeoJieer OuyangYongjun XuWei Shi
Proposes a multi-agent reinforcement learning framework that partitions external retrieval databases to isolate relevant memories and eliminate noise, boosting language model performance by up to 12% across summarization, translation, and dialogue generation tasks.
Retrieval-Augmented Generation (RAG) is widely used to ground Large Language Models (LLMs) in factual external information, reducing hallucinations and improving response accuracy. However, conventional systems query an entire monolithic database at once, which frequently introduces irrelevant noise, increases computational latency, and dilutes the model's focus on essential information. As enterprise databases expand, performing coarse-grained searches over massive datasets creates critical performance bottlenecks and quality degradations across automated generative workflows.
The article introduces and evaluates a multiple partition paradigm for RAG, known as M-RAG. The objective is to demonstrate how dividing an external database into specialized sub-partitions and utilizing multi-agent reinforcement learning can optimize fine-grained memory retrieval and enhance downstream language generation performance without requiring fine-tuning of the underlying language models.
To evaluate this framework, the authors conducted comprehensive experiments across seven benchmark datasets spanning three language generation tasks: text summarization, machine translation, and multi-turn dialogue. They benchmarked the approach against leading retrieval methods across five diverse language model architectures, including Mixtral 8x7B, Llama 2 13B, Gemma 7B, Mistral 7B, and Phi-2 2.7B. The approach structures the retrieval pipeline using two lightweight, collaboratively trained reinforcement learning agents: Agent-S, which dynamically routes incoming queries to the most suitable partition, and Agent-R, which iteratively refines and evaluates retrieved demonstration memories before final text generation.
The findings confirm substantial performance improvements across all evaluated domains. When compared against the strongest baseline methods, M-RAG improved text summarization metrics by up to 11%, machine translation quality by approximately 8%, and dialogue generation relevance by 12%. The experiments demonstrated that querying partitioned subsets consistently yields better generation outcomes than querying a single monolithic database. Furthermore, index construction for smaller partitions proved significantly faster than indexing an entire database, while retrieval overhead remained minimal compared to the overall text generation latency.
These results demonstrate that database architecture and fine-grained retrieval routing are critical levers for improving generative AI accuracy. By keeping the core language model frozen and utilizing lightweight external agents to manage memory selection, organizations can achieve higher quality outputs without incurring the massive computational and financial costs of model re-training. Additionally, maintaining partitioned database structures naturally aligns with enterprise data privacy controls, access rights management, and distributed cloud computing systems.
Organizations implementing retrieval-augmented workflows should consider shifting from single-index vector stores to modular, partitioned architectures matched to specific domains, such as category-based or graph-indexed partitioning. Implementation teams should adopt intelligent routing and memory verification mechanisms before feeding retrieved context into generation models. Decision-makers should note that while M-RAG maintains fast runtime during live inference, initial training requires multiple queries to the language model to optimize agent policies. The reported evaluations were conducted using 4-bit quantized model weights, providing high confidence in the architectural trends while suggesting that deployment teams conduct task-specific indexing pilots to establish optimal partition counts.
- Paper: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, Patrick Lewis et al. (2020). This seminal paper introduces the fundamental retrieval-augmented generation paradigm that M-RAG restructures into a multi-partition architecture.
- Paper: Retrieval-Augmented Generation for Large Language Models: A Survey, Yunfan Gao et al. (2023). This survey provides a comprehensive taxonomy of advanced and modular RAG paradigms, framing the exact retrieval bottlenecks that M-RAG addresses.
- Paper: In-Context Retrieval-Augmented Language Models, Ori Ram et al. (2023). This work establishes how frozen language models can be augmented in-context with retrieved passages, laying the foundation for M-RAG's non-fine-tuned setup.
- Paper: Active Retrieval Augmented Generation, Zhengbao Jiang et al. (2023). This paper presents active decision-making strategies for retrieval during generation, contextualizing M-RAG's agentic retrieval routing.
- Paper: Evidentiality-guided Generation for Knowledge-Intensive NLP Tasks, Akari Asai et al. (2022). This study analyzes evidentiality and relevance filtering of retrieved context, which directly relates to the memory refinement role of Agent-R in M-RAG.
- Paper: Improving language models by retrieving from trillions of tokens, Sebastian Borgeaud et al. (2022). This paper examines the scaling and architectural challenges of querying multi-trillion token retrieval datastores, motivating M-RAG's partitioned database approach.
- Paper: When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories, Alex Troy Mallen et al. (2022). This empirical investigation defines when parametric memory fails and non-parametric retrieval is required, underpinning the core motivation of M-RAG.
- Paper: RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs, Yue Yu et al. (2024). RankRAG unifies context ranking and generation directly within the language model's instruction tuning, presenting an alternative to external routing agents.
- Paper: From Local to Global: A Graph RAG Approach to Query-Focused Summarization, Darren Edge et al. (2024). GraphRAG extends modular context organization by building hierarchical community-partitioned knowledge graphs for global sensemaking and summarization.
- Paper: Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning, Bowen Jin et al. (2025). Search-R1 applies reinforcement learning directly to teach language models autonomous search-and-reason trajectories, advancing the agentic retrieval concept.
- Paper: General Agentic Memory Via Deep Research, B. Y. Yan et al. (2025). General Agentic Memory utilizes dual-agent architectures to manage and research structured external memory over long-horizon tasks.
- Paper: Enhancing Noise Robustness of Retrieval-Augmented Language Models with Adaptive Adversarial Training, Feiteng Fang et al. (2024). This work explores adversarial training to make models resilient to noisy or irrelevant retrieved contexts, complementing partition-based noise reduction.
- Paper: REFRAG: Rethinking RAG based Decoding, Xiaoqiang Lin et al. (2025). REFRAG investigates context compression techniques to alleviate the decoding latency bottleneck incurred by long retrieved passages.
- Paper: Plan-on-Graph: Self-Correcting Adaptive Planning of Large Language Model on Knowledge Graphs, Liyi Chen et al. (2024). Plan-on-Graph builds upon structured knowledge routing by adaptively navigating graph partitions with self-correcting planning.
