MAIN-RAG: Multi-Agent Filtering Retrieval-Augmented Generation
Chia-Yuan ChangZhimeng JiangVineeth RakeshMenghai PanChin-Chia Michael YehGuanchu WangMingzhi HuZhichao XuYan ZhengMahashweta Das
Proposes a training-free multi-agent framework that uses dynamic score thresholds to filter out noisy retrieved documents, boosting question-answering accuracy by up to 11% without model fine-tuning.
Large language models often struggle with generating outdated or factually inaccurate information, known as hallucinations. Retrieval-augmented generation addresses this issue by pulling relevant external knowledge to ground responses. However, existing systems frequently retrieve irrelevant or noisy documents. These low-quality inputs mislead the model, degrade response accuracy, and increase unnecessary computational overhead.
The article's main objective is to introduce and evaluate MAIN-RAG, a training-free framework that uses multiple collaborative software agents to score, filter, and rank retrieved documents before generating a final answer.
To evaluate this approach, the authors tested MAIN-RAG across four standard question-answering benchmarks, including scientific reasoning, open-domain knowledge retrieval, and long-form answer generation. The framework deploys three roles powered by standard pre-trained language models without any model fine-tuning or extra training data. An initial predictor agent drafts answers based on retrieved documents, a judge agent assigns relevance scores to each document-query-answer pairing using statistical token probabilities, and an adaptive threshold dynamically filters out low-scoring documents before a final predictor generates the response.
The findings show that MAIN-RAG consistently outperforms traditional training-free baselines across all four benchmark datasets, increasing answer accuracy by 2% to 11%. Improvements were especially prominent on queries involving rare, long-tail knowledge—such as in the PopQA dataset—where retrieval noise is typically high. Furthermore, MAIN-RAG closed the performance gap with computationally expensive, fine-tuned models and occasionally exceeded them in text quality and fluency metrics. The analysis also confirmed that sorting filtered documents in descending order of relevance yields significantly higher and more stable accuracy than random or ascending order.
These results indicate that organizations can substantially improve artificial intelligence reliability and reduce hallucination risks without investing in expensive model retraining or custom data labeling. Dynamically filtering noise at inference time protects output quality and optimizes computational resources by feeding only high-value context into the final model.
Organizations implementing retrieval pipelines should adopt multi-agent filtering and dynamic scoring as a cost-effective alternative to model fine-tuning. Before wide-scale deployment, teams should conduct pilot testing on their domain-specific datasets to assess latency trade-offs, as multi-agent processing requires several sequential model calls. Future work should explore integrating fine-grained scoring thresholds and human feedback mechanisms.
While confidence in the benchmark gains is high, the evaluations were confined to open-source models across four specific benchmarks, meaning performance may vary under different organizational workflows or proprietary architectures. Additionally, leaders should consider that multiple sequential model inferences increase total computing consumption and environmental impact, requiring careful architectural optimization in high-volume production environments.
- Paper: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, Patrick Lewis et al. (2020). Read this foundational RAG paper first to understand the retrieve-then-generate pipeline that MAIN-RAG modifies by filtering retrieved context before generation.
- Paper: Re2G: Retrieve, Rerank, Generate, Michael R. Glass et al. (2022). Re2G establishes the retrieve-rerank-generate pattern, making its intermediate passage-ranking stage a direct precursor to MAIN-RAG’s filtering and ranking pipeline.
- Paper: Evidentiality-guided Generation for Knowledge-Intensive NLP Tasks, Akari Asai et al. (2022). Its evidentiality-guided approach shows how judging whether passages support an answer can improve generation, a key premise behind MAIN-RAG’s document scoring.
- Paper: Chain-of-Note: Enhancing Robustness in Retrieval-Augmented Language Models, Wenhao Yu et al. (2024). Chain-of-Note introduces evaluating retrieved documents for relevance and credibility before answering, preparing readers for MAIN-RAG’s pre-generation filtering step.
- Paper: Improving Passage Retrieval with Zero-Shot Question Generation, Devendra Singh Sachan et al. (2022). UPR demonstrates using language-model token likelihoods to rerank retrieved passages, clarifying the probability-based scoring idea used by MAIN-RAG’s judge agent.
No sufficiently relevant recommendations were found.
