MAIN-RAG: Multi-Agent Filtering Retrieval-Augmented Generation

Chia-Yuan ChangZhimeng JiangVineeth RakeshMenghai PanChin-Chia Michael YehGuanchu WangMingzhi HuZhichao XuYan ZhengMahashweta Das

article2025ACL52 citations

Proposes a training-free multi-agent framework that uses dynamic score thresholds to filter out noisy retrieved documents, boosting question-answering accuracy by up to 11% without model fine-tuning.

Listen

Large language models often struggle with generating outdated or factually inaccurate information, known as hallucinations. Retrieval-augmented generation addresses this issue by pulling relevant external knowledge to ground responses. However, existing systems frequently retrieve irrelevant or noisy documents. These low-quality inputs mislead the model, degrade response accuracy, and increase unnecessary computational overhead.

The article's main objective is to introduce and evaluate MAIN-RAG, a training-free framework that uses multiple collaborative software agents to score, filter, and rank retrieved documents before generating a final answer.

To evaluate this approach, the authors tested MAIN-RAG across four standard question-answering benchmarks, including scientific reasoning, open-domain knowledge retrieval, and long-form answer generation. The framework deploys three roles powered by standard pre-trained language models without any model fine-tuning or extra training data. An initial predictor agent drafts answers based on retrieved documents, a judge agent assigns relevance scores to each document-query-answer pairing using statistical token probabilities, and an adaptive threshold dynamically filters out low-scoring documents before a final predictor generates the response.

The findings show that MAIN-RAG consistently outperforms traditional training-free baselines across all four benchmark datasets, increasing answer accuracy by 2% to 11%. Improvements were especially prominent on queries involving rare, long-tail knowledge—such as in the PopQA dataset—where retrieval noise is typically high. Furthermore, MAIN-RAG closed the performance gap with computationally expensive, fine-tuned models and occasionally exceeded them in text quality and fluency metrics. The analysis also confirmed that sorting filtered documents in descending order of relevance yields significantly higher and more stable accuracy than random or ascending order.

These results indicate that organizations can substantially improve artificial intelligence reliability and reduce hallucination risks without investing in expensive model retraining or custom data labeling. Dynamically filtering noise at inference time protects output quality and optimizes computational resources by feeding only high-value context into the final model.

Organizations implementing retrieval pipelines should adopt multi-agent filtering and dynamic scoring as a cost-effective alternative to model fine-tuning. Before wide-scale deployment, teams should conduct pilot testing on their domain-specific datasets to assess latency trade-offs, as multi-agent processing requires several sequential model calls. Future work should explore integrating fine-grained scoring thresholds and human feedback mechanisms.

While confidence in the benchmark gains is high, the evaluations were confined to open-source models across four specific benchmarks, meaning performance may vary under different organizational workflows or proprietary architectures. Additionally, leaders should consider that multiple sequential model inferences increase total computing consumption and environmental impact, requiring careful architectural optimization in high-volume production environments.

arXiv: 2501.00332

No sufficiently relevant recommendations were found.

Cover for MAIN-RAG: Multi-Agent Filtering Retrieval-Augmented Generation

Abstract

Large Language Models (LLMs) are becoming essential tools for various natural language processing tasks but often suffer from generating outdated or incorrect information. Retrieval-Augmented Generation (RAG) addresses this issue by incorporating external, real-time information retrieval to ground LLM responses. However, the existing RAG systems frequently struggle with the quality of retrieval documents, as irrelevant or noisy documents degrade performance, increase computational overhead, and undermine response reliability. To tackle this problem, we propose Multi-Agent Filtering Retrieval-Augmented Generation (MAIN-RAG), a training-free RAG framework that leverages multiple LLM agents to collaboratively filter and score retrieved documents. Specifically, MAIN-RAG introduces an adaptive filtering mechanism that dynamically adjusts the relevance filtering threshold based on score distributions, effectively minimizing noise while maintaining high recall of relevant documents. The proposed approach leverages inter-agent consensus to ensure robust document selection without requiring additional training data or fine-tuning. Experimental results across four QA benchmarks demonstrate that MAIN-RAG consistently outperforms traditional RAG approaches, achieving a 2–11% improvement in answer accuracy while reducing the number of irrelevant retrieved documents. Quantitative analysis further reveals that our approach achieves superior response consistency and answer accuracy over baseline methods, offering a competitive and practical alternative to training-based solutions.

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 2.1 Notations and Objectives
  • 2.2 Impact of Noisy Retrieval Documents
  • 2.3 Related Work
  • 3 Multi-Agent Filtering RAG (MAIN-RAG)
  • 3.1 Definition of LLM Agents in MAIN-RAG
  • 3.2 Relevance Judgment Quantification
  • 3.3 Adaptive Judge Bar τ q
  • 4 Experiments
  • 4.1 Tasks and Datasets
  • 4.2 Baselines
  • 4.3 Experimental Settings
  • 4.4 Quantitative Analysis (RQ1)
  • 4.5 Ablation Studies on Adaptive Judge Bar τ q for Filtering and Ranking (RQ2)
  • 4.6 Case Studies of Different Adaptive Judge Bar τ q (RQ3)
  • 5 Conclusion and Future Work
  • 6 Limitations
  • Acknowledgments
  • References
  • Appendix
  • A Computation Infrastructure
  • B Performance Comparison among MAIN-RAG and Its Variant Baselines
  • C System Instructions of Agent-1 (Predictor), Agent-2 (Judge), and Agent-3 (Final-Predictor)
  • D Case Studies of Different Adaptive Judge Bar τ q in MAIN-RAG

Knowls

  1. Knowl 1 — Three-agent post-retrieval filtering and generation

    model/method

    MAIN-RAG is a training-free RAG framework that filters and ranks retrieved documents after retrieval, before producing the final answer. For a query and each retrieved document, Agent-1 (Predictor) generates an answer using that document, forming a document–query–answer triplet. Agent-2 (Judge) evaluates each triplet for whether the document supplies information relevant to the query and whether the generated answer answers the query based on that document. Its judgments are converted into relevance scores; documents below a query-specific threshold are removed, and retained documents are ordered by score. Agent-3 (Final-Predictor) answers the query using the filtered, ordered document list. The agents can be instantiated with pretrained LLMs without additional training or fine-tuning.

  2. Knowl 2 — Adaptive query-specific document filtering threshold

    model/method

    For a query with NN retrieved documents, let rir_i be the relevance score assigned to document ii, let rˉq\bar r_q be the mean of the NN scores, and let σq\sigma_q be their standard deviation. MAIN-RAG uses the score distribution to set a query-specific filtering threshold: the unadjusted threshold is the mean score, and a more permissive variant lowers it by a multiple of the standard deviation, τq=rˉq−nσq\tau_q=\bar r_q-n\sigma_q. Here nn is a nonnegative adjustment parameter; n=0n=0 gives the unadjusted mean threshold. The framework retains documents whose scores satisfy ri≥τqr_i\geq\tau_q and removes the rest. Lowering the threshold is intended to retain more potentially useful documents when recall is a concern. The paper identifies nn as MAIN-RAG’s only hyperparameter.

  3. Knowl 3 — Log-odds score for judging document support

    equation

    Agent-2 is prompted to make a binary judgment for each document–query–answer triplet: 'Yes' indicates that the document provides specific information for answering the query and that the generated answer directly answers the query based on that document; 'No' indicates that at least one condition is unmet. MAIN-RAG turns this judgment into a sortable document score by taking the difference between the LLM’s log probabilities for the two answer tokens: ri=log⁡pi(Yes)−log⁡pi(No)r_i=\log p_i(\mathrm{Yes})-\log p_i(\mathrm{No}). Here rir_i is the relevance score for document ii, and pi(Yes)p_i(\mathrm{Yes}) and pi(No)p_i(\mathrm{No}) are the probabilities Agent-2 assigns to those tokens under the triplet and judging prompt. A higher score favors the supportive-document judgment; the scores are used for both filtering and ordering.

  4. Knowl 4 — Score distributions motivate the adaptive threshold

    empirical result

    In an analysis of RGB benchmark document scores using Mistral-7B as the judging LLM, scores for related documents tended to be concentrated toward higher values with relatively small spread, whereas scores for noisy documents were more broadly distributed and had larger spread. The authors interpret this pattern as greater judge confidence for related documents and less reliable judgments for noisy ones. It motivates using the mean score for a query as a distribution-adaptive filtering threshold, but the observed distributions do not guarantee that the threshold will classify every document correctly.

  5. Knowl 5 — Evaluation tasks and default experimental configuration

    experimental setup

    The evaluation used four question-answering benchmarks in zero-shot settings. ARC-Challenge was evaluated on its test set using accuracy. TriviaQA-unfiltered used 11,313 evaluation queries, and PopQA used a long-tail subset of 1,399 rare-entity queries whose entities had fewer than 100 monthly Wikipedia page views; for these two datasets, a prediction was counted as correct when the generation included a gold answer. ALCE-ASQA was evaluated using the reported correctness metrics str-em and ROUGE, and fluency using MAUVE. The default retriever was pretrained Contriever-MS MARCO, which retrieved up to 20 documents per query. All three agents were instantiated with either pretrained Mistral-7B or pretrained Llama-3-8B, without further tuning, and generation was greedy.

  6. Knowl 6 — MAIN-RAG benchmark performance

    data/table

    On the four evaluated tasks, MAIN-RAG improved on standard training-free RAG using the same base LLM for both tested agent configurations. The values below are reported in the paper; TriviaQA, PopQA, and ARC-C are accuracy, while the three ASQA values are str-em, ROUGE, and MAUVE, respectively.

    Standard RAG with Mistral-7B: TriviaQA 69.4, PopQA 55.5, ARC-C 57.1, ASQA 32.4 / 34.8 / 54.3. MAIN-RAG with Mistral-7B: 71.0, 58.9, 58.9, and 35.7 / 36.2 / 60.0.

    Standard RAG with Llama-3-8B: TriviaQA 73.1, PopQA 61.8, ARC-C 55.6, ASQA 37.1 / 36.5 / 63.0. MAIN-RAG with Llama-3-8B: 74.1, 64.0, 61.9, and 39.2 / 42.0 / 70.6.

    Thus, MAIN-RAG exceeded its same-model standard-RAG baseline on all reported metrics. It did not outperform every training-based baseline on every metric: for example, Self-RAG scored 67.3 on ARC-C accuracy and 74.3 on ASQA MAUVE, compared with 61.9 and 70.6 for MAIN-RAG with Llama-3-8B.

  7. Knowl 7 — Threshold and document-order ablations

    data/table

    The ablation compared the default mean-score threshold and descending score order with ascending order and thresholds lowered by multiples of the score standard deviation. Values are accuracy percentages, listed as TriviaQA / PopQA / ARC-C.

    With Mistral-7B: descending default, 71.0 / 58.9 / 58.9; ascending default, 70.2 / 53.5 / 57.4; threshold lowered by 0.5σ0.5\sigma, 71.2 / 58.6 / 59.0; lowered by 1.0σ1.0\sigma, 70.8 / 58.0 / 58.5; lowered by 1.5σ1.5\sigma, 70.4 / 58.4 / 57.7.

    With Llama-3-8B: descending default, 74.1 / 64.0 / 61.9; ascending default, 73.6 / 63.5 / 60.7; threshold lowered by 0.5σ0.5\sigma, 74.1 / 64.0 / 58.6; lowered by 1.0σ1.0\sigma, 74.1 / 63.3 / 58.9; lowered by 1.5σ1.5\sigma, 74.3 / 64.0 / 57.2.

    Descending order outperformed ascending order on all three datasets for both LLM configurations. Lowered thresholds produced occasional best results, but the unadjusted threshold was at least second-best across the three datasets and two configurations, supporting its use as the default.

  8. Knowl 8 — Retrieved-document order affects RAG accuracy

    empirical result

    On the RGB benchmark, the authors evaluated RAG with 10 retrieved documents by randomly shuffling their order 10 times under each of three noise conditions: 4, 6, or 8 noisy documents out of 10. The reported box plots show variation in exact-match performance across orderings, with the best orderings performing substantially better than the worst. This experiment supports the motivation for assigning relevance scores and ordering documents rather than passing an arbitrary retrieval order to the answer generator.

  9. Knowl 9 — Illustrative successes and a filtering failure

    empirical result

    The paper’s PopQA case studies show that MAIN-RAG can retain useful evidence at different threshold values, but can also pass irrelevant evidence and yield an incorrect answer. With Mistral-7B, a threshold of 9.575 retained a document stating that Montxu Miranda was born in Santurce, and the answer was correct; a threshold of -8.425 retained evidence supporting Maniowy as the capital of Gmina Czorsztyn, and the answer was also correct. At a threshold of 0.4875, documents about other people named Arcangelo were retained for a question about Arcangelo Ghisleri, and the generated occupation was wrong. These examples illustrate possible behavior rather than establishing a general relationship between threshold value and accuracy.

  10. Knowl 10 — Evaluation scope and acknowledged environmental limitation

    limitation

    The paper’s experiments cover four datasets and two pretrained LLM architectures, with a focus on inference using retrieved external documents. The authors acknowledge that running LLM inference in a RAG workflow contributes to carbon emissions and represents an environmental risk. They identify finer-grained adaptive thresholds, evaluation beyond question answering, and incorporating human feedback or tuning-based approaches as future directions.

Coverage note — Background and related work are omitted; the generic Agent-1 and Agent-3 prompt wording and hardware inventory are not separate methodological contributions, while the Judge’s operative criteria are included.

References

  1. 1.Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi. 2024. Self-rag: Learning to retrieve, generate, and critique through self-reflection. In The Twelfth International Conference on Learning Representations.
  2. 2.Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. 2021. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258.
  3. 3.Tom B Brown. 2020. Language models are few-shot learners. arXiv preprint arXiv:2005.14165.
  4. 4.Jiawei Chen, Hongyu Lin, Xianpei Han, and Le Sun. 2024. Benchmarking large language models in retrieval-augmented generation. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 17754–17762.
  5. 5.Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457.
  6. 6.Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783.
  7. 7.Yann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy S Liang, and Tatsunori B Hashimoto. 2024. Alpacafarm: A simulation framework for methods that learn from human feedback. Advances in Neural Information Processing Systems, 36.
  8. 8.Tianyu Gao, Howard Yen, Jiatong Yu, and Danqi Chen. 2023. Enabling large language models to generate text with citations. arXiv preprint arXiv:2305.14627.
  9. 9.Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. 2020. Retrieval augmented language model pre-training. In International conference on machine learning, pages 3929–3938. PMLR.
  10. 10.Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bojanowski, Armand Joulin, and Edouard Grave. 2021. Unsupervised dense information retrieval with contrastive learning. arXiv preprint arXiv:2112.09118.
  11. 11.Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12):1–38.
  12. 12.Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023a. Mistral 7b. arXiv preprint arXiv:2310.06825.
  13. 13.Zhengbao Jiang, Frank Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi-Yu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023b. Active retrieval augmented generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 7969–7992, Singapore. Association for Computational Linguistics.
  14. 14.Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer. 2017. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. arXiv preprint arXiv:1705.03551.
  15. 15.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. arXiv preprint arXiv:2004.04906.
  16. 16.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing Systems, 33:9459–9474.
  17. 17.Xiaonan Li, Changtai Zhu, Linyang Li, Zhangyue Yin, Tianxiang Sun, and Xipeng Qiu. 2023. Llatrieval: Llm-verified retrieval for verifiable generation. arXiv preprint arXiv:2311.07838.
  18. 18.Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. TruthfulQA: Measuring how models mimic human falsehoods. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3214–3252, Dublin, Ireland. Association for Computational Linguistics.
  19. 19.Nelson F Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024. Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics, 11:157–173.
  20. 20.Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, and Hannaneh Hajishirzi. 2022. When not to trust language models: Investigating effectiveness of parametric and non-parametric memories. arXiv preprint arXiv:2212.10511.
  21. 21.Krishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun, Sean Welleck, Yejin Choi, and Zaid Harchaoui. 2021. Mauve: Measuring the gap between neural text and human text using divergence frontiers. Advances in Neural Information Processing Systems, 34:4816–4828.
  22. 22.Ori Ram, Yoav Levine, Itay Dalmedigos, Dor Muhlgay, Amnon Shashua, Kevin Leyton-Brown, and Yoav Shoham. 2023. In-context retrieval-augmented language models. Transactions of the Association for Computational Linguistics, 11:1316–1331.
  23. 23.Victor Sanh, Albert Webson, Colin Raffel, Stephen Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Arun Raja, Manan Dey, et al. 2022. Multitask prompted training enables zero-shot task generalization. In International Conference on Learning Representations.
  24. 24.Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2024. Toolformer: Language models can teach themselves to use tools. Advances in Neural Information Processing Systems, 36.
  25. 25.Ivan Stelmakh, Yi Luan, Bhuwan Dhingra, and Ming-Wei Chang. 2022. Asqa: Factoid questions meet long-form answers. arXiv preprint arXiv:2204.06092.
  26. 26.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  27. 27.A Vaswani. 2017. Attention is all you need. Advances in Neural Information Processing Systems.
  28. 28.Yile Wang, Peng Li, Maosong Sun, and Yang Liu. 2023. Self-knowledge guided retrieval augmentation for large language models. arXiv preprint arXiv:2310.05002.
  29. 29.Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. 2021. Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652.
  30. 30.Yue Yu, Wei Ping, Zihan Liu, Boxin Wang, Jiaxuan You, Chao Zhang, Mohammad Shoeybi, and Bryan Catanzaro. 2024. RankRAG: Unifying context ranking with retrieval-augmented generation in LLMs. In The Thirty-eighth Annual Conference on Neural Information Processing Systems.
  31. 31.Rowan Zellers, Ari Holtzman, Hannah Rashkin, Yonatan Bisk, Ali Farhadi, Franziska Roesner, and Yejin Choi. 2019. Defending against neural fake news. Advances in neural information processing systems, 32.
  32. 32.Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, et al. 2023. Siren’s song in the ai ocean: a survey on hallucination in large language models. arXiv preprint arXiv:2309.01219.

Citation

MLA
Chang, C.-Y., et al. “MAIN-RAG: Multi-Agent Filtering Retrieval-Augmented Generation”. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2025, pp. 2607–22, https://doi.org/10.18653/v1/2025.acl-long.131.
APA
Chang, C.-Y., Jiang, Z., Rakesh, V., Pan, M., Yeh, C.-C. M., Wang, G., Hu, M., Xu, Z., Zheng, Y., Das, M., & Zou, N. (2025). MAIN-RAG: Multi-Agent Filtering Retrieval-Augmented Generation. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2607–2622. https://doi.org/10.18653/v1/2025.acl-long.131
Chicago
Chang, C.-Y., Z. Jiang, V. Rakesh, et al. 2025. “MAIN-RAG: Multi-Agent Filtering Retrieval-Augmented Generation”. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2607–22. https://doi.org/10.18653/v1/2025.acl-long.131.
Harvard
Chang, C.-Y. et al. (2025) “MAIN-RAG: Multi-Agent Filtering Retrieval-Augmented Generation”, Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 2607–2622. Available at: https://doi.org/10.18653/v1/2025.acl-long.131.
Vancouver
1. Chang C-Y, Jiang Z, Rakesh V, et al (2025) MAIN-RAG: Multi-Agent Filtering Retrieval-Augmented Generation. In: Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 2607–2622

BibTeX

@inproceedings{chang-etal-2025-main,
    title = "{MAIN}-{RAG}: Multi-Agent Filtering Retrieval-Augmented Generation",
    author = "Chang, Chia-Yuan  and
      Jiang, Zhimeng  and
      Rakesh, Vineeth  and
      Pan, Menghai  and
      Yeh, Chin-Chia Michael  and
      Wang, Guanchu  and
      Hu, Mingzhi  and
      Xu, Zhichao  and
      Zheng, Yan  and
      Das, Mahashweta  and
      Zou, Na",
    editor = "Che, Wanxiang  and
      Nabende, Joyce  and
      Shutova, Ekaterina  and
      Pilehvar, Mohammad Taher",
    booktitle = "Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2025",
    address = "Vienna, Austria",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.acl-long.131/",
    doi = "10.18653/v1/2025.acl-long.131",
    pages = "2607--2622",
    ISBN = "979-8-89176-251-0"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/