CompAct: Compressing Retrieved Documents Actively for Question Answering

Chanwoong YoonTaewhoo LeeHyeon HwangMinbyul JeongJaewoo Kang

article2024EMNLP77 citations

Presents CompAct, an active context compression framework that dynamically integrates multi-hop evidence across retrieved documents and applies early termination, achieving up to 47x compression while boosting reader accuracy on complex question-answering benchmarks.

Listen

Modern question-answering systems increasingly rely on retrieval-augmented generation, providing language models with external documents to ground responses in facts. However, real-world deployment faces a major bottleneck: supplying extensive, lengthy documents introduces significant noise and overwhelms language models, reducing their accuracy when answering multi-hop questions that require synthesizing clues scattered across multiple sources. Simply passing raw documents to large models also dramatically increases computational overhead and operating expenses.

The main objective of the article is to demonstrate and evaluate COMPACT, an active context-compression framework designed to condense large volumes of retrieved documents into concise summaries while preserving essential multi-document reasoning information.

To achieve this, the approach segments retrieved documents and processes them iteratively. Rather than compressing all documents at once, the system sequentially updates a compressed context by jointly analyzing previous notes with incoming text segments, while using an early termination evaluation to stop processing as soon as sufficient evidence is collected. The framework was developed by fine-tuning an open-source 7-billion parameter language model using synthetic demonstrations generated by advanced models, and it was evaluated across five single-document and multi-document benchmark datasets.

The evaluation revealed several key findings. First, the framework achieved extreme context reduction, reaching compression rates between 37x and 51x—condensing roughly 3,000 document tokens into fewer than 200 tokens. Second, on multi-document benchmarks like HotpotQA, it improved accuracy by 7.0 F1 points over existing compression baselines and outperformed long-context language models that process entire raw documents. Third, it demonstrated seamless plug-and-play compatibility across diverse standard retrievers and reader models. Fourth, using this method as a front-end filter for proprietary commercial models reduced API operational costs by up to 90–97% while maintaining or improving overall answer accuracy.

These results demonstrate that larger context windows and raw text feeding are not necessary for complex information retrieval. In fact, providing raw or poorly filtered documents degrades reader performance due to distracting noise. By acting as an efficient, modular bridge, the active compression method significantly lowers financial costs and infrastructure load, making advanced question-answering practical even for models with small input capacities.

Decision-makers and engineering teams should consider adopting active context compression modules as standard intermediary layers between information retrieval systems and downstream reader models to optimize cost-performance trade-offs. Further work should focus on piloting the framework across varied domain-specific environments, testing its performance on model backbones of different sizes, and optimizing the iterative compression speed to reduce runtime latency during document ingestion.

arXiv: 2407.09014dmis-lab/CompAct
  • Paper: REFRAG: Rethinking RAG based Decoding, Xiaoqiang Lin et al. (2025). REFRAG carries retrieved-context compression further by encoding chunks into compact representations to accelerate RAG decoding while retaining answer quality.
  • Paper: Prompt Compression for Large Language Models: A Survey, Zongqian Li et al. (2025). This later survey organizes prompt-compression methods and trade-offs, placing CompAct’s active QA-focused approach within the broader field.
Cover for CompAct: Compressing Retrieved Documents Actively for Question Answering

Abstract

Retrieval-augmented generation supports language models to strengthen their factual groundings by providing external contexts. However, language models often face challenges when given extensive information, diminishing their effectiveness in solving questions. Context compression tackles this issue by filtering out irrelevant information, but current methods still struggle in realistic scenarios where crucial information cannot be captured with a single-step approach. To overcome this limitation, we introduce CompAct, a novel framework that employs an active strategy to condense extensive documents without losing key information. Our experiments demonstrate that CompAct brings significant improvements in both performance and compression rate on multi-hop question-answering benchmarks. CompAct flexibly operates as a cost-efficient plug-in module with various off-the-shelf retrievers or readers, achieving exceptionally high compression rates (47x).

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 2.1 Multi-Document Question Answering
  • 2.2 Multi-hop Information-Seeking
  • 2.3 Context Compression
  • 2.4 Task Formulation
  • 3 COMPACT
  • 3.1 Active Compression
  • 3.2 Early Termination
  • 3.3 Dataset Construction
  • 4 Experiment
  • 4.1 Experimental Setup
  • 4.2 Datasets
  • 4.3 Baselines
  • 4.4 Results
  • 5 Analysis
  • 5.1 Compressor as a Plug-in Module
  • 5.2 Component Effectiveness
  • 5.3 Cost Efficiency
  • 5.4 Computational Efficiency
  • 6 Conclusion
  • Limitations
  • Ethics Statement
  • Acknowledgments
  • References
  • A Practicality of Compressing Contexts
  • B Additional Comparison
  • C Length of Compressed Text
  • D Implementation Details
  • D.1 Training & Inference
  • D.2 Baselines

Knowls

  1. Knowl 1 — COMPACT sequentially updates a query-focused context

    algorithm

    COMPACT is a trained compressor placed between a retriever and a question-answering reader. It processes retrieved documents in batches, repeatedly rewriting a compact context so that information in each new batch can be integrated with information retained from earlier batches. The output context is intended to preserve evidence relevant to question qq, rather than to answer the question from the compressor’s parametric knowledge.

    Let Dk=(d1,…,dk)D_k=(d_1,\ldots,d_k) be the kk retrieved documents in rank order, and let jj be the number of documents per segment. Segment StS_t contains documents with ranks (t−1)j+1(t-1)j+1 through tjtj. At iteration tt, the compressor π\pi receives question qq, segment StS_t, and the previous compressed context Ct−1C_{t-1}, and produces an updated context CtC_t and an evaluation EtE_t:

    (Ct,Et)=π(q,St,Ct−1).(C_t,E_t)=\pi(q,S_t,C_{t-1}).

    Here, C0C_0 is empty, and EtE_t includes an assessment of whether the accumulated context is sufficient to answer the question. In the default setup, j=5j=5, the retriever supplies the top 30 documents, and thus there are at most six iterations per question.

    Input: question qq, ranked retrieved documents DkD_k, segment size j=5j=5, trained compressor π\pi
    Output: compressed context CC
    Set C0C_0 to an empty context
    For each consecutive segment StS_t of jj documents in DkD_k, in retrieval order:
        Generate (Ct,Et)=π(q,St,Ct−1)(C_t, E_t)=\pi(q, S_t, C_{t-1})
        If EtE_t indicates [COMPLETE]:
            Return CtC_t
        Set the previous context to CtC_t
    Return the last generated context
  2. Knowl 2 — Evidence-based early termination limits unnecessary compression steps

    model/method

    At every COMPACT iteration, the compressor evaluates whether the newly updated context contains enough information to answer the question. The evaluation consists of a rationale and one of two condition tokens: [COMPLETE] or [INCOMPLETE]. The judgment is based on the question and the available segment together with the previous compressed context. [COMPLETE] stops processing further segments; [INCOMPLETE] causes COMPACT to continue if more retrieved documents remain. If the documents are exhausted, the latest compressed context is returned.

    This stopping rule is intended to avoid processing redundant segments and to adapt the number of iterations to the evidence available for each question. It is especially relevant when questions differ in complexity or when supporting information appears at different ranks in the retrieved results.

  3. Knowl 3 — Synthetic supervision trains compression and completeness judgments

    experimental setup

    The authors constructed 28,800 training instances from a subset of HotpotQA training questions using GPT-4o (API version dated 2024-05-13). Instances were balanced across four categories, with 7,200 examples each: [COMPLETE] at the first iteration, [COMPLETE] at a subsequent iteration, [INCOMPLETE] at the first iteration, and [INCOMPLETE] at a subsequent iteration.

    For each example, the data-generation process selected sentences containing direct answer clues or useful context, summarized the selected material without attempting to answer the question or making unsupported assumptions, and judged whether the summary alone contained sufficient information. Subsequent-iteration examples included both the previous summary and a new document segment. Data came from realistic retrieval results and from distractor scenarios with predefined contexts containing the supporting facts; the latter helped supply examples where completeness was reached only after an initial iteration. The collection setup used Contriever over the 2018 Wikipedia corpus, 30 retrieved documents per question, and five documents per segment, allowing at most six iterations.

    COMPACT was supervised-fine-tuned from instruction-tuned Mistral-7B-Instruct-v0.2 on this dataset. Training used four 80-GB Nvidia A100 GPUs, Adam with learning rate 2×10−62\times10^{-6}, batch size 64, warmup ratio 0.1, and seven epochs. Inference used greedy decoding with temperature 0 and top-p=1.0p=1.0.

  4. Knowl 4 — COMPACT improves compression and answer quality on five QA benchmarks

    empirical result

    The main comparison used LLaMA3-8B as the reader and the top 30 retrieved documents. Compression rate is the number of tokens in retrieved documents divided by the number of tokens in the compressed context; larger values indicate more compression. EM and F1 measure answer quality. COMPACT achieved the highest compression rate among the listed compressors on all five datasets and the highest compressor F1 on each dataset. Its TriviaQA EM, however, was below RECOMP and the raw-document result.

    Method HotpotQA MuSiQue 2WikiMQA NQ TriviaQA
    Comp. EM F1 Comp. EM F1 Comp. EM F1 Comp. EM F1 Comp. EM F1
    Oracle 10.8x 39.9 51.2 10.3x 14.2 23.6 11.0x 37.4 43.2 - - - - - -
    Raw document 1x 29.4 40.3 1x 6.5 15.6 1x 25.4 31.2 1x 39.0 51.3 1x 68.9 77.1
    AutoCompressors 35.4x 18.4 28.4 34.7x 3.9 11.9 36.2x 19.0 24.5 34.4x 17.3 31.8 34.5x 55.3 64.3
    LongLLMLingua 3.4x 25.6 35.3 3.4x 4.8 13.5 3.6x 27.9 32.9 3.5x 27.7 40.6 3.3x 64.0 70.8
    RECOMP (extractive) 34.3x 29.7 39.9 32.7x 6.7 15.7 35.9x 29.9 34.9 32.7x 34.6 45.1 39.2x 67.6 74.1
    COMPACT 47.6x 35.5 46.9 37.2x 8.7 18.1 51.2x 31.0 37.1 48.5x 38.4 50.0 49.4x 65.4 74.9

    On the three multi-document benchmarks, COMPACT’s F1 scores were 46.9 on HotpotQA, 18.1 on MuSiQue, and 37.1 on 2WikiMQA, exceeding the listed compressor baselines while using substantially fewer tokens than the raw-document setup. The authors trained on a HotpotQA subset; evaluations on MuSiQue, 2WikiMQA, NQ, and TriviaQA did not use those datasets’ training sets.

  5. Knowl 5 — The compressed contexts work across retrievers and readers

    empirical result

    COMPACT was evaluated as a plug-in compressor with different retrieval and reader models, using 500 random HotpotQA development examples for the retriever comparisons. The retrievers included Contriever and BM25, and the comparisons included raw documents, gold documents, and RECOMP. Tests varied the retrieved-document count up to 40. With Contriever, which often missed relevant documents at high ranks, COMPACT made use of information from lower-ranked documents; with BM25, its performance remained comparatively consistent through the top 40 and showed a saturation trend similar to the gold-document condition. In both retriever settings, the reported curves placed COMPACT above raw documents and RECOMP, while remaining below gold documents in some settings.

    The authors also assessed the same compressed texts with GPT-3.5-Turbo, LLaMA2-13B, and LLaMA3-8B readers. COMPACT maintained performance as more documents were included, whereas raw-document performance degraded at higher document counts for GPT-3.5-Turbo. The reported results support using COMPACT with replaceable off-the-shelf retrievers and readers; they do not depend on a single fixed retriever-reader pair.

  6. Knowl 6 — Ablation separates the effects of compressed text and rationale

    empirical result

    An ablation on 500 random examples per dataset compared using only the rationale, only the compressed text (CT), and CT together with the rationale. The reported columns are compression rate and reader F1; the readers were LLaMA3-8B, LLaMA2-13B, and GPT-3.5-Turbo. Rationale-only outputs were much shorter but generally had lower F1. Adding the rationale to CT usually reduced F1 relative to CT alone, with no consistent benefit. The authors suggest that rationale content may distract the reader from answering from the compressed evidence.

    Reader Component HotpotQA MuSiQue 2WikiMQA
    Comp. F1 Comp. F1 Comp. F1
    LLaMA3-8B Rationale 130.8x 41.6 120.0x 15.9 141.3x 32.3
    LLaMA3-8B CT 47.5x 48.3 36.5x 19.1 52.2x 36.2
    LLaMA3-8B CT + rationale 33.6x 47.3 27.1x 19.0 36.4x 35.6
    LLaMA2-13B Rationale 141.8x 41.8 129.2x 16.9 152.4x 30.8
    LLaMA2-13B CT 48.1x 48.5 37.0x 18.6 52.7x 35.6
    LLaMA2-13B CT + rationale 34.6x 47.3 28.0x 18.6 37.4x 34.2
    GPT-3.5-Turbo Rationale 135.2x 38.0 123.5x 13.8 146.2x 24.0
    GPT-3.5-Turbo CT 48.1x 49.2 37.0x 20.9 53.0x 34.0
    GPT-3.5-Turbo CT + rationale 33.9x 47.0 27.4x 18.5 36.7x 36.5

    CT means the compressed text without the evaluation rationale. CT alone had the highest F1 in eight of the nine reader-dataset comparisons; the exception was GPT-3.5-Turbo on 2WikiMQA, where CT plus rationale scored 36.5 versus 34.0 for CT.

  7. Knowl 7 — COMPACT lowers API inference costs with proprietary readers

    empirical result

    The authors measured reader API costs in US dollars on 500 HotpotQA development examples using four proprietary readers. For every reader in this comparison, COMPACT had the lowest reported cost among raw documents, RECOMP, LongLLMLingua (labeled LINGUA*), and COMPACT. Its F1 was also higher than the raw-document and compressor alternatives in all four rows except that the table’s F1 values are specific to these reader-model runs.

    Raw RECOMP LINGUA* COMPACT
    Reader Cost F1 Cost F1 Cost F1 Cost F1
    GPT-3.5-Turbo 1.09 44.5 0.04 40.1 0.33 38.4 0.04 49.2
    GPT-4o 10.75 55.8 0.43 48.1 3.31 47.6 0.28 56.0
    Claude-3.5 6.45 36.0 0.26 37.0 1.99 30.2 0.17 42.2
    Gemini-1.5-pro 7.54 52.0 0.31 41.7 2.36 40.1 0.20 44.8

    Costs are in USD. LINGUA* denotes LongLLMLingua. For GPT-3.5-Turbo, RECOMP and COMPACT both cost 0.04 USD at the reported precision, while COMPACT’s F1 was 49.2 compared with 40.1 for RECOMP.

  8. Knowl 8 — Compute overhead depends on QA task complexity

    empirical result

    Average inference computation was compared between raw retrieved contexts and COMPACT using LLaMA3-8B as reader with the top 30 documents. TFLOPs are normalized by the number of instances. COMPACT used more TFLOPs on the three multi-document QA datasets, while producing higher F1; on NQ and TriviaQA it used fewer TFLOPs and had similar F1 to the raw-context condition. The results indicate that iterative processing can add computation when evidence takes multiple steps to gather, while early termination can reduce computation on the single-document datasets tested.

    Dataset Raw TFLOPs COMPACT TFLOPs Raw F1 COMPACT F1
    HotpotQA 34.1 35.8 40.0 48.3
    MuSiQue 33.6 49.3 16.2 19.0
    2WikiMQA 35.9 42.4 29.5 37.2
    NQ 32.9 26.7 52.9 53.8
    TriviaQA 33.5 24.6 78.5 77.3

    The largest reported overhead was on MuSiQue, where COMPACT used 49.3 TFLOPs versus 33.6 for raw contexts and improved F1 from 16.2 to 19.0. On TriviaQA, COMPACT reduced computation from 33.5 to 24.6 TFLOPs, with F1 changing from 78.5 to 77.3.

  9. Knowl 9 — A separate LLaMA2-7B comparison shows gains over ICAE and QGC

    empirical result

    A supplementary comparison followed the setup used by the QGC study and used LLaMA2-7B as reader. It reported compression rate and dataset-specific accuracy or F1. COMPACT had the highest reported HotpotQA F1 and TriviaQA EM among the listed methods, while its NQ accuracy was below QGC. The authors note that COMPACT used one set of weights rather than dataset-specific fine-tuning and did not rely on an external reranker in this comparison.

    Method NQ Comp. NQ Acc. TriviaQA Comp. TriviaQA EM HotpotQA Comp. HotpotQA F1
    Oracle 59.2x 73.5 - - 42.2x 57.7
    ICAE 21.5x 53.3 10.2x 48.9 9.5x 34.5
    QGC 15.2x 60.9 7.9x 57.5 8.8x 51.6
    QGC (ϵ=0.42\epsilon=0.42) 20.6x 57.6 10.9x 57.1 12.1x 51.2
    COMPACT 14.6x 57.0 10.9x 64.4 12.2x 59.0

    This comparison is distinct from the main LLaMA3-8B evaluation: its reader, comparison methods, and compression rates differ. In particular, COMPACT’s NQ accuracy of 57.0 was lower than QGC’s 60.9, despite higher TriviaQA EM and HotpotQA F1 in the reported rows.

  10. Knowl 10 — The evaluation identifies inference and model-scale limitations

    limitation

    COMPACT takes longer to infer than other compressors because it processes retrieved documents iteratively. Its early-termination decisions can also be wrong: the authors report that even GPT-4o sometimes misjudged whether a context was complete, leaving possible errors in the automatically constructed training data despite filtering. Finally, resource constraints limited the trained compressor to Mistral-7B-Instruct-v0.2; the paper does not establish whether the approach performs similarly with models smaller or larger than 7B parameters.

Coverage note — No other substantial contributed material was omitted; background comparisons and non-contribution appendices were excluded.

References

  1. 1.Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783.
  2. 2.Shengnan An, Zexiong Ma, Zeqi Lin, Nanning Zheng, and Jian-Guang Lou. 2024. Make your llm fully utilize the context. arXiv preprint arXiv:2404.16811.
  3. 3.Anthropic. 2024. claude-3.5-sonnet.
  4. 4.Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, et al. 2016. Ms marco: A human generated machine reading comprehension dataset. arXiv preprint arXiv:1611.09268.
  5. 5.Zheng Cai, Maosong Cao, Haojiong Chen, Kai Chen, Keyu Chen, Xin Chen, Xun Chen, Zehui Chen, Zhi Chen, Pei Chu, et al. 2024. InternLM2 technical report. arXiv preprint arXiv:2403.17297.
  6. 6.Zhiwei Cao, Qian Cao, Yu Lu, Ningxin Peng, Luyang Huang, Shanbo Cheng, and Jinsong Su. 2024. Retaining key information under high compression ratios: Query-guided compressor for llms. arXiv preprint arXiv:2406.02376.
  7. 7.Wenhu Chen, Hanwen Zha, Zhiyu Chen, Wenhan Xiong, Hong Wang, and William Yang Wang. 2020. Hybridqa: A dataset of multi-hop question answering over tabular and textual data. In Findings of the Association for Computational Linguistics: EMNLP 2020.
  8. 8.Xin Cheng, Xun Wang, Xingxing Zhang, Tao Ge, Si-Qing Chen, Furu Wei, Huishuai Zhang, and Dongyan Zhao. 2024. xrag: Extreme context compression for retrieval-augmented generation with one token. arXiv preprint arXiv:2405.13792.
  9. 9.Alexis Chevalier, Alexander Wettig, Anirudh Ajith, and Danqi Chen. 2023. Adapting language models to compress contexts. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing.
  10. 10.Tao Ge, Hu Jing, Lei Wang, Xun Wang, Si-Qing Chen, and Furu Wei. 2024. In-context autoencoder for context compression in a large language model. In The Twelfth International Conference on Learning Representations.
  11. 11.Google. 2024. gemini-1.5-pro.
  12. 12.Bernal Jiménez Gutiérrez, Yiheng Shu, Yu Gu, Michihiro Yasunaga, and Yu Su. 2024. Hipporag: Neurobiologically inspired long-term memory for large language models. arXiv preprint arXiv:2405.14831.
  13. 13.Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara, and Akiko Aizawa. 2020a. Constructing a multi-hop QA dataset for comprehensive evaluation of reasoning steps. In Proceedings of the 28th International Conference on Computational Linguistics, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  14. 14.Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara, and Akiko Aizawa. 2020b. Constructing a multi-hop qa dataset for comprehensive evaluation of reasoning steps. In Proceedings of the 28th International Conference on Computational Linguistics.
  15. 15.Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bojanowski, Armand Joulin, and Edouard Grave. 2022. Unsupervised dense information retrieval with contrastive learning. Transactions on Machine Learning Research.
  16. 16.Gautier Izacard, Patrick Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel, and Edouard Grave. 2023. Atlas: Few-shot learning with retrieval augmented language models. Journal of Machine Learning Research.
  17. 17.Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023a. Mistral 7b. arXiv preprint arXiv:2310.06825.
  18. 18.Huiqiang Jiang, Qianhui Wu, Chin-Yew Lin, Yuqing Yang, and Lili Qiu. 2023b. LLMLingua: Compressing prompts for accelerated inference of large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Singapore. Association for Computational Linguistics.
  19. 19.Huiqiang Jiang, Qianhui Wu, Xufang Luo, Dongsheng Li, Chin-Yew Lin, Yuqing Yang, and Lili Qiu. 2023c. Longllmlingua: Accelerating and enhancing llms in long context scenarios via prompt compression. arXiv preprint arXiv:2310.06839.
  20. 20.Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017. TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics.
  21. 21.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020a. Dense passage retrieval for open-domain question answering. arXiv preprint arXiv:2004.04906.
  22. 22.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020b. Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). Association for Computational Linguistics.
  23. 23.Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. 2019. Generalization through memorization: Nearest neighbor language models. In International Conference on Learning Representations.
  24. 24.Diederik Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In International Conference on Learning Representations (ICLR), San Diega, CA, USA.
  25. 25.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. 2019. Natural questions: a benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:453–466.
  26. 26.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing Systems.
  27. 27.Tianle Li, Ge Zhang, Quy Duc Do, Xiang Yue, and Wenhu Chen. 2024. Long-context llms struggle with long in-context learning. arXiv preprint arXiv:2404.02060.
  28. 28.Yucheng Li, Bo Dong, Frank Guerin, and Chenghua Lin. 2023. Compressing context to enhance inference efficiency of large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Singapore. Association for Computational Linguistics.
  29. 29.Nelson F Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024. Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics.
  30. 30.Vaibhav Mavi, Anubhav Jangra, and Adam Jatowt. 2022. A survey on multi-hop question answering and generation. arXiv preprint arXiv:2204.09140.
  31. 31.Jesse Mu, Xiang Lisa Li, and Noah Goodman. 2024. Learning to compress prompts with gist tokens. In Proceedings of the 37th International Conference on Neural Information Processing Systems, Red Hook, NY, USA. Curran Associates Inc.
  32. 32.OpenAI. 2023. Chatgpt.
  33. 33.OpenAI. 2024. Gpt-4o.
  34. 34.Zhuoshi Pan, Qianhui Wu, Huiqiang Jiang, Menglin Xia, Xufang Luo, Jue Zhang, Qingwei Lin, Victor Rühle, Yuqing Yang, Chin-Yew Lin, et al. 2024. Llmlingua-2: Data distillation for efficient and faithful task-agnostic prompt compression. arXiv preprint arXiv:2403.12968.
  35. 35.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. 2019. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems.
  36. 36.Hongjin Qian, Zheng Liu, Peitian Zhang, Kelong Mao, Yujia Zhou, Xu Chen, and Zhicheng Dou. 2024. Are long-llms a necessity for long-context tasks? arXiv preprint arXiv:2405.15318.
  37. 37.Stephen Robertson, Hugo Zaragoza, et al. 2009. The probabilistic relevance framework: Bm25 and beyond. Foundations and Trends® in Information Retrieval.
  38. 38.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  39. 39.Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2022. Musique: Multi-hop questions via single-hop question composition. Transactions of the Association for Computational Linguistics.
  40. 40.Lewis Tunstall, Edward Beeching, Nathan Lambert, Nazneen Rajani, Shengyi Huang, Kashif Rasul, Alexander M. Rush, and Thomas Wolf. 2023. The alignment handbook. https://github.com/huggingface/alignment-handbook.
  41. 41.Zhiruo Wang, Jun Araki, Zhengbao Jiang, Md Rizwan Parvez, and Graham Neubig. 2023. Learning to filter context for retrieval-augmented generation. arXiv preprint arXiv:2311.08377.
  42. 42.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al. 2019. Huggingface’s transformers: State-of-the-art natural language processing. arXiv preprint arXiv:1910.03771.
  43. 43.Fangyuan Xu, Weijia Shi, and Eunsol Choi. 2024. RECOMP: Improving retrieval-augmented LMs with context compression and selective augmentation. In The Twelfth International Conference on Learning Representations.
  44. 44.Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D Manning. 2018. Hotpotqa: A dataset for diverse, explainable multi-hop question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing.
  45. 45.Yusen Zhang, Ruoxi Sun, Yanfei Chen, Tomas Pfister, Rui Zhang, and Sercan Ö. Arik. 2024. Chain of agents: Large language models collaborating on long-context tasks. Preprint, arXiv:2406.02818.
  46. 46.Jared Rasley, Samyam Rajbhandari, Oshrat Ruwase, and Yuxiong He. 2020. DeepSpeed: System optimizations enable training deep learning models with over 100 billion parameters. Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 3505–3506. https://doi.org/10.1145/3394486.3406703.
  47. 47.Mohamed Abdin, Sebastian Jacobs, Adeel Awan, Jatin Aneja, Ahmed Awadallah, Hany Hassan Awadalla, Nguyen Bach, Mohit Bahree, Ahmad Bakhtiari, Harsh Behl, et al. 2024. Phi-3 technical report: A highly capable language model locally on your phone. arXiv preprint arXiv:2404.14219.
  48. 48.01.AI, Andrew Young, Bin Chen, Chao Li, Chen Huang, Guodong Zhang, Guang Zhang, Haotian Li, Jiaming Zhu, Jin Chen, Jiawei Chang, Kaifeng Yu, Pengfei Liu, Qi Liu, Shang Yue, Shuai Yang, Shuo Yang, Tao Yu, Wei Xie, Wei Huang, Xiaoyi Hu, Xudong Ren, Xinting Niu, Ping Nie, Yihan Xu, Yufei Liu, Yida Wang, Yuxin Cai, Zheng Gu, Zhenghao Liu, and Zhilin Dai. 2024. Yi: Open foundation models by 01.AI. arXiv preprint arXiv:2404.14219.
  49. 49.Yu Wang, Nedim Lipka, Ryan A. Rossi, Alexa Siu, Ruiyi Zhang, and Tyler Derr. 2024. Knowledge graph prompting for multi-document question answering. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, number 17, pages 19206–19214.
  50. 50.Howard Chen, Ramakanth Pasunuru, Jason Weston, and Asli Celikyilmaz. 2023. Walking down the memory maze: Beyond context limit through interactive reading. arXiv preprint arXiv:2310.05029.
  51. 51.Kuang-Huei Lee, Xinyun Chen, Hiroki Furuta, John Canny, and Ian Fischer. 2024. A human-inspired reading agent with gist memory of very long contexts. In Proceedings of the Forty-first International Conference on Machine Learning.

Citation

MLA
Yoon, C., et al. “CompAct: Compressing Retrieved Documents Actively for Question Answering”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 21424–39, https://doi.org/10.18653/v1/2024.emnlp-main.1194.
APA
Yoon, C., Lee, T., Hwang, H., Jeong, M., & Kang, J. (2024). CompAct: Compressing Retrieved Documents Actively for Question Answering. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 21424–21439. https://doi.org/10.18653/v1/2024.emnlp-main.1194
Chicago
Yoon, C., T. Lee, H. Hwang, M. Jeong, and J. Kang. 2024. “CompAct: Compressing Retrieved Documents Actively for Question Answering”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 21424–39. https://doi.org/10.18653/v1/2024.emnlp-main.1194.
Harvard
Yoon, C. et al. (2024) “CompAct: Compressing Retrieved Documents Actively for Question Answering”, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 21424–21439. Available at: https://doi.org/10.18653/v1/2024.emnlp-main.1194.
Vancouver
1. Yoon C, Lee T, Hwang H, Jeong M, Kang J (2024) CompAct: Compressing Retrieved Documents Actively for Question Answering. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 21424–21439

BibTeX

@inproceedings{yoon-etal-2024-compact,
    title = "{C}omp{A}ct: Compressing Retrieved Documents Actively for Question Answering",
    author = "Yoon, Chanwoong  and
      Lee, Taewhoo  and
      Hwang, Hyeon  and
      Jeong, Minbyul  and
      Kang, Jaewoo",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.1194/",
    doi = "10.18653/v1/2024.emnlp-main.1194",
    pages = "21424--21439"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/