CompAct: Compressing Retrieved Documents Actively for Question Answering
Chanwoong YoonTaewhoo LeeHyeon HwangMinbyul JeongJaewoo Kang
Presents CompAct, an active context compression framework that dynamically integrates multi-hop evidence across retrieved documents and applies early termination, achieving up to 47x compression while boosting reader accuracy on complex question-answering benchmarks.
Modern question-answering systems increasingly rely on retrieval-augmented generation, providing language models with external documents to ground responses in facts. However, real-world deployment faces a major bottleneck: supplying extensive, lengthy documents introduces significant noise and overwhelms language models, reducing their accuracy when answering multi-hop questions that require synthesizing clues scattered across multiple sources. Simply passing raw documents to large models also dramatically increases computational overhead and operating expenses.
The main objective of the article is to demonstrate and evaluate COMPACT, an active context-compression framework designed to condense large volumes of retrieved documents into concise summaries while preserving essential multi-document reasoning information.
To achieve this, the approach segments retrieved documents and processes them iteratively. Rather than compressing all documents at once, the system sequentially updates a compressed context by jointly analyzing previous notes with incoming text segments, while using an early termination evaluation to stop processing as soon as sufficient evidence is collected. The framework was developed by fine-tuning an open-source 7-billion parameter language model using synthetic demonstrations generated by advanced models, and it was evaluated across five single-document and multi-document benchmark datasets.
The evaluation revealed several key findings. First, the framework achieved extreme context reduction, reaching compression rates between 37x and 51x—condensing roughly 3,000 document tokens into fewer than 200 tokens. Second, on multi-document benchmarks like HotpotQA, it improved accuracy by 7.0 F1 points over existing compression baselines and outperformed long-context language models that process entire raw documents. Third, it demonstrated seamless plug-and-play compatibility across diverse standard retrievers and reader models. Fourth, using this method as a front-end filter for proprietary commercial models reduced API operational costs by up to 90–97% while maintaining or improving overall answer accuracy.
These results demonstrate that larger context windows and raw text feeding are not necessary for complex information retrieval. In fact, providing raw or poorly filtered documents degrades reader performance due to distracting noise. By acting as an efficient, modular bridge, the active compression method significantly lowers financial costs and infrastructure load, making advanced question-answering practical even for models with small input capacities.
Decision-makers and engineering teams should consider adopting active context compression modules as standard intermediary layers between information retrieval systems and downstream reader models to optimize cost-performance trade-offs. Further work should focus on piloting the framework across varied domain-specific environments, testing its performance on model backbones of different sizes, and optimizing the iterative compression speed to reduce runtime latency during document ingestion.
- Paper: LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models, Huiqiang Jiang et al. (2023). LLMLingua establishes the earlier prompt-compression approach that makes CompAct’s task-specific, iterative compression strategy easier to understand.
- Paper: LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression, Huiqiang Jiang et al. (2024). LongLLMLingua provides a closely related predecessor for compressing retrieved prompts to preserve question-relevant evidence and improve long-context QA.
- Paper: Compressing Context to Enhance Inference Efficiency of Large Language Models, Yucheng Li et al. (2023). Selective Context introduces model-agnostic pruning of redundant input text, a useful prerequisite for understanding CompAct’s learned document-compression alternative.
- Paper: Adapting Language Models to Compress Contexts, Alexis Chevalier et al. (2023). AutoCompressors develops recursive compression across text segments, providing context for CompAct’s iterative accumulation of compressed notes.
- Paper: REFRAG: Rethinking RAG based Decoding, Xiaoqiang Lin et al. (2025). REFRAG carries retrieved-context compression further by encoding chunks into compact representations to accelerate RAG decoding while retaining answer quality.
- Paper: Prompt Compression for Large Language Models: A Survey, Zongqian Li et al. (2025). This later survey organizes prompt-compression methods and trade-offs, placing CompAct’s active QA-focused approach within the broader field.
