DYLE: Dynamic Latent Extraction for Abstractive Long-Input Summarization
Ziming MaoChen Henry WuAnsong NiYusen ZhangRui ZhangTao YuBudhaditya DebChenguang ZhuAhmed Hassan AwadallahDragomir R. Radev
Proposes a dynamic latent extraction framework that jointly trains an extractor and generator using snippet-level dynamic weights and consistency loss, substantially outperforming prior methods on long-document and long-dialogue summarization benchmarks.
Modern language models excel at summarizing short passages, but they face steep memory and computational hurdles when processing lengthy documents and multi-turn meeting transcripts. Existing approaches typically compromise by restricting attention spans, dividing documents into disconnected segments, or using multi-step pipelines where errors in selecting relevant text severely degrade final summary quality. The article addresses this bottleneck by introducing and evaluating Dynamic Latent Extraction for Abstractive Summarization (DYLE), a framework designed to summarize long texts efficiently, accurately, and with clear interpretability.
The framework jointly trains a text extractor and an abstractive generator. The extractor divides long source text into manageable chunks to identify key snippets, which remain latent during generation. As the generator drafts each word of the summary, it dynamically assigns attention weights across the extracted snippets based on the words generated so far. Training is stabilized through reference-based target snippets and a novel consistency loss, which encourages the extractor to mirror the generator's attention patterns. The authors evaluated the system across three standard benchmarks: GovReport (approximately 19,500 government reports averaging 9,400 words), QMSum (long multi-party meeting transcripts averaging 9,070 words), and arXiv (lengthy scientific papers).
The evaluation revealed substantial performance improvements. On GovReport, the framework achieved 61.01 ROUGE-1 and 28.83 ROUGE-2, surpassing the prior state-of-the-art by 4.15 and 6.21 points respectively. On QMSum, it set a new benchmark standard with 34.42 ROUGE-1, outperforming specialized dialogue models. Ablation analyses proved that the consistency loss and hybrid training supervision were critical; removing either caused noticeable performance drops across tasks. Furthermore, the dynamic attention weights successfully highlighted which source snippets drove specific summary statements, providing actionable transparency into the generation process.
These findings demonstrate that joint extraction and generation can overcome the computational and memory limits of standard transformers without sacrificing context or incurring cascading pipeline errors. For organizations managing extensive documentation, compliance filings, or meeting records, this architecture offers high-quality automated summarization and built-in traceability. The authors recommend adopting joint dynamic-weight architectures for long-context tasks and exploring their expansion into related domains such as enterprise question answering and multi-turn dialogue analysis.
Decision-makers should note that the model was slightly outperformed on the arXiv dataset, where source articles require broader cross-sentence synthesis. The approach also remains subject to minor information loss inherent in snippet selection and hardware constraints on the maximum number of extracted snippets. Nevertheless, given the consistent gains across diverse benchmark domains, there is high confidence in the framework's effectiveness for structured documents and dialogue records.
- Paper: Longformer: The Long-Document Transformer, Iz Beltagy et al. (2020). Longformer establishes efficient attention for long documents and its LED variant applies that design to summarization, grounding DYLE’s attempt to handle long inputs.
- Paper: PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization, Jingqing Zhang et al. (2020). PEGASUS provides a key pretrained abstractive summarization baseline whose generation approach helps situate DYLE’s joint extractor-generator design.
- Paper: Text Summarization with Pretrained Encoders, Yang Liu et al. (2019). BERTSUM combines document-level encoding with extractive and abstractive summarization, clarifying the earlier architecture that DYLE extends with dynamic snippet selection.
- Paper: Get To The Point: Summarization with Pointer-Generator Networks, Abigail See et al. (2017). The pointer-generator work establishes copying and coverage mechanisms for abstractive summaries, useful context for DYLE’s dynamically attended source snippets.
- Paper: A Deep Reinforced Model for Abstractive Summarization, Romain Paulus et al. (2017). This earlier abstractive model develops attention and reinforcement-learning techniques for coherent summaries, providing background for DYLE’s generation component.
No sufficiently relevant recommendations were found.
