DYLE: Dynamic Latent Extraction for Abstractive Long-Input Summarization

Ziming MaoChen Henry WuAnsong NiYusen ZhangRui ZhangTao YuBudhaditya DebChenguang ZhuAhmed Hassan AwadallahDragomir R. Radev

article2022ACL68 citations

Proposes a dynamic latent extraction framework that jointly trains an extractor and generator using snippet-level dynamic weights and consistency loss, substantially outperforming prior methods on long-document and long-dialogue summarization benchmarks.

Listen

Modern language models excel at summarizing short passages, but they face steep memory and computational hurdles when processing lengthy documents and multi-turn meeting transcripts. Existing approaches typically compromise by restricting attention spans, dividing documents into disconnected segments, or using multi-step pipelines where errors in selecting relevant text severely degrade final summary quality. The article addresses this bottleneck by introducing and evaluating Dynamic Latent Extraction for Abstractive Summarization (DYLE), a framework designed to summarize long texts efficiently, accurately, and with clear interpretability.

The framework jointly trains a text extractor and an abstractive generator. The extractor divides long source text into manageable chunks to identify key snippets, which remain latent during generation. As the generator drafts each word of the summary, it dynamically assigns attention weights across the extracted snippets based on the words generated so far. Training is stabilized through reference-based target snippets and a novel consistency loss, which encourages the extractor to mirror the generator's attention patterns. The authors evaluated the system across three standard benchmarks: GovReport (approximately 19,500 government reports averaging 9,400 words), QMSum (long multi-party meeting transcripts averaging 9,070 words), and arXiv (lengthy scientific papers).

The evaluation revealed substantial performance improvements. On GovReport, the framework achieved 61.01 ROUGE-1 and 28.83 ROUGE-2, surpassing the prior state-of-the-art by 4.15 and 6.21 points respectively. On QMSum, it set a new benchmark standard with 34.42 ROUGE-1, outperforming specialized dialogue models. Ablation analyses proved that the consistency loss and hybrid training supervision were critical; removing either caused noticeable performance drops across tasks. Furthermore, the dynamic attention weights successfully highlighted which source snippets drove specific summary statements, providing actionable transparency into the generation process.

These findings demonstrate that joint extraction and generation can overcome the computational and memory limits of standard transformers without sacrificing context or incurring cascading pipeline errors. For organizations managing extensive documentation, compliance filings, or meeting records, this architecture offers high-quality automated summarization and built-in traceability. The authors recommend adopting joint dynamic-weight architectures for long-context tasks and exploring their expansion into related domains such as enterprise question answering and multi-turn dialogue analysis.

Decision-makers should note that the model was slightly outperformed on the arXiv dataset, where source articles require broader cross-sentence synthesis. The approach also remains subject to minor information loss inherent in snippet selection and hardware constraints on the maximum number of extracted snippets. Nevertheless, given the consistent gains across diverse benchmark domains, there is high confidence in the framework's effectiveness for structured documents and dialogue records.

Mao et al (2022).pdf

No sufficiently relevant recommendations were found.

Cover for DYLE: Dynamic Latent Extraction for Abstractive Long-Input Summarization

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Our Approach
  • 3.1 Extractor-Generator Framework
  • 3.2 Extractor for Long Inputs
  • 3.3 Generator with Dynamic Weights
  • 3.4 Leveraging Extractive Oracles
  • 3.5 Training Objective
  • 4 Experiment Setups
  • 4.1 Datasets
  • 4.2 Baselines and Implementation
  • 4.3 Implementation Details
  • 5 Experiment Results
  • 5.1 Main Results
  • 5.2 Evaluation of Auxiliary Optimizations
  • 6 Analysis and Discussion
  • 7 Conclusions
  • Acknowledgment
  • References
  • A Additional Dynamic Weight Visualization

Knowls

  1. Knowl 1 — DYLE uses a chunked top-K extractor to supply latent snippets to a generator

    model/method

    DYLE represents a long input as snippets X=(x_1,ldots,x_L)—sentences for documents and utterances for dialogues—with an optional query qq. A shared RoBERTa encoder processes consecutive snippets in separate chunks, with the query concatenated to each chunk; an MLP maps each snippet representation to an extractor score si=Eη(q,xi)s_i=E_\eta(q,x_i). The extractor selects the KK highest-scoring snippets, denoted XKX_K. The generator conditions on these selected snippets rather than the full input, and they remain latent in the sense that DYLE jointly trains the extractor and generator without treating extracted snippets as a separately fixed intermediate output. The top-KK selection is nondifferentiable, so the generation loss is not backpropagated through that selection into the extractor; DYLE instead trains the extractor with oracle and consistency losses. The method reduces the input presented to the generator, but the paper notes that joint training does not eliminate information loss.

  2. Knowl 2 — The generator marginalizes over snippets with decoding-dependent weights

    equation

    At output step tt, DYLE computes a token distribution separately from each selected snippet and mixes those distributions using weights that depend on the already generated tokens. Let yty_t be the next summary token, y<ty_{<t} the preceding tokens, qq an optional query, and XKX_K the set of KK extracted snippets. For each x∈XKx\in X_K, the generator obtains a contextual representation htxh_t^x from (q,x,y<t)(q,x,y_{<t}), uses a language-model head to compute Pθ(yt∣q,x,y<t)P_\theta(y_t\mid q,x,y_{<t}), and maps htxh_t^x to a weight logit. A softmax over the logits for XKX_K defines Pθ(x∣q,XK,y<t)P_\theta(x\mid q,X_K,y_{<t}). The resulting token and sequence probabilities are

    Pθ(yt∣q,XK,y<t)=∑x∈XKPθ(yt∣q,x,y<t)Pθ(x∣q,XK,y<t),Pθ(y∣q,XK)=∏t=1TPθ(yt∣q,XK,y<t),P_\theta(y_t\mid q,X_K,y_{<t})=\sum_{x\in X_K}P_\theta(y_t\mid q,x,y_{<t})P_\theta(x\mid q,X_K,y_{<t}),\qquad P_\theta(y\mid q,X_K)=\prod_{t=1}^{T}P_\theta(y_t\mid q,X_K,y_{<t}),

    where TT is the summary length and θ\theta denotes generator parameters. Because the snippet weights depend on y<ty_{<t}, the generator can change which extracted information it uses at different decoding steps. Snippets are encoded independently, so they do not attend to one another.

  3. Knowl 3 — Consistency loss aligns extractor scores with average generator weights

    equation

    DYLE trains the extractor to match its distribution over selected snippets to the generator's average dynamic-weight distribution. For a summary of length TT, the consistency loss is

    Lconsistη=KL ⁣(1T∑t=1TPθ(⋅∣q,XK,y<t)  ∥  softmaxx∈XK(Eη(q,x))),L_{\mathrm{consist}}^\eta= \mathrm{KL}\!\left( \frac{1}{T}\sum_{t=1}^{T}P_\theta(\cdot\mid q,X_K,y_{<t}) \;\middle\|\; \mathrm{softmax}_{x\in X_K}\big(E_\eta(q,x)\big) \right),

    where Pθ(⋅∣q,XK,y<t)P_\theta(\cdot\mid q,X_K,y_{<t}) is the generator's distribution over the KK snippets at step tt, Eη(q,x)E_\eta(q,x) is the extractor score for snippet xx, and KL(p∥r)\mathrm{KL}(p\|r) is Kullback–Leibler divergence from distribution pp to rr. The loss is used to update extractor parameters η\eta, not generator parameters θ\theta, and gradients are not propagated through top-KK selection.

  4. Knowl 4 — Greedy ROUGE oracles supervise extractor salience scores

    model/method

    DYLE constructs extractive oracles XoX_o during training by starting with an empty set and greedily adding an input snippet whose addition to the already selected snippets maximizes the mean of ROUGE-1, ROUGE-2, and ROUGE-L against the gold summary. The oracle loss trains the extractor to assign probability mass to these oracle snippets:

    Loracleη=−1∣Xo∣∑x∈Xolog⁡exp⁡(Eη(q,x))∑xi∈Xexp⁡(Eη(q,xi)),L_{\mathrm{oracle}}^\eta=-\frac{1}{|X_o|}\sum_{x\in X_o}\log\frac{\exp(E_\eta(q,x))}{\sum_{x_i\in X}\exp(E_\eta(q,x_i))},

    where XX is the full set of input snippets, qq is an optional query, and Eη(q,x)E_\eta(q,x) is the extractor score. Thus each oracle snippet is a positive target in a softmax over all input snippets. Oracles provide training supervision only; DYLE does not use them at test time.

  5. Knowl 5 — Hybrid training combines oracle snippets with extractor-ranked snippets

    algorithm

    During training, DYLE forms the generator's KK-snippet input using both oracle and model-selected snippets. If the oracle set XoX_o contains at least KK snippets, the input XKX_K is the first KK oracle snippets in the greedy selection order. If ∣Xo∣<K|X_o|<K, DYLE includes every oracle snippet and fills the remaining positions with the extractor's highest-ranked snippets that are not already in XoX_o. This gives the generator oracle-informed inputs while exposing it to extractor selections that may be useful despite being absent from the oracle. At test time, DYLE uses extractor-ranked snippets, not oracle snippets.

  6. Knowl 6 — DYLE separates generator and extractor optimization across three losses

    equation

    The total training objective is

    Lθ,η=λgLgenθ+λoLoracleη+λcLconsistη,L^{\theta,\eta}=\lambda_gL_{\mathrm{gen}}^\theta+\lambda_oL_{\mathrm{oracle}}^\eta+\lambda_cL_{\mathrm{consist}}^\eta,

    where Lgenθ=−log⁡Pθ(y∣q,XK)L_{\mathrm{gen}}^\theta=-\log P_\theta(y\mid q,X_K) is the negative log-likelihood of the gold summary, LoracleηL_{\mathrm{oracle}}^\eta supervises oracle snippets, LconsistηL_{\mathrm{consist}}^\eta aligns extractor and generator snippet distributions, and the λ\lambda values weight the losses. The generator is optimized only with the generation loss; the extractor is optimized only with the oracle and consistency losses. The paper used (λg,λo,λc)=(1,1,1)(\lambda_g,\lambda_o,\lambda_c)=(1,1,1) for QMSum, (0.5,1,1)(0.5,1,1) for GovReport, and (0.5,1,5)(0.5,1,5) for arXiv.

  7. Knowl 7 — Evaluation covers long-document and query-based long-dialogue summarization

    experimental setup

    DYLE was evaluated on GovReport and arXiv for long-document summarization and QMSum for query-based meeting summarization. The paper reports source and target lengths of 9,409 and 553 for GovReport, 6,030 and 273 for arXiv, and 9,070 and 69 for QMSum; QMSum has a query, while the two document datasets do not. The extractor was initialized from RoBERTa-base and the generator from BART-large. Training used Adam, learning rates 5×10−55\times10^{-5} for the extractor and 5×10−65\times10^{-6} for the generator, and an effective batch size of 8. The largest tested extraction count was K=25K=25, and performance was evaluated with ROUGE-1, ROUGE-2, and ROUGE-L.

  8. Knowl 8 — DYLE leads on GovReport and QMSum but trails the strongest arXiv baseline

    empirical result

    Against previously reported results, DYLE obtained ROUGE-1/2/L scores of 61.01/28.83/57.82 on GovReport, exceeding the best listed prior system, Sinkhorn with HEPOS at maximum input length 10,240 (56.86/22.62/53.82), by 4.15/6.21/4.00 points. On QMSum, DYLE scored 34.42/9.71/30.10, above DialogLM's 34.02/9.19/29.77 and DialogLM-Sparse's 33.69/9.32/30.01 on all three metrics. On arXiv, DYLE scored 46.41/17.95/41.54; it did not match LSH's 48.24/20.26 ROUGE-1/2 scores or HAT-BART's 42.17 ROUGE-L score. The authors suggest that DYLE's weaker comparison with LSH on arXiv may reflect arXiv's shorter inputs and more abstractive summaries, which can make individual snippets less suitable generation units.

  9. Knowl 9 — Ablations show gains from oracle supervision, consistency, and hybrid inputs

    empirical result

    Removing any of the three auxiliary optimizations lowered ROUGE-1/2/L on both GovReport and QMSum. The full model scored 61.01/28.83/57.82 on GovReport and 34.42/9.71/30.10 on QMSum. Without hybrid training, scores were 60.89/28.28/57.31 and 31.77/8.33/28.37; without consistency loss, 60.59/28.48/57.49 and 32.51/8.77/28.94; and without oracle loss, 57.57/25.92/53.14 and 32.13/8.38/28.63, respectively. Removing oracle supervision caused the largest GovReport drop, while removing hybrid training caused the largest QMSum drop. In the no-hybrid ablation, only oracle snippets train the generator, and consistency is computed over the oracle set; in the no-consistency ablation, extractor and generator are trained independently with oracle-based inputs.

  10. Knowl 10 — Dynamic weights expose snippet use, while extraction favors recall over precision

    empirical result

    Visualizations of DYLE's dynamic snippet weights show multiple consecutive high-weight regions when generating actual summaries, consistent with alignments between snippets and generated content; weights for random summaries are more uniformly distributed. The visualizations show fewer snippets receiving weight on query-based QMSum than on general-summary GovReport, which the authors attribute to queried information being concentrated in fewer snippets. ROUGE precision–recall decompositions also show that extracted snippets tend to have higher recall and lower precision than generated summaries. For ROUGE-1 on GovReport, extracted snippets had precision 48.98, recall 73.40, and F1 57.56, versus 63.16, 61.61, and 61.01 for generated summaries. On QMSum, the corresponding extracted-snippet values were 4.25, 76.90, and 7.74, versus 29.78, 45.64, and 34.42 for generated summaries. These diagnostics identify information coverage in extraction and selection accuracy in generation as distinct opportunities for improvement.

Coverage note — The separate sweep over extraction counts and the generator performance using gold extractive oracles are omitted as secondary diagnostics; the method, principal benchmark comparisons, ablations, and snippet-use analyses are included.

References

  1. 1.Sanghwan Bae, Taeuk Kim, Jihoon Kim, and Sanggoo Lee. 2019. Summary level training of sentence rewriting for abstractive summarization. In Proceedings of the 2nd Workshop on New Frontiers in Summarization, pages 10–20.
  2. 2.Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015. Neural machine translation by jointly learning to align and translate. In ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings.
  3. 3.Ahsaas Bajaj, Pavitra Dangati, Kalpesh Krishna, Pradhiksha Ashok Kumar, Rheeya Uppaal, Bradford Windsor, Eliot Brenner, Dominic Dotterrer, Rajarshi Das, and Andrew McCallum. 2021. Long document summarization in a low resource setting using pretrained language models. In Proceedings of the ACL-IJCNLP 2021 Student Research Workshop, ACL 2021, Online, JUli 5-10, 2021, pages 71–80.
  4. 4.Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020. Longformer: The long-document transformer. CoRR, abs/2004.05150.
  5. 5.Arthur Bražinskas, Mirella Lapata, and Ivan Titov. 2021. Learning opinion summarizers by selecting informative reviews. arXiv e-prints, pages arXiv–2109.
  6. 6.Jean Carletta, Simone Ashby, Sebastien Bourban, Mike Flynn, Mael Guillemot, Thomas Hain, Jaroslav Kadlec, Vasilis Karaiskos, Wessel Kraaij, Melissa Kronenthal, et al. 2005. The ami meeting corpus: A pre-announcement. In International workshop on machine learning for multimodal interaction, pages 28–39. Springer.
  7. 7.Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin. 2016. Training deep nets with sublinear memory cost. CoRR, abs/1604.06174.
  8. 8.Yen-Chun Chen and Mohit Bansal. 2018. Fast abstractive summarization with reinforce-selected sentence rewriting. In Proceedings of ACL 2018, Long Papers, pages 675–686.
  9. 9.Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019. Generating long sequences with sparse transformers. CoRR, abs/1904.10509.
  10. 10.Arman Cohan, Franck Dernoncourt, Doo Soon Kim, Trung Bui, Seokhwan Kim, Walter Chang, and Nazli Goharian. 2018. A discourse-aware attention model for abstractive summarization of long documents. In Proceedings of NAACL-HLT 2018, Short Papers, pages 615–621, New Orleans, Louisiana.
  11. 11.Peng Cui and Le Hu. 2021. Sliding selector network with dynamic memory for extractive summarization of long documents. In Proceedings of NAACL-HLT 2021, pages 5881–5891.
  12. 12.Alexios Gidiotis and Grigorios Tsoumakas. 2020. A divide-and-conquer approach to the summarization of long documents. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 28.
  13. 13.Quentin Grail, Julien Perez, and Eric Gaussier. 2021. Globalizing BERT-based transformer architectures for long document summarization. In Proceedings of EACL 2021, pages 1792–1810, Online.
  14. 14.Luyang Huang, Shuyang Cao, Nikolaus Nova Parulian, Heng Ji, and Lu Wang. 2021. Efficient attentions for long document summarization. In Proceedings of NAACL-HLT 2021, Online, June 6-11, 2021, pages 1419–1436.
  15. 15.Adam Janin, Don Baron, Jane Edwards, Dan Ellis, David Gelbart, Nelson Morgan, Barbara Peskin, Thilo Pfau, Elizabeth Shriberg, Andreas Stolcke, et al. 2003. The icsi meeting corpus. In ICASSP 2003., volume 1, pages I–I. IEEE.
  16. 16.Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya. 2020. Reformer: The efficient transformer. In ICLR.
  17. 17.Farshad Kiyoumarsi. 2015. Evaluation of automatic text summarizations based on human summaries. Procedia-Social and Behavioral Sciences, 192:83–91.
  18. 18.Anastassia Kornilova and Vladimir Eidelman. 2019. Billsum: A corpus for automatic summarization of us legislation. In Proceedings of the 2nd Workshop on New Frontiers in Summarization, pages 48–56.
  19. 19.Logan Lebanoff, Kaiqiang Song, Franck Dernoncourt, Doo Soon Kim, Seokhwan Kim, Walter Chang, and Fei Liu. 2019. Scoring sentence singletons and pairs for abstractive summarization. In Proceedings of ACL 2019, pages 2175–2189, Florence, Italy.
  20. 20.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020a. BART: denoising sequence-to-sequence pretraining for natural language generation, translation, and comprehension. In Proceedings of ACL 2020, Online, July 5-10, 2020, pages 7871–7880.
  21. 21.Patrick S. H. Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020b. Retrieval-augmented generation for knowledge-intensive NLP tasks. In NeurIPS.
  22. 22.Yang Liu, Chenguang Zhu, and Michael Zeng. 2021. End-to-end segmentation-based news summarization. arXiv preprint arXiv:2110.07850.
  23. 23.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized BERT pretraining approach. CoRR, abs/1907.11692.
  24. 24.Shashi Narayan, Shay B Cohen, and Mirella Lapata. 2018. Ranking sentences for extractive summarization with reinforcement learning. In NAACL-HLT.
  25. 25.Ansong Ni, Zhangir Azerbayev, Mutethia Mutuma, Troy Feng, Yusen Zhang, Tao Yu, Ahmed Hassan Awadallah, and Dragomir Radev. 2021a. SummerTime: Text summarization toolkit for non-experts. In Proceedings of EMNLP 2021: System Demonstrations, pages 329–338.
  26. 26.Ansong Ni, Matt Gardner, and Pradeep Dasigi. 2021b. Mitigating false-negative contexts in multi-document question answering with retrieval marginalization. In Proceedings of EMNLP 2021, pages 6149–6161.
  27. 27.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21:140:1–140:67.
  28. 28.Tobias Rohde, Xiaoxia Wu, and Yinhan Liu. 2021. Hierarchical learning for generation with long source sequences. arXiv preprint arXiv:2104.07545.
  29. 29.Eva Sharma, Chen Li, and Lu Wang. 2019. Bigpatent: A large-scale dataset for abstractive and coherent summarization. In Proceedings of ACL 2019, pages 2204–2213.
  30. 30.Xiaofei Sun, Chun Fan, Zijun Sun, Yuxian Meng, Fei Wu, and Jiwei Li. 2020. Summarize, outline, and elaborate: Long-text generation via hierarchical supervision from extractive summaries. arXiv preprint arXiv:2010.07074.
  31. 31.Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler. 2020a. Long range arena: A benchmark for efficient transformers. arXiv preprint arXiv:2011.04006.
  32. 32.Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler. 2020b. Efficient transformers: A survey. CoRR, abs/2009.06732.
  33. 33.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In NIPS 2017, 4-9 December 2017, Long Beach, CA, USA, pages 5998–6008.
  34. 34.Ronald J. Williams. 1992. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Mach. Learn., 8:229–256.
  35. 35.Jeff Wu, Long Ouyang, Daniel M Ziegler, Nissan Stiennon, Ryan Lowe, Jan Leike, and Paul Christiano. 2021. Recursively summarizing books with human feedback. arXiv preprint arXiv:2109.10862.
  36. 36.Wen Xiao and Giuseppe Carenini. 2019. Extractive summarization of long documents by combining global and local context. In Proceedings of EMNLP-IJCNLP 2019, pages 3011–3021.
  37. 37.Jiacheng Xu and Greg Durrett. 2019. Neural extractive text summarization with syntactic compression. In Proceedings of EMNLP-IJCNLP 2019.
  38. 38.Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al. 2020. Big bird: Transformers for longer sequences. In NeurIPS.
  39. 39.Haoyu Zhang, Jingjing Cai, Jianjun Xu, and Ji Wang. 2019. Pretraining-based natural language generation for text summarization. In Proceedings of the 23rd Conference on Computational Natural Language Learning (CoNLL).
  40. 40.Jingqing Zhang, Yao Zhao, Mohammad Saleh, and Peter Liu. 2020. Pegasus: Pre-training with extracted gap-sentences for abstractive summarization. In International Conference on Machine Learning, pages 11328–11339. PMLR.
  41. 41.Yusen Zhang, Ansong Ni, Ziming Mao, Chen Henry Wu, Chenguang Zhu, Budhaditya Deb, Ahmed Hassan Awadallah, Dragomir R. Radev, and Rui Zhang. 2021a. Summˆn: A multi-stage summarization framework for long input dialogues and documents. CoRR, abs/2110.10150.
  42. 42.Yusen Zhang, Ansong Ni, Tao Yu, Rui Zhang, Chenguang Zhu, Budhaditya Deb, Asli Celikyilmaz, Ahmed Hassan Awadallah, and Dragomir Radev. 2021b. An exploratory study on long dialogue summarization: What works and what’s next. In EMNLP 2021: Findings.
  43. 43.Ming Zhong, Yang Liu, Yichong Xu, Chenguang Zhu, and Michael Zeng. 2021a. Dialoglm: Pre-trained model for long dialogue understanding and summarization. arXiv preprint arXiv:2109.02492.
  44. 44.Ming Zhong, Da Yin, Tao Yu, Ahmad Zaidi, Mutethia Mutuma, Rahul Jha, Ahmed Hassan Awadallah, Asli Celikyilmaz, Yang Liu, Xipeng Qiu, and Dragomir R. Radev. 2021b. Qmsum: A new benchmark for query-based multi-domain meeting summarization. In Proceedings of NAACL-HLT 2021, Online, June 6-11, 2021.
  45. 45.Chenguang Zhu, Ruochen Xu, Michael Zeng, and Xuedong Huang. 2020. A hierarchical network for abstractive meeting summarization with cross-domain pretraining. In Proceedings of EMNLP 2020: Findings.

Citation

MLA
Mao, Z., et al. “DYLE: Dynamic Latent Extraction for Abstractive Long-Input Summarization”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 1687–98, https://doi.org/10.18653/v1/2022.acl-long.118.
APA
Mao, Z., Wu, C. H., Ni, A., Zhang, Y., Zhang, R., Yu, T., Deb, B., Zhu, C., Awadallah, A., & Radev, D. (2022). DYLE: Dynamic Latent Extraction for Abstractive Long-Input Summarization. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1687–1698. https://doi.org/10.18653/v1/2022.acl-long.118
Chicago
Mao, Z., C. H. Wu, A. Ni, et al. 2022. “DYLE: Dynamic Latent Extraction for Abstractive Long-Input Summarization”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1687–98. https://doi.org/10.18653/v1/2022.acl-long.118.
Harvard
Mao, Z. et al. (2022) “DYLE: Dynamic Latent Extraction for Abstractive Long-Input Summarization”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 1687–1698. Available at: https://doi.org/10.18653/v1/2022.acl-long.118.
Vancouver
1. Mao Z, Wu CH, Ni A, Zhang Y, Zhang R, Yu T, Deb B, Zhu C, Awadallah A, Radev D (2022) DYLE: Dynamic Latent Extraction for Abstractive Long-Input Summarization. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 1687–1698

BibTeX

@inproceedings{mao-etal-2022-dyle,
    title = "{DYLE}: Dynamic Latent Extraction for Abstractive Long-Input Summarization",
    author = "Mao, Ziming  and
      Wu, Chen Henry  and
      Ni, Ansong  and
      Zhang, Yusen  and
      Zhang, Rui  and
      Yu, Tao  and
      Deb, Budhaditya  and
      Zhu, Chenguang  and
      Awadallah, Ahmed  and
      Radev, Dragomir",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.118/",
    doi = "10.18653/v1/2022.acl-long.118",
    pages = "1687--1698"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/