Built independently by an author, for readers. Read the story and support ChapterPal

keyword

multi-document summarization

Multi-document summarization is a natural language processing task that involves automatically generating a concise and coherent summary from a collection of multiple documents covering the same or related topics. Unlike single-document summarization, this process requires identifying and aggregating key information distributed across distinct sources, synthesizing shared themes, and reconciling differing perspectives while filtering out redundant or contradictory statements. The resulting summary can be produced using extractive methods that select the most salient sentences directly from the source texts, or abstractive methods that paraphrase and generate new phrasing to create a unified, structured overview.

8 items

Re-Examining System-Level Correlations of Automatic Summarization Evaluation Metrics

Re-Examining System-Level Correlations of Automatic Summarization Evaluation Metrics

Daniel Deutsch, Rotem Dror, Dan Roth

OrganizationsUniversity of Pennsylvania

Why you should read this

Reveals critical flaws in standard summarization metric evaluations by demonstrating that common metrics like ROUGE show near-zero correlation with human judgments when discriminating between similarly performing systems.

How reliably an automatic summarization evaluation metric replicates human judgments of summary quality is quantified by system-level correlations. We identify two ways in which the definition of the system-level correlation is inconsistent with how metrics are used to evaluate systems in practice and propose changes to rectify this disconnect. First, we calculate the system score for an automatic metric using the full test set instead of the subset of summaries judged by humans, which is currently standard practice. We demonstrate how this small change leads to more precise estimates of system-level correlations. Second, we propose to calculate correlations only on pairs of systems that are separated by small differences in automatic scores which are commonly observed in practice. This allows us to demonstrate that our best estimate of the correlation of ROUGE to human judgments is near 0 in realistic scenarios. The results from the analyses point to the need to collect more high-quality human judgments and to improve automatic metrics when differences in system scores are small.¹

Added

2026-10-03

LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, Yuxiao Dong, Jie Tang, Juanzi Li

OrganizationsInstitute of Automation, Chinese Academy of SciencesTsinghua UniversityZhipu AI

Why you should read this

Presents LongBench, the first bilingual multi-task benchmark spanning 21 datasets across six task categories to systematically evaluate and compare how large language models comprehend long-context inputs in English and Chinese.

Although large language models (LLMs) demonstrate impressive performance for many language tasks, most of them can only handle texts a few thousand tokens long, limiting their applications on longer sequence inputs, such as books, reports, and codebases. Recent works have proposed methods to improve LLMs' long context capabilities by extending context windows and more sophisticated memory mechanisms. However, comprehensive benchmarks tailored for evaluating long context understanding are lacking. In this paper, we introduce LongBench, the first bilingual, multi-task benchmark for long context understanding, enabling a more rigorous evaluation of long context understanding. LongBench comprises 21 datasets across 6 task categories in both English and Chinese, with an average length of 6,711 words (English) and 13,386 characters (Chinese). These tasks cover key long-text application areas including single-doc QA, multi-doc QA, summarization, few-shot learning, synthetic tasks, and code completion. All datasets in LongBench are standardized into a unified format, allowing for effortless automatic evaluation of LLMs. Upon comprehensive evaluation of 8 LLMs on LongBench, we find that: (1) Commercial model (GPT-3.5-Turbo-16k) outperforms other open-sourced models, but still struggles on longer contexts. (2) Scaled position embedding and fine-tuning on longer sequences lead to substantial improvement on long context understanding. (3) Context compression technique such as retrieval brings improvement for model with weak ability on long contexts, but the performance still lags behind models that have strong long context understanding capability. The code and datasets are available at this https URL.

Added

2026-09-24

PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization

PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization

Jingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. Liu

OrganizationsGoogleImperial College London

Why you should read this

Proposes a gap-sentence pre-training objective specifically designed for abstractive text summarization, setting state-of-the-art performance across 12 diverse datasets while demonstrating remarkable sample efficiency with as few as 1,000 fine-tuning examples.

Recent work pre-training Transformers with self-supervised objectives on large text corpora has shown great success when fine-tuned on downstream NLP tasks including text summarization. However, pre-training objectives tailored for abstractive text summarization have not been explored. Furthermore there is a lack of systematic evaluation across diverse domains. In this work, we propose pre-training large Transformer-based encoder-decoder models on massive text corpora with a new self-supervised objective. In PEGASUS, important sentences are removed/masked from an input document and are generated together as one output sequence from the remaining sentences, similar to an extractive summary. We evaluated our best PEGASUS model on 12 downstream summarization tasks spanning news, science, stories, instructions, emails, patents, and legislative bills. Experiments demonstrate it achieves state-of-the-art performance on all 12 downstream datasets measured by ROUGE scores. Our model also shows surprising performance on low-resource summarization, surpassing previous state-of-the-art results on 6 datasets with only 1000 examples. Finally we validated our results using human evaluation and show that our model summaries achieve human performance on multiple datasets.

Added

2026-09-14

LexRank: Graph-based Lexical Centrality as Salience in Text Summarization

LexRank: Graph-based Lexical Centrality as Salience in Text Summarization

Günes Erkan, Dragomir R. Radev

OrganizationsUniversity of Michigan

Why you should read this

Introduces LexRank, a foundational graph-based algorithm that determines sentence importance using eigenvector centrality on lexical similarity graphs, consistently outperforming centroid methods in extractive text summarization.

We introduce a stochastic graph-based method for computing relative importance of textual units for Natural Language Processing. We test the technique on the problem of Text Summarization (TS). Extractive TS relies on the concept of sentence salience to identify the most important sentences in a document or set of documents. Salience is typically defined in terms of the presence of particular important words or in terms of similarity to a centroid pseudo-sentence. We consider a new approach, LexRank, for computing sentence importance based on the concept of eigenvector centrality in a graph representation of sentences. In this model, a connectivity matrix based on intra-sentence cosine similarity is used as the adjacency matrix of the graph representation of sentences. Our system, based on LexRank ranked in first place in more than one task in the recent DUC 2004 evaluation. In this paper we present a detailed analysis of our approach and apply it to a larger data set including data from earlier DUC evaluations. We discuss several methods to compute centrality using the similarity graph. The results show that degree-based methods (including LexRank) outperform both centroid-based methods and other systems participating in DUC in most of the cases. Furthermore, the LexRank with threshold method outperforms the other degree-based techniques including continuous LexRank. We also show that our approach is quite insensitive to the noise in the data that may result from an imperfect topical clustering of documents.

Added

2026-09-11