Built independently by an author, for readers. Read the story and support ChapterPal

keyword

BERTS2S

BERTS2S, short for BERT sequence-to-sequence, is a neural network architecture designed for conditional text generation tasks such as abstractive document summarization and machine translation. Built on the standard Transformer encoder-decoder framework, the architecture leverages transfer learning by initializing both the encoder and decoder with pretrained Bidirectional Encoder Representations from Transformers checkpoints rather than initializing them randomly. In this setup, the encoder and autoregressive decoder typically share pretrained parameter weights, while the intervening encoder-decoder cross-attention layers are initialized randomly and learned during fine-tuning on the target generation task. By adapting rich bidirectional language representations to a sequence-to-sequence pipeline, BERTS2S substantially improves output fluency, semantic coverage, and factual consistency compared to standard sequence models trained from scratch.

1 item

On Faithfulness and Factuality in Abstractive Summarization

On Faithfulness and Factuality in Abstractive Summarization

Joshua Maynez, Shashi Narayan, Bernd Bohnet, Ryan T. McDonald

OrganizationsGoogle

Why you should read this

Reveals widespread factual hallucinations across neural abstractive summarization models through human evaluation and establishes textual entailment as a superior alternative to standard ROUGE metrics for measuring faithfulness.

It is well known that the standard likelihood training and approximate decoding objectives in neural text generation models lead to less human-like responses for open-ended tasks such as language modeling and story generation. In this paper we have analyzed limitations of these models for abstractive document summarization and found that these models are highly prone to hallucinate content that is unfaithful to the input document. We conducted a large scale human evaluation of several neural abstractive summarization systems to better understand the types of hallucinations they produce. Our human annotators found substantial amounts of hallucinated content in all model generated summaries. However, our analysis does show that pretrained models are better summarizers not only in terms of raw metrics, i.e., ROUGE, but also in generating faithful and factual summaries as evaluated by humans. Furthermore, we show that textual entailment measures better correlate with faithfulness than standard metrics, potentially leading the way to automatic evaluation metrics as well as training and decoding criteria.

Added

2026-09-24