SummaReranker: A Multi-Task Mixture-of-Experts Re-ranking Framework for Abstractive Summarization

Mathieu RavautShafiq R. JotyNancy F. Chen

article2022ACL122 citations

Proposes a multi-task mixture-of-experts re-ranking framework that jointly optimizes across multiple evaluation metrics to select superior candidate summaries from standard generation models, setting new state-of-the-art results across several standard abstractive summarization benchmarks.

Listen

Modern automated text summarization models frequently fail to output their best possible summary. Standard systems generate text word by word using search algorithms like beam search, which pick a single summary based on sequence probabilities. However, this process often discards other candidate summaries that are significantly higher in quality, coherence, and factual accuracy. Analysis indicates that the theoretical best summary among a pool of generated options—known as the oracle score—can exceed standard outputs by up to 30.5% in standard overlap metrics (ROUGE-1). This massive gap demonstrates that underlying text generation models are currently underutilized due to suboptimal candidate selection.

The article demonstrates and evaluates a second-stage re-ranking system called SummaReranker. The objective is to design a lightweight, multi-task framework that re-evaluates a pool of summary candidates generated by an underlying language model and selects the highest-quality output.

To achieve this, the authors implemented a mixture-of-experts model built on top of RoBERTa-large. The re-ranker evaluates candidate summaries jointly with the original source document and estimates the probability that a candidate is the best choice across multiple standard evaluation metrics simultaneously (including word-overlap metrics like ROUGE and embedding-based metrics like BERTScore and BARTScore). The approach was evaluated across three distinct datasets representing news and social media domains (CNN-DailyMail, XSum, and Reddit TIFU) using two state-of-the-art base language models (PEGASUS and BART).

The evaluation produced several key findings. First, SummaReranker established a new state of the art in summarization, improving standard overlap performance by 5.44% on CNN-DailyMail, 1.31% on XSum, and 9.34% on Reddit TIFU compared to standard generation. Second, combining diverse candidate generation methods (such as standard beam search and diverse beam search) consistently improved re-ranking quality. Third, base models trained on only 50% of the training data and augmented with SummaReranker outperformed models trained on 100% of the data without re-ranking. Finally, both automated recall analyses and human evaluations confirmed that the re-ranker consistently selects summaries that are more abstractive, complete, and faithful than baseline outputs.

These findings indicate that organizations deploying automated summarization can achieve substantial quality improvements without the massive computational expense of retraining foundational generation models. Training the re-ranker requires only a fraction of the compute of the primary model (approximately four days on a single graphics processing unit for standard news datasets). While scoring candidates adds a modest latency during live generation (roughly 38 milliseconds per candidate), practical trade-off analysis shows that scoring as few as six to eight candidates captures most performance gains efficiently.

For practical implementation, organizations should adopt second-stage re-ranking when summarization quality and factual consistency are critical, using a balanced pool of six to eight candidates generated via diverse search strategies. Before deploying the system on lengthy documents (such as technical manuals or legal contracts), technical teams must conduct further development, as the current model is constrained by a 512-token input limit. Users should also note that performance gains vary depending on how abstractive the baseline domain already is, showing dramatic gains on standard text but narrower improvements on highly condensed single-sentence summaries.

ntunlp/SummaRerankerRavaut et al (2022).pdf
Cover for SummaReranker: A Multi-Task Mixture-of-Experts Re-ranking Framework for Abstractive Summarization

Abstract

Sequence-to-sequence neural networks have recently achieved great success in abstractive summarization, especially through fine-tuning large pre-trained language models on the downstream dataset. These models are typically decoded with beam search to generate a unique summary. However, the search space is very large, and with the exposure bias, such decoding is not optimal. In this paper, we show that it is possible to directly train a second-stage model performing re-ranking on a set of summary candidates. Our mixture-of-experts SummaReranker learns to select a better candidate and consistently improves the performance of the base model. With a base PEGASUS, we push ROUGE scores by 5.44% on CNN-DailyMail (47.16 ROUGE-1), 1.31% on XSum (48.12 ROUGE-1) and 9.34% on Reddit TIFU (29.83 ROUGE-1), reaching a new state-of-the-art. Our code and checkpoints will be available at https://github.com/ntunlp/SummaReranker.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Model
  • 3.1 Re-ranking Framework
  • 3.2 Model Architecture
  • 3.3 Tackling Training and Inference Gap
  • 4 Experiments
  • 4.1 Scope & Datasets
  • 4.2 Training & Inference Details
  • 4.3 Base Setup Results
  • 4.4 Transfer Setup Results
  • 4.5 Ranking Evaluation
  • 4.6 Qualitative Evaluation
  • 5 Discussion
  • 6 Conclusion
  • Acknowledgements
  • References
  • A Hyper Parameters & Packages
  • B Oracle Scores
  • C Unique Candidates Scores
  • D Identical Candidates Scores
  • E Metrics Correlation
  • F Base Setup Results
  • G Recall Curves
  • H Human Evaluation
  • I Candidate Selection
  • J Speed/Performance Trade-off
  • K Re-ranking Examples
  • CNN/DM
  • Reddit TIFU

Knowls

  1. Knowl 1 — SummaReranker Multi-Task Mixture-of-Experts Architecture for Summary Re-ranking

    model/method

    SummaReranker is a second-stage re-ranking system designed to select the highest quality summary from a pool of mm summary candidates C={C1,…,Cm}\mathcal{C} = \{C_1, \ldots, C_m\} generated by a base abstractive sequence-to-sequence model for a source document SS. The framework formulates re-ranking as multi-task binary classification across NN automatic evaluation metrics M={μ1,…,μN}\mathcal{M} = \{\mu_1, \ldots, \mu_N\} (such as ROUGE-1, ROUGE-2, ROUGE-L, BERTScore, and BARTScore).

    For each candidate summary CiC_i, the input sequence is formed by concatenating the source document and the candidate: [CLS] Source [SEP] Candidate. This sequence is passed to a pre-trained RoBERTa-large encoder. The final hidden state of the [CLS] token is passed through a shared bottom multi-layer perceptron (MLP) consisting of two fully connected layers with ReLU activations to produce a joint representation x∈Rd\boldsymbol{x} \in \mathbb{R}^d.

    To optimize over multiple correlated evaluation metrics simultaneously, the joint representation x\boldsymbol{x} is processed by a Multi-Gate Mixture-of-Experts (MMoE) layer with E=2NE = 2N experts E1,…,EE\mathcal{E}_1, \ldots, \mathcal{E}_E, where each expert is a two-layer MLP with ReLU activations. During training, an expert dropout rate of 50%50\% is applied, ensuring that the expected number of active experts equals the number of tasks NN. For each metric task k∈{1,…,N}k \in \{1, \ldots, N\}, a task-specific gating matrix Wk∈RE×dW_k \in \mathbb{R}^{E \times d} computes softmax routing weights over the experts. The resulting expert mixture is fed into a task-specific prediction tower TkT_k (a single-layer MLP) followed by a sigmoid activation to output the probability pθμk(Ci)p_\theta^{\mu_k}(C_i) that candidate CiC_i is optimal under metric μk\mu_k.

  2. Knowl 2 — Multi-Task Binary Cross-Entropy Loss for Summary Re-ranking

    equation

    For a source document SS with a candidate pool C={C1,…,Cm}\mathcal{C} = \{C_1, \ldots, C_m\} and a specific evaluation metric μ∈M\mu \in \mathcal{M}, the oracle candidate Cμ∗C^*_\mu is defined as:

    Cμ∗=arg⁡max⁡Ci∈Cμ(Ci)C^*_\mu = \arg\max_{C_i \in \mathcal{C}} \mu(C_i)

    The ground-truth binary label yi∈{0,1}y_i \in \{0, 1\} for candidate CiC_i with respect to metric μ\mu is:

    yi={1if Ci=Cμ∗0otherwisey_i = \begin{cases} 1 & \text{if } C_i = C^*_\mu \\ 0 & \text{otherwise} \end{cases}

    The re-ranking model parameterized by θ\theta predicts the probability pθμ(Ci)p_\theta^\mu(C_i) that CiC_i is the optimal candidate, trained using binary cross-entropy loss:

    Lμ=−yilog⁡pθμ(Ci)−(1−yi)log⁡(1−pθμ(Ci))\mathcal{L}_\mu = - y_i \log p_\theta^\mu(C_i) - (1 - y_i) \log (1 - p_\theta^\mu(C_i))

    To jointly optimize over NN evaluation metrics M={μ1,…,μN}\mathcal{M} = \{\mu_1, \ldots, \mu_N\}, the total multi-task loss L\mathcal{L} is the unweighted average across all metric-specific losses:

    L=1N∑μ∈MLμ\mathcal{L} = \frac{1}{N} \sum_{\mu \in \mathcal{M}} \mathcal{L}_\mu

  3. Knowl 3 — Multi-Gate Mixture-of-Experts Prediction and Inference Candidate Selection

    equation

    In the SummaReranker architecture, given the joint source-candidate representation x∈Rd\boldsymbol{x} \in \mathbb{R}^d, EE expert subnetworks E1,…,EE\mathcal{E}_1, \ldots, \mathcal{E}_E, a metric-specific gating weight matrix Wk∈RE×dW_k \in \mathbb{R}^{E \times d}, and a metric-specific prediction tower TkT_k, the pre-activation prediction fθk(x)f_\theta^k(\boldsymbol{x}) for metric μ\mu indexed by task k∈{1,…,N}k \in \{1, \ldots, N\} is:

    fθk(x)=Tk(∑i=1E[softmax(Wkx)]iEi(x))f_\theta^k(\boldsymbol{x}) = T_k\left( \sum_{i=1}^E \left[\text{softmax}(W_k \boldsymbol{x})\right]_i \mathcal{E}_i(\boldsymbol{x}) \right)

    where [softmax(Wkx)]i=exp⁡((Wkx)i)∑j=1Eexp⁡((Wkx)j)\left[\text{softmax}(W_k \boldsymbol{x})\right]_i = \frac{\exp((W_k \boldsymbol{x})_i)}{\sum_{j=1}^E \exp((W_k \boldsymbol{x})_j)} denotes the routing weight assigned to expert Ei\mathcal{E}_i for metric task kk.

    The predicted probability pθμ(Ci)p_\theta^\mu(C_i) that candidate CiC_i maximizes metric μ\mu is computed via the sigmoid function:

    pθμ(Ci)=σ(fθk(x))=11+exp⁡(−fθk(x))p_\theta^\mu(C_i) = \sigma\left(f_\theta^k(\boldsymbol{x})\right) = \frac{1}{1 + \exp\left(-f_\theta^k(\boldsymbol{x})\right)}

    At inference time, to select a single summary from the candidate pool C={C1,…,Cm}\mathcal{C} = \{C_1, \ldots, C_m\}, SummaReranker selects the candidate maximizing the sum of predicted probabilities over all target evaluation metrics M\mathcal{M}:

    C∗=arg⁡max⁡Ci∈C∑μ∈Mpθμ(Ci)C^* = \arg\max_{C_i \in \mathcal{C}} \sum_{\mu \in \mathcal{M}} p_\theta^\mu(C_i)

  4. Knowl 4 — Two-Fold Cross-Inference Training and Extreme Candidate Sampling Strategy

    model/method

    To prevent train-test distribution mismatch caused by base generation models overfitting their own training sets, SummaReranker employs a two-fold data partitioning strategy:

    1. Two-fold Split Training: The training set is shuffled and partitioned into two equal disjoint halves (AA and BB). A base sequence-to-sequence model is fine-tuned independently on half AA and on half BB. The model fine-tuned on AA performs decoding on half BB, and the model fine-tuned on BB decodes on half AA. These out-of-fold generated candidate sets are combined to form the re-ranker's training dataset.
    2. Extreme Candidate Sampling (mtop=1,mbottom=1m_{\text{top}}=1, m_{\text{bottom}}=1): Generated candidates are ranked by decreasing sum of normalized scores across the evaluation metrics. During training, only the top candidate (mtop=1m_{\text{top}}=1) and bottom candidate (mbottom=1m_{\text{bottom}}=1) are retained as positive and negative training instances per document. This reduces per-step training complexity from O(m)\mathcal{O}(m) to O(mtop+mbottom)=O(1)\mathcal{O}(m_{\text{top}} + m_{\text{bottom}}) = \mathcal{O}(1) while providing maximum discriminative contrast.
    3. Evaluation Modes:
      • Base setup: The base model fine-tuned on 50% data generates test candidates, which are ranked by SummaReranker.
      • Transfer setup: The base model fine-tuned on 100% data generates test candidates, which are ranked by SummaReranker (trained only on the 50% split outputs).
  5. Knowl 5 — Transfer Setup Summarization Performance on CNN/DailyMail, XSum, and Reddit TIFU

    data/table

    In the transfer setup, SummaReranker (SR) is applied to re-rank candidate pools generated by base PEGASUS-large and BART-large models fine-tuned on the full training sets. R-1, R-2, and R-L denote ROUGE-1, ROUGE-2, and ROUGE-L F1 scores; BS and BaS denote BERTScore and BARTScore. Decoding methods include beam search {1}\{1\} and diverse beam search {2}\{2\}.

    Model Setting Dtrain Dtest m R-1 R-2 R-L
    CNN/DailyMail
    PEGASUS baseline (our setup) {2} {2} 15 44.56 20.90 41.58
    GSum + RefSum (Liu et al., 2021) {1} {1} 4 46.18 22.36 42.91
    BART + SimCLS (Liu and Liu, 2021) {2} {2} 16 46.67 22.15 43.54
    BART + SR (Optimized: R-1/2/L) {1, 2} {1, 2} 30 46.62 22.39 43.59
    PEGASUS + SR (Optimized: R-1/2/L) {1, 2} {1, 2} 30 47.16 22.55 43.87
    XSum
    PEGASUS baseline (our setup) {1} {1} 15 47.33 24.75 39.43
    GSum + RefSum (Liu et al., 2021) {1} {1} 4 47.45 24.55 39.41
    PEGASUS + SimCLS (Liu and Liu, 2021) {2} {2} 16 47.61 24.57 39.44
    PEGASUS + SR (Optimized: R-1/2/L) {1, 2} {1} 15 48.12 24.95 40.00
    Reddit TIFU
    PEGASUS baseline (our setup) {1} {1} 15 26.28 9.01 21.52
    BART baseline (our setup) {1} {1} 15 27.42 9.53 22.10
    BART + SR (Optimized: R-1/2/L) {1, 2} {1} 15 28.99 9.82 22.96
    PEGASUS + SR (Optimized: R-1/2/L) {1, 2} {1, 2} 30 29.83 9.50 23.47

    PEGASUS + SummaReranker sets a new state-of-the-art on CNN/DailyMail (47.16 R-1, +5.44% relative gain over diverse beam search baseline) and on XSum (48.12 R-1, +1.31% relative gain). On Reddit TIFU, PEGASUS + SummaReranker achieves 29.83 R-1 (+9.34% relative gain).

  6. Knowl 6 — Base Setup Summarization Performance on CNN/DailyMail

    data/table

    In the base setup, PEGASUS and BART base models are fine-tuned on only 50% of the CNN/DailyMail training set. Candidate summaries are decoded via beam search {1}\{1\} (15 candidates) and diverse beam search {2}\{2\} (15 candidates). SummaReranker (SR) is trained on out-of-fold candidate sets and jointly optimized for ROUGE-1 (R-1), ROUGE-2 (R-2), and ROUGE-L (R-L).

    Model Decoding Methods R-1 R-2 R-L Mean Gain (%)
    PEGASUS - 1st half {1} 42.23 19.62 38.90 -
    PEGASUS - 1st half {2} 42.50 19.75 39.55 -
    PEGASUS - 2nd half {1} 42.46 19.95 39.19 -
    PEGASUS - 2nd half {2} 42.75 19.93 39.86 -
    BART - 1st half {1} 42.79 20.25 39.66 -
    BART - 1st half {2} 40.70 18.99 37.88 -
    BART - 2nd half {1} 42.93 20.36 39.73 -
    BART - 2nd half {2} 41.93 19.79 39.06 -
    PEGASUS - 1st half + SR {1} 44.02 20.97 40.68 5.23
    PEGASUS - 1st half + SR {2} 45.66 21.31 42.51 7.61
    PEGASUS - 2nd half + SR {1} 44.11 21.08 40.82 4.57
    PEGASUS - 2nd half + SR {2} 45.73 21.31 42.62 6.94
    BART - 1st half + SR {1} 44.23 21.23 41.09 3.94
    BART - 1st half + SR {2} 45.05 21.47 42.12 11.65
    BART - 2nd half + SR {1} 44.51 21.52 41.29 4.44
    BART - 2nd half + SR {2} 45.61 21.78 42.62 9.32
    PEGASUS - 1st half + SR {1, 2} 46.12 21.97 42.84 9.36
    PEGASUS - 2nd half + SR {1, 2} 46.19 22.02 42.92 8.70
    BART - 1st half + SR {1, 2} 45.76 22.14 42.71 7.99
    BART - 2nd half + SR {1, 2} 45.96 22.18 42.88 7.98

    When scoring 30 candidates generated from both beam search and diverse beam search {1,2}\{1, 2\}, SummaReranker boosts the performance of models trained on only 50% of the training set to 46.19 R-1 (PEGASUS) and 45.96 R-1 (BART). These numbers exceed the baseline performance of standard PEGASUS and BART models trained on 100% of the training set (44.16 R-1 each) as well as the full-data single-stage model GSum (45.94 R-1).

  7. Knowl 7 — Best Candidate Recall at Threshold $k$ with Discrete Score Ties

    equation

    When evaluating how effectively a re-ranker selects the best candidate from a candidate pool C\mathcal{C} of size mm, discrete evaluation metrics (such as ROUGE) can assign identical maximum scores to multiple candidates. Let mm denote the total candidate pool size, k∈{1,…,m}k \in \{1, \ldots, m\} the rank threshold, and mbest∈{1,…,m}m_{\text{best}} \in \{1, \ldots, m\} the number of candidates that share the oracle maximum metric score. The random uniform ranking baseline recall R@k\text{R@}k (the probability that at least one of the mbestm_{\text{best}} optimal candidates is included in a randomly drawn subset of size kk without replacement) is:

    R@k=(mmbest)−(m−kmbest)(mmbest)\text{R@}k = \frac{\binom{m}{m_{\text{best}}} - \binom{m - k}{m_{\text{best}}}}{\binom{m}{m_{\text{best}}}}

    where (ab)=0\binom{a}{b} = 0 when a<ba < b. When exactly one candidate attains the maximum score (mbest=1m_{\text{best}} = 1), this reduces to the conventional formula R@k=km\text{R@}k = \frac{k}{m}.

  8. Knowl 8 — Best Candidate Recall and Candidate Overlap Improvements

    empirical result

    Evaluating SummaReranker's ability to rank the oracle candidate within the top kk positions out of m=15m = 15 diverse beam search candidates with PEGASUS demonstrates substantial gains over both the base generation log-probability ranking and the random baseline across thresholds k∈{1,…,5}k \in \{1, \ldots, 5\}:

    • CNN/DailyMail: SummaReranker achieves Recall@1 of 14.97% (vs. 8.57% base PEGASUS, 6.75% random) and Recall@5 of 50.84% (vs. 35.94% base PEGASUS, 33.60% random), an absolute improvement of +14.90% at k=5k=5.
    • XSum: Recall@1 reaches 16.57% (vs. 14.60% base PEGASUS, 8.05% random) and Recall@5 reaches 56.71% (vs. 47.17% base PEGASUS, 37.72% random), an absolute gain of +9.54% at k=5k=5.
    • Reddit TIFU: Recall@1 reaches 16.70% (vs. 14.54% base PEGASUS, 11.39% random) and Recall@5 reaches 53.34% (vs. 48.11% base PEGASUS, 46.70% random), an absolute gain of +5.23% at k=5k=5.

    In addition, SummaReranker retains the base model's top candidate in only 2.75%–22.23% of instances, while identifying and selecting an oracle-optimal candidate in up to 32.88% of cases.

  9. Knowl 9 — Abstractiveness Analysis of Re-ranked Candidate Summaries

    empirical result

    Abstractiveness measured by the percentage of novel nn-grams (n∈{1,2,3,4}n \in \{1, 2, 3, 4\}) not present in the source text shows that SummaReranker systematically selects more abstractive candidates than the base generation model without modifying the underlying sequence-to-sequence architecture or decoding algorithm:

    • On CNN/DailyMail, where base PEGASUS generation under beam search produces summaries with ~16% novel 2-grams, ~24% novel 3-grams, and ~33% novel 4-grams, applying SummaReranker increases novel 2-grams to ~20%, 3-grams to ~29%, and 4-grams to ~36%. When applied to diverse beam search, novel nn-grams increase from ~27% to ~39% (2-grams), ~39% to ~51% (3-grams), and ~47% to ~60% (4-grams), narrowing the abstractiveness gap relative to reference summaries (~49% 2-grams, ~68% 3-grams, ~80% 4-grams).
    • On Reddit TIFU, re-ranking increases novel nn-grams across all n∈{1,2,3,4}n \in \{1, 2, 3, 4\} for both beam search and diverse beam search.
    • On XSum, where candidate summaries are already highly abstractive (~27% novel 1-grams, ~72% novel 2-grams), SummaReranker yields slight increases in higher-order novel nn-grams (n∈{2,3,4}n \in \{2, 3, 4\}).
  10. Knowl 10 — Human Evaluation of Faithfulness on Selected Summaries

    empirical result

    A blind human evaluation was conducted on 50 randomly sampled test documents per dataset, where three graduate student raters with professional English proficiency compared top beam search summaries from base PEGASUS-large against the corresponding candidate chosen by SummaReranker for factual consistency and faithfulness to the source document:

    • CNN/DailyMail: SummaReranker candidate was preferred in 49.33% of comparisons (std = 12.20%), base PEGASUS in 32.00% (std = 6.00%), and 18.67% were ties (std = 9.50%).
    • XSum: SummaReranker candidate was preferred in 30.00% of comparisons (std = 7.12%), base PEGASUS in 28.00% (std = 10.20%), and 42.00% were ties (std = 16.33%).
    • Reddit TIFU: SummaReranker candidate was preferred in 58.00% of comparisons (std = 4.32%), base PEGASUS in 28.00% (std = 2.82%), and 16.00% were ties (std = 4.32%).

    Across all domains, human annotators preferred the candidate selected by SummaReranker more often than the base model's default top beam summary.

  11. Knowl 11 — Inference Latency and Candidate Sampling Size Trade-off

    empirical result

    SummaReranker inference time scales linearly with candidate pool size O(m)\mathcal{O}(m), requiring approximately 38 ms per candidate on an NVIDIA RTX 2080 Ti GPU. Evaluating performance across randomly sub-sampled candidate pools k∈{1,…,15}k \in \{1, \ldots, 15\} indicates:

    • On CNN/DailyMail, re-ranking as few as k=2k = 2 candidates outperforms the single-stage PEGASUS beam search baseline.
    • On XSum, re-ranking k∈[3,8]k \in [3, 8] candidates is required to exceed the base model.
    • On Reddit TIFU, k∈[3,4]k \in [3, 4] candidates are required to surpass the baseline.
    • Scoring pools beyond 30 candidates (e.g., adding top-kk and top-pp sampling to beam search and diverse beam search to reach 60 candidates) leads to performance saturation (CNN/DailyMail R-1 is 47.16 with 30 candidates vs. 47.04 with 60 candidates), despite higher theoretical oracle scores for 60 candidates. Evaluating 6 to 8 candidates provides an effective empirical trade-off between inference throughput and summarization performance.
  12. Knowl 12 — Context Truncation and Fixed Multi-Task Loss Weighting Limitations

    limitation

    SummaReranker is subject to two main structural limitations:

    1. Context Window Constraint: SummaReranker jointly encodes the full source document concatenated with the candidate summary within a single sequence [CLS] Source [SEP] Candidate. Because RoBERTa-large is bounded by a maximum positional encoding limit of 512 tokens, long source documents must be truncated, restricting the framework's direct application to long-document summarization domains (e.g., scientific papers or legal texts) without incorporating extended context architectures.
    2. Equal Multi-Task Loss Weighting: The training objective weights all metric loss heads equally with fixed coefficients 1N\frac{1}{N}. When jointly optimizing over lexical overlap metrics (ROUGE-1, ROUGE-2, ROUGE-L) and embedding-based model metrics (BERTScore, BARTScore), joint 5-task training yields lower ROUGE scores (46.59 R-1 on CNN/DM) than training solely on ROUGE (47.16 R-1), demonstrating task interference and the need for adaptive or Pareto multi-task loss weighting.

Coverage note — Detailed qualitative re-ranking text demonstrations (full summary examples in Appendix K) and base model fine-tuning hyperparameter listings (Tables 7 and 8) were omitted as they represent illustrative examples and standard configuration details rather than standalone conceptual contributions.

References

  1. 1.Armen Aghajanyan, Akshat Shrivastava, Anchit Gupta, Naman Goyal, Luke Zettlemoyer, and Sonal Gupta. 2020. Better fine-tuning by reducing representational collapse. arXiv preprint arXiv:2008.03156.
  2. 2.Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer. 2015. Scheduled sampling for sequence prediction with recurrent neural networks. In Proceedings of the 28th International Conference on Neural Information Processing Systems - Volume 1, NIPS'15, page 1171–1179, Cambridge, MA, USA. MIT Press.
  3. 3.Sumanta Bhattacharyya, Amirmohammad Rooshenas, Subhajit Naskar, Simeng Sun, Mohit Iyyer, and Andrew McCallum. 2021. Energy-based reranking: Improving neural machine translation using energy-based models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 4528–4537, Online. Association for Computational Linguistics.
  4. 4.Eugene Charniak and Mark Johnson. 2005. Coarse-to-fine n-best parsing and maxent discriminative reranking. In Proceedings of the 43rd Annual Meeting of the Association for Computational Linguistics (ACL'05), pages 173–180.
  5. 5.Arman Cohan, Franck Dernoncourt, Doo Soon Kim, Trung Bui, Seokhwan Kim, Walter Chang, and Nazli Goharian. 2018. A discourse-aware attention model for abstractive summarization of long documents. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 615–621, New Orleans, Louisiana. Association for Computational Linguistics.
  6. 6.Michael Collins and Terry Koo. 2005. Discriminative reranking for natural language parsing. Computational Linguistics, 31(1):25–70.
  7. 7.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  8. 8.Zi-Yi Dou, Pengfei Liu, Hiroaki Hayashi, Zhengbao Jiang, and Graham Neubig. 2021. GSum: A general framework for guided neural abstractive summarization. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4830–4842, Online. Association for Computational Linguistics.
  9. 9.Alexander R Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2021. Summeval: Re-evaluating summarization evaluation. Transactions of the Association for Computational Linguistics, 9:391–409.
  10. 10.Angela Fan, Mike Lewis, and Yann Dauphin. 2018. Hierarchical neural story generation. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 889–898, Melbourne, Australia. Association for Computational Linguistics.
  11. 11.Xavier Glorot, Antoine Bordes, and Yoshua Bengio. 2011. Deep sparse rectifier neural networks. In Proceedings of the fourteenth international conference on artificial intelligence and statistics, pages 315–323. JMLR Workshop and Conference Proceedings.
  12. 12.Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015. Teaching machines to read and comprehend. Advances in neural information processing systems, 28:1693–1701.
  13. 13.Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2019. The curious case of neural text degeneration. arXiv preprint arXiv:1904.09751.
  14. 14.Srinivasan Iyer, Sewon Min, Yashar Mehdad, and Wentau Yih. 2021. RECONSIDER: Improved re-ranking using span-focused cross-attention for open domain question answering. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1280–1287, Online. Association for Computational Linguistics.
  15. 15.Byeongchang Kim, Hyunwoo Kim, and Gunhee Kim. 2019. Abstractive summarization of Reddit posts with multi-level memory networks. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 2519–2531, Minneapolis, Minnesota. Association for Computational Linguistics.
  16. 16.Fajri Koto, Jey Han Lau, and Timothy Baldwin. 2021. Evaluating the efficacy of summarization evaluation across languages. arXiv preprint arXiv:2106.01478.
  17. 17.Bernhard Kratzwald, Anna Eigenmann, and Stefan Feuerriegel. 2019. RankQA: Neural question answering with answer re-ranking. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 6076–6085, Florence, Italy. Association for Computational Linguistics.
  18. 18.Bernhard Kratzwald and Stefan Feuerriegel. 2018. Adaptive document retrieval for deep question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 576–581, Brussels, Belgium. Association for Computational Linguistics.
  19. 19.Wojciech Kryscinski, Nitish Shirish Keskar, Bryan McCann, Caiming Xiong, and Richard Socher. 2019. Neural text summarization: A critical evaluation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 540–551, Hong Kong, China. Association for Computational Linguistics.
  20. 20.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online. Association for Computational Linguistics.
  21. 21.Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74–81, Barcelona, Spain. Association for Computational Linguistics.
  22. 22.Chin-Yew Lin and Eduard Hovy. 2003. Automatic evaluation of summaries using n-gram co-occurrence statistics. In Proceedings of the 2003 Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics, pages 150–157.
  23. 23.Xi Lin, Hui-Ling Zhen, Zhenhua Li, Qing-Fu Zhang, and Sam Kwong. 2019. Pareto multi-task learning. Advances in neural information processing systems, 32:12060–12070.
  24. 24.Xiang Lin, Simeng Han, and Shafiq Joty. 2021. Straight to the gradient: Learning to use novel tokens for neural text generation. In Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 6642–6653. PMLR.
  25. 25.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
  26. 26.Yixin Liu, Zi-Yi Dou, and Pengfei Liu. 2021. RefSum: Refactoring neural summarization. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1437–1448, Online. Association for Computational Linguistics.
  27. 27.Yixin Liu and Pengfei Liu. 2021. SimCLS: A simple framework for contrastive learning of abstractive summarization. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pages 1065–1072, Online. Association for Computational Linguistics.
  28. 28.Ramesh Nallapati. 2004. Discriminative models for information retrieval. In Proceedings of the 27th annual international ACM SIGIR conference on Research and development in information retrieval, pages 64–71.
  29. 29.Ramesh Nallapati, Bowen Zhou, Cicero dos Santos, Çağlar Gülçehre, and Bing Xiang. 2016. Abstractive text summarization using sequence-to-sequence RNNs and beyond. In Proceedings of The 20th SIGNLL Conference on Computational Natural Language Learning, pages 280–290, Berlin, Germany. Association for Computational Linguistics.
  30. 30.Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018. Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 1797–1807, Brussels, Belgium. Association for Computational Linguistics.
  31. 31.Rodrigo Nogueira and Kyunghyun Cho. 2019. Passage re-ranking with bert. arXiv preprint arXiv:1901.04085.
  32. 32.Vinay Pandramish and Dipti Misra Sharma. 2020. Checkpoint reranking: An approach to select better hypothesis for neural machine translation systems. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: Student Research Workshop, pages 286–291, Online. Association for Computational Linguistics.
  33. 33.Weizhen Qi, Yeyun Gong, Yu Yan, Can Xu, Bolun Yao, Bartuer Zhou, Biao Cheng, Daxin Jiang, Jiusheng Chen, Ruofei Zhang, Houqiang Li, and Nan Duan. 2021. ProphetNet-X: Large-scale pre-training models for English, Chinese, multi-lingual, dialog, and code generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing: System Demonstrations, pages 232–239, Online. Association for Computational Linguistics.
  34. 34.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2019. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv:1910.10683.
  35. 35.Abigail See, Peter J. Liu, and Christopher D. Manning. 2017. Get to the point: Summarization with pointer-generator networks. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1073–1083, Vancouver, Canada. Association for Computational Linguistics.
  36. 36.Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. 2017. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. arXiv preprint arXiv:1701.06538.
  37. 37.Noam Shazeer and Mitchell Stern. 2018. Adafactor: Adaptive learning rates with sublinear memory cost. In International Conference on Machine Learning, pages 4596–4604. PMLR.
  38. 38.Shichao Sun and Wenjie Li. 2021. Alleviating exposure bias via contrastive learning for abstractive text summarization. arXiv preprint arXiv:2108.11846.
  39. 39.Ashwin K Vijayakumar, Michael Cogswell, Ramprasath R Selvaraju, Qing Sun, Stefan Lee, David Crandall, and Dhruv Batra. 2016. Diverse beam search: Decoding diverse solutions from neural sequence models. arXiv preprint arXiv:1610.02424.
  40. 40.Ronald J Williams and David Zipser. 1989. A learning algorithm for continually running fully recurrent neural networks. Neural computation, 1(2):270–280.
  41. 41.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  42. 42.Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021. Bartscore: Evaluating generated text as text generation. arXiv preprint arXiv:2106.11520.
  43. 43.Jingqing Zhang, Yao Zhao, Mohammad Saleh, and Peter Liu. 2020. Pegasus: Pre-training with extracted gap-sentences for abstractive summarization. In International Conference on Machine Learning, pages 11328–11339. PMLR.
  44. 44.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019a. Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675.
  45. 45.Zhengyan Zhang, Xu Han, Zhiyuan Liu, Xin Jiang, Maosong Sun, and Qun Liu. 2019b. ERNIE: Enhanced language representation with informative entities. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 1441–1451, Florence, Italy. Association for Computational Linguistics.
  46. 46.Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, and Steffen Eger. 2019a. MoverScore: Text generation evaluating with contextualized embeddings and earth mover distance. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 563–578, Hong Kong, China. Association for Computational Linguistics.
  47. 47.Zhe Zhao, Lichan Hong, Li Wei, Jilin Chen, Aniruddh Nath, Shawn Andrews, Aditee Kumthekar, Maheswaran Sathiamoorthy, Xinyang Yi, and Ed Chi. 2019b. Recommending what video to watch next: a multitask ranking system. In Proceedings of the 13th ACM Conference on Recommender Systems, pages 43–51.

Citation

MLA
Ravaut, M., et al. “SummaReranker: A Multi-Task Mixture-of-Experts Re-ranking Framework for Abstractive Summarization”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 4504–24, https://doi.org/10.18653/v1/2022.acl-long.309.
APA
Ravaut, M., Joty, S., & Chen, N. (2022). SummaReranker: A Multi-Task Mixture-of-Experts Re-ranking Framework for Abstractive Summarization. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 4504–4524. https://doi.org/10.18653/v1/2022.acl-long.309
Chicago
Ravaut, M., S. Joty, and N. Chen. 2022. “SummaReranker: A Multi-Task Mixture-of-Experts Re-ranking Framework for Abstractive Summarization”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 4504–24. https://doi.org/10.18653/v1/2022.acl-long.309.
Harvard
Ravaut, M., Joty, S. and Chen, N. (2022) “SummaReranker: A Multi-Task Mixture-of-Experts Re-ranking Framework for Abstractive Summarization”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 4504–4524. Available at: https://doi.org/10.18653/v1/2022.acl-long.309.
Vancouver
1. Ravaut M, Joty S, Chen N (2022) SummaReranker: A Multi-Task Mixture-of-Experts Re-ranking Framework for Abstractive Summarization. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 4504–4524

BibTeX

@inproceedings{ravaut-etal-2022-summareranker,
    title = "{S}umma{R}eranker: A Multi-Task Mixture-of-Experts Re-ranking Framework for Abstractive Summarization",
    author = "Ravaut, Mathieu  and
      Joty, Shafiq  and
      Chen, Nancy",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.309/",
    doi = "10.18653/v1/2022.acl-long.309",
    pages = "4504--4524"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/