HEGEL: Hypergraph Transformer for Long Document Summarization

Haopeng ZhangXiao LiuJiawei Zhang

article2022EMNLP59 citations

Proposes a hypergraph transformer architecture that captures high-order cross-sentence dependencies across section structures, latent topics, and keyword coreferences to improve extractive summarization for long documents.

Listen

Automated text summarization systems perform reliably on short news articles, but they struggle with complex, long-form documents such as scientific research papers. Long texts contain thousands of words across diverse topics and intricate section structures, leading to severe computational bottlenecks and an inability to track relationships between sentences spaced far apart. Standard graph-based solutions attempt to address this by linking pairs of sentences, yet they fail to capture multi-sentence relationships or integrate diverse structural and semantic contexts simultaneously.

The article develops and evaluates HEGEL (Hypergraph transformer for Extractive Long document summarization), a new neural network architecture designed to model complex, multi-sentence dependencies and extract the most informative sentences from long documents.

To overcome the limitations of pairwise graph models, the approach represents documents as hypergraphs—mathematical structures where a single connection, or hyperedge, can link multiple sentences at once. HEGEL integrates three distinct relation types: document section structure (local context), latent topics (global semantic themes), and shared keywords (coreference). Specialized hypergraph transformer attention layers then propagate information across these connections to score sentence importance. The article evaluates this approach across two large-scale scientific benchmarks, PubMed (over 112,000 training papers) and arXiv (over 201,000 training papers), comparing performance against standard extractive and abstractive baselines.

HEGEL consistently outperformed all unsupervised, extractive, and abstractive baseline models on both benchmark datasets in standard overlap and fluency metrics. Analysis shows that local section structure is the single most critical dependency; removing section hyperedges caused the largest performance drop, and the model allocates more than half of its attention to section-level connections. In addition, incorporating global topic and keyword connections, alongside hierarchical sentence positioning, measurably enhanced the accuracy of salient sentence selection. Crucially, the model achieved these gains while maintaining high parameter efficiency, using approximately 50% fewer parameters than standard heterogeneous graph models and 90% fewer parameters than large-scale long-document language models.

These findings demonstrate that capturing multi-sentence relationships from multiple viewpoints is more effective for summarizing long documents than relying solely on pairwise graphs or massive language models. From an operational perspective, the high parameter efficiency translates to lower computational costs, reduced infrastructure requirements, and faster processing timelines without sacrificing summarization quality.

Organizations handling high volumes of structured, long-form technical or scientific documents should consider adopting hypergraph-based architectures over traditional pairwise graph or heavy transformer models. Technical teams looking to extend this approach can incorporate additional relation types, such as syntactic dependencies or domain-specific metadata, to further refine sentence selection.

Confidence in the reported results is high for academic literature, as the evaluations are grounded in extensive testing on established public benchmarks. However, leaders should note two main limitations before deployment in broader production settings: the model relies on external keyword and topic extraction tools during data pre-processing, and evaluation has so far been restricted to scientific papers. Further pilot testing is advisable when applying the framework to other document types, such as legal or financial filings.

arXiv: 2210.04126
Cover for HEGEL: Hypergraph Transformer for Long Document Summarization

Abstract

Extractive summarization for long documents is challenging due to the extended structured input context. The long-distance sentence dependency hinders cross-sentence relations modeling, the critical step of extractive summarization. This paper proposes HEGEL, a hypergraph neural network for long document summarization by capturing high-order cross-sentence relations. HEGEL updates and learns effective sentence representations with hypergraph transformer layers and fuses different types of sentence dependencies, including latent topics, keywords coreference, and section structure. We validate HEGEL by conducting extensive experiments on two benchmark datasets, and experimental results demonstrate the effectiveness and efficiency of HEGEL.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 2.1 Scientific Paper Summarization
  • 2.2 Graph based summarization
  • 3 Method
  • 3.1 Document as a Hypergraph
  • 3.1.1 Node Representation
  • 3.1.2 Hyperedge Construction
  • 3.2 Hypergraph Transformer Layer
  • 3.2.1 Hypergraph Attention
  • 3.2.2 Hypergraph Transformer
  • 3.3 Training Objective
  • 4 Experiment
  • 4.1 Experiment Setup
  • 4.2 Implementation Details
  • 4.3 Experiment Results
  • 5 Analysis
  • 5.1 Ablation Study
  • 5.2 Hyperedge Analysis
  • 5.3 Embedding Analysis
  • 5.4 Case Study
  • 6 Conclusion
  • Limitations
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — HEGEL hypergraph model for long-document extraction

    model/method

    HEGEL (Hypergraph Transformer for Extractive Long-document summarization) converts a document into a sentence-level hypergraph and predicts the salience of every sentence. Each sentence is a node, while a hyperedge can connect more than two sentences that share section structure, latent topic, or keyword information. The model initializes sentence nodes with fixed Sentence-BERT embeddings plus hierarchical positional information, applies stacked hypergraph Transformer layers to propagate information between nodes and hyperedges, and uses a multilayer perceptron with a sigmoid output to estimate whether each sentence should be selected for an extractive summary. Unlike an ordinary graph, the hypergraph directly represents triadic and higher-order sentence relations and combines local structural context with global semantic context.

  2. Knowl 2 — Formal hypergraph and hierarchical sentence representation

    definition

    For a document D={s1,…,sn}D=\{s_1,\ldots,s_n\} with nn sentences, HEGEL defines a hypergraph G=(V,E)G=(V,E), where V={v1,…,vn}V=\{v_1,\ldots,v_n\} contains one node viv_i for each sentence sis_i, and E={e1,…,em}E=\{e_1,\ldots,e_m\} contains mm hyperedges. Each hyperedge satisfies ej⊆Ve_j\subseteq V and ∣ej∣≥2|e_j|\geq 2. Its incidence matrix A∈{0,1}n×mA\in\{0,1\}^{n\times m} is

    Aij={1,vi∈ej,0,vi∉ej.A_{ij}=\begin{cases} 1,&v_i\in e_j,\\ 0,&v_i\notin e_j. \end{cases}

    The fixed Sentence-BERT embedding of sentence sis_i is xi∈Rdmodelx_i\in\mathbb{R}^{d_{\mathrm{model}}}. If pisecp_i^{\mathrm{sec}} is its section index and pisenp_i^{\mathrm{sen}} is its index within that section, HEGEL adds hierarchical positional encoding

    HPE⁡(si)=γ1PE⁡(pisec)+γ2PE⁡(pisen),\operatorname{HPE}(s_i)=\gamma_1\operatorname{PE}(p_i^{\mathrm{sec}})+\gamma_2\operatorname{PE}(p_i^{\mathrm{sen}}),

    where γ1\gamma_1 and γ2\gamma_2 are scalar rescaling weights. For position pos⁡\operatorname{pos} and coordinate index rr, the sinusoidal encoding is

    PE⁡(pos⁡,2r)=sin⁡(pos⁡100002r/dmodel),PE⁡(pos⁡,2r+1)=cos⁡(pos⁡100002r/dmodel).\operatorname{PE}(\operatorname{pos},2r)=\sin\left(\frac{\operatorname{pos}}{10000^{2r/d_{\mathrm{model}}}}\right),\qquad \operatorname{PE}(\operatorname{pos},2r+1)=\cos\left(\frac{\operatorname{pos}}{10000^{2r/d_{\mathrm{model}}}}\right).

    The initial node representation is hi(0)=xi+HPE⁡(si)h_i^{(0)}=x_i+\operatorname{HPE}(s_i). This preserves both semantic sentence content and the document's section-level and within-section order.

  3. Knowl 3 — Section, topic, and keyword hyperedge construction

    model/method

    HEGEL constructs three incidence submatrices and concatenates them into one hypergraph structure. If the document has qq sections, each section becomes one section hyperedge containing all sentences in that section:

    Aijsec={1,si belongs to section j,0,otherwise.A^{\mathrm{sec}}_{ij}=\begin{cases}1,&s_i\text{ belongs to section }j,\\0,&\text{otherwise.}\end{cases}

    Latent Dirichlet Allocation produces pp latent topics. Each topic becomes a topic hyperedge containing the sentences assigned to that topic:

    Aijtopic={1,si belongs to latent topic j,0,otherwise.A^{\mathrm{topic}}_{ij}=\begin{cases}1,&s_i\text{ belongs to latent topic }j,\\0,&\text{otherwise.}\end{cases}

    KeyBERT extracts kk document keywords. Each keyword becomes a keyword hyperedge containing every sentence that contains the keyword, regardless of sentence distance:

    Aijkw={1,si contains keyword j,0,otherwise.A^{\mathrm{kw}}_{ij}=\begin{cases}1,&s_i\text{ contains keyword }j,\\0,&\text{otherwise.}\end{cases}

    The complete incidence matrix is

    A=Asec∥Atopic∥Akw,m=q+p+k,A=A^{\mathrm{sec}}\Vert A^{\mathrm{topic}}\Vert A^{\mathrm{kw}}, \qquad m=q+p+k,

    where ∥\Vert denotes column-wise concatenation. Section hyperedges provide local discourse connectivity and ensure that every sentence is connected to its section, whereas topic and keyword hyperedges provide long-distance global connections. The resulting hypergraph therefore fuses structural, topical, and lexical dependencies rather than relying on one sentence-relation type.

  4. Knowl 4 — Two-phase multi-head hypergraph Transformer

    model/method

    A HEGEL hypergraph Transformer layer performs attention in two phases: it first aggregates sentence nodes into hyperedge representations, then sends selected hyperedge information back to the sentence nodes. Let H(l−1)={hi(l−1)}i=1nH^{(l-1)}=\{h_i^{(l-1)}\}_{i=1}^n be the node representations entering layer ll, let AA be the incidence matrix, and let WhW_h, WeW_e, wahw_{ah}, and waew_{ae} be trainable matrices or vectors.

    For hyperedge eje_j, first define uk=LeakyReLU⁡(Whhk(l−1))u_k=\operatorname{LeakyReLU}(W_hh_k^{(l-1)}). The node attention inside that hyperedge and its updated representation are

    αjk=exp⁡(wahTuk)∑vp∈ejexp⁡(wahTup),gj(l)=LeakyReLU⁡(∑vk∈ejαjkWhhk(l−1)).\alpha_{jk}=\frac{\exp(w_{ah}^{\mathsf T}u_k)}{\sum_{v_p\in e_j}\exp(w_{ah}^{\mathsf T}u_p)}, \qquad g_j^{(l)}=\operatorname{LeakyReLU}\left(\sum_{v_k\in e_j}\alpha_{jk}W_hh_k^{(l-1)}\right).

    For a node viv_i connected to hyperedge eke_k, define

    zki=LeakyReLU⁡([Wegk(l)∥Whhi(l−1)]),z_{ki}=\operatorname{LeakyReLU}\left([W_eg_k^{(l)}\Vert W_hh_i^{(l-1)}]\right),

    where [⋅∥⋅][\cdot\Vert\cdot] is vector concatenation. The hyperedge attention for node viv_i and the updated node representation are

    βki=exp⁡(waeTzki)∑eq∋viexp⁡(waeTzqi),hi(l)=LeakyReLU⁡(∑ek∋viβkiWegk(l)).\beta_{ki}=\frac{\exp(w_{ae}^{\mathsf T}z_{ki})}{\sum_{e_q\ni v_i}\exp(w_{ae}^{\mathsf T}z_{qi})}, \qquad h_i^{(l)}=\operatorname{LeakyReLU}\left(\sum_{e_k\ni v_i}\beta_{ki}W_eg_k^{(l)}\right).

    HEGEL uses multiple independent attention heads, concatenates their outputs, and projects them with an output matrix. The resulting multi-head hypergraph attention is combined with the incoming node representation by residual connections and layer normalization, followed by a position-wise feed-forward network and a second residual-plus-normalization operation:

    H′(l)=LN⁡(MH-HGA⁡(H(l−1),A)+H(l−1)),H'^{(l)}=\operatorname{LN}(\operatorname{MH\text{-}HGA}(H^{(l-1)},A)+H^{(l-1)}), H(l)=LN⁡(FFN⁡(H′(l))+H′(l)).H^{(l)}=\operatorname{LN}(\operatorname{FFN}(H'^{(l)})+H'^{(l)}).

    This node-to-hyperedge-to-node message passing allows each sentence to combine information from all sentences sharing any of its section, topic, or keyword hyperedges.

  5. Knowl 5 — Sentence salience prediction and training objective

    equation

    After LL hypergraph Transformer layers, HEGEL predicts a confidence score for each sentence sis_i from its final representation hi(L)h_i^{(L)}. With trainable projection matrices Wp1W_{p1} and Wp2W_{p2}, the prediction is

    zi=LeakyReLU⁡(Wp1hi(L)),y^i=sigmoid⁡(Wp2zi),z_i=\operatorname{LeakyReLU}(W_{p1}h_i^{(L)}), \qquad \hat y_i=\operatorname{sigmoid}(W_{p2}z_i),

    where y^i∈(0,1)\hat y_i\in(0,1) is the predicted probability that sentence sis_i belongs in the extractive summary. Given a binary ground-truth label yi∈{0,1}y_i\in\{0,1\}, the model is optimized with document- and sentence-averaged binary cross-entropy:

    L=−1N∑d=1N1Nd∑i=1Nd[yilog⁡y^i+(1−yi)log⁡(1−y^i)],\mathcal{L}=-\frac{1}{N}\sum_{d=1}^{N}\frac{1}{N_d}\sum_{i=1}^{N_d}\left[y_i\log\hat y_i+(1-y_i)\log(1-\hat y_i)\right],

    where NN is the number of training documents and NdN_d is the number of sentences in document dd. At inference time, sentences are ranked or selected according to their predicted confidence scores. Extractive labels are constructed by greedily optimizing ROUGE against each document's gold abstract.

  6. Knowl 6 — Datasets and HEGEL training configuration

    experimental setup

    HEGEL is evaluated on the original train, validation, and test splits of the ArXiv and PubMed scientific-paper summarization datasets. PubMed contains biomedical papers, whereas ArXiv contains papers from multiple scientific fields. The dataset sizes and average lengths are:

    ArXiv PubMed
    # train 201,427 112,291
    # validation 6,431 6,402
    # test 6,436 6,449
    average document length 4,938 3,016
    average summary length 203 220

    The initial sentence encoder is the fixed all-mpnet-base-v2 Sentence-BERT checkpoint, with embedding dimension 768; the model input-layer dimension is 1024. HEGEL uses two hypergraph Transformer layers, eight attention heads per layer, hidden dimension 128, and output-layer hidden dimension 4096. At most 100 topics are generated per document. Topic and keyword hyperedges are discarded when they connect fewer than 5 or more than 25 sentence nodes. Both positional rescaling weights are set to γ1=γ2=0.001\gamma_1=\gamma_2=0.001.

    The model is trained with Adam at learning rate 0.00010.0001 and dropout rate 0.30.3 for at most 20 epochs on an RTX A6000 GPU. ROUGE-1 F-score on the validation set selects checkpoints, and early stopping uses patience 3. Evaluation reports ROUGE-1, ROUGE-2, and ROUGE-L F-scores; ROUGE-1 and ROUGE-2 measure informativeness, while ROUGE-L measures sequence overlap related to summary fluency.

  7. Knowl 7 — Benchmark comparison on PubMed and ArXiv

    data/table

    The benchmark compares unsupervised extractive systems, neural extractive systems, and abstractive systems using ROUGE F-scores. HEGEL obtains the highest ROUGE-1 and ROUGE-2 scores among the listed systems on both datasets. On PubMed it also obtains the highest ROUGE-L score. On ArXiv, HEGEL's ROUGE-L is 39.89, below HiStruct+'s 40.16, although its ROUGE-1 and ROUGE-2 scores are higher. The results show that the proposed hypergraph model is particularly effective for informativeness-oriented metrics and is competitive with both extractive and abstractive baselines.

    Model PubMed ArXiv
    R-1 R-2 R-L R-1 R-2 R-L
    ORACLE 55.05 27.48 49.11 53.88 23.05 46.54
    LEAD 35.63 12.28 25.17 33.66 8.94 22.19
    LexRank (2004) 39.19 13.89 34.59 33.85 10.73 28.99
    PACSUM (2019) 39.79 14.00 36.09 38.57 10.93 34.33
    HIPORANK (2021) 43.58 17.00 39.31 39.34 12.56 34.89
    ChengLapata (2016) 43.89 18.53 30.17 42.24 15.97 27.88
    SummaRuNNer (2016) 43.89 18.78 30.36 42.81 16.52 28.23
    ExtSum-LG (2019) 44.85 19.70 31.43 43.62 17.36 29.14
    SentCLF (2020) 45.01 19.91 41.16 34.01 8.71 30.41
    SentPTR (2020) 43.30 17.92 39.47 42.32 15.63 38.06
    ExtSum-LG + RdLoss (2021) 45.30 20.42 40.95 44.01 17.79 39.09
    ExtSum-LG + MMR (2021) 45.39 20.37 40.99 43.87 17.50 38.97
    HiStruct+ (2022) 46.59 20.39 42.11 45.22 17.67 40.16
    PGN (2017) 35.86 10.22 29.69 32.06 9.04 25.16
    DiscourseAware (2018) 38.93 15.37 35.21 35.80 11.05 31.80
    TLM-I+E (2020) 42.13 16.27 39.21 41.62 14.69 38.03
    DANCER-LSTM (2020) 44.09 17.69 40.27 41.87 15.92 37.61
    DANCER-RUM (2020) 43.98 17.65 40.25 42.70 16.54 38.44
    HEGEL 47.13 21.00 42.18 46.41 18.17 39.89

    On PubMed, HEGEL reaches 47.13/21.00/42.18 for ROUGE-1/2/L, compared with 46.59/20.39/42.11 for HiStruct+. On ArXiv, it reaches 46.41/18.17/39.89, compared with HiStruct+'s 45.22/17.67/40.16.

  8. Knowl 8 — Contribution of positional encoding and hyperedge types

    empirical result

    Ablation experiments on PubMed remove hierarchical position encoding or one of the three hyperedge types from HEGEL. Every removed component lowers performance relative to the complete model, demonstrating that sentence order and all three relation types contribute to summary selection.

    Model ROUGE-1 ROUGE-2 ROUGE-L
    full HEGEL 47.13 21.00 42.18
    without position 46.86 20.05 41.91
    without keyword hyperedges 46.92 20.71 42.03
    without topic hyperedges 46.35 20.30 41.48
    without section hyperedges 45.63 19.30 40.71

    Removing section hyperedges produces the largest performance deterioration. The paper attributes this to the loss of local section context and to increased hypergraph sparsity, which weakens information aggregation. Topic and keyword hyperedges still provide measurable gains by supplying global semantic and lexical links, while hierarchical positional encoding contributes sequential-order information.

  9. Knowl 9 — Attention, embedding, and qualitative analyses

    empirical result

    The learned hyperedge attention patterns on PubMed show that section hyperedges have the largest average degree and receive more than half of the attention assigned to the three hyperedge types. Topic hyperedges are the most numerous on average, while keyword hyperedges receive the least attention. This pattern is consistent with the ablation results: local section context is especially important for aggregating sentence information, even though topic and keyword links supply complementary global relations.

    For 100 PubMed test documents, the final sentence embeddings are projected to two dimensions with t-SNE. Ground-truth summary sentences form visible clusters, including a concentration in the lower-left region of the visualization, whereas non-ground-truth sentences are more dispersed. This provides qualitative evidence that the final node representations separate salient from non-salient sentences.

    A qualitative output contains five selected sentences drawn from sections 1 through 4 of a paper about MERS-CoV infections in dromedaries imported from Oman to the United Arab Emirates. The selected sentences cover the genomic-analysis method, asymptomatic human infections, genomic similarity to viruses associated with hospital-acquired infections in Saudi Arabia and South Korea, and evidence that related viruses circulate on the Arabian Peninsula. Although these sentences are far apart in the source document, they are connected through shared latent-topic and keyword relations, illustrating how multi-node hyperedges can support long-distance, higher-order selection.

  10. Knowl 10 — Limitations and parameter-efficiency claim

    limitation

    HEGEL depends on external preprocessing models—LDA for latent topics and KeyBERT for keywords—to construct part of its hypergraph, so errors or domain mismatch in those models can affect the sentence relations available to HEGEL. The evaluation is also limited to academic-paper datasets, specifically ArXiv and PubMed, and does not establish performance on other kinds of long documents.

    The paper reports that these preprocessing operations do not increase the model's computation complexity because they are performed before the end-to-end neural model. It also reports that HEGEL uses 50% fewer parameters than the heterogeneous graph model of Wang et al. (2020) and 90% fewer parameters than Longformer-base, attributing this parameter efficiency to using hyperedge-based cross-sentence attention.

Coverage note — No substantial contributed material was omitted; minor implementation and baseline-description details were consolidated into the experimental setup and benchmark-result knowls.

References

  1. 1.Iz Beltagy, Matthew E Peters, and Arman Cohan. 2020. Longformer: The long-document transformer. arXiv preprint arXiv:2004.05150.
  2. 2.David M Blei, Andrew Y Ng, and Michael I Jordan. 2003. Latent dirichlet allocation. Journal of machine Learning research, 3(Jan):993–1022.
  3. 3.Jianpeng Cheng and Mirella Lapata. 2016. Neural summarization by extracting sentences and words. arXiv preprint arXiv:1603.07252.
  4. 4.Arman Cohan, Franck Dernoncourt, Doo Soon Kim, Trung Bui, Seokhwan Kim, Walter Chang, and Nazli Goharian. 2018. A discourse-aware attention model for abstractive summarization of long documents. arXiv preprint arXiv:1804.05685.
  5. 5.Peng Cui and Le Hu. 2021. Topic-guided abstractive multi-document summarization. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 1463–1472, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  6. 6.Peng Cui, Le Hu, and Yuanchao Liu. 2020. Enhancing extractive text summarization with topic-aware graph neural networks. arXiv preprint arXiv:2010.06253.
  7. 7.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805.
  8. 8.Kaize Ding, Jianling Wang, Jundong Li, Dingcheng Li, and Huan Liu. 2020. Be more with less: Hypergraph attention networks for inductive text classification. arXiv preprint arXiv:2011.00387.
  9. 9.Yue Dong, Andrei Mircea, and Jackie CK Cheung. 2020. Discourse-aware unsupervised summarization of long scientific documents. arXiv preprint arXiv:2005.00513.
  10. 10.Günes Erkan and Dragomir R Radev. 2004. Lexrank: Graph-based lexical centrality as salience in text summarization. Journal of artificial intelligence research, 22:457–479.
  11. 11.Alexios Gidiotis, Stefanos Stefanidis, and Grigorios Tsoumakas. 2020. AUTH @ CLSciSumm 20, LaySumm 20, LongSumm 20. In Proceedings of the First Workshop on Scholarly Document Processing, pages 251–260, Online. Association for Computational Linguistics.
  12. 12.Maarten Grootendorst. 2020. Keybert: Minimal keyword extraction with bert.
  13. 13.Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015. Teaching machines to read and comprehend. In Advances in neural information processing systems, pages 1693–1701.
  14. 14.Baoyu Jing, Zeyu You, Tao Yang, Wei Fan, and Hanghang Tong. 2021. Multiplex graph neural network for extractive text summarization. arXiv preprint arXiv:2108.12870.
  15. 15.Jiaxin Ju, Ming Liu, Huan Yee Koh, Yuan Jin, Lan Du, and Shirui Pan. 2021. Leveraging information bottleneck for scientific document summarization. arXiv preprint arXiv:2110.01280.
  16. 16.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. 2019. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. arXiv preprint arXiv:1910.13461.
  17. 17.Haoran Li, Junnan Zhu, Jiajun Zhang, Chengqing Zong, and Xiaodong He. 2020. Keywords-guided abstractive sentence summarization. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 8196–8203.
  18. 18.Chin-Yew Lin and Eduard Hovy. 2003. Automatic evaluation of summaries using n-gram co-occurrence statistics. In Proceedings of the 2003 Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics, pages 150–157.
  19. 19.Yang Liu and Mirella Lapata. 2019. Text summarization with pretrained encoders. arXiv preprint arXiv:1908.08345.
  20. 20.Ilya Loshchilov and Frank Hutter. 2017. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101.
  21. 21.Ramesh Nallapati, Feifei Zhai, and Bowen Zhou. 2016a. Summarunner: A recurrent neural network based sequence model for extractive summarization of documents. arXiv preprint arXiv:1611.04230.
  22. 22.Ramesh Nallapati, Bowen Zhou, Caglar Gulcehre, Bing Xiang, et al. 2016b. Abstractive text summarization using sequence-to-sequence rnns and beyond. arXiv preprint arXiv:1602.06023.
  23. 23.Shashi Narayan, Shay B Cohen, and Mirella Lapata. 2018. Ranking sentences for extractive summarization with reinforcement learning. arXiv preprint arXiv:1802.08636.
  24. 24.Dragomir R Radev, Hongyan Jing, Małgorzata Stys, and Daniel Tam. 2004. Centroid-based summarization of multiple documents. Information Processing & Management, 40(6):919–938.
  25. 25.Nils Reimers and Iryna Gurevych. 2019. Sentence-bert: Sentence embeddings using siamese bert-networks. arXiv preprint arXiv:1908.10084.
  26. 26.Qian Ruan, Malte Ostendorff, and Georg Rehm. 2022. Histruct+: Improving extractive text summarization with hierarchical structure information. arXiv preprint arXiv:2203.09629.
  27. 27.Evan Sandhaus. 2008. The new york times annotated corpus. Linguistic Data Consortium, Philadelphia, 6(12):e26752.
  28. 28.Abigail See, Peter J Liu, and Christopher D Manning. 2017. Get to the point: Summarization with pointer-generator networks. arXiv preprint arXiv:1704.04368.
  29. 29.Sandeep Subramanian, Raymond Li, Jonathan Pilault, and Christopher Pal. 2019. On extractive and abstractive neural document summarization with transformer language models. arXiv preprint arXiv:1909.03186.
  30. 30.Frederick Suppe. 1998. The structure of a scientific paper. Philosophy of Science, 65(3):381–405.
  31. 31.Laurens van der Maaten and Geoffrey Hinton. 2008. Visualizing data using t-sne. Journal of Machine Learning Research, 9(86):2579–2605.
  32. 32.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008.
  33. 33.Petar Velickovi c, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Lio, and Yoshua Bengio. 2017. Graph attention networks. arXiv preprint arXiv:1710.10903.
  34. 34.Danqing Wang, Pengfei Liu, Yining Zheng, Xipeng Qiu, and Xuanjing Huang. 2020. Heterogeneous graph neural networks for extractive document summarization. arXiv preprint arXiv:2004.12393.
  35. 35.Lu Wang and Claire Cardie. 2013. Domain-independent abstract generation for focused meeting summarization. In Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1395–1405.
  36. 36.Wen Xiao and Giuseppe Carenini. 2019. Extractive summarization of long documents by combining global and local context. arXiv preprint arXiv:1909.08089.
  37. 37.Wen Xiao and Giuseppe Carenini. 2020. Systematically exploring redundancy reduction in summarizing long documents. arXiv preprint arXiv:2012.00052.
  38. 38.Jiacheng Xu, Zhe Gan, Yu Cheng, and Jingjing Liu. 2019. Discourse-aware neural extractive model for text summarization. arXiv preprint arXiv:1910.14142.
  39. 39.Michihiro Yasunaga, Rui Zhang, Kshitijh Meelu, Ayush Pareek, Krishnan Srinivasan, and Dragomir Radev. 2017. Graph-based neural multi-document summarization. arXiv preprint arXiv:1706.06681.
  40. 40.Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng, Guolin Ke, Di He, Yanming Shen, and Tie-Yan Liu. 2021. Do transformers really perform badly for graph representation? Advances in Neural Information Processing Systems, 34.
  41. 41.Haopeng Zhang, Semih Yavuz, Wojciech Kryscinski, Kazuma Hashimoto, and Yingbo Zhou. 2022. Improving the faithfulness of abstractive summarization via entity coverage control. In Findings of the Association for Computational Linguistics: NAACL 2022, pages 528–535.
  42. 42.Haopeng Zhang and Jiawei Zhang. 2020. Text graph transformer for document classification. In Conference on Empirical Methods in Natural Language Processing (EMNLP).
  43. 43.Jiawei Zhang, Haopeng Zhang, Li Sun, and Congying Xia. 2020. Graph-bert: Only attention is needed for learning graph representations. arXiv preprint arXiv:2001.05140.
  44. 44.Hao Zheng and Mirella Lapata. 2019. Sentence centrality revisited for unsupervised summarization. arXiv preprint arXiv:1906.03508.

Citation

MLA
Zhang, H., et al. “HEGEL: Hypergraph Transformer for Long Document Summarization”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 10167–76, https://doi.org/10.18653/v1/2022.emnlp-main.692.
APA
Zhang, H., Liu, X., & Zhang, J. (2022). HEGEL: Hypergraph Transformer for Long Document Summarization. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 10167–10176. https://doi.org/10.18653/v1/2022.emnlp-main.692
Chicago
Zhang, H., X. Liu, and J. Zhang. 2022. “HEGEL: Hypergraph Transformer for Long Document Summarization”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 10167–76. https://doi.org/10.18653/v1/2022.emnlp-main.692.
Harvard
Zhang, H., Liu, X. and Zhang, J. (2022) “HEGEL: Hypergraph Transformer for Long Document Summarization”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 10167–10176. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.692.
Vancouver
1. Zhang H, Liu X, Zhang J (2022) HEGEL: Hypergraph Transformer for Long Document Summarization. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 10167–10176

BibTeX

@inproceedings{zhang-etal-2022-hegel,
    title = "{HEGEL}: Hypergraph Transformer for Long Document Summarization",
    author = "Zhang, Haopeng  and
      Liu, Xiao  and
      Zhang, Jiawei",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.692/",
    doi = "10.18653/v1/2022.emnlp-main.692",
    pages = "10167--10176"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/