FactKB: Generalizable Factuality Evaluation using Language Models Enhanced with Factual Knowledge

Shangbin FengVidhisha BalachandranYuyang BaiYulia Tsvetkov

article2023EMNLP69 citations

Proposes FactKB, a factuality evaluation framework that pretrains language models on structured knowledge base facts to effectively detect entity and relation errors in generated summaries across diverse domains.

Listen

Automated text summarization systems are increasingly deployed across news, social media, and scientific domains, yet they frequently generate inaccurate statements that distort factual information. Evaluating the factual consistency of these machine-generated summaries is vital for ensuring reliable and safe artificial intelligence deployment. However, existing evaluation metrics struggle with domain shifts and fail to catch subtle factual distortions—such as swapping entities, misattributing actions, or misrepresenting relationships—often scoring poorly when tested on content outside their narrow training domains.

The article introduces and evaluates FACTKB, a novel evaluation framework designed to robustly detect factual errors in machine-generated summaries across diverse domains. The core objective is to establish whether pretraining language models on structured factual knowledge improves their ability to verify entity and relation consistency without requiring complex document preprocessing.

To achieve this, the authors pretrained language models using structured facts extracted from external knowledge bases via three distinct strategies: direct entity-level facts (Entity Wiki), context-supported facts paired with Wikipedia descriptions (Evidence Extraction), and multi-hop entity pathways (Knowledge Walk). The pretrained models were then fine-tuned on human-annotated factual error detection data. The evaluation benchmarked FACTKB against several leading metrics on both in-domain news datasets (FactCollect and FRANK) and out-of-domain scientific and biomedical verification benchmarks (CovidFact, HealthVer, and SciFact).

The article demonstrates several significant findings. First, FACTKB established state-of-the-art performance on in-domain news evaluation, outperforming baseline models by an average of 3.8 balanced accuracy points on FactCollect and boosting correlation with human judgments on the FRANK benchmark by 5 to 15 correlation points. Second, in zero-shot cross-domain transfers to scientific literature, where baseline metrics degraded nearly to random chance (scoring around 50% balanced accuracy), FACTKB maintained robust accuracy and outperformed baselines by an average of 4.1 balanced accuracy points. Third, error-specific diagnostics revealed that the model's primary advantage stems from its superior ability to detect semantic frame errors—involving who did what to whom—which constitute over half of real-world summarization errors. Finally, the approach proved lightweight and broadly compatible across six major language model architectures and six knowledge base sources.

These findings suggest that factual pretraining effectively anchors language models against subtle entity confusions and hallucinated relations, addressing a primary vulnerability of automated summarization systems. For organizations deploying text generation, adopting a knowledge-enhanced evaluation metric substantially reduces the operational risk of publishing or relying on unfaithful summaries, particularly in high-stakes technical or regulatory environments. Unlike alternative graph- or question-answering-based metrics that require cumbersome preprocessing pipelines, FACTKB operates via standard sequence classification, minimizing computational overhead and deployment latency.

Organizations evaluating large-scale text generation workflows should consider integrating knowledge-pretrained factual classifiers to establish consistent quality benchmarks across both standard and specialized text domains. Practitioners should utilize moderate pretraining volumes and step lengths to prevent catastrophic forgetting of base language modeling capabilities. Further work is recommended to explore domain-tailored knowledge base pairings and multi-model ensembling to maximize fact-checking precision in specialized operational settings.

Confidence in these findings is supported by consistent, statistically significant improvements across multiple random seeds, diverse model architectures, and several independent benchmarks. However, leaders should note key limitations: the framework currently produces a binary factual decision rather than localized word-level error highlighting, relies on a two-stage training process that complicates hyperparameter tuning, and has primarily been validated on scientific literature beyond news, leaving domains like conversational and social media text for future verification.

No sufficiently relevant recommendations were found.

Cover for FactKB: Generalizable Factuality Evaluation using Language Models Enhanced with Factual Knowledge

Abstract

Evaluating the factual consistency of automatically generated summaries is essential for the progress and adoption of reliable summarization systems. Despite recent advances, existing factuality evaluation models are not robust, being especially prone to entity and relation errors in new domains. We propose FACTKB—a simple new approach to factuality evaluation that is generalizable across domains, in particular with respect to entities and relations. FACTKB is based on language models pretrained using facts extracted from external knowledge bases. We introduce three types of complementary factuality pretraining objectives based on entity-specific facts, facts extracted from auxiliary knowledge about entities, and facts constructed compositionally through knowledge base walks. The resulting factuality evaluation model achieves state-of-the-art performance on two in-domain news summarization benchmarks as well as on three out-of-domain scientific literature datasets. Further analysis of FACTKB shows improved ability to detect erroneous entities and relations in summaries and is robust and easily generalizable across domains. Code and data are available at https://github.com/BunsenFeng/FactKB.

Table of Contents

  • 1 Introduction
  • 2 FACTKB Methodology
  • 2.1 Factuality Pretraining
  • 2.2 FACTKB Training
  • 3 Data and Experiment Settings
  • 3.1 Training
  • 3.2 Evaluation
  • 4 Results
  • 5 Analysis and Discussion
  • 5.1 Where did FACTKB Improve?
  • 5.2 KB and LM Compatibility
  • 5.3 Simplicity Study
  • 5.4 Parameter Analysis
  • 6 Related Work
  • 7 Conclusion
  • Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • A Merging the Three Strategies
  • B Qualitative Analysis
  • C Dataset Details
  • D Baseline Details
  • E LM and KB Details
  • F Statistical Significance Test Details
  • G Computational Resources
  • H Scientific Artifacts

Knowls

  1. Knowl 1 — FACTKB uses knowledge-base pretraining before factuality classification

    model/method

    FACTKB is an entailment-style factuality evaluator for a generated summary and its source document. It first continues pretraining an encoder language model on text derived from a knowledge base (KB), using one of three entity- and relation-focused masked-language-modeling objectives. It then fine-tunes the resulting model on human-labeled factuality examples as a sequence classifier. At evaluation time, the input is SUMMARY [SEP] DOCUMENT, and the classifier predicts FACTUAL or NON-FACTUAL from the [CLS] representation. The intended benefit of the KB stage is stronger representations of entities, relations, and multi-hop facts, improving detection of factual errors that shift across domains.

  2. Knowl 2 — Entity Wiki pretraining reconstructs one-hop facts around each entity

    model/method

    For every entity in a KB, Entity Wiki gathers its one-hop neighboring entities and serializes the entity names and connecting relation names as a short collection of facts, separated by [SEP]. This produces at most one serialized example per KB entity. The language model is trained on this corpus with masked language modeling: entities and relations are randomly masked with probability pp, and the model predicts the masked text from the surrounding facts. This objective exposes the model to direct facts about entities and their relations; its corpus size is bounded by the number of KB entities.

  3. Knowl 3 — Evidence Extraction teaches prediction from entity-description context

    model/method

    Evidence Extraction samples a KB triple consisting of a head entity, a relation, and a tail entity. It forms an example from the head entity name, relation name, a masked tail, and the first paragraph of the head entity's Wikipedia description. The language model is trained to predict the masked tail using that auxiliary description as evidence. Repeating the procedure for NN sampled triples yields a corpus whose size is bounded by the number of KB triples. The objective is designed to encourage use of relevant surrounding evidence when judging whether a factual claim is supported.

  4. Knowl 4 — Knowledge Walk pretraining represents compositional multi-hop facts

    model/method

    Knowledge Walk begins at a randomly selected KB entity, repeatedly chooses a neighboring entity, and records the relation for each traversed edge. It serializes the resulting path as an ordered sequence of entity and relation names, producing a compositional statement spanning KK edges. Entities or relations in the path are randomly masked, and the language model learns to predict them from the remaining path context with masked language modeling. This objective targets multi-hop claims rather than only isolated triples. FACTKB's main experiments used walk length K=5K=5 and generated N=105N=10^5 walks.

  5. Knowl 5 — Training configuration and evaluation protocol

    experimental setup

    The main FACTKB experiments initialized from RoBERTa-base and used YAGO to construct the three pretraining corpora. Each objective was tested separately. The reported corpora contained 5.4 million tokens for Entity Wiki, 12.2 million for Evidence Extraction, and 2.7 million for Knowledge Walk. Pretraining used five epochs, learning rate 2×10−52\times10^{-5}, batch size 32, and Adam; fine-tuning used learning rate 10−410^{-4}, batch size 32, RAdam, a maximum of 50 epochs, and early stopping. The default masking probability was 0.15. Models were fine-tuned on FactCollect, a news-domain dataset of summary–article pairs with binary human factuality labels. For zero-shot scientific-domain evaluation, models trained on FactCollect were tested on CovidFact, HealthVer, and SciFact; examples labeled “NOT ENOUGH INFORMATION” were excluded from HealthVer and SciFact to make the evaluation binary.

  6. Knowl 6 — FACTKB improves balanced accuracy and F1 on FactCollect

    empirical result

    On FactCollect, FACTKB's three separately pretrained variants outperformed the prior factuality-evaluation baselines in the full test set and in the CNN/DM and XSUM subsets. The table gives mean scores and standard deviations across five random seeds. The strongest prior baseline shown for the full set and CNN/DM is FactGraph-adapters; for XSUM, the best prior balanced accuracy is FactGraph-adapters' and the best prior F1 is SummaC's. The paper reports the FACTKB improvements over existing methods as statistically significant in the marked comparisons.

    Setting Model BACC F1
    All data FactGraph-adapters (prior baseline) 87.6 ±\pm 0.7 87.8 ±\pm 0.7
    All data FACTKB-WIKI 89.3 ±\pm 0.4 89.5 ±\pm 0.5
    All data FACTKB-EVIDENCE 89.4 ±\pm 0.2 89.5 ±\pm 0.3
    All data FACTKB-WALK 89.1 ±\pm 0.4 89.3 ±\pm 0.5
    CNN/DM FactGraph-adapters (prior baseline) 76.0 ±\pm 2.8 87.5 ±\pm 0.4
    CNN/DM FACTKB-WIKI 77.3 ±\pm 0.3 88.2 ±\pm 0.6
    CNN/DM FACTKB-EVIDENCE 77.7 ±\pm 1.4 87.9 ±\pm 0.7
    CNN/DM FACTKB-WALK 78.3 ±\pm 1.2 87.7 ±\pm 0.4
    XSUM FACTKB-WIKI 77.3 ±\pm 1.3 91.8 ±\pm 1.2
    XSUM FACTKB-EVIDENCE 76.8 ±\pm 1.9 90.8 ±\pm 0.8
    XSUM FACTKB-WALK 76.4 ±\pm 0.3 90.4 ±\pm 1.4

    For XSUM, the strongest prior balanced accuracy was 69.9 ±\pm 2.3 from FactGraph-adapters, while the strongest prior F1 was 90.4 from SummaC. The results show that all three KB-based pretraining strategies are useful, although which variant leads depends on the dataset subset and metric.

  7. Knowl 7 — FACTKB correlates more strongly with FRANK human judgments

    empirical result

    On FRANK, which evaluates agreement with human judgments of summary factuality, FACTKB achieved the highest correlation in five of the six reported Pearson and Spearman settings. The values below are Pearson ρ\rho and Spearman rr for the full dataset, CNN/DM, and XSUM; the paper reports displayed pp-values of .00 for all FACTKB entries. FactGraph is included as the strongest prior classification-based comparator in the full and CNN/DM settings; on XSUM its Spearman correlation is higher than the FACTKB variants'.

    Setting Model Pearson ρ\rho Spearman rr
    All data FactGraph .35 .42
    All data FACTKB-WIKI .46 .52
    All data FACTKB-EVIDENCE .43 .49
    All data FACTKB-WALK .47 .52
    CNN/DM FactGraph .45 .34
    CNN/DM FACTKB-WIKI .57 .49
    CNN/DM FACTKB-EVIDENCE .53 .45
    CNN/DM FACTKB-WALK .57 .45
    XSUM FactGraph .30 .49
    XSUM FACTKB-WIKI .29 .39
    XSUM FACTKB-EVIDENCE .31 .37
    XSUM FACTKB-WALK .35 .36

    The strongest gains occur on the full FRANK set and CNN/DM. On XSUM, FACTKB-EVIDENCE has the highest Pearson correlation among the listed FACTKB variants, while FactGraph has the highest Spearman correlation.

  8. Knowl 8 — FACTKB transfers better to scientific factuality datasets

    empirical result

    The authors trained FACTKB on news-domain FactCollect and evaluated it zero-shot on three scientific or biomedical claim-verification datasets. The table reports balanced accuracy (BACC) and F1; RoBERTa is the non-KB-initialized reference, and each FACTKB row corresponds to a separate pretraining objective. The paper reports an average improvement of 4.1 BACC points over existing factuality metrics across the three datasets. Entity Wiki is strongest on CovidFact and HealthVer BACC, while Knowledge Walk has the highest SciFact BACC; Evidence Extraction has the highest SciFact F1. Thus, the results support transfer to these tested domains but do not imply that every FACTKB variant leads on every metric.

    Dataset Model BACC F1
    CovidFact RoBERTa 59.0 ±\pm 3.2 46.4 ±\pm 4.3
    CovidFact FACTKB-WIKI 64.8 ±\pm 0.3 54.4 ±\pm 0.7
    CovidFact FACTKB-EVIDENCE 63.9 ±\pm 0.6 53.3 ±\pm 1.7
    CovidFact FACTKB-WALK 63.7 ±\pm 1.0 53.1 ±\pm 1.6
    HealthVer RoBERTa 55.0 ±\pm 2.2 50.0 ±\pm 3.9
    HealthVer FACTKB-WIKI 60.1 ±\pm 0.4 71.6 ±\pm 2.9
    HealthVer FACTKB-EVIDENCE 59.0 ±\pm 1.0 70.8 ±\pm 0.9
    HealthVer FACTKB-WALK 58.5 ±\pm 0.5 68.7 ±\pm 1.7
    SciFact RoBERTa 58.1 ±\pm 4.0 71.3 ±\pm 3.5
    SciFact FACTKB-WIKI 62.9 ±\pm 0.4 72.3 ±\pm 1.1
    SciFact FACTKB-EVIDENCE 61.4 ±\pm 0.5 74.1 ±\pm 1.6
    SciFact FACTKB-WALK 63.1 ±\pm 1.1 67.6 ±\pm 4.1

    The models were evaluated on the test sets, with five-seed averages and standard deviations reported for the FACTKB and RoBERTa rows. The out-of-domain chart and results indicate a substantial decline for many prior metrics relative to their news-domain performance, whereas FACTKB retains stronger performance.

  9. Knowl 9 — The clearest error-category gain is on entity and relation mistakes

    empirical result

    Using FRANK's error categories—semantic frame, discourse, and content verifiability—the authors assessed how removing each category changed correlation with human judgments. FACTKB showed a significant advantage over existing metrics for semantic-frame errors, which concern entities and the relations among them. Its performance on discourse and content-verifiability errors was described as slightly better than or comparable to other metrics. This category-level analysis is consistent with the intended role of KB pretraining: improving detection of entity and relation errors without sacrificing coverage of other factuality errors.

  10. Knowl 10 — The method generally transfers across language models and knowledge bases

    empirical result

    The authors tested FACTKB with six language-model checkpoints—RoBERTa, ELECTRA, BART, DeBERTa, ALBERT, and DistilRoBERTa—and six knowledge bases—YAGO, Wikidata, ConceptNet, ATOMIC, KGAP, and UMLS—using the FactCollect balanced-accuracy evaluation. Across these combinations, the KB-based factuality pretraining generally improved error detection over the corresponding vanilla language-model checkpoint, supporting FACTKB as a method rather than a RoBERTa–YAGO-only system. The experiments also showed variation by choice of components: RoBERTa and DeBERTa among the tested language models, and YAGO, KGAP, and UMLS among the tested KBs, tended to perform better. The study does not establish a universally optimal LM–KB pairing.

  11. Knowl 11 — Moderate pretraining amounts and five-edge walks performed best in the tested sweeps

    empirical result

    Parameter sweeps found that Evidence Extraction and Knowledge Walk generally worked well with corpus sizes of 10410^4 or 10510^5 examples; substantially larger corpora could reduce performance. The authors suggest catastrophic forgetting as one possible explanation, not as a demonstrated mechanism. Across the tested pretraining durations, 1–10 epochs were generally preferable to very long continued pretraining. For Knowledge Walk, the tested lengths were 1, 2, 5, 10, and 50 edges, and K=5K=5 gave the best performance, which the authors interpret as a useful balance in compositionality. The main experiments therefore used N=105N=10^5, five pretraining epochs, and K=5K=5.

  12. Knowl 12 — FACTKB's demonstrated scope and diagnostic capabilities remain limited

    limitation

    The evidence for out-of-domain generalization is concentrated on scientific and biomedical datasets; performance on other domains, including social media, remains untested in this study. FACTKB gives a factual/non-factual classification but does not localize the error or provide the fine-grained diagnosis offered by some more heavily processed systems. Its two-stage pretraining-plus-fine-tuning design also leaves multiple hyperparameters to tune, and the best language-model and knowledge-base combination is not known in advance. These limits qualify the paper's claims of broad generalizability and ease of use.

Coverage note — The attempted combined-pretraining variant is omitted because it was not significantly better than a single strategy; qualitative examples and the usage-complexity comparison are illustrative or secondary rather than independent findings.

References

  1. 1.Oshin Agarwal, Heming Ge, Siamak Shakeri, and Rami Al-Rfou. 2021. Knowledge graph based synthetic corpus generation for knowledge-enhanced language model pre-training. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3554–3565.
  2. 2.Roee Aharoni, Shashi Narayan, Joshua Maynez, Jonathan Herzig, Elizabeth Clark, and Mirella Lapata. 2023. Multilingual summarization with factual consistency evaluation. In Findings of the Association for Computational Linguistics: ACL 2023, pages 3562–3591.
  3. 3.Alfonso Amayuelas, Shuai Zhang, Xi Susie Rao, and Ce Zhang. 2021. Neural methods for logical reasoning over knowledge graphs. In International Conference on Learning Representations.
  4. 4.Vidhisha Balachandran, Hannaneh Hajishirzi, William Cohen, and Yulia Tsvetkov. 2022. Correcting diverse factual errors in abstractive summarization via post-editing and language model infilling. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing.
  5. 5.Vidhisha Balachandran, Artidoro Pagnoni, Jay Yoon Lee, Dheeraj Rajagopal, Jaime Carbonell, and Yulia Tsvetkov. 2021. StructSum: Summarization via structured representations. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 2575–2585.
  6. 6.Jacob Benesty, Jingdong Chen, Yiteng Huang, and Israel Cohen. 2009. Pearson correlation coefficient. In Noise reduction in speech processing, pages 1–4. Springer.
  7. 7.Abhik Bhattacharjee, Tahmid Hasan, Wasi Uddin Ahmad, Yuan-Fang Li, Yong-Bin Kang, and Rifat Shahriyar. 2023. CrossSum: Beyond English-centric cross-lingual summarization for 1,500+ language pairs. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2541–2564.
  8. 8.Steven Bird, Ewan Klein, and Edward Loper. 2009. Natural language processing with Python: analyzing text with the natural language toolkit. " O’Reilly Media, Inc.".
  9. 9.Antoine Bordes, Nicolas Usunier, Alberto Garcia-Duran, Jason Weston, and Oksana Yakhnenko. 2013. Translating embeddings for modeling multi-relational data. Advances in neural information processing systems, 26.
  10. 10.Antoine Bosselut, Ronan Le Bras, and Yejin Choi. 2021. Dynamic neuro-symbolic knowledge graph construction for zero-shot commonsense question answering. In Proceedings of the 35th AAAI Conference on Artificial Intelligence (AAAI).
  11. 11.Isabel Cachola, Kyle Lo, Arman Cohan, and Daniel S Weld. 2020. Tldr: Extreme summarization of scientific documents. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 4766–4777.
  12. 12.Ziqiang Cao, Furu Wei, Wenjie Li, and Sujian Li. 2018. Faithful to the original: Fact aware neural abstractive summarization. In thirty-second AAAI conference on artificial intelligence.
  13. 13.Wenhu Chen, Yu Su, Xifeng Yan, and William Yang Wang. 2020. Kgpt: Knowledge-grounded pre-training for data-to-text generation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 8635–8648.
  14. 14.Xiuying Chen, Guodong Long, Chongyang Tao, Mingzhe Li, Xin Gao, Chengqi Zhang, and Xiangliang Zhang. 2023a. Improving the robustness of summarization systems with dual augmentation. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6846–6857.
  15. 15.Yulong Chen, Yang Liu, Ruochen Xu, Ziyi Yang, Chenguang Zhu, Michael Zeng, and Yue Zhang. 2023b. UniSumm and SummZoo: Unified model and diverse benchmark for few-shot summarization. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 12833–12855.
  16. 16.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186.
  17. 17.Pierre Dognin, Inkit Padhi, Igor Melnyk, and Payel Das. 2021. ReGen: Reinforcement learning for text and knowledge base generation using pretrained language models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 1084–1099.
  18. 18.Hady Elsahar and Matthias Gallé. 2019. To annotate or not? predicting performance drop under domain shift. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2163–2173.
  19. 19.Matan Eyal, Tal Baumel, and Michael Elhadad. 2019. Question answering as an automatic evaluation metric for news article summarization. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 3938–3948.
  20. 20.Alexander Fabbri, Chien-Sheng Wu, Wenhao Liu, and Caiming Xiong. 2022. QAFactEval: Improved QA-based factual consistency evaluation for summarization. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2587–2601.
  21. 21.William Falcon and The PyTorch Lightning team. 2019. PyTorch Lightning.
  22. 22.Yair Feldman and Ran El-Yaniv. 2019. Multi-hop paragraph retrieval for open-domain question answering. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2296–2309.
  23. 23.Shangbin Feng, Zilong Chen, Wenqian Zhang, Qingyao Li, Qinghua Zheng, Xiaojun Chang, and Minnan Luo. 2021. Kgap: Knowledge graph augmented political perspective detection in news media. arXiv preprint arXiv:2108.03861.
  24. 24.Shangbin Feng, Zhaoxuan Tan, Wenqian Zhang, Zhenyu Lei, and Yulia Tsvetkov. 2023. KALM: Knowledge-aware integration of local, document, and global contexts for long document understanding. In Proceedings of ACL 2023, pages 2116–2138.
  25. 25.Yue Feng, Zhen Han, Mingming Sun, and Ping Li. 2022. Multi-hop open-domain question answering over structured and unstructured knowledge. In Findings of the Association for Computational Linguistics: NAACL 2022, pages 151–156, Seattle, United States.
  26. 26.Joseph Fisher, Arpit Mittal, Dave Palfrey, and Christos Christodoulopoulos. 2020. Debiasing knowledge graph embeddings. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7332–7345.
  27. 27.Nicolas Garneau and Luc Lamontagne. 2021. Trainable ranking models to evaluate the semantic accuracy of data-to-text neural generator. In Proceedings of the 2nd Workshop on Evaluation and Comparison of NLP Systems, pages 51–61.
  28. 28.Tomas Goldsack, Zhihao Zhang, Chenghua Lin, and Carolina Scarton. 2022. Making science simple: Corpora for the lay summarisation of scientific literature. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing.
  29. 29.Tanya Goyal and Greg Durrett. 2020. Evaluating factuality in generation with dependency-level entailment. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3592–3603.
  30. 30.Tanya Goyal and Greg Durrett. 2021. Annotating and modeling fine-grained factuality in summarization. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies.
  31. 31.Suchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A Smith. 2020. Don’t stop pretraining: Adapt language models to domains and tasks. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8342–8360.
  32. 32.Charles R Harris, K Jarrod Millman, Stéfan J Van Der Walt, Ralf Gommers, Pauli Virtanen, David Cournapeau, Eric Wieser, Julian Taylor, Sebastian Berg, Nathaniel J Smith, et al. 2020. Array programming with numpy. Nature, 585(7825):357–362.
  33. 33.Pengcheng He, Baolin Peng, Song Wang, Yang Liu, Ruochen Xu, Hany Hassan, Yu Shi, Chenguang Zhu, Wayne Xiong, Michael Zeng, Jianfeng Gao, and Xuedong Huang. 2023. Z-code++: A pre-trained language model optimized for abstractive summarization. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5095–5112.
  34. 34.Ruifang He, Liangliang Zhao, and Huanyu Liu. 2020. TWEETSUM: Event oriented social summarization dataset. In Proceedings of the 28th International Conference on Computational Linguistics, pages 5731–5736.
  35. 35.Yu-Jung Heo, Eun-Sol Kim, Woo Suk Choi, and Byoung-Tak Zhang. 2022. Hypergraph transformer: Weakly-supervised multi-hop reasoning for knowledge-based visual question answering. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 373–390.
  36. 36.Linmei Hu, Tianchi Yang, Luhao Zhang, Wanjun Zhong, Duyu Tang, Chuan Shi, Nan Duan, and Ming Zhou. 2021. Compare to the knowledge: Graph neural fake news detection with external knowledge. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 754–763.
  37. 37.Yong-Ho Jung, Jun-Hyung Park, Joon-Young Choi, Mingyu Lee, Junho Kim, Kang-Min Kim, and SangKeun Lee. 2022. Learning from missing relations: Contrastive learning with commonsense knowledge graphs for commonsense inference. In Findings of the Association for Computational Linguistics: ACL 2022, pages 1514–1523.
  38. 38.Ambedkar Kanapala, Sukomal Pal, and Rajendra Pamula. 2019. Text summarization from legal documents: a survey. Artificial Intelligence Review, 51(3):371–402.
  39. 39.Ryuji Kano, Yasuhide Miura, Motoki Taniguchi, Yan-Ying Chen, Francine Chen, and Tomoko Ohkuma. 2018. Harnessing popularity in social media for extractive summarization of online conversations. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 1139–1145.
  40. 40.Yu Jin Kim, Beong-woo Kwak, Youngwook Kim, Reinald Kim Amplayo, Seung-won Hwang, and Jinyoung Yeo. 2022. Modularized transfer learning with multiple knowledge graphs for zero-shot commonsense reasoning. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2244–2257.
  41. 41.Wojciech Kryściński, Bryan McCann, Caiming Xiong, and Richard Socher. 2020. Evaluating the factual consistency of abstractive text summarization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9332–9346.
  42. 42.Philippe Laban, Tobias Schnabel, Paul N. Bennett, and Marti A. Hearst. 2022. SummaC: Re-Visiting NLI-based Models for Inconsistency Detection in Summarization. Transactions of the Association for Computational Linguistics, 10:163–177.
  43. 43.Timothée Lacroix, Guillaume Obozinski, and Nicolas Usunier. 2019. Tensor decompositions for temporal knowledge base completion. In International Conference on Learning Representations.
  44. 44.Egoitz Laparra, Steven Bethard, and Timothy A Miller. 2020. Rethinking domain adaptation for machine learning over clinical language. JAMIA open, 3(2):146–150.
  45. 45.Guy Lev, Michal Shmueli-Scheuer, Jonathan Herzig, Achiya Jerbi, and David Konopnicki. 2019. TalkSumm: A dataset and scalable annotation method for scientific paper summarization based on conference talks. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2125–2131.
  46. 46.Chen Li, Zhongyu Wei, Yang Liu, Yang Jin, and Fei Huang. 2016. Using relevant public posts to enhance news article summarization. In Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: Technical Papers, pages 557–566.
  47. 47.Shaobo Li, Xiaoguang Li, Lifeng Shang, Chengjie Sun, Bingquan Liu, Zhenzhou Ji, Xin Jiang, and Qun Liu. 2022. Pre-training language models with deterministic factual knowledge. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing.
  48. 48.Paul Pu Liang, Chiyu Wu, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2021. Towards understanding and mitigating social biases in language models. In International Conference on Machine Learning, pages 6565–6576. PMLR.
  49. 49.Jiacheng Liu, Alisa Liu, Ximing Lu, Sean Welleck, Peter West, Ronan Le Bras, Yejin Choi, and Hannaneh Hajishirzi. 2022a. Generated knowledge prompting for commonsense reasoning. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3154–3169.
  50. 50.Xiao Liu, Shiyu Zhao, Kai Su, Yukuo Cen, Jiezhong Qiu, Mengdi Zhang, Wei Wu, Yuxiao Dong, and Jie Tang. 2022b. Mask and reason: Pre-training knowledge graph transformers for complex logical queries. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages 1120–1130.
  51. 51.Yang Liu and Mirella Lapata. 2019. Text summarization with pretrained encoders. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3730–3740.
  52. 52.Yang Liu, Chenguang Zhu, and Michael Zeng. 2022c. End-to-end segmentation-based news summarization. In Findings of the Association for Computational Linguistics: ACL 2022, pages 544–554.
  53. 53.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
  54. 54.Yixin Liu, Budhaditya Deb, Milagro Teruel, Aaron Halfa ker, Dragomir Radev, and Ahmed Hassan Awadallah. 2023a. On improving summarization factual consistency from natural language feedback. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 15144–15161.
  55. 55.Yixin Liu, Alex Fabbri, Pengfei Liu, Yilun Zhao, Linyong Nan, Ruilin Han, Simeng Han, Shafiq Joty, Chien-Sheng Wu, Caiming Xiong, and Dragomir Radev. 2023b. Revisiting the gold standard: Grounding summarization evaluation with robust human evaluation. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 4140–4170.
  56. 56.Yongtai Liu, Joshua Maynez, Gonçalo Simões, and Shashi Narayan. 2022d. Data augmentation for low-resource dialogue summarization. In Findings of the Association for Computational Linguistics: NAACL 2022, pages 703–710.
  57. 57.Zheheng Luo, Qianqian Xie, and Sophia Ananiadou. 2023. Chatgpt as a factual inconsistency evaluator for abstractive text summarization. arXiv preprint arXiv:2303.15621.
  58. 58.Kaixin Ma, Hao Cheng, Xiaodong Liu, Eric Nyberg, and Jianfeng Gao. 2022. Open domain question answering with a unified knowledge interface. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1605–1620, Dublin, Ireland.
  59. 59.Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020. On faithfulness and factuality in abstractive summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1906–1919.
  60. 60.Ninareh Mehrabi, Pei Zhou, Fred Morstatter, Jay Pujara, Xiang Ren, and Aram Galstyan. 2021. Lawyers are dishonest? quantifying representational harms in commonsense knowledge resources. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 5016–5033.
  61. 61.Kevin Meng, David Bau, Alex J Andonian, and Yonatan Belinkov. 2022. Locating and editing factual associations in gpt. In Advances in Neural Information Processing Systems.
  62. 62.Sayantan Mitra, Roshni Ramnani, and Shubhashis Sengupta. 2022. Constraint-based multi-hop question answering with knowledge graph. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: Industry Track, pages 280–288.
  63. 63.Leann Myers and Maria J Sirois. 2004. Spearman correlation coefficients, differences between. Encyclopedia of statistical sciences, 12.
  64. 64.Moin Nadeem, Anna Bethke, and Siva Reddy. 2021. Stereoset: Measuring stereotypical bias in pretrained language models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 5356–5371.
  65. 65.Feng Nan, Cicero dos Santos, Henghui Zhu, Patrick Ng, Kathleen Mckeown, Ramesh Nallapati, Dejiao Zhang, Zhiguo Wang, Andrew O Arnold, and Bing Xiang. 2021. Improving factual consistency of abstractive summarization via question answering. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6881–6894.
  66. 66.Shashi Narayan, Shay Cohen, and Maria Lapata. 2018. Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization. In 2018 Conference on Empirical Methods in Natural Language Processing, pages 1797–1807. Association for Computational Linguistics.
  67. 67.Shashi Narayan, Yao Zhao, Joshua Maynez, Gonçalo Simões, Vitaly Nikolaev, and Ryan McDonald. 2021. Planning with learned entity prompts for abstractive summarization. Transactions of the Association for Computational Linguistics, 9:1475–1492.
  68. 68.Barlas Oguz, Xilun Chen, Vladimir Karpukhin, Stan Peshterliev, Dmytro Okhonko, Michael Schlichtkrull, Sonal Gupta, Yashar Mehdad, and Scott Yih. 2022. UniK-QA: Unified representations of structured and unstructured knowledge for open-domain question answering. In Findings of the Association for Computational Linguistics: NAACL 2022, pages 1535–1546, Seattle, United States.
  69. 69.Artidoro Pagnoni, Vidhisha Balachandran, and Yulia Tsvetkov. 2021. Understanding factuality in abstractive summarization with frank: A benchmark for factuality metrics. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4812–4829.
  70. 70.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. 2019. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems, 32.
  71. 71.F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, R. Dubourg, J. Vanderplas, J. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. 2011. Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12:2825–2830.
  72. 72.Thomas Pellissier Tanon, Gerhard Weikum, and Fabian Suchanek. 2020. Yago 4: A reason-able knowledge base. In European Semantic Web Conference, pages 583–596. Springer.
  73. 73.Xutan Peng, Yi Zheng, Chenghua Lin, and Advaith Siddharthan. 2021. Summarising historical text in modern languages. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 3123–3142.
  74. 74.Seth Polsley, Pooja Jhunjhunwala, and Ruihong Huang. 2016. Casesummarizer: a system for automated summarization of legal texts. In Proceedings of COLING 2016, the 26th international conference on Computational Linguistics: System Demonstrations, pages 258–262.
  75. 75.Ofir Press, Muru Zhang, Sewon Min, Ludwig Schmidt, Noah A Smith, and Mike Lewis. 2022. Measuring and narrowing the compositionality gap in language models. arXiv preprint arXiv:2210.03350.
  76. 76.Vinay Venkatesh Ramasesh, Aitor Lewkowycz, and Ethan Dyer. 2021. Effect of scale on catastrophic forgetting in neural networks. In International Conference on Learning Representations.
  77. 77.Leonardo F. R. Ribeiro, Mengwen Liu, Iryna Gurevych, Markus Dreyer, and Mohit Bansal. 2022. FactGraph: Evaluating factuality in summarization with semantic graph representations. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3238–3253.
  78. 78.Md Rashad Al Hasan Rony, Ricardo Usbeck, and Jens Lehmann. 2022. DialoKG: Knowledge-structure aware task-oriented dialogue generation. In Findings of the Association for Computational Linguistics: NAACL 2022, pages 2557–2571, Seattle, United States.
  79. 79.Corby Rosset, Chenyan Xiong, Minh Hieu Phan, Xia Song, Paul Bennett, and Saurabh Tiwary. 2020. Knowledge-aware language model pretraining. ArXiv, abs/2007.00655.
  80. 80.Sascha Rothe, Joshua Maynez, and Shashi Narayan. 2021. A thorough evaluation of task-specific pre-training for summarization. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 140–145.
  81. 81.Arkadiy Saakyan, Tuhin Chakrabarty, and Smaranda Muresan. 2021. Covid-fact: Fact extraction and verification of real-world claims on covid-19 pandemic. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 2116–2129.
  82. 82.Mourad Sarrouti, Asma Ben Abacha, Yassine M’rabet, and Dina Demner-Fushman. 2021. Evidence-based fact-checking of health-related claims. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 3499–3512.
  83. 83.Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano, Alex Wang, and Patrick Gallinari. 2021. Questeval: Summarization asks for fact-based evaluation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6594–6604.
  84. 84.Omar Shaikh, Hongxin Zhang, William Held, Michael Bernstein, and Diyi Yang. 2023. On second thought, let’s not think step by step! bias and toxicity in zero-shot reasoning. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 4454–4470.
  85. 85.Siamak Shakeri, Cicero dos Santos, Henghui Zhu, Patrick Ng, Feng Nan, Zhiguo Wang, Ramesh Nallapati, and Bing Xiang. 2020. End-to-end synthetic data generation for domain adaptation of question answering systems. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5445–5460.
  86. 86.Robyn Speer, Joshua Chin, and Catherine Havasi. 2017. Conceptnet 5.5: An open multilingual graph of general knowledge. In Thirty-first AAAI conference on artificial intelligence.
  87. 87.Shahbaz Syed, Michael Völske, Nedim Lipka, Benno Stein, Hinrich Schütze, and Martin Potthast. 2019. Towards summarization for social media - results of the TL;DR challenge. In Proceedings of the 12th International Conference on Natural Language Generation, pages 523–528.
  88. 88.Masato Takatsuka, Tetsunori Kobayashi, and Yoshihiko Hayashi. 2022. Phrase-level localization of inconsistency errors in summarization by weak supervision. In Proceedings of the 29th International Conference on Computational Linguistics, pages 6151–6164.
  89. 89.Yi Chern Tan and L Elisa Celis. 2019. Assessing social and intersectional biases in contextualized word representations. Advances in Neural Information Processing Systems, 32.
  90. 90.Liyan Tang, Tanya Goyal, Alex Fabbri, Philippe Laban, Jiacheng Xu, Semih Yavuz, Wojciech Kryscinski, Justin Rousseau, and Greg Durrett. 2023. Understanding factual errors in summarization: Errors, summarizers, datasets, error detectors. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 11626–11644.
  91. 91.Xiangru Tang, Arjun Nair, Borui Wang, Bingyao Wang, Jai Desai, Aaron Wade, Haoran Li, Asli Celikyilmaz, Yashar Mehdad, and Dragomir Radev. 2022. CONFIT: Toward faithful dialogue summarization with linguistically-informed contrastive fine-tuning. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5657–5668.
  92. 92.Prasetya Utama, Joshua Bambrick, Nafise Moosavi, and Iryna Gurevych. 2022. Falsesum: Generating document-level NLI examples for recognizing factual inconsistency in summarization. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2763–2776.
  93. 93.Shikhar Vashishth, Soumya Sanyal, Vikram Nitin, and Partha Talukdar. 2019. Composition-based multi-relational graph convolutional networks. In International Conference on Learning Representations.
  94. 94.Ellen Voorhees, Tasmeer Alam, Steven Bedrick, Dina Demner-Fushman, William R Hersh, Kyle Lo, Kirk Roberts, Ian Soboroff, and Lucy Lu Wang. 2021. Trec-covid: constructing a pandemic information retrieval test collection. In ACM SIGIR Forum, volume 54, pages 1–12. ACM New York, NY, USA.
  95. 95.Denny Vrandecic and Markus Krötzsch. 2014. Wikidata: a free collaborative knowledgebase. Communications of the ACM, 57(10):78–85.
  96. 96.David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi. 2020. Fact or fiction: Verifying scientific claims. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7534–7550.
  97. 97.David Wadden, Kyle Lo, Lucy Lu Wang, Arman Cohan, Iz Beltagy, and Hannaneh Hajishirzi. 2022. MultiVerS: Improving scientific claim verification with weak supervision and full-document context. In Findings of the Association for Computational Linguistics: NAACL 2022, pages 61–76.
  98. 98.Alex Wang, Kyunghyun Cho, and Mike Lewis. 2020a. Asking and answering questions to evaluate the factual consistency of summaries. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5008–5020.
  99. 99.Lucy Lu Wang, Kyle Lo, Yoganand Chandrasekhar, Russell Reas, Jiangjiang Yang, Doug Burdick, Darrin Eide, Kathryn Funk, Yannis Katsis, Rodney Michael Kinney, et al. 2020b. Cord-19: The covid-19 open research dataset. In Proceedings of the 1st Workshop on NLP for COVID-19 at ACL 2020.
  100. 100.Wenya Wang and Sinno Pan. 2022. Deep inductive logic reasoning for multi-hop reading comprehension. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 4999–5009.
  101. 101.Xiaozhi Wang, Tianyu Gao, Zhaocheng Zhu, Zhengyan Zhang, Zhiyuan Liu, Juanzi Li, and Jian Tang. 2021. Kepler: A unified model for knowledge embedding and pre-trained language representation. Transactions of the Association for Computational Linguistics, 9:176–194.
  102. 102.Peter West, Chandra Bhagavatula, Jack Hessel, Jena Hwang, Liwei Jiang, Ronan Le Bras, Ximing Lu, Sean Welleck, and Yejin Choi. 2022. Symbolic knowledge distillation: from general language models to commonsense models. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4602–4625.
  103. 103.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 conference on empirical methods in natural language processing: system demonstrations, pages 38–45.
  104. 104.Yuexiang Xie, Fei Sun, Yang Deng, Yaliang Li, and Bolin Ding. 2021. Factual consistency evaluation for text summarization via counterfactual estimation. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 100–110.
  105. 105.Xinnuo Xu, Ondřej Dušek, Jingyi Li, Verena Rieser, and Ioannis Konstas. 2020. Fact-based content weighting for evaluating abstractive summarisation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5071–5081.
  106. 106.Michihiro Yasunaga, Antoine Bosselut, Hongyu Ren, Xikun Zhang, Christopher D Manning, Percy S Liang, and Jure Leskovec. 2022. Deep bidirectional language-knowledge graph pretraining. Advances in Neural Information Processing Systems, 35:37309–37323.
  107. 107.Michihiro Yasunaga, Hongyu Ren, Antoine Bosselut, Percy Liang, and Jure Leskovec. 2021. Qa-gnn: Reasoning with language models and knowledge graphs for question answering. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 535–546.
  108. 108.Wenhao Yu, Meng Jiang, Zhiting Hu, Qingyun Wang, Heng Ji, and Nazneen Rajani. 2021. Knowledge-enriched natural language generation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing: Tutorial Abstracts, pages 11–16.
  109. 109.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019. Bertscore: Evaluating text generation with bert. In International Conference on Learning Representations.
  110. 110.Wenqian Zhang, Shangbin Feng, Zilong Chen, Zhenyu Lei, Jundong Li, and Minnan Luo. 2022a. KCD: Knowledge walks and textual cues enhanced political perspective detection in news media. In Proceedings of NAACL 2022, pages 4129–4140.
  111. 111.X Zhang, A Bosselut, M Yasunaga, H Ren, P Liang, C Manning, and J Leskovec. 2022b. Greaselm: Graph reasoning enhanced language models for question answering. In International Conference on Representation Learning (ICLR).

Citation

MLA
Feng, S., et al. “FactKB: Generalizable Factuality Evaluation Using Language Models Enhanced with Factual Knowledge”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 933–52, https://doi.org/10.18653/v1/2023.emnlp-main.59.
APA
Feng, S., Balachandran, V., Bai, Y., & Tsvetkov, Y. (2023). FactKB: Generalizable Factuality Evaluation using Language Models Enhanced with Factual Knowledge. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 933–952. https://doi.org/10.18653/v1/2023.emnlp-main.59
Chicago
Feng, S., V. Balachandran, Y. Bai, and Y. Tsvetkov. 2023. “FactKB: Generalizable Factuality Evaluation Using Language Models Enhanced with Factual Knowledge”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 933–52. https://doi.org/10.18653/v1/2023.emnlp-main.59.
Harvard
Feng, S. et al. (2023) “FactKB: Generalizable Factuality Evaluation using Language Models Enhanced with Factual Knowledge”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 933–952. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.59.
Vancouver
1. Feng S, Balachandran V, Bai Y, Tsvetkov Y (2023) FactKB: Generalizable Factuality Evaluation using Language Models Enhanced with Factual Knowledge. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 933–952

BibTeX

@inproceedings{feng-etal-2023-factkb,
    title = "{F}act{KB}: Generalizable Factuality Evaluation using Language Models Enhanced with Factual Knowledge",
    author = "Feng, Shangbin  and
      Balachandran, Vidhisha  and
      Bai, Yuyang  and
      Tsvetkov, Yulia",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.59/",
    doi = "10.18653/v1/2023.emnlp-main.59",
    pages = "933--952"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/