FactKB: Generalizable Factuality Evaluation using Language Models Enhanced with Factual Knowledge
Shangbin FengVidhisha BalachandranYuyang BaiYulia Tsvetkov
Proposes FactKB, a factuality evaluation framework that pretrains language models on structured knowledge base facts to effectively detect entity and relation errors in generated summaries across diverse domains.
Automated text summarization systems are increasingly deployed across news, social media, and scientific domains, yet they frequently generate inaccurate statements that distort factual information. Evaluating the factual consistency of these machine-generated summaries is vital for ensuring reliable and safe artificial intelligence deployment. However, existing evaluation metrics struggle with domain shifts and fail to catch subtle factual distortions—such as swapping entities, misattributing actions, or misrepresenting relationships—often scoring poorly when tested on content outside their narrow training domains.
The article introduces and evaluates FACTKB, a novel evaluation framework designed to robustly detect factual errors in machine-generated summaries across diverse domains. The core objective is to establish whether pretraining language models on structured factual knowledge improves their ability to verify entity and relation consistency without requiring complex document preprocessing.
To achieve this, the authors pretrained language models using structured facts extracted from external knowledge bases via three distinct strategies: direct entity-level facts (Entity Wiki), context-supported facts paired with Wikipedia descriptions (Evidence Extraction), and multi-hop entity pathways (Knowledge Walk). The pretrained models were then fine-tuned on human-annotated factual error detection data. The evaluation benchmarked FACTKB against several leading metrics on both in-domain news datasets (FactCollect and FRANK) and out-of-domain scientific and biomedical verification benchmarks (CovidFact, HealthVer, and SciFact).
The article demonstrates several significant findings. First, FACTKB established state-of-the-art performance on in-domain news evaluation, outperforming baseline models by an average of 3.8 balanced accuracy points on FactCollect and boosting correlation with human judgments on the FRANK benchmark by 5 to 15 correlation points. Second, in zero-shot cross-domain transfers to scientific literature, where baseline metrics degraded nearly to random chance (scoring around 50% balanced accuracy), FACTKB maintained robust accuracy and outperformed baselines by an average of 4.1 balanced accuracy points. Third, error-specific diagnostics revealed that the model's primary advantage stems from its superior ability to detect semantic frame errors—involving who did what to whom—which constitute over half of real-world summarization errors. Finally, the approach proved lightweight and broadly compatible across six major language model architectures and six knowledge base sources.
These findings suggest that factual pretraining effectively anchors language models against subtle entity confusions and hallucinated relations, addressing a primary vulnerability of automated summarization systems. For organizations deploying text generation, adopting a knowledge-enhanced evaluation metric substantially reduces the operational risk of publishing or relying on unfaithful summaries, particularly in high-stakes technical or regulatory environments. Unlike alternative graph- or question-answering-based metrics that require cumbersome preprocessing pipelines, FACTKB operates via standard sequence classification, minimizing computational overhead and deployment latency.
Organizations evaluating large-scale text generation workflows should consider integrating knowledge-pretrained factual classifiers to establish consistent quality benchmarks across both standard and specialized text domains. Practitioners should utilize moderate pretraining volumes and step lengths to prevent catastrophic forgetting of base language modeling capabilities. Further work is recommended to explore domain-tailored knowledge base pairings and multi-model ensembling to maximize fact-checking precision in specialized operational settings.
Confidence in these findings is supported by consistent, statistically significant improvements across multiple random seeds, diverse model architectures, and several independent benchmarks. However, leaders should note key limitations: the framework currently produces a binary factual decision rather than localized word-level error highlighting, relies on a two-stage training process that complicates hyperparameter tuning, and has primarily been validated on scientific literature beyond news, leaving domains like conversational and social media text for future verification.
- Paper: TRUE: Re-evaluating Factual Consistency Evaluation, Or Honovich et al. (2022). TRUE establishes the unified binary factual-consistency benchmark and evaluation setting that FactKB builds on to assess its verifier.
- Paper: ERNIE: Enhanced Language Representation with Informative Entities, Zhengyan Zhang et al. (2019). ERNIE introduces injecting structured knowledge into language-model pretraining, a core approach FactKB adapts for factuality evaluation.
No sufficiently relevant recommendations were found.
