Knowledge vault: a web-scale approach to probabilistic knowledge fusion
Xin Luna DongEvgeniy GabrilovichGeremy HeitzWilko HornNi LaoKevin MurphyThomas StrohmannShaohua SunWei Zhang
Proposes an automated web-scale framework that builds an extensive knowledge base by probabilistically fusing noisy web extractions across multiple data modalities with graph-based prior knowledge to accurately estimate fact correctness.
Large-scale knowledge repositories are critical for modern search and question-answering systems, but existing repositories remain highly incomplete due to their reliance on human curation, structured data, and text-only extraction methods. For example, major bases lack basic facts such as place of birth or nationality for over 70% of indexed people. Unsupervised web extraction can expand coverage, but traditional pipelines introduce massive noise and false positives. The article introduces and evaluates Knowledge Vault, an automated, web-scale framework designed to build a probabilistic knowledge base by fusing diverse web extractions with predictive graph priors learned from existing knowledge.
The research demonstrates that combining noisy multi-source web extractions with graph-based prior models significantly increases fact coverage while maintaining rigorous quality control. The approach extracts 1.6 billion candidate subject-predicate-object triples across 4,469 relation types using four web extraction streams: free text, page structure trees, relational web tables, and human page annotations. These extractions are combined with two link-prediction graph models—a path-ranking algorithm and a neural network multi-layer perceptron—trained on Freebase data. Supervised machine learning algorithms fuse these distinct evidence sources and calibrate probability scores to produce reliable certainty metrics.
The evaluation yielded several major findings. First, Knowledge Vault extracted 271 million high-confidence facts with a predicted truth probability of at least 90%, making it approximately 38 times larger than previous comparable automatically constructed systems. Second, about one-third of these high-confidence facts represented entirely new knowledge not previously documented in Freebase. Third, fusing graph priors directly with web extractors increased the total yield of high-confidence facts from 100 million to 271 million, representing a 171% expansion while simultaneously shrinking uncertainty and filtering false positives. Fourth, probabilistic calibration proved highly accurate, ensuring that facts predicted at a 90% confidence level were true approximately 90% of the time. Finally, page structure extractors delivered the highest volume of high-confidence extractions among web sources, whereas tabular extractions yielded the least due to limited schema overlap.
These findings indicate that organizations can scalably build massive, reliable knowledge bases without solely depending on manual curation or clean infoboxes. Fusing prior graph relationships acts as an effective noise filter, dramatically lowering operational verification costs and reducing the risk of surfacing hallucinated or incorrect data in downstream applications. Furthermore, the framework's calibrated probability scores allow downstream systems to dynamically set confidence thresholds tailored to their specific risk and precision requirements.
To build upon this framework, future engineering should extend the model beyond independent binary assumptions to account for mutual exclusion constraints and numerical relationship correlations. Additionally, systems should incorporate sophisticated source-copy detection, temporal validity tracking for changing facts, automated discovery of unseen entity types, and richer ontological abstractions beyond standard relational triples. While the local closed world assumption used for training introduces minor label noise compared to human evaluation, the results remain robust and justify full-scale deployment for web-scale knowledge expansion.
- Paper: Reasoning With Neural Tensor Networks for Knowledge Base Completion, Richard Socher et al. (2013). Its Neural Tensor Network provides a concrete precursor to the neural link-prediction priors that Knowledge Vault fuses with web-extracted facts.
No sufficiently relevant recommendations were found.
