Knowledge vault: a web-scale approach to probabilistic knowledge fusion

Xin Luna DongEvgeniy GabrilovichGeremy HeitzWilko HornNi LaoKevin MurphyThomas StrohmannShaohua SunWei Zhang

article2014KDD1,843 citations

Proposes an automated web-scale framework that builds an extensive knowledge base by probabilistically fusing noisy web extractions across multiple data modalities with graph-based prior knowledge to accurately estimate fact correctness.

Listen

Large-scale knowledge repositories are critical for modern search and question-answering systems, but existing repositories remain highly incomplete due to their reliance on human curation, structured data, and text-only extraction methods. For example, major bases lack basic facts such as place of birth or nationality for over 70% of indexed people. Unsupervised web extraction can expand coverage, but traditional pipelines introduce massive noise and false positives. The article introduces and evaluates Knowledge Vault, an automated, web-scale framework designed to build a probabilistic knowledge base by fusing diverse web extractions with predictive graph priors learned from existing knowledge.

The research demonstrates that combining noisy multi-source web extractions with graph-based prior models significantly increases fact coverage while maintaining rigorous quality control. The approach extracts 1.6 billion candidate subject-predicate-object triples across 4,469 relation types using four web extraction streams: free text, page structure trees, relational web tables, and human page annotations. These extractions are combined with two link-prediction graph models—a path-ranking algorithm and a neural network multi-layer perceptron—trained on Freebase data. Supervised machine learning algorithms fuse these distinct evidence sources and calibrate probability scores to produce reliable certainty metrics.

The evaluation yielded several major findings. First, Knowledge Vault extracted 271 million high-confidence facts with a predicted truth probability of at least 90%, making it approximately 38 times larger than previous comparable automatically constructed systems. Second, about one-third of these high-confidence facts represented entirely new knowledge not previously documented in Freebase. Third, fusing graph priors directly with web extractors increased the total yield of high-confidence facts from 100 million to 271 million, representing a 171% expansion while simultaneously shrinking uncertainty and filtering false positives. Fourth, probabilistic calibration proved highly accurate, ensuring that facts predicted at a 90% confidence level were true approximately 90% of the time. Finally, page structure extractors delivered the highest volume of high-confidence extractions among web sources, whereas tabular extractions yielded the least due to limited schema overlap.

These findings indicate that organizations can scalably build massive, reliable knowledge bases without solely depending on manual curation or clean infoboxes. Fusing prior graph relationships acts as an effective noise filter, dramatically lowering operational verification costs and reducing the risk of surfacing hallucinated or incorrect data in downstream applications. Furthermore, the framework's calibrated probability scores allow downstream systems to dynamically set confidence thresholds tailored to their specific risk and precision requirements.

To build upon this framework, future engineering should extend the model beyond independent binary assumptions to account for mutual exclusion constraints and numerical relationship correlations. Additionally, systems should incorporate sophisticated source-copy detection, temporal validity tracking for changing facts, automated discovery of unseen entity types, and richer ontological abstractions beyond standard relational triples. While the local closed world assumption used for training introduces minor label noise compared to human evaluation, the results remain robust and justify full-scale deployment for web-scale knowledge expansion.

No sufficiently relevant recommendations were found.

Cover for Knowledge vault: a web-scale approach to probabilistic knowledge fusion

Abstract

Recent years have witnessed a proliferation of large-scale knowledge bases, including Wikipedia, Freebase, YAGO, Microsoft's Satori, and Google's Knowledge Graph. To increase the scale even further, we need to explore automatic methods for constructing knowledge bases. Previous approaches have primarily focused on text-based extraction, which can be very noisy. Here we introduce Knowledge Vault, a Web-scale probabilistic knowledge base that combines extractions from Web content (obtained via analysis of text, tabular data, page structure, and human annotations) with prior knowledge derived from existing knowledge repositories. We employ supervised machine learning methods for fusing these distinct information sources. The Knowledge Vault is substantially bigger than any previously published structured knowledge repository, and features a probabilistic inference system that computes calibrated probabilities of fact correctness. We report the results of multiple studies that explore the relative utility of the different information sources and extraction methods.

Table of Contents

  • 1. INTRODUCTION
  • 2. OVERVIEW
  • 2.1 Evaluation protocol
  • 2.2 Local closed world assumption (LCWA)
  • 3. FACT EXTRACTION FROM THE WEB
  • 3.1 Extraction methods
  • 3.1.1 Text documents (TXT)
  • 3.1.2 HTML trees (DOM)
  • 3.1.3 HTML tables (TBL)
  • 3.1.4 Human Annotated pages (ANO)
  • 3.2 Fusing the extractors
  • 3.3 Calibration of the probability estimates
  • 3.4 Comparison of the methods
  • 3.5 The beneficial effects of adding more evidence
  • 4. GRAPH-BASED PRIORS
  • 4.1 Path ranking algorithm (PRA)
  • 4.2 Neural network model (MLP)
  • 4.3 Fusing the priors
  • 5. FUSING EXTRACTORS AND PRIORS
  • 6. EVALUATING LCWA
  • 7. RELATED WORK
  • 8. DISCUSSION
  • 9. CONCLUSIONS
  • Acknowledgments
  • 10. REFERENCES

Knowls

  1. Knowl 1 — Probabilistic Knowledge Fusion of Web Extractors and Graph Priors

    model/method

    The Knowledge Vault (KV) constructs a Web-scale probabilistic knowledge base by fusing noisy extractions from Web sources with graph-based prior models trained on existing structured knowledge bases (such as Freebase).

    The task is formalized as predicting the probability Pr⁡(G(s,p,o)=1)\Pr(G(s, p, o) = 1) that a directed labeled edge of relation predicate p∈{1,…,P}p \in \{1, \dots, P\} exists between subject entity s∈{1,…,E}s \in \{1, \dots, E\} and object entity o∈{1,…,E}o \in \{1, \dots, E\} in a knowledge graph GG.

    The overall fusion pipeline operates in three stages:

    1. Extraction: Independent extractors process text documents, DOM trees, Web tables, and annotated HTML to propose candidate triples with source-level extraction confidence scores.
    2. Graph-based Priors: Statistical relational learning models (such as the Path Ranking Algorithm and multi-layer perceptron tensor embeddings) compute prior existence probabilities for triples conditioned on existing edges in the knowledge graph, excluding the target edge.
    3. Probabilistic Fusion: A supervised classifier fuses the multi-modal extraction signals and the prior predictions. For each relation predicate pp, a distinct model (such as boosted decision stumps) is trained on labeled triples, followed by Platt scaling to produce well-calibrated posterior probabilities Pr⁡(G(s,p,o)=1∣extractions,priors)\Pr(G(s, p, o) = 1 \mid \text{extractions}, \text{priors}).
  2. Knowl 2 — Local Closed World Assumption for Distant Supervision and Evaluation

    assumption

    In incomplete knowledge bases, unobserved triples cannot be assumed false under a standard closed-world assumption without introducing high false-negative rates. The Local Closed World Assumption (LCWA) provides a labeling heuristic to generate binary training and evaluation targets for candidate triples (s,p,o)(s, p, o).

    Let O(s,p)={o′∣(s,p,o′)∈KB}\mathcal{O}(s, p) = \{o' \mid (s, p, o') \in \text{KB}\} denote the set of known object values for a given subject ss and predicate pp in the existing knowledge base. For any candidate triple (s,p,o)(s, p, o):

    • If (s,p,o)∈O(s,p)(s, p, o) \in \mathcal{O}(s, p), the triple is labeled correct (y=1y = 1).
    • If (s,p,o)∉O(s,p)(s, p, o) \notin \mathcal{O}(s, p) and ∣O(s,p)∣>0|\mathcal{O}(s, p)| > 0, the knowledge base is assumed locally complete for the pair (s,p)(s, p), and the candidate triple is labeled incorrect (y=0y = 0).
    • If ∣O(s,p)∣=0|\mathcal{O}(s, p)| = 0, the triple is treated as unlabeled and excluded from the training and test sets.
  3. Knowl 3 — Multi-Layer Perceptron Neural Tensor Model for Knowledge Graph Link Prediction

    model/method

    To model graph-based priors for link prediction in knowledge graphs, entity and predicate tokens are embedded into a continuous latent semantic space. Given a candidate triple (s,p,o)(s, p, o), latent embedding vectors u⃗s∈RK\vec{u}_s \in \mathbb{R}^K, w⃗p∈RK\vec{w}_p \in \mathbb{R}^K, and v⃗o∈RK\vec{v}_o \in \mathbb{R}^K represent the subject entity, predicate relation, and object entity, respectively (with K=60K = 60).

    The prior probability of edge existence is predicted via a multi-layer perceptron (MLP) defined as:

    Pr⁡(G(s,p,o)=1)=σ(β⃗Tf(A[u⃗s,w⃗p,v⃗o]))\Pr(G(s, p, o) = 1) = \sigma\left(\vec{\beta}^T f\left(\mathbf{A} [\vec{u}_s, \vec{w}_p, \vec{v}_o]\right)\right)

    where:

    • [u⃗s,w⃗p,v⃗o]∈R3K[\vec{u}_s, \vec{w}_p, \vec{v}_o] \in \mathbb{R}^{3K} is the concatenation of the subject, predicate, and object embeddings.
    • A∈RL×3K\mathbf{A} \in \mathbb{R}^{L \times 3K} is the first-layer weight matrix mapping the concatenated embeddings to LL hidden units (with L=60L = 60).
    • f(⋅)f(\cdot) is an elementwise non-linear activation function (such as tanh⁡\tanh).
    • β⃗∈RL\vec{\beta} \in \mathbb{R}^L is the second-layer output weight vector.
    • σ(z)=11+e−z\sigma(z) = \frac{1}{1 + e^{-z}} is the logistic sigmoid function.

    This formulation requires O(L+LK+KE+KP)O(L + LK + KE + KP) parameters for EE entities and PP predicates, avoiding the O(KE+K2MP)O(KE + K^2 M P) parameter complexity of tensor-per-relation architectures while capturing non-linear interactions among entities and predicates.

  4. Knowl 4 — Heterogeneous Web Extraction Pipelines across Text, DOM, Tables, and Annotations

    model/method

    Knowledge Vault employs four distinct extraction pipelines to extract structured RDF triples (s,p,o)(s, p, o) from Web data:

    1. Text Documents (TXT): Uses an NLP pipeline (named entity recognition, part-of-speech tagging, dependency parsing, coreference resolution, and entity linking). Relation extractors are trained via distant supervision using seed pairs from Freebase. Features from surface text and dependency paths are used to train independent binary logistic regression classifiers per predicate in parallel via MapReduce.
    2. HTML DOM Trees (DOM): Extracts relations from semi-structured web pages and deep web form results. Feature vectors are constructed from the lexicalized paths in the HTML DOM tree between linked entity pairs, scored using supervised classifiers.
    3. HTML Tables (TBL): Performs entity linking on table cells, and applies schema matching heuristics to identify relations between columns based on Freebase predicate type constraints. Ambiguous columns are discarded, and extraction scores reflect entity linking confidence.
    4. Human Annotated Pages (ANO): Extracts Schema.org semantic markups embedded in web pages, maps them to corresponding Freebase predicates, and assigns confidence based on entity linking.
  5. Knowl 5 — Fact Extractor Fusion and Confidence Calibration

    model/method

    Extractions from heterogeneous extractors are aggregated into a single confidence score for each candidate triple t=(s,p,o)t = (s, p, o) through the following procedure:

    1. Feature Construction: For each extractor i∈{TXT,DOM,TBL,ANO}i \in \{\text{TXT}, \text{DOM}, \text{TBL}, \text{ANO}\}, a two-dimensional feature is computed: [ni,sˉi][\sqrt{n_i}, \bar{s}_i], where nin_i is the number of distinct domains (counting at most once per domain to mitigate web-source duplication/copying) containing tt, and sˉi\bar{s}_i is the mean extraction score across those sources (set to 0 if the extractor did not produce tt).
    2. Predicate-Specific Classification: Boosted decision stumps are trained per predicate on the combined 8-dimensional feature vector f⃗(t)\vec{f}(t) with labels assigned under the Local Closed World Assumption.
    3. Probability Calibration via Platt Scaling: Raw classification scores are converted into well-calibrated probabilities via Platt scaling, fitting a 1D logistic regression model on a holdout validation set such that a predicted probability pp corresponds to an empirical true positive rate of pp.
  6. Knowl 6 — Path Ranking Algorithm for Graph-Based Prior Link Prediction

    model/method

    The Path Ranking Algorithm (PRA) performs link prediction on the existing knowledge graph by learning path-based inference rules.

    For a target predicate pp:

    1. Random walks of bounded length are executed starting from subject nodes ss to reach object nodes oo.
    2. For each unique relation path sequence (e.g., X→parentOfZ←parentOfYX \xrightarrow{\text{parentOf}} Z \xleftarrow{\text{parentOf}} Y), the feature value for pair (s,o)(s, o) is the probability of reaching oo from ss via that path.
    3. A binary logistic regression classifier is trained independently per predicate over these path probability features using labels derived from the Local Closed World Assumption.

    At inference time, path probabilities along the selected paths are computed over the known knowledge graph (excluding the target edge) and evaluated by the logistic regression classifier to produce a prior probability for the existence of (s,p,o)(s, p, o).

  7. Knowl 7 — Scale Comparison of Knowledge Vault with Existing Knowledge Bases

    data/table

    Knowledge Vault (KV) substantially scales the number of confident extracted facts compared to prior extraction-based and human-curated knowledge bases. A confident fact is defined as an extracted triple with a calibrated correctness probability P≥0.9P \ge 0.9.

    Name # Entity types # Entity instances # Relation types # Confident facts
    Knowledge Vault (KV) 1,100 45M 4,469 271M
    DeepDive 4 2.7M 34 7M
    NELL 271 5.19M 306 0.435M
    PROSPERA 11 N/A 14 0.1M
    YAGO2 350,000 9.8M 100 4M
    Freebase 1,500 40M 35,000 637M
    Knowledge Graph (KG) 1,500 570M 35,000 18,000M

    KV, DeepDive, NELL, and PROSPERA rely strictly on automatic Web extraction; Freebase and KG rely on human curation and structured databases; YAGO2 combines both. Among extraction-based systems, KV yields 271M confident facts (approximately 38 times larger than DeepDive's 7M facts) across 4,469 relation types.

  8. Knowl 8 — Extraction Yield and Precision Across Heterogeneous Web Extractors

    data/table

    The performance and output volume of individual extraction modalities and their fusion applied to a web corpus demonstrate substantial variance in yield and precision:

    System # Triples # > 0.7 # > 0.9 Frac. > 0.9 AUC
    TBL 9.4M 3.8M 0.59M 0.06 0.856
    ANO 140M 2.4M 0.25M 0.002 0.920
    TXT 330M 20M 7.1M 0.02 0.867
    DOM 1,200M 150M 94M 0.08 0.928
    FUSED-EX. 1,600M 160M 100M 0.06 0.927

    DOM tree extraction generates the largest volume of candidate triples (1.2B) and confident triples (94M at P>0.9P > 0.9) with the highest individual AUC (0.928). HTML Web Tables (TBL) produce the fewest triples (9.4M) due to low schema match rates between table columns and Freebase predicates. Fusing all extractors yields 100M confident triples (P>0.9P > 0.9), achieving an overall extractor AUC of 0.927.

  9. Knowl 9 — Empirical Performance Gains from Fusing Extractors with Graph Priors

    empirical result

    Combining graph priors with Web extractions provides substantial empirical improvements in ranking quality, confidence, and knowledge base yield:

    • Classification AUC: On a balanced test set evaluated under the Local Closed World Assumption, fused extractors achieve an AUC of 0.927, fused priors achieve an AUC of 0.911 (individual PRA achieves 0.884, MLP achieves 0.882), and joint fusion (extractors + priors) achieves an AUC of 0.947.
    • Yield of Confident Facts: Integrating priors increases the number of high-confidence triples (P≥0.9P \ge 0.9) from 100M (extractors alone) to 271M (extractors + priors), a 2.7-fold increase. Approximately 33% of these 271M confident triples were novel facts not present in the original Freebase graph.
    • Uncertainty Reduction: Joint fusion sharpens the probability distribution across triples by reducing the proportion of candidate facts falling into the ambiguous probability window [0.3,0.7][0.3, 0.7], pushing true facts closer to 1.01.0 and filtering out extraction errors by driving false positives close to 0.00.0.
  10. Knowl 10 — Evaluation of Local Closed World Assumption Against Human Judgment

    data/table

    To evaluate the validity of the Local Closed World Assumption (LCWA) as a proxy for ground truth, an in-house evaluation was conducted on a sample of 1,000 triples across 10 predicates from the balanced test set (yielding 695 verified true/false triples after discarding 305 unknown ratings):

    Evaluation Labels Fused Prior AUC Fused Extractor AUC Prior + Extractor AUC
    LCWA Labels 0.943 0.872 0.959
    Human Labels 0.843 0.852 0.869

    While models trained and tested on LCWA show higher absolute AUC scores due to local completeness assumptions, the relative performance trends across the prior, extractor, and fused models remain consistent when evaluated against human-annotated ground truth, supporting the use of LCWA for large-scale training.

  11. Knowl 11 — Failure of Strict Mutual Exclusion Under Multi-Granularity Mentions

    limitation

    Enforcing strict mutual exclusion constraints on single-valued/functional relations (e.g., forcing ∑oiPr⁡(s,p,oi)=1\sum_{o_i} \Pr(s, p, o_i) = 1 for a subject ss and functional predicate pp) fails in web-scale extraction due to multi-granularity entity representations.

    Extracted object values often refer to valid geographical or conceptual abstractions of the same underlying entity at different levels of granularity (e.g., Barack Obama born in Honolulu, Hawaii, and USA). Treating distinct candidate values as mutually exclusive causes valid, non-conflicting extractions to penalize one another's probability mass unless hierarchical subsumption and granularity compatibility are explicitly modeled.

Coverage note — Deliberately omitted qualitative nearest-neighbor lists for specific predicates in embedding space and exploratory discussions of soft numerical Gaussian priors, as they serve as illustrative examples rather than standalone core contributions.

References

  1. 1.AKBC-WEKEX. The Knowledge Extraction Workshop at NAACL-HLT, 2012.
  2. 2.G. Angeli and C. Manning. Philosophers are mortal: Inferring the truth of unseen facts. In CoNLL, 2013.
  3. 3.S. Auer, C. Bizer, G. Kobilarov, J. Lehmann, R. Cyganiak, and Z. Ives. DBpedia: A nucleus for a web of open data. In The semantic web, pages 722–735, 2007.
  4. 4.K. Bollacker, C. Evans, P. Paritosh, T. Sturge, and J. Taylor. Freebase: a collaboratively created graph database for structuring human knowledge. In SIGMOD, pages 1247–1250. ACM, 2008.
  5. 5.A. Bordes, X. Glorot, J. Weston, and Y. Bengio. Joint learning of words and meaning representations for open-text semantic parsing. In AI/Statistics, 2012.
  6. 6.M. Cafarella, A. Halevy, Z. D. Wang, E. Wu, and Y. Zhang. WebTables: Exploring the Power of Tables on the Web. VLDB, 1(1):538–549, 2008.
  7. 7.M. J. Cafarella, A. Y. Halevy, and J. Madhavan. Structured data on the web. Commun. ACM, 54(2):72–79, 2011.
  8. 8.A. Carlson, J. Betteridge, B. Kisiel, B. Settles, E. H. Jr., and T. Mitchell. Toward an architecture for never-ending language learning. In AAAI, 2010.
  9. 9.O. Deshpande, D. Lambda, M. Tourn, S. Das, S. Subramaniam, A. Rajaraman, V. Harinarayan, and A. Doan. Building, maintaing and using knowledge bases: A report from the trenches. In SIGMOD, 2013.
  10. 10.X. L. Dong, L. Berti-Equille, and D. Srivastatva. Integrating conflicting data: the role of source dependence. In VLDB, 2009.
  11. 11.L. Drumond, S. Rendle, and L. Schmidt-Thieme. Predicting RDF Triples in Incomplete Knowledge Bases with Tensor Factorization. In 10th ACM Intl. Symp. on Applied Computing, 2012.
  12. 12.A. Fader, S. Soderland, and O. Etzioni. Identifying relations for open information extraction. In EMNLP, 2011.
  13. 13.J. Fan, D. Ferrucci, D. Gondek, and A. Kalyanpur. Prismatic: Inducing knowledge from a large scale lexicalized relation resource. In First Intl. Workshop on Formalisms and Methodology for Learning by Reading, pages 122–127. Association for Computational Linguistics, 2010.
  14. 14.T. Franz, A. Schultz, S. Sizov, and S. Staab. TripleRank: Ranking Semantic Web Data by Tensor Decomposition. In ISWC, 2009.
  15. 15.L. A. Gal´arraga, C. Teflioudi, K. Hose, and F. Suchanek. Amie: association rule mining under incomplete evidence in ontological knowledge bases. In WWW, pages 413–422, 2013.
  16. 16.R. Grishman. Information extraction: Capabilities and challenges. Technical report, NYU Dept. CS, 2012.
  17. 17.R. Gupta, A. Halevy, X. Wang, S. Whang, and F. Wu. Biperpedia: An Ontology for Search Applications. In VLDB, 2014.
  18. 18.B. Hachey, W. Radford, J. Nothman, M. Honnibal, and J. Curran. Evaluating entity linking with wikipedia. Artificial Intelligence, 194:130–150, 2013.
  19. 19.J. Hoffart, F. M. Suchanek, K. Berberich, and G. Weikum. YAGO2: A Spatially and Temporally Enhanced Knowledge Base from Wikipedia. Artificial Intelligence Journal, 2012.
  20. 20.R. Jenatton, N. L. Roux, A. Bordes, and G. Obozinski. A latent factor model for highly multi-relational data. In NIPS, 2012.
  21. 21.H. Ji, T. Cassidy, Q. Li, and S. Tamang. Tackling Representation, Annotation and Classification Challenges for Temporal Knowledge Base Population. Knowledge and Information Systems, pages 1–36, August 2013.
  22. 22.H. Ji and R. Grishman. Knowledge base population: successful approaches and challenges. In Proc. ACL, 2011.
  23. 23.S. Jiang, D. Lowd, and D. Dou. Learning to refine an automatically extracted knowledge base using markov logic. In Intl. Conf. on Data Mining, 2012.
  24. 24.N. Lao, T. Mitchell, and W. Cohen. Random walk inference and learning in a large scale knowledge base. In EMNLP, 2011.
  25. 25.X. Li and R. Grishman. Confidence estimation for knowledge base population. In Recent Advances in NLP, 2013.
  26. 26.Mausam, M. Schmitz, R. Bart, S. Soderland, and O. Etzioni. Open language learning for information extraction. In EMNLP, 2012.
  27. 27.T. Mikolov, K. Chen, G. Corrado, and J. Dean. Efficient estimation of word representations in vector space. In ICLR, 2013.
  28. 28.B. Min, R. Grishman, L. Wan, C. Wang, and D. Gondek. Distant supervision for relation extraction with an incomplete knowledge base. In NAACL, 2013.
  29. 29.M. Mintz, S. Bills, R. Snow, and D. Jurafksy. Distant supervision for relation extraction without labeled data. In Prof. Conf. Recent Advances in NLP, 2009.
  30. 30.N. Nakashole, M. Theobald, and G. Weikum. Scalable knowledge harvesting with high precision and high recall. In WSDM, pages 227–236, 2011.
  31. 31.M. Nickel, V. Tresp, and H.-P. Kriegel. Factorizing YAGO: scalable machine learning for linked data. In WWW, 2012.
  32. 32.F. Niu, C. Zhang, and C. Re. Elementary: Large-scale Knowledge-base Construction via Machine Learning and Statistical Inference. Intl. J. On Semantic Web and Information Systems, 2012.
  33. 33.J. Platt. Probabilities for SV machines. In A. Smola, P. Bartlett, B. Schoelkopf, and D. Schuurmans, editors, Advances in Large Margin Classifiers. MIT Press, 2000.
  34. 34.J. Pujara, H. Miao, L. Getoor, and W. Cohen. Knowledge graph identification. In International Semantic Web Conference (ISWC), 2013.
  35. 35.L. Reyzin and R. Schapire. How boosting the margin can also boost classifier complexity. In Intl. Conf. on Machine Learning, 2006.
  36. 36.A. Ritter, L. Zettlemoyer, Mausam, and O. Etzioni. Modeling missing data in distant supervision for information extraction. Trans. Assoc. Comp. Linguistics, 1, 2013.
  37. 37.R. Socher, D. Chen, C. Manning, and A. Ng. Reasoning with Neural Tensor Networks for Knowledge Base Completion. In NIPS, 2013.
  38. 38.R. Speer and C. Havasi. Representing general relational knowledge in conceptnet 5. In Proc. of LREC Conference, 2012.
  39. 39.F. Suchanek, G. Kasneci, and G. Weikum. YAGO - A Core of Semantic Knowledge. In WWW, 2007.
  40. 40.D. Suciu, D. Olteanu, C. Re, and C. Koch. Probabilistic Databases. Morgan & Claypool, 2011.
  41. 41.B. Suh, G. Convertino, E. H. Chi, and P. Pirolli. The singularity is not near: slowing growth of wikipedia. In Proceedings of the 5th International Symposium on Wikis and Open Collaboration, WikiSym ’09, pages 8:1–8:10, 2009.
  42. 42.P. Venetis, A. Halevy, J. Madhavan, M. Pasca, W. Shen, F. Wu, G. Miao, and C. Wi. Recovering semantics of tables on the web. In Proc. of the VLDB Endowment, 2012.
  43. 43.D. Z. Wang, E. Michelakis, M. Garofalakis, and J. Hellerstein. BayesStore: Managing Large, Uncertain Data Repositories with Probabilistic Graphical Models. In VLDB, 2008.
  44. 44.G. Weikum and M. Theobald. From information to knowledge: harvesting entities and relationships from web sources. In Proceedings of the twenty-ninth ACM SIGMOD-SIGACT-SIGART symposium on Principles of database systems, pages 65–76. ACM, 2010.
  45. 45.M. Wick, S. Singh, A. Kobren, and A. McCallum. Assessing confidence of knowledge base content with an experimental study in entity resolution. In AKBC workshop, 2013.
  46. 46.M. Wick, S. Singh, H. Pandya, and A. McCallum. A Joint Model for Discovering and Linking Entities. In AKBC Workshop, 2013.
  47. 47.W. Wu, H. Li, H. Wang, and K. Q. Zhu. Probase: A probabilistic taxonomy for text understanding. In SIGMOD, pages 481–492. ACM, 2012.
  48. 48.Z. Xu, V. Tresp, K. Yu, and H.-P. Kriegel. Infinite hidden relational models. In UAI, 2006.

Citation

MLA
Dong, X., et al. “Knowledge Vault”. Proceedings of the 20th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2014, pp. 601–10, https://doi.org/10.1145/2623330.2623623.
APA
Dong, X., Gabrilovich, E., Heitz, G., Horn, W., Lao, N., Murphy, K., Strohmann, T., Sun, S., & Zhang, W. (2014). Knowledge vault. Proceedings of the 20th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 601–610. https://doi.org/10.1145/2623330.2623623
Chicago
Dong, X., E. Gabrilovich, G. Heitz, et al. 2014. “Knowledge Vault”. Proceedings of the 20th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 601–10. https://doi.org/10.1145/2623330.2623623.
Harvard
Dong, X. et al. (2014) “Knowledge vault”, Proceedings of the 20th ACM SIGKDD international conference on Knowledge discovery and data mining. ACM, pp. 601–610. Available at: https://doi.org/10.1145/2623330.2623623.
Vancouver
1. Dong X, Gabrilovich E, Heitz G, Horn W, Lao N, Murphy K, Strohmann T, Sun S, Zhang W (2014) Knowledge vault. In: Proceedings of the 20th ACM SIGKDD international conference on Knowledge discovery and data mining. ACM, pp 601–610

BibTeX

@inproceedings{Dong_2014, series={KDD ’14}, title={Knowledge vault: a web-scale approach to probabilistic knowledge fusion}, url={http://dx.doi.org/10.1145/2623330.2623623}, DOI={10.1145/2623330.2623623}, booktitle={Proceedings of the 20th ACM SIGKDD international conference on Knowledge discovery and data mining}, publisher={ACM}, author={Dong, Xin and Gabrilovich, Evgeniy and Heitz, Geremy and Horn, Wilko and Lao, Ni and Murphy, Kevin and Strohmann, Thomas and Sun, Shaohua and Zhang, Wei}, year={2014}, month=Aug, pages={601–610}, collection={KDD ’14} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF