NodePiece: Compositional and Parameter-Efficient Representations of Large Knowledge Graphs

Mikhail GalkinEtienne G. DenisJiapeng WuWilliam L. Hamilton

article2022ICLR114 citations

Introduces NodePiece, an anchor-based compositional tokenization framework that reduces knowledge graph parameter footprints by up to 70x and generalizes to unseen entities while matching or outperforming standard models on large-scale benchmarks.

Listen

Knowledge graphs organize vast amounts of interconnected real-world information, serving as core infrastructure for search engines, recommendation systems, and enterprise data discovery. However, standard methods for learning graph representations assign an independent mathematical embedding vector to every individual entity. As enterprise graphs expand to millions or billions of entities, this traditional shallow approach causes memory consumption and hardware costs to grow unsustainably. Furthermore, conventional models struggle to generalize to newly arriving entities without undergoing expensive, full-scale retraining.

The article introduces and evaluates NodePiece, a compositional and parameter-efficient representation method designed to scale knowledge graph modeling. The main objective of the article is to demonstrate that a graph can be effectively represented using a small, fixed-size vocabulary of sub-graph units, thereby drastically reducing parameter counts while maintaining competitive accuracy across multiple graph reasoning tasks.

Inspired by subword tokenization in modern language models, the approach builds a compact vocabulary comprising a small fraction of selected anchor nodes alongside all known relation types. Instead of storing an individual embedding for every node, each entity is represented as a structured sequence consisting of its nearest anchor nodes, its topological distance to those anchors, and its immediate outgoing relation types. A lightweight neural encoder, such as a multi-layer perceptron or a transformer, then processes this sequence to generate the final entity representation. The authors evaluated this framework across several established benchmarks, including link prediction, relation prediction, out-of-sample inference, and node classification on graphs containing up to 2.5 million entities and 17 million connections.

The findings show that NodePiece achieves substantial memory efficiency without significant performance degradation. First, on a large-scale Wikidata benchmark with 2.5 million nodes, NodePiece outperformed traditional high-performing shallow embedding models while requiring approximately 70 times fewer parameters and operating within standard single-GPU hardware limits. Second, across standard transductive link prediction benchmarks, the model retained 80% to 90% of state-of-the-art accuracy while using less than 10% of total nodes as anchors and roughly 10 times fewer parameters overall. Third, in semi-supervised node classification, the compositional approach improved hard exact-match accuracy threefold over baseline graph neural networks, demonstrating that smaller parameter footprints can prevent severe model overfitting. Finally, the framework successfully performed inference on completely unseen entities and disjoint graphs without requiring task-specific architectural modifications.

These results demonstrate that massive, dedicated node embedding tables are often unnecessary for effective graph representation. By shifting the computational burden from linear memory storage to fixed-size vocabularies paired with expressive neural encoders, organizations can substantially reduce infrastructure costs, minimize GPU memory requirements, and deploy models on larger graphs using standard hardware. The architecture also naturally accommodates dynamic enterprise data streams, allowing immediate representation of newly created users, products, or entities without costly retraining pipelines.

Organizations operating large-scale graph machine learning systems should consider piloting anchor-based compositional tokenization as a drop-in replacement for traditional shallow embedding lookups, particularly in memory-constrained deployment environments. For relation-rich graphs, teams can leverage relational context to maintain accuracy even under aggressive vocabulary compression. Before broad operational rollout, engineering teams should evaluate graph density and structural connectivity, as highly sparse graphs with few relation types may require larger anchor budgets or tuned distance encodings to maintain fine-grained precision.

The primary limitations noted in the article center on sparse or highly regular graphs, where small anchor counts can reduce top-tier precision (such as exact top-one ranking metrics) or increase the likelihood of hash collisions between neighboring entities. Nevertheless, the experimental results demonstrate high reliability across diverse benchmarks, confirming that compositional tokenization offers a robust, scalable alternative for enterprise knowledge graph learning.

No sufficiently relevant recommendations were found.

Cover for NodePiece: Compositional and Parameter-Efficient Representations of Large Knowledge Graphs

Abstract

Conventional representation learning algorithms for knowledge graphs (KG) map each entity to a unique embedding vector. Such a shallow lookup results in a linear growth of memory consumption for storing the embedding matrix and incurs high computational costs when working with real-world KGs. Drawing parallels with subword tokenization commonly used in NLP, we explore the landscape of more parameter-efficient node embedding strategies with possibly sublinear memory requirements. To this end, we propose NodePiece, an anchor-based approach to learn a fixed-size entity vocabulary. In NodePiece, a vocabulary of subword/sub-entity units is constructed from anchor nodes in a graph with known relation types. Given such a fixed-size vocabulary, it is possible to bootstrap an encoding and embedding for any entity, including those unseen during training. Experiments show that NodePiece performs competitively in node classification, link prediction, and relation prediction tasks while retaining less than 10% of explicit nodes in a graph as anchors and often having 10x fewer parameters. To this end, we show that a NodePiece-enabled model outperforms existing shallow models on a large OGB WikiKG 2 graph having 70x fewer parameters.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 NodePiece Vocabulary Construction
  • 3.1 Anchor Selection
  • 3.2 Node Tokenization
  • 3.3 Encoding
  • 4 Experiments
  • 4.1 Transductive Link Prediction
  • 4.1.1 OGB WikiKG 2
  • 4.2 Inductive Link Prediction
  • 4.3 Node Classification
  • 4.4 Out-of-sample Link Prediction
  • 5 Conclusion
  • References
  • A Implementation & Hyperparameters
  • A.1 Datasets
  • A.2 Transductive Link Prediction
  • A.3 Relation Prediction
  • A.4 Node Classification
  • A.5 Out-of-sample Link Prediction
  • A.6 Deployment in Real-world Dynamic Knowledge Graphs
  • B Limitations and Future Work
  • C Anchor Selection Strategies
  • D Embedding Visualizations
  • E Transductive Link Prediction Results: MRR
  • F Datasets Construction
  • F.1 Node Classification: WD50K NC
  • F.2 Ouf-of-sample Link Prediction: oYAGO 3-10
  • G Node Classification: Training Curves
  • H Proofs
  • I Relation Prediction
  • J Inductive Link Prediction
  • K Uniqueness of Node Hashes

Knowls

  1. Knowl 1 — NodePiece Entity Tokenization Scheme

    model/method

    NodePiece represents entities in a multi-relational directed knowledge graph G=(N,E,R)G = (N, E, R) using a fixed-size vocabulary V=A∪RV = A ∪ R, where A⊂NA \subset N is a pre-selected subset of anchor nodes (∣A∣≪∣N∣|A| \ll |N|) and R=Rdirect∪RinverseR = R_{\text{direct}} \cup R_{\text{inverse}} is the set of all direct and inverse relation types (∣R∣≪∣N∣|R| \ll |N|).

    To tokenize any target node n∈Nn \in N, the algorithm constructs a discrete token sequence hash(n)\text{hash}(n) of fixed capacity:

    hash(n)=({ai}i=1k,{zai}i=1k,{rj}j=1m)\text{hash}(n) = \left( \{a_i\}_{i=1}^k, \{z_{a_i}\}_{i=1}^k, \{r_j\}_{j=1}^m \right)

    where:

    • {ai}i=1k⊆A\{a_i\}_{i=1}^k \subseteq A is a sequence of kk anchor nodes (e.g., the kk nearest anchors found via breadth-first search in the ll-hop neighborhood of nn).
    • {zai}i=1k⊂{0,1,…,diameter(G)}\{z_{a_i}\}_{i=1}^k \subset \{0, 1, \dots, \text{diameter}(G)\} are discrete shortest-path hop distances in GG from anchor aia_i to nn. When tokenizing an anchor node aj∈Aa_j \in A, its nearest anchor is aja_j itself with distance zaj=0z_{a_j} = 0.
    • {rj}j=1m⊆Nr(n)\{r_j\}_{j=1}^m \subseteq N_r(n) is a relational context of mm immediate unique outgoing relation types sampled from the set of relation types leaving nn, denoted Nr(n)N_r(n). If ∣Nr(n)∣<m|N_r(n)| < m, the sequence is padded to length mm with an auxiliary [PAD]\text{[PAD]} token.
    • If a node resides in a disconnected component, it is assigned an auxiliary [DISCONNECTED]\text{[DISCONNECTED]} token or converted into an anchor.
  2. Knowl 2 — NodePiece Vectorized Hash Representation and Encoders

    model/method

    Given the tokenized sequence hash(n)=({ai}i=1k,{zai}i=1k,{rj}j=1m)\text{hash}(n) = (\{a_i\}_{i=1}^k, \{z_{a_i}\}_{i=1}^k, \{r_j\}_{j=1}^m), NodePiece maps each atom to a dd-dimensional embedding using learnable vocabulary matrices:

    • Anchor embeddings an∈Rk×d\mathbf{a}_n \in \mathbb{R}^{k \times d} from anchor embedding matrix EA∈R∣A∣×d\mathbf{E}_A \in \mathbb{R}^{|A| \times d}.
    • Relation embeddings rn∈Rm×d\mathbf{r}_n \in \mathbb{R}^{m \times d} from relation embedding matrix ER∈R∣R∣×d\mathbf{E}_R \in \mathbb{R}^{|R| \times d}.
    • Discrete distance embeddings zan∈Rk×d\mathbf{z}_{a_n} \in \mathbb{R}^{k \times d} mapped via a learnable lookup matrix Z∈R(diameter(G)+1)×d\mathbf{Z} \in \mathbb{R}^{(\text{diameter}(G) + 1) \times d}.

    The distance embeddings serve as positional encodings and are added elementwise to the corresponding anchor embeddings, forming distance-augmented anchor representations a^n=an+zan\hat{\mathbf{a}}_n = \mathbf{a}_n + \mathbf{z}_{a_n}. Concatenating these with the relational context produces the vectorized hash:

    h(n)=[a^n,rn]=[an+zan,rn]∈R(k+m)×d\mathbf{h}(n) = [\hat{\mathbf{a}}_n, \mathbf{r}_n] = [\mathbf{a}_n + \mathbf{z}_{a_n}, \mathbf{r}_n] \in \mathbb{R}^{(k+m) \times d}

    A parameterized encoder function enc:R(k+m)×d→Rd\text{enc}: \mathbb{R}^{(k+m) \times d} \to \mathbb{R}^d computes the final embedding for node nn:

    1. MLP Encoder: Flattens h(n)\mathbf{h}(n) to a vector in R1×(k+m)d\mathbb{R}^{1 \times (k+m)d} and passes it through a 2-layer Multi-Layer Perceptron to project it to Rd\mathbb{R}^d.
    2. Transformer Encoder: Processes h(n)\mathbf{h}(n) as a sequence of k+mk+m tokens using multi-head self-attention layers followed by average pooling across sequence tokens to produce Rd\mathbb{R}^d.
  3. Knowl 3 — Janossy Pooling Equivalence of Nearest-Anchor Encoding

    theoretical result

    Let H∪=A∪RH^\cup = A \cup R denote the union of all anchor and relation tokens. A permutation-invariant Janossy function f:N×H∪×Rd→Ff: N \times H^\cup \times \mathbb{R}^d \to F constructed from a permutation-sensitive base function f∗f^* parameterized by θ(f)∈Rd\theta^{(f)} \in \mathbb{R}^d on a sequence h\mathbf{h} is defined by:

    f(∣h∣,h;θ(f))=1∣h∣!∑π∈Π∣h∣f∗(∣h∣,hπ;θ(f))f(|\mathbf{h}|, \mathbf{h}; \theta^{(f)}) = \frac{1}{|\mathbf{h}|!} \sum_{\pi \in \Pi_{|\mathbf{h}|}} f^*(|\mathbf{h}|, \mathbf{h}_\pi; \theta^{(f)})

    where Π∣h∣\Pi_{|\mathbf{h}|} is the symmetric group of permutations of length ∣h∣|\mathbf{h}|.

    A kk-ary Janossy function truncates the sequence using a projection ↓k(h)\downarrow_k(\mathbf{h}) to the first kk elements:

    f(∣h∣,h;θ(f))=1∣h∣!∑π∈Π∣h∣f∗(∣h∣,↓k(hπ);θ(f))f(|\mathbf{h}|, \mathbf{h}; \theta^{(f)}) = \frac{1}{|\mathbf{h}|!} \sum_{\pi \in \Pi_{|\mathbf{h}|}} f^*(|\mathbf{h}|, \downarrow_k(\mathbf{h}_\pi); \theta^{(f)})

    Proposition: The nearest-anchor encoder with (∣A∣k)\binom{|A|}{k} anchors and ∣m∣|m| subsampled relations is a π\pi-SGD approximation of (k+∣m∣)(k + |m|)-ary Janossy pooling with a canonical ordering induced by shortest-path anchor distances.

    Because equidistant anchors break absolute canonical ordering uniqueness, random ordering of equidistant anchors during training acts as uniform permutation sampling (π\pi-SGD), enabling permutation-invariant node representations to be learned with standard permutation-sensitive architectures like MLPs.

  4. Knowl 4 — Anchor Selection Policies and Distance Distribution Dynamics

    model/method

    NodePiece supports stochastic and deterministic anchor selection strategies to construct the anchor vocabulary A⊂NA \subset N:

    1. Random Anchor Selection: Selects ∣A∣|A| nodes uniformly at random from NN. To guarantee that the number of unique combinations of size kk exceeds the total node count ∣N∣|N|, parameters must satisfy the combinatorial bound:

    (∣A∣k)≥∣N∣\binom{|A|}{k} \ge |N|

    For example, (5020)≈4.7×1013\binom{50}{20} \approx 4.7 \times 10^{13} possible combinations, which is larger than the entity count of open knowledge graphs.

    1. Deterministic Mix Strategy: For downstream benchmarks, a non-overlapping split is defined where:
    • 40% of anchors are nodes with the highest Personalized PageRank (PPR) scores.
    • 40% of anchors are nodes with the highest vertex degree.
    • 20% of anchors are sampled uniformly at random.

    Anchor nodes selected by higher-priority criteria (PPR) are skipped if encountered in lower-priority criteria (degree).

    Centrality-based selections (PPR, degree, mixed) significantly skew the shortest-path anchor distance distribution towards smaller hop values compared to purely uniform random selection, increasing the proportion of anchors found within 2 to 3 hops of target nodes.

  5. Knowl 5 — Large-Scale Transductive Link Prediction on OGB WikiKG 2

    data/table

    On the OGB WikiKG 2 benchmark (2,500,604 entities, 535 relation types, 17,137,181 directed edges), NodePiece selects a vocabulary of 20,000 anchor nodes (<1%<1\% of all entities) with k=20k=20 nearest anchors and relational context size m=12m=12. A 2-layer MLP encoder is paired with an AutoSF scoring decoder.

    Model #Params MRR
    NodePiece + AutoSF 6.9M 0.570±0.0030.570 \pm 0.003
    - rel. context 5.9M 0.592±0.0030.592 \pm 0.003
    - anc. dists 6.9M 0.570±0.0040.570 \pm 0.004
    - no anchors (rels only) 1.3M 0.476±0.0010.476 \pm 0.001
    AutoSF 500M 0.546±0.0050.546 \pm 0.005
    PairRE 500M 0.521±0.0030.521 \pm 0.003
    RotatE 1250M 0.433±0.0020.433 \pm 0.002
    TransE 1250M 0.426±0.0030.426 \pm 0.003

    NodePiece + AutoSF achieves a higher Mean Reciprocal Rank (MRR 0.5700.570) than full-vocabulary AutoSF (0.5460.546) and RotatE (0.4330.433) while utilizing approximately 72×72\times and 180×180\times fewer parameters, respectively. In the extreme ablation using zero anchors (only relational context), NodePiece maintains 1.3M parameters and an MRR of 0.4760.476, outperforming full-table RotatE and TransE.

  6. Knowl 6 — Inductive Link Prediction on Disjoint Knowledge Graphs

    data/table

    In the inductive link prediction benchmark where training and inference graphs have disjoint entity sets (Vtrain∩Vtest=∅V_{\text{train}} \cap V_{\text{test}} = \emptyset) but share relation types, learned anchor nodes from training cannot transfer to test graphs. NodePiece operates in an anchor-free mode (k=0k=0), tokenizing every node solely by its outgoing relational context sequence of size mm.

    NodePiece token vectors are passed to a CompGCN graph neural network and scored using the RotatE function. Performance is evaluated using filtered Hits@10 ranked against 50 negative samples per triple across 4 splits of FB15k-237, WN18RR, and NELL-995:

    FB15k-237 WN18RR NELL-995
    Class Method V1 V2 V3 V4 V1 V2 V3 V4 V1 V2 V3 V4
    Path Neural LP 0.529 0.589 0.529 0.559 0.744 0.689 0.462 0.671 0.408 0.787 0.827 0.806
    DRUM 0.529 0.587 0.529 0.559 0.744 0.689 0.462 0.671 0.194 0.786 0.827 0.806
    RuleN 0.498 0.778 0.877 0.856 0.809 0.782 0.534 0.716 0.535 0.818 0.773 0.614
    GNN GraIL 0.642 0.818 0.828 0.893 0.825 0.787 0.584 0.734 0.595 0.933 0.914 0.732
    NBFNet 0.834 0.949 0.951 0.960 0.948 0.905 0.893 0.890 - - - -
    NP + CompGCN 0.873 0.939 0.944 0.949 0.830 0.886 0.785 0.807 0.890 0.901 0.936 0.893

    On relation-rich graphs (FB15k-237 and NELL-995), anchor-free NodePiece + CompGCN outperforms GraIL across splits (e.g., 0.8730.873 vs. 0.6420.642 on FB15k-237 V1) and performs competitively with NBFNet, demonstrating that relational context alone provides a strong basis for inductive representation.

  7. Knowl 7 — Semi-Supervised Node Classification Generalization on WD50K NC

    data/table

    On the WD50K NC dataset (46,164 entities, 526 relation types, 222,563 triples, 465 class labels extracted from Wikidata), models are evaluated in transductive semi-supervised settings with 5% and 10% labeled training nodes. NodePiece employs 50 anchors total (∣A∣=50|A|=50), k=10k=10 nearest anchors per node, and relational context m=5m=5, feeding bootstrapped embeddings into a 3-layer CompGCN.

    WD50K (5% labeled) WD50K (10% labeled)
    Model ∣V∣|V| #P (M) ROC-AUC PRC-AUC Hard Acc ROC-AUC PRC-AUC Hard Acc
    MLP 46k + 1k 4.1 0.503 0.016 0.001 0.510 0.017 0.002
    CompGCN 46k + 1k 4.4 0.836 0.280 0.176 0.834 0.265 0.161
    NodePiece + GNN 50 + 1k 0.75 0.981 0.443 0.513 0.981 0.450 0.516
    - no rel. context 50 + 1k 0.64 0.982 0.446 0.534 0.982 0.449 0.530
    - no distances 50 + 1k 0.74 0.981 0.448 0.516 0.981 0.448 0.513
    - no anchors (rels only) 0 + 1k 0.54 0.984 0.453 0.532 0.984 0.456 0.533

    Full-lookup CompGCN severely overfits due to learning 46k individual entity vectors on limited labeled examples. NodePiece reduces the parameter count from 4.4M to 0.75M and improves Hard Accuracy from 0.1760.176 to 0.5130.513 (3x gain) and PRC-AUC from 0.2800.280 to 0.4430.443 at 5% labeling, displaying near-zero generalization gap between training and validation curves.

  8. Knowl 8 — Out-of-Sample Link Prediction on Dynamic Graph Splits

    data/table

    In out-of-sample link prediction, new entities not present during training arrive at validation/test time connected to seen nodes via a few edges. NodePiece uses a 2-layer Transformer encoder to tokenize unseen nodes on the fly via the learned anchor and relation vocabulary VV, scoring triples with DistMult.

    oFB15k-237 oYAGO 3-10 (117k)
    Model ∣V∣|V| #P (M) MRR H@10 ∣V∣|V| #P (M) MRR H@10
    oDistMult-ERAvg 11k + 0.5k 2.4 0.256 0.420 117k + 74 23.4 OOM OOM
    NodePiece + DistMult 1k + 0.5k 1.0 0.206 0.372 10k + 74 2.7 0.133 0.261
    - no rel. context 1k + 0.5k 1.0 0.173 0.329 10k + 74 2.7 0.125 0.245
    - no distances 1k + 0.5k 1.0 0.208 0.372 10k + 74 2.7 0.133 0.260
    - no anchors (rels only) 0 + 0.5k 0.8 0.069 0.127 0 + 74 0.7 0.015 0.017

    NodePiece retains 88%88\% of the Hits@10 score of oDistMult-ERAvg on oFB15k-237 with less than half the parameters. On the larger oYAGO 3-10 split (117k entities), oDistMult fails with an out-of-memory error on a 256 GB RAM host, whereas NodePiece trains in 40 epochs with 2.7M parameters.

  9. Knowl 9 — Transductive Link Prediction Benchmark Comparison

    data/table

    Transductive link prediction benchmarks evaluated across FB15k-237, WN18RR, CoDEx-L, and YAGO 3-10 compare full shallow RotatE embedding lookups against NodePiece + RotatE (using a 2-layer MLP encoder and ≤10%\le 10\% total nodes as anchors):

    Dataset Model ∣V∣|V| #P (M) MRR Hits@10 % H@10
    FB15k-237 RotatE 15k + 0.5k 29 0.338 0.533 100%
    NodePiece + RotatE 1k + 0.5k 3.2 0.256 0.420 79%
    WN18RR RotatE 40k + 22 41 0.476 0.571 100%
    NodePiece + RotatE 500 + 22 4.4 0.403 0.515 90%
    CoDEx-L RotatE (500d) 77k + 138 77 0.258 0.387 100%
    RotatE (20d) 77k + 138 3.8 0.196 0.322 83%
    NodePiece + RotatE 7k + 138 3.6 0.190 0.313 81%
    YAGO 3-10 RotatE (500d) 123k + 74 123 0.495 0.670 100%
    RotatE (20d) 123k + 74 4.8 0.121 0.262 39%
    NodePiece + RotatE 10k + 74 4.1 0.247 0.488 73%

    NodePiece sustains 80% to 90% of the Hits@10 of 10x larger full-lookup RotatE models. When baseline RotatE is constrained to a parameter budget matching NodePiece on YAGO 3-10 by reducing the embedding dimension from 500d to 20d, its performance collapses to 0.262 Hits@10, whereas NodePiece achieves 0.488 Hits@10 (+22.6 absolute points) because its parameter budget is invested in higher-capacity anchor dimensions (100d) and a shared MLP encoder.

  10. Knowl 10 — Topological Impact on Anchor and Relational Token Necessity

    empirical result

    The necessity of anchor nodes versus relational context in NodePiece is heavily governed by the density and relational diversity of the knowledge graph:

    1. Dense, Relation-Rich Graphs (FB15k-237, YAGO 3-10, WD50K): Removing anchor tokens completely (relying solely on mm sampled outgoing relations) produces minimal degradation in link prediction (e.g., dropping only 7 points Hits@10 on FB15k-237) and maintains or improves performance on node classification (0.984 ROC-AUC on WD50K NC). The high cardinality and unique combinations of edge types around a node provide sufficient discriminative capacity.

    2. Sparse, Low-Relation Graphs (WN18RR): WN18RR has a large diameter (23 hops), average node-anchor distance of ~6 hops, and only 11 relation types (22 with inverses). In this setting, removing anchors causes link prediction Hits@10 to collapse from 0.515 to 0.019, and removing distance encodings reduces MRR from 0.403 to 0.266. Sparser graphs require larger anchor budgets (∣A∣≥500|A| \ge 500, k≥20k \ge 20) and distance encodings to resolve structural locality.

Coverage note — None was omitted; all primary methodological components, theoretical formulations, and empirical findings across transductive, inductive, out-of-sample link prediction, and node classification are covered.

References

  1. 1.Marjan Albooyeh, Rishab Goel, and Seyed Mehran Kazemi. Out-of-sample representation learning for knowledge graphs. In Findings of the Association for Computational Linguistics: EMNLP 2020, pp. 2657–2666, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.findings-emnlp.241. URL https://www.aclweb.org/anthology/2020.findings-emnlp.241.
  2. 2.Mehdi Ali, Max Berrendorf, Charles Tapley Hoyt, Laurent Vermue, Mikhail Galkin, Sahand Sharifzadeh, Asja Fischer, Volker Tresp, and Jens Lehmann. Bringing light into the dark: A large-scale evaluation of knowledge graph embedding models under a unified framework. CoRR, abs/2006.13365, 2020.
  3. 3.Mehdi Ali, Max Berrendorf, Charles Tapley Hoyt, Laurent Vermue, Sahand Sharifzadeh, Volker Tresp, and Jens Lehmann. PyKEEN 1.0: A Python Library for Training and Evaluating Knowledge Graph Embeddings. Journal of Machine Learning Research, 22(82):1–6, 2021. URL http://jmlr.org/papers/v22/20-825.html.
  4. 4.Jinheon Baek, Dong Bok Lee, and Sung Ju Hwang. Learning to extrapolate knowledge: Transductive few-shot out-of-graph link prediction. Advances in Neural Information Processing Systems, 33, 2020.
  5. 5.Ivana Balazevic, Carl Allen, and Timothy M. Hospedales. Multi-relational poincaré graph embeddings. In Hanna M. Wallach, Hugo Larochelle, Alina Beygelzimer, Florence d’Alché-Buc, Emily B. Fox, and Roman Garnett (eds.), Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pp. 4465–4475, 2019.
  6. 6.Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. Enriching word vectors with subword information. Transactions of the Association for Computational Linguistics, 5:135–146, 2017. doi: 10.1162/tacl_a_00051. URL https://www.aclweb.org/anthology/Q17-1010.
  7. 7.Antoine Bordes, Nicolas Usunier, Alberto García-Durán, Jason Weston, and Oksana Yakhnenko. Translating embeddings for modeling multi-relational data. In Christopher J. C. Burges, Léon Bottou, Zoubin Ghahramani, and Kilian Q. Weinberger (eds.), Advances in Neural Information Processing Systems 26: 27th Annual Conference on Neural Information Processing Systems 2013. Proceedings of a meeting held December 5-8, 2013, Lake Tahoe, Nevada, United States, pp. 2787–2795, 2013.
  8. 8.Ines Chami, Adva Wolf, Da-Cheng Juan, Frederic Sala, Sujith Ravi, and Christopher Ré. Low-dimensional hyperbolic knowledge graph embeddings. In Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel R. Tetreault (eds.), Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pp. 6901–6914. Association for Computational Linguistics, 2020.
  9. 9.Mingyang Chen, Wen Zhang, Wei Zhang, Qiang Chen, and Huajun Chen. Meta relational learning for few-shot link prediction in knowledge graphs. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pp. 4208–4217, 2019.
  10. 10.Louis Clouâtre, Philippe Trempe, Amal Zouaq, and Sarath Chandar. MLMLM: link prediction with mean likelihood masked language model. In Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli (eds.), Findings of the Association for Computational Linguistics: ACL/IJCNLP 2021, Online Event, August 1-6, 2021, volume ACL/IJCNLP 2021 of Findings of ACL, pp. 4321–4331. Association for Computational Linguistics, 2021.
  11. 11.Daniel Daza, Michael Cochez, and Paul Groth. Inductive entity representations from text via link prediction. In Jure Leskovec, Marko Grobelnik, Marc Najork, Jie Tang, and Leila Zia (eds.), WWW ’21: The Web Conference 2021, Virtual Event / Ljubljana, Slovenia, April 19-23, 2021, pp. 798–808. ACM / IW3C2, 2021.
  12. 12.Tim Dettmers, Pasquale Minervini, Pontus Stenetorp, and Sebastian Riedel. Convolutional 2d knowledge graph embeddings. In Sheila A. McIlraith and Kilian Q. Weinberger (eds.), Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, (AAAI-18), the 30th innovative Applications of Artificial Intelligence (IAAI-18), and the 8th AAAI Symposium on Educational Advances in Artificial Intelligence (EAAI-18), New Orleans, Louisiana, USA, February 2-7, 2018, pp. 1811–1818. AAAI Press, 2018.
  13. 13.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: pre-training of deep bidirectional transformers for language understanding. In Jill Burstein, Christy Doran, and Thamar Solorio (eds.), Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers), pp. 4171–4186. Association for Computational Linguistics, 2019.
  14. 14.Bahare Fatemi, Perouz Taslakian, David Vázquez, and David Poole. Knowledge hypergraphs: Prediction beyond binary relations. In Christian Bessiere (ed.), Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence, IJCAI 2020, pp. 2191–2197. ijcai.org, 2020.
  15. 15.Matthias Fey and Jan E. Lenssen. Fast graph representation learning with PyTorch Geometric. In ICLR Workshop on Representation Learning on Graphs and Manifolds, 2019.
  16. 16.Mikhail Galkin, Priyansh Trivedi, Gaurav Maheshwari, Ricardo Usbeck, and Jens Lehmann. Message passing for hyper-relational knowledge graphs. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 7346–7359. Association for Computational Linguistics, 2020.
  17. 17.Takuo Hamaguchi, Hidekazu Oiwa, Masashi Shimbo, and Yuji Matsumoto. Knowledge transfer for out-of-knowledge-base entities : A graph neural network approach. In Carles Sierra (ed.), Proceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence, IJCAI 2017, Melbourne, Australia, August 19-25, 2017, pp. 1802–1808. ijcai.org, 2017.
  18. 18.Weihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong, Hongyu Ren, Bowen Liu, Michele Catasta, and Jure Leskovec. Open graph benchmark: Datasets for machine learning on graphs. In Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin (eds.), Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, 2020.
  19. 19.Shaoxiong Ji, Shirui Pan, Erik Cambria, Pekka Marttinen, and Philip S. Yu. A survey on knowledge graphs: Representation, acquisition and applications. CoRR, abs/2002.00388, 2020.
  20. 20.Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020.
  21. 21.Timothée Lacroix, Nicolas Usunier, and Guillaume Obozinski. Canonical tensor decomposition for knowledge base completion. In Jennifer G. Dy and Andreas Krause (eds.), Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsmässan, Stockholm, Sweden, July 10-15, 2018, volume 80 of Proceedings of Machine Learning Research, pp. 2869–2878. PMLR, 2018.
  22. 22.Adam Lerer, Ledell Wu, Jiajun Shen, Timothee Lacroix, Luca Wehrstedt, Abhijit Bose, and Alex Peysakhovich. PyTorch-BigGraph: A Large-scale Graph Embedding System. In Proceedings of the 2nd SysML Conference, Palo Alto, CA, USA, 2019.
  23. 23.Paul Pu Liang, Manzil Zaheer, Yuan Wang, and Amr Ahmed. Anchor & transform: Learning sparse embeddings for large vocabularies. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=Vd7lCMvtLqg.
  24. 24.Farzaneh Mahdisoltani, Joanna Biega, and Fabian M. Suchanek. YAGO3: A knowledge base from multilingual wikipedias. In Seventh Biennial Conference on Innovative Data Systems Research, CIDR 2015, Asilomar, CA, USA, January 4-7, 2015, Online Proceedings. www.cidrdb.org, 2015.
  25. 25.Tharun Medini, Beidi Chen, and Anshumali Shrivastava. {SOLAR}: Sparse orthogonal learned and random embeddings. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=fw-BHZ1KjxJ.
  26. 26.Christian Meilicke, Manuel Fink, Yanjie Wang, Daniel Ruffinelli, Rainer Gemulla, and Heiner Stuckenschmidt. Fine-grained evaluation of rule- and embedding-based systems for knowledge graph completion. In The Semantic Web - ISWC 2018 - 17th International Semantic Web Conference, Monterey, CA, USA, October 8-12, 2018, Proceedings, Part I, volume 11136 of Lecture Notes in Computer Science, pp. 3–20. Springer, 2018.
  27. 27.Tomás Mikolov, Ilya Sutskever, Kai Chen, Gregory S. Corrado, and Jeffrey Dean. Distributed representations of words and phrases and their compositionality. In Christopher J. C. Burges, Léon Bottou, Zoubin Ghahramani, and Kilian Q. Weinberger (eds.), Advances in Neural Information Processing Systems 26: 27th Annual Conference on Neural Information Processing Systems 2013. Proceedings of a meeting held December 5-8, 2013, Lake Tahoe, Nevada, United States, pp. 3111–3119, 2013.
  28. 28.Ryan L. Murphy, Balasubramaniam Srinivasan, Vinayak Rao, and Bruno Ribeiro. Janossy pooling: Learning deep permutation-invariant functions for variable-size inputs, 2019.
  29. 29.Maximilian Nickel, Kevin Murphy, Volker Tresp, and Evgeniy Gabrilovich. A review of relational machine learning for knowledge graphs. Proc. IEEE, 104(1):11–33, 2016.
  30. 30.Lawrence Page, Sergey Brin, Rajeev Motwani, and Terry Winograd. The pagerank citation ranking: Bringing order to the web. Technical Report 1999-66, Stanford InfoLab, November 1999. Previous number = SIDL-WP-1999-0120.
  31. 31.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. Pytorch: An imperative style, high-performance deep learning library. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pp. 8024–8035, 2019.
  32. 32.Jeffrey Pennington, Richard Socher, and Christopher D. Manning. Glove: Global vectors for word representation. In Alessandro Moschitti, Bo Pang, and Walter Daelemans (eds.), Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing, EMNLP 2014, October 25-29, 2014, Doha, Qatar, A meeting of SIGDAT, a Special Interest Group of the ACL, pp. 1532–1543. ACL, 2014.
  33. 33.Xipeng Qiu, Tianxiang Sun, Yige Xu, Yunfan Shao, Ning Dai, and Xuanjing Huang. Pre-trained models for natural language processing: A survey. Science China Technological Sciences, pp. 1–26, 2020.
  34. 34.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  35. 35.Paolo Rosso, Dingqi Yang, and Philippe Cudré-Mauroux. Beyond triplets: Hyper-relational knowledge graph embedding for link prediction. In Proceedings of The Web Conference 2020, pp. 1885–1896, 2020.
  36. 36.Mrinmaya Sachan. Knowledge graph embedding compression. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 2681–2691, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.238. URL https://www.aclweb.org/anthology/2020.acl-main.238.
  37. 37.Ali Sadeghian, Mohammadreza Armandpour, Patrick Ding, and Daisy Zhe Wang. DRUM: end-to-end differentiable rule mining on knowledge graphs. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pp. 15321–15331, 2019.
  38. 38.Tara Safavi and Danai Koutra. CoDEx: A Comprehensive Knowledge Graph Completion Benchmark. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 8328–8350, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-main.669. URL https://www.aclweb.org/anthology/2020.emnlp-main.669.
  39. 39.Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter. CoRR, abs/1910.01108, 2019. URL http://arxiv.org/abs/1910.01108.
  40. 40.Michael Sejr Schlichtkrull, Thomas N. Kipf, Peter Bloem, Rianne van den Berg, Ivan Titov, and Max Welling. Modeling relational data with graph convolutional networks. In The Semantic Web - 15th International Conference, ESWC 2018, Heraklion, Crete, Greece, June 3-7, 2018, Proceedings, volume 10843 of Lecture Notes in Computer Science, pp. 593–607. Springer, 2018.
  41. 41.Mike Schuster and Kaisuke Nakajima. Japanese and korean voice search. In 2012 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2012, Kyoto, Japan, March 25-30, 2012, pp. 5149–5152. IEEE, 2012. doi: 10.1109/ICASSP.2012.6289079. URL https://doi.org/10.1109/ICASSP.2012.6289079.
  42. 42.Rico Sennrich, Barry Haddow, and Alexandra Birch. Neural machine translation of rare words with subword units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 1715–1725, Berlin, Germany, August 2016. Association for Computational Linguistics. doi: 10.18653/v1/P16-1162. URL https://www.aclweb.org/anthology/P16-1162.
  43. 43.Zhiqing Sun, Zhi-Hong Deng, Jian-Yun Nie, and Jian Tang. Rotate: Knowledge graph embedding by relational rotation in complex space. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net, 2019.
  44. 44.Komal K. Teru, Etienne Denis, and Will Hamilton. Inductive relation prediction by subgraph reasoning. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event, volume 119 of Proceedings of Machine Learning Research, pp. 9448–9457. PMLR, 2020.
  45. 45.Kristina Toutanova and Danqi Chen. Observed versus latent features for knowledge base and text inference. In Proceedings of the 3rd Workshop on Continuous Vector Space Models and their Compositionality, pp. 57–66, Beijing, China, July 2015. Association for Computational Linguistics. doi: 10.18653/v1/W15-4007. URL https://www.aclweb.org/anthology/W15-4007.
  46. 46.Shikhar Vashishth, Soumya Sanyal, Vikram Nitin, and Partha Talukdar. Composition-based multi-relational graph convolutional networks. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=BylA_C4tPr.
  47. 47.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergus, S. V. N. Vishwanathan, and Roman Garnett (eds.), Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pp. 5998–6008, 2017.
  48. 48.Denny Vrandecic and Markus Krötzsch. Wikidata: a free collaborative knowledgebase. Commun. ACM, 57(10):78–85, 2014.
  49. 49.Kai Wang, Yu Liu, Qian Ma, and Quan Z Sheng. Mulde: Multi-teacher knowledge distillation for low-dimensional knowledge graph embeddings. arXiv preprint arXiv:2010.07152, 2020.
  50. 50.Peifeng Wang, Jialong Han, Chenliang Li, and Rong Pan. Logic attention based neighborhood aggregation for inductive knowledge graph embedding. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pp. 7152–7159, 2019.
  51. 51.Xiaozhi Wang, Tianyu Gao, Zhaocheng Zhu, Zhengyan Zhang, Zhiyuan Liu, Juanzi Li, and Jian Tang. KEPLER: A unified model for knowledge embedding and pre-trained language representation. Trans. Assoc. Comput. Linguistics, 9:176–194, 2021.
  52. 52.Fan Yang, Zhilin Yang, and William W. Cohen. Differentiable learning of logical rules for knowledge base reasoning. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pp. 2319–2328, 2017.
  53. 53.Liang Yao, Chengsheng Mao, and Yuan Luo. KG-BERT: BERT for knowledge graph completion. CoRR, abs/1909.03193, 2019. URL http://arxiv.org/abs/1909.03193.
  54. 54.Chuxu Zhang, Huaxiu Yao, Chao Huang, Meng Jiang, Zhenhui Li, and Nitesh V Chawla. Few-shot knowledge graph completion. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pp. 3041–3048, 2020a.
  55. 55.Yongqi Zhang, Quanming Yao, Wenyuan Dai, and Lei Chen. Autosf: Searching scoring functions for knowledge graph embedding. In 36th IEEE International Conference on Data Engineering, ICDE 2020, Dallas, TX, USA, April 20-24, 2020, pp. 433–444. IEEE, 2020b.
  56. 56.Yongqi Zhang, Quanming Yao, Wenyuan Dai, and Lei Chen. Autosf: Searching scoring functions for knowledge graph embedding. In 2020 IEEE 36th International Conference on Data Engineering (ICDE), pp. 433–444. IEEE, 2020c.
  57. 57.Yushan Zhu, Wen Zhang, Hui Chen, Xu Cheng, Wei Zhang, and Huajun Chen. Distile: Distiling knowledge graph embeddings for faster and cheaper reasoning. arXiv preprint arXiv:2009.05912, 2020.
  58. 58.Zhaocheng Zhu, Zuobai Zhang, Louis-Pascal Xhonneux, and Jian Tang. Neural bellman-ford networks: A general graph neural network framework for link prediction. In Neural Information Processing Systems, NeurIPS, 2021.

Citation

MLA
Galkin, M., et al. “NodePiece: Compositional and Parameter-Efficient Representations of Large Knowledge Graphs”. arXiv, 2021, http://arxiv.org/abs/2106.12144v2.
APA
Galkin, M., Denis, E., Wu, J., & Hamilton, W. L. (2021). NodePiece: Compositional and Parameter-Efficient Representations of Large Knowledge Graphs. arXiv. http://arxiv.org/abs/2106.12144v2
Chicago
Galkin, M., E. Denis, J. Wu, and W. L. Hamilton. 2021. “NodePiece: Compositional and Parameter-Efficient Representations of Large Knowledge Graphs”. arXiv. http://arxiv.org/abs/2106.12144v2.
Harvard
Galkin, M. et al. (2021) “NodePiece: Compositional and Parameter-Efficient Representations of Large Knowledge Graphs”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2106.12144v2.
Vancouver
1. Galkin M, Denis E, Wu J, Hamilton WL (2021) NodePiece: Compositional and Parameter-Efficient Representations of Large Knowledge Graphs. arXiv

BibTeX

@article{galkin2021nodepiece,
  title = {NodePiece: Compositional and Parameter-Efficient Representations of Large Knowledge Graphs},
  author = {Galkin, Mikhail and Denis, Etienne and Wu, Jiapeng and Hamilton, William L.},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2106.12144v2},
  eprint = {2106.12144}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors