Poincaré Embeddings for Learning Hierarchical Representations

Maximilian NickelDouwe Kiela

article2017NeurIPS1,832 citations

Proposes embedding symbolic data into hyperbolic Poincaré space via Riemannian optimization, enabling compact representations that capture both hierarchical structure and similarity far more effectively than standard Euclidean vector spaces.

Listen

Modern artificial intelligence and machine learning applications rely heavily on learning mathematical vector representations of symbolic data, such as text concepts, taxonomies, and network nodes. Standard techniques map these symbols into flat Euclidean spaces; however, real-world data like social networks and language vocabularies naturally contain latent tree-like or hierarchical structures. Because Euclidean space cannot efficiently represent branching trees without expanding into high dimensions, current models suffer from high memory requirements, slow computation, and a risk of overfitting.

The article aims to demonstrate that embedding symbolic data into hyperbolic space—specifically using the Poincaré ball model—enables compact representations that simultaneously capture both semantic similarity and hierarchical depth in an unsupervised manner. To achieve this, the authors develop a scalable optimization algorithm based on Riemannian stochastic gradient descent and evaluate it across multiple benchmarks: reconstructing and predicting links on the WordNet noun hierarchy (over 82,000 concepts and 740,000 relations), predicting connections across four academic collaboration networks, and scoring lexical entailment on the HyperLex benchmark.

The analysis reveals three critical findings. First, Poincaré embeddings achieve superior representation accuracy at significantly lower dimensions: on WordNet reconstruction, a 5-dimensional Poincaré embedding achieves a mean average precision of 0.823 and an average rank of 4.9, radically outperforming a 200-dimensional Euclidean baseline (precision of 0.168, rank of 1157.3) and translational baselines. Second, in social network link prediction, the hyperbolic model substantially outperforms Euclidean alternatives in low-dimensional settings (e.g., reaching 0.660 precision on the GRQC dataset at dimension 10 versus 0.438 for Euclidean). Third, without task-specific training, a 5-dimensional Poincaré model sets a state-of-the-art Spearman rank correlation of 0.512 on HyperLex, surpassing standard WordNet-based metrics that score between 0.214 and 0.283.

These findings indicate that moving from Euclidean to hyperbolic geometry provides a major boost in computational efficiency and model quality. Organizations can compress large hierarchical knowledge bases and network graphs by an order of magnitude or more without losing structural fidelity, leading to lower storage footprints, reduced computational costs, and better generalization when inferring missing relationships.

Technical leaders and practitioners working with hierarchical, relational, or network data should consider adopting Poincaré embeddings for graph-based knowledge retrieval and link prediction tasks to reduce dimensionality and improve accuracy. Future work should focus on extending hyperbolic models to multi-relational graphs and specialized natural language word embeddings, as well as refining optimization methods to accelerate convergence.

While the results demonstrate strong performance across several standard datasets, hyperbolic embeddings are primarily advantageous for data exhibiting hierarchical or tree-like latent properties; datasets lacking intrinsic hierarchy may not see comparable gains. In addition, optimization stability depends on careful parameter initialization and early training learning rate adjustments, requiring thoughtful implementation during deployment.

arXiv: 1705.08039

No sufficiently relevant recommendations were found.

  • Paper: The Numerical Stability of Hyperbolic Representation Learning, Gal Mishne et al. (2023). This later study directly examines the Poincaré model’s numerical limits and optimization behavior, extending the source’s embedding method into a focused analysis of its stability.
  • Paper: Geom-GCN: Geometric Graph Convolutional Networks, Hongbin Pei et al. (2020). Geom-GCN carries hyperbolic geometry into graph convolution, showing how geometric representations can support structural neighborhood aggregation and downstream node classification.
Cover for Poincaré Embeddings for Learning Hierarchical Representations

Abstract

Representation learning has become an invaluable approach for learning from symbolic data such as text and graphs. However, while complex symbolic datasets often exhibit a latent hierarchical structure, state-of-the-art methods typically learn embeddings in Euclidean vector spaces, which do not account for this property. For this purpose, we introduce a new approach for learning hierarchical representations of symbolic data by embedding them into hyperbolic space -- or more precisely into an n-dimensional Poincaré ball. Due to the underlying hyperbolic geometry, this allows us to learn parsimonious representations of symbolic data by simultaneously capturing hierarchy and similarity. We introduce an efficient algorithm to learn the embeddings based on Riemannian optimization and show experimentally that Poincaré embeddings outperform Euclidean embeddings significantly on data with latent hierarchies, both in terms of representation capacity and in terms of generalization ability.

Table of Contents

  • 1 Introduction
  • 2 Embeddings and Hyperbolic Geometry
  • 3 Poincaré Embeddings
  • 3.1 Optimization
  • 3.2 Training Details
  • 4 Evaluation
  • 4.1 Embedding Taxonomies
  • 4.2 Network Embeddings
  • 4.3 Lexical Entailment
  • 5 Discussion and Future Work
  • References

Knowls

  1. Knowl 1 — Poincaré Ball Model and Hyperbolic Distance for Hierarchical Representations

    model/method

    The Poincaré ball model represents dd-dimensional hyperbolic space Hd\mathbb{H}^d as the open unit ball Bd={x∈Rd:∥x∥<1}\mathcal{B}^d = \{x \in \mathbb{R}^d : \|x\| < 1\} equipped with the Riemannian metric tensor

    gx=(21−∥x∥2)2gEg_x = \left(\frac{2}{1 - \|x\|^2}\right)^2 g^E

    where x∈Bdx \in \mathcal{B}^d, ∥⋅∥\|\cdot\| is the standard Euclidean norm, and gEg^E denotes the Euclidean metric tensor. The hyperbolic distance between two points u,v∈Bdu, v \in \mathcal{B}^d is given by

    d(u,v)=arcosh⁡(1+2∥u−v∥2(1−∥u∥2)(1−∥v∥2))d(u, v) = \operatorname{arcosh}\left(1 + 2 \frac{\|u - v\|^2}{(1 - \|u\|^2)(1 - \|v\|^2)}\right)

    The boundary ∂B=Sd−1\partial \mathcal{B} = \mathcal{S}^{d-1} represents points at infinite hyperbolic distance. Geodesics in Bd\mathcal{B}^d are circular arcs orthogonal to the boundary ∂B\partial \mathcal{B} as well as straight diameters passing through the origin.

    Because the distance grows exponentially as points approach the boundary ∂B\partial \mathcal{B}, the Poincaré ball naturally captures both hierarchical generality and semantic similarity:

    1. The Euclidean norm ∥x∥\|x\| reflects the hierarchical depth of an object (the root node of a tree is positioned near the origin ∥x∥≈0\|x\| \approx 0, while leaves reside close to the boundary ∥x∥→1\|x\| \to 1).
    2. The hyperbolic distance d(u,v)d(u, v) reflects semantic dissimilarity between objects uu and vv.
  2. Knowl 2 — Riemannian Optimization for Poincaré Embeddings

    algorithm

    Poincaré embeddings are learned by solving the constrained optimization problem min⁡ΘL(Θ)\min_{\Theta} \mathcal{L}(\Theta) subject to ∥θi∥<1\|\theta_i\| < 1 for all θi∈Θ={θi}i=1n⊂Bd\theta_i \in \Theta = \{\theta_i\}_{i=1}^n \subset \mathcal{B}^d, using Riemannian Stochastic Gradient Descent (RSGD). Because Bd\mathcal{B}^d is a conformal model, the Riemannian gradient ∇RL(θ)\nabla_R \mathcal{L}(\theta) is obtained by multiplying the Euclidean gradient ∇E=∂L∂θ\nabla_E = \frac{\partial \mathcal{L}}{\partial \theta} by the inverse metric tensor gθ−1=(1−∥θ∥2)24Idg_\theta^{-1} = \frac{(1 - \|\theta\|^2)^2}{4} I_d. Parameter updates apply the retraction Rθ(v)=θ+v\mathcal{R}_\theta(v) = \theta + v followed by a projection onto the interior of Bd\mathcal{B}^d:

    θt+1=proj⁡(θt−ηt(1−∥θt∥2)24∇EL(θt))\theta_{t+1} = \operatorname{proj}\left(\theta_t - \eta_t \frac{(1 - \|\theta_t\|^2)^2}{4} \nabla_E \mathcal{L}(\theta_t)\right)

    where the projection operator is defined as

    proj⁡(θ)={θ∥θ∥−ϵif ∥θ∥≥1θotherwise\operatorname{proj}(\theta) = \begin{cases} \frac{\theta}{\|\theta\|} - \epsilon & \text{if } \|\theta\| \ge 1 \\ \theta & \text{otherwise} \end{cases}

    with numerical stability constant ϵ=10−5\epsilon = 10^{-5}.

    Input: Training relation set D\mathcal{D}, initial learning rate η\eta, burn-in factor c=10c = 10, burn-in epochs Tburn=10T_{\text{burn}} = 10, total epochs TT, embedding dimension dd
    Output: Poincaré embeddings Θ={θi}i=1n⊂Bd\Theta = \{\theta_i\}_{i=1}^n \subset \mathcal{B}^d
    Initialize each θi∼U(−0.001,0.001)d\theta_i \sim \mathcal{U}(-0.001, 0.001)^d independently for all i∈{1,…,n}i \in \{1, \dots, n\}
    for epoch t=1t = 1 to TT do
        if t≤Tburnt \le T_{\text{burn}} then
            ηt←η/c\eta_t \leftarrow \eta / c
        else
            ηt←η\eta_t \leftarrow \eta
        for each observed pair (u,v)∈D(u, v) \in \mathcal{D} and sampled negatives do
            Compute Euclidean gradient ∇E=∂L∂θ\nabla_E = \frac{\partial \mathcal{L}}{\partial \theta}
            θ←θ−ηt(1−∥θ∥2)24∇E\theta \leftarrow \theta - \eta_t \frac{(1 - \|\theta\|^2)^2}{4} \nabla_E
            if ∥θ∥≥1\|\theta\| \ge 1 then
                θ←θ∥θ∥−ϵ\theta \leftarrow \frac{\theta}{\|\theta\|} - \epsilon
    return Θ\Theta

    The computational and memory complexity of each embedding update is O(d)\mathcal{O}(d), scaling linearly with the embedding dimensionality.

  3. Knowl 3 — Euclidean Gradient of the Poincaré Distance

    equation

    For points θ,x∈Bd\theta, x \in \mathcal{B}^d, define scalar quantities:

    α=1−∥θ∥2,β=1−∥x∥2,γ=1+2αβ∥θ−x∥2\alpha = 1 - \|\theta\|^2, \quad \beta = 1 - \|x\|^2, \quad \gamma = 1 + \frac{2}{\alpha \beta} \|\theta - x\|^2

    The partial derivative of the Poincaré distance d(θ,x)=arcosh⁡(γ)d(\theta, x) = \operatorname{arcosh}(\gamma) with respect to the embedding vector θ\theta is given by

    ∂d(θ,x)∂θ=4βγ2−1(∥x∥2−2⟨θ,x⟩+1α2θ−xα)\frac{\partial d(\theta, x)}{\partial \theta} = \frac{4}{\beta \sqrt{\gamma^2 - 1}} \left( \frac{\|x\|^2 - 2\langle \theta, x \rangle + 1}{\alpha^2} \theta - \frac{x}{\alpha} \right)

    By symmetry of the Poincaré metric, ∂d(x,θ)∂θ\frac{\partial d(x, \theta)}{\partial \theta} is computed analogously. The Euclidean gradient of a scalar loss L\mathcal{L} with respect to θ\theta is obtained via the chain rule as ∇E=∂L∂d(θ,x)∂d(θ,x)∂θ\nabla_E = \frac{\partial \mathcal{L}}{\partial d(\theta, x)} \frac{\partial d(\theta, x)}{\partial \theta}.

  4. Knowl 4 — Soft Ranking Loss for Taxonomy Embeddings

    model/method

    To learn embeddings of directed relations without explicit hierarchy supervision, observed pairs D={(u,v)}\mathcal{D} = \{(u, v)\} (e.g., hypernymy relations) are trained by minimizing a soft ranking cross-entropy loss:

    L(Θ)=∑(u,v)∈Dlog⁡e−d(u,v)∑v′∈N(u)e−d(u,v′)\mathcal{L}(\Theta) = \sum_{(u, v) \in \mathcal{D}} \log \frac{e^{-d(u, v)}}{\sum_{v' \in \mathcal{N}(u)} e^{-d(u, v')}}

    where N(u)={v′∣(u,v′)∉D}∪{u}\mathcal{N}(u) = \{v' \mid (u, v') \notin \mathcal{D}\} \cup \{u\} denotes the candidate set of negative examples for uu. For each positive relation (u,v)∈D(u, v) \in \mathcal{D}, 10 negative examples are sampled randomly from N(u)\mathcal{N}(u). This loss function penalizes observed relations for being further apart than negative relations, without pushing distinct subtrees arbitrarily far apart.

  5. Knowl 5 — Reconstruction and Link Prediction Results on WordNet Noun Hierarchy

    data/table

    Evaluations on the transitive closure of the WordNet noun hierarchy (comprising 82,115 nouns and 743,241 hypernymy edges) compare Poincaré embeddings against Euclidean (d(u,v)=∥u−v∥2d(u, v) = \|u - v\|^2) and Translational (d(u,v)=∥u−v+r∥2d(u, v) = \|u - v + r\|^2) embeddings across varying dimensions. Metrics reported are Mean Rank and Mean Average Precision (MAP).

    Dimensionality
    5 10 20 50 100 200
    Reconstruction
    Euclidean Rank 3542.3 2286.9 1685.9 1281.7 1187.3 1157.3
    MAP 0.024 0.059 0.087 0.140 0.162 0.168
    Translational Rank 205.9 179.4 95.3 92.8 92.7 91.0
    MAP 0.517 0.503 0.563 0.566 0.562 0.565
    Poincaré Rank 4.9 4.02 3.84 3.98 3.9 3.83
    MAP 0.823 0.851 0.855 0.860 0.857 0.870
    Link Prediction
    Euclidean Rank 3311.1 2199.5 952.3 351.4 190.7 81.5
    MAP 0.024 0.059 0.176 0.286 0.428 0.490
    Translational Rank 65.7 56.6 52.1 47.2 43.2 40.4
    MAP 0.545 0.554 0.554 0.560 0.562 0.559
    Poincaré Rank 5.7 4.3 4.9 4.6 4.6 4.6
    MAP 0.825 0.852 0.861 0.863 0.856 0.855

    A 5-dimensional Poincaré embedding achieves a Reconstruction MAP of 0.823 and Link Prediction MAP of 0.825, substantially outperforming a 200-dimensional Euclidean embedding (MAP 0.168 / 0.490) and Translational embedding (MAP 0.565 / 0.559).

  6. Knowl 6 — Fermi-Dirac Edge Probability Model for Hyperbolic Network Embeddings

    model/method

    For undirected network datasets, the connection probability between nodes uu and vv given their embeddings Θ\Theta is modeled via the Fermi-Dirac distribution:

    P((u,v)=1∣Θ)=1e(d(u,v)−r)/t+1P((u, v) = 1 \mid \Theta) = \frac{1}{e^{(d(u, v) - r)/t} + 1}

    where d(u,v)d(u, v) is the distance function (Poincaré or Euclidean), r>0r > 0 is a target radius parameter defining the threshold neighborhood within which edges are likely, and t>0t > 0 is a temperature parameter that specifies the steepness of the logistic function and influences average clustering and degree distribution. The model is trained using binary cross-entropy with negative sampling.

  7. Knowl 7 — Empirical Performance on Social and Collaboration Network Graphs

    data/table

    Reconstruction and link prediction performance (Mean Average Precision) on four scientific collaboration network datasets across embedding dimensions d∈{10,20,50,100}d \in \{10, 20, 50, 100\}:

    Reconstruction Link Prediction
    Dataset Method 10 20 50 100 10 20 50 100
    ASTROPH Euclidean 0.376 0.788 0.969 0.989 0.508 0.815 0.946 0.960
    (N=18,772;E=198,110N=18,772; E=198,110) Poincaré 0.703 0.897 0.982 0.990 0.671 0.860 0.977 0.988
    CONDMAT Euclidean 0.356 0.860 0.991 0.998 0.308 0.617 0.725 0.736
    (N=23,133;E=93,497N=23,133; E=93,497) Poincaré 0.799 0.963 0.996 0.998 0.539 0.718 0.756 0.758
    GRQC Euclidean 0.522 0.931 0.994 0.998 0.438 0.584 0.673 0.683
    (N=5,242;E=14,496N=5,242; E=14,496) Poincaré 0.990 0.999 0.999 0.999 0.660 0.691 0.695 0.697
    HEPPH Euclidean 0.434 0.742 0.937 0.966 0.642 0.749 0.779 0.783
    (N=12,008;E=118,521N=12,008; E=118,521) Poincaré 0.811 0.960 0.994 0.997 0.683 0.743 0.770 0.774

    Poincaré embeddings consistently outperform Euclidean embeddings in the low-dimensional regime (e.g., d=10d=10), demonstrating superior capacity to model latent network hierarchies with compact representations.

  8. Knowl 8 — Graded Lexical Entailment Scoring Function

    equation

    To determine the directional hypernymy relation is-a(u,v)\text{is-a}(u, v) (asserting that uu is a type of vv) from continuous Poincaré embeddings, the asymmetric scoring function is defined as

    score(is-a(u,v))=−(1+α(∥v∥−∥u∥))d(u,v)\text{score}(\text{is-a}(u, v)) = -(1 + \alpha(\|v\| - \|u\|)) d(u, v)

    where d(u,v)d(u, v) is the Poincaré distance, and α>0\alpha > 0 is a penalty scaling factor (set to α=103\alpha = 10^3). The term α(∥v∥−∥u∥)\alpha(\|v\| - \|u\|) penalizes configurations where vv is placed deeper in the hierarchy than uu (i.e., ∥v∥>∥u∥\|v\| > \|u\|), enforcing that broader/higher-level concepts must have smaller norms (closer to the origin) than their hyponyms.

  9. Knowl 9 — Benchmark Results on Graded and Non-Graded Lexical Entailment

    data/table

    Evaluation of Poincaré embeddings (d=5d = 5, trained unsupervised on WordNet) on the HyperLex dataset (2,163 rated noun pairs evaluated using Spearman's rank correlation ρ\rho) and WBLESS (non-graded lexical entailment):

    Model FR SLQS-Sim WN-Basic WN-WuP WN-LCh Vis-ID Euclidean Poincaré
    Spearman's ρ\rho 0.283 0.229 0.240 0.214 0.214 0.253 0.389 0.512

    The 5-dimensional Poincaré embeddings achieve ρ=0.512\rho = 0.512 on HyperLex, surpassing the Euclidean baseline (ρ=0.389\rho = 0.389) and previous WordNet-based metrics (WN-Basic ρ=0.240\rho = 0.240, WN-WuP ρ=0.214\rho = 0.214). On the non-graded lexical entailment dataset WBLESS, the same embeddings obtain a classification accuracy of 0.86.

Coverage note — No substantial contributed material was omitted; the knowls cover the mathematical definition of Poincaré embeddings, Riemannian SGD optimization with projections, distance gradient calculations, loss functions, network link prediction models, asymmetric entailment scoring, and all experimental results.

References

  1. 1.Aaron B Adcock, Blair D Sullivan, and Michael W Mahoney. Tree-like structure in large social and information networks. In Data Mining (ICDM), 2013 IEEE 13th International Conference on, pages 1–10. IEEE, 2013.
  2. 2.Shun-ichi Amari. Natural gradient works efficiently in learning. Neural Computation, 10(2): 251–276, 1998.
  3. 3.M Boguñá, F Papadopoulos, and D Krioukov. Sustaining the internet with hyperbolic mapping. Nature communications, 1:62, 2010.
  4. 4.Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. Enriching word vectors with subword information. arXiv preprint arXiv:1607.04606, 2016.
  5. 5.Silvere Bonnabel. Stochastic gradient descent on riemannian manifolds. IEEE Trans. Automat. Contr., 58(9):2217–2229, 2013.
  6. 6.Antoine Bordes, Nicolas Usunier, Alberto García-Durán, Jason Weston, and Oksana Yakhnenko. Translating embeddings for modeling multi-relational data. In Advances in Neural Information Processing Systems 26, pages 2787–2795, 2013.
  7. 7.Guillaume Bouchard, Sameer Singh, and Theo Trouillon. On approximate reasoning capabilities of low-rank vector spaces. AAAI Spring Syposium on Knowledge Representation and Reasoning (KRR): Integrating Symbolic and Neural Approaches, 2015.
  8. 8.Aaron Clauset, Cristopher Moore, and Mark EJ Newman. Hierarchical structure and the prediction of missing links in networks. Nature, 453(7191):98–101, 2008.
  9. 9.John Rupert Firth. A synopsis of linguistic theory, 1930-1955. Studies in linguistic analysis, 1957.
  10. 10.Mikhael Gromov. Hyperbolic groups. In Essays in group theory, pages 75–263. Springer, 1987.
  11. 11.Aditya Grover and Jure Leskovec. node2vec: Scalable feature learning for networks. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pages 855–864. ACM, 2016.
  12. 12.Zellig S Harris. Distributional structure. Word, 10(2-3):146–162, 1954.
  13. 13.Peter D Hoff, Adrian E Raftery, and Mark S Handcock. Latent space approaches to social network analysis. Journal of the american Statistical association, 97(460):1090–1098, 2002.
  14. 14.Douwe Kiela, Laura Rimell, Ivan Vulić, and Stephen Clark. Exploiting image generality for lexical entailment detection. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics (ACL 2015), pages 119–124. ACL, 2015.
  15. 15.Robert Kleinberg. Geographic routing using hyperbolic space. In INFOCOM 2007. 26th IEEE International Conference on Computer Communications. IEEE, pages 1902–1909. IEEE, 2007.
  16. 16.Dmitri Krioukov, Fragkiskos Papadopoulos, Maksim Kitsak, Amin Vahdat, and Marián Boguñá. Hyperbolic geometry of complex networks. Physical Review E, 82(3):036106, 2010.
  17. 17.Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg Corrado, and Jeffrey Dean. Distributed representations of words and phrases and their compositionality. CoRR, abs/1310.4546, 2013.
  18. 18.George Miller and Christiane Fellbaum. Wordnet: An electronic lexical database, 1998.
  19. 19.Maximilian Nickel, Volker Tresp, and Hans-Peter Kriegel. A three-way model for collective learning on multi-relational data. In Proceedings of the 28th International Conference on Machine Learning, ICML, pages 809–816, 2011.
  20. 20.Maximilian Nickel, Xueyan Jiang, and Volker Tresp. Reducing the rank in relational factorization models by including observable patterns. In Advances in Neural Information Processing Systems 27, pages 1179–1187, 2014.
  21. 21.Maximilian Nickel, Lorenzo Rosasco, and Tomaso A. Poggio. Holographic embeddings of knowledge graphs. In Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence, pages 1955–1961, 2016.
  22. 22.Alberto Paccanaro and Geoffrey E. Hinton. Learning distributed representations of concepts using linear relational embedding. IEEE Trans. Knowl. Data Eng., 13(2):232–244, 2001.
  23. 23.Jeffrey Pennington, Richard Socher, and Christopher D Manning. Glove: Global vectors for word representation. In EMNLP, volume 14, pages 1532–1543, 2014.
  24. 24.Bryan Perozzi, Rami Al-Rfou, and Steven Skiena. Deepwalk: Online learning of social representations. In Proceedings of the 20th ACM SIGKDD international conference on Knowledge discovery and data mining, pages 701–710. ACM, 2014.
  25. 25.Erzsébet Ravasz and Albert-László Barabási. Hierarchical organization in complex networks. Physical Review E, 67(2):026112, 2003.
  26. 26.Benjamin Recht, Christopher Ré, Stephen J. Wright, and Feng Niu. Hogwild: A lock-free approach to parallelizing stochastic gradient descent. In Advances in Neural Information Processing Systems 24, pages 693–701, 2011.
  27. 27.Sebastian Riedel, Limin Yao, Andrew McCallum, and Benjamin M Marlin. Relation extraction with matrix factorization and universal schemas. In Proceedings of NAACL-HLT, pages 74–84, 2013.
  28. 28.Mark Steyvers and Joshua B Tenenbaum. The large-scale structure of semantic networks: Statistical analyses and a model of semantic growth. Cognitive science, 29(1):41–78, 2005.
  29. 29.Théo Trouillon, Johannes Welbl, Sebastian Riedel, Éric Gaussier, and Guillaume Bouchard. Complex embeddings for simple link prediction. In Proceedings of the 33nd International Conference on Machine Learning, ICML 2016, New York City, NY, USA, June 19-24, 2016, pages 2071–2080, 2016.
  30. 30.Ivan Vendrov, Ryan Kiros, Sanja Fidler, and Raquel Urtasun. Order-embeddings of images and language. arXiv preprint arXiv:1511.06361, 2015.
  31. 31.Luke Vilnis and Andrew McCallum. Word representations via gaussian embedding. In International Conference on Learning Representations (ICLR), 2015.
  32. 32.Ivan Vulić, Daniela Gerz, Douwe Kiela, Felix Hill, and Anna Korhonen. Hyperlex: A large-scale evaluation of graded lexical entailment. arXiv preprint arXiv:1608.02117, 2016.
  33. 33.Julie Weeds, Daoud Clarke, Jeremy Reffin, David Weir, and Bill Keller. Learning to distinguish hypernyms and co-hyponyms. In Proceedings of the 25th International Conference on Computational Linguistics COLING, pages 2249–2259. Dublin City University and Association for Computational Linguistics, 2014.
  34. 34.Hongyi Zhang, Sashank J. Reddi, and Suvrit Sra. Riemannian SVRG: fast stochastic optimization on riemannian manifolds. In Advances in Neural Information Processing Systems 29, pages 4592–4600, 2016.
  35. 35.George Kingsley Zipf. Human Behaviour and the Principle of Least Effort: an Introduction to Human Ecology. Addison-Wesley, 1949.

Citation

MLA
Nickel, M., and D. Kiela. “Poincaré Embeddings for Learning Hierarchical Representations”. arXiv, 2017, http://arxiv.org/abs/1705.08039v2.
APA
Nickel, M., & Kiela, D. (2017). Poincaré Embeddings for Learning Hierarchical Representations. arXiv. http://arxiv.org/abs/1705.08039v2
Chicago
Nickel, M., and D. Kiela. 2017. “Poincaré Embeddings for Learning Hierarchical Representations”. arXiv. http://arxiv.org/abs/1705.08039v2.
Harvard
Nickel, M. and Kiela, D. (2017) “Poincaré Embeddings for Learning Hierarchical Representations”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1705.08039v2.
Vancouver
1. Nickel M, Kiela D (2017) Poincaré Embeddings for Learning Hierarchical Representations. arXiv

BibTeX

@article{nickel2017poincare,
  title = {Poincaré Embeddings for Learning Hierarchical Representations},
  author = {Nickel, Maximilian and Kiela, Douwe},
  year = {2017},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1705.08039v2},
  eprint = {1705.08039}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors