keyword
hierarchical softmax
Hierarchical softmax is an efficient, tree-based approximation of the standard softmax function used in neural network models for large-vocabulary classification and representation learning tasks. Instead of evaluating the probability across every possible output class in a flat distribution, hierarchical softmax organizes the classes as leaf nodes in a tree structure, commonly a binary Huffman tree. The probability of selecting a target class is computed as the product of the probabilities of binary routing decisions made at each internal node along the path from the root to that leaf. By decomposing the multi-class prediction into a sequence of independent binary decisions, this technique reduces the computational time complexity for training and evaluating models with large vocabularies from linear to logarithmic with respect to the total number of classes.
7 items

Deep Graph Kernels
Pinar Yanardag, S.V.N. Vishwanathan
Why you should read this
Introduces a unified framework that adapts deep language modeling techniques to learn continuous sub-structure embeddings, overcoming diagonal dominance and feature sparsity to boost graph classification accuracy across multiple standard kernels.
In this paper, we present Deep Graph Kernels, a unified frame-work to learn latent representations of sub-structures for graphs, inspired by latest advancements in language modeling and deep learning. Our framework leverages the dependency information between sub-structures by learning their latent representations. We demonstrate instances of our framework on three popular graph kernels, namely Graphlet kernels, Weisfeiler-Lehman subtree kernels, and Shortest-Path graph kernels. Our experiments on several benchmark datasets show that Deep Graph Kernels achieve significant improvements in classification accuracy over state-of-the-art graph kernels.
Added
2026-09-25

GraRep: Learning Graph Representations with Global Structural Information
Shaosheng Cao, Wei Lu, Qiongkai Xu
Why you should read this
Proposes GraRep, a graph representation learning model that captures high-order relational information by directly factorizing distinct k-step probability transition matrices to preserve global graph structure across separate subspaces without sampling.
In this paper, we present GraRep, a novel model for learning vertex representations of weighted graphs. This model learns low dimensional vectors to represent vertices appearing in a graph and, unlike existing work, integrates global structural information of the graph into the learning process. We also formally analyze the connections between our work and several previous research efforts, including the DeepWalk model of Perozzi et al. [20] as well as the skip-gram model with negative sampling of Mikolov et al. [18] We conduct experiments on a language network, a social network as well as a citation network and show that our learned global representations can be effectively used as features in tasks such as clustering, classification and visualization. Empirical results demonstrate that our representation significantly outperforms other state-of-the-art methods in such tasks.
Added
2026-09-24

Language Modeling with Gated Convolutional Networks
Yann Dauphin, Angela Fan, Michael Auli, David Grangier
Why you should read this
Introduces a gated convolutional architecture for language modeling that achieves state-of-the-art results on WikiText-103 and reduces evaluation latency by an order of magnitude compared to recurrent baselines through parallelized computation.
The pre-dominant approach to language modeling to date is based on recurrent neural networks. Their success on this task is often linked to their ability to capture unbounded context. In this paper we develop a finite context approach through stacked convolutions, which can be more efficient since they allow parallelization over sequential tokens. We propose a novel simplified gating mechanism that outperforms Oord et al (2016) and investigate the impact of key architectural decisions. The proposed approach achieves state-of-the-art on the WikiText-103 benchmark, even though it features long-term dependencies, as well as competitive results on the Google Billion Words benchmark. Our model reduces the latency to score a sentence by an order of magnitude compared to a recurrent baseline. To our knowledge, this is the first time a non-recurrent approach is competitive with strong recurrent models on these large scale language tasks.
Added
2026-09-13

Bag of Tricks for Efficient Text Classification
Armand Joulin, Edouard Grave, Piotr Bojanowski, Tomas Mikolov
Why you should read this
Presents fastText, a lightweight linear model that matches the accuracy of deep neural networks on text classification benchmarks while training orders of magnitude faster on standard CPU hardware.
This paper explores a simple and efficient baseline for text classification. Our experiments show that our fast text classifier fastText is often on par with deep learning classifiers in terms of accuracy, and many orders of magnitude faster for training and evaluation. We can train fastText on more than one billion words in less than ten minutes using a standard multicore~CPU, and classify half a million sentences among~312K classes in less than a minute.
Added
2026-09-10


DeepWalk: online learning of social representations
Bryan Perozzi, Rami Al-Rfou, Steven Skiena
Why you should read this
Proposes DeepWalk, a scalable algorithm that applies language modeling techniques to truncated random walks to learn continuous node embeddings, significantly improving network classification performance on sparse labeled data.
We present DeepWalk, a novel approach for learning latent representations of vertices in a network. These latent representations encode social relations in a continuous vector space, which is easily exploited by statistical models. DeepWalk generalizes recent advancements in language modeling and unsupervised feature learning (or deep learning) from sequences of words to graphs. DeepWalk uses local information obtained from truncated random walks to learn latent representations by treating walks as the equivalent of sentences. We demonstrate DeepWalk's latent representations on several multi-label network classification tasks for social networks such as BlogCatalog, Flickr, and YouTube. Our results show that DeepWalk outperforms challenging baselines which are allowed a global view of the network, especially in the presence of missing information. DeepWalk's representations can provide scores up to 10% higher than competing methods when labeled data is sparse. In some experiments, DeepWalk's representations are able to outperform all baseline methods while using 60% less training data. DeepWalk is also scalable. It is an online learning algorithm which builds useful incremental results, and is trivially parallelizable. These qualities make it suitable for a broad class of real world applications such as network classification, and anomaly detection.
Added
2026-09-06

Distributed Representations of Words and Phrases and their Compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg Corrado, Jeffrey Dean
Why you should read this
Demonstrates how to create simple yet powerful word vector representations that capture meaningful relationships between words and phrases, enabling basic arithmetic operations to produce linguistically sensible results while training significantly faster than previous methods.
The recently introduced continuous Skip-gram model is an efficient method for learning high-quality distributed vector representations that capture a large number of precise syntactic and semantic word relationships. In this paper we present several extensions that improve both the quality of the vectors and the training speed. By subsampling of the frequent words we obtain significant speedup and also learn more regular word representations. We also describe a simple alternative to the hierarchical softmax called negative sampling. An inherent limitation of word representations is their indifference to word order and their inability to represent idiomatic phrases. For example, the meanings of "Canada" and "Air" cannot be easily combined to obtain "Air Canada". Motivated by this example, we present a simple method for finding phrases in text, and show that learning good vector representations for millions of phrases is possible.
Added
2026-02-21

Efficient Estimation of Word Representations in Vector Space
Tomáš Mikolov, Kai Chen, Greg Corrado, Jeffrey Dean
Why you should read this
Proposes the continuous bag-of-words and skip-gram architectures that establish the fundamental paradigm of efficiently learning dense semantic vectors from text corpora.
We propose two novel model architectures for computing continuous vector representations of words from very large data sets. The quality of these representations is measured in a word similarity task, and the results are compared to the previously best performing techniques based on different types of neural networks. We observe large improvements in accuracy at much lower computational cost, i.e. it takes less than a day to learn high quality word vectors from a 1.6 billion words data set. Furthermore, we show that these vectors provide state-of-the-art performance on our test set for measuring syntactic and semantic word similarities.
Added
2025-09-29
License
Published with permission
