Built independently by an author, for readers. Read the story and support ChapterPal

keyword

hierarchical softmax

Hierarchical softmax is an efficient, tree-based approximation of the standard softmax function used in neural network models for large-vocabulary classification and representation learning tasks. Instead of evaluating the probability across every possible output class in a flat distribution, hierarchical softmax organizes the classes as leaf nodes in a tree structure, commonly a binary Huffman tree. The probability of selecting a target class is computed as the product of the probabilities of binary routing decisions made at each internal node along the path from the root to that leaf. By decomposing the multi-class prediction into a sequence of independent binary decisions, this technique reduces the computational time complexity for training and evaluating models with large vocabularies from linear to logarithmic with respect to the total number of classes.

7 items

DeepWalk: online learning of social representations

DeepWalk: online learning of social representations

Bryan Perozzi, Rami Al-Rfou, Steven Skiena

OrganizationsStony Brook University

Why you should read this

Proposes DeepWalk, a scalable algorithm that applies language modeling techniques to truncated random walks to learn continuous node embeddings, significantly improving network classification performance on sparse labeled data.

We present DeepWalk, a novel approach for learning latent representations of vertices in a network. These latent representations encode social relations in a continuous vector space, which is easily exploited by statistical models. DeepWalk generalizes recent advancements in language modeling and unsupervised feature learning (or deep learning) from sequences of words to graphs. DeepWalk uses local information obtained from truncated random walks to learn latent representations by treating walks as the equivalent of sentences. We demonstrate DeepWalk's latent representations on several multi-label network classification tasks for social networks such as BlogCatalog, Flickr, and YouTube. Our results show that DeepWalk outperforms challenging baselines which are allowed a global view of the network, especially in the presence of missing information. DeepWalk's representations can provide F1F_1 scores up to 10% higher than competing methods when labeled data is sparse. In some experiments, DeepWalk's representations are able to outperform all baseline methods while using 60% less training data. DeepWalk is also scalable. It is an online learning algorithm which builds useful incremental results, and is trivially parallelizable. These qualities make it suitable for a broad class of real world applications such as network classification, and anomaly detection.

Added

2026-09-06

Distributed Representations of Words and Phrases and their Compositionality

Distributed Representations of Words and Phrases and their Compositionality

Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg Corrado, Jeffrey Dean

OrganizationsGoogle

Why you should read this

Demonstrates how to create simple yet powerful word vector representations that capture meaningful relationships between words and phrases, enabling basic arithmetic operations to produce linguistically sensible results while training significantly faster than previous methods.

The recently introduced continuous Skip-gram model is an efficient method for learning high-quality distributed vector representations that capture a large number of precise syntactic and semantic word relationships. In this paper we present several extensions that improve both the quality of the vectors and the training speed. By subsampling of the frequent words we obtain significant speedup and also learn more regular word representations. We also describe a simple alternative to the hierarchical softmax called negative sampling. An inherent limitation of word representations is their indifference to word order and their inability to represent idiomatic phrases. For example, the meanings of "Canada" and "Air" cannot be easily combined to obtain "Air Canada". Motivated by this example, we present a simple method for finding phrases in text, and show that learning good vector representations for millions of phrases is possible.

Added

2026-02-21

Efficient Estimation of Word Representations in Vector Space

Efficient Estimation of Word Representations in Vector Space

Tomáš Mikolov, Kai Chen, Greg Corrado, Jeffrey Dean

OrganizationsGoogle

Why you should read this

Proposes the continuous bag-of-words and skip-gram architectures that establish the fundamental paradigm of efficiently learning dense semantic vectors from text corpora.

We propose two novel model architectures for computing continuous vector representations of words from very large data sets. The quality of these representations is measured in a word similarity task, and the results are compared to the previously best performing techniques based on different types of neural networks. We observe large improvements in accuracy at much lower computational cost, i.e. it takes less than a day to learn high quality word vectors from a 1.6 billion words data set. Furthermore, we show that these vectors provide state-of-the-art performance on our test set for measuring syntactic and semantic word similarities.

Added

2025-09-29

License

Published with permission