Built independently by an author, for readers. Read the story and support ChapterPal

keyword

hierarchical ConvNet

A hierarchical ConvNet is a deep learning architecture that processes sequential data, such as natural language text, by extracting and aggregating features across multiple levels of abstraction using stacked convolutional layers. Instead of relying solely on the output of the final layer, the network computes intermediate representations at each successive convolutional stage, typically by applying max-pooling over the resulting feature maps. These multi-level features capture linguistic information at varying granularities, ranging from local word patterns and short phrases at lower layers to broader semantic and syntactic structures at higher layers. By concatenating the pooled representations from each layer into a unified, fixed-size vector, a hierarchical ConvNet forms a comprehensive sentence representation that simultaneously preserves both fine-grained local context and high-level global meaning.

1 item

Supervised Learning of Universal Sentence Representations from Natural Language Inference Data

Supervised Learning of Universal Sentence Representations from Natural Language Inference Data

Alexis Conneau, Douwe Kiela, Holger Schwenk, Loïc Barrault, Antoine Bordes

OrganizationsMetaUniversité du Maine

Why you should read this

Demonstrates that training sentence encoders on supervised natural language inference data yields universal embeddings that consistently outperform unsupervised baselines across diverse transfer tasks, establishing inference as an effective pretraining objective for natural language processing.

Many modern NLP systems rely on word embeddings, previously trained in an unsupervised manner on large corpora, as base features. Efforts to obtain embeddings for larger chunks of text, such as sentences, have however not been so successful. Several attempts at learning unsupervised representations of sentences have not reached satisfactory enough performance to be widely adopted. In this paper, we show how universal sentence representations trained using the supervised data of the Stanford Natural Language Inference datasets can consistently outperform unsupervised methods like SkipThought vectors on a wide range of transfer tasks. Much like how computer vision uses ImageNet to obtain features, which can then be transferred to other tasks, our work tends to indicate the suitability of natural language inference for transfer learning to other NLP tasks. Our encoder is publicly available.

Added

2026-09-16