Convolutional Neural Network Architectures for Matching Natural Language Sentences
Baotian HuZhengdong LuHang LiQingcai Chen
Proposes general convolutional neural network architectures that capture both hierarchical sentence structure and multi-level semantic interactions without relying on language-specific prior knowledge, outperforming traditional baselines across diverse sentence-matching tasks.
Semantic matching—determining the relevance, coherence, or similarity between two sentences—is a fundamental requirement for applications such as search engines, machine translation, and automated dialogue systems. Natural language sentences possess complex hierarchical and sequential structures, yet conventional matching techniques often rely on simplistic keyword overlaps or rigid parsers that fail to capture subtle linguistic interactions across diverse domains.
The article evaluates whether deep convolutional neural network architectures can effectively model both the internal composition of individual sentences and the rich, multi-level interaction patterns between sentence pairs. It demonstrates two novel matching models: Architecture-I, which builds separate hierarchical representations for each sentence before comparing them, and Architecture-II, which directly models localized word and phrase interactions across sentences from the earliest layers.
To evaluate these models, the researchers conducted extensive empirical experiments across three diverse matching benchmarks: an English sentence-completion task using 3 million training triples from Reuters news data, a Chinese social media response-matching task using 45 million training triples from Weibo, and a standard English paraphrase identification benchmark. The models relied on unsupervised word embeddings without requiring specialized linguistic tools or pre-parsed grammatical trees.
The empirical findings demonstrate that Architecture-II consistently outperforms both Architecture-I and established competitor models. In the sentence-completion task, Architecture-II achieved a top-choice precision of 49.62%, compared to 47.51% for Architecture-I and 25.76% to 41.56% for baseline approaches. In the social media response-matching task, Architecture-II secured a top-choice accuracy of 61.95%, surpassing Architecture-I (59.18%) and alternative methods (49.85% to 56.48%). On the smaller paraphrase benchmark, the generic models achieved competitive accuracy (69.90%) and F1 scores (80.91%), performing on par with classical systems despite having no task-specific tailoring.
These results indicate that convolutional models provide a scalable, language-independent foundation for text matching that eliminates the need for expensive, brittle linguistic feature engineering. Allowing sentence segments to interact early in the neural network architecture preserves critical sequential context and localized semantic dependencies, leading to higher accuracy in automated ranking and dialogue applications.
Organizations implementing automated matching or retrieval pipelines should adopt joint interaction-based convolutional architectures (such as Architecture-II) where large-scale training data is available. Before deployment on smaller, niche datasets (fewer than 10,000 instances), practitioners should conduct pilot analyses using regularization techniques like dropout to prevent overfitting, or consider fine-tuning underlying word embeddings to maximize performance.
Confidence in the reported architectures is high for data-rich matching environments, as evidenced by large margin gains on datasets containing hundreds of thousands to millions of instances. However, decision-makers should exercise caution when applying these models to data-constrained domains or tasks that depend strictly on deep, global synonymy, where specialized domain features or alternative representation methods may still be required.
- Paper: Natural Language Processing (almost) from Scratch, Ronan Collobert et al. (2011). It introduces foundational convolutional architectures and lookup tables for sentence processing from raw text without hand-crafted features.
- Paper: A unified architecture for natural language processing: deep neural networks with multitask learning, Ronan Collobert et al. (2008). It lays the initial groundwork for applying deep convolutional neural networks with pooling directly to natural language processing tasks.
- Paper: Efficient Estimation of Word Representations in Vector Space, Tomáš Mikolov et al. (2013). It establishes the word-vector continuous representations that serve as the fundamental inputs for neural sentence-matching architectures.
- Paper: Linguistic Regularities in Continuous Space Word Representations, Tomáš Mikolov et al. (2013). It explains how continuous vector spaces capture semantic and syntactic regularities needed to compute sentence compositionality.
- Paper: Semantic Compositionality through Recursive Matrix-Vector Spaces, Richard Socher et al. (2012). It formulates earlier recursive neural approaches for modeling semantic compositionality in phrases and sentences.
- Paper: Corpus-based and Knowledge-based Measures of Text Semantic Similarity, Rada Mihalcea et al. (2006). It defines classic semantic similarity formulation and evaluation metrics for sentence-level text matching.
- Paper: Recurrent Convolutional Neural Networks for Text Classification, Siwei Lai et al. (2015). It extends convolutional sentence modeling by combining convolutional layers with recurrent structures to capture broader contextual dependencies.
- Paper: A Decomposable Attention Model for Natural Language Inference, Ankur P. Parikh et al. (2016). It advances sentence pair modeling by proposing a lightweight, attention-driven decomposable framework for natural language inference.
- Paper: Supervised Learning of Universal Sentence Representations from Natural Language Inference Data, Alexis Conneau et al. (2017). It systematically compares convolutional and other sentence encoder architectures for learning universal sentence representations on inference data.
- Paper: Convolutional Sequence to Sequence Learning, Jonas Gehring et al. (2017). It scales purely convolutional neural architectures from sentence representation and matching to full sequence-to-sequence generation.
- Paper: Improved Semantic Representations From Tree-Structured Long Short-Term Memory Networks, Kai Sheng Tai et al. (2015). It explores tree-structured recurrent neural networks as an alternative paradigm to CNNs for semantic sentence matching and classification.
- Paper: Skip-Thought Vectors, Ryan Kiros et al. (2015). It extends sentence representation learning to unsupervised generic encoder-decoder settings for downstream semantic relatedness and matching.
- Paper: SemEval-2017 Task 1: Semantic Textual Similarity Multilingual and Crosslingual Focused Evaluation, Daniel Cer et al. (2017). It provides a multilingual and cross-lingual benchmark evaluation for measuring sentence semantic textual similarity.
- Paper: GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding, Alex Wang et al. (2018). It establishes a comprehensive multi-task benchmark encompassing sentence matching, textual similarity, and natural language understanding.
