Convolutional Neural Network Architectures for Matching Natural Language Sentences

Baotian HuZhengdong LuHang LiQingcai Chen

article2014NeurIPS1,382 citations

Proposes general convolutional neural network architectures that capture both hierarchical sentence structure and multi-level semantic interactions without relying on language-specific prior knowledge, outperforming traditional baselines across diverse sentence-matching tasks.

Listen

Semantic matching—determining the relevance, coherence, or similarity between two sentences—is a fundamental requirement for applications such as search engines, machine translation, and automated dialogue systems. Natural language sentences possess complex hierarchical and sequential structures, yet conventional matching techniques often rely on simplistic keyword overlaps or rigid parsers that fail to capture subtle linguistic interactions across diverse domains.

The article evaluates whether deep convolutional neural network architectures can effectively model both the internal composition of individual sentences and the rich, multi-level interaction patterns between sentence pairs. It demonstrates two novel matching models: Architecture-I, which builds separate hierarchical representations for each sentence before comparing them, and Architecture-II, which directly models localized word and phrase interactions across sentences from the earliest layers.

To evaluate these models, the researchers conducted extensive empirical experiments across three diverse matching benchmarks: an English sentence-completion task using 3 million training triples from Reuters news data, a Chinese social media response-matching task using 45 million training triples from Weibo, and a standard English paraphrase identification benchmark. The models relied on unsupervised word embeddings without requiring specialized linguistic tools or pre-parsed grammatical trees.

The empirical findings demonstrate that Architecture-II consistently outperforms both Architecture-I and established competitor models. In the sentence-completion task, Architecture-II achieved a top-choice precision of 49.62%, compared to 47.51% for Architecture-I and 25.76% to 41.56% for baseline approaches. In the social media response-matching task, Architecture-II secured a top-choice accuracy of 61.95%, surpassing Architecture-I (59.18%) and alternative methods (49.85% to 56.48%). On the smaller paraphrase benchmark, the generic models achieved competitive accuracy (69.90%) and F1 scores (80.91%), performing on par with classical systems despite having no task-specific tailoring.

These results indicate that convolutional models provide a scalable, language-independent foundation for text matching that eliminates the need for expensive, brittle linguistic feature engineering. Allowing sentence segments to interact early in the neural network architecture preserves critical sequential context and localized semantic dependencies, leading to higher accuracy in automated ranking and dialogue applications.

Organizations implementing automated matching or retrieval pipelines should adopt joint interaction-based convolutional architectures (such as Architecture-II) where large-scale training data is available. Before deployment on smaller, niche datasets (fewer than 10,000 instances), practitioners should conduct pilot analyses using regularization techniques like dropout to prevent overfitting, or consider fine-tuning underlying word embeddings to maximize performance.

Confidence in the reported architectures is high for data-rich matching environments, as evidenced by large margin gains on datasets containing hundreds of thousands to millions of instances. However, decision-makers should exercise caution when applying these models to data-constrained domains or tasks that depend strictly on deep, global synonymy, where specialized domain features or alternative representation methods may still be required.

arXiv: 1503.03244
Cover for Convolutional Neural Network Architectures for Matching Natural Language Sentences

Abstract

Semantic matching is of central importance to many natural language tasks \cite{bordes2014semantic,RetrievalQA}. A successful matching algorithm needs to adequately model the internal structures of language objects and the interaction between them. As a step toward this goal, we propose convolutional neural network models for matching two sentences, by adapting the convolutional strategy in vision and speech. The proposed models not only nicely represent the hierarchical structures of sentences with their layer-by-layer composition and pooling, but also capture the rich matching patterns at different levels. Our models are rather generic, requiring no prior knowledge on language, and can hence be applied to matching tasks of different nature and in different languages. The empirical study on a variety of matching tasks demonstrates the efficacy of the proposed model on a variety of matching tasks and its superiority to competitor models.

Table of Contents

  • 1 Introduction
  • 2 Convolutional Sentence Model
  • 2.1 Some Analysis on the Convolutional Architecture
  • 3 Convolutional Matching Models
  • 3.1 Architecture-I (Arc-I)
  • 3.2 Architecture-II (Arc-II)
  • 3.3 Some Analysis on Arc-II
  • 4 Training
  • 5 Experiments
  • 5.1 Competitor Methods
  • 5.2 Experiment I: Sentence Completion
  • 5.3 Experiment II: Matching A Response to A Tweet
  • 5.4 Experiment III: Paraphrase Identification
  • 5.5 Discussions
  • 6 Related Work
  • 7 Conclusion
  • References

Knowls

  1. Knowl 1 — Architecture-II (ARC-II) 2D Convolutional Matching Model

    model/method

    Architecture-II (ARC-II) is a deep convolutional sentence matching model that operates directly on the interaction space between two sentences SXS_X and SYS_Y, allowing low-level segments of both sentences to interact before sentence-level representations are formed.

    Let the input sentences be represented by sequences of word embeddings x=(x1,…,x∣SX∣)x = (x_1, \dots, x_{|S_X|}) and y=(y1,…,y∣SY∣)y = (y_1, \dots, y_{|S_Y|}), zero-padded to fixed lengths.

    1. Layer-1 (1D Convolution on Segment Pairs): For sliding window position ii in SXS_X and jj in SYS_Y with window width k1k_1, the concatenated segment vector is z^i,j(0)=[xi:i+k1−1⊤,yj:j+k1−1⊤]⊤∈R2k1D0\hat{z}_{i,j}^{(0)} = [x_{i:i+k_1-1}^\top, y_{j:j+k_1-1}^\top]^\top \in \mathbb{R}^{2k_1 D_0}. The output for feature map f∈{1,…,F1}f \in \{1, \dots, F_1\} is computed as: zi,j(1,f)=g(z^i,j(0))⋅σ(w(1,f)z^i,j(0)+b(1,f))z_{i,j}^{(1,f)} = g(\hat{z}_{i,j}^{(0)}) \cdot \sigma(w^{(1,f)} \hat{z}_{i,j}^{(0)} + b^{(1,f)}) where g(v)=0g(v) = 0 if all elements in vector vv equal 0 and g(v)=1g(v) = 1 otherwise, and σ(⋅)\sigma(\cdot) is an activation function (e.g., ReLU).

    2. Layer-2 (2D Max-Pooling): Non-overlapping 2×22 \times 2 pooling is applied across both spatial dimensions: zi,j(2,f)=max⁡({z2i−1,2j−1(1,f),z2i−1,2j(1,f),z2i,2j−1(1,f),z2i,2j(1,f)})z_{i,j}^{(2,f)} = \max(\{z_{2i-1,2j-1}^{(1,f)}, z_{2i-1,2j}^{(1,f)}, z_{2i,2j-1}^{(1,f)}, z_{2i,2j}^{(1,f)}\})

    3. Deeper Layers (2D Convolutions and 2D Max-Pooling): Subsequent odd layers perform 2D convolutions over kℓ×kℓk_\ell \times k_\ell local patches of the preceding feature maps: zi,j(ℓ)=g(z^i,j(ℓ−1))⋅σ(W(ℓ)z^i,j(ℓ−1)+b(ℓ)),ℓ=3,5,…z_{i,j}^{(\ell)} = g(\hat{z}_{i,j}^{(\ell-1)}) \cdot \sigma(W^{(\ell)} \hat{z}_{i,j}^{(\ell-1)} + b^{(\ell)}), \quad \ell = 3, 5, \dots where z^i,j(ℓ−1)\hat{z}_{i,j}^{(\ell-1)} concatenates vectors from the 2D receptive field in Layer-(ℓ−1)(\ell-1). Subsequent even layers perform 2D max-pooling.

    4. Final Scoring: The activations from the final pooling layer are flattened and fed into a multi-layer perceptron (MLP) to output a scalar matching score s(x,y)s(x, y).

  2. Knowl 2 — Architecture-I (ARC-I) Siamese Convolutional Matching Model

    model/method

    Architecture-I (ARC-I) is a Siamese convolutional sentence matching model that encodes two input sentences SXS_X and SYS_Y independently into fixed-length dense vectors using identical convolutional sentence models, and subsequently computes a matching score via a multi-layer perceptron (MLP).

    For a sentence xx, the convolutional sentence model applies:

    1. 1D Convolution: Over sliding windows of width kℓk_\ell across layer ℓ−1\ell-1 representations: zi(ℓ,f)(x)=g(z^i(ℓ−1))⋅σ(w(ℓ,f)z^i(ℓ−1)+b(ℓ,f)),f=1,…,Fℓz_i^{(\ell, f)}(x) = g(\hat{z}_i^{(\ell-1)}) \cdot \sigma(w^{(\ell, f)} \hat{z}_i^{(\ell-1)} + b^{(\ell, f)}), \quad f = 1, \dots, F_\ell where z^i(0)=[xi⊤,xi+1⊤,…,xi+k1−1⊤]⊤\hat{z}_i^{(0)} = [x_i^\top, x_{i+1}^\top, \dots, x_{i+k_1-1}^\top]^\top concatenates word embeddings of window size k1k_1, σ(⋅)\sigma(\cdot) is the activation function (e.g., ReLU), and g(v)g(v) is a zero-gating function (g(v)=0g(v) = 0 if v=0v = \mathbf{0}, else 11) to eliminate padding artifacts.
    2. 1D Max-Pooling: In non-overlapping 2-unit windows for every feature map ff: zi(ℓ,f)=max⁡(z2i−1(ℓ−1,f),z2i(ℓ−1,f)),ℓ=2,4,…z_i^{(\ell, f)} = \max(z_{2i-1}^{(\ell-1, f)}, z_{2i}^{(\ell-1, f)}), \quad \ell = 2, 4, \dots

    After alternating convolution and pooling layers reach a fixed-length vector representation for SXS_X and SYS_Y, the two vectors are concatenated and passed through an MLP to produce the scalar matching score s(x,y)s(x, y). In ARC-I, sentence representations are formed without interaction until the final MLP.

  3. Knowl 3 — Theoretical Subsumption of ARC-I by ARC-II

    theoretical result

    The ARC-II architecture subsumes ARC-I as a special case.

    If the parameters in ARC-II are constrained such that each feature map filter in the first convolutional layer connects exclusively to the sliding window of either sentence SXS_X or sentence SYS_Y (by zeroing out cross-sentence weights in W(1)W^{(1)}):

    1. The output matrix for each filter ff, denoted z1:n,1:n(1,f)z_{1:n, 1:n}^{(1, f)}, becomes rank-one, carrying the exact representation of a 1D convolution performed independently on SXS_X or SYS_Y.
    2. The subsequent 2×22 \times 2 2D max-pooling operation decouples into independent 1D max-pooling operations along the individual dimensions of SXS_X and SYS_Y.
    3. Restricting the convolutional weights in subsequent layers to maintain this separation allows ARC-II to preserve independent abstraction hierarchies on each sentence until the terminal MLP, exactly reproducing ARC-I.
  4. Knowl 4 — Order Preservation in ARC-II

    theoretical result

    In ARC-II, the combined 2D convolution and 2D max-pooling operations preserve the relative sequential order of words from both sentences in a conditional sense: feature unit zi,j(ℓ)z_{i,j}^{(\ell)} contains representations of tokens in sentence SXS_X that precede the tokens represented in zi+1,j(ℓ)z_{i+1,j}^{(\ell)} (and similarly for sentence SYS_Y along index jj).

    When trained on triples (SX,SY,S~Y)(S_X, S_Y, \tilde{S}_Y) where S~Y\tilde{S}_Y is formed by randomly shuffling the words of true response SYS_Y, ARC-II learns to distinguish correctly ordered sentences from shuffled ones with approximately 60% accuracy under contrastive negative sampling.

  5. Knowl 5 — Margin-Ranking Loss and Training Procedure for Sentence Matching

    model/method

    The matching models (ARC-I and ARC-II) are trained discriminatively using a margin-ranking loss over triples (x,y+,y−)(x, y^+, y^-), where xx is a source sentence, y+y^+ is a correct/positive match, and y−y^- is an incorrect/negative match:

    e(x,y+,y−;Θ)=max⁡(0,1+s(x,y−)−s(x,y+))e(x, y^+, y^-; \Theta) = \max(0, 1 + s(x, y^-) - s(x, y^+))

    where s(x,y)s(x, y) is the predicted scalar matching score and Θ\Theta encompasses all convolutional filter weights, biases, and MLP parameters.

    Key optimization and regularization details:

    • Optimized via stochastic gradient descent (SGD) with mini-batch sizes of 100--200.
    • Gradients for zero-padded units turned off by the gate g(v)=0g(v) = 0 are discounted during backpropagation.
    • Early stopping is sufficient regularization for large datasets (>500K>500\text{K} triples).
    • Dropout combined with early stopping is required for small datasets (<10K<10\text{K} triples) to prevent overfitting.
    • Activation function: ReLU across all convolution and MLP layers.
  6. Knowl 6 — Sentence Completion Benchmark Results on Reuters

    data/table

    The sentence completion task evaluates a model's ability to recover the correct second clause SYS_Y of a two-clause Reuters sentence given the first clause SXS_X. Models were trained on 3 million triples generated from 600K positive pairs and tested on 50K positive pairs, each evaluated against 4 hard negative clauses having cosine similarity of 0.7--0.8 with the original second clause. Performance is measured by Precision at 1 (P@1 in %).

    Model P@1 (%)
    Random Guess 20.00
    URAE+MLP 25.76
    DEEPMATCH 32.50
    SENMLP 36.14
    WORDEMBED 37.63
    SENNA+MLP 41.56
    ARC-I 47.51
    ARC-II 49.62

    ARC-II achieved the highest precision (49.62%), outperforming ARC-I (47.51%) and shallow/non-convolutional baselines. URAE+MLP performed worst among learned models (25.76%), attributed to domain mismatch and clause splitting disrupting syntactic parsing.

  7. Knowl 7 — Short-Text Conversation Response Matching Results on Weibo

    data/table

    The tweet-response matching task evaluates selection of the true response to a microblog post among negative alternatives on a Chinese Weibo dataset. Models were trained on 45 million triples (4.5 million positive pairs, each with 10 random negative responses) and evaluated on 300K original test pairs, each paired with 4 random negative responses. Performance is measured by Precision at 1 (P@1 in %).

    Model P@1 (%)
    Random Guess 20.00
    DEEPMATCH 49.85
    SENMLP 52.22
    WORDEMBED 54.31
    SENNA+MLP 56.48
    ARC-I 59.18
    ARC-II 61.95

    ARC-II achieved the best performance (61.95%), beating ARC-I (59.18%) and the baseline methods by leveraging early 2D interaction modeling on loose, local matching patterns.

  8. Knowl 8 — Paraphrase Identification Results on the MSRP Benchmark

    data/table

    Paraphrase identification on the Microsoft Research Paraphrase (MSRP) dataset tests matching of homogeneous sentences (4,076 training pairs, 1,725 test pairs). Performance is measured by classification Accuracy (Acc. in %) and F1 score (in %).

    Model Acc. (%) F1 (%)
    Baseline 66.5 79.90
    SENMLP 68.4 79.50
    SENNA+MLP 68.4 79.70
    WORDEMBED 68.7 80.49
    ARC-I 69.6 80.27
    ARC-II 69.9 80.91
    Rus et al. (2008) 70.6 80.50

    On this small dataset, ARC-II achieved 69.9% accuracy and 80.91% F1, closely matching hand-crafted feature baselines (Rus et al., 2008) without task-specific engineering, though remaining below specialized systems (e.g., Unfolding-RAE at 76.8%/83.6%).

Coverage note — None was omitted; all key architectural descriptions, theoretical analyses, training protocols, and empirical evaluations across the three target datasets have been fully represented.

References

  1. 1.O. Abdel-Hamid, A. Mohamed, H. Jiang, and G. Penn. Applying convolutional neural networks concepts to hybrid nn-hmm model for speech recognition. In Proceedings of ICASSP, 2012.
  2. 2.B. Antoine, X. Glorot, J. Weston, and Y. Bengio. A semantic matching energy function for learning with multi-relational data. Machine Learning, 94(2):233–259, 2014.
  3. 3.Y. Bengio. Learning deep architectures for ai. Found. Trends Mach. Learn., 2(1):1–127, 2009.
  4. 4.Y. Bengio, J. Louradourand, R. Collobert, and J. Weston. Curriculum learning. In Proceedings of ICML, 2009.
  5. 5.P. F. Brown, S. A. D. Pietra, V. J. D. Pietra, and R. L. Mercer. The mathematics of statistical machine translation: Parameter estimation. Computational linguistics, 19(2):263–311, 1993.
  6. 6.R. Collobert, J. Weston, L. Bottou, M. Karlen, K. Kavukcuoglu, and P. Kuksa. Natural language processing (almost) from scratch. Journal of Machine Learning Research, 12:2493–2537, 2011.
  7. 7.G. E. Dahl, T. N. Sainath, and G. E. Hinton. Improving deep neural networks for lvcsr using rectified linear units and dropout. In Proceedings of ICASSP, 2013.
  8. 8.G. E. Hinton, N. Srivastava, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov. Improving neural networks by preventing co-adaptation of feature detectors. CoRR, abs/1207.0580, 2012.
  9. 9.N. Kalchbrenner, E. Grefenstette, and P. Blunsom. A convolutional neural network for modelling sentences. In Proceedings of ACL, Baltimore and USA, 2014.
  10. 10.Y. Kim. Convolutional neural networks for sentence classification. In Proceedings of EMNLP, 2014.
  11. 11.Y. LeCun and Y. Bengio. Convolutional networks for images, speech and time series. The Handbook of Brain Theory and Neural Networks, 3361, 1995.
  12. 12.Y. Lewis, David D.and Yang, T. G. Rose, and F. Li. Rcv1: A new benchmark collection for text categorization research. Journal of Machine Learning Research, 5:361–397, 2004.
  13. 13.Z. Lu and H. Li. A deep architecture for matching short texts. In Advances in NIPS, 2013.
  14. 14.T. Mikolov, K. Chen, G. Corrado, and J. Dean. Efficient estimation of word representations in vector space. CoRR, abs/1301.3781, 2013.
  15. 15.T. Mikolov and M. Karafi'at. Recurrent neural network based language model. In Proceedings of INTERSPEECH, 2010.
  16. 16.C. Rich, L. Steve, and G. Lee. Overfitting in neural nets: Backpropagation, conjugate gradient, and early stopping. In Advances in NIPS, 2000.
  17. 17.V. Rus, P. M. McCarthy, M. C. Lintean, D. S. McNamara, and A. C. Graesser. Paraphrase identification with lexico-syntactic graph subsumption. In Proceedings of FLAIRS Conference, 2008.
  18. 18.Y. Shen, X. He, J. Gao, L. Deng, and G. Mesnil. Learning semantic representations using convolutional neural networks for web search. In Proceedings of WWW, 2014.
  19. 19.R. Socher, E. H. Huang, and A. Y. Ng. Dynamic pooling and unfolding recursive autoencoders for paraphrase detection. In Advances in NIPS, 2011.
  20. 20.R. Socher, C. C. Lin, A. Y. Ng, and C. D. Manning. Parsing Natural Scenes and Natural Language with Recursive Neural Networks. In Proceedings of ICML, 2011.
  21. 21.R. Socher, J. Pennington, E. H. Huang, A. Y. Ng, and C. D. Manning. Semi-supervised recursive autoencoders for predicting sentiment distributions. In Proceedings of EMNLP, 2011.
  22. 22.Y. Song and D. Roth. On dataless hierarchical text classification. In Proceedings of AAAI, 2014.
  23. 23.Y. Sun, X. Wang, and X. Tang. Hybrid deep learning for face verification. In Proceedings of ICCV, 2013.
  24. 24.S. V. N. Vishwanathan, N. N. Schraudolph, R. Kondor, and K. M. Borgwardt. Graph kernels. Journal of Machine Learning Research(JMLR), 11:1201–1242, 2010.
  25. 25.B. Wang, X. Wang, C. Sun, B. Liu, and L. Sun. Modeling semantic relevance for question-answer pairs in web social communities. In Proceedings of ACL, 2010.
  26. 26.H. Wang, Z. Lu, H. Li, and E. Chen. A dataset for research on short-text conversations. In Proceedings of EMNLP, Seattle, Washington, USA, 2013.
  27. 27.W. Wu, Z. Lu, and H. Li. Learning bilinear model for matching queries and documents. The Journal of Machine Learning Research, 14(1):2519–2548, 2013.
  28. 28.X. Xue, J. Jiwoon, and C. W. Bruce. Retrieval models for question and answer archives. In Proceedings of SIGIR ’08, New York, NY, USA, 2008.

Citation

MLA
Hu, B., et al. “Convolutional Neural Network Architectures for Matching Natural Language Sentences”. arXiv, 2015, http://arxiv.org/abs/1503.03244v1.
APA
Hu, B., Lu, Z., Li, H., & Chen, Q. (2015). Convolutional Neural Network Architectures for Matching Natural Language Sentences. arXiv. http://arxiv.org/abs/1503.03244v1
Chicago
Hu, B., Z. Lu, H. Li, and Q. Chen. 2015. “Convolutional Neural Network Architectures for Matching Natural Language Sentences”. arXiv. http://arxiv.org/abs/1503.03244v1.
Harvard
Hu, B. et al. (2015) “Convolutional Neural Network Architectures for Matching Natural Language Sentences”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1503.03244v1.
Vancouver
1. Hu B, Lu Z, Li H, Chen Q (2015) Convolutional Neural Network Architectures for Matching Natural Language Sentences. arXiv

BibTeX

@article{hu2015convolutional,
  title = {Convolutional Neural Network Architectures for Matching Natural Language Sentences},
  author = {Hu, Baotian and Lu, Zhengdong and Li, Hang and Chen, Qingcai},
  year = {2015},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1503.03244v1},
  eprint = {1503.03244}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors