A Sensitivity Analysis of (and Practitioners’ Guide to) Convolutional Neural Networks for Sentence Classification

Ye ZhangByron C. Wallace

article2015International Joint Conference on Natural Language Processing1,314 citationsBest Student Paper Award

Establishes practical guidelines for configuring convolutional neural networks in sentence classification by isolating which hyperparameter choices critically influence model accuracy and which can be safely neglected.

Listen

Convolutional neural networks have emerged as high-performing tools for text categorization, offering an attractive modern baseline to replace traditional models like support vector machines and logistic regression. However, configuring these neural networks requires tuning numerous architectural components and hyperparameters, such as filter region sizes, feature map quantities, pooling strategies, and regularization terms. Because training these models is computationally expensive, performing an exhaustive search over every potential configuration is impractical in real-world deployments.

The article evaluates the sensitivity of single-layer convolutional neural networks across various hyperparameter settings and architectural choices. Its main objective is to distinguish between design decisions that critically impact classification performance and those that are relatively inconsequential, thereby establishing practical default settings and search ranges for practitioners.

To establish robust conclusions, the authors conducted extensive empirical sensitivity analyses across nine benchmark sentence classification datasets, spanning sentiment analysis, question classification, subjectivity, and irony detection. Rather than relying on single-run cross-validation means—which mask significant stochastic variance from random parameter initializations and optimization steps—the study executed replicated 10-fold cross-validations, repeating runs up to 100 times to measure true mean performance and performance ranges.

The analysis yielded several key findings. First, pre-trained word representations that are updated during model training uniformly outperform fixed representations, whereas simple one-hot encodings perform poorly on short sentence datasets. Second, filter region size and the number of feature maps strongly influence classification accuracy; optimal single filter sizes generally range from 1 to 10 for standard sentences, and setting feature maps between 100 and 600 delivers strong results before hitting diminishing returns. Third, 1-max pooling consistently and decisively outperforms local pooling and average pooling strategies across all tested datasets. Fourth, common activation functions such as ReLU and hyperbolic tangent consistently achieve top performance, and even linear identity functions remain competitive. Finally, standard regularization methods, including dropout and weight norm constraints, showed surprisingly little effect on overall performance, only offering modest improvements when models were scaled to hundreds of feature maps.

These findings provide clear practical implications for project timelines, computing costs, and model optimization workflows. Engineering teams do not need to waste valuable computational resources exploring complex pooling methods or extensive regularization tuning on simple single-layer architectures. Instead, development efforts and budget should focus on tuning filter sizes, selecting effective pre-trained embeddings, and setting adequate feature map capacities.

Based on the empirical evidence, practitioners should adopt a standardized optimization strategy: initialize models with task-tuned embeddings (such as word2vec or GloVe), lock the architecture to 1-max pooling with ReLU or hyperbolic tangent activations, and perform a focused search over filter region sizes (initially 1 to 10, or larger for longer texts) followed by combining nearby filter sizes. When scaling feature maps up to 600, teams should maintain small dropout rates (between 0.0 and 0.5) to avoid overfitting. Crucially, evaluation protocols must incorporate repeated cross-validation runs to account for stochastic variance of 1.5 to 3.4 percentage points before drawing conclusions about model superiority.

Confidence in these guidelines is high for single-layer text classification tasks on standard sentence lengths. However, stakeholders should note that these findings are bounded by short-text datasets and may not transfer directly to multi-layer deep architectures, very large document classification tasks, or scenarios with massive training corpora where one-hot encodings or complex semi-supervised networks may behave differently.

arXiv: 1510.03820
Cover for A Sensitivity Analysis of (and Practitioners’ Guide to) Convolutional Neural Networks for Sentence Classification

Abstract

Convolutional Neural Networks (CNNs) have recently achieved remarkably strong performance on the practically important task of sentence classification (kim 2014, kalchbrenner 2014, johnson 2014). However, these models require practitioners to specify an exact model architecture and set accompanying hyperparameters, including the filter region size, regularization parameters, and so on. It is currently unknown how sensitive model performance is to changes in these configurations for the task of sentence classification. We thus conduct a sensitivity analysis of one-layer CNNs to explore the effect of architecture components on model performance; our aim is to distinguish between important and comparatively inconsequential design decisions for sentence classification. We focus on one-layer CNNs (to the exclusion of more complex models) due to their comparative simplicity and strong empirical performance, which makes it a modern standard baseline method akin to Support Vector Machine (SVMs) and logistic regression. We derive practical advice from our extensive empirical results for those interested in getting the most out of CNNs for sentence classification in real world settings.

Table of Contents

  • 1 Introduction
  • 2 Background and Preliminaries
  • 2.1 CNN Architecture
  • 3 Datasets
  • 4 Baseline Models
  • 4.1 Baseline Configuration
  • 4.2 Effect of input word vectors
  • 4.3 Effect of filter region size
  • 4.4 Effect of number of feature maps for each filter region size
  • 4.5 Effect of activation function
  • 4.6 Effect of pooling strategy
  • 4.7 Effect of regularization
  • 5 Conclusions
  • 5.1 Summary of Main Empirical Findings
  • 5.2 Specific advice to practitioners
  • 6 Acknowledgments
  • References

Knowls

  1. Knowl 1 — One-Layer CNN Architecture for Sentence Classification

    model/method

    The standard one-layer Convolutional Neural Network (CNN) architecture for sentence classification processes an input tokenized sentence of length ss represented as a dense matrix A∈Rs×dA \in \mathbb{R}^{s \times d}, where row ii corresponds to a dd-dimensional word embedding vector for the ii-th token.

    Convolution is performed using linear filters of weight matrix w∈Rh×dw \in \mathbb{R}^{h \times d}, where hh denotes the filter region size (the height of the filter spanning hh consecutive word rows) and the filter width is fixed to the embedding dimension dd. Let A[i:j]A[i : j] denote the sub-matrix of AA from row ii to row jj. Applying the filter repeatedly across all valid contiguous word windows yields an intermediate sequence o∈Rs−h+1o \in \mathbb{R}^{s - h + 1}, with elements:

    oi=w⋅A[i:i+h−1]=∑j=1h∑k=1dwj,kA[i+j−1,k]o_i = w \cdot A[i : i + h - 1] = \sum_{j=1}^h \sum_{k=1}^d w_{j,k} A[i + j - 1, k]

    where ⋅\cdot denotes the Frobenius inner product (element-wise multiplication followed by summation) for i∈{1,…,s−h+1}i \in \{1, \dots, s - h + 1\}. A bias term b∈Rb \in \mathbb{R} and an element-wise activation function ff are applied to induce a feature map c∈Rs−h+1c \in \mathbb{R}^{s - h + 1}:

    ci=f(oi+b)c_i = f(o_i + b)

    To handle variable sentence lengths ss and produce a fixed-size representation, 1-max pooling is applied across each feature map cc, extracting the single maximum scalar value c^=max⁡ici\hat{c} = \max_i c_i. Outputs from all feature maps across multiple region sizes and multiple filters per region size are concatenated to form a penultimate feature vector z∈Rmz \in \mathbb{R}^m. The vector zz is passed to a fully connected softmax classification layer, regularized via dropout with probability pp and an l2l_2 norm constraint threshold cnormc_{\text{norm}} on the softmax weight vector. Parameters are estimated using stochastic gradient descent (specifically ADADELTA) minimizing categorical cross-entropy loss, with word vectors either held static or updated (non-static) during backpropagation.

  2. Knowl 2 — Stochastic Estimation Variance in Sentence CNN Evaluation

    empirical result

    Parameter estimation in sentence classification CNNs is subject to substantial stochastic variance stemming from random weight initialization, dropout masks, and stochastic gradient descent (SGD/ADADELTA) mini-batch orderings.

    When evaluating a fixed one-layer CNN architecture on fixed 10-fold cross-validation splits across 100 independent replications on standard sentence classification benchmark datasets, the mean 10-fold cross-validation accuracy exhibits a spread of up to 1.5 percentage points purely due to optimization randomness. On imbalanced datasets measured by Area Under the ROC Curve (AUC), such as Reddit irony detection, the 10-fold CV mean score fluctuates by up to 3.4 AUC points across identical-fold replications.

    Consequently, reporting single cross-validation runs or single mean accuracies without replication variance risks drawing spurious conclusions regarding model architecture and hyperparameter superiority. Multiple cross-validation replications reporting mean, minimum, and maximum scores are required for reliable comparison.

  3. Knowl 3 — Effect and Optimization of Filter Region Sizes in Sentence CNNs

    empirical result

    Filter region size hh (the filter height spanning adjacent words) significantly impacts classification accuracy in one-layer CNNs:

    1. Single Region Size: The optimal single filter region size is dataset-dependent. For datasets with typical sentence lengths (3636 to 5656 words), the optimal single filter height lies in the range h∈[3,7]h \in [3, 7]. For datasets containing longer sentences (such as Customer Reviews, with maximum sentence length of 105 words), optimal region sizes are larger (h≥10h \ge 10). Line-searching across h∈[1,10]h \in [1, 10] serves as an effective heuristic.

    2. Combining Multiple Region Sizes: Combining several filters whose region sizes are close to the optimal single region size (for instance, combining (5,6,7)(5, 6, 7) or (7,8,9)(7, 8, 9) when the optimal single size is 77) improves accuracy compared to standard baseline configurations like (3,4,5)(3, 4, 5). Conversely, combining filter sizes that are far from the optimal region size (e.g., (14,15,16)(14, 15, 16) when the optimal size is 33) degrades classification performance below that of using only the single best region size.

    3. Single vs. Multiple Sizes: Using a single optimal region size with an increased number of feature maps often matches or outperforms naive multi-size combinations.

  4. Knowl 4 — Effect and Search Range of Feature Map Capacity

    empirical result

    The number of feature maps allocated to each filter region size directly affects sentence classification performance and computational complexity:

    • Increasing the number of feature maps per region size from 1010 up to approximately 100–600100\text{--}600 consistently improves model accuracy across benchmark datasets.
    • Increasing capacity beyond 600600 feature maps per region size yields diminishing or negligible performance gains and frequently degrades test performance due to overfitting.
    • Model training and inference times scale linearly with the total number of feature maps.
    • The recommended search range for hyperparameter tuning of feature map count is 100100 to 600600 per filter region size.
  5. Knowl 5 — Superiority of Global 1-Max Pooling over Alternative Pooling Strategies

    empirical result

    Empirical evaluation of pooling strategies across sentence classification datasets demonstrates that global 1-max pooling uniformly outperforms all tested alternative pooling methods:

    1. Local Max Pooling: Partitioning feature maps into local equal-sized windows of size 3,10,20,3, 10, 20, or 3030 and taking the maximum within each window underperforms global 1-max pooling across all benchmark datasets.
    2. Global kk-Max Pooling: Extracting the top kk maximum activation values preserving temporal order with k∈{5,10,15,20}k \in \{5, 10, 15, 20\} consistently yields lower classification accuracy than 11-max pooling (k=1k = 1).
    3. Local and Global Average Pooling: Taking average activations over local regions of sizes 3,10,20,303, 10, 20, 30 or over the full feature map degrades accuracy substantially compared to max pooling and incurs higher runtime.

    This behavior occurs because in sentence classification, the exact position of predictive n-grams is largely invariant, and individual key phrases provide a stronger predictive signal than the global sentence average.

  6. Knowl 6 — Activation Function Selection in Sentence Classification CNNs

    empirical result

    Evaluation of seven activation functions in the convolutional layer across nine sentence classification datasets reveals distinct performance tiers:

    • Rectified Linear Unit (ReLU) and Hyperbolic Tangent (tanh) achieve the strongest overall classification accuracy across almost all datasets.
    • Identity (Iden, i.e., linear convolution with no non-linear activation) is highly competitive on one-layer CNNs, achieving the best accuracy on several datasets (such as Movie Reviews and Customer Reviews), showing that a linear transformation of dense word vectors is often sufficient for sentence classification.
    • SoftPlus is generally suboptimal, achieving the top score on only 1 of the 9 datasets (MPQA).
    • Sigmoid, Cube (f(x)=x3f(x) = x^3), and tanh-cube (f(x)=tanh⁡(x3)f(x) = \tanh(x^3)) consistently underperform ReLU, tanh, and Identity.
  7. Knowl 7 — Regularization Sensitivity to Dropout and L2 Weight Constraints

    empirical result

    The effect of regularization in one-layer sentence CNNs is relatively modest compared to deeper vision models:

    • Penultimate Layer Dropout: In standard configurations (e.g., 100 feature maps per region size), varying the dropout rate p∈[0.0,0.5]p \in [0.0, 0.5] yields minor performance differences. High dropout rates (p≥0.7p \ge 0.7) degrade accuracy significantly, with p=0.9p = 0.9 causing drops of up to 10%10\%.
    • Dropout under High Capacity: When feature map counts are increased to large numbers (e.g., 500 feature maps per region size), higher dropout rates (p=0.7p = 0.7) become beneficial on datasets prone to overfitting (such as SST-1).
    • Convolutional Layer Dropout: Applying dropout directly to input sentence matrices (randomly zeroing elements with probability pp) provides negligible benefit for small pp and sharply hurts accuracy as pp grows.
    • L2L_2 Norm Constraint: Rescaling the softmax weight vector to a maximum norm constraint threshold cc generally does not improve classification performance and can degrade accuracy on certain datasets.
  8. Knowl 8 — Impact of Word Representation Choices on Sentence CNNs

    empirical result

    Input word representation design significantly affects sentence CNN accuracy:

    • Non-static vs. Static: Non-static word embeddings (fine-tuning embeddings during training via backpropagation) uniformly outperform static (frozen) word embeddings across all datasets.
    • Word2Vec vs. GloVe: Neither pre-trained 300-dimensional Google word2vec nor 300-dimensional GloVe uniformly dominates the other; relative performance depends on the specific dataset and task.
    • Embedding Concatenation: Concatenating 300-dimensional word2vec and 300-dimensional GloVe vectors into 600-dimensional inputs does not reliably improve classification accuracy over using either individual pre-trained embedding.
    • One-Hot Encodings: Direct one-hot bag-of-words input representations into a one-layer CNN perform poorly on sentence classification compared to dense pre-trained embeddings due to sentence sparsity and limited training sample size.
  9. Knowl 9 — Comparative Performance of Sentence Classification CNNs Against SVM Baselines

    data/table

    Across nine sentence classification benchmarks, one-layer CNNs with pre-trained word vectors consistently outperform traditional linear and kernel Support Vector Machine (SVM) baselines, including Bag-of-Words SVM with unigram and bigram features (bowSVM), RBF kernel SVM trained on averaged 300-dimensional word2vec representations (wvSVM), and linear SVM combining Bag-of-Words with averaged word2vec vectors (bowwvSVM).

    Dataset bowSVM wvSVM bowwvSVM Non-static word2vec-CNN
    MR 78.24 78.53 79.67 81.24 (80.69, 81.56)
    SST-1 37.92 44.34 43.15 47.08 (46.42, 48.01)
    SST-2 80.54 81.97 83.30 85.49 (85.03, 85.90)
    Subj 89.13 90.94 91.74 93.20 (92.97, 93.45)
    TREC 87.95 83.61 87.33 91.54 (91.15, 91.92)
    CR 80.21 80.79 81.31 83.92 (82.95, 84.56)
    MPQA 85.38 89.27 89.70 89.32 (88.84, 89.73)
    Opi 61.81 62.46 62.25 64.93 (64.23, 65.58)
    Irony 65.74 65.58 66.74 67.07 (65.60, 69.00)

    All table entries report 10-fold cross-validation accuracy percentages, except Irony, which reports Area Under the ROC Curve (AUC) percentage. CNN values report the mean (and minimum, maximum range) calculated across 10 independent replications of 10-fold cross-validation.

  10. Knowl 10 — Practitioner Hyperparameter Tuning Protocol for Sentence CNNs

    model/method

    Based on sensitivity analyses across nine datasets, practitioners deploying one-layer CNNs for sentence classification should follow this hyperparameter configuration and tuning protocol:

    1. Word Embeddings: Initialize with pre-trained 300-dimensional embeddings (Google word2vec or GloVe) in non-static mode (allowing fine-tuning during training). Avoid one-hot encodings for standard sentence classification datasets.
    2. Filter Region Size: Perform a 1D line-search over single region sizes h∈[1,10]h \in [1, 10] (extending to larger sizes for datasets with long average sentence length). Once the single best size h∗h^* is identified, evaluate combining a small set of filter region sizes tightly clustered around h∗h^* (e.g., {h∗−1,h∗,h∗+1}\{h^*-1, h^*, h^*+1\}).
    3. Number of Feature Maps: Search within the range of 100100 to 600600 feature maps per filter region size. Increase capacity toward or above 600600 only if validation accuracy peaks near the upper search boundary.
    4. Activation Function: Test ReLU and tanh as primary candidates; consider Identity (linear) activation for single-layer models.
    5. Pooling: Standardize on global 1-max pooling; do not expend computational budget evaluating local max, kk-max, or average pooling.
    6. Regularization: Set penultimate dropout rate p∈[0.0,0.5]p \in [0.0, 0.5] and use a large l2l_2 norm threshold. Increase dropout rate (p>0.5p > 0.5) only when scaling to high feature map counts that show evidence of overfitting.
    7. Cross-Validation: Run multiple replications of cross-validation to report mean and range statistics, accounting for SGD and initialization stochasticity.

Coverage note — No substantial contributed material was omitted; all hyperparameter sensitivity analyses, empirical findings across 9 datasets, baseline comparisons against SVMs and logistic regression, and practical tuning guidelines are fully represented.

References

  1. 1.Yoshua Bengio, Rejean Ducharme, Pascal Vincent, and Christian Janvin. 2003. A neural probabilistic language model. The Journal of Machine Learning Research, 3:1137–1155.
  2. 2.Yoshua Bengio. 2009. Learning deep architectures for ai. Foundations and trends in Machine Learning, 2(1):1–127.
  3. 3.Yoshua Bengio. 2012. Practical recommendations for gradient-based training of deep architectures. In Neural Networks: Tricks of the Trade, pages 437–478. Springer.
  4. 4.James Bergstra, Daniel Yamins, and David Daniel Cox. 2013. Making a science of model search: Hyperparameter optimization in hundreds of dimensions for vision architectures.
  5. 5.Y-Lan Boureau, Francis Bach, Yann LeCun, and Jean Ponce. 2010a. Learning mid-level features for recognition. In Computer Vision and Pattern Recognition (CVPR), 2010 IEEE Conference on, pages 2559–2566. IEEE.
  6. 6.Y-Lan Boureau, Jean Ponce, and Yann LeCun. 2010b. A theoretical analysis of feature pooling in visual recognition. In Proceedings of the 27th International Conference on Machine Learning (ICML-10), pages 111–118.
  7. 7.Y-Lan Boureau, Nicolas Le Roux, Francis Bach, Jean Ponce, and Yann LeCun. 2011. Ask the locals: multi-way local pooling for image recognition. In Computer Vision (ICCV), 2011 IEEE International Conference on, pages 2651–2658. IEEE.
  8. 8.Thomas M Breuel. 2015. The effects of hyperparameters on sgd training of neural networks. arXiv preprint arXiv:1508.02788.
  9. 9.Danqi Chen and Christopher D Manning. 2014. A fast and accurate dependency parser using neural networks. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), volume 1, pages 740–750.
  10. 10.Adam Coates, Andrew Y Ng, and Honglak Lee. 2011. An analysis of single-layer networks in unsupervised feature learning. In International conference on artificial intelligence and statistics, pages 215–223.
  11. 11.Ronan Collobert and Jason Weston. 2008. A unified architecture for natural language processing: Deep neural networks with multitask learning. In Proceedings of the 25th international conference on Machine learning, pages 160–167. ACM.
  12. 12.Ronan Collobert, Jason Weston, Leon Bottou, Michael Karlen, Koray Kavukcuoglu, and Pavel Kuksa. 2011. Natural language processing (almost) from scratch. The Journal of Machine Learning Research, 12:2493–2537.
  13. 13.Charles Dugas, Yoshua Bengio, Francois Belisle, Claude Nadeau, and Ren e Garcia. 2001. Incorporating second-order functional knowledge for better option pricing. Advances in Neural Information Processing Systems, pages 472–478.
  14. 14.Kavita Ganesan, ChengXiang Zhai, and Jiawei Han. 2010. Opinosis: a graphbased approach to abstractive summarization of highly redundant opinions. In Proceedings of the 23rd International Conference on Computational Linguistics, pages 340–348. Association for Computational Linguistics.
  15. 15.Yoav Goldberg. 2015. A primer on neural network models for natural language processing. arXiv preprint arXiv:1510.00726.
  16. 16.Geoffrey E Hinton, Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever, and Ruslan R Salakhutdinov. 2012. Improving neural networks by preventing co-adaptation of feature detectors. arXiv preprint arXiv:1207.0580.
  17. 17.Minqing Hu and Bing Liu. 2004. Mining and summarizing customer reviews. In Proceedings of the tenth ACM SIGKDD international conference on Knowledge discovery and data mining, pages 168–177. ACM.
  18. 18.Mohit Iyyer, Varun Manjunatha, Jordan Boyd-Graber, and Hal Daume III. 2015. Deep unordered composition rivals syntactic methods for text classification.
  19. 19.Thorsten Joachims. 1998. Text categorization with support vector machines: Learning with many relevant features. Springer.
  20. 20.Rie Johnson and Tong Zhang. 2014. Effective use of word order for text categorization with convolutional neural networks. arXiv preprint arXiv:1412.1058.
  21. 21.Rie Johnson and Tong Zhang. 2015. Semi-supervised convolutional neural networks for text categorization via region embedding. In Advances in Neural Information Processing Systems, pages 919–927.
  22. 22.Nal Kalchbrenner, Edward Grefenstette, and Phil Blunsom. 2014. A convolutional neural network for modelling sentences. In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 655–665, Baltimore, Maryland, June. Association for Computational Linguistics.
  23. 23.Yoon Kim. 2014. Convolutional neural networks for sentence classification. arXiv preprint arXiv:1408.5882.
  24. 24.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012. Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems, pages 1097–1105.
  25. 25.Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. 2015. Deep learning. Nature, 521(7553):436–444.
  26. 26.Xin Li and Dan Roth. 2002. Learning question classifiers. In Proceedings of the 19th international conference on Computational linguistics-Volume 1, pages 1–7. Association for Computational Linguistics.
  27. 27.Andrew L Maas, Awni Y Hannun, and Andrew Y Ng. 2013. Rectifier nonlinearities improve neural network acoustic models. In Proc. ICML, volume 30.
  28. 28.Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013. Distributed representations of words and phrases and their compositionality. In Advances in neural information processing systems, pages 3111–3119.
  29. 29.Bo Pang and Lillian Lee. 2005. Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales. In Proceedings of the ACL.
  30. 30.F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. 2011. Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12:2825–2830.
  31. 31.Wenzhe Pei, Tao Ge, and Baobao Chang. 2015. An effective neural network model for graph-based dependency parsing. In Proc. of ACL.
  32. 32.Jeffrey Pennington, Richard Socher, and Christopher D Manning. 2014. Glove: Global vectors for word representation. Proceedings of the Empiricial Methods in Natural Language Processing (EMNLP 2014), 12:1532–1543.
  33. 33.David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams. 1988. Learning representations by back-propagating errors. Cognitive modeling, 5:3.
  34. 34.Richard Socher, Alex Perelygin, Jean Y Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the conference on empirical methods in natural language processing (EMNLP), volume 1631, page 1642. Citeseer.
  35. 35.Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014. Dropout: A simple way to prevent neural networks from overfitting. The Journal of Machine Learning Research, 15(1):1929–1958.
  36. 36.Byron C Wallace, Kevin Small, Carla E Brodley, and Thomas A Trikalinos. 2011. Class imbalance, redux. In Data Mining (ICDM), 2011 IEEE 11th International Conference on, pages 754–763. IEEE.
  37. 37.Byron C Wallace, Laura Kertz Do Kook Choe, and Eugene Charniak. 2014. Humans require context to infer ironic intent (so computers probably do, too). In Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL), pages 512–516.
  38. 38.Peng Wang, Jiaming Xu, Bo Xu, Chenglin Liu, Heng Zhang, Fangyuan Wang, and Hongwei Hao. 2015. Semantic clustering and convolutional neural network for short text categorization. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pages 352–357, Beijing, China, July. Association for Computational Linguistics.
  39. 39.Janyce Wiebe, Theresa Wilson, and Claire Cardie. 2005. Annotating expressions of opinions and emotions in language. Language resources and evaluation, 39(2-3):165–210.
  40. 40.Dani Yogatama and Noah A Smith. 2015. Bayesian optimization of text representations. arXiv preprint arXiv:1503.00693.
  41. 41.Matthew D Zeiler. 2012. Adadelta: An adaptive learning rate method. arXiv preprint arXiv:1212.5701.

Citation

MLA
Zhang, Y., and B. Wallace. “A Sensitivity Analysis of (and Practitioners' Guide To) Convolutional Neural Networks for Sentence Classification”. arXiv, 2015, http://arxiv.org/abs/1510.03820v4.
APA
Zhang, Y., & Wallace, B. (2015). A Sensitivity Analysis of (and Practitioners' Guide to) Convolutional Neural Networks for Sentence Classification. arXiv. http://arxiv.org/abs/1510.03820v4
Chicago
Zhang, Y., and B. Wallace. 2015. “A Sensitivity Analysis of (and Practitioners' Guide To) Convolutional Neural Networks for Sentence Classification”. arXiv. http://arxiv.org/abs/1510.03820v4.
Harvard
Zhang, Y. and Wallace, B. (2015) “A Sensitivity Analysis of (and Practitioners' Guide to) Convolutional Neural Networks for Sentence Classification”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1510.03820v4.
Vancouver
1. Zhang Y, Wallace B (2015) A Sensitivity Analysis of (and Practitioners' Guide to) Convolutional Neural Networks for Sentence Classification. arXiv

BibTeX

@article{zhang2015sensitivity,
  title = {A Sensitivity Analysis of (and Practitioners' Guide to) Convolutional Neural Networks for Sentence Classification},
  author = {Zhang, Ye and Wallace, Byron},
  year = {2015},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1510.03820v4},
  eprint = {1510.03820}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF