Deep Convolutional Neural Networks for Sentiment Analysis of Short Texts

Cícero Nogueira dos SantosMaíra Gatti

article2014COLING1,328 citations

Proposes CharSCNN, a deep convolutional neural network that jointly extracts character- and sentence-level representations to improve sentiment classification performance on short texts across movie review and Twitter benchmarks.

Listen

Analyzing sentiment in short texts such as social media posts and single sentences is an essential capability for modern digital intelligence, brand monitoring, and customer feedback management. However, short texts present significant analytical challenges due to their limited contextual information, non-standard language, hashtags, and complex syntactic nuances such as negation. Traditional keyword and bag-of-words techniques struggle to capture these subtleties, creating a need for robust automated approaches that can extract meaning across multiple linguistic levels without requiring costly, handcrafted grammatical engineering.

The article demonstrates a novel deep learning framework, named the Character to Sentence Convolutional Neural Network (CharSCNN), designed to perform sentiment analysis on short texts. The primary objective is to evaluate how effectively a multi-level feature extraction model—spanning from individual character shapes to full sentence structures—can classify sentiment across distinct domains without relying on external syntactic parsers.

To evaluate the system, the authors conducted empirical experiments across two benchmark datasets: the Stanford Sentiment Treebank, comprising formal movie review sentences, and the Stanford Twitter Sentiment corpus, containing informal social media messages. The architecture processes texts by first generating character-level and word-level representations, the latter initialized using unsupervised pre-training on a large Wikipedia text collection of 1.75 billion tokens. These combined vectors are processed through two successive convolutional layers to capture contextual features across words and complete sentences, followed by classification scoring optimized through stochastic gradient descent.

The findings show that the proposed approach establishes state-of-the-art accuracy across both test domains. On the movie review dataset, the model attained 85.7% accuracy in binary positive/negative classification and 48.3% in fine-grained five-class classification, outperforming previous recursive neural network baselines by 2.6 percentage points on the fine-grained task. On the social media dataset, the model achieved a record accuracy of 86.4%, with character-level features providing an absolute accuracy gain of 1.2 percentage points over word-only models. Additionally, unsupervised pre-training of word vectors proved crucial, improving accuracy by 1.5 percentage points on formal reviews and 4.5 percentage points on social media posts. The internal feature distributions also confirmed that the convolutional architecture robustly identifies and adjusts for sentiment reversals caused by negation without needing explicit grammatical parse trees.

These results demonstrate that organizations can achieve superior text classification performance using feed-forward convolutional models that do not depend on complex linguistic parsing tools or manually engineered rules. This reduces the computational pipeline complexity and lower deployment costs while increasing operational processing speeds. The strong performance of character-level modeling specifically mitigates the risks of misclassifying informal slang, morphological variations, and hashtag-heavy communications.

Organizations handling high-volume text analytics should consider adopting multi-level convolutional architectures and leveraging unsupervised pre-training on large background datasets to boost classification performance. For operational decision-making, the source supports evaluating model variants based on text domain: character-level embeddings should be prioritized for informal social data, whereas simpler word-level architectures may suffice for well-formed formal prose. Further research and domain-specific pilot testing are recommended to examine the impact of pre-training language models directly on specialized industry corpora rather than general encyclopedic text.

The primary limitation of the study is that character embeddings showed minimal impact on formal review texts, proving beneficial primarily in noisy social media contexts. Additionally, training on social data utilized a 5% sample of available records rather than the full corpus. Nevertheless, the consistent performance across two distinct benchmark datasets supports high confidence in the architecture's effectiveness for short-text sentiment classification.

Santos et al (2014).pdf
Cover for Deep Convolutional Neural Networks for Sentiment Analysis of Short Texts

Abstract

Sentiment analysis of short texts such as single sentences and Twitter messages is challenging because of the limited contextual information that they normally contain. Effectively solving this task requires strategies that combine the small text content with prior knowledge and use more than just bag-of-words. In this work we propose a new deep convolutional neural network that exploits from character- to sentence-level information to perform sentiment analysis of short texts. We apply our approach for two corpora of two different domains: the Stanford Sentiment Treebank (SSTb), which contains sentences from movie reviews; and the Stanford Twitter Sentiment corpus (STS), which contains Twitter messages. For the SSTb corpus, our approach achieves state-of-the-art results for single sentence sentiment prediction in both binary positive/negative classification, with 85.7% accuracy, and fine-grained classification, with 48.3% accuracy. For the STS corpus, our approach achieves a sentiment prediction accuracy of 86.4%.

Table of Contents

  • 1 Introduction
  • 2 Neural Network Architecture
  • 2.1 Initial Representation Levels
  • 2.1.1 Word-Level Embeddings
  • 2.1.2 Character-Level Embeddings
  • 2.2 Sentence-Level Representation and Scoring
  • 2.3 Network Training
  • 3 Related Work
  • 4 Experimental Setup and Results
  • 4.1 Sentiment Analysis Datasets
  • 4.2 Unsupervised Learning of Word-Level Embeddings
  • 4.3 Model Setup
  • 4.4 Results for SSTb Corpus
  • 4.5 Results for STS Corpus
  • 4.6 Sentence-level features
  • 5 Conclusions
  • References

Knowls

  1. Knowl 1 — CharSCNN Architecture for Short Text Sentiment Classification

    model/method

    The Character to Sentence Convolutional Neural Network (CharSCNN) is a deep feed-forward architecture designed for sentiment analysis of short texts of arbitrary length without requiring syntactic parse trees.

    Given an input sentence x=(w1,w2,…,wN)x = (w_1, w_2, \dots, w_N) of NN words, the network extracts hierarchical representations through the following sequence of stages:

    1. Word-Level Representation: Each word wnw_n is mapped via embedding matrix lookup Wwrd∈Rdwrd×∣Vwrd∣W^{wrd} \in \mathbb{R}^{d^{wrd} \times |V^{wrd}|} to a dense vector rwrd∈Rdwrdr^{wrd} \in \mathbb{R}^{d^{wrd}}.
    2. Character-Level Convolutional Representation: The sequence of characters of wnw_n is processed through a character embedding lookup Wchr∈Rdchr×∣Vchr∣W^{chr} \in \mathbb{R}^{d^{chr} \times |V^{chr}|}, a 1D convolution over sliding character windows of size kchrk^{chr}, and global max-pooling over the word to produce a morphological/shape vector rwch∈Rclu0r^{wch} \in \mathbb{R}^{cl^0_u}.
    3. Joint Word-Character Representation: Each word token is represented by concatenating its word and character vectors: un=[rwrd;rwch]∈Rdwrd+clu0u_n = [r^{wrd}; r^{wch}] \in \mathbb{R}^{d^{wrd} + cl^0_u}.
    4. Sentence-Level Convolution: A second convolutional layer applies sliding window filters of size kwrdk^{wrd} over the sequence of joint embeddings (u1,u2,…,uN)(u_1, u_2, \dots, u_N), followed by global max-pooling across the full sentence length to generate a fixed-size sentence representation rxsent∈Rclu1r^{sent}_x \in \mathbb{R}^{cl^1_u}.
    5. Scoring and Classification: The vector rxsentr^{sent}_x is passed through a fully connected hidden layer with hyperbolic tangent (tanh⁡\tanh) activation and a linear output layer to produce raw scores s(x)∈R∣T∣s(x) \in \mathbb{R}^{|T|} across the set of target sentiment labels TT.
  2. Knowl 2 — Character-Level Convolutional Feature Extraction for Word Embeddings

    model/method

    To capture intra-word morphological patterns, word shapes, capitalization, suffixes (e.g., "-ly"), and sub-tokens in hashtags (e.g., "#SoSad"), CharSCNN computes a fixed-size character-level word embedding rwch∈Rclu0r^{wch} \in \mathbb{R}^{cl^0_u} for each word ww composed of MM characters (c1,c2,…,cM)(c_1, c_2, \dots, c_M).

    First, each character cmc_m is converted to a vector rmchr∈Rdchrr^{chr}_m \in \mathbb{R}^{d^{chr}} via character embedding matrix Wchr∈Rdchr×∣Vchr∣W^{chr} \in \mathbb{R}^{d^{chr} \times |V^{chr}|}:

    rmchr=Wchrvcmr^{chr}_m = W^{chr} v^{c_m}

    where vcmv^{c_m} is a one-hot vector of size ∣Vchr∣|V^{chr}|. A sliding context window of size kchrk^{chr} centered at character mm is formed by concatenating the character embedding with its (kchr−1)/2(k^{chr}-1)/2 left and right neighbors (using boundary padding when needed):

    zm=[rm−(kchr−1)/2chr,…,rm+(kchr−1)/2chr]T∈Rdchrkchrz_m = \left[ r^{chr}_{m-(k^{chr}-1)/2}, \dots, r^{chr}_{m+(k^{chr}-1)/2} \right]^T \in \mathbb{R}^{d^{chr} k^{chr}}

    The jj-th element of the resulting word's character-level embedding rwch∈Rclu0r^{wch} \in \mathbb{R}^{cl^0_u} is obtained by applying a linear convolutional transformation followed by global max-pooling across all character window positions m∈{1,…,M}m \in \{1, \dots, M\}:

    [rwch]j=max⁡1≤m≤M[W0zm+b0]j[r^{wch}]_j = \max_{1 \le m \le M} \left[ W^0 z_m + b^0 \right]_j

    where W0∈Rclu0×dchrkchrW^0 \in \mathbb{R}^{cl^0_u \times d^{chr} k^{chr}} is the convolutional filter weight matrix and b0∈Rclu0b^0 \in \mathbb{R}^{cl^0_u} is the bias vector.

  3. Knowl 3 — Sentence-Level Convolution and Sentiment Scoring

    model/method

    Given a sentence x=(w1,…,wN)x = (w_1, \dots, w_N) with joint word-level representations un=[rwrd;rwch]∈Rdwrd+clu0u_n = [r^{wrd}; r^{wch}] \in \mathbb{R}^{d^{wrd} + cl^0_u} for each word n∈{1,…,N}n \in \{1, \dots, N\}, CharSCNN uses a sentence-level convolutional layer to handle variable sentence lengths and capture shift-invariant local syntactic/semantic phrases.

    A context window of size kwrdk^{wrd} centered at the nn-th word token is formed by concatenation (with sentence boundary padding tokens as needed):

    zn=[un−(kwrd−1)/2,…,un+(kwrd−1)/2]T∈R(dwrd+clu0)kwrdz_n = \left[ u_{n-(k^{wrd}-1)/2}, \dots, u_{n+(k^{wrd}-1)/2} \right]^T \in \mathbb{R}^{(d^{wrd} + cl^0_u) k^{wrd}}

    The sentence-level feature vector rxsent∈Rclu1r^{sent}_x \in \mathbb{R}^{cl^1_u} is extracted via a 1D convolution and global max-pooling across all word window positions:

    [rxsent]j=max⁡1≤n≤N[W1zn+b1]j[r^{sent}_x]_j = \max_{1 \le n \le N} \left[ W^1 z_n + b^1 \right]_j

    where W1∈Rclu1×(dwrd+clu0)kwrdW^1 \in \mathbb{R}^{cl^1_u \times (d^{wrd} + cl^0_u) k^{wrd}} and b1∈Rclu1b^1 \in \mathbb{R}^{cl^1_u} are trainable parameters.

    The global sentence feature vector rxsentr^{sent}_x is then fed into a two-layer feed-forward network to compute sentiment score vector s(x)∈R∣T∣s(x) \in \mathbb{R}^{|T|} over sentiment classes TT:

    s(x)=W3h(W2rxsent+b2)+b3s(x) = W^3 h(W^2 r^{sent}_x + b^2) + b^3

    where W2∈Rhlu×clu1W^2 \in \mathbb{R}^{hlu \times cl^1_u}, b2∈Rhlub^2 \in \mathbb{R}^{hlu}, W3∈R∣T∣×hluW^3 \in \mathbb{R}^{|T| \times hlu}, and b3∈R∣T∣b^3 \in \mathbb{R}^{|T|} are trainable parameters, hluhlu is the number of hidden units, and h(⋅)=tanh⁡(⋅)h(\cdot) = \tanh(\cdot) is the hyperbolic tangent activation function.

  4. Knowl 4 — CharSCNN Training Objective and Optimization

    equation

    CharSCNN is trained in a supervised manner by minimizing the negative log-likelihood over a labeled training corpus D={(x,y)}D = \{(x, y)\}, where xx is a short text and y∈Ty \in T is its true sentiment category.

    The model scores sθ(x)∈R∣T∣s_\theta(x) \in \mathbb{R}^{|T|} parameterized by θ={Wwrd,Wchr,W0,b0,W1,b1,W2,b2,W3,b3}\theta = \{W^{wrd}, W^{chr}, W^0, b^0, W^1, b^1, W^2, b^2, W^3, b^3\} are transformed into a normalized conditional probability distribution using the softmax function:

    p(τ∣x,θ)=esθ(x)τ∑i∈Tesθ(x)ip(\tau \mid x, \theta) = \frac{e^{s_\theta(x)_\tau}}{\sum_{i \in T} e^{s_\theta(x)_i}}

    The resulting conditional log-probability for a class label τ\tau is:

    log⁡p(τ∣x,θ)=sθ(x)τ−log⁡(∑i∈Tesθ(x)i)\log p(\tau \mid x, \theta) = s_\theta(x)_\tau - \log \left( \sum_{i \in T} e^{s_\theta(x)_i} \right)

    Parameters θ\theta are optimized via stochastic gradient descent (SGD) to minimize empirical risk:

    θ↦∑(x,y)∈D−log⁡p(y∣x,θ)\theta \mapsto \sum_{(x, y) \in D} -\log p(y \mid x, \theta)

    Gradients with respect to all layer parameters and lookup matrices are computed using backpropagation.

  5. Knowl 5 — Unsupervised Pre-training and Character Embedding Initialization

    experimental setup

    To provide rich semantic and syntactic prior knowledge, word-level embeddings are pre-trained on the December 2013 English Wikipedia corpus (~1.75 billion tokens after preprocessing). Preprocessing involves removing non-English paragraphs, substituting non-Western characters, lowercasing all words, converting numeric digits to 0, and filtering sentences shorter than 20 characters or with fewer than 5 tokens.

    Pre-training is performed using the word2vec skip-gram model with:

    • Context window size: 99
    • Minimum word frequency cutoff: 1010
    • Vocabulary size ∣Vwrd∣|V^{wrd}|: 870,214870,214 words
    • Embedding dimension dwrdd^{wrd}: 3030

    In contrast, character-level embeddings are not pre-trained; they are initialized uniformly at random from U(−r,r)\mathcal{U}(-r, r), where:

    r=6∣Vchr∣+dchrr = \sqrt{\frac{6}{|V^{chr}| + d^{chr}}}

    Character vocabularies preserve casing to capture capitalization cues, yielding ∣Vchr∣=94|V^{chr}| = 94 for SSTb and ∣Vchr∣=453|V^{chr}| = 453 for STS.

  6. Knowl 6 — CharSCNN Hyperparameter Configurations

    experimental setup

    Model hyperparameters were tuned on development validation sets for the Stanford Sentiment Treebank (SSTb) and Stanford Twitter Sentiment (STS) datasets:

    Parameter Parameter Description SSTb STS
    dwrdd^{wrd} Word embedding dimension 30 30
    kwrdk^{wrd} Word context window size 5 5
    dchrd^{chr} Character embedding dimension 5 5
    kchrk^{chr} Character context window size 3 3
    clu0cl^0_u Character convolutional units (size of rwchr^{wch}) 10 50
    clu1cl^1_u Word convolutional units (size of rsentr^{sent}) 300 300
    hluhlu Fully-connected hidden units 300 300
    λ\lambda Learning rate 0.02 0.01

    Training converges within 5 to 10 epochs. The only hyperparameter differences across domains are the learning rate λ\lambda and the character convolution capacity clu0cl^0_u, which is increased for noisy microblogging text.

  7. Knowl 7 — Sentiment Classification Accuracy on Stanford Sentiment Treebank (SSTb)

    empirical result

    CharSCNN was evaluated on the SSTb test set of complete sentences across fine-grained (5 classes: very negative, negative, neutral, positive, very positive) and binary (positive/negative) classification settings, comparing against word-only convolutional networks (SCNN) and baseline models:

    Model Trained on Phrases Fine-Grained Accuracy (%) Positive/Negative Accuracy (%)
    CharSCNN yes 48.3 85.7
    SCNN yes 48.3 85.5
    CharSCNN no 43.5 82.3
    SCNN no 43.5 82.0
    RNTN (Socher et al., 2013b) yes 45.7 85.4
    MV-RNN (Socher et al., 2013b) yes 44.4 82.9
    RNN (Socher et al., 2013b) yes 43.2 82.4
    NB (Socher et al., 2013b) yes 41.0 81.8
    SVM (Socher et al., 2013b) yes 40.7 79.4

    Key empirical findings on SSTb:

    1. Training with labeled sub-phrases substantially improves full-sentence test accuracy (by +4.8%+4.8\% in fine-grained and +3.4%+3.4\% to +3.5%+3.5\% in binary settings) even though the convolutional network does not take syntactic parse trees as input.
    2. CharSCNN achieved state-of-the-art fine-grained accuracy (48.3%48.3\%), outperforming Recursive Neural Tensor Networks (RNTN, 45.7%45.7\%) by +2.6%+2.6\% absolute.
    3. Character-level features yield negligible gain over word-only SCNN on clean movie review prose (48.3%48.3\% vs. 48.3%48.3\% fine-grained; 85.7%85.7\% vs. 85.5%85.5\% binary).
    4. Unsupervised word embedding initialization provides approximately +1.5%+1.5\% absolute accuracy gain over random initialization on SSTb.
  8. Knowl 8 — Sentiment Classification Accuracy on Stanford Twitter Sentiment (STS)

    empirical result

    CharSCNN and SCNN were evaluated on the binary (positive/negative) test set of the Stanford Twitter Sentiment (STS) corpus, using a training subset of 80,000 tweets and development set of 16,000 tweets:

    Model Accuracy (%) [Unsup. Pre-training] Accuracy (%) [Random Embeddings]
    CharSCNN 86.4 81.9
    SCNN 85.2 82.2
    LProp (Speriosu et al., 2011) 84.7 –
    MaxEnt (Go et al., 2009) 83.0 –
    NB (Go et al., 2009) 82.7 –
    SVM (Go et al., 2009) 82.2 –

    Key empirical findings on STS:

    1. CharSCNN established a new state-of-the-art accuracy on STS of 86.4%86.4\%, outperforming prior methods including Label Propagation (LProp, 84.7%84.7\%), Maximum Entropy (83.0%83.0\%), Naive Bayes (82.7%82.7\%), and SVM (82.2%82.2\%).
    2. Character-level representations provide a +1.2%+1.2\% absolute improvement over word-only SCNN (86.4%86.4\% vs. 85.2%85.2\%) on Twitter text, demonstrating the value of sub-word features for hashtags, slang, and morphological variations.
    3. Unsupervised pre-training of word embeddings provides a +4.5%+4.5\% absolute accuracy gain compared to random word vector initialization (86.4%86.4\% vs. 81.9%81.9\%).
  9. Knowl 9 — Negation Mechanism and Feature Attribution in Sentence Max-Pooling

    empirical result

    Analysis of the 300 sentence-level feature activations selected by the global max-pooling operator demonstrates that CharSCNN handles semantic negation compositionally without parse tree inputs:

    1. In the positive sentence "I liked every single minute of this film.", max-pooled features concentrate primarily on the sentiment carrier "liked" (~68 features) and the topic "film" (~40 features).
    2. In the negated counterpart "I did n't like a single minute of this film.", the impact of "like" drops sharply to ~4 features, while the negation phrase "did n't" dominates feature extraction (~50 features for "did" and ~60 for "n't").
    3. In the strongly negative phrase "It 's just incredibly dull", the expression "incredibly dull" accounts for 69%69\% of all extracted max-pooled features.
    4. In its negated counterpart "It 's definitely not dull", the phrase "definitely not dull" accounts for 77%77\% of all selected features, shifting the sentence score to positive.

    This confirms that convolutional filters paired with max-pooling dynamically reallocate feature saliency to negation modifiers and suppress the isolated polarity of modified terms.

Coverage note — None was omitted; all key contributions—including model architecture, mathematical formulation, training objective, unsupervised embedding initialization, hyperparameter configurations, SSTb and STS empirical results, and negation feature analysis—have been captured.

References

  1. 1.Andrei Alexandrescu and Katrin Kirchhoff. 2006. Factored neural language models. In Proceedings of the Human Language Technology Conference of the NAACL, pages 1–4, New York City, USA, June.
  2. 2.Luciano Barbosa and Junlan Feng. 2010. Robust sentiment detection on twitter from biased and noisy data. In Proceedings of the 23rd International Conference on Computational Linguistics, pages 36–44.
  3. 3.James Bergstra, Olivier Breuleux, Frédéric Bastien, Pascal Lamblin, Razvan Pascanu, Guillaume Desjardins, Joseph Turian, David Warde-Farley, and Yoshua Bengio. 2010. Theano: a CPU and GPU math expression compiler. In Proceedings of the Python for Scientific Computing Conference (SciPy).
  4. 4.Grzegorz Chrupala. 2013. Text segmentation with character-level text embeddings. In Proceedings of the ICML workshop on Deep Learning for Audio, Speech and Language Processing.
  5. 5.R. Collobert, J. Weston, L. Bottou, M. Karlen, K. Kavukcuoglu, and P. Kuksa. 2011. Natural language processing (almost) from scratch. Journal of Machine Learning Research, 12:2493–2537.
  6. 6.R. Collobert. 2011. Deep learning for efficient discriminative parsing. In Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics (AISTATS), pages 224–232.
  7. 7.Cícero Nogueira dos Santos and Bianca Zadrozny. 2014. Learning character-level representations for part-of-speech tagging. In Proceedings of the 31st International Conference on Machine Learning, JMLR: W&CP volume 32, Beijing, China.
  8. 8.Alec Go, Richa Bhayani, and Lei Huang. 2009. Twitter sentiment classification using distant supervision. Technical report, Stanford University.
  9. 9.Angeliki Lazaridou, Marco Marelli, Roberto Zamparelli, and Marco Baroni. 2013. Compositional–ly derived representations of morphologically complex words in distributional semantics. In Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (ACL), pages 1517–1526.
  10. 10.Yann Lecun, Lon Bottou, Yoshua Bengio, and Patrick Haffner. 1998. Gradient-based learning applied to document recognition. In Proceedings of the IEEE, pages 2278–2324.
  11. 11.Minh-Thang Luong, Richard Socher, and Christopher D. Manning. 2013. Better word representations with recursive neural networks for morphology. In Proceedings of the Conference on Computational Natural Language Learning, Sofia, Bulgaria.
  12. 12.Christopher D. Manning. 2011. Part-of-speech tagging from 97% to 100%: Is it time for some linguistics? In Proceedings of the 12th International Conference on Computational Linguistics and Intelligent Text Processing, CICLing’11, pages 171–189.
  13. 13.Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient estimation of word representations in vector space. In Proceedings of Workshop at International Conference on Learning Representations.
  14. 14.Preslav Nakov, Sara Rosenthal, Zornitsa Kozareva, Veselin Stoyanov, Alan Ritter, and Theresa Wilson. 2013. Semeval-2013 task 2: Sentiment analysis in twitter. In Second Joint Conference on Lexical and Computational Semantics (*SEM), Volume 2: Proceedings of the Seventh International Workshop on Semantic Evaluation (SemEval 2013), pages 312–320, Atlanta, Georgia, USA, June. Association for Computational Linguistics.
  15. 15.Richard Socher, Jeffrey Pennington, Eric H. Huang, Andrew Y. Ng, and Christopher D. Manning. 2011. Semi-supervised recursive autoencoders for predicting sentiment distributions. In Proceedings of the Conference on Empirical Methods in Natural Language Processing, pages 151–161.
  16. 16.Richard Socher, Brody Huval, Christopher D. Manning, and Andrew Y. Ng. 2012. Semantic compositionality through recursive matrix-vector spaces. In Proceedings of theConference on Empirical Methods in Natural Language Processing, pages 1201–1211.
  17. 17.Richard Socher, John Bauer, Christopher D. Manning, and Andrew Y. Ng. 2013a. Parsing with compositional vector grammars. In Proceedings of the Annual Meeting of the Association for Computational Linguistics.
  18. 18.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Y. Ng, and Christopher Potts. 2013b. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the Conference on Empirical Methods in Natural Language Processing, pages 1631–1642.
  19. 19.Michael Speriosu, Nikita Sudan, Sid Upadhyay, and Jason Baldridge. 2011. Twitter polarity classification with label propagation over lexical links and the follower graph. In Proceedings of the First Workshop on Unsupervised Learning in NLP, EMNLP, pages 53–63.
  20. 20.A. Waibel, T. Hanazawa, G. Hinton, K. Shikano, and K. J. Lang. 1989. Phoneme recognition using time-delay neural networks. IEEE Transactions on Acoustics, Speech and Signal Processing, 37(3):328–339.
  21. 21.Xiaoqing Zheng, Hanyang Chen, and Tianyu Xu. 2013. Deep learning for chinese word segmentation and pos tagging. In Proceedings of the Conference on Empirical Methods in NLP, pages 647–657.

Citation

MLA
Santos, C. dos ., and M. Gatti. “Deep Convolutional Neural Networks for Sentiment Analysis of Short Texts”. Proceedings of COLING 2014, the 25th International Conference on Computational Linguistics: Technical Papers, 2014, pp. 69–78, https://aclanthology.org/C14-1008/.
APA
Santos, C. dos ., & Gatti, M. (2014). Deep Convolutional Neural Networks for Sentiment Analysis of Short Texts. Proceedings of COLING 2014, the 25th International Conference on Computational Linguistics: Technical Papers, 69–78. https://aclanthology.org/C14-1008/
Chicago
Santos, C. dos ., and M. Gatti. 2014. “Deep Convolutional Neural Networks for Sentiment Analysis of Short Texts”. Proceedings of COLING 2014, the 25th International Conference on Computational Linguistics: Technical Papers, 69–78. https://aclanthology.org/C14-1008/.
Harvard
Santos, C. dos . and Gatti, M. (2014) “Deep Convolutional Neural Networks for Sentiment Analysis of Short Texts”, Proceedings of COLING 2014, the 25th International Conference on Computational Linguistics: Technical Papers. Association for Computational Linguistics, pp. 69–78. Available at: https://aclanthology.org/C14-1008/.
Vancouver
1. Santos C dos, Gatti M (2014) Deep Convolutional Neural Networks for Sentiment Analysis of Short Texts. In: Proceedings of COLING 2014, the 25th International Conference on Computational Linguistics: Technical Papers. Association for Computational Linguistics, pp 69–78

BibTeX

@inproceedings{dos-santos-gatti-2014-deep,
    title = "Deep Convolutional Neural Networks for Sentiment Analysis of Short Texts",
    author = "dos Santos, C{\'i}cero  and
      Gatti, Ma{\'i}ra",
    editor = "Tsujii, Junichi  and
      Hajic, Jan",
    booktitle = "Proceedings of {COLING} 2014, the 25th International Conference on Computational Linguistics: Technical Papers",
    month = aug,
    year = "2014",
    address = "Dublin, Ireland",
    publisher = "Dublin City University and Association for Computational Linguistics",
    url = "https://aclanthology.org/C14-1008/",
    pages = "69--78"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/