Document Modeling with Gated Recurrent Neural Network for Sentiment Classification

Duyu TangBing QinTing Liu

article2015EMNLP1,501 citations

Proposes a hierarchical neural architecture that combines sentence-level convolutional or recurrent encoders with gated recurrent networks to model inter-sentence relationships, outperforming traditional baselines on large-scale document sentiment classification benchmarks.

Listen

Understanding customer sentiment from full-text online reviews is critical for modern business intelligence, yet automated document-level classification remains challenging. Traditional machine learning tools rely on sparse word counts that ignore how adjacent sentences relate to one another, while standard neural sequence approaches often struggle to retain context across long multi-sentence texts.

The article demonstrates a bottom-up neural network framework that learns continuous document representations to accurately predict overall review sentiment. The primary objective is to evaluate whether hierarchical composition—encoding word meanings into sentences, and then chaining sentences through gated neural connections—significantly improves classification performance over established baselines.

The researchers evaluated their approach across four large-scale review datasets: three Yelp restaurant collections (from 2013, 2014, and 2015) and a 10-class IMDB movie review dataset, encompassing over 3.3 million documents in total. The model first builds sentence representations using either convolutional filters or long short-term memory units, and then processes those sentence vectors sequentially with a gated recurrent neural network. This end-to-end framework was benchmarked against strong support vector machine classifiers, existing neural embeddings, and standard recurrent networks using classification accuracy and mean squared error.

The findings show that the proposed gated neural approach consistently outperforms all baseline models across every dataset. In particular, pairing sentence-level long short-term memory with a document-level gated recurrent network achieved the highest classification accuracy (up to 67.6% on Yelp and 45.3% on IMDB) and substantially lower prediction error. Additionally, gated networks dramatically outperformed standard recurrent networks, which degraded severely due to information loss over long sentence sequences. Across configurations, sentence-level long short-term memory modestly outperformed convolutional encoders.

These results confirm that capturing inter-sentence relationships without manual feature engineering yields superior predictive accuracy. For engineering teams and decision-makers, adopting hierarchical gated neural architectures reduces reliance on expensive, hand-crafted feature pipelines while improving the fidelity of customer sentiment analytics. Standard recurrent architectures should be avoided for document-level modeling due to severe decay in long-term sequence processing.

Organizations implementing document-level sentiment systems should transition to hierarchical neural pipelines that combine sentence encoders with gated recurrent networks. Future research should explore integrating explicit sentiment-sensitive discourse relations or tree-structured models to further capture complex document structures at scale. Because the models were tested exclusively on English restaurant and movie reviews with reliable numerical rating labels, stakeholders should validate performance when applying the framework to domains with unmatched ratings, informal short texts, or different languages.

Tang et al (2015).pdf
Cover for Document Modeling with Gated Recurrent Neural Network for Sentiment Classification

Abstract

Document level sentiment classification remains a challenge: encoding the intrinsic relations between sentences in the semantic meaning of a document. To address this, we introduce a neural network model to learn vector-based document representation in a unified, bottom-up fashion. The model first learns sentence representation with convolutional neural network or long short-term memory. Afterwards, semantics of sentences and their relations are adaptively encoded in document representation with gated recurrent neural network. We conduct document level sentiment classification on four large-scale review datasets from IMDB and Yelp Dataset Challenge. Experimental results show that: (1) our neural model shows superior performances over several state-of-the-art algorithms; (2) gated recurrent neural network dramatically outperforms standard recurrent neural network in document modeling for sentiment classification.

Table of Contents

  • 1 Introduction
  • 2 The Approach
  • 2.1 Sentence Composition
  • 2.2 Document Composition with Gated Recurrent Neural Network
  • 2.3 Sentiment Classification
  • 3 Experiment
  • 3.1 Experimental Setting
  • 3.2 Baseline Methods
  • 3.3 Comparison to Other Methods
  • 3.4 Model Analysis
  • 4 Related Work
  • 5 Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Gated Recurrent Neural Network for Document Composition

    equation

    To compose a sequence of sentence representations (s1,s2,…,sN)(s_1, s_2, \dots, s_N) into a fixed-length document vector while avoiding the vanishing/exploding gradient problems of standard recurrent neural networks, a gated recurrent neural network (GRNN) is applied sequentially across sentences. At sentence step tt, given the previous hidden state vector ht−1∈Rlhh_{t-1} \in \mathbb{R}^{l_h} and the current sentence representation vector st∈Rlocs_t \in \mathbb{R}^{l_{oc}}, the transition equations are:

    it=σ(Wi⋅[ht−1;st]+bi)i_t = \sigma(W_i \cdot [h_{t-1}; s_t] + b_i)

    ft=σ(Wf⋅[ht−1;st]+bf)f_t = \sigma(W_f \cdot [h_{t-1}; s_t] + b_f)

    gt=tanh⁡(Wr⋅[ht−1;st]+br)g_t = \tanh(W_r \cdot [h_{t-1}; s_t] + b_r)

    ht=tanh⁡(it⊙gt+ft⊙ht−1)h_t = \tanh(i_t \odot g_t + f_t \odot h_{t-1})

    where σ(⋅)\sigma(\cdot) denotes the element-wise sigmoid function, tanh⁡(⋅)\tanh(\cdot) denotes hyperbolic tangent, ⊙\odot represents element-wise multiplication, and [ht−1;st]∈Rlh+loc[h_{t-1}; s_t] \in \mathbb{R}^{l_h + l_{oc}} is the concatenation of the previous hidden state and the current sentence input. The learnable parameters are weight matrices Wi,Wf,Wr∈Rlh×(lh+loc)W_i, W_f, W_r \in \mathbb{R}^{l_h \times (l_h + l_{oc})} and bias vectors bi,bf,br∈Rlhb_i, b_f, b_r \in \mathbb{R}^{l_h}.

    The input gate iti_t and forget gate ftf_t adaptively regulate the incorporation of current sentence semantics and the retention of historical context. This formulation functions as an LSTM without an output gate (i.e., with the output gate permanently active), ensuring that no accumulated sentence semantics are suppressed during composition.

    From the sequence of hidden vectors (h1,…,hN)(h_1, \dots, h_N), the final document representation vector ddocd_{\text{doc}} can be constructed in three ways:

    1. Standard Sequential (GatedNN): Taking the last hidden vector, ddoc=hNd_{\text{doc}} = h_N.
    2. Averaged GatedNN (GatedNN Avg): Taking the mean across all hidden states, ddoc=1N∑t=1Nhtd_{\text{doc}} = \frac{1}{N} \sum_{t=1}^N h_t.
    3. Bidirectional Averaged GatedNN (Bi GatedNN Avg): Computing forward hidden vectors h⃗t\vec{h}_t and backward hidden vectors h←t\overleftarrow{h}_t, concatenating them, and averaging across all time steps.
  2. Knowl 2 — Hierarchical Bottom-Up Document Sentiment Modeling Framework

    model/method

    The document sentiment classification model operates in a two-stage hierarchical bottom-up process based on the principle of compositionality:

    1. Sentence Composition Stage: Each sentence in a document is mapped from its sequence of word embeddings into a continuous, fixed-length sentence vector using either a Convolutional Neural Network (Conv-GRNN) with multiple filter granularities or a Long Short-Term Memory network (LSTM-GRNN).
    2. Document Composition Stage: The resulting sequence of sentence vectors is fed into a Gated Recurrent Neural Network (GRNN) that adaptively models inter-sentence semantic relations (e.g., contrast and cause) and long-distance dependencies across the document to generate a fixed-length document representation.

    The document representation is directly mapped to class probabilities via a linear transformation and a softmax layer, allowing the entire model (word embeddings, sentence encoder, document composition GRNN, and classifier) to be trained end-to-end via stochastic gradient descent using cross-entropy loss.

  3. Knowl 3 — Multi-Granularity Convolutional Sentence Composition

    model/method

    To produce a continuous sentence vector from variable-length word sequences without relying on external syntactic parsers, a Convolutional Neural Network (CNN) with multiple filter widths is used.

    Let a sentence of nn words be represented by its word embeddings (e1,e2,…,en)(e_1, e_2, \dots, e_n), where each ei∈Rde_i \in \mathbb{R}^d and dd is the embedding dimension. For a convolutional filter of window width lcl_c, the input window at position ii is the concatenated embedding vector Ic=[ei;ei+1;… ;ei+lc−1]∈Rd⋅lcI_c = [e_i; e_{i+1}; \dots; e_{i+l_c-1}] \in \mathbb{R}^{d \cdot l_c}. The linear transformation produces a local feature map:

    Oc=Wc⋅Ic+bcO_c = W_c \cdot I_c + b_c

    where Wc∈Rloc×(d⋅lc)W_c \in \mathbb{R}^{l_{oc} \times (d \cdot l_c)}, bc∈Rlocb_c \in \mathbb{R}^{l_{oc}}, and locl_{oc} is the output feature dimension.

    To capture global sentence semantics, the linear layer outputs across all valid positions in the sentence are aggregated using an average pooling layer, followed by a pointwise tanh⁡\tanh activation function. To capture local semantics across multiple nn-gram granularities simultaneously, three distinct convolutional filters with widths lc∈{1,2,3}l_c \in \{1, 2, 3\} (corresponding to unigrams, bigrams, and trigrams) are computed in parallel, and their pooled non-linear outputs are averaged to yield the final sentence vector s∈Rlocs \in \mathbb{R}^{l_{oc}}.

  4. Knowl 4 — Document Sentiment Classification and Loss Function

    equation

    Given a document representation vector ddoc∈Rlhd_{\text{doc}} \in \mathbb{R}^{l_h} produced by the document composition component, the model projects it to the class space using a linear classification layer followed by a softmax function. The predicted probability distribution over CC sentiment classes is given by:

    Pi(d)=exp⁡(xi)∑j=1Cexp⁡(xj)P_i(d) = \frac{\exp(x_i)}{\sum_{j=1}^C \exp(x_j)}

    where x=Wsoftmax⋅ddoc+bsoftmax∈RCx = W_{\text{softmax}} \cdot d_{\text{doc}} + b_{\text{softmax}} \in \mathbb{R}^C, with weight matrix Wsoftmax∈RC×lhW_{\text{softmax}} \in \mathbb{R}^{C \times l_h} and bias vector bsoftmax∈RCb_{\text{softmax}} \in \mathbb{R}^C.

    The model is trained in a supervised setting to minimize the cross-entropy loss between the true one-hot distribution Pg(d)P^g(d) and the predicted distribution P(d)P(d) across all training documents TT:

    loss=−∑d∈T∑i=1CPig(d)log⁡(Pi(d))\text{loss} = -\sum_{d \in T} \sum_{i=1}^C P_i^g(d) \log(P_i(d))

    where Pig(d)=1P_i^g(d) = 1 if class ii is the ground-truth sentiment label for document dd, and 00 otherwise. Optimization is performed end-to-end with respect to all network parameters θ=[Wc;bc;Wi;bi;Wf;bf;Wr;br;Wsoftmax;bsoftmax]\theta = [W_c; b_c; W_i; b_i; W_f; b_f; W_r; b_r; W_{\text{softmax}}; b_{\text{softmax}}] using stochastic gradient descent.

  5. Knowl 5 — Experimental Datasets and Setup for Document Sentiment Classification

    experimental setup

    The models are evaluated on four large-scale review datasets:

    • Yelp 2013: 335,018 reviews, average 8.90 sentences/doc, average 151.6 words/doc, vocabulary size 211,245, 5 sentiment classes (class distribution: .09/.09/.14/.33/.36).
    • Yelp 2014: 1,125,457 reviews, average 9.22 sentences/doc, average 156.9 words/doc, vocabulary size 476,191, 5 sentiment classes (class distribution: .10/.09/.15/.30/.36).
    • Yelp 2015: 1,569,264 reviews, average 8.97 sentences/doc, average 151.9 words/doc, vocabulary size 612,636, 5 sentiment classes (class distribution: .10/.09/.14/.30/.37).
    • IMDB: 348,415 movie reviews, average 14.02 sentences/doc, average 325.6 words/doc, vocabulary size 115,831, 10 sentiment classes (class distribution: .07/.04/.05/.05/.08/.11/.15/.17/.12/.18).

    Data Splits and Preprocessing: Yelp datasets are partitioned into 80% training, 10% development, and 10% testing. The IMDB dataset uses the standard split from Diao et al. (2014). Tokenization and sentence splitting are performed using Stanford CoreNLP.

    Hyperparameters and Training Configuration:

    • Word representations: 200-dimensional embeddings trained separately on each corpus using SkipGram (d=200d = 200).
    • CNN sentence encoder: Three convolutional filters of widths 1, 2, and 3; output length loc=50l_{oc} = 50.
    • Initialization: Parameters other than pre-trained word embeddings are randomly initialized from a uniform distribution U(−0.01,0.01)\mathcal{U}(-0.01, 0.01).
    • Optimization: Stochastic Gradient Descent (SGD) with a learning rate of 0.03.

    Evaluation Metrics:

    • Accuracy: Standard overall classification accuracy.
    • Mean Squared Error (MSE): Measures the divergence between predicted and ground-truth sentiment labels, accounting for rating scale distance: MSE=∑i=1N(goldi−predictedi)2N\text{MSE} = \frac{\sum_{i=1}^N (\text{gold}_i - \text{predicted}_i)^2}{N}
  6. Knowl 6 — Sentiment Classification Performance Across Benchmark Datasets

    data/table

    Document-level sentiment classification performance of the proposed Conv-GRNN and LSTM-GRNN models against competitive baselines (majority class baseline, linear SVMs with bag-of-unigrams, bag-of-bigrams, hand-crafted text features, averaged SkipGram embeddings, and sentiment-specific word embeddings SSWE; JMARS; Paragraph Vector; and flat CNN) on Yelp 2013, Yelp 2014, Yelp 2015, and IMDB.

    Method Yelp 2013 Yelp 2014 Yelp 2015 IMDB
    Accuracy MSE Accuracy MSE Accuracy MSE Accuracy MSE
    Majority 0.356 3.06 0.361 3.28 0.369 3.30 0.179 17.46
    SVM + Unigrams 0.589 0.79 0.600 0.78 0.611 0.75 0.399 4.23
    SVM + Bigrams 0.576 0.75 0.616 0.65 0.624 0.63 0.409 3.74
    SVM + TextFeatures 0.598 0.68 0.618 0.63 0.624 0.60 0.405 3.56
    SVM + AverageSG 0.543 1.11 0.557 1.08 0.568 1.04 0.319 5.57
    SVM + SSWE 0.535 1.12 0.543 1.13 0.554 1.11 0.262 9.16
    JMARS N/A – N/A – N/A – N/A 4.97
    Paragraph Vector 0.577 0.86 0.592 0.70 0.605 0.61 0.341 4.69
    Convolutional NN 0.597 0.76 0.610 0.68 0.615 0.68 0.376 3.30
    Conv-GRNN 0.637 0.56 0.655 0.51 0.660 0.50 0.425 2.71
    LSTM-GRNN 0.651 0.50 0.671 0.48 0.676 0.49 0.453 3.00

    Key Observations:

    1. Both Conv-GRNN and LSTM-GRNN outperform all baseline methods (including strong n-gram SVMs and Paragraph Vector) across all four datasets in both Accuracy and MSE.
    2. LSTM-GRNN consistently achieves the highest classification accuracy on all datasets (0.651 on Yelp 2013, 0.671 on Yelp 2014, 0.676 on Yelp 2015, and 0.453 on IMDB), demonstrating that LSTM sentence encoders capture intra-sentence composition slightly better than multi-filter CNNs.
    3. Simple word embedding averaging (SVM + AverageSG) performs poorly (e.g., 0.543 on Yelp 2013, 0.319 on IMDB), confirming that compositional semantic modeling is necessary for document-level sentiment classification.
  7. Knowl 7 — Ablation Analysis of Document Composition Mechanisms and Recurrent Gating

    data/table

    To evaluate the specific impact of document composition strategies, seven variants built on top of CNN sentence representations were compared on Yelp 2013, Yelp 2014, Yelp 2015, and IMDB:

    • Average: Element-wise average of all sentence vectors in the document.
    • Recurrent: Standard RNN transition ht=tanh⁡(Wr[ht−1;st]+br)h_t = \tanh(W_r [h_{t-1}; s_t] + b_r) where document representation is the final hidden state hNh_N.
    • Recurrent Avg: Standard RNN with document representation as the average of all hidden states 1N∑t=1Nht\frac{1}{N}\sum_{t=1}^N h_t.
    • Bi Recurrent Avg: Bidirectional standard RNN with averaged hidden states.
    • GatedNN: Gated RNN taking the last hidden state hNh_N.
    • GatedNN Avg: Gated RNN taking the average of all hidden states.
    • Bi GatedNN Avg: Bidirectional Gated RNN taking the average of forward and backward hidden states.
    Model Setting Yelp 2013 Yelp 2014 Yelp 2015 IMDB
    Accuracy MSE Accuracy MSE Accuracy MSE Accuracy MSE
    Average 0.598 0.65 0.605 0.75 0.614 0.67 0.366 3.91
    Recurrent 0.377 1.37 0.306 1.75 0.383 1.67 0.176 12.29
    Recurrent Avg 0.582 0.69 0.591 0.70 0.597 0.74 0.344 3.71
    Bi Recurrent Avg 0.587 0.73 0.597 0.73 0.577 0.82 0.372 3.32
    GatedNN 0.636 0.58 0.656 0.52 0.651 0.51 0.430 2.95
    GatedNN Avg 0.635 0.57 0.659 0.52 0.657 0.56 0.416 2.78
    Bi GatedNN Avg 0.637 0.56 0.655 0.51 0.660 0.50 0.425 2.71

    Key Findings:

    1. Standard RNN Failure: Standard sequential RNN performs drastically worse than simple sentence vector averaging (e.g., accuracy drops from 0.598 to 0.377 on Yelp 2013, and from 0.366 to 0.176 on IMDB). This is caused by severe vanishing gradients over long sequences, which erase semantic contributions from earlier sentences.
    2. Mitigation via Hidden State Averaging: Averaging standard RNN hidden states (Recurrent Avg) substantially improves upon the last hidden state (from 0.377 to 0.582 on Yelp 2013), but still fails to outperform naive sentence averaging (0.598).
    3. Effectiveness of Gating: Incorporating input and forget gates (GatedNN, GatedNN Avg, Bi GatedNN Avg) dramatically resolves the gradient vanishing issue, increasing accuracy by 25–35 percentage points over standard Recurrent models and outperforming simple sentence averaging by 4–6% across all datasets.

Coverage note — No substantial contributed material was omitted.

References

  1. 1.Marco Baroni, Georgiana Dinu, and German´ Kruszewski. 2014. Don’t count, predict! a systematic comparison of context-counting vs. context-predicting semantic vectors. In ACL, pages 238–247.
  2. 2.Yoshua Bengio, Patrice Simard, and Paolo Frasconi. 1994. Learning long-term dependencies with gradient descent is difficult. Neural Networks, IEEE Transactions on, 5(2):157–166.
  3. 3.Yoshua Bengio, Rejean Ducharme, Pascal Vincent, and ´ Christian Janvin. 2003. A neural probabilistic language model. Journal of Machine Learning Research, 3:1137–1155.
  4. 4.Yoshua Bengio, Ian J. Goodfellow, and Aaron Courville. 2015. Deep learning. Book in preparation for MIT Press.
  5. 5.Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014. Learning phrase representations using rnn encoder–decoder for statistical machine translation. In EMNLP, pages 1724–1734.
  6. 6.Junyoung Chung, Caglar Gulcehre, Kyunghyun Cho, and Yoshua Bengio. 2015. Gated feedback recurrent neural networks. ICML.
  7. 7.Misha Denil, Alban Demiraj, Nal Kalchbrenner, Phil Blunsom, and Nando de Freitas. 2014. Modelling, visualising and summarising documents with a single convolutional neural network. arXiv preprint:1406.3830.
  8. 8.Qiming Diao, Minghui Qiu, Chao-Yuan Wu, Alexander J Smola, Jing Jiang, and Chong Wang. 2014. Jointly modeling aspects, ratings and sentiments for movie recommendation (jmars). In SIGKDD, pages 193–202. ACM.
  9. 9.Li Dong, Furu Wei, Ming Zhou, and Ke Xu. 2014. Adaptive multi-compositionality for recursive neural models with applications to sentiment analysis. In AAAI, pages 1537–1543.
  10. 10.Rong-En Fan, Kai-Wei Chang, Cho-Jui Hsieh, Xiang-Rui Wang, and Chih-Jen Lin. 2008. Liblinear: A library for large linear classification. JMLR.
  11. 11.Gottlob Frege. 1892. On sense and reference. Ludlow (1997), pages 563–584.
  12. 12.Gayatree Ganu, Noemie Elhadad, and Amelie Marian. ´ 2009. Beyond the stars: Improving rating predictions using review text content. In WebDB.
  13. 13.Xavier Glorot, Antoine Bordes, and Yoshua Bengio. 2011. Domain adaptation for large-scale sentiment classification: A deep learning approach. In ICML, pages 513–520.
  14. 14.Andrew B Goldberg and Xiaojin Zhu. 2006. Seeing stars when there aren’t many stars: graph-based semi-supervised learning for sentiment categorization. In GraphBased Method for NLP, pages 45–52.
  15. 15.Alex Graves, Navdeep Jaitly, and A-R Mohamed. 2013. Hybrid speech recognition with deep bidirectional lstm. In Automatic Speech Recognition and Understanding (ASRU), 2013 IEEE Workshop on, pages 273–278. IEEE.
  16. 16.Karl Moritz Hermann and Phil Blunsom. 2013. The role of syntax in vector space models of compositional semantics. In ACL, pages 894–904.
  17. 17.Sepp Hochreiter and Jurgen Schmidhuber. 1997. ¨ Long short-term memory. Neural computation, 9(8):1735–1780.
  18. 18.Ozan Irsoy and Claire Cardie. 2014. Deep recursive neural networks for compositionality in language. In NIPS, pages 2096–2104.
  19. 19.Rie Johnson and Tong Zhang. 2015. Effective use of word order for text categorization with convolutional neural networks. NAACL.
  20. 20.Dan Jurafsky and James H Martin. 2000. Speech & language processing. Pearson Education India.
  21. 21.Nal Kalchbrenner, Edward Grefenstette, and Phil Blunsom. 2014. A convolutional neural network for modelling sentences. In ACL, pages 655–665.
  22. 22.Yoon Kim. 2014. Convolutional neural networks for sentence classification. In EMNLP, pages 1746–1751.
  23. 23.Svetlana Kiritchenko, Xiaodan Zhu, and Saif M Mohammad. 2014. Sentiment analysis of short informal texts. Journal of Artificial Intelligence Research, pages 723–762.
  24. 24.Igor Labutov and Hod Lipson. 2013. Re-embedding words. In Annual Meeting of the Association for Computational Linguistics.
  25. 25.Quoc V. Le and Tomas Mikolov. 2014. Distributed representations of sentences and documents. In ICML, pages 1188–1196.
  26. 26.Jiwei Li, Dan Jurafsky, and Eudard Hovy. 2015a. When are tree structures necessary for deep learning of representations? arXiv preprint arXiv:1503.00185.
  27. 27.Jiwei Li, Minh-Thang Luong, and Dan Jurafsky. 2015b. A hierarchical neural autoencoder for paragraphs and documents. arXiv preprint arXiv:1506.01057.
  28. 28.Jiwei Li. 2014. Feature weight tuning for recursive neural networks. Arxiv preprint, 1412.3714.
  29. 29.Bing Liu. 2012. Sentiment analysis and opinion mining. Synthesis Lectures on Human Language Technologies, 5(1):1–167.
  30. 30.Andrew L Maas, Raymond E Daly, Peter T Pham, Dan Huang, Andrew Y Ng, and Christopher Potts. 2011. Learning word vectors for sentiment analysis. In ACL, pages 142–150.
  31. 31.Christopher D Manning and Hinrich Schutze. 1999. ¨ Foundations of statistical natural language processing. MIT press.
  32. 32.Christopher Manning, Mihai Surdeanu, John Bauer, Jenny Finkel, Steven Bethard, and David McClosky. 2014. The stanford corenlp natural language processing toolkit. In ACL, pages 55–60.
  33. 33.Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013. Distributed representations of words and phrases and their compositionality. In NIPS, pages 3111–3119.
  34. 34.Jeff Mitchell and Mirella Lapata. 2010. Composition in distributional models of semantics. Cognitive Science, 34(8):1388–1429.
  35. 35.Georgios Paltoglou and Mike Thelwall. 2010. A study of information retrieval weighting schemes for sentiment analysis. In Proceedings of Annual Meeting of the Association for Computational Linguistics, pages 1386–1395.
  36. 36.Bo Pang and Lillian Lee. 2005. Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales. In ACL, pages 115–124.
  37. 37.Bo Pang and Lillian Lee. 2008. Opinion mining and sentiment analysis. Foundations and trends in information retrieval, 2(1-2):1–135.
  38. 38.Bo Pang, Lillian Lee, and Shivakumar Vaithyanathan. 2002. Thumbs up?: sentiment classification using machine learning techniques. In EMNLP, pages 79–86.
  39. 39.Romain Paulus, Richard Socher, and Christopher D Manning. 2014. Global belief recursive neural networks. In NIPS, pages 2888–2896.
  40. 40.Jeffrey Pennington, Richard Socher, and Christopher D Manning. 2014. Glove: Global vectors for word representation. In EMNLP, pages 1532–1543.
  41. 41.Lizhen Qu, Georgiana Ifrim, and Gerhard Weikum. 2010. The bag-of-opinions method for review rating prediction from sparse text patterns. In COLING, pages 913–921.
  42. 42.Richard Socher, John Bauer, Christopher D. Manning, and Andrew Y. Ng. 2013a. Parsing with compositional vector grammars. In ACL.
  43. 43.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013b. Recursive deep models for semantic compositionality over a sentiment treebank. In EMNLP, pages 1631–1642.
  44. 44.Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014. Sequence to sequence learning with neural networks. In NIPS, pages 3104–3112.
  45. 45.Kai Sheng Tai, Richard Socher, and Christopher D Manning. 2015. Improved semantic representations from tree-structured long short-term memory networks. In ACL.
  46. 46.Duyu Tang, Furu Wei, Nan Yang, Ming Zhou, Ting Liu, and Bing Qin. 2014. Learning sentiment-specific word embedding for twitter sentiment classification. In ACL, pages 1555–1565.
  47. 47.Duyu Tang, Bing Qin, and Ting Liu. 2015. Learning semantic representations of users and products for document level sentiment classification. In ACL, pages 1014–1023.
  48. 48.Peter D Turney. 2002. Thumbs up or thumbs down?: semantic orientation applied to unsupervised classification of reviews. In ACL, pages 417–424.
  49. 49.Sida Wang and Christopher D Manning. 2012. Baselines and bigrams: Simple, good sentiment and topic classification. In ACL, pages 90–94.
  50. 50.Rui Xia and Chengqing Zong. 2010. Exploring the use of word relation features for sentiment classification. In COLING, pages 1336–1344.
  51. 51.Liheng Xu, Kang Liu, and Jun Zhao. 2014. Joint opinion relation detection using one-class deep neural network. In COLING, pages 677–687.
  52. 52.Kelvin Xu, Jimmy Ba, Ryan Kiros, Aaron Courville, Ruslan Salakhutdinov, Richard Zemel, and Yoshua Bengio. 2015. Show, attend and tell: Neural image caption generation with visual attention. ICML.
  53. 53.Ainur Yessenalina and Claire Cardie. 2011. Compositional matrix-space models for sentiment analysis. In EMNLP, pages 172–182.
  54. 54.Wojciech Zaremba and Ilya Sutskever. 2014. Learning to execute. arXiv preprint arXiv:1410.4615.
  55. 55.Yongfeng Zhang, Haochen Zhang, Min Zhang, Yiqun Liu, and Shaoping Ma. 2014. Do users rate or review?: boost phrase-level sentiment labeling with review-level sentiment classification. In SIGIR, pages 1027–1030. ACM.
  56. 56.Han Zhao, Zhengdong Lu, and Pascal Poupart. 2015. Self-adaptive hierarchical sentence model. In IJCAI.
  57. 57.Lanjun Zhou, Binyang Li, Wei Gao, Zhongyu Wei, and Kam-Fai Wong. 2011. Unsupervised discovery of discourse relations for eliminating intra-sentence polarity ambiguities. In EMNLP, pages 162–171, .
  58. 58.Xiaodan Zhu, Parinaz Sobhani, and Hongyu Guo. 2015. Long short-term memory over tree structures. ICML.

Citation

MLA
Tang, D., et al. “Document Modeling with Gated Recurrent Neural Network for Sentiment Classification”. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, 2015, pp. 1422–32, https://doi.org/10.18653/v1/D15-1167.
APA
Tang, D., (秦兵), B. Q., & Liu, T. (2015). Document Modeling with Gated Recurrent Neural Network for Sentiment Classification. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, 1422–1432. https://doi.org/10.18653/v1/D15-1167
Chicago
Tang, D., B. Q. (秦兵), and T. Liu. 2015. “Document Modeling with Gated Recurrent Neural Network for Sentiment Classification”. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, 1422–32. https://doi.org/10.18653/v1/D15-1167.
Harvard
Tang, D., (秦兵), B.Q. and Liu, T. (2015) “Document Modeling with Gated Recurrent Neural Network for Sentiment Classification”, Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 1422–1432. Available at: https://doi.org/10.18653/v1/D15-1167.
Vancouver
1. Tang D, (秦兵) BQ, Liu T (2015) Document Modeling with Gated Recurrent Neural Network for Sentiment Classification. In: Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 1422–1432

BibTeX

@inproceedings{tang-etal-2015-document,
    title = "Document Modeling with Gated Recurrent Neural Network for Sentiment Classification",
    author = "Tang, Duyu  and
      Qin, Bing  and
      Liu, Ting",
    editor = "M{\`a}rquez, Llu{\'i}s  and
      Callison-Burch, Chris  and
      Su, Jian",
    booktitle = "Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing",
    month = sep,
    year = "2015",
    address = "Lisbon, Portugal",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/D15-1167/",
    doi = "10.18653/v1/D15-1167",
    pages = "1422--1432"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by-nc-sa/4.0/