Semi-Supervised Recursive Autoencoders for Predicting Sentiment Distributions

Richard SocherJeffrey PenningtonEric H. HuangAndrew Y. NgChristopher D. Manning

article2011EMNLP1,360 citations

Proposes semi-supervised recursive autoencoders that learn hierarchical vector representations of variable-length phrases directly from text to predict complex multi-label sentiment distributions without relying on handcrafted lexicons or parsers.

Listen

Understanding human sentiment in user-generated text—such as social media posts, blogs, and customer feedback—is increasingly vital for decision-makers. Traditional automated sentiment analysis models rely heavily on simple bag-of-words methods, which ignore word order and context, or depend on labor-intensive, hand-crafted linguistic resources such as sentiment dictionaries and rule-based parsing systems. Furthermore, most existing systems simplify sentiment into one-dimensional positive or negative ratings, failing to capture the rich, multi-dimensional emotional reactions present in everyday human communication.

The main objective of the article is to demonstrate a semi-supervised recursive autoencoder model that automatically learns phrase and sentence structures directly from text. It evaluates this framework on both standard binary sentiment tasks and the prediction of complex, multi-dimensional sentiment distributions without requiring pre-defined sentiment lexicons or hand-crafted linguistic rules.

The researchers developed a neural network architecture that begins with continuous word vector embeddings and greedily constructs hierarchical sentence representations by minimizing reconstruction error. To evaluate performance across different domains, the team tested the model on standard binary sentiment benchmarks—including movie reviews and the MPQA opinion dataset—as well as a newly analyzed dataset of over 31,000 anonymous personal confessions from the Experience Project. This confession dataset features user votes across five distinct emotional reactions: sympathy, approval, amusement, empathy, and shock. The model was evaluated on its ability to classify binary polarity, select the single most frequent emotional reaction, and accurately predict the complete proportional distribution across all five emotional categories.

The findings show that the proposed framework consistently outperforms competitive baseline models and established state-of-the-art approaches. First, on the Experience Project dataset, the model achieved 50.1% accuracy in predicting the top emotional reaction, surpassing a heavily engineered feature baseline using external sentiment lexicons and spelling normalizers by about 3.1 percentage points. Second, the model more accurately predicted full emotional probability distributions, lowering the average divergence error relative to word vector and bag-of-words baselines. Third, on binary benchmarks, the model reached 77.7% accuracy on movie reviews and 86.4% on opinion polarity, outperforming previous tree-based conditional random field models while eliminating reliance on external parsers or sentiment rules. Finally, the analysis revealed that balancing supervised classification with unsupervised structure reconstruction prevented overfitting, with optimal performance occurring when reconstruction error was weighted at twenty percent.

These results demonstrate that organizations can accurately extract nuanced, multi-faceted human emotions from text without investing substantial time and capital into building expensive, language-specific dictionaries and grammar rules. Because the framework learns semantic representations directly from raw text, it significantly reduces development costs, mitigates the risk of missing context-dependent meaning, and speeds up deployment across diverse domains. Operational efficiency is further supported by reasonable computational requirements: the model trained in 3 to 12 hours on standard 4-core hardware and performed inference on hundreds of texts within seconds.

Organizations seeking to analyze complex customer feedback or user sentiment should consider deploying recursive neural models over traditional keyword-matching and bag-of-words systems. For practical implementation, technical teams should leverage unsupervised pre-trained word embeddings and retain a modest reconstruction loss weighting to ensure model stability. If initial domain vocabularies contain rare or unseen terms, incorporating domain-specific lexicons into training can yield slight performance gains. Further work should explore applying this architecture across additional languages and multi-sentence document structures.

While the model delivers robust predictive performance across diverse benchmarks, stakeholders should note certain limitations. Performance on the confession dataset was evaluated specifically on entries with at least four user votes (6,129 entries) to ensure reliable ground truth, meaning predictions on extremely sparse or unvoted text may exhibit greater variance. Additionally, the greedy tree-construction algorithm approximates linguistic hierarchy for speed rather than strict syntactic grammar. Confidence in the reported results is high, as the findings are validated across multiple public datasets using cross-validation and standard statistical evaluation metrics.

Socher et al (2011).pdf
Cover for Semi-Supervised Recursive Autoencoders for Predicting Sentiment Distributions

Abstract

We introduce a novel machine learning framework based on recursive autoencoders for sentence-level prediction of sentiment label distributions. Our method learns vector space representations for multi-word phrases. In sentiment prediction tasks these representations outperform other state-of-the-art approaches on commonly used datasets, such as movie reviews, without using any pre-defined sentiment lexica or polarity shifting rules. We also evaluate the model's ability to predict sentiment distributions on a new dataset based on confessions from the experience project. The dataset consists of personal user stories annotated with multiple labels which, when aggregated, form a multinomial distribution that captures emotional reactions. Our algorithm can more accurately predict distributions over such labels compared to several competitive baselines.

Table of Contents

  • 1 Introduction
  • 2 Semi-Supervised Recursive Autoencoders
  • 2.1 Neural Word Representations
  • 2.2 Traditional Recursive Autoencoders
  • 2.3 Unsupervised Recursive Autoencoder for Structure Prediction
  • 2.4 Semi-Supervised Recursive Autoencoders
  • 3 Learning
  • 4 Experiments
  • 4.1 EP Dataset: The Experience Project
  • 4.2 EP: Predicting the Label with Most Votes
  • 4.3 EP: Predicting Sentiment Distributions
  • 4.4 Binary Polarity Classification
  • 4.5 Reconstruction vs. Classification Error
  • 5 Related Work
  • 5.1 Autoencoders and Deep Learning
  • 5.2 Sentiment Analysis
  • 6 Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Semi-Supervised Recursive Autoencoder Unit Architecture

    model/method

    A Semi-Supervised Recursive Autoencoder (RAE) unit computes a fixed-length parent vector representation from two nn-dimensional child vectors c1,c2∈Rnc_1, c_2 \in \mathbb{R}^n, reconstructs the children, and predicts a sentiment distribution over KK classes.

    The parent vector p∈Rnp \in \mathbb{R}^n is computed via an encoding layer followed by unit length normalization:

    p=f(W(1)[c1;c2]+b(1))∥f(W(1)[c1;c2]+b(1))∥p = \frac{f(W^{(1)}[c_1; c_2] + b^{(1)})}{\|f(W^{(1)}[c_1; c_2] + b^{(1)})\|}

    where [c1;c2]∈R2n[c_1; c_2] \in \mathbb{R}^{2n} is the concatenation of the two children, W(1)∈Rn×2nW^{(1)} \in \mathbb{R}^{n \times 2n} is the encoding weight matrix, b(1)∈Rnb^{(1)} \in \mathbb{R}^n is the bias term, and ff is an element-wise activation function such as tanh⁡\tanh. Normalizing pp to unit length prevents the network from artificially minimizing reconstruction error by shrinking the magnitude of hidden representations at higher levels of the tree.

    The reconstruction layer attempts to reproduce the original children from pp:

    [c1′;c2′]=W(2)p+b(2)[c'_1; c'_2] = W^{(2)}p + b^{(2)}

    where W(2)∈R2n×nW^{(2)} \in \mathbb{R}^{2n \times n} and b(2)∈R2nb^{(2)} \in \mathbb{R}^{2n}.

    A softmax classification layer attached to the parent vector pp outputs a predicted conditional multinomial label distribution d(p;θ)∈RKd(p; \theta) \in \mathbb{R}^K:

    d(p;θ)=softmax(Wlabelp)d(p; \theta) = \text{softmax}(W^{\text{label}} p)

    where Wlabel∈RK×nW^{\text{label}} \in \mathbb{R}^{K \times n} is the label parameter matrix, and ∑k=1Kdk(p;θ)=1\sum_{k=1}^K d_k(p; \theta) = 1.

  2. Knowl 2 — Joint Objective Function for Semi-Supervised Recursive Autoencoders

    equation

    The semi-supervised objective function JJ over a corpus of NN pairs of sentences and target label distributions (x,t)(x, t) combines tree reconstruction error and supervised cross-entropy error with L2L_2 parameter regularization:

    J=1N∑(x,t)E(x,t;θ)+λ2∥θ∥2J = \frac{1}{N} \sum_{(x, t)} E(x, t; \theta) + \frac{\lambda}{2} \|\theta\|^2

    where θ={W(1),b(1),W(2),b(2),Wlabel,L}\theta = \{W^{(1)}, b^{(1)}, W^{(2)}, b^{(2)}, W^{\text{label}}, L\} is the complete set of model parameters (including the word embedding matrix LL), λ\lambda is the regularization hyperparameter, and E(x,t;θ)E(x, t; \theta) sums the combined error over all non-terminal nodes ss of the tree T(RAEθ(x))\mathcal{T}(\text{RAE}_\theta(x)) constructed for sentence xx:

    E(x,t;θ)=∑s∈T(RAEθ(x))[αErec([c1;c2]s;θ)+(1−α)EcE(ps,t;θ)]E(x, t; \theta) = \sum_{s \in \mathcal{T}(\text{RAE}_\theta(x))} \left[ \alpha E_{\text{rec}}([c_1; c_2]_s; \theta) + (1 - \alpha) E_{cE}(p_s, t; \theta) \right]

    The reconstruction error ErecE_{\text{rec}} is weighted by the number of words n1n_1 and n2n_2 spanning underneath children c1c_1 and c2c_2, giving larger importance to subtrees that cover more words:

    Erec([c1;c2];θ)=n1n1+n2∥c1−c1′∥2+n2n1+n2∥c2−c2′∥2E_{\text{rec}}([c_1; c_2]; \theta) = \frac{n_1}{n_1 + n_2} \|c_1 - c'_1\|^2 + \frac{n_2}{n_1 + n_2} \|c_2 - c'_2\|^2

    where c1′,c2′c'_1, c'_2 are the reconstructed child vectors. The cross-entropy error EcEE_{cE} against the gold target multinomial distribution t∈RKt \in \mathbb{R}^K (where ∑k=1Ktk=1\sum_{k=1}^K t_k = 1) is defined as:

    EcE(p,t;θ)=−∑k=1Ktklog⁡dk(p;θ)E_{cE}(p, t; \theta) = -\sum_{k=1}^K t_k \log d_k(p; \theta)

    where dk(p;θ)=[softmax(Wlabelp)]kd_k(p; \theta) = [\text{softmax}(W^{\text{label}} p)]_k, and α∈[0,1]\alpha \in [0, 1] balances the unsupervised reconstruction and supervised cross-entropy objectives.

  3. Knowl 3 — Greedy Binary Tree Construction for Recursive Autoencoders

    algorithm

    To construct a hierarchical tree representation for an mm-word sentence without requiring a predefined syntactic parser, the unsupervised RAE greedily collapses adjacent constituents by minimizing the span-weighted reconstruction error at each step.

    Input: Sequence of continuous word vectors (x_1, x_2, ..., x_m) of length m with word-span counts (n_1 = 1, ..., n_m = 1)
    Output: Hierarchical binary tree T and root vector representation p_root
    Initialize active constituent list C = [(x_1, 1), (x_2, 1), ..., (x_m, 1)]
    Initialize tree node set T = empty
    while length of C > 1 do
        min_error = infinity
        best_index = -1
        best_parent = null
        best_span = 0
        
        for i = 1 to (length of C) - 1 do
            (c_1, span_1) = C[i]
            (c_2, span_2) = C[i + 1]
            
            p_candidate = f(W^(1) * [c_1; c_2] + b^(1))
            p_candidate = p_candidate / norm(p_candidate)
            [c'_1; c'_2] = W^(2) * p_candidate + b^(2)
            
            error = (span_1 / (span_1 + span_2)) * norm(c_1 - c'_1)^2 + (span_2 / (span_1 + span_2)) * norm(c_2 - c'_2)^2
            
            if error < min_error then
                min_error = error
                best_index = i
                best_parent = p_candidate
                best_span = span_1 + span_2
            end if
        end for
        
        Add triplet (best_parent -> C[best_index], C[best_index + 1]) to T
        Replace elements C[best_index] and C[best_index + 1] in C with (best_parent, best_span)
    end while
    return T, C[1].vector
  4. Knowl 4 — Training Semi-Supervised RAEs via Backpropagation Through Structure and L-BFGS

    model/method

    Optimization of the parameter set θ=(W(1),b(1),W(2),b(2),Wlabel,L)\theta = (W^{(1)}, b^{(1)}, W^{(2)}, b^{(2)}, W^{\text{label}}, L) is performed over batch training data.

    For each training sentence, the tree structure T(RAEθ(x))\mathcal{T}(\text{RAE}_\theta(x)) is first constructed using the greedy autoencoder algorithm. Derivatives of the joint objective JJ with respect to the network weights and word vectors LL are then computed using Backpropagation Through Structure (BPTS).

    Because greedy tree construction depends on the encoder weights W(1)W^{(1)} and word representations LL, the objective function is not strictly continuous with respect to θ\theta. However, optimizing the batch objective with the L-BFGS quasi-Newton method converges smoothly in practice and reliably discovers high-quality representations.

  5. Knowl 5 — Experience Project (EP) Sentiment Distribution Dataset

    experimental setup

    The Experience Project (EP) dataset consists of personal user confessions annotated with five user reaction categories:

    1. Sorry, Hugs (condolence)
    2. You Rock (approval / congratulations)
    3. Teehee (amusement)
    4. I Understand (empathy)
    5. Wow, Just Wow (surprise / shock)

    The full dataset contains 31,676 confession entries and 74,859 total votes, with an average of 2.4 votes per confession (variance of 33) and an average length of 129 words per entry. The distribution over total votes across the five categories is approximately [0.22,0.20,0.11,0.37,0.10][0.22, 0.20, 0.11, 0.37, 0.10].

    To ensure reliable reaction profiles, experiments use the subset EP≥4\text{EP}_{\ge 4} comprising 6,129 entries that received at least 4 votes. The dataset is split into training (49%), development (21%), and test (30%) sets. For multi-sentence entries, predicted label distributions across sentences are averaged.

  6. Knowl 6 — Dominant Emotion and Sentiment Distribution Prediction on Experience Project Confessions

    empirical result

    The semi-supervised RAE was evaluated on the Experience Project subset (extEP≥4 ext{EP}_{\ge 4}) on two tasks: predicting the single label with the most votes (accuracy) and predicting the full 5-class multinomial emotion distribution (evaluated by average Kullback-Leibler divergence KL(g∥p)=∑igilog⁡(gi/pi)\text{KL}(g \parallel p) = \sum_i g_i \log(g_i / p_i) between gold distribution gg and predicted distribution pp, where lower is better).

    Method Dominant Label Accuracy (%) Average KL Divergence
    Random 20.0 1.20
    Most Frequent (I Understand) 38.1 –
    Average Training Distribution – 0.83
    Baseline 1: Binary BoW (Logistic Regression) 46.4 0.81
    Baseline 2: Features (SVM + tf-idf + LIWC/Inquirer + WordNet) 47.0 0.72
    Baseline 3: Word Vectors (Word-level Softmax + SVM) 45.5 0.73
    RAE (our method) 50.1 0.70

    The RAE outperforms hand-engineered sentiment lexica and bag-of-words baselines by over 3% in accuracy and achieves the lowest KL divergence, demonstrating the effectiveness of compositionality for complex emotional reaction distributions.

  7. Knowl 7 — Binary Sentiment Polarity Classification on Movie Reviews and MPQA

    empirical result

    The semi-supervised RAE was evaluated using 10-fold cross-validation on two binary sentiment classification datasets: Movie Reviews (MR, 10,662 instances) and the MPQA opinion dataset (10,624 instances).

    Method MR Accuracy (%) MPQA Accuracy (%)
    Voting with two lexica 63.1 81.7
    Rule-based reversal on trees 62.9 82.8
    Bag of features with reversal 76.4 84.1
    Tree-CRF (Nakagawa et al., 2010) 77.3 86.1
    RAE (random word init.) 76.8 85.7
    RAE (pre-trained word init.) 77.7 86.4

    On the MR dataset, RAE achieves 77.7% accuracy without any pre-defined sentiment lexica or parsers. On MPQA, adding the sentiment lexicon used by Nakagawa et al. (2010) to cover words unseen in training improved RAE accuracy from 86.0% to 86.4%. Pre-trained word embeddings provide an improvement of less than 1% over random initialization.

  8. Knowl 8 — Effect of Loss Weighting Parameter on Sentiment Classification Accuracy

    empirical result

    In the joint objective E=αErec+(1−α)EcEE = \alpha E_{\text{rec}} + (1 - \alpha) E_{cE}, the hyperparameter α\alpha governs the trade-off between unsupervised reconstruction loss and supervised cross-entropy loss.

    Evaluating across values of α∈[0,1]\alpha \in [0, 1] on the development split of the Movie Reviews (MR) dataset showed that while placing the majority of weight on the supervised classification objective is crucial (purely unsupervised α=1.0\alpha = 1.0 yields low sentiment accuracy ∼84.2%\sim 84.2\%, and purely supervised α=0.0\alpha = 0.0 yields ∼86.0%\sim 86.0\%), setting α=0.2\alpha = 0.2 achieves the peak performance of approximately 87.2%87.2\%. Incorporating reconstruction error acts as a regularizer that prevents overfitting to sentence labels.

Coverage note — None was omitted; all key contributions (word vector lookups, RAE tree architecture, objective functions, greedy parsing algorithm, training regime, dataset introductions, and empirical evaluations across EP, MR, and MPQA benchmarks) have been fully captured.

References

  1. 1.P. Beineke, T. Hastie, C. D. Manning, and S. Vaithyanathan. 2004. Exploring sentiment summarization. In Proceedings of the AAAI Spring Symposium on Exploring Attitude and Affect in Text: Theories and Applications.
  2. 2.Y. Bengio, R. Ducharme, P. Vincent, and C. Janvin. 2003. A neural probabilistic language model. Journal of Machine Learning Research, 3:1137–1155.
  3. 3.D. M. Blei, A. Y. Ng, and M. I. Jordan. 2003. Latent dirichlet allocation. Journal of Machine Learning Research., 3:993–1022.
  4. 4.Y. Choi and C. Cardie. 2008. Learning with compositional semantics as structural inference for subsentential sentiment analysis. In EMNLP.
  5. 5.R. Collobert and J. Weston. 2008. A unified architecture for natural language processing: deep neural networks with multitask learning. In Proceedings of ICML, pages 160–167.
  6. 6.S. Das and M. Chen. 2001. Yahoo! for Amazon: Extracting market sentiment from stock message boards. In Proceedings of the Asia Pacific Finance Association Annual Conference (APFA).
  7. 7.K. Dave, S. Lawrence, and D. M. Pennock. 2003. Mining the peanut gallery: Opinion extraction and semantic classification of product reviews. In Proceedings of WWW, pages 519–528.
  8. 8.X. Ding, B. Liu, and P. S. Yu. 2008. A holistic lexicon-based approach to opinion mining. In Proceedings of the Conference on Web Search and Web Data Mining (WSDM).
  9. 9.J. L. Elman. 1991. Distributed representations, simple recurrent networks, and grammatical structure. Machine Learning, 7(2-3):195–225.
  10. 10.A. Esuli and F. Sebastiani. 2007. Pageranking wordnet synsets: An application to opinion mining. In Proceedings of the Association for Computational Linguistics (ACL).
  11. 11.C. Goller and A. Küchler. 1996. Learning task-dependent distributed representations by backpropagation through structure. In Proceedings of the International Conference on Neural Networks (ICNN-96).
  12. 12.G. Grefenstette, Y. Qu, J. G. Shanahan, and D. A. Evans. 2004. Coupling niche browsers and affect analysis for an opinion mining application. In Proceedings of Recherche d’Information Assistee par Ordinateur (RIAO).
  13. 13.D. Ikeda, H. Takamura, L. Ratinov, and M. Okumura. 2008. Learning to shift the polarity of words for sentiment classification. In IJCNLP.
  14. 14.S. Kim and E. Hovy. 2007. Crystal: Analyzing predictive opinions on the web. In EMNLP-CoNLL.
  15. 15.A. L. Maas, R. E. Daly, P. T. Pham, D. Huang, A. Y. Ng, and C. Potts. 2011. Learning accurate, compact, and interpretable tree annotation. In Proceedings of ACL.
  16. 16.Y. Mao and G. Lebanon. 2007. Isotonic Conditional Random Fields and Local Sentiment Flow. In NIPS.
  17. 17.P. Mirowski, M. Ranzato, and Y. LeCun. 2010. Dynamic auto-encoders for semantic indexing. In Proceedings of the NIPS 2010 Workshop on Deep Learning.
  18. 18.T. Nakagawa, K. Inui, and S. Kurohashi. 2010. Dependency tree-based sentiment classification using CRFs with hidden variables. In NAACL, HLT.
  19. 19.B. Pang and L. Lee. 2004. A sentimental education: Sentiment analysis using subjectivity summarization based on minimum cuts. In ACL.
  20. 20.B. Pang and L. Lee. 2005. Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales. In ACL, pages 115–124.
  21. 21.B. Pang and L. Lee. 2008. Opinion mining and sentiment analysis. Foundations and Trends in Information Retrieval, 2(1-2):1–135.
  22. 22.B. Pang, L. Lee, and S. Vaithyanathan. 2002. Thumbs up? Sentiment classification using machine learning techniques. In EMNLP.
  23. 23.J. W. Pennebaker, R.J. Booth, and M. E. Francis. 2007. Linguistic inquiry and word count: Liwc2007 operators manual. University of Texas.
  24. 24.L. Polanyi and A. Zaenen. 2006. Contextual valence shifters.
  25. 25.J. B. Pollack. 1990. Recursive distributed representations. Artificial Intelligence, 46:77–105, November.
  26. 26.C. Potts. 2010. On the negativity of negation. In David Lutz and Nan Li, editors, Proceedings of Semantics and Linguistic Theory 20. CLC Publications, Ithaca, NY.
  27. 27.B. Snyder and R. Barzilay. 2007. Multiple aspect ranking using the Good Grief algorithm. In HLT-NAACL.
  28. 28.R. Socher, C. D. Manning, and A. Y. Ng. 2010. Learning continuous phrase representations and syntactic parsing with recursive neural networks. In Proceedings of the NIPS-2010 Deep Learning and Unsupervised Feature Learning Workshop.
  29. 29.R. Socher, C. C. Lin, A. Y. Ng, and C. D. Manning. 2011. Parsing Natural Scenes and Natural Language with Recursive Neural Networks. In ICML.
  30. 30.P. J. Stone. 1966. The General Inquirer: A Computer Approach to Content Analysis. The MIT Press.
  31. 31.J. Turian, L. Ratinov, and Y. Bengio. 2010. Word representations: a simple and general method for semi-supervised learning. In Proceedings of ACL, pages 384–394.
  32. 32.P. Turney. 2002. Thumbs up or thumbs down? Semantic orientation applied to unsupervised classification of reviews. In ACL.
  33. 33.L. Velikovich, S. Blair-Goldensohn, K. Hannan, and R. McDonald. 2010. The viability of web-derived polarity lexicons. In NAACL, HLT.
  34. 34.T. Voegtlin and P. Dominey. 2005. Linear Recursive Distributed Representations. Neural Networks, 18(7).
  35. 35.J. Wiebe, T. Wilson, and C. Cardie. 2005. Annotating expressions of opinions and emotions in language. Language Resources and Evaluation, 39.
  36. 36.T. Wilson, J. Wiebe, and P. Hoffmann. 2005. Recognizing contextual polarity in phrase-level sentiment analysis. In HLT/EMNLP.
  37. 37.H. Yu and V. Hatzivassiloglou. 2003. Towards answering opinion questions: Separating facts from opinions and identifying the polarity of opinion sentences. In EMNLP.

Citation

MLA
Socher, R., et al. “Semi-Supervised Recursive Autoencoders for Predicting Sentiment Distributions”. Proceedings of the 2011 Conference on Empirical Methods in Natural Language Processing, 2011, pp. 151–61, https://aclanthology.org/D11-1014/.
APA
Socher, R., Pennington, J., Huang, E. H., Ng, A. Y., & Manning, C. D. (2011). Semi-Supervised Recursive Autoencoders for Predicting Sentiment Distributions. Proceedings of the 2011 Conference on Empirical Methods in Natural Language Processing, 151–161. https://aclanthology.org/D11-1014/
Chicago
Socher, R., J. Pennington, E. H. Huang, A. Y. Ng, and C. D. Manning. 2011. “Semi-Supervised Recursive Autoencoders for Predicting Sentiment Distributions”. Proceedings of the 2011 Conference on Empirical Methods in Natural Language Processing, 151–61. https://aclanthology.org/D11-1014/.
Harvard
Socher, R. et al. (2011) “Semi-Supervised Recursive Autoencoders for Predicting Sentiment Distributions”, Proceedings of the 2011 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 151–161. Available at: https://aclanthology.org/D11-1014/.
Vancouver
1. Socher R, Pennington J, Huang EH, Ng AY, Manning CD (2011) Semi-Supervised Recursive Autoencoders for Predicting Sentiment Distributions. In: Proceedings of the 2011 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 151–161

BibTeX

@inproceedings{socher-etal-2011-semi,
    title = "Semi-Supervised Recursive Autoencoders for Predicting Sentiment Distributions",
    author = "Socher, Richard  and
      Pennington, Jeffrey  and
      Huang, Eric H.  and
      Ng, Andrew Y.  and
      Manning, Christopher D.",
    editor = "Barzilay, Regina  and
      Johnson, Mark",
    booktitle = "Proceedings of the 2011 Conference on Empirical Methods in Natural Language Processing",
    month = jul,
    year = "2011",
    address = "Edinburgh, Scotland, UK.",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/D11-1014/",
    pages = "151--161"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by-nc-sa/4.0/