Domain Adaptation for Large-Scale Sentiment Classification: A Deep Learning Approach

Xavier GlorotAntoine BordesYoshua Bengio

article2011ICML1,865 citations

Proposes an unsupervised deep learning framework based on stacked denoising auto-encoders that extracts domain-invariant representations to achieve state-of-the-art cross-domain sentiment classification on large-scale Amazon product reviews.

Listen

Modern businesses increasingly rely on automated sentiment classification to monitor customer feedback and identify market opportunities across a rapidly expanding range of products and services. However, manually labeling training data for every new category is prohibitively expensive and time-consuming, while models trained on one product category often perform poorly on another due to differing vocabularies. The article evaluates whether an unsupervised deep learning approach can discover shared, high-level feature representations across multiple domains to enable effective sentiment classification transfer with minimal human supervision.

The authors implemented a two-step framework using Stacked Denoising Auto-encoders with sparse rectifier units. The system first learns non-linear data representations from raw review text across all available categories without using labels, reconstructs corrupted inputs to capture meaningful underlying concepts, and then trains a standard linear classifier on labeled examples from only a single source category. This approach was benchmarked against existing techniques on a balanced four-domain Amazon review dataset and evaluated on an industrial-scale, unbalanced dataset covering 22 distinct product domains with over 340,000 reviews.

The evaluation produced four key findings: First, on the four-domain benchmark, the proposed method outperformed competing state-of-the-art methods in 11 out of 12 cross-domain transfer tasks, reducing the average transfer error rate from 24.3% (raw baseline) and 21.3% (prior best method) down to 16.7%. Second, pre-training the unsupervised model across all available domains simultaneously yielded better transfer performance than training only on paired source-target domains. Third, on the 22-domain dataset, the system achieved the lowest average transfer error rate (10.9% with three stacked layers compared to 14.5% for the baseline and 13.9% for a standard neural network), demonstrating that deeper architectures extract progressively better abstractions for large-scale transfer. Fourth, diagnostic analysis showed that the learned representations effectively disentangled domain-specific vocabulary from universal sentiment expressions.

These findings indicate that organizations can significantly reduce data annotation costs and operational deployment timelines by training a single unsupervised feature extraction pipeline across all review data, followed by lightweight sentiment classifiers trained on limited labeled subsets. The source evidence supports deploying multi-layer stacked auto-encoders trained jointly across all target domains rather than maintaining siloed classifiers. While results show strong confidence across the tested e-commerce datasets, practitioners should note that the evaluation was confined to bag-of-words input encodings and consumer product reviews, suggesting that additional pilot validation is advisable before extending the system to structurally different text formats.

Glorot et al (2011).pdf
Cover for Domain Adaptation for Large-Scale Sentiment Classification: A Deep Learning Approach

Abstract

The exponential increase in the availability of online reviews and recommendations makes sentiment classification an interesting topic in academic and industrial research. Reviews can span so many different domains that it is difficult to gather annotated training data for all of them. Hence, this paper studies the problem of domain adaptation for sentiment classifiers, hereby a system is trained on labeled reviews from one source domain but is meant to be deployed on another. We propose a deep learning approach which learns to extract a meaningful representation for each review in an unsupervised fashion. Sentiment classifiers trained with this high-level feature representation clearly outperform state-of-the-art methods on a benchmark composed of reviews of 4 types of Amazon products. Furthermore, this method scales well and allowed us to successfully perform domain adaptation on a larger industrial-strength dataset of 22 domains.

Table of Contents

  • 1. Introduction
  • 2. Domain Adaptation
  • 2.1. Related Work
  • 2.2. Applications to Sentiment Classification
  • 3. Deep Learning Approach
  • 3.1. Background
  • 3.2. Stacked Denoising Auto-encoders
  • 3.3. Proposed Protocol
  • 3.4. Discussion
  • 4. Empirical Evaluation
  • 4.1. Experimental Setup
  • 4.2. Metrics
  • 4.3. Benchmark Experiments
  • 4.4. Large-Scale Experiments
  • 5. Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Stacked Denoising Autoencoder Architecture for Cross-Domain Sentiment Classification

    model/method

    A two-step deep learning protocol is used for cross-domain sentiment classification when unlabeled text data is available from multiple domains and labeled data is available only for a source domain.

    In the first step, high-level representations are learned in an unsupervised manner using a Stacked Denoising Autoencoder (SDA) trained greedily layer-wise on bag-of-words representations (binary indicators of the 5,000 most frequent unigrams and bigrams) from all available domains:

    1. First layer: Given input vector x∈{0,1}dx \in \{0, 1\}^d, stochastically corrupt xx to x~\tilde{x} using masking noise where each active input (xi=1x_i = 1) is independently set to 00 with probability PP (typically P=0.8P = 0.8). The encoder computes code activations using the rectifier non-linearity h(x)=max⁡(0,Wx+b)h(x) = \max(0, W x + b). The decoder uses the logistic sigmoid activation function σ(z)=(1+exp⁡(−z))−1\sigma(z) = (1 + \exp(-z))^{-1} to produce reconstruction r(x~)=σ(W′h(x~)+c)r(\tilde{x}) = \sigma(W' h(\tilde{x}) + c). The layer is trained by stochastic gradient descent minimizing the Kullback-Leibler divergence between xx and r(x~)r(\tilde{x}).
    2. Upper layers: For subsequent autoencoder layers, the code layer continues to use rectifier non-linearities h(x)=max⁡(0,Wx+b)h(x) = \max(0, W x + b). The corruption process is additive Gaussian noise applied prior to the rectifier non-linearity of the input to maintain representation sparsity. The decoder uses the softplus activation function g(z)=log⁡(1+exp⁡(z))g(z) = \log(1 + \exp(z)). The reconstruction criterion is the squared error ∥x−r(x~)∥2\|x - r(\tilde{x})\|^2.

    In the second step, the rectifier activations of the trained SDA code layers serve as transformed feature representations. A linear Support Vector Machine (SVM) with squared hinge loss is trained on the transformed labeled examples of the source domain and evaluated directly on the transformed instances of the target domain.

  2. Knowl 2 — Representation Disentanglement vs Proxy A-Distance in Domain Transfer

    empirical result

    In domain adaptation theory (Ben-David et al., 2007), cross-domain generalization error is bounded in part by the A\mathcal{A}-distance between source distribution pSp_S and target distribution pTp_T. The proxy A\mathcal{A}-distance is estimated as: d^A=2(1−2ϵ)\hat{d}_A = 2(1 - 2\epsilon) where ϵ\epsilon is the generalization error of a linear SVM trained to discriminate between instances from domain SS and domain TT. Lower d^A\hat{d}_A corresponds to less distinguishable domain distributions.

    When input representations are transformed via a Stacked Denoising Autoencoder (SDA) trained on multi-domain sentiment data:

    1. The proxy A\mathcal{A}-distance d^A\hat{d}_A between all domain pairs increases compared to raw bag-of-words features (making domains easier to linearly distinguish).
    2. Despite the increase in d^A\hat{d}_A, cross-domain sentiment transfer error decreases substantially.

    This behavior occurs because unsupervised SDA feature extraction disentangles domain-specific factors from sentiment polarity factors rather than projecting them into domain-invariant mixtures. When L1L_1-regularized linear SVMs are trained across 6 domain discrimination tasks and 5 sentiment classification tasks, raw features are heavily reused across both domain and sentiment tasks, whereas SDA-transformed features show near-orthogonal partitioning (individual features are selected either for domain discrimination or for sentiment classification, but rarely both).

  3. Knowl 3 — Multi-Domain Transfer Evaluation Metrics: Transfer Loss, Transfer Ratio, and In-Domain Ratio

    definition

    Let e(S,T)e(S, T) denote the test classification error of a model trained on labeled data from source domain SS and evaluated on target domain TT, where e(T,T)e(T, T) is the in-domain test error. Let eb(T,T)e_b(T, T) denote the baseline in-domain test error obtained by a linear SVM trained and tested on raw features of domain TT.

    Three metrics evaluate domain adaptation across multiple domains:

    • Transfer Loss (tt): The difference between the transfer error and the baseline in-domain error: t(S,T)=e(S,T)−eb(T,T)t(S, T) = e(S, T) - e_b(T, T) A negative transfer loss (t(S,T)<0t(S, T) < 0) indicates that a classifier trained on a different domain using transformed features outperforms a baseline classifier trained directly on target in-domain data.

    • Transfer Ratio (QQ): The average relative error ratio across all nn ordered source-target pairs (S,T)(S, T) with S≠TS \neq T: Q=1n∑(S,T),S≠Te(S,T)eb(T,T)Q = \frac{1}{n} \sum_{(S,T), S \neq T} \frac{e(S, T)}{e_b(T, T)} Replacing differences with ratios reduces sensitivity to variations in intrinsic task difficulty across heterogeneous domains.

    • In-Domain Ratio (II): The average relative in-domain performance of the transformed representation across all mm domains: I=1m∑Te(T,T)eb(T,T)I = \frac{1}{m} \sum_{T} \frac{e(T, T)}{e_b(T, T)}

  4. Knowl 4 — Cross-Domain Sentiment Classification Benchmark Results on 4 Amazon Domains

    empirical result

    On the 4-domain Amazon sentiment benchmark (comprising Books, DVDs, Electronics, and Kitchen appliances, with 12 distinct source-target transfer pairs), a linear SVM trained on representations extracted by a 1-layer, 5,000-unit Stacked Denoising Autoencoder with multi-domain shared pre-training (SDAsh\text{SDA}_{sh}) achieves state-of-the-art transfer performance:

    • Average Transfer Error: Baseline (linear SVM on raw features) achieves 24.3%24.3\%, Spectral Feature Alignment (SFA) achieves 21.3%21.3\%, and SDAsh\text{SDA}_{sh} achieves 16.7%16.7\%.
    • Pairwise Comparison: SDAsh\text{SDA}_{sh} achieves the lowest transfer loss on 11 out of 12 source-target pairs compared against Structural Correspondence Learning (SCL), SFA, and Multi-label Consensus Training (MCT), with SCL performing slightly better only on the Kitchen →\to Electronics pair.
    • Negative Transfer Loss: For every target domain, SDAsh\text{SDA}_{sh} produces at least one source domain transfer that achieves a negative transfer loss, indicating that training on transferred source features surpasses the in-domain baseline on raw target features.
    • Shared vs Pairwise Unsupervised Training: SDAsh\text{SDA}_{sh} (unsupervised pre-training using unlabeled data across all 4 domains simultaneously) achieves a lower transfer ratio QQ than SDA trained strictly on the source-target domain pair, demonstrating that incorporating unrelated domains during unsupervised feature learning improves transfer.
  5. Knowl 5 — Large-Scale Domain Adaptation Across 22 Amazon Product Domains

    empirical result

    On an industrial-scale dataset comprising 22 Amazon product domains and over 340,000 reviews with diverse domain sizes and class imbalances, deep autoencoder representations outperform supervised and shallow baselines:

    • Averaged Transfer Generalization Errors across all transfer pairs:

      • Baseline (linear SVM on raw bag-of-words): 14.5%14.5\%
      • Multi-Layer Perceptron (MLP; 1 hidden layer with 5,000 tanh⁡\tanh units + softmax): 13.9%13.9\%
      • 1-layer SDA with shared pre-training (SDAsh1\text{SDA}_{sh}1; 5,000 rectifier units): 11.5%11.5\%
      • 3-layer SDA with shared pre-training (SDAsh3\text{SDA}_{sh}3; 5,000 rectifier units per layer): 10.9%10.9\%
    • Impact of Depth: Stacking 3 layers of denoising autoencoders yields superior transfer compared to a single layer (10.9%10.9\% vs 11.5%11.5\%). Furthermore, the relative improvement of SDAsh3\text{SDA}_{sh}3 over the baseline is larger for transfer ratio QQ than for in-domain ratio II, indicating that deeper hierarchical feature discovery specifically aids cross-domain generalization.

  6. Knowl 6 — Amazon Multi-Domain Sentiment Classification Dataset Statistics

    data/table

    The Amazon sentiment dataset contains review text represented as bag-of-words binary vectors for unigrams and bigrams. It exists in two versions: a large-scale, heterogeneous, unbalanced 22-domain dataset and a balanced 4-domain benchmark dataset.

    Domain Train size Test size Unlab. size % Neg. ex
    Complete (large-scale) data set
    Toys 6318 2527 3791 19.63%
    Software 1032 413 620 37.77%
    Apparel 4470 1788 2682 14.49%
    Video 8694 3478 5217 13.63%
    Automotive 362 145 218 20.69%
    Books 10625 10857 32845 12.08%
    Jewelry 982 393 589 15.01%
    Grocery 1238 495 743 13.54%
    Camera 2652 1061 1591 16.31%
    Baby 2046 818 1227 21.39%
    Magazines 1195 478 717 22.59%
    Cell 464 186 279 37.10%
    Electronics 10196 4079 6118 21.94%
    DVDs 10625 9218 26245 14.16%
    Outdoor 729 292 437 20.55%
    Health 3254 1301 1952 21.21%
    Music 10625 24872 88865 8.33%
    Videogame 720 288 432 17.01%
    Kitchen 9233 3693 5540 20.96%
    Beauty 1314 526 788 15.78%
    Sports 2679 1072 1607 18.75%
    Food 691 277 415 13.36%
    (Smaller-scale) benchmark
    Books 1600 400 4465 50%
    Kitchen 1600 400 5945 50%
    Electronics 1600 400 5681 50%
    DVDs 1600 400 3586 50%

    The 22-domain dataset features large variations in domain size (from 362 training reviews in Automotive to 10,625 in Books/DVDs/Music) and negative class proportions (from 8.33%8.33\% in Music to 37.77%37.77\% in Software). The 4-domain benchmark is balanced with 1,000 positive and 1,000 negative examples per domain (split into 1,600 train and 400 test).

Coverage note — None was omitted; all contributed models, empirical findings on both the 4-domain benchmark and 22-domain large-scale dataset, metrics, and theoretical analyses of representation disentanglement are covered.

References

  1. 1.Ben-David, S., Blitzer, J., Crammer, K., and Sokolova, P. M. (2007). Analysis of representations for domain adaptation. In Advances in Neural Information Processing Systems 20 (NIPS’07).
  2. 2.Bengio, Y. (2009). Learning deep architectures for AI. Foundations and Trends in Machine Learning, 2(1), 1–127. Also published as a book. Now Publishers, 2009.
  3. 3.Bengio, Y., Lamblin, P., Popovici, D., and Larochelle, H. (2006). Greedy layer-wise training of deep networks. In Adv. in Neural Inf. Proc. Syst. 19 , pages 153–160.
  4. 4.Bergstra, J., Breuleux, O., Bastien, F., Lamblin, P., Pascanu, R., Desjardins, G., Turian, J., Warde-Farley, D., and Bengio, Y. (2010). Theano: a CPU and GPU math expression compiler. In Proceedings of the Python for Scientific Computing Conference (SciPy). Oral.
  5. 5.Blitzer, J., McDonald, R., and Pereira, F. (2006). Domain adaptation with structural correspondence learning. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP’06).
  6. 6.Blitzer, J., Dredze, M., and Pereira, F. (2007). Biographies, bollywood, boom-boxes and blenders: Domain adaptation for sentiment classification. In Proceedings of Association for Computational Linguistics (ACL’07).
  7. 7.Collobert, R. and Weston, J. (2008). A unified architecture for natural language processing: Deep neural networks with multitask learning. In Proceedings of Internationnal Conference on Machine Learning 2008 , pages 160–167.
  8. 8.Dai, W., Xue, G.-R., Yang, Q., and Yu, Y. (2007). Transferring naive Bayes classifiers for text classification. In Proc. of Assoc. for the Adv. of Art. Int. (AAAI’07).
  9. 9.Daumé III, H. and Marcu, D. (2006). Domain adaptation for statistical classifiers. Journal of Artificial Intelligence Research, 26, 101–126.
  10. 10.Glorot, X., Bordes, A., and Bengio, Y. (2011). Deep sparse rectifier neural networks. In Proceeding of the Conference on Artificial Intelligence and Statistics.
  11. 11.Goodfellow, I., Le, Q., Saxe, A., and Ng, A. (2009). Measuring invariances in deep networks. In Advances in Neural Information Processing Systems 22 , pages 646–654.
  12. 12.Hinton, G. E. and Salakhutdinov, R. (2006). Reducing the dimensionality of data with neural networks. Science, 313(5786), 504–507.
  13. 13.Hinton, G. E., Osindero, S., and Teh, Y. (2006). A fast learning algorithm for deep belief nets. Neural Computation, 18, 1527–1554.
  14. 14.Jiang, J. and Zhai, C. (2007). Instance weighting for domain adaptation in nlp. In Proceedings of Association for Computational Linguistics (ACL’07).
  15. 15.Li, S. and Zong, C. (2008). Multi-domain adaptation for sentiment classification: Using multiple classifier combining methods. In Proc. of the Conference on Natural Language Processing and Knowledge Engineering.
  16. 16.Nair, V. and Hinton, G. E. (2010). Rectified linear units improve restricted boltzmann machines. In Proceedings of the International Conference on Machine Learning.
  17. 17.Pan, S. J., Ni, X., Sun, J.-T., Yang, Q., and Chen, Z. (2010). Cross-domain sentiment classification via spectral feature alignment. In Proceedings of the International World Wide Web Conference (WWW’10).
  18. 18.Pang, B. and Lee, L. (2008). Opinion mining and sentiment analysis. Foundations and Trends in Information Retrieval, 2(1-2), 1–135.
  19. 19.Pang, B., Lee, L., and Vaithyanathan, S. (2002). Thumbs up? Sentiment classification using machine learning techniques. In Proceedings of the Conference on Empirical Methods in Natural Language Processing.
  20. 20.Raina, R., Battle, A., Lee, H., Packer, B., and Ng, A. Y. (2007). Self-taught learning: transfer learning from unlabeled data. In Proceedings of the International Conference on Machine Learning, pages 759–766.
  21. 21.Ranzato, M. and Szummer, M. (2008). Semi-supervised learning of compact document representations with deep networks. In Proceedings of the International Conference on Machine Learning (ICML’08), pages 792–799.
  22. 22.Salakhutdinov, R. and Hinton, G. E. (2007). Semantic hashing. In Proceedings of the 2007 Workshop on Information Retrieval and applications of Graphical Models (SIGIR 2007), Amsterdam. Elsevier.
  23. 23.Sindhwani, V. and Keerthi, S. S. (2006). Large scale semi-supervised linear svms. In Proc. of the 2006 Workshop on Inf. Retrieval and Applications of Graphical Models .
  24. 24.Snyder, B. and Barzilay, R. (2007). Multiple aspect ranking using the Good Grief algorithm. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics.
  25. 25.Thomas, M., Pang, B., and Lee, L. (2006). Get out the vote: Determining support or opposition from Congressional floor-debate transcripts. In Empirical Methods in Natural Language Processing (EMNLP’06).
  26. 26.Vincent, P. (2011). A connection between score matching and denoising autoencoders. Neural Computation, to appear.
  27. 27.Vincent, P., Larochelle, H., Bengio, Y., and Manzagol, P.-A. (2008). Extracting and composing robust features with denoising autoencoders. In International Conference on Machine Learning, pages 1096–1103.

Citation

MLA
Glorot, X., et al. “Domain Adaptation for Large-Scale Sentiment Classification: A Deep Learning Approach”. HAL (Le Centre Pour La Communication Scientifique Directe), 2011, http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.231.3442.
APA
Glorot, X., Bordes, A., & Bengio, Y. (2011). Domain Adaptation for Large-Scale Sentiment Classification: A Deep Learning Approach. HAL (Le Centre Pour La Communication Scientifique Directe). http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.231.3442
Chicago
Glorot, X., A. Bordes, and Y. Bengio. 2011. “Domain Adaptation for Large-Scale Sentiment Classification: A Deep Learning Approach”. HAL (Le Centre Pour La Communication Scientifique Directe). http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.231.3442.
Harvard
Glorot, X., Bordes, A. and Bengio, Y. (2011) “Domain Adaptation for Large-Scale Sentiment Classification: A Deep Learning Approach”, HAL (Le Centre pour la Communication Scientifique Directe) [Preprint]. Available at: http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.231.3442.
Vancouver
1. Glorot X, Bordes A, Bengio Y (2011) Domain Adaptation for Large-Scale Sentiment Classification: A Deep Learning Approach. HAL (Le Centre pour la Communication Scientifique Directe)

BibTeX

@article{glorot2011domain,
  title = {Domain Adaptation for Large-Scale Sentiment Classification: A Deep Learning Approach},
  author = {Glorot, Xavier and Bordes, Antoine and Bengio, Yoshua},
  year = {2011},
  journal = {HAL (Le Centre pour la Communication Scientifique Directe)},
  url = {http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.231.3442}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission