Parsing Natural Scenes and Natural Language with Recursive Neural Networks

Richard SocherCliff Chiung-Yu LinAndrew Y. NgChristopher D. Manning

article2011ICML1,490 citations

Presents a unified recursive neural network architecture that jointly predicts hierarchical parse trees and semantic representations across both visual scenes and natural language, setting a new benchmark for scene segmentation and classification while maintaining competitive syntactic parsing accuracy.

Listen

Real-world visual and linguistic data naturally exhibit nested, hierarchical relationships, such as smaller image regions forming parts of larger objects, or words forming phrases within sentences. Traditional computational approaches often rely on separate, highly specialized models for computer vision and natural language processing, frequently treating images as flat collections of regions rather than structured wholes. The article develops and evaluates a unified deep learning architecture based on recursive neural networks that can automatically discover and predict these recursive structures across both visual scenes and textual data.

The approach uses a single recursive neural network framework paired with a maximum-margin structured prediction objective. For visual tasks, images are segmented into small regions with standard visual features, which are mapped into a shared semantic vector space. The network iteratively scores and merges adjacent regions into larger super-segments until a full hierarchical tree representing the entire image is formed, simultaneously predicting category labels at each node. For text parsing, the exact same core architecture embeds words into a continuous vector space and greedily merges them into syntactic phrases. The models were evaluated on standard benchmark datasets, including the Stanford background dataset for image segmentation and scene classification, and the Wall Street Journal section of the Penn Treebank for natural language parsing.

The key findings demonstrate that this unified architecture achieves top-tier performance across multiple tasks. First, on pixel-level semantic image segmentation, the model achieved 78.1% accuracy on the Stanford background dataset, establishing a new state of the art and outperforming standard Markov random fields and conditional random fields. Second, using the learned hierarchical tree features for whole-scene classification yielded an accuracy of 88.1%, outperforming standard global baseline descriptors by roughly 4 percentage points. Third, when applied to natural language parsing for sentences up to 15 words, the model achieved an unlabeled bracketing accuracy of 90.29%, performing competitively within 1.3 percentage points of specialized parsers despite operating entirely on continuous learned representations without explicit grammar rules.

These results indicate that a single recursive deep learning approach can replace multiple domain-specific, manually engineered systems, substantially lowering the architectural complexity and development overhead of multi-modal AI systems. The learned continuous representations successfully capture contextual meaning and compositionality, enabling automated systems to reason about part-whole relationships in both visual scenes and language without hand-crafted symbolic rules.

Organizations developing computer vision or multi-modal analysis pipelines should consider adopting unified recursive architectures to jointly handle segmentation, annotation, and classification. However, decision-makers should note certain limitations: the natural language evaluations were restricted to shorter sentences of up to 15 words, and the visual parsing relied on a greedy merging search. Further pilot evaluations on longer, more complex sentences and larger-scale, highly diverse image datasets are recommended before full-scale deployment in production environments.

Socher et al (2011).pdf
Cover for Parsing Natural Scenes and Natural Language with Recursive Neural Networks

Abstract

Recursive structure is commonly found in the inputs of different modalities such as natural scene images or natural language sentences. Discovering this recursive structure helps us to not only identify the units that an image or sentence contains but also how they interact to form a whole. We introduce a max-margin structure prediction architecture based on recursive neural networks that can successfully recover such structure both in complex scene images as well as sentences. The same algorithm can be used both to provide a competitive syntactic parser for natural language sentences from the Penn Treebank and to outperform alternative approaches for semantic scene segmentation, annotation and classification. For segmentation and annotation our algorithm obtains a new level of state-of-the-art performance on the Stanford background dataset (78.1%). The features from the image parse tree outperform Gist descriptors for scene classification by 4%.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Mapping Segments and Words into Syntactico-Semantic Space
  • 3.1. Input Representation of Scene Images
  • 3.2. Input Representation for Natural Language Sentences
  • 4. Recursive Neural Networks for Structure Prediction
  • 4.1. Max-Margin Estimation
  • 4.2. Greedy Structure Predicting RNNs
  • 4.3. Category Classifiers in the Tree
  • 4.4. Improvements for Language Parsing
  • 5. Learning
  • 6. Experiments
  • 6.1. Scene Understanding: Segmentation and Annotation
  • 6.2. Scene Classification
  • 6.3. Nearest Neighbor Scene Subtrees
  • 6.4. Supervised Parsing
  • 6.5. Nearest Neighbor Phrases
  • 7. Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Recursive Neural Network Architecture for Hierarchical Structure Prediction

    model/method

    The Recursive Neural Network (RNN) recursively computes parent representations and merge scores for adjacent pairs of constituent nodes. Given two child feature vectors c1,c2∈Rnc_1, c_2 \in \mathbb{R}^n, the network produces a parent feature representation p∈Rnp \in \mathbb{R}^n and a scalar merge score s∈Rs \in \mathbb{R} via:

    p=f(W[c1;c2]+b)p = f(W [c_1; c_2] + b) s=Wscoreps = W^{\text{score}} p

    where W∈Rn×2nW \in \mathbb{R}^{n \times 2n} is the learned parameter matrix for child combination, b∈Rnb \in \mathbb{R}^n is the bias vector, [c1;c2]∈R2n[c_1; c_2] \in \mathbb{R}^{2n} is the column concatenation of the two child vectors, ff is an element-wise activation function (specifically the standard sigmoid function f(z)=1/(1+e−z)f(z) = 1 / (1 + e^{-z})), and Wscore∈R1×nW^{\text{score}} \in \mathbb{R}^{1 \times n} is a learned row vector that projects the parent vector to a scalar score.

    For a proposed binary tree structure y^\hat{y} over input xx, the total tree score s(RNN(θ,x,y^))s(\text{RNN}(\theta, x, \hat{y})) is defined as the sum of all local merge scores over the set of non-terminal nodes N(y^)N(\hat{y}):

    s(RNN(θ,x,y^))=∑d∈N(y^)sds(\text{RNN}(\theta, x, \hat{y})) = \sum_{d \in N(\hat{y})} s_d

    where sds_d is the scalar score evaluated at non-terminal node dd, and θ\theta denotes the full collection of model parameters.

  2. Knowl 2 — Mapping Image Segments and Text Tokens to Syntactico-Semantic Space

    model/method

    Before recursive composition, input elements from images or sentences are projected into a shared nn-dimensional continuous semantic vector space:

    1. Scene Image Segments: An image is oversegmented into superpixels (segments). For each segment i∈{1,…,Nsegs}i \in \{1, \dots, N_{\text{segs}}\}, an initial feature vector Fi∈R119F_i \in \mathbb{R}^{119} is extracted, containing color and texture features, boosted pixel classifier scores, appearance, and shape descriptors. The segment's initial representation ai∈Rna_i \in \mathbb{R}^n is computed as:

    ai=f(WsemFi+bsem)a_i = f(W^{\text{sem}} F_i + b^{\text{sem}})

    where Wsem∈Rn×119W^{\text{sem}} \in \mathbb{R}^{n \times 119} is a learned weight matrix, bsem∈Rnb^{\text{sem}} \in \mathbb{R}^n is a bias vector, and ff is the element-wise sigmoid function.

    1. Natural Language Words: For an ordered sequence of NwordsN_{\text{words}} words in a sentence, each word ii with vocabulary index k∈{1,…,∣V∣}k \in \{1, \dots, |V|\} is mapped to its semantic vector ai∈Rna_i \in \mathbb{R}^n via a projection layer using a learned word embedding matrix L∈Rn×∣V∣L \in \mathbb{R}^{n \times |V|}:

    ai=Leka_i = L e_k

    where ek∈{0,1}∣V∣e_k \in \{0, 1\}^{|V|} is a binary one-hot vector that is 1 at index kk and 0 elsewhere, and ∣V∣|V| is the vocabulary size.

  3. Knowl 3 — Structured Max-Margin Objective and Margin Loss for Tree Parsing

    model/method

    The parameters θ\theta of the Recursive Neural Network (RNN) are optimized using a structured max-margin objective. Given NN training examples with inputs xix_i and ground-truth element labels lil_i, the regularized objective function is:

    J(θ)=1N∑i=1Nri(θ)+λ2∥θ∥22J(\theta) = \frac{1}{N} \sum_{i=1}^N r_i(\theta) + \frac{\lambda}{2} \|\theta\|_2^2

    where λ\lambda is the L2L_2 regularization hyperparameter, and the structured risk ri(θ)r_i(\theta) for instance ii is:

    ri(θ)=max⁡y^∈T(xi)(s(RNN(θ,xi,y^))+Δ(xi,li,y^))−max⁡yi∈Y(xi,li)s(RNN(θ,xi,yi))r_i(\theta) = \max_{\hat{y} \in \mathcal{T}(x_i)} \left( s(\text{RNN}(\theta, x_i, \hat{y})) + \Delta(x_i, l_i, \hat{y}) \right) - \max_{y_i \in \mathcal{Y}(x_i, l_i)} s(\text{RNN}(\theta, x_i, y_i))

    Here, T(xi)\mathcal{T}(x_i) is the set of all possible binary trees for input xix_i, Y(xi,li)\mathcal{Y}(x_i, l_i) is the set of valid ground-truth trees, s(RNN(θ,xi,y))s(\text{RNN}(\theta, x_i, y)) is the total RNN score for tree yy, and Δ(xi,li,y^)\Delta(x_i, l_i, \hat{y}) is the structured loss.

    For visual scene parsing, a tree is considered correct (Y(xi,li)\mathcal{Y}(x_i, l_i)) if all adjacent superpixels with the same semantic label are merged together into homogeneous super-segments before any merge occurs between super-segments of different classes. The structured loss penalizes incorrect merges:

    Δ(x,l,y^)=κ∑d∈N(y^)1{subTree(d)∉Y(x,l)}\Delta(x, l, \hat{y}) = \kappa \sum_{d \in N(\hat{y})} \mathbf{1}\{\text{subTree}(d) \notin \mathcal{Y}(x, l)\}

    where N(y^)N(\hat{y}) is the set of non-terminal nodes in proposed tree y^\hat{y}, subTree(d)\text{subTree}(d) denotes the subtree rooted at node dd, 1{⋅}\mathbf{1}\{\cdot\} is an indicator function, and κ\kappa is a penalty scaling factor. For natural language sentences, the ground-truth set contains a single tree (Y(x)={y}\mathcal{Y}(x) = \{y\}), the second maximization in ri(θ)r_i(\theta) evaluates directly on yy, and the margin loss Δ\Delta equals the count of incorrect phrase spans.

  4. Knowl 4 — Greedy Bottom-Up Parsing Algorithm for Natural Scene Images

    algorithm

    For image parsing, an input image consists of segment representations {a1,…,aNsegs}\{a_1, \dots, a_{N_{\text{segs}}}\} in Rn\mathbb{R}^n and a symmetric adjacency matrix A∈{0,1}Nsegs×NsegsA \in \{0, 1\}^{N_{\text{segs}} \times N_{\text{segs}}} where A(i,j)=1A(i, j) = 1 if segments ii and jj share a spatial boundary. Because the space of binary trees over arbitrary planar graphs cannot be searched with dynamic programming, parsing is performed via greedy agglomeration:

    Input: Initial segment activations a1,…,aNsegs{a_1, \dots, a_{N_{\text{segs}}}}, symmetric adjacency matrix AA, parameters (W,b,Wscore)(W, b, W^{\text{score}})
    Output: A binary parse tree over all segments with parent activations and scores
    Initialize candidate set C←{[ai,aj]:A(i,j)=1}C \leftarrow \{ [a_i, a_j] : A(i, j) = 1 \}
    for each candidate pair [ai,aj]∈C[a_i, a_j] \in C do
        Compute potential parent p(i,j)←f(W[ai;aj]+b)p_{(i, j)} \leftarrow f(W [a_i; a_j] + b)
        Compute score s(i,j)←Wscorep(i,j)s_{(i, j)} \leftarrow W^{\text{score}} p_{(i, j)}
    end for
    while size of active segment set >1> 1 do
        Select (i∗,j∗)←arg⁡max⁡(i,j):[ai,aj]∈Cs(i,j)(i^*, j^*) \leftarrow \arg\max_{(i, j) : [a_i, a_j] \in C} s_{(i, j)}
        Create parent node p(i∗,j∗)←f(W[ai∗;aj∗]+b)p_{(i^*, j^*)} \leftarrow f(W [a_{i^*}; a_{j^*}] + b)
        Remove all pairs containing ai∗a_{i^*} or aj∗a_{j^*} from CC:
            C←C∖{[au,av]∈C:u∈{i∗,j∗} or v∈{i∗,j∗}}C \leftarrow C \setminus \{ [a_u, a_v] \in C : u \in \{i^*, j^*\} \text{ or } v \in \{i^*, j^*\} \}
        Define merged super-segment k∗=(i∗,j∗)k^* = (i^*, j^*)
        Update adjacency matrix AA to include k∗k^*, where A(k∗,m)=1A(k^*, m) = 1 for all active segments mm where A(i∗,m)=1A(i^*, m) = 1 or A(j∗,m)=1A(j^*, m) = 1
        for each active segment mm adjacent to k∗k^* do
            Compute candidate parents p(k∗,m)←f(W[pk∗;am]+b)p_{(k^*, m)} \leftarrow f(W [p_{k^*}; a_m] + b) and p(m,k∗)←f(W[am;pk∗]+b)p_{(m, k^*)} \leftarrow f(W [a_m; p_{k^*}] + b)
            Compute scores s(k∗,m)←Wscorep(k∗,m)s_{(k^*, m)} \leftarrow W^{\text{score}} p_{(k^*, m)} and s(m,k∗)←Wscorep(m,k∗)s_{(m, k^*)} \leftarrow W^{\text{score}} p_{(m, k^*)}
            Add [pk∗,am][p_{k^*}, a_m] and [am,pk∗][a_m, p_{k^*}] to CC
        end for
    end while
    return root node activation and hierarchical merge tree
  5. Knowl 5 — Continuous-Representation Beam Search Parsing for Sentences

    model/method

    For natural language sentences, the input adjacency structure is strictly linear (each word token ii only neighbors token i−1i-1 and token i+1i+1). Parsing is conducted using a bottom-up chart-based search analogous to the Cocke-Younger-Kasami (CKY) algorithm:

    • Unlike standard Context-Free Grammars (CFGs) in Chomsky Normal Form, every phrase constituent is represented by a continuous vector p∈Rnp \in \mathbb{R}^n rather than a discrete non-terminal category.
    • Because representations are continuous vectors, exact category-equality pruning is inapplicable.
    • In each span cell of the chart, beam search retains the single highest-scoring candidate constituent representation (keeping kk-best constituents per cell showed no empirical improvement over keeping the single best constituent).
  6. Knowl 6 — Category Classification at Recursive Neural Network Tree Nodes

    model/method

    Every node in the parse tree (both terminal leaves and recursively formed non-terminal parent nodes) has an associated nn-dimensional feature vector p∈Rnp \in \mathbb{R}^n. Semantic categories (such as object classes for image regions or syntactic phrase labels for sentence constituents) are predicted at each node by passing pp through a softmax classification layer:

    labelp=softmax(Wlabelp)\text{label}_p = \text{softmax}(W^{\text{label}} p)

    where Wlabel∈RC×nW^{\text{label}} \in \mathbb{R}^{C \times n} is a learned classification matrix and CC is the number of target classes.

    During training, the cross-entropy error of this softmax layer is minimized simultaneously with the parsing objective. The error gradients backpropagate through the tree structure, jointly updating the label classifier weights WlabelW^{\text{label}}, the compositional RNN parameters (W,b,Wscore)(W, b, W^{\text{score}}), and the input projection parameters (Wsem,bsemW^{\text{sem}}, b^{\text{sem}} for images or embedding lookup table LL for text).

  7. Knowl 7 — Optimization via Subgradient Descent and Backpropagation Through Structure

    model/method

    To optimize the non-differentiable structured max-margin objective J(θ)J(\theta), the subgradient method is used in conjunction with L-BFGS. Let θ=(Wsem,W,Wscore,Wlabel)\theta = (W^{\text{sem}}, W, W^{\text{score}}, W^{\text{label}}) (with WsemW^{\text{sem}} replaced by word lookup table LL for natural language parsing). The subgradient ∂J∂θ\frac{\partial J}{\partial \theta} is given by:

    ∂J∂θ=1N∑i=1N(∂s(y^i)∂θ−∂s(yi)∂θ)+λθ\frac{\partial J}{\partial \theta} = \frac{1}{N} \sum_{i=1}^N \left( \frac{\partial s(\hat{y}_i)}{\partial \theta} - \frac{\partial s(y_i)}{\partial \theta} \right) + \lambda \theta

    where:

    • y^i=arg⁡max⁡y^∈T(xi)(s(RNN(θ,xi,y^))+Δ(xi,li,y^))\hat{y}_i = \arg\max_{\hat{y} \in \mathcal{T}(x_i)} \left( s(\text{RNN}(\theta, x_i, \hat{y})) + \Delta(x_i, l_i, \hat{y}) \right) is the highest-scoring tree under the loss-augmented score.
    • yi=arg⁡max⁡y∈Y(xi,li)s(RNN(θ,xi,y))y_i = \arg\max_{y \in \mathcal{Y}(x_i, l_i)} s(\text{RNN}(\theta, x_i, y)) is the highest-scoring ground-truth tree.

    The gradient of a tree score ∂s(y)∂θ\frac{\partial s(y)}{\partial \theta} is computed using Backpropagation Through Structure (BPTS), in which incoming error signals at each non-terminal parent node are split and recursively propagated to its left and right child nodes.

  8. Knowl 8 — Pixel-Level Semantic Segmentation Accuracy on the Stanford Background Dataset

    data/table

    The Recursive Neural Network (RNN) framework was evaluated on the Stanford Background Dataset (containing 8 semantic classes: sky, tree, road, grass, water, building, mountain, foreground object) using 5-fold cross-validation. Superpixels were classified based on the multinomial distribution from the softmax layer at the tree leaves after end-to-end RNN training. The RNN model achieves state-of-the-art pixel accuracy:

    Method Semantic Pixel Accuracy (%)
    Pixel CRF (Gould et al., 2009) 74.3
    Logistic Regression on Superpixel Features 75.9
    Region-based energy (Gould et al., 2009) 76.4
    Local Labeling (Tighe Lazebnik, 2010) 76.9
    Superpixel MRF (Tighe Lazebnik, 2010) 77.5
    Simultaneous MRF (Tighe Lazebnik, 2010) 77.5
    RNN (Proposed method) 78.1

    A flat baseline using a single neural network layer followed by a softmax layer on individual superpixels (without tree structure) achieved approximately 76.1% accuracy (2% below the full RNN). On a 2.6 GHz laptop, the Matlab implementation required 16 seconds to parse and segment 143 test images.

  9. Knowl 9 — Scene Classification via Tree-Averaged Feature Representations

    empirical result

    Images from the Stanford Background Dataset were categorized into three global scene types: city, countryside, and sea-side. A linear Support Vector Machine (SVM) was trained on features extracted from the Recursive Neural Network parse trees:

    • All-Node Averaged RNN Features: Averaging the activation vectors p∈Rnp \in \mathbb{R}^n across all nodes in the parsed tree yielded a classification accuracy of 88.1%, outperforming standard Gist descriptors (84.0%).
    • Top-Node Only Baseline: Using solely the root node activation vector of the parse tree yielded an accuracy of 71.0% (substantially above the 33.3% random baseline, but 17.1% lower than averaging all nodes, showing that pooling across all hierarchical constituent levels preserves important sub-scene information).
  10. Knowl 10 — Natural Language Syntactic Parsing Performance on the Penn Treebank

    empirical result

    The Recursive Neural Network (RNN) parser was evaluated on the Wall Street Journal (WSJ) section of the Penn Treebank using standard splits (sections 2–21 for training, section 22 for development, and section 23 for testing) on sentences of length at most 15 words, with 100-dimensional vector representations (n=100n=100):

    • WSJ Test Set (Section 23, sentences ≤15\le 15 words):
      • RNN Parser Unlabeled Bracketing F1: 90.29%
      • Berkeley Parser (Petrov et al., 2006) Unlabeled Bracketing F1: 91.63%
    • WSJ Development Set (Section 22, sentences ≤15\le 15 words):
      • RNN Parser Unlabeled Bracketing F1: 92.06%
      • Berkeley Parser Unlabeled Bracketing F1: 92.08%

    The RNN achieves performance competitive with state-of-the-art parsers despite not receiving discrete syntactic category labels of child nodes during parsing decisions. On a 2.6 GHz laptop, the Matlab implementation parsed 421 sentences in 72 seconds.

Coverage note — All primary contributions—the RNN formulation for recursive parsing, input projection layers for images and sentences, the structured max-margin loss, greedy and chart-based parsing algorithms, backpropagation through structure, and empirical results on segmentation, scene classification, and WSJ syntactic parsing—are included. Qualitative nearest-neighbor visualizations were summarized under representation properties and omitted as standalone knowls.

References

  1. 1.Aude, O. and Torralba, A. Modeling the Shape of the Scene: A Holistic Representation of the Spatial Envelope. IJCV, 42, 2001.
  2. 2.Bengio, Y., Ducharme, R., Vincent, P., and Janvin, C. A neural probabilistic language model. JMLR, 3, 2003.
  3. 3.Collobert, R. and Weston, J. A unified architecture for natural language processing: deep neural networks with multitask learning. In ICML, 2008.
  4. 4.Comaniciu, D. and Meer, P. Mean shift: a robust approach toward feature space analysis. IEEE PAMI, 24(5):603–619, May 2002.
  5. 5.Goller, C. and K"uchler, A. Learning task-dependent distributed representations by backpropagation through structure. In ICNN, 1996.
  6. 6.Gould, S., Fulton, R., and Koller, D. Decomposing a Scene into Geometric and Semantically Consistent Regions. In ICCV, 2009.
  7. 7.Gupta, A. and Davis, L. S. Beyond nouns: Exploiting prepositions and comparative adjectives for learning visual classifiers. In ECCV, 2008.
  8. 8.Henderson, J. Neural network probability estimation for broad coverage parsing. In EACL, 2003.
  9. 9.Hinton, G. E. and Salakhutdinov, R. R. Reducing the dimensionality of data with neural networks. Science, 313, 2006.
  10. 10.Hoiem, D., Efros, A.A., and Hebert, M. Putting Objects in Perspective. CVPR, 2006.
  11. 11.Lee, H., Grosse, R., Ranganath, R., and Ng, A. Convolutional deep belief networks for scalable unsupervised learning of hierarchical representations. In ICML, 2009.
  12. 12.Li, L-J., Socher, R., and Fei-Fei, L. Towards total scene understanding:classification, annotation and segmentation in an automatic framework. In CVPR, 2009.
  13. 13.Manning, C. D. and Sch"utze, H. Foundations of Statistical Natural Language Processing. The MIT Press, Cambridge, Massachusetts, 1999.
  14. 14.Petrov, S., Barrett, L., Thibaux, R., and Klein, D. Learning accurate, compact, and interpretable tree annotation. In ACL, 2006.
  15. 15.Rabinovich, A., Vedaldi, A., Galleguillos, C., Wiewiora, E., and Belongie, S. Objects in context. In ICCV, 2007.
  16. 16.Ratliff, N., Bagnell, J. A., and Zinkevich, M. (Online) subgradient methods for structured prediction. In AIStats, 2007.
  17. 17.Schmid, Cordelia. Beyond bags of features: Spatial pyramid matching for recognizing natural scene categories. In CVPR, 2006.
  18. 18.Shotton, J., Winn, J., Rother, C., and Criminisi, A. Textonboost: Joint appearance, shape and context modeling for multi-class object recognition and segmentation. In ECCV, 2006.
  19. 19.Siskind, J. M., J. Sherman, Jr, Pollak, I., Harper, M. P., and Bouman, C. A. Spatial Random Tree Grammars for Modeling Hierarchal Structure in Images with Regions of Arbitrary Shape. IEEE PAMI, 29, 2007.
  20. 20.Socher, R. and Fei-Fei, L. Connecting modalities: Semisupervised segmentation and annotation of images using unaligned text corpora. In CVPR, 2010.
  21. 21.Socher, R., Manning, C. D., and Ng, A. Y. Learning continuous phrase representations and syntactic parsing with recursive neural networks. In Deep Learning and Unsupervised Feature Learning Workshop, 2010.
  22. 22.Taskar, B., Klein, D., Collins, M., Koller, D., and Manning, C. Max-margin parsing. In EMNLP, 2004.
  23. 23.Tighe, Joseph and Lazebnik, Svetlana. Superparsing: scalable nonparametric image parsing with superpixels. In ECCV, 2010.
  24. 24.Zhu, Long, Chen, Yuanhao, Torralba, Antonio, Freeman, William T., and Yuille, Alan L. Part and appearance sharing: Recursive Compositional Models for multiview. In CVPR, 2010.
  25. 25.Zhu, Song C. and Mumford, David. A stochastic grammar of images. Found. Trends. Comput. Graph. Vis., 2(4): 259–362, 2006.

Citation

MLA
Socher, R., et al. “Parsing Natural Scenes and Natural Language with Recursive Neural Networks”. International Conference on Machine Learning, 2011, pp. 129–36, http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.221.4910.
APA
Socher, R., Lin, C. C.-Y., Manning, C. D., & Ng, A. Y. (2011). Parsing Natural Scenes and Natural Language with Recursive Neural Networks. International Conference on Machine Learning, 129–136. http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.221.4910
Chicago
Socher, R., C. C.-Y. Lin, C. D. Manning, and A. Y. Ng. 2011. “Parsing Natural Scenes and Natural Language with Recursive Neural Networks”. International Conference on Machine Learning, 129–36. http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.221.4910.
Harvard
Socher, R. et al. (2011) “Parsing Natural Scenes and Natural Language with Recursive Neural Networks”, International Conference on Machine Learning, pp. 129–136. Available at: http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.221.4910.
Vancouver
1. Socher R, Lin CC-Y, Manning CD, Ng AY (2011) Parsing Natural Scenes and Natural Language with Recursive Neural Networks. International Conference on Machine Learning 129–136

BibTeX

@article{socher2011parsing,
  title = {Parsing Natural Scenes and Natural Language with Recursive Neural Networks},
  author = {Socher, Richard and Lin, Cliff Chiung-Yu and Manning, Christopher D. and Ng, Andrew Y.},
  year = {2011},
  journal = {International Conference on Machine Learning},
  pages = {129-136},
  url = {http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.221.4910}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors