Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models

Ryan KirosRuslan SalakhutdinovRichard S. Zemel

article2014arXiv1,467 citations

Proposes an encoder-decoder framework that unifies joint visual-semantic embeddings with a structure-content language model, enabling both bidirectional image-sentence retrieval and novel caption generation through multimodal vector space arithmetic.

Listen

Automatically generating natural language descriptions for visual content is a longstanding challenge that requires integrating computer vision, spatial reasoning, and natural language processing. Recent advances in deep neural networks have made accurate object recognition viable, but producing fluent, grammatically correct, and contextually precise captions remains difficult. Solving this problem is critical for scaling content-based image retrieval systems and building systems capable of visual question answering.

The article demonstrates an end-to-end framework that unifies visual-semantic embeddings with a novel language model. It aims to evaluate how effectively this pipeline can rank cross-modal data and generate accurate image descriptions from scratch by framing caption generation as a translation problem.

The authors implemented an encoder-decoder architecture evaluated across standard benchmark datasets, including Flickr8K, Flickr30K, Microsoft COCO, and the SBU Captioned Photo dataset. For the encoder, the pipeline maps image features extracted from deep convolutional networks into a shared multimodal space alongside text representations generated by a long short-term memory recurrent neural network. The decoder introduces a structure-content neural language model that separates sentence structure (such as parts of speech) from content representations. Performance was measured using standard ranking retrieval metrics, such as recall and median rank, alongside qualitative assessments of generated captions and vector space arithmetic.

The evaluation yielded several key findings. First, the long short-term memory encoder matched or surpassed prior state-of-the-art models on Flickr8K and Flickr30K benchmarks without requiring explicit, compute-heavy object detections. Second, pairing the encoder with a 19-layer Oxford convolutional network set new benchmark records, achieving top-1 image annotation recall of 18.0% on Flickr8K and 23.0% on Flickr30K, while reducing the median rank to 5 on Flickr30K. Third, the structure-content decoder successfully trained purely on text data while retaining the ability to generate captions directly from image vectors at test time. Finally, simpler linear encoders demonstrated multimodal arithmetic regularities (for example, modifying an image vector by subtracting and adding color words retrieved the modified image concept), though linear encoders proved inferior for ranking performance.

These findings indicate that explicit multimodal embedding spaces offer significant operational and computational advantages over perplexity-based language scoring methods. Because retrieval relies on fast matrix multiplication of pre-computed vectors, the approach dramatically improves scalability for large image databases. Additionally, enabling the language model to train on uncaptioned text reduces dependency on expensive, manually labeled image-caption datasets.

Organizations developing large-scale image search and automated metadata systems should adopt unified embedding architectures to lower computational costs while improving retrieval accuracy. As immediate next steps, the article recommends incorporating attention mechanisms to dynamically focus on specific image regions during generation, deploying deeper bidirectional encoders, and exploring whether integrating object detections can further refine descriptive fidelity.

Readers should note certain limitations: the generation scoring weights were tuned manually based on qualitative reviews due to the unreliability of automated metrics like BLEU. In addition, semantic vector arithmetic requires linear encoders, which trade off retrieval accuracy. However, confidence remains high in the core ranking and retrieval results due to consistent gains across standard benchmark datasets.

Cover for Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models

Abstract

Inspired by recent advances in multimodal learning and machine translation, we introduce an encoder-decoder pipeline that learns (a): a multimodal joint embedding space with images and text and (b): a novel language model for decoding distributed representations from our space. Our pipeline effectively unifies joint image-text embedding models with multimodal neural language models. We introduce the structure-content neural language model that disentangles the structure of a sentence to its content, conditioned on representations produced by the encoder. The encoder allows one to rank images and sentences while the decoder can generate novel descriptions from scratch. Using LSTM to encode sentences, we match the state-of-the-art performance on Flickr8K and Flickr30K without using object detections. We also set new best results when using the 19-layer Oxford convolutional network. Furthermore we show that with linear encoders, the learned embedding space captures multimodal regularities in terms of vector space arithmetic e.g. image of a blue car - "blue" + "red" is near images of red cars. Sample captions generated for 800 images are made available for comparison.

Table of Contents

  • 1 Introduction
  • 1.1 Multimodal representation learning
  • 1.2 Generating descriptions of images
  • 1.3 Encoder-decoder methods for machine translation
  • 2 An encoder-decoder model for ranking and generation
  • 2.1 Long short-term memory RNNs
  • 2.2 Multimodal distributed representations
  • 2.3 Log-bilinear neural language models
  • 2.4 Multiplicative neural language models
  • 2.5 Structure-content neural language models
  • 3 Experiments
  • 3.1 Image-sentence ranking
  • 3.1.1 Results
  • 3.2 Multimodal linguistic regularities
  • 3.3 Image caption generation
  • 4 Discussion
  • References
  • 5 Supplementary material: Additional experimentation and details
  • 5.1 Multimodal linguistic regularities
  • 5.2 Image description generation

Knowls

  1. Knowl 1 — Structure-Content Neural Language Model Architecture

    model/method

    The Structure-Content Neural Language Model (SC-NLM) is a conditional recurrent language model designed to generate image descriptions by disentangling sentence structure from semantic content. The model conditions the generation of each word wnw_n on two separate signals: a sequence of forward structure variables tn:n+k=(tn,ext...,tn+k)t_{n:n+k} = (t_n, ext{...}, t_{n+k}) over a context window of size kk (such as part-of-speech tags acting as soft syntactic templates) and a multimodal content embedding vector u∈RKu \in \mathbb{R}^K (produced by the visual-semantic encoder).

    The structure tokens ti∈RKt_i \in \mathbb{R}^K are retrieved from a learned lookup table. The combined structure-content attribute vector u^∈RG\hat{u} \in \mathbb{R}^G is computed as:

    u^=[∑i=nn+kT(i)ti+T(u)u+b]+\hat{u} = \left[ \sum_{i=n}^{n+k} T^{(i)} t_i + T^{(u)} u + b \right]_+

    where [z]+=max⁡(0,z)[z]_+ = \max(0, z) denotes the rectified linear unit (ReLU) activation function, T(i)∈RG×GT^{(i)} \in \mathbb{R}^{G \times G} are structure context matrices for each relative position i∈{n,...,n+k}i \in \{n, \text{...}, n+k\}, T(u)∈RG×KT^{(u)} \in \mathbb{R}^{G \times K} is a linear mapping matrix for the content embedding uu, and b∈RGb \in \mathbb{R}^G is a learnable bias vector. In typical configurations, the embedding dimensions are set to G=K=300G = K = 300.

  2. Knowl 2 — Factored Multiplicative Word Prediction in the SC-NLM

    equation

    In the Structure-Content Neural Language Model (SC-NLM), next-word probabilities are computed using a 3-way factored tensor interaction that modulates word embeddings by the combined structure-content attribute vector u^∈RG\hat{u} \in \mathbb{R}^G. Given context words (w1,...,wn−1)(w_1, \text{...}, w_{n-1}), forward structure tags tn:n+kt_{n:n+k}, and content vector uu, the probability of the nn-th word wn=i∈{1,...,V}w_n = i \in \{1, \text{...}, V\} (for vocabulary size VV) is defined by:

    P(wn=i∣w1:n−1,tn:n+k,u)=exp⁡(Wfv(:,i)⊤f+bi)∑j=1Vexp⁡(Wfv(:,j)⊤f+bj)P(w_n = i \mid w_{1:n-1}, t_{n:n+k}, u) = \frac{\exp\left( W_{fv}(:, i)^\top f + b_i \right)}{\sum_{j=1}^V \exp\left( W_{fv}(:, j)^\top f + b_j \right)}

    where bi∈Rb_i \in \mathbb{R} is the word output bias, Wfv(:,i)∈RFW_{fv}(:, i) \in \mathbb{R}^F is the ii-th column of factor matrix Wfv∈RF×VW_{fv} \in \mathbb{R}^{F \times V}, and f∈RFf \in \mathbb{R}^F is the factor vector defined as:

    f=(Wfkr^)⊙(Wfdu^)f = (W_{fk} \hat{r}) \odot (W_{fd} \hat{u})

    Here ⊙\odot represents component-wise multiplication, Wfk∈RF×KW_{fk} \in \mathbb{R}^{F \times K} and Wfd∈RF×GW_{fd} \in \mathbb{R}^{F \times G} are factor parameter matrices parameterized with F=100F = 100 factors, u^∈RG\hat{u} \in \mathbb{R}^G is the combined structure-content attribute vector, and r^∈RK\hat{r} \in \mathbb{R}^K is the linear context prediction:

    r^=∑j=1n−1C(j)E(:,wj)\hat{r} = \sum_{j=1}^{n-1} C^{(j)} E(:, w_j)

    where C(j)∈RK×KC^{(j)} \in \mathbb{R}^{K \times K} are context parameter matrices and E=(Wfk)⊤Wfv∈RK×VE = (W_{fk})^\top W_{fv} \in \mathbb{R}^{K \times V} is the folded word embedding matrix.

  3. Knowl 3 — Joint Visual-Semantic Embedding with LSTM Sentence Encoder

    model/method

    The visual-semantic encoder maps image feature vectors and natural language sentences into a shared KK-dimensional multimodal embedding space RK\mathbb{R}^K (K=300K = 300).

    Given an image feature vector q∈RDq \in \mathbb{R}^D extracted from the penultimate layer of a deep convolutional network (e.g., D=4096D = 4096 from AlexNet or a 19-layer OxfordNet/VGG), the image embedding is computed via a learned linear projection x=WIq∈RKx = W_I q \in \mathbb{R}^K, where WI∈RK×DW_I \in \mathbb{R}^{K \times D}.

    A sentence description S=(w1,...,wN)S = (w_1, \text{...}, w_N) consisting of words with pre-computed continuous bag-of-words (CBOW) embeddings wi∈RKw_i \in \mathbb{R}^K is encoded by a single-layer Long Short-Term Memory (LSTM) recurrent neural network. The LSTM updates its input gate ItI_t, forget gate FtF_t, memory cell CtC_t, output gate OtO_t, and hidden state MtM_t across time steps t=1,...,Nt = 1, \text{...}, N. The sentence representation v∈RKv \in \mathbb{R}^K is defined as the final hidden state of the LSTM:

    v=MNv = M_N

    The alignment score between an image embedding xx and a sentence embedding vv is their cosine similarity:

    s(x,v)=x∥x∥⋅v∥v∥s(x, v) = \frac{x}{\|x\|} \cdot \frac{v}{\|v\|}

  4. Knowl 4 — Pairwise Ranking Loss for Image-Sentence Embeddings

    equation

    The parameters θ={WI,LSTM weights}\theta = \{W_I, \text{LSTM weights}\} of the visual-semantic embedding model are learned by minimizing a symmetric margin-based pairwise ranking loss across all training image embeddings xx and sentence embeddings vv:

    min⁡θ∑x∑kmax⁡{0,α−s(x,v)+s(x,vk)}+∑v∑kmax⁡{0,α−s(v,x)+s(v,xk)}\min_{\theta} \sum_x \sum_k \max\{0, \alpha - s(x, v) + s(x, v_k)\} + \sum_v \sum_k \max\{0, \alpha - s(v, x) + s(v, x_k)\}

    where s(x,v)=x⊤v∥x∥∥v∥s(x, v) = \frac{x^\top v}{\|x\| \|v\|} is the cosine similarity score, vkv_k is a contrastive (non-descriptive) sentence embedding for image xx, xkx_k is a contrastive image embedding for sentence vv, and α=0.2\alpha = 0.2 is the margin hyperparameter. Word embeddings WT∈RK×VW_T \in \mathbb{R}^{K \times V} are kept fixed during training, and contrastive samples are randomly selected from the training set and resampled at each epoch.

  5. Knowl 5 — Decoupled Training of the Structure-Content Decoder via Multimodal Embeddings

    model/method

    The Structure-Content Neural Language Model (SC-NLM) decoder is trained purely on text corpora without needing paired image-text instances during decoder optimization. During training, the content conditioning vector u∈RKu \in \mathbb{R}^K is provided by the sentence representation v=LSTM(S)v = \text{LSTM}(S) produced by the LSTM sentence encoder on caption SS.

    Because the encoder projects both images and sentences into a common multimodal space where matching image embeddings xx and description embeddings vv are closely aligned (x≈vx \approx v), the decoder can be conditioned directly on image embeddings (u=xu = x) at inference time. This decoupling allows the language model decoder to be trained or scaled on large monolingual text datasets lacking paired visual data.

  6. Knowl 6 — Image Caption Generation and Candidate Rescoring Algorithm

    algorithm

    The caption generation pipeline maps an image into the shared multimodal space, generates candidate descriptions via structure-conditioned sampling from the SC-NLM, and re-ranks candidates using an ensemble scoring function.

    Input: Image feature vector q∈RDq \in \mathbb{R}^D, projection matrix WIW_I, LSTM encoder, SC-NLM decoder, Kneser-Ney trigram language model LMLM, set of training POS sequences TPOS\mathcal{T}_{POS} with lengths in {4,…,12}\{4, \dots, 12\}, sample count M=1000M = 1000, nearest-neighbor count N=5N = 5
    Output: Ranked list of top 5 generated captions
    Compute unit-normalized image embedding x=WIq/∥WIq∥x = W_I q / \|W_I q\|
    Retrieve top NN nearest words and NN nearest sentences to xx in embedding space via cosine similarity
    Compute mean concept vector uconcept=12N(∑j=1Nw(j)+∑j=1Nv(j))u_{concept} = \frac{1}{2N} (\sum_{j=1}^N w^{(j)} + \sum_{j=1}^N v^{(j)})
    Form candidate conditioning set U={x,uconcept}\mathcal{U} = \{x, u_{concept}\}
    Initialize candidate pool C=∅\mathcal{C} = \emptyset
    for m=1m = 1 to MM do
        Sample conditioning vector u∈Uu \in \mathcal{U} uniformly
        Sample POS sequence T=(t1,…,tL)∈TPOST = (t_1, \dots, t_L) \in \mathcal{T}_{POS} uniformly
        Generate candidate sentence Sm=(w1,…,wL)S_m = (w_1, \dots, w_L) by MAP decoding from SC-NLM conditioned on uu and TT
        Compute sentence embedding vm=LSTM(Sm)v_m = \text{LSTM}(S_m) and unit-normalize vmv_m
        Compute translation relevance score strans(Sm)=x⊤vms_{trans}(S_m) = x^\top v_m
        Apply multiplicative repetition penalty to strans(Sm)s_{trans}(S_m) for repeated non-stopwords in SmS_m
        Compute fluency score sLM(Sm)=log⁡PLM(Sm)s_{LM}(S_m) = \log P_{LM}(S_m) under Kneser-Ney trigram model
        Compute combined score score(Sm)=wtrans⋅strans(Sm)+wLM⋅sLM(Sm)score(S_m) = w_{trans} \cdot s_{trans}(S_m) + w_{LM} \cdot s_{LM}(S_m)
        Add (Sm,score(Sm))(S_m, score(S_m)) to C\mathcal{C}
    end for
    Sort candidate descriptions in C\mathcal{C} in descending order of combined score
    return Top 5 descriptions from C\mathcal{C}
  7. Knowl 7 — Multimodal Vector Space Arithmetic with Linear Sentence Encoders

    model/method

    When a multimodal joint embedding model is trained with an additive linear sentence encoder v=∑i=1Nwiv = \sum_{i=1}^N w_i (where wi∈RKw_i \in \mathbb{R}^K are word embeddings) with unit normalization rather than an LSTM, the learned cross-modal space exhibits linear compositional vector arithmetic across images and text.

    For instance, given the embedding of an image with a blue car Ibcar≈vblue+vcarI_{bcar} \approx v_{blue} + v_{car} and the embeddings of attribute words vbluev_{blue} and vredv_{red}, subtracting the negative word vector and adding the positive word vector yields an approximation of the target image embedding: Ircar≈Ibcar−vblue+vredI_{rcar} \approx I_{bcar} - v_{blue} + v_{red}.

    Given a query image feature vector qq, a negative word vector wnw_n, and a positive word vector wpw_p (all unit normalized), the target image x∗x^* is retrieved via:

    x∗=argmax⁡x(q−wn+wp)⊤x∥q−wn+wp∥x^* = \operatorname{argmax}_x \frac{(q - w_n + w_p)^\top x}{\|q - w_n + w_p\|}

    To remove spurious nearest-neighbor results, a re-ranking step retrieves the top NN candidates and re-sorts them by Euclidean distance to their empirical mean vector.

  8. Knowl 8 — Cross-Modal Retrieval Performance on Flickr8K

    data/table

    Image-sentence ranking experiments on Flickr8K (8,000 images, 5 descriptions per image; 1,000 test images) evaluate Image Annotation (retrieving sentences given an image query) and Image Search (retrieving images given a sentence query). Evaluation metrics are Recall@K (R@K: percentage of queries where a correct match is within the top KK) and Median Rank (Med r) of the first ground-truth item.

    Image Annotation Image Search
    Model R@1 R@5 R@10 Med r R@1 R@5 R@10 Med r
    Random Ranking 0.1 0.6 1.1 631 0.1 0.5 1.0 500
    SDT-RNN 4.5 18.0 28.6 32 6.1 18.5 29.0 29
    DeViSE ()† 4.8 16.5 27.3 28 5.9 20.1 29.6 29
    SDT-RNN ()† 6.0 22.7 34.0 23 6.6 21.6 31.7 25
    DeFrag 5.9 19.2 27.3 34 5.2 17.6 26.5 32
    DeFrag ()† 12.6 32.9 44.0 14 9.7 29.6 42.5 15
    m-RNN 14.5 37.2 48.5 11 11.5 31.0 42.4 15
    Our model (Toronto ConvNet) 13.5 36.2 45.7 13 10.4 31.0 43.7 14
    Our model (OxfordNet) 18.0 40.9 55.0 8 12.5 37.0 51.5 10

    A dagger (\dag) indicates the use of R-CNN object detection features in addition to full-frame features. The LSTM sentence encoder outperforms tree-based models and approaches that use object detections, while replacing the 4096-dimensional Toronto ConvNet features with 19-layer OxfordNet features establishes new state-of-the-art results across all retrieval metrics.

  9. Knowl 9 — Cross-Modal Retrieval Performance on Flickr30K

    data/table

    Cross-modal retrieval evaluation on Flickr30K (30,000 images, 5 descriptions per image; 1,000 test images) comparing full-frame and object-detection-based embedding methods against the LSTM sentence encoder.

    Image Annotation Image Search
    Model R@1 R@5 R@10 Med r R@1 R@5 R@10 Med r
    Random Ranking 0.1 0.6 1.1 631 0.1 0.5 1.0 500
    DeViSE ()† 4.5 18.1 29.2 26 6.7 21.9 32.7 25
    SDT-RNN ()† 9.6 29.8 41.1 16 8.9 29.8 41.1 16
    DeFrag ()† 14.2 37.7 51.3 10 10.2 30.8 44.2 14
    DeFrag + Finetune CNN ()† 16.4 40.2 54.7 8 10.3 31.4 44.5 13
    m-RNN 18.4 40.2 50.9 10 12.6 31.2 41.5 16
    Our model (Toronto ConvNet) 14.8 39.2 50.9 10 11.8 34.0 46.3 13
    Our model (OxfordNet) 23.0 50.7 62.9 5 16.8 42.0 56.5 8

    A dagger (\dag) denotes methods incorporating R-CNN fragment/object detections. With OxfordNet features, the LSTM encoder achieves R@1 of 23.0% on image annotation (Median rank 5) and R@1 of 16.8% on image search (Median rank 8), significantly outperforming all prior models.

  10. Knowl 10 — Trade-off Between Multimodal Vector Arithmetic and Cross-Modal Retrieval Accuracy

    empirical result

    There is a fundamental trade-off between semantic vector space arithmetic and retrieval accuracy depending on sentence encoder architecture:

    1. Linear bag-of-words sentence encoders (v=∑i=1Nwiv = \sum_{i=1}^N w_i) exhibit linear multimodal vector arithmetic (such as image of a blue car −- blue ++ red ≈\approx image of a red car). However, linear encoders discard word order and syntax, leading to poor retrieval performance (e.g., DeViSE achieves only R@1 of 4.5% on Flickr30K Image Annotation and 6.7% on Image Search).
    2. Recurrent LSTM encoders capture sequential dependencies and sentence compositionality, yielding substantially higher cross-modal retrieval performance (Flickr30K Image Annotation R@1 of 14.8% with Toronto ConvNet and 23.0% with OxfordNet). However, the resulting non-linear vector space does not preserve linear additive/subtractive arithmetic regularities.

Coverage note — None; all primary contributions, including the visual-semantic LSTM encoder, the SC-NLM factored multiplicative decoder, caption generation and scoring algorithm, multimodal vector arithmetic, and empirical benchmark evaluations on Flickr8K and Flickr30K, are covered.

References

  1. 1.Sepp Hochreiter and Jürgen Schmidhuber. Long short-term memory. Neural computation, 1997.
  2. 2.Ryan Kiros, Richard S Zemel, and Ruslan Salakhutdinov. Multimodal neural language models. ICML, 2014.
  3. 3.Micah Hodosh, Peter Young, and Julia Hockenmaier. Framing image description as a ranking task: Data, models and evaluation metrics. JAIR, 2013.
  4. 4.Jason Weston, Samy Bengio, and Nicolas Usunier. Large scale image annotation: learning to rank with joint word-image embeddings. Machine learning, 2010.
  5. 5.Andrea Frome, Greg S Corrado, Jon Shlens, Samy Bengio, Jeffrey Dean, and Tomas Mikolov MarcAurelio Ranzato. Devise: A deep visual-semantic embedding model. NIPS, 2013.
  6. 6.Richard Socher, Q Le, C Manning, and A Ng. Grounded compositional semantics for finding and describing images with sentences. In TACL, 2014.
  7. 7.Junhua Mao, Wei Xu, Yi Yang, Jiang Wang, and Alan L Yuille. Explain images with multimodal recurrent neural networks. arXiv preprint arXiv:1410.1090, 2014.
  8. 8.Nal Kalchbrenner and Phil Blunsom. Recurrent continuous translation models. In EMNLP, 2013.
  9. 9.Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. Learning phrase representations using rnn encoder-decoder for statistical machine translation. EMNLP, 2014.
  10. 10.Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473, 2014.
  11. 11.Ilya Sutskever, Oriol Vinyals, and Quoc V Le. Sequence to sequence learning with neural networks. NIPS, 2014.
  12. 12.Tomas Mikolov, Wen-tau Yih, and Geoffrey Zweig. Linguistic regularities in continuous space word representations. In NAACL-HLT, 2013.
  13. 13.Karl Moritz Hermann and Phil Blunsom. Multilingual distributed representations without word alignment. ICLR, 2014.
  14. 14.Karl Moritz Hermann and Phil Blunsom. Multilingual models for compositional distributional semantics. In ACL, 2014.
  15. 15.Andrej Karpathy, Armand Joulin, and Li Fei-Fei. Deep fragment embeddings for bidirectional image sentence mapping. NIPS, 2014.
  16. 16.Nitish Srivastava and Ruslan Salakhutdinov. Multimodal learning with deep boltzmann machines. In NIPS, 2012.
  17. 17.Jiquan Ngiam, Aditya Khosla, Mingyu Kim, Juhan Nam, Honglak Lee, and Andrew Ng. Multimodal deep learning. In ICML, 2011.
  18. 18.Yangqing Jia, Mathieu Salzmann, and Trevor Darrell. Learning cross-modality similarity for multinomial data. In ICCV, 2011.
  19. 19.Yunchao Gong, Liwei Wang, Micah Hodosh, Julia Hockenmaier, and Svetlana Lazebnik. Improving image-sentence embeddings using large weakly annotated photo collections. In ECCV. 2014.
  20. 20.Phil Blunsom, Nando de Freitas, Edward Grefenstette, Karl Moritz Hermann, et al. A deep architecture for semantic parsing. In ACL 2014 Workshop on Semantic Parsing, 2014.
  21. 21.Girish Kulkarni, Visruth Premraj, Sagnik Dhar, Siming Li, Yejin Choi, Alexander C Berg, and Tamara L Berg. Baby talk: Understanding and generating simple image descriptions. In CVPR, 2011.
  22. 22.Ali Farhadi, Mohsen Hejrati, Mohammad Amin Sadeghi, Peter Young, Cyrus Rashtchian, Julia Hockenmaier, and David Forsyth. Every picture tells a story: Generating sentences from images. In ECCV. 2010.
  23. 23.Siming Li, Girish Kulkarni, Tamara L Berg, Alexander C Berg, and Yejin Choi. Composing simple image descriptions using web-scale n-grams. In CONLL, 2011.
  24. 24.Yezhou Yang, Ching Lik Teo, Hal Daumé III, and Yiannis Aloimonos. Corpus-guided sentence generation of natural images. In EMNLP, 2011.
  25. 25.Margaret Mitchell, Xufeng Han, Jesse Dodge, Alyssa Mensch, Amit Goyal, Alex Berg, Kota Yamaguchi, Tamara Berg, Karl Stratos, and Hal Daumé III. Midge: Generating image descriptions from computer vision detections. In EACL, 2012.
  26. 26.Polina Kuznetsova, Vicente Ordonez, Alexander C Berg, Tamara L Berg, and Yejin Choi. Collective generation of natural image descriptions. ACL, 2012.
  27. 27.Polina Kuznetsova, Vicente Ordonez, Tamara L. Berg, and Yejin Choi. Treetalk : Composition and compression of trees for image descriptions. TACL, 2014.
  28. 28.Marcus Rohrbach, Wei Qiu, Ivan Titov, Stefan Thater, Manfred Pinkal, and Bernt Schiele. Translating video content to natural language descriptions. In ICCV, 2013.
  29. 29.Andriy Mnih and Geoffrey Hinton. Three new graphical models for statistical language modelling. In ICML, pages 641–648, 2007.
  30. 30.Ryan Kiros, Richard S Zemel, and Ruslan Salakhutdinov. A multiplicative model for learning distributed text-based attribute representations. NIPS, 2014.
  31. 31.Alex Graves, Marcus Liwicki, Santiago Fernández, Roman Bertolami, Horst Bunke, and Jürgen Schmidhuber. A novel connectionist system for unconstrained handwriting recognition. TPAMI, 2009.
  32. 32.Alex Graves. Generating sequences with recurrent neural networks. arXiv preprint arXiv:1308.0850, 2013.
  33. 33.Alex Graves, Navdeep Jaitly, and Abdel-rahman Mohamed. Hybrid speech recognition with deep bidirectional lstm. In IEEE Workshop on ASRU, 2013.
  34. 34.Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. Dropout: A simple way to prevent neural networks from overfitting. JMLR, 2014.
  35. 35.Wojciech Zaremba, Ilya Sutskever, and Oriol Vinyals. Recurrent neural network regularization. arXiv preprint arXiv:1409.2329, 2014.
  36. 36.Alex Krizhevsky, Ilya Sutskever, and Geoff Hinton. Imagenet classification with deep convolutional neural networks. In NIPS, 2012.
  37. 37.Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781, 2013.
  38. 38.Roland Memisevic and Geoffrey Hinton. Unsupervised learning of image transformations. In CVPR, pages 1–8, 2007.
  39. 39.Alex Krizhevsky, Geoffrey E Hinton, et al. Factored 3-way restricted boltzmann machines for modeling natural images. In AISTATS, pages 621–628, 2010.
  40. 40.Vicente Ordonez, Girish Kulkarni, and Tamara L Berg. Im2text: Describing images using 1 million captioned photographs. In NIPS, 2011.
  41. 41.Jacob Devlin, Rabih Zbib, Zhongqiang Huang, Thomas Lamar, Richard Schwartz, and John Makhoul. Fast and robust neural network joint models for statistical machine translation. ACL, 2014.
  42. 42.Peter Young Alice Lai Micah Hodosh and Julia Hockenmaier. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. TACL, 2014.
  43. 43.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014.
  44. 44.Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. CVPR, 2014.
  45. 45.Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Janvin. A neural probabilistic language model. JMLR, 2003.
  46. 46.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. arXiv preprint arXiv:1405.0312, 2014.

Citation

MLA
Kiros, R., et al. “Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models”. arXiv, 2014, http://arxiv.org/abs/1411.2539v1.
APA
Kiros, R., Salakhutdinov, R., & Zemel, R. S. (2014). Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models. arXiv. http://arxiv.org/abs/1411.2539v1
Chicago
Kiros, R., R. Salakhutdinov, and R. S. Zemel. 2014. “Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models”. arXiv. http://arxiv.org/abs/1411.2539v1.
Harvard
Kiros, R., Salakhutdinov, R. and Zemel, R.S. (2014) “Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1411.2539v1.
Vancouver
1. Kiros R, Salakhutdinov R, Zemel RS (2014) Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models. arXiv

BibTeX

@article{kiros2014unifying,
  title = {Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models},
  author = {Kiros, Ryan and Salakhutdinov, Ruslan and Zemel, Richard S.},
  year = {2014},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1411.2539v1},
  eprint = {1411.2539}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors