From captions to visual concepts and back

Hao FangSaurabh GuptaF. IandolaR. SrivastavaL. DengPiotr DollárJianfeng GaoXiaodong HeMargaret MitchellJohn C. Platt

article2014CVPR1,363 citations

Presents an image captioning pipeline that learns visual concept detectors directly from weakly-supervised captions via multiple instance learning, generates candidate descriptions with a maximum-entropy language model, and selects the best description using a deep multimodal similarity model.

Listen

Automatically generating accurate, natural language descriptions of images is a fundamental challenge in artificial intelligence, with applications ranging from digital accessibility to automated media indexing. Traditional captioning systems rely heavily on expensive, hand-labeled bounding box annotations and struggle to identify abstract visual concepts or construct fluent, commonsense descriptions. The article addresses this challenge by evaluating an automated image captioning pipeline that learns visual concepts, language statistics, and cross-modal relevance directly from paired image and caption datasets without requiring manual bounding boxes.

The authors develop a three-stage framework. First, they train visual detectors for a 1,000-word vocabulary across various parts of speech using weakly supervised multiple instance learning applied to sub-regions of images via deep convolutional neural networks. Second, a maximum-entropy statistical language model takes these detected words to generate candidate sentences through a beam-search optimization process. Finally, a deep multimodal similarity model maps images and text into a shared semantic vector space, re-ranking candidate sentences to select the best caption. The framework was evaluated on the benchmark Microsoft Common Objects in Context (COCO) dataset, spanning over 80,000 training images and an official test set of over 40,000 images, alongside the PASCAL sentence dataset.

The experimental findings show that the proposed approach achieves state-of-the-art results across standard benchmarks. On the official Microsoft COCO test server, the system achieved a BLEU-4 score of 29.1% (compared to 21.7% for human-written captions) and equaled or surpassed human performance benchmarks on 12 of 14 evaluation metrics, including CIDEr. In human subjective assessments, human evaluators judged the system's generated captions to be of equal or better quality than human-written captions 34% of the time. Additionally, training visual detectors via multiple instance learning on image sub-regions systematically outperformed whole-image classification baselines, achieving an average precision of 34.0% across all vocabulary categories compared to 30.8% for standard image classifiers.

These results demonstrate that systems can learn rich, salient visual concepts—including verbs and adjectives—directly from raw caption data without costly bounding box annotations. By combining statistical language models with global multimodal re-ranking, the pipeline filters out visual noise and preserves commonsense semantics. This substantially lowers data labeling costs while delivering caption quality suitable for real-world deployment.

Organizations implementing automated image description should adopt weakly supervised region-based visual detection paired with multimodal semantic re-ranking rather than relying solely on end-to-end language models or costly hand-annotated object boxes. Decision-makers should note that while automatic metric scores frequently exceed human reference numbers, human judges still prefer human-written descriptions in roughly two-thirds of cases, meaning automatic metrics should not be treated as a complete replacement for human evaluation. Further work should explore refining abstract relationship detection, expanding vocabulary coverage beyond frequent words, and conducting pilot testing on domain-specific imagery before deploying in high-risk operational environments.

arXiv: 1411.4952
Cover for From captions to visual concepts and back

Abstract

This paper presents a novel approach for automatically generating image descriptions: visual detectors, language models, and multimodal similarity models learnt directly from a dataset of image captions. We use multiple instance learning to train visual detectors for words that commonly occur in captions, including many different parts of speech such as nouns, verbs, and adjectives. The word detector outputs serve as conditional inputs to a maximum-entropy language model. The language model learns from a set of over 400,000 image descriptions to capture the statistics of word usage. We capture global semantics by re-ranking caption candidates using sentence-level features and a deep multimodal similarity model. Our system is state-of-the-art on the official Microsoft COCO benchmark, producing a BLEU-4 score of 29.1%. When human judges compare the system captions to ones written by other people on our held-out test set, the system captions have equal or better quality 34% of the time.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Word Detection
  • 3.1 Training Word Detectors
  • 3.2 Generating Word Scores for a Test Image
  • 4 Language Generation
  • 4.1 Statistical Model
  • 4.2 Generation Process
  • 5 Sentence Re-Ranking
  • 5.1 Deep Multimodal Similarity Model
  • 6 Experimental Results
  • 6.1 Datasets
  • 6.2 Word Detection
  • 6.3 Caption Generation
  • 7 Conclusion
  • References

Knowls

  1. Knowl 1 — Pipeline Architecture for Image Caption Generation

    model/method

    The image captioning framework decomposes sentence generation from an input image into a three-stage pipeline:

    1. Visual Word Detection: A convolutional neural network (CNN) combined with Multiple Instance Learning (MIL) predicts a set of salient vocabulary words (spanning nouns, verbs, and adjectives) associated with sub-regions of the image, without requiring bounding box annotations.
    2. Constrained Language Generation: A Maximum Entropy Language Model (ME LM), conditioned on the set of detected visual words that have not yet been mentioned, generates candidate high-likelihood sentences using left-to-right beam search.
    3. Global Sentence Re-Ranking: Candidate sentences are re-ranked using Minimum Error Rate Training (MERT) over a linear combination of sentence-level linguistic features and a global semantic compatibility score computed by a Deep Multimodal Similarity Model (DMSM).
  2. Knowl 2 — Weakly-Supervised Word Detection via Noisy-OR Multiple Instance Learning

    model/method

    To detect vocabulary words from an image without bounding box annotations, each training image ii is treated as a bag bib_i of overlapping sub-regions bijb_{ij}. For each target word w∈Vw \in \mathcal{V} in a vocabulary V\mathcal{V} (the 1000 most common words in caption data), the probability piwp_i^w that image ii contains word ww is modeled using a Noisy-OR Multiple Instance Learning (MIL) formulation:

    piw=1−∏j∈bi(1−pijw)p_i^w = 1 - \prod_{j \in b_i} (1 - p_{ij}^w)

    where pijwp_{ij}^w is the instance-level probability that sub-region jj corresponds to word ww. The instance probability is parameterized by applying a logistic sigmoid to the fc7 feature representation ϕ(bij)\phi(b_{ij}) of a convolutional neural network:

    pijw=11+exp⁡(−(vwTϕ(bij)+uw))p_{ij}^w = \frac{1}{1 + \exp\left(-\left(v_w^T \phi(b_{ij}) + u_w\right)\right)}

    with learnable weight vector vwv_w and bias uwu_w.

    The CNN fully connected layers (fc6, fc7, fc8) are expressed as convolutions to yield a fully convolutional network. When an image is resized so its longer side is 565 pixels, forward propagation produces a 12×1212 \times 12 spatial response map at fc8 (equivalent to sliding a 224×224224 \times 224 window with a stride of 32). The CNN is optimized end-to-end with binary cross-entropy loss using stochastic gradient descent for 3 epochs. At test time, raw image-level probabilities piwp_i^w are calibrated into precisions using a held-out subset of training data, and all words meeting or exceeding a precision threshold τ\tau form the candidate detection set V~\tilde{\mathcal{V}} alongside their detector scores.

  3. Knowl 3 — Maximum Entropy Language Model Conditioned on Remaining Visual Detections

    model/method

    Candidate sentences are generated by a Maximum Entropy Language Model (ME LM) that estimates the probability of emitting the next word wlw_l conditioned on preceding words w1,…,wl−1w_1, \dots, w_{l-1} and the subset of detected visual words V~l−1⊆V~\tilde{\mathcal{V}}_{l-1} \subseteq \tilde{\mathcal{V}} that have not yet appeared in the sentence:

    Pr⁡(wl=wˉl∣wˉl−1,…,wˉ1,⟨s⟩,V~l−1)=exp⁡(∑k=1Kλkfk(wˉl,wˉl−1,…,wˉ1,⟨s⟩,V~l−1))∑v∈V∪{⟨/s⟩}exp⁡(∑k=1Kλkfk(v,wˉl−1,…,wˉ1,⟨s⟩,V~l−1))\Pr(w_l = \bar{w}_l \mid \bar{w}_{l-1}, \dots, \bar{w}_1, \langle s \rangle, \tilde{\mathcal{V}}_{l-1}) = \frac{\exp\left(\sum_{k=1}^K \lambda_k f_k(\bar{w}_l, \bar{w}_{l-1}, \dots, \bar{w}_1, \langle s \rangle, \tilde{\mathcal{V}}_{l-1})\right)}{\sum_{v \in \mathcal{V} \cup \{\langle /s \rangle\}} \exp\left(\sum_{k=1}^K \lambda_k f_k(v, \bar{w}_{l-1}, \dots, \bar{w}_1, \langle s \rangle, \tilde{\mathcal{V}}_{l-1})\right)}

    where ⟨s⟩\langle s \rangle and ⟨/s⟩\langle /s \rangle denote the start-of-sentence and end-of-sentence tokens, λk\lambda_k is the weight of feature fkf_k, and wˉj∈V∪{⟨/s⟩}\bar{w}_j \in \mathcal{V} \cup \{\langle /s \rangle\}. The top 15 most frequent closed-class words (e.g., a, on, of, the, in, with, and, is, to, an, at, are, next, that, it) are excluded from V~\tilde{\mathcal{V}} to avoid trivial generation.

    The feature functions fkf_k include:

    • Attribute (0/10/1): 11 if wˉl∈V~l−1\bar{w}_l \in \tilde{\mathcal{V}}_{l-1} (predicted word is an unmentioned detected word).
    • N-gram+ (0/10/1): 11 if the NN-gram ending in wˉl\bar{w}_l matches pattern κ\kappa and wˉl∈V~l−1\bar{w}_l \in \tilde{\mathcal{V}}_{l-1} (NN up to 4).
    • N-gram- (0/10/1): 11 if the NN-gram ending in wˉl\bar{w}_l matches pattern κ\kappa and wˉl∉V~l−1\bar{w}_l \notin \tilde{\mathcal{V}}_{l-1}.
    • End (0/10/1): 11 if wˉl=⟨/s⟩\bar{w}_l = \langle /s \rangle and all detected attributes have been mentioned (V~l−1=∅\tilde{\mathcal{V}}_{l-1} = \emptyset).
    • Score (R\mathbb{R}): The visual detector log-probability log⁡piwˉl\log p_i^{\bar{w}_l} when wˉl∈V~l−1\bar{w}_l \in \tilde{\mathcal{V}}_{l-1}, and 00 otherwise.

    Training maximizes conditional log-likelihood and is accelerated using Noise Contrastive Estimation (NCE) with 15 negative samples per token.

  4. Knowl 4 — Left-to-Right Beam Search Sentence Generation with Target Attribute Coverage

    algorithm

    Sentence generation searches for high-likelihood word sequences covering detected visual concepts using a left-to-right beam search, followed by forming an MM-best candidate list based on target attribute coverage TT.

    Input: Detected attribute set V~\tilde{\mathcal{V}}, maximum sentence length LL, beam width kk, target candidate count MM, initial target attribute count TT
    Output: MM-best candidate sentence list M\mathcal{M}
    Initialize stack $S_0 \leftarrow \{(\text{path}=[\langle s \rangle], \text{score}=0.0, \text{rem\_attrs}=\tilde{\mathcal{V}})\}
    Initialize completed list C←∅\mathcal{C} \leftarrow \emptyset
    for l=0l = 0 to L−1L-1:
        Initialize candidate paths Pl+1←∅P_{l+1} \leftarrow \emptyset
        for each hypothesis h∈Slh \in S_l:
            wlast←w_{\text{last}} \leftarrow last word in h.pathh.\text{path}
            candidates←{⟨/s⟩}∪(top 100 frequent words)∪h.rem_attrs∪(observed bigram followers of wlast)\text{candidates} \leftarrow \{\langle /s \rangle\} \cup (\text{top 100 frequent words}) \cup h.\text{rem\_attrs} \cup (\text{observed bigram followers of } w_{\text{last}})
            for each word w∈candidatesw \in \text{candidates}:
                scorel+1←h.score+log⁡Pr⁡(w∣h.path,h.rem_attrs)\text{score}_{l+1} \leftarrow h.\text{score} + \log \Pr(w \mid h.\text{path}, h.\text{rem\_attrs})
                new_rem←h.rem_attrs∖{w}\text{new\_rem} \leftarrow h.\text{rem\_attrs} \setminus \{w\}
                new_path←h.path+[w]\text{new\_path} \leftarrow h.\text{path} + [w]
                if w==⟨/s⟩w == \langle /s \rangle:
                    C←C∪{(path=new_path,score=scorel+1,rem_attrs=new_rem)}\mathcal{C} \leftarrow \mathcal{C} \cup \{(\text{path}=\text{new\_path}, \text{score}=\text{score}_{l+1}, \text{rem\_attrs}=\text{new\_rem})\}
                else:
                    Pl+1←Pl+1∪{(path=new_path,score=scorel+1,rem_attrs=new_rem)}P_{l+1} \leftarrow P_{l+1} \cup \{(\text{path}=\text{new\_path}, \text{score}=\text{score}_{l+1}, \text{rem\_attrs}=\text{new\_rem})\}
        Sl+1←S_{l+1} \leftarrow top kk hypotheses in Pl+1P_{l+1} sorted descending by score\text{score}
    repeat:
        M←\mathcal{M} \leftarrow all sentences in C\mathcal{C} covering at least TT detected attributes, sorted descending by log-likelihood score
        if ∣M∣<M|\mathcal{M}| < M:
            T←T−1T \leftarrow T - 1
    until ∣M∣≥M|\mathcal{M}| \ge M or T<0T < 0
    M←\mathcal{M} \leftarrow top MM sentences from M\mathcal{M}
    return M\mathcal{M}
  5. Knowl 5 — Deep Multimodal Similarity Model (DMSM) for Image-Sentence Alignment

    model/method

    The Deep Multimodal Similarity Model (DMSM) embeds an image QQ and a text candidate DD into a shared continuous semantic vector space where their global semantic relevance is measured by cosine similarity:

    R(Q,D)=cosine(yQ,yD)=yQTyD∥yQ∥∥yD∥R(Q, D) = \text{cosine}(y_Q, y_D) = \frac{y_Q^T y_D}{\|y_Q\| \|y_D\|}

    The model consists of two subnetworks:

    1. Image Network: The fc7 layer representation of a CNN (AlexNet or VGG fine-tuned on full-image multi-label word prediction) is passed through three stacked fully connected layers with hyperbolic tangent (tanh⁡\tanh) activation functions to produce image semantic vector yQy_Q.
    2. Text Network: Words in a candidate caption are converted to letter-trigram count vectors to capture character n-gram statistics and provide robustness against rare or misspelled words. The representations are passed through a deep convolutional neural network to produce text semantic vector yDy_D.

    Given an image query QQ, a matching caption D+D^+, and N=50N = 50 randomly chosen non-matching captions D−D^-, the posterior probability of the correct caption is defined by a softmax over relevance scores:

    P(D+∣Q)=exp⁡(γR(Q,D+))∑D′∈{D+}∪{D−}exp⁡(γR(Q,D′))P(D^+ \mid Q) = \frac{\exp(\gamma R(Q, D^+))}{\sum_{D' \in \{D^+\} \cup \{D^-\}} \exp(\gamma R(Q, D'))}

    where γ=10\gamma = 10 is a scaling factor tuned on validation data. The parameters Λ\Lambda of both networks are trained jointly by minimizing the negative log posterior loss across image-caption pairs:

    L(Λ)=−∑(Q,D+)log⁡P(D+∣Q)\mathcal{L}(\Lambda) = -\sum_{(Q, D^+)} \log P(D^+ \mid Q)

  6. Knowl 6 — Sentence Re-Ranking using Minimum Error Rate Training (MERT)

    model/method

    To select the final caption from the MM-best candidate list M\mathcal{M} produced by beam search, candidate sentences are re-ranked using a linear combination of sentence-level features:

    Score(S,Q)=∑m=1Dwmhm(S,Q)\text{Score}(S, Q) = \sum_{m=1}^D w_m h_m(S, Q)

    where hm(S,Q)h_m(S, Q) are sentence features and wmw_m are weights optimized via Minimum Error Rate Training (MERT) directly on the validation set using the BLEU metric.

    The feature set comprises:

    1. The log-likelihood of the sequence under the Maximum Entropy LM.
    2. The word length of the sentence.
    3. The average log-probability per word (normalized log-likelihood).
    4. The logarithm of the sentence's rank in log-likelihood among the MM-best list.
    5. 11 binary indicator features indicating whether the number of mentioned visual objects is xx for each x∈{0,1,…,10}x \in \{0, 1, \dots, 10\}.
    6. The DMSM cosine similarity score R(Q,S)R(Q, S) between the image vector and the candidate sentence vector.
  7. Knowl 7 — Microsoft COCO Validation/Test Split Caption Generation Performance

    empirical result

    Performance of seven pipeline variants evaluated on a held-out test split of 20,444 images from the Microsoft COCO validation set using Perplexity (PPLX), BLEU-4, METEOR (against 4 references), and human preference studies on Amazon Mechanical Turk:

    System PPLX BLEU METEOR ≈\approx human >> human ≥\ge human
    Unconditioned LM 24.1 1.2% 6.8% – – –
    Shuffled Human – 1.7% 7.3% – – –
    Baseline (AlexNet) 20.9 16.9% 18.9% 9.9% (±\pm1.5%) 2.4% (±\pm0.8%) 12.3% (±\pm1.6%)
    Baseline + Score 20.2 20.1% 20.5% 16.9% (±\pm2.0%) 3.9% (±\pm1.0%) 20.8% (±\pm2.2%)
    Baseline + Score + DMSM 20.2 21.1% 20.7% 18.7% (±\pm2.1%) 4.6% (±\pm1.1%) 23.3% (±\pm2.3%)
    Baseline + Score + DMSM + ft 19.2 23.3% 22.2% – – –
    VGG + Score + ft 18.1 23.6% 22.8% – – –
    VGG + Score + DMSM + ft 18.1 25.7% 23.6% 26.2% (±\pm2.1%) 7.8% (±\pm1.3%) 34.0% (±\pm2.5%)
    Human-written captions – 19.3% 24.1% – – –

    Adding visual detector scores to the ME LM improves BLEU from 16.9% to 20.1%. Introducing DMSM re-ranking further raises BLEU to 21.1%, and CNN fine-tuning (ft) alongside VGG features increases BLEU to 25.7%, exceeding the human reference score of 19.3%. In human pairwise evaluation, the full model (VGG + Score + DMSM + ft) is judged equal to or better than human captions 34.0% of the time. Replacing MIL with whole-image classification under the same VGG + Score + DMSM + ft setting degraded performance to PPLX = 18.9, BLEU = 21.9%, and METEOR = 21.4%.

  8. Knowl 8 — Official Microsoft COCO Captioning Challenge Benchmark Evaluation

    empirical result

    Evaluation on the 40,775 unseen test images of the official Microsoft COCO Captioning Challenge benchmark server using 5 reference captions and 40 reference captions (human reference scores shown in parentheses):

    References CIDEr BLEU-4 BLEU-1 ROUGE-L METEOR
    5 references 0.912 (0.854) 0.291 (0.217) 0.695 (0.663) 0.519 (0.484) 0.247 (0.252)
    40 references 0.925 (0.910) 0.567 (0.471) 0.880 (0.880) 0.662 (0.626) 0.331 (0.335)

    The system achieved state-of-the-art results across all 14 official challenge metrics and matched or surpassed human performance on 12 of the 14 metrics (surpassing humans on CIDEr, BLEU-1 through BLEU-4, and ROUGE-L, while remaining close on METEOR). It was the only system evaluated to exceed human performance on the CIDEr consensus metric (0.912 vs 0.854 for 5 references, and 0.925 vs 0.910 for 40 references).

  9. Knowl 9 — Visual Word Detection Accuracy by Part of Speech

    empirical result

    Average Precision (AP) and Precision at Human Recall (PHR) across the 1000-word vocabulary grouped by part of speech, evaluated on the MS COCO dataset. Noisy-OR Multiple Instance Learning (MIL) on region features is compared against whole-image classification baselines and human agreement:

    Metric / Method NN VB JJ DT PRP IN Others All
    Word Count 616 176 119 10 11 38 30 1000
    Average Precision (AP)
    Chance 2.0 2.3 2.5 23.6 4.7 11.9 7.7 2.9
    Classification (AlexNet) 32.4 16.7 20.7 31.6 16.8 21.4 15.6 27.1
    Classification (VGG) 37.0 19.4 22.5 32.9 19.4 22.5 16.9 30.8
    MIL (AlexNet) 36.9 18.0 22.9 31.7 16.8 21.4 15.2 30.4
    MIL (VGG) 41.4 20.7 24.9 32.4 19.1 22.8 16.3 34.0
    Human Agreement 63.8 35.0 35.9 43.1 32.5 34.3 31.6 52.8
    Precision at Human Recall
    Classification (AlexNet) 39.0 27.7 37.0 37.3 26.2 31.5 25.0 35.9
    Classification (VGG) 45.3 31.0 37.1 40.2 29.6 33.9 25.5 40.6
    MIL (AlexNet) 46.0 29.4 40.1 37.9 25.9 31.5 21.6 40.8
    MIL (VGG) 51.6 33.3 44.3 39.2 29.4 34.3 23.9 45.7
    Human Agreement 63.8 35.0 35.9 43.1 32.5 34.3 31.6 52.8

    MIL consistently improves over whole-image classification across all parts of speech (overall AP improves from 30.8% to 34.0% with VGG, and overall PHR improves from 40.6% to 45.7%), with the largest improvements occurring on nouns (+4.4% AP) and adjectives (+2.4% AP) due to sub-region spatial localization.

  10. Knowl 10 — Performance Comparison on PASCAL Sentence Dataset

    empirical result

    On the PASCAL sentence benchmark (evaluated on the 847 test images used by prior systems), the weakly-supervised MIL word detector and ME LM generation pipeline achieves 17.6% BLEU and 19.2% METEOR. Under identical test images, the rule-based syntactic generation system Midge achieves 2.0% BLEU and 9.2% METEOR.

Coverage note — No substantial contributed material was omitted.

References

  1. 1.S. Banerjee and A. Lavie. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, 2005.
  2. 2.A. L. Berger, S. A. D. Pietra, and V. J. D. Pietra. A maximum entropy approach to natural language processing. Computational Linguistics, 1996.
  3. 3.A. Carlson, J. Betteridge, B. Kisiel, B. Settles, E. R. Hruschka Jr, and T. M. Mitchell. Toward an architecture for never-ending language learning. In AAAI, 2010.
  4. 4.X. Chen, H. Fang, T. Lin, R. Vedantam, S. Gupta, P. Dollár, and C. L. Zitnick. Microsoft coco captions: Data collection and evaluation server. arXiv preprint arXiv:1504.00325, 2015.
  5. 5.X. Chen, A. Shrivastava, and A. Gupta. Neil: Extracting visual knowledge from web data. In ICCV, 2013.
  6. 6.X. Chen and C. L. Zitnick. Mind’s eye: A recurrent visual representation for image caption generation. CVPR, 2015.
  7. 7.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. ImageNet: A large-scale hierarchical image database. In CVPR, 2009.
  8. 8.S. Divvala, A. Farhadi, and C. Guestrin. Learning everything about anything: Webly-supervised visual concept learning. In CVPR, 2014.
  9. 9.J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell. Long-term recurrent convolutional networks for visual recognition and description. CVPR, 2015.
  10. 10.D. Elliott and F. Keller. Comparing automatic evaluation measures for image description. In ACL, 2014.
  11. 11.M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman. The PASCAL visual object classes (VOC) challenge. IJCV, 88(2):303–338, June 2010.
  12. 12.A. Farhadi, M. Hejrati, M. A. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. Forsyth. Every picture tells a story: Generating sentences from images. In ECCV, 2010.
  13. 13.R. Girshick, J. Donahue, T. Darrell, and J. Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. In CVPR, 2014.
  14. 14.G. Gkioxari, B. Hariharan, R. Girshick, and J. Malik. Using k-poselets for detecting people and localizing their keypoints. In CVPR, 2014.
  15. 15.M. Hodosh, P. Young, and J. Hockenmaier. Framing image description as a ranking task: Data, models and evaluation metrics. JAIR, 47:853–899, 2013.
  16. 16.P. Huang, X. He, J. Gao, L. Deng, A. Acero, and L. Heck. Learning deep structured semantic models for web search using clickthrough data. In CIKM, 2013.
  17. 17.Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell. Caffe: Convolutional architecture for fast feature embedding. arXiv preprint arXiv:1408.5093, 2014.
  18. 18.A. Karpathy and L. Fei-Fei. Deep visual-semantic alignments for generating image descriptions. CVPR, 2015.
  19. 19.A. Karpathy, A. Joulin, and L. Fei-Fei. Deep fragment embeddings for bidirectional image sentence mapping. arXiv preprint arXiv:1406.5679, 2014.
  20. 20.R. Kiros, R. Zemel, and R. Salakhutdinov. Multimodal neural language models. In NIPS Deep Learning Workshop, 2013.
  21. 21.A. Krizhevsky, I. Sutskever, and G. E. Hinton. ImageNet classification with deep convolutional neural networks. In NIPS, 2012.
  22. 22.G. Kulkarni, V. Premraj, S. Dhar, S. Li, Y. Choi, A. C. Berg, and T. L. Berg. Baby talk: Understanding and generating simple image descriptions. In CVPR, 2011.
  23. 23.P. Kuznetsova, V. Ordonez, A. C. Berg, T. L. Berg, and Y. Choi. Collective generation of natural image descriptions. In ACL, 2012.
  24. 24.R. Lau, R. Rosenfeld, and S. Roukos. Trigger-based language models: A maximum entropy approach. In ICASSP, 1993.
  25. 25.R. Lebret, P. O. Pinheiro, and R. Collobert. Phrase-based image captioning. arXiv preprint arXiv:1502.03671, 2015.
  26. 26.S. Li, G. Kulkarni, T. L. Berg, A. C. Berg, and Y. Choi. Composing simple image descriptions using web-scale n-grams. In CoNLL, 2011.
  27. 27.C.-Y. Lin and F. J. Och. Automatic evaluation of machine translation quality using longest common subsequence and skip-bigram statistics. In Proceedings of the 42nd Annual Meeting on Association for Computational Linguistics, ACL ’04, Stroudsburg, PA, USA, 2004. Association for Computational Linguistics.
  28. 28.T. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick. Microsoft COCO: Common objects in context. In ECCV, 2014.
  29. 29.J. Mao, W. Xu, Y. Yang, J. Wang, and A. L. Yuille. Explain images with multimodal recurrent neural networks. arXiv preprint arXiv:1410.1090, 2014.
  30. 30.O. Maron and T. Lozano-Pérez. A framework for multiple-instance learning. NIPS, 1998.
  31. 31.T. Mikolov, A. Deoras, D. Povey, L. Burget, and J. Cernocky. Strategies for training large scale neural network language models. In ASRU, 2011.
  32. 32.M. Mitchell, X. Han, J. Dodge, A. Mensch, A. Goyal, A. Berg, K. Yamaguchi, T. Berg, K. Stratos, and H. Daumé III. Midge: Generating image descriptions from computer vision detections. In EACL, 2012.
  33. 33.A. Mnih and G. Hinton. Three new graphical models for statistical language modelling. In ICML, 2007.
  34. 34.A. Mnih and Y. W. Teh. A fast and simple algorithm for training neural probabilistic language models. In ICML, 2012.
  35. 35.F. J. Och. Minimum error rate training in statistical machine translation. In ACL, 2003.
  36. 36.V. Ordonez, G. Kulkarni, and T. L. Berg. Im2text: Describing images using 1 million captioned photographs. In NIPS, 2011.
  37. 37.K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu. Bleu: a method for automatic evaluation of machine translation. In ACL, 2002.
  38. 38.C. Rashtchian, P. Young, M. Hodosh, and J. Hockenmaier. Collecting image annotations using Amazon’s mechanical turk. In NAACL HLT Workshop Creating Speech and Language Data with Amazon’s Mechanical Turk, 2010.
  39. 39.A. Ratnaparkhi. Trainable methods for surface natural language generation. In NAACL, 2000.
  40. 40.A. Ratnaparkhi. Trainable approaches to surface natural language generation and their application to conversational dialog systems. Computer Speech & Language, 16(3):435–455, 2002.
  41. 41.Y. Shen, X. He, J. Gao, L. Deng, and G. Mesnil. A latent semantic model with convolutional-pooling structure for information retrieval. In CIKM, 2014.
  42. 42.K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. CoRR, abs/1409.1556, 2014.
  43. 43.R. Socher, Q. Le, C. Manning, and A. Ng. Grounded compositional semantics for finding and describing images with sentences. In NIPS Deep Learning Workshop, 2013.
  44. 44.R. Vedantam, C. L. Zitnick, and D. Parikh. Cider: Consensus-based image description evaluation. CoRR, abs/1411.5726, 2014.
  45. 45.O. Vinyals, A. Toshev, S. Bengio, and D. Erhan. Show and tell: A neural image caption generator. CVPR, 2015.
  46. 46.K. Xu, J. Ba, R. Kiros, A. Courville, R. Salakhutdinov, R. Zemel, and Y. Bengio. Show, attend and tell: Neural image caption generation with visual attention. arXiv preprint arXiv:1502.03044, 2015.
  47. 47.Y. Yang, C. L. Teo, H. Daumé III, and Y. Aloimonos. Corpus-guided sentence generation of natural images. In EMNLP, 2011.
  48. 48.B. Z. Yao, X. Yang, L. Lin, M. W. Lee, and S.-C. Zhu. I2T: Image parsing to text description. Proceedings of the IEEE, 98(8):1485–1508, 2010.
  49. 49.C. Zhang, J. C. Platt, and P. A. Viola. Multiple instance boosting for object detection. In NIPS, 2005.
  50. 50.C. L. Zitnick and P. Dollár. Edge boxes: Locating object proposals from edges. In ECCV, 2014.
  51. 51.C. L. Zitnick and D. Parikh. Bringing semantics into focus using visual abstraction. In CVPR, 2013.

Citation

MLA
Fang, H., et al. “From Captions to Visual Concepts and Back”. arXiv, 2014, http://arxiv.org/abs/1411.4952v3.
APA
Fang, H., Gupta, S., Iandola, F., Srivastava, R., Deng, L., Dollár, P., Gao, J., He, X., Mitchell, M., Platt, J. C., Zitnick, C. L., & Zweig, G. (2014). From Captions to Visual Concepts and Back. arXiv. http://arxiv.org/abs/1411.4952v3
Chicago
Fang, H., S. Gupta, F. Iandola, et al. 2014. “From Captions to Visual Concepts and Back”. arXiv. http://arxiv.org/abs/1411.4952v3.
Harvard
Fang, H. et al. (2014) “From Captions to Visual Concepts and Back”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1411.4952v3.
Vancouver
1. Fang H, Gupta S, Iandola F, et al (2014) From Captions to Visual Concepts and Back. arXiv

BibTeX

@article{fang2014from,
  title = {From Captions to Visual Concepts and Back},
  author = {Fang, Hao and Gupta, Saurabh and Iandola, Forrest and Srivastava, Rupesh and Deng, Li and Dollár, Piotr and Gao, Jianfeng and He, Xiaodong and Mitchell, Margaret and Platt, John C. and Zitnick, C. Lawrence and Zweig, Geoffrey},
  year = {2014},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1411.4952v3},
  eprint = {1411.4952}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE