Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models

Bryan A. PlummerLiwei WangChris M. CervantesJuan C. CaicedoJulia HockenmaierSvetlana Lazebnik

article2015IJCV2,703 citations

Introduces the Flickr30k Entities benchmark by grounding over 240,000 caption phrases to image bounding boxes, establishing a standard dataset and baseline for phrase localization and vision-language grounding.

Listen

The article addresses the lack of explicit links between phrases in image captions and specific regions in images, a gap that limits progress on grounded language understanding and compositional image description models. Existing benchmarks like Flickr30k and MSCOCO pair images with sentences but provide no region-to-phrase correspondences, forcing models to treat such mappings as latent or rely on unrelated detectors.

The work set out to create the first large-scale dataset supplying these correspondences and to demonstrate their value for new and existing tasks. Researchers augmented the Flickr30k collection of 31,783 images and 158,915 captions by adding 244,035 coreference chains across captions and 275,775 bounding boxes tied to entity mentions.

A multi-stage crowdsourcing pipeline on Mechanical Turk first resolved cross-caption coreference through binary link judgments followed by verification, then collected bounding boxes via requirement, drawing, quality, and coverage checks. Experiments trained a Canonical Correlation Analysis embedding on region-phrase pairs for text-to-image reference resolution and tested both training-time and test-time use of the correspondences for bidirectional image-sentence retrieval.

The resulting dataset contains an average of 7.7 coreference chains and 8.7 boxes per image. Localization performance reached 11.22 mean average precision and 25.3 percent Recall@1 on 100 proposals per image, with notable variation across entity types. Adding region-phrase data to whole-image training improved retrieval Recall@1 by roughly 23 points over a strong baseline, and test-time region-phrase matching yielded a further 2-point gain.

These findings indicate that explicit grounding supervision can measurably strengthen retrieval models while exposing the remaining difficulty of accurate phrase localization. The annotations matter because they enable direct training and evaluation of compositional models that must associate specific textual mentions with image locations rather than producing generic captions.

The dataset should be used to develop models that incorporate spatial constraints and object interactions, to benchmark cross-caption coreference, and to distinguish visual from non-visual text. Further gains will require larger-scale experiments, better region proposals for small or rare entities, and methods that handle the roughly 8 percent of images containing residual annotation errors.

Cover for Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models

Abstract

The Flickr30k dataset has become a standard benchmark for sentence-based image description. This paper presents Flickr30k Entities, which augments the 158k captions from Flickr30k with 244k coreference chains, linking mentions of the same entities across different captions for the same image, and associating them with 276k manually annotated bounding boxes. Such annotations are essential for continued progress in automatic image description and grounded language understanding. They enable us to define a new benchmark for localization of textual entity mentions in an image. We present a strong baseline for this task that combines an image-text embedding, detectors for common objects, a color classifier, and a bias towards selecting larger objects. While our baseline rivals in accuracy more complex state-of-the-art models, we show that its gains cannot be easily parlayed into improvements on such tasks as image-sentence retrieval, thus underlining the limitations of current methods and the need for further research.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Datasets with Region-Level Descriptions
  • 2.2 Grounded Language Understanding
  • 3 Annotation Process
  • 3.1 Coreference Resolution
  • 3.1.1 Binary Coreference Link Annotation
  • 3.1.2 Coreference Chain Verification
  • 3.2 Bounding Box Annotations
  • 3.2.1 Box Requirement
  • 3.2.2 Box Drawing
  • 3.2.3 Box Quality
  • 3.2.4 Box Coverage
  • 3.3 Quality Control
  • 3.3.1 Additional Review
  • 3.3.2 Box and Coreference Chain Merging
  • 3.3.3 Error Analysis
  • 3.4 Dataset Statistics
  • 4 Experimental Evaluation
  • 4.1 Phrase Localization
  • 4.1.1 Region-Phrase Model
  • 4.1.2 Evaluation Protocol
  • 4.1.3 Phrase Localization Experiments
  • 4.1.4 Phrase Localization Discussion
  • 4.2 Image-Sentence Retrieval
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — Flickr30k Entities Dataset Statistics and Structure

    data/table

    The Flickr30k Entities dataset augments the 31,783 images and 158,915 English captions (5 per image) of the Flickr30k corpus with cross-caption coreference chains and spatial bounding box groundings for mentioned entities. It contains 513,644 entity or scene mentions (averaging 3.2 mentions per caption and 16.6 mentions per image), clustered into 244,035 coreference chains (7.7 chains per image), and 275,775 manually annotated bounding boxes (8.7 boxes per image).

    Entity Type #Chains Mentions/Chain Boxes/Chain
    people 59,766 3.17 1.95
    clothing 42,380 1.76 1.44
    body parts 12,809 1.50 1.42
    animals 5,086 3.63 1.44
    vehicles 5,561 2.77 1.21
    instruments 1,827 2.85 1.61
    scene 46,919 2.03 0.62
    other 82,098 1.94 1.04
    total 244,035 2.10 1.13

    Aggregated across all five captions per image, people appear in 94.2% of images, scene mentions in 79.7%, clothing in 69.9%, body parts in 28.0%, vehicles in 13.8%, animals in 12.0%, instruments in 4.3%, and other objects in 91.8%. Furthermore, 48.6% of coreference chains contain multiple mentions across different captions. In terms of spatial grounding, 59.1% of chains link to a single bounding box, 20.0% link to multiple bounding boxes (frequent in collective mentions such as groups or families), and 20.9% link to no bounding box (non-visual entities, abstract concepts, or global scenes).

  2. Knowl 2 — Crowdsourcing Pipeline for Coreference Resolution and Grounding

    model/method

    The annotation process decomposes the collection of region-to-phrase correspondences into two sequential stages:

    1. Cross-Caption Coreference Resolution:

      • Mentions are candidate noun-phrase (NP) chunks (averaging 2.35 words), excluding personal pronouns (he, she, they) and non-visual terms (background, air).
      • Pairwise link querying is reduced from O(M2)\mathcal{O}(|M|^2) to O(MC)\mathcal{O}(|M||C|) (where MM is the set of mentions in the 5 captions and CC is the set of identified chains) by exploiting transitivity: a new mention mm is only queried against one representative mention from each existing chain cCc \in C.
      • Comparisons are pruned with two constraints: mentions from the same caption cannot corefer, and mentions classified into different coarse dictionary types (people, body parts, animals, clothing/color, instruments, vehicles, scene, other) cannot corefer.
      • Multi-mention chains undergo a verification step where workers confirm whether all phrases refer to the same entities; if rejected, the chain is split into subsets sharing the same head noun.
    2. Bounding Box Annotation Workflow:

      • Box Requirement: Workers determine whether a representative mention requires 1\ge 1 bounding box, describes a scene/place, or is non-visual. Mentions in people, clothing, and body parts categories bypass this task directly to drawing.
      • Box Drawing: Workers annotate individual bounding boxes for each distinguishable entity, or a single enclosing box for indistinguishable groups (e.g., crowds).
      • Box Quality: Drawn boxes are validated for tight spatial coverage without redundancy.
      • Box Coverage: Workers verify whether all referenced entities have been boxed; if incomplete, the item returns to drawing.
    3. Post-Processing Merging:

      • Fragmented boxes with an Intersection-over-Union (IoU\text{IoU}) 0.8\ge 0.8 (or 0.9\ge 0.9 for 'other') are merged, excluding incompatible category pairs (e.g., clothing and people). Coreference chains that point to the exact same set of boxes are subsequently unified.
  3. Knowl 3 — Text-to-Image Reference Resolution Baseline via Normalized CCA

    model/method

    The baseline model for grounding entity phrases to image regions maps visual regions and text phrases into a shared latent semantic space using Normalized Canonical Correlation Analysis (CCA).

    • Visual Features: Candidate image regions are generated using the top 100 proposals from EdgeBoxes. Each region is encoded as a 4096-dimensional activation vector extracted from the 19-layer VGG network computed from a single crop.
    • Text Features: Each word in an NP chunk is represented by a 300-dimensional word2vec embedding. A Hybrid Gaussian-Laplacian Mixture Model (HGLMM) Fisher Vector codebook with 30 mixture centers captures first- and second-order statistics, producing an 18,000-dimensional phrase vector (300×30×2300 \times 30 \times 2).
    • Normalized CCA Embedding: CCA projection matrices are learned between region features and phrase features. The projection matrices are scaled by their corresponding canonical correlation eigenvalues, and projected vectors are normalized to unit 2\ell_2 length. Region candidates are ranked for a given query phrase by cosine distance in this shared space.
    • Phrase Frequency Resampling: To counteract class imbalance where frequent phrases (such as 'a man') dominate rare phrases, the training set of 423,134 ground-truth NP chunks is resampled by capping each unique phrase to at most NN exemplars. Setting N=10N=10 (137,133 training pairs) yields optimal localization performance compared to N=1N=1 (70,759 pairs) or using all chunks without resampling.
  4. Knowl 4 — Phrase Localization Performance on Flickr30k Entities

    data/table

    Text-to-image reference resolution evaluates the ability to localize phrases extracted from image captions within candidate region proposals (the top 100 EdgeBoxes per image). A predicted region is deemed correct if its Intersection over Union with the ground-truth entity bounding box is IoU0.5\text{IoU} \ge 0.5. Performance is measured by Recall@KK (percentage of queries where a correct box is ranked K\le K, with K=100K=100 serving as the proposal upper bound) and Mean Average Precision after Non-Maximum Suppression (AP-NMS).

    Evaluated on 1,000 test images (14,558 test phrase instances) using the resampled (N=10N=10) CCA model:

    Metric people clothing bodyparts animals vehicles instruments scene other mAP / overall
    #Instances 5656 2306 523 518 400 162 1619 3374 14558
    AP-NMS 13.16 11.48 4.85 13.84 13.67 11.42 10.86 10.50 11.22
    R@1 (%) 29.58 24.20 10.52 33.40 34.75 35.80 20.20 20.75 25.30
    R@10 (%) 71.25 52.99 29.83 66.99 76.75 61.11 57.44 47.24 59.66
    R@100 (%) 89.36 66.48 39.39 84.56 91.00 69.75 75.05 67.40 76.91

    The model attains higher localization recall on distinct, prominent objects such as vehicles (R@100 = 91.00%) and people (R@100 = 89.36%), while smaller or part-level categories like body parts (R@100 = 39.39%) suffer from a lack of high-quality proposals and localized discriminability.

  5. Knowl 5 — Weighted Region-to-Phrase Distance for Image-Sentence Retrieval

    equation

    To incorporate fine-grained phrase-to-region localization into bidirectional image-sentence retrieval at test time, the distance between a candidate sentence (consisting of LL noun phrases) and an image (represented as a set of candidate regions) is defined as:

    DRP=1Lγi=1Lpir(pi)22D_{RP} = \frac{1}{L^\gamma} \sum_{i=1}^{L} \|p_i - r(p_i)\|_2^2

    where pip_i is the feature representation of the ii-th phrase in the sentence, r(pi)r(p_i) is the feature of the highest-ranked matching image region for phrase pip_i in the shared CCA space, and γ1\gamma \ge 1 is a length penalty attenuation exponent. Setting γ=1.5\gamma = 1.5 compensates for the fact that detailed sentences with more phrases accumulate higher cumulative distances.

    The overall matching distance D^IS\hat{D}_{IS} between the image and the sentence is a linear combination:

    D^IS=αDIS+(1α)DRP\hat{D}_{IS} = \alpha D_{IS} + (1 - \alpha) D_{RP}

    where DISD_{IS} is the squared Euclidean distance between the global image feature and global sentence feature in the whole-image-sentence CCA space, and α[0,1]\alpha \in [0, 1] is a weighting parameter set to α=0.7\alpha = 0.7.

  6. Knowl 6 — Bidirectional Image-Sentence Retrieval Performance with Region Grounding

    data/table

    Bidirectional image-sentence retrieval performance on the 1,000 test images and 5,000 test sentences of Flickr30k. Image Annotation refers to using an image to retrieve candidate sentences; Image Search refers to using a sentence to retrieve candidate images. Performance is evaluated using Recall@KK (K{1,5,10}K \in \{1, 5, 10\}), the percentage of queries where at least one correct ground-truth match is ranked in the top KK.

    Image Annotation Image Search
    Method R@1 (%) R@5 (%) R@10 (%) R@1 (%) R@5 (%) R@10 (%)
    BRNN 22.2 48.2 61.4 15.2 37.7 50.5
    MNLM 23.0 50.7 62.9 16.8 42.0 56.5
    m-RNN 35.4 63.8 73.7 22.8 50.7 63.1
    GMM+HGLMM FV 33.3 62.0 74.7 25.6 53.2 66.8
    Whole image-sentence CCA (HGLMM FV) 36.5 62.2 73.3 24.7 53.4 66.8
    SAE (N=10 CCA auxiliary) 38.3 63.7 75.5 25.4 54.2 67.4
    Weighted Distance (α=0.7,γ=1.5\alpha=0.7, \gamma=1.5) 39.1 64.8 76.4 26.9 56.2 69.7

    Incorporating region-phrase supervision during training via Stacked Auxiliary Embeddings (SAE) improves R@1 by 1.8% on image annotation and 0.7% on image search over the whole image-sentence CCA baseline. Combining global distances with test-time region-to-phrase correspondence distances (Weighted Distance) achieves the best overall performance, reaching 39.1% R@1 on image annotation and 26.9% R@1 on image search.

  7. Knowl 7 — Stacked Auxiliary Embedding for Grounded Image-Sentence Retrieval

    model/method

    Directly merging whole-image/sentence pairs and region/phrase pairs into a single training set is problematic because whole-image features (averaged over 10 crops) have different statistical density and scale compared to single-crop region features, and whole sentences have different vocabulary and length distributions than noun phrases.

    To leverage fine-grained region-phrase correspondences without data mixing artifacts, the Stacked Auxiliary Embedding (SAE) framework is applied:

    1. An auxiliary CCA model is trained strictly on region-phrase pairs (with N=10N=10 resampling per phrase).
    2. For global image-sentence retrieval, auxiliary feature projections generated by this region-phrase model are concatenated with the base global image and sentence feature vectors.
    3. A primary CCA embedding is trained on these stacked representations to align full images with complete captions, thereby transferring granular object-level semantic associations to global retrieval.
  8. Knowl 8 — Limitations of Proposal-Based Reference Resolution and CCA Grounding

    limitation

    The baseline approach for phrase localization exhibits several fundamental limitations:

    1. Retrieval vs. Localization Mismatch: CCA optimizes global projection alignment rather than tight bounding box regression. High similarity scores can be assigned to loose candidate boxes containing context rather than precise boundaries enclosing the entire entity.
    2. Missing Proposals for Small Entities: Small objects (e.g., hats, balls) and body parts frequently fail to receive candidate proposals with IoU0.5\text{IoU} \ge 0.5 from bottom-up edge-based proposal generators (EdgeBoxes), leading to low proposal recall ceilings (e.g., 39.39% for body parts).
    3. Absence of Spatial and Relational Reasoning: The model ranks candidate regions independently for each phrase, ignoring spatial configuration and syntactic relations. For instance, when distinguishing multiple interacting entities of the same type (e.g., determining which child is on a swing vs. standing nearby), independent CCA matching fails to resolve referential ambiguity without explicit spatial and interaction modeling.

Coverage note — Table 3 (localization results for the 24 most common individual phrases) was omitted in favor of the complete category-level and overall benchmark metrics in Table 2 (Knowl 4).

References

  1. 1.X. Chen and C. L. Zitnick. Learning a recurrent visual representation for image caption generation. arXiv:1411.5654, 2014.
  2. 2.J. Devlin, H. Cheng, H. Fang, S. Gupta, L. Deng, X. He, G. Zweig, and M. Mitchell. Language models for image captioning: The quirks and what works. In arxiv.org:1505.01809, 2015.
  3. 3.J. Dodge, A. Goyal, X. Han, A. Mensch, M. Mitchell, K. Stratos, K. Yamaguchi, Y. Choi, H. D. III, A. C. Berg, and T. L. Berg. Detecting visual text. In NAACL, 2012.
  4. 4.J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell. Long-term recurrent convolutional networks for visual recognition and description. arXiv:1411.4389, 2014.
  5. 5.M. Everingham, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman. The Pascal visual object classes (VOC) challenge. IJCV, 88(2):303–338, 2010.
  6. 6.H. Fang, S. Gupta, F. Iandola, R. Srivastava, L. Deng, P. Dollár, J. Gao, X. He, M. Mitchell, J. Platt, et al. From captions to visual concepts and back. arXiv:1411.4952, 2014.
  7. 7.A. Farhadi, S. Hejrati, A. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. A. Forsyth. Every picture tells a story: Generating sentences from images. In ECCV. 2010.
  8. 8.S. Fidler, A. Sharma, and R. Urtasun. A sentence is worth a thousand pixels. In CVPR, 2013.
  9. 9.Y. Gong, Q. Ke, M. Isard, and S. Lazebnik. A multi-view embedding space for modeling internet images, tags, and their semantics. IJCV, 106(2):210–233, 2014.
  10. 10.Y. Gong, L. Wang, M. Hodosh, J. Hockenmaier, and S. Lazebnik. Improving image-sentence embeddings using large weakly annotated photo collections. In ECCV, 2014.
  11. 11.M. Grubinger, P. Clough, H. Müller, and T. Deselaers. The iapr tc-12 benchmark: A new evaluation resource for visual information systems. In International Workshop OntoImage, pages 13–23, 2006.
  12. 12.M. Hodosh, P. Young, and J. Hockenmaier. Framing image description as a ranking task: Data, models and evaluation metrics. JAIR, 2013.
  13. 13.M. Hodosh, P. Young, C. Rashtchian, and J. Hockenmaier. Crosscaption coreference resolution for automatic image understanding. In CoNLL, pages 162–171. ACL, 2010.
  14. 14.H. Hotelling. Relations between two sets of variates. Biometrika, pages 321–377, 1936.
  15. 15.J. Johnson, R. Krishna, M. Stark, L.-J. Li, D. A. Shamma, M. Bernstein, and L. Fei-Fei. Image retrieval using scene graphs. In CVPR, 2015.
  16. 16.A. Karpathy and L. Fei-Fei. Deep visual-semantic alignments for generating image descriptions. arXiv:1412.2306, 2014.
  17. 17.A. Karpathy, A. Joulin, and L. Fei-Fei. Deep fragment embeddings for bidirectional image sentence mapping. In NIPS, 2014.
  18. 18.S. Kazemzadeh, V. Ordonez, M. Matten, and T. Berg. Referitgame: Referring to objects in photographs of natural scenes. In EMNLP, 2014.
  19. 19.R. Kiros, R. Salakhutdinov, and R. S. Zemel. Unifying visualsemantic embeddings with multimodal neural language models. arXiv:1411.2539, 2014.
  20. 20.B. Klein, G. Lev, G. Sadeh, and L. Wolf. Fisher vectors derived from hybrid gaussian-laplacian mixture models for image annotation. CVPR, 2015.
  21. 21.C. Kong, D. Lin, M. Bansal, R. Urtasun, and S. Fidler. What are you talking about? text-to-image coreference. In CVPR, 2014.
  22. 22.G. Kulkarni, V. Premraj, S. Dhar, S. Li, Y. Choi, A. C. Berg, and T. L. Berg. Baby talk: Understanding and generating image descriptions. In CVPR, 2011.
  23. 23.R. Lebret, P. O. Pinheiro, and R. Collobert. Phrase-based image captioning. arXiv:1502.03671, 2015.
  24. 24.T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick. Microsoft COCO: Common objects in context. In ECCV, 2014.
  25. 25.J. Mao, W. Xu, Y. Yang, J. Wang, and A. Yuille. Deep captioning with multimodal recurrent neural networks (m-rnn). arXiv:1412.6632, 2014.
  26. 26.J. F. McCarthy and W. G. Lehnert. Using decision trees for coreference resolution. arXiv cmp-lg/9505043, 1995.
  27. 27.T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean. Distributed representations of words and phrases and their compositionality. In NIPS, 2013.
  28. 28.V. Ordonez, G. Kulkarni, and T. L. Berg. Im2Text: Describing images using 1 million captioned photographs. NIPS, 2011.
  29. 29.F. Perronnin, J. Sánchez, and T. Mensink. Improving the fisher kernel for large-scale image classification. In ECCV, 2010.
  30. 30.V. Ramanathan, A. Joulin, P. Liang, and L. Fei-Fei. Linking people in videos with "their" names using coreference resolution. In ECCV, 2014.
  31. 31.C. Rashtchian, P. Young, M. Hodosh, and J. Hockenmaier. Collecting image annotations using Amazon's mechanical turk. In NAACL HLT Workshop on Creating Speech and Language Data with Amazon's Mechanical Turk, pages 139–147. ACL, 2010.
  32. 32.K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv:1409.1556, 2014.
  33. 33.R. Socher, J. Bauer, C. D. Manning, and A. Y. Ng. Parsing With Compositional Vector Grammars. In ACL, 2013.
  34. 34.W. M. Soon, H. T. Ng, and D. C. Y. Lim. A machine learning approach to coreference resolution of noun phrases. Computational Linguistics, 27(4):521–544, 2001.
  35. 35.A. Sorokin and D. Forsyth. Utility data annotation with Amazon Mechanical Turk. Internet Vision Workshop, 2008.
  36. 36.H. Su, J. Deng, and L. Fei-Fei. Crowdsourcing annotations for visual object detection. In AAAI Technical Report, 4th Human Computation Workshop, 2012.
  37. 37.O. Vinyals, A. Toshev, S. Bengio, and D. Erhan. Show and tell: A neural image caption generator. arXiv:1411.4555, 2014.
  38. 38.K. Xu, J. Ba, R. Kiros, A. Courville, R. Salakhutdinov, R. Zemel, and Y. Bengio. Show, attend and tell: Neural image caption generation with visual attention. arXiv:1502.03044, 2015.
  39. 39.B. Yao, X. Yang, L. Lin, M. W. Lee, and S.-C. Zhu. I2T: Image parsing to text description. Proc. IEEE, 98(8):1485 – 1508, 2010.
  40. 40.P. Young, A. Lai, M. Hodosh, and J. Hockenmaier. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. TACL, 2:67–78, 2014.
  41. 41.C. L. Zitnick and P. Dollár. Edge boxes: Locating object proposals from edges. In ECCV, 2014.

Citation

MLA
Plummer, B. A., et al. “Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models”. International Journal of Computer Vision, vol. 123, no. 1, 2016, pp. 74–93, https://doi.org/10.1007/s11263-016-0965-7.
APA
Plummer, B. A., Wang, L., Cervantes, C. M., Caicedo, J. C., Hockenmaier, J., & Lazebnik, S. (2016). Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models. International Journal of Computer Vision, 123(1), 74–93. https://doi.org/10.1007/s11263-016-0965-7
Chicago
Plummer, B. A., L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik. 2016. “Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models”. International Journal of Computer Vision 123 (1): 74–93. https://doi.org/10.1007/s11263-016-0965-7.
Harvard
Plummer, B.A. et al. (2016) “Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models”, International Journal of Computer Vision, 123(1), pp. 74–93. Available at: https://doi.org/10.1007/s11263-016-0965-7.
Vancouver
1. Plummer BA, Wang L, Cervantes CM, Caicedo JC, Hockenmaier J, Lazebnik S (2016) Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models. International Journal of Computer Vision 123:74–93

BibTeX

@article{Plummer_2016, title={Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models}, volume={123}, ISSN={1573-1405}, url={http://dx.doi.org/10.1007/s11263-016-0965-7}, DOI={10.1007/s11263-016-0965-7}, number={1}, journal={International Journal of Computer Vision}, publisher={Springer Science and Business Media LLC}, author={Plummer, Bryan A. and Wang, Liwei and Cervantes, Chris M. and Caicedo, Juan C. and Hockenmaier, Julia and Lazebnik, Svetlana}, year={2016}, month=Oct, pages={74–93} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF