keyword
Image descriptions
Image descriptions are natural language texts that explain or summarize the visual content, objects, attributes, actions, and spatial relationships depicted within an image. In computer vision, natural language processing, and multimodal artificial intelligence, these descriptions range from comprehensive scene captions that summarize entire visual events to localized referring expressions that uniquely identify individual entities or regions. They serve as a fundamental bridge between visual perception and linguistic meaning, enabling automated systems to perform tasks such as image captioning, cross-modal retrieval, visual question answering, accessibility support, and semantic inference.
4 items

From captions to visual concepts and back
Hao Fang, Saurabh Gupta, F. Iandola, R. Srivastava, L. Deng, Piotr Dollár, Jianfeng Gao, Xiaodong He, Margaret Mitchell, John C. Platt, C. L. Zitnick, G. Zweig
Why you should read this
Presents an image captioning pipeline that learns visual concept detectors directly from weakly-supervised captions via multiple instance learning, generates candidate descriptions with a maximum-entropy language model, and selects the best description using a deep multimodal similarity model.
This paper presents a novel approach for automatically generating image descriptions: visual detectors, language models, and multimodal similarity models learnt directly from a dataset of image captions. We use multiple instance learning to train visual detectors for words that commonly occur in captions, including many different parts of speech such as nouns, verbs, and adjectives. The word detector outputs serve as conditional inputs to a maximum-entropy language model. The language model learns from a set of over 400,000 image descriptions to capture the statistics of word usage. We capture global semantics by re-ranking caption candidates using sentence-level features and a deep multimodal similarity model. Our system is state-of-the-art on the official Microsoft COCO benchmark, producing a BLEU-4 score of 29.1%. When human judges compare the system captions to ones written by other people on our held-out test set, the system captions have equal or better quality 34% of the time.
Added
2026-09-25

Generation and Comprehension of Unambiguous Object Descriptions
Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan Yuille, Kevin Murphy
Why you should read this
Presents a unified deep learning framework for generating and comprehending unambiguous referring expressions in images by accounting for visual context, accompanied by a large-scale MS-COCO benchmark dataset.
We propose a method that can generate an unambiguous description (known as a referring expression) of a specific object or region in an image, and which can also comprehend or interpret such an expression to infer which object is being described. We show that our method outperforms previous methods that generate descriptions of objects without taking into account other potentially ambiguous objects in the scene. Our model is inspired by recent successes of deep learning methods for image captioning, but while image captioning is difficult to evaluate, our task allows for easy objective evaluation. We also present a new large-scale dataset for referring expressions, based on MS-COCO. We have released the dataset and a toolbox for visualization and evaluation, see this https URL
Added
2026-09-20

ReferItGame: Referring to Objects in Photographs of Natural Scenes
Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, Tamara L. Berg
Why you should read this
Presents a crowdsourced two-player game that simultaneously collects and verifies a large-scale real-world dataset of natural language referring expressions for objects in photographs, accompanied by an optimization-based generation model.
In this paper we introduce a new game to crowd-source natural language referring expressions. By designing a two player game, we can both collect and verify referring expressions directly within the game. To date, the game has produced a dataset containing 130,525 expressions, referring to 96,654 distinct objects, in 19,894 photographs of natural scenes. This dataset is larger and more varied than previous REG datasets and allows us to study referring expressions in real-world scenes. We provide an in depth analysis of the resulting dataset. Based on our findings, we design a new optimization based model for generating referring expressions and perform experimental evaluations on 3 test sets.
Added
2026-09-18

From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, J. Hockenmaier
Why you should read this
Proposes a method for grounding semantic inference in visual denotation graphs constructed from paired image captions, yielding novel denotational similarity metrics that outperform standard distributional models in recognizing textual entailment and semantic textual similarity.
We propose to use the visual denotations of linguistic expressions (i.e. the set of images they describe) to define novel denotational similarity metrics, which we show to be at least as beneficial as distributional similarities for two tasks that require semantic inference. To compute these denotational similarities, we construct a denotation graph, i.e. a subsumption hierarchy over constituents and their denotations, based on a large corpus of 30K images and 150K descriptive captions.
Added
2026-09-11
