Built independently by an author, for readers. Read the story and support ChapterPal

keyword

Image descriptions

Image descriptions are natural language texts that explain or summarize the visual content, objects, attributes, actions, and spatial relationships depicted within an image. In computer vision, natural language processing, and multimodal artificial intelligence, these descriptions range from comprehensive scene captions that summarize entire visual events to localized referring expressions that uniquely identify individual entities or regions. They serve as a fundamental bridge between visual perception and linguistic meaning, enabling automated systems to perform tasks such as image captioning, cross-modal retrieval, visual question answering, accessibility support, and semantic inference.

4 items

From captions to visual concepts and back

From captions to visual concepts and back

Hao Fang, Saurabh Gupta, F. Iandola, R. Srivastava, L. Deng, Piotr Dollár, Jianfeng Gao, Xiaodong He, Margaret Mitchell, John C. Platt, C. L. Zitnick, G. Zweig

OrganizationsDalle Molle Institute for Artificial Intelligence ResearchGoogleMetaMicrosoftUniversity of California BerkeleyUniversity of Washington

Why you should read this

Presents an image captioning pipeline that learns visual concept detectors directly from weakly-supervised captions via multiple instance learning, generates candidate descriptions with a maximum-entropy language model, and selects the best description using a deep multimodal similarity model.

This paper presents a novel approach for automatically generating image descriptions: visual detectors, language models, and multimodal similarity models learnt directly from a dataset of image captions. We use multiple instance learning to train visual detectors for words that commonly occur in captions, including many different parts of speech such as nouns, verbs, and adjectives. The word detector outputs serve as conditional inputs to a maximum-entropy language model. The language model learns from a set of over 400,000 image descriptions to capture the statistics of word usage. We capture global semantics by re-ranking caption candidates using sentence-level features and a deep multimodal similarity model. Our system is state-of-the-art on the official Microsoft COCO benchmark, producing a BLEU-4 score of 29.1%. When human judges compare the system captions to ones written by other people on our held-out test set, the system captions have equal or better quality 34% of the time.

Added

2026-09-25