Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning

Piyush SharmaNan DingSebastian GoodmanRadu Soricut

article2018ACL3,070 citations

Introduces Conceptual Captions, a 3.3-million-example dataset harvested and hypernymed from web alt-text, and demonstrates that training vision-language models on this large-scale data substantially reduces object hallucinations and improves open-domain image description quality.

Listen

A new dataset called Conceptual Captions addresses the limited size and variety of existing resources for training automatic image captioning systems. Current collections such as MS-COCO contain roughly 120,000 images with human-written captions and favor a narrow range of everyday scenes, which restricts model generalization and encourages hallucinations of objects that are statistically common in the training data rather than present in a given image. The work extracts candidate image-description pairs at web scale from alt-text attributes, applies successive filters for image quality, text well-formedness, and visual-textual overlap, then replaces specific names and details with appropriate hypernyms to produce clean, learnable captions.

The authors constructed a pipeline that processed billions of webpages and retained approximately 3.3 million image-caption pairs after rigorous filtering and transformation steps. They trained both RNN-based and Transformer-based captioning models on this collection, using an Inception-ResNet-v2 network to extract image features, and compared performance against identical models trained on COCO data. Evaluation covered in-domain and out-of-domain test sets, including the Flickr 1K set, with both automatic metrics and human ratings from professional annotators.

Human evaluations on the out-of-domain Flickr test set showed that Transformer models trained on Conceptual Captions received majority “good” ratings in 50.6 percent of cases, compared with 36.2 percent for the same architecture trained on COCO. These models also produced fewer hallucinations, handled cartoons and product images without collapse, and generated more precise terms such as “graduates” or “cloister of the cathedral.” Automatic metrics confirmed strong in-domain results for Conceptual Captions models but diverged from human judgments on out-of-domain data, underscoring known weaknesses in current evaluation measures.

The findings indicate that larger, web-harvested datasets with controlled generalization can materially improve caption quality and robustness without requiring manual annotation. Practical systems for image retrieval or accessibility tools would therefore benefit from training on Conceptual Captions or similar resources. The authors recommend further exploration of Transformer architectures, which deliver higher accuracy with lower computational cost than RNN alternatives.

The main limitations are the low overall yield of the filtering pipeline (0.2 percent of candidates) and the reliance on automatic metrics that under-penalize hallucinations. Readers should treat reported automatic scores with caution when comparing across domains.

Cover for Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning

Abstract

We present a new dataset of image caption annotations, Conceptual Captions, which contains an order of magnitude more images than the MS-COCO dataset (Lin et al., 2014) and represents a wider variety of both images and image caption styles. We achieve this by extracting and filtering image caption annotations from billions of webpages. We also present quantitative evaluations of a number of image captioning models and show that a model architecture based on Inception-ResNet-v2 (Szegedy et al., 2016) for image-feature extraction and Transformer (Vaswani et al., 2017) for sequence modeling achieves the best performance when trained on the Conceptual Captions dataset.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Conceptual Captions Dataset Creation
  • 4 Image Captioning Models
  • 4.1 RNN-based Models
  • 4.2 Transformer Model
  • 5 Experimental Results
  • 5.1 Dataset Details
  • 5.2 Experimental Setup
  • 5.3 Qualitative Results
  • 5.4 Quantitative Results
  • 5.4.1 Human Evaluation Results
  • 5.4.2 Automatic Evaluation Results
  • 6 Conclusions
  • References

Knowls

  1. Knowl 1 — Multi-Stage Web-Scale Extraction and Filtering Pipeline for Conceptual Captions

    algorithm

    The Conceptual Captions dataset is constructed from billions of raw web pages via a distributed Flume processing pipeline that extracts and filters candidate image-Alt-text pairs:

    1. Image-Based Filtering: Discards images based on encoding, dimensions, aspect ratio, and offensive content detectors. It retains only JPEG images where both width and height exceed 400 pixels, the aspect ratio (larger dimension divided by smaller dimension) is at most 2.0, and profanity/pornography detectors are not triggered. This stage discards more than 65% of raw candidates.

    2. Text-Based Filtering: Extracts the HTML alt attribute and analyzes it using Google Cloud Natural Language APIs (part-of-speech tagging, sentiment polarity, profanity detectors):

    • Syntactic requirements: discards candidates lacking a determiner, a noun, or a preposition, as well as candidates with an excessively high noun ratio or high token repetition rates.
    • Well-formedness heuristics: enforces sentence capitalization (first word capitalized, bounded ratio of capitalized words).
    • Lexical vocabulary filter: verifies candidate tokens against a vocabulary VWV_W of 10910^9 token types that occur at least 5 times in English Wikipedia; any candidate with an out-of-vocabulary token is rejected.
    • Sentiment and profanity filtering: rejects candidates with extreme sentiment polarity or profanity flags.
    • Boilerplate removal: strips common boilerplate prefixes/suffixes (e.g., "click to enlarge picture", "stock photo") and patterns (e.g., "profile photo", "embedded image permalink"). Only approximately 3% of candidates survive text filtering.
    1. Image-Text Cross-Filtering: Predicts visual labels for each image using Google Cloud Vision APIs from an ontology of ≈105\approx 10^5 classes (all covered by VWV_W), assigning 5 to 20 labels per image. Candidate text tokens and vision labels are matched using morphological stemming. If there is zero overlap between text tokens and predicted image labels, the pair is discarded (eliminating ≈60%\approx 60\% of remaining candidates).

    Across the pipeline, only approximately 0.2% of candidate pairs from over 5 billion images pass all criteria.

  2. Knowl 2 — Text Transformation and Hypernymization Pipeline for Image Captions

    algorithm

    Raw web Alt-text frequently contains specific proper names, dates, and locations that are difficult for vision models to predict directly from image pixels. The hypernymization and text transformation pipeline converts raw Alt-text into learnable, generalized conceptual descriptions:

    1. Syntactic and Modifier Pruning:
    • Identifies named entities and syntactic dependencies using Google Cloud Natural Language APIs.
    • Strips noun modifiers of specific types, including proper nouns, numerical counts (e.g., "Two"), and measurement units.
    • Removes temporal expressions (dates, durations) and prepositional location phrases (e.g., "in Los Angeles").
    1. Knowledge Graph Hypernym Substitution:
    • Queries the Google Knowledge Graph (KG) Search API with extracted named entities.
    • Replaces entity surface tokens with their KG hypernym category (for example, replacing "Harrison Ford" and "Calista Flockhart" with "actor", or "British Airways Airbus A319 aircraft" with "aircraft").
    1. Syntactic Repair and Sentence Filtering:
    • Resolves coordinate noun phrases that share the same hypernym head into a single pluralized form (e.g., "actor and actor" →\to "actors").
    • Discards any sentences that become too short, fragmented, or syntactically inconsistent (discarding ≈20%\approx 20\% of transformed candidates).
    1. Low-Count Concept Pruning:
    • Runs entity resolution across the whole dataset and clusters all resolved entity types.
    • Discards candidate pairs unless every detected entity concept belongs to a class with dataset frequency >100> 100 (discarding ≈55%\approx 55\% of candidates and retaining ≈16,000\approx 16,000 well-represented entity types).
  3. Knowl 3 — Conceptual Captions Dataset Statistics and Human Quality Benchmark

    data/table

    The Conceptual Captions dataset comprises approximately 3.3 million image-caption pairs derived from web Alt-text.

    Split Examples Unique Tokens Tokens / Caption
    Mean StdDev Median
    Train 3,318,333 51,201 10.3 4.5 9.0
    Validation 28,355 13,063 10.3 4.6 9.0
    Test 22,530 11,731 10.1 4.5 9.0

    The test set contains 22,530 examples that were verified by human raters (retaining pairs receiving at least 2 out of 3 GOOD annotations), while the training and validation splits were generated fully automatically. Caption lengths remain consistent across splits with a mean of 10.1–10.3 tokens and a median of 9.0 tokens.

    Human evaluation on a random sample of 4,000 test examples under double-blind conditions with 3 raters per image-caption pair yielded the following quality distribution:

    Dataset Sample GOOD judgments (out of 3)
    1+ 2+ 3
    Conceptual Captions 96.9% 90.3% 78.5%

    Over 90% of the automatically produced captions received a majority (2+2+) of positive ratings, demonstrating the cleanliness and precision of the automated pipeline.

  4. Knowl 4 — Transformer Architecture for Image Captioning

    model/method

    The Transformer-based image captioning architecture adapts multi-head self-attention and cross-attention networks to generate natural language descriptions from visual features.

    1. Visual Representation: An image is preprocessed by an Inception-ResNet-v2 CNN, producing a sequence of spatial embeddings X=(x1,x2,…,xL)∈RL×dX = (x_1, x_2, \dots, x_L) \in \mathbb{R}^{L \times d} with feature dimension d=512d = 512, evaluated as either a single global embedding (L=1L=1) or an 8×88 \times 8 grid of spatial partition embeddings (L=64L=64). In the 8×88 \times 8 setting, 1D sinusoidal positional encodings across positions j∈{0,…,63}j \in \{0, \dots, 63\} are added to the visual embeddings.

    2. Encoder: The encoder contains a stack of N=6N=6 identical layers. For layer n∈{0,…,N−1}n \in \{0, \dots, N-1\} with layer input Xn={xn,1,…,xn,L}X_n = \{x_{n,1}, \dots, x_{n,L}\} where X0=XX_0 = X and encoder output H=XNH = X_N: xn,j′=ATTN(xn,j,Xn;Wqe,Wke,Wve)=softmax(⟨xn,jWqe,XnWke⟩)XnWvex'_{n,j} = \text{ATTN}(x_{n,j}, X_n; W_q^e, W_k^e, W_v^e) = \text{softmax}\left(\langle x_{n,j} W_q^e, X_n W_k^e \rangle\right) X_n W_v^e xn+1,j=FFN(xn,j′;Wfe)x_{n+1,j} = \text{FFN}(x'_{n,j}; W_f^e) where Wqe,Wke,WveW_q^e, W_k^e, W_v^e are query, key, and value projection matrices in the self-attention sub-layer, and WfeW_f^e parameterizes the position-wise feedforward network.

    3. Decoder: The decoder contains a stack of N=6N=6 identical layers that processes target token sequence Y=(y1,…,yT)Y = (y_1, \dots, y_T). For layer n∈{0,…,N−1}n \in \{0, \dots, N-1\} with input Zn={zn,1,…,zn,T}Z_n = \{z_{n,1}, \dots, z_{n,T}\} where Z0=YZ_0 = Y:

    • Causal masked self-attention over preceding generated tokens: zn,j′=ATTN(zn,j,Zn,1:j;Wqd,Wkd,Wvd)z'_{n,j} = \text{ATTN}(z_{n,j}, Z_{n,1:j}; W_q^d, W_k^d, W_v^d)
    • Cross-attention attending over top-layer visual encoder representations HH: zn,j′′=ATTN(zn,j′,H;Wqc,Wkc,Wvc)z''_{n,j} = \text{ATTN}(z'_{n,j}, H; W_q^c, W_k^c, W_v^c)
    • Feedforward layer: zn+1,j=FFN(zn,j′′;Wfd)z_{n+1,j} = \text{FFN}(z''_{n,j}; W_f^d) where Wqd,Wkd,WvdW_q^d, W_k^d, W_v^d and Wqc,Wkc,WvcW_q^c, W_k^c, W_v^c denote decoder self-attention and cross-attention weight matrices, and WfdW_f^d is the decoder feedforward weight matrix.

    The model utilizes 8 attention heads per layer, and the word embedding matrix (dimension 512) is weight-tied to the output projection layer.

  5. Knowl 5 — Inception-ResNet-v2 RNN-Based Captioning Baseline

    model/method

    The RNN-based image captioning baseline adapts the Show-and-Tell sequence-to-sequence model using an Inception-ResNet-v2 feature extractor:

    1. Visual Encoding: The CNN outputs visual feature vectors X=(x1,…,xL)X = (x_1, \dots, x_L), evaluated either as a single image embedding (1×11 \times 1, L=1L=1) or serialized across an 8×88 \times 8 spatial partition grid (L=64L=64).

    2. Recurrent Encoder: A 1-layer, 512-dimensional LSTM encoder consumes the visual features sequentially: hl=RNNenc(xl,hl−1),with H=hLh_l = \text{RNN}_{\text{enc}}(x_l, h_{l-1}), \quad \text{with } H = h_L

    3. Recurrent Decoder: A separate 1-layer, 512-dimensional LSTM decoder generates text representations ztz_t conditioned on word embeddings yty_t, initialized with the final image encoder state: zt=RNNdec(yt,zt−1),where z0=Hz_t = \text{RNN}_{\text{dec}}(y_t, z_{t-1}), \quad \text{where } z_0 = H

    This architecture does not employ cross-attention mechanisms, as experiments showed that Show-Attend-Tell style attention yielded inferior performance in this setup.

  6. Knowl 6 — Experimental Training and Inference Configuration

    experimental setup

    Captioning models are trained and evaluated under the following standardized configuration:

    • Image Preprocessing: Random cropping and distortion with scale ratios from 50% to 100% to prevent pixel-level overfitting.
    • Text Handling: Captions are truncated to a maximum of 15 tokens. Tokens with frequency <4< 4 are replaced with <UNK>, resulting in a vocabulary of ≈9,000\approx 9,000 tokens for MS-COCO and ≈25,000\approx 25,000 for Conceptual Captions. Word embedding size is 512, tied to the output projection matrix.
    • Optimization: Maximum Likelihood Estimation (MLE) cross-entropy loss optimized with Adagrad at learning rate 0.010.01 and mini-batch size 25. Models are trained for 5×1065 \times 10^6 steps with batch updates asynchronously distributed across 40 workers.
    • Checkpoint Selection: Final checkpoints are selected according to the best CIDEr score on the validation set of the respective training domain.
    • Inference: Autoregressive decoding with beam search using a beam size of 4.
  7. Knowl 7 — Human Evaluation of Captioning Models on the Flickr 1K Benchmark

    empirical result

    To evaluate generalization without in-domain bias, models trained on MS-COCO versus Conceptual Captions were evaluated on the Flickr 1K test set under double-blind human judgment. Three professional raters scored each image-caption pair as GOOD or BAD based on common-sense visual accuracy:

    Model Training Dataset 1+ GOOD 2+ GOOD 3+ GOOD
    RNN8×8\text{RNN}_{8\times 8} COCO 0.390 0.276 0.173
    T2T8×8\text{T2T}_{8\times 8} COCO 0.478 0.362 0.275
    RNN8×8\text{RNN}_{8\times 8} Conceptual 0.571 0.418 0.277
    T2T8×8\text{T2T}_{8\times 8} Conceptual 0.659 0.506 0.355
    1. Training on Conceptual Captions produces substantially better generalization than training on MS-COCO across all model architectures: for the T2T8×8\text{T2T}_{8\times 8} model, majority positive rating (2+2+ GOOD) increases from 36.2%36.2\% to 50.6%50.6\% (a 14.4 percentage point gain).
    2. Transformer-based models (T2T8×8\text{T2T}_{8\times 8}) outperform RNN-based models (RNN8×8\text{RNN}_{8\times 8}) by over 8 percentage points in majority approval under both COCO training (36.2%36.2\% vs. 27.6%27.6\%) and Conceptual Captions training (50.6%50.6\% vs. 41.8%41.8\%).
  8. Knowl 8 — Cross-Dataset Automated Evaluation Metrics Comparison

    data/table

    Automated evaluation metrics (CIDEr, ROUGE-L, METEOR, SPICE) were measured across models trained on MS-COCO versus Conceptual Captions and evaluated on three test sets:

    COCO C40 Test Set
    Model Training Dataset CIDEr ROUGE-L METEOR
    RNN1×1\text{RNN}_{1\times 1} COCO 1.021 0.694 0.348
    RNN8×8\text{RNN}_{8\times 8} COCO 1.044 0.698 0.354
    T2T1×1\text{T2T}_{1\times 1} COCO 1.032 0.700 0.358
    T2T8×8\text{T2T}_{8\times 8} COCO 1.032 0.700 0.356
    RNN1×1\text{RNN}_{1\times 1} Conceptual 0.403 0.445 0.191
    RNN8×8\text{RNN}_{8\times 8} Conceptual 0.410 0.437 0.189
    T2T1×1\text{T2T}_{1\times 1} Conceptual 0.348 0.403 0.171
    T2T8×8\text{T2T}_{8\times 8} Conceptual 0.345 0.400 0.170
    22.5K Conceptual Captions Test Set
    Model Training Dataset CIDEr ROUGE-L SPICE
    RNN1×1\text{RNN}_{1\times 1} COCO 0.183 0.149 0.062
    RNN8×8\text{RNN}_{8\times 8} COCO 0.191 0.152 0.065
    T2T1×1\text{T2T}_{1\times 1} COCO 0.184 0.148 0.062
    T2T8×8\text{T2T}_{8\times 8} COCO 0.190 0.151 0.064
    RNN1×1\text{RNN}_{1\times 1} Conceptual 1.351 0.326 0.235
    RNN8×8\text{RNN}_{8\times 8} Conceptual 1.401 0.330 0.240
    T2T1×1\text{T2T}_{1\times 1} Conceptual 1.588 0.331 0.254
    T2T8×8\text{T2T}_{8\times 8} Conceptual 1.676 0.336 0.257
    Flickr 1K Test Set
    Model Training Dataset CIDEr ROUGE-L SPICE
    RNN1×1\text{RNN}_{1\times 1} COCO 0.340 0.414 0.101
    RNN8×8\text{RNN}_{8\times 8} COCO 0.356 0.413 0.103
    T2T1×1\text{T2T}_{1\times 1} COCO 0.341 0.404 0.101
    T2T8×8\text{T2T}_{8\times 8} COCO 0.359 0.416 0.103
    RNN1×1\text{RNN}_{1\times 1} Conceptual 0.269 0.310 0.076
    RNN8×8\text{RNN}_{8\times 8} Conceptual 0.275 0.309 0.076
    T2T1×1\text{T2T}_{1\times 1} Conceptual 0.226 0.280 0.068
    T2T8×8\text{T2T}_{8\times 8} Conceptual 0.227 0.277 0.066

    When evaluated in-domain, models achieve strong metric performance (COCO models score 1.02–1.04 CIDEr on COCO C40; Conceptual-trained T2T8×8\text{T2T}_{8\times 8} scores 1.676 CIDEr on Conceptual Captions). When tested cross-domain, scores drop sharply (COCO models drop <0.20< 0.20 CIDEr on Conceptual Captions; Conceptual models drop to 0.34–0.410.34–0.41 CIDEr on COCO).

  9. Knowl 9 — Suppression of Object Hallucination and Expressive Vocabulary via Conceptual Captions Training

    empirical result

    Training captioning models on Conceptual Captions mitigates key failure modes observed in models trained on MS-COCO:

    1. Hallucination Suppression: MS-COCO models frequently hallucinate objects due to spurious co-occurrences in the training annotations (e.g., hallucinating "cake" whenever a child sits at a table, or predicting "clock and two doors" in architectural hallways). Conceptual Captions data exhibits lower co-occurrence bias across its 3.3M diverse samples, allowing models to disentangle visual concepts without inserting ungrounded objects.

    2. Semantic Expressiveness: Conceptual-trained models generate more specific and informative domain terminology than COCO models (e.g., identifying "graduates" instead of "group of men", or "the cloister of the cathedral" instead of "a narrow hallway").

    3. Robustness to Diverse Image Modalities: While MS-COCO models fail on non-photorealistic images by hallucinating natural objects (e.g., describing a cartoon as "a stuffed animal" or "a picture of a fish on the side of a car"), Conceptual-trained models successfully recognize and describe non-natural images (e.g., describing vector graphics as "a cartoon businessman asking for help").

  10. Knowl 10 — Failure of Automated Captioning Metrics to Penalize Object Hallucinations

    limitation

    Standard automated evaluation metrics for image captioning (CIDEr, ROUGE-L, METEOR, SPICE) exhibit severe divergence from human quality judgments on out-of-domain evaluation benchmarks:

    1. Metric–Human Contradiction: On the Flickr 1K test set, automated n-gram metrics rate COCO-trained models higher than Conceptual Captions-trained models (e.g., CIDEr scores of ≈0.36\approx 0.36 vs. ≈0.23\approx 0.23) and rank RNNs above Transformers. In direct contrast, double-blind human evaluations show that Conceptual Captions-trained Transformer models outperform COCO-trained RNNs by a large margin (50.6%50.6\% vs. 27.6%27.6\% majority approval).

    2. Under-Penalization of Hallucinations: Standard metrics calculate n-gram precision/recall against reference captions, imposing only a minor precision penalty for tokens that do not match the groundtruth. In contrast, human annotators penalize object hallucinations severely, marking descriptions with non-existent objects as completely BAD. Because COCO-trained models reproduce common stylistic phrasing found in human-written Flickr captions despite hallucinating visual entities, automated metrics artificially reward them.

Coverage note — None was omitted.

References

  1. 1.Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. 2016. SPICE: semantic propositional image caption evaluation. In ECCV.
  2. 2.D. Bahdanau, K. Cho, and Y. Bengio. 2015. Neural machine translation by jointly learning to align and translate. In Proceedings of ICLR.
  3. 3.Satanjeev Banerjee and Alon Lavie. 2005. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the ACL Workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization.
  4. 4.Yoshua Bengio. 2009. Learning deep architectures for ai. Found. Trends Mach. Learn. 2(1):1–127.
  5. 5.Raffaella Bernardi, Ruket Cakici, Desmond Elliott, Aykut Erdem, Erkut Erdem, Nazli Ikizler-Cinbis, Frank Keller, Adrian Muscat, and Barbara Plank. 2016. Automatic description generation from images: A survey of models, datasets, and evaluation measures. JAIR 55.
  6. 6.Craig Chambers, Ashish Raniwala, Frances Perry, Stephen Adams, Robert Henry, Robert Bradshaw, and Nathan. 2010. Flumejava: Easy, efficient data-parallel pipelines. In ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI). 2 Penn Plaza, Suite 701 New York, NY 10121-0701, pages 363–375. http://dl.acm.org/citation.cfm?id=1806638.
  7. 7.Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio. 2014. Empirical evaluation of gated recurrent neural networks on sequence modeling. arXiv preprint arXiv:1412.3555 .
  8. 8.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. 2009. ImageNet: A large-scale hierarchical image database. In CVPR.
  9. 9.Nan Ding and Radu Soricut. 2017. Cold-start reinforcement learning with softmax policy gradients. In NIPS.
  10. 10.Jeff Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell. 2014. Long-term recurrent convolutional networks for visual recognition and description. In Proc. of IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  11. 11.John Duchi, Elad Hazan, and Yoram Singer. 2011. Adaptive subgradient methods for online learning and stochastic optimization. Journal of Machine Learning Research 12(Jul):2121–2159.
  12. 12.Hao Fang, Saurabh Gupta, Forrest Iandola, Rupesh Srivastava, Li Deng, Piotr Dollár, Jianfeng Gao, Xiaodong He, Margaret Mitchell, John Platt, et al. 2015. From captions to visual concepts and back. In Proc. of IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  13. 13.Sepp Hochreiter and Jürgen Schmidhuber. 1997. Long short-term memory. Neural computation 9(8):1735–1780.
  14. 14.Micah Hodosh, Peter Young, and Julia Hockenmaier. 2013. Framing image description as a ranking task: Data, models and evaluation metrics. JAIR .
  15. 15.Jonathan Huang, Vivek Rathod, Chen Sun, Menglong Zhu, Anoop Korattikara, Alireza Fathi, Ian Fischer, Zbigniew Wojna, Yang Song, Sergio Guadarrama, and Kevin Murphy. 2016. Speed/accuracy trade-offs for modern convolutional object detectors. CoRR abs/1611.10012.
  16. 16.Andrej Karpathy and Li Fei-Fei. 2015. Deep visual-semantic alignments for generating image descriptions. In Proc. of IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  17. 17.Ryan Kiros, Ruslan Salakhutdinov, and Richard S Zemel. 2015. Unifying visual-semantic embeddings with multimodal neural language models. Transactions of the Association for Computational Linguistics .
  18. 18.A. Krizhevsky, I. Sutskever, and G. Hinton. 2012. Imagenet classification with deep convolutional neural networks. In NIPS.
  19. 19.Chin-Yew Lin and Franz Josef Och. 2004. Automatic evaluation of machine translation quality using longest common subsequence and skip-bigram statistics. In Proceedings of ACL.
  20. 20.Tsung-Yi Lin, Michael Maire, Serge J. Belongie, Lubomir D. Bourdev, Ross B. Girshick, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. 2014. Microsoft COCO: common objects in context. CoRR abs/1405.0312.
  21. 21.Siqi Liu, Zhenhai Zhu, Ning Ye, Sergio Guadarrama, and Kevin Murphy. 2017. Optimization of image description metrics using policy gradient methods. In International Conference on Computer Vision (ICCV).
  22. 22.Junhua Mao, Jiajing Xu, Yushi Jing, and Alan Yuille. 2016. Training and evaluating multimodal word embeddings with large-scale web annotated images. In NIPS.
  23. 23.Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba. 2015. Sequence level training with recurrent neural networks. CoRR abs/1511.06732.
  24. 24.Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014. Sequence to sequence learning with neural networks. In Advances in neural information processing systems. pages 3104–3112.
  25. 25.Christian Szegedy, Sergey Ioffe, and Vincent Vanhoucke. 2016. Inception-v4, inception-resnet and the impact of residual connections on learning. CoRR abs/1602.07261.
  26. 26.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems.
  27. 27.Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. 2015. Cider: Consensus-based image description evaluation. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  28. 28.Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. 2015a. Show and tell: A neural image caption generator. In Proceedings of the IEEE conference on computer vision and pattern recognition. pages 3156–3164.
  29. 29.Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. 2015b. Show and tell: A neural image caption generator. In Proc. of IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  30. 30.Kelvin Xu, Jimmy Ba, Ryan Kiros, Aaron Courville, Ruslan Salakhutdinov, Richard Zemel, and Yoshua Bengio. 2015. Show, attend and tell: Neural image caption generation with visual attention. In Proc. of the 32nd International Conference on Machine Learning (ICML).
  31. 31.Z. Yang, Y. Yuan, Y. Wu, R. Salakhutdinov, and W. W. Cohen. 2016. Review networks for caption generation. In NIPS.
  32. 32.Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. 2014. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. TACL 2:67–78.

Citation

MLA
Sharma, P., et al. “Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning”. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2018, pp. 2556–65, https://doi.org/10.18653/v1/P18-1238.
APA
Sharma, P., Ding, N., Goodman, S., & Soricut, R. (2018). Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2556–2565. https://doi.org/10.18653/v1/P18-1238
Chicago
Sharma, P., N. Ding, S. Goodman, and R. Soricut. 2018. “Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning”. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2556–65. https://doi.org/10.18653/v1/P18-1238.
Harvard
Sharma, P. et al. (2018) “Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning”, Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 2556–2565. Available at: https://doi.org/10.18653/v1/P18-1238.
Vancouver
1. Sharma P, Ding N, Goodman S, Soricut R (2018) Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning. In: Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 2556–2565

BibTeX

@inproceedings{sharma-etal-2018-conceptual,
    title = "Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning",
    author = "Sharma, Piyush  and
      Ding, Nan  and
      Goodman, Sebastian  and
      Soricut, Radu",
    editor = "Gurevych, Iryna  and
      Miyao, Yusuke",
    booktitle = "Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2018",
    address = "Melbourne, Australia",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/P18-1238/",
    doi = "10.18653/v1/P18-1238",
    pages = "2556--2565"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/