VisualGPT: Data-efficient Adaptation of Pretrained Language Models for Image Captioning

Jun ChenHan GuoKai YiBoyang LiMohamed Elhoseiny

article2022CVPR303 citations

Proposes a data-efficient adaptation framework for image captioning that balances pretrained language model knowledge with visual features via a self-resurrecting activation unit, outperforming baselines under extreme low-data regimes.

Listen

Deploying machine learning models for image captioning typically requires hundreds of thousands of paired image and text examples. Manually curating and annotating such datasets is costly and labor-intensive, while web-scraped alternatives often introduce low-quality or inaccurate data. In specialized fields, such as medical diagnostic reporting and low-resource languages, gathering large collections of paired data is often impossible. Consequently, building automated captioning systems that learn effectively from minimal labeled data has become a critical operational objective.

The article demonstrates and evaluates VisualGPT, a framework designed to efficiently generate image captions using only tiny fractions of domain training data. The main objective is to adapt large, unimodal pretrained language models to visual captioning tasks without overwriting their rich linguistic knowledge during fine-tuning.

To achieve this, the authors combine a randomly initialized visual encoder with a caption decoder initialized from a pretrained language model. They introduce a self-resurrecting attention mechanism that selectively gates visual and textual inputs. This mechanism produces sparse activations to shield learned language structures from disruption while retaining the ability to reactivate dormant gating pathways during training. The authors evaluated the model across two standard benchmark image datasets—using sample splits as small as 0.1%, 0.5%, and 1% of the training data—as well as a specialized medical dataset comprising chest radiographs and clinical reports.

The primary findings show that VisualGPT significantly outperforms competing architectures in low-data regimes. When trained on only 0.1% to 1.0% of standard benchmark data, the system exceeded the performance of baseline models by up to 10.0% CIDEr score on one benchmark and 17.9% on another. In the medical domain, the model established a new state of the art on the chest radiograph dataset without requiring large-scale multimodal pretraining. Furthermore, human evaluations confirmed that human raters consistently preferred VisualGPT captions over alternatives by approximately 37% to 39% of total votes, while reporting significantly reduced rates of object hallucination and omission.

These results demonstrate that organizations can deploy high-performing image description models without incurring the massive financial and temporal costs of large-scale dataset curation. By effectively transferring unimodal language knowledge into cross-modal tasks, technical teams can bypass expensive multimodal pretraining steps and mitigate the risk of deploying models in niche or data-scarce domains.

Organizations operating in data-constrained domains should pilot this selective gating approach to adapt existing text-based models rather than collecting massive paired datasets from scratch. Development teams should use the provided open-source implementation to evaluate performance on their internal image workflows before making costly investments in manual annotation.

Readers should note that the performance advantage of this method diminishes as in-domain training data scales to full dataset capacity, where abundant paired examples reduce the relative value of transferred language priors. Confidence in these results is high for low-resource and data-restricted settings, though performance remains bounded by the vocabulary diversity of the underlying language model and the quality of initial object extraction.

arXiv: 2102.10407
Cover for VisualGPT: Data-efficient Adaptation of Pretrained Language Models for Image Captioning

Abstract

The limited availability of annotated data often hinders real-world applications of machine learning. To efficiently learn from small quantities of multimodal data, we leverage the linguistic knowledge from a large pre-trained language model (PLM) and quickly adapt it to new domains of image captioning. To effectively utilize a pretrained model, it is critical to balance the visual input and prior linguistic knowledge from pretraining. We propose VisualGPT, which employs a novel self-resurrecting encoder-decoder attention mechanism to quickly adapt the PLM with a small amount of in-domain image-text data. The proposed self-resurrecting activation unit produces sparse activations that prevent accidental overwriting of linguistic knowledge. When trained on 0.1%, 0.5% and 1% of the respective training sets, VisualGPT surpasses the best baseline by up to 10.0% CIDEr on MS COCO [43] and 17.9% CIDEr on Conceptual Captions [63]. Furthermore, VisualGPT achieves the state-of-the-art result on IU X-ray [15], a medical report generation dataset. Our code is available at https://github.com/Vision-CAIR/VisualGPT.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Preliminaries: Transformer for Captioning
  • 4. VisualGPT
  • 4.1. Self-Resurrecting Activation Unit
  • 4.2. The Architecture and Training of VisualGPT
  • 5. Experiments
  • 5.1. Datasets and Evaluation Metrics
  • 5.2. Experimental Settings
  • 5.3. Quantitative Results
  • 5.4. Ablation Studies
  • 5.5. Human Study
  • 5.6. Analysis
  • 5.7. Limitation
  • 6. Conclusions
  • References

Knowls

  1. Knowl 1 — VisualGPT Architecture for Data-Efficient Image Captioning

    model/method

    VisualGPT adapts an autoregressive pretrained language model (PLM), specifically GPT-2, to the task of image captioning under scarce in-domain multimodal data. The architecture comprises:

    1. Visual Encoder: A Transformer encoder consisting of KK layers. An off-the-shelf object detector extracts regional visual features and spatial location embeddings from an input image, producing an encoder representation I∈RS×O×KI \in \mathbb{R}^{S \times O \times K}, where OO is the number of detected object regions, SS is the hidden state dimensionality, and KK is the number of encoder layers.
    2. Pretrained Caption Decoder: A Transformer decoder with MM layers initialized directly from a pretrained autoregressive language model (such as GPT-2).
    3. Meshed Encoder-Decoder Cross-Attention: Encoder-decoder attention layers are newly inserted between the masked self-attention and feed-forward sublayers of the decoder. Connections are formed between all decoder layers and all encoder layers (meshed connections). Within each cross-attention module, the Self-Resurrecting Activation Unit (SRAU) balances visual features from the encoder and linguistic states from previous decoder layers to prevent destructive updates to pretrained linguistic weights.
  2. Knowl 2 — Self-Resurrecting Activation Unit

    model/method

    The Self-Resurrecting Activation Unit (SRAU) is a gating module designed to balance visual representations and linguistic context in the cross-attention layer of VisualGPT. Given the decoder state H∈Rt×SH \in \mathbb{R}^{t \times S} (at generation step tt with hidden dimension SS) and the visual encoder representation I∈RO×SI \in \mathbb{R}^{O \times S}, cross-attention is computed as:

    EncDecAttn(H,I)=softmax((WqH)(WkI)⊤D)WvI\text{EncDecAttn}(H, I) = \text{softmax}\left(\frac{(W^q H)(W^k I)^\top}{\sqrt{D}}\right) W^v I

    where Wq,Wk,WvW^q, W^k, W^v are trainable projection matrices and DD is the scaling dimension.

    The SRAU combines the attended visual feature and the decoder linguistic state using complementary element-wise gating matrices BvisB^{\text{vis}} and BlanB^{\text{lan}}:

    Output=Bvis⊗EncDecAttn(H,I)+Blan⊗H\text{Output} = B^{\text{vis}} \otimes \text{EncDecAttn}(H, I) + B^{\text{lan}} \otimes H

    where ⊗\otimes denotes element-wise multiplication. The gates are computed element-wise for coordinate (i,j)(i, j) via rectified sigmoid activations:

    Bvis[i,j]=σ(H[i,j])⋅I(σ(H[i,j])>τ)B^{\text{vis}}[i, j] = \sigma(H[i, j]) \cdot \mathbb{I}(\sigma(H[i, j]) > \tau)

    Blan[i,j]=(1−σ(H[i,j]))⋅I(1−σ(H[i,j])>τ)B^{\text{lan}}[i, j] = (1 - \sigma(H[i, j])) \cdot \mathbb{I}(1 - \sigma(H[i, j]) > \tau)

    where σ(⋅)\sigma(\cdot) is the standard logistic sigmoid function, I(⋅)\mathbb{I}(\cdot) is the indicator function, and τ∈[0,1)\tau \in [0, 1) is a sparsity threshold hyperparameter (default τ=0.2\tau = 0.2).

    When gate values drop below τ\tau, the indicator function zeroes the activation, enforcing sparsity and shielding pretrained linguistic weights from noisy gradient updates. Because the gates are asymmetric, when one gate outputs zero with zero gradient, the gradient through the non-zero complementary gate can still update HH, allowing the zeroed gate to reactivate (resurrect) during subsequent optimization.

  3. Knowl 3 — Two-Stage Training Procedure for VisualGPT

    model/method

    VisualGPT is trained in two successive stages using a small set of in-domain paired image-caption examples:

    1. Supervised Cross-Entropy Training: The model parameters—including the pretrained GPT-2 decoder and randomly initialized encoder and cross-attention parameters—are optimized to minimize the autoregressive cross-entropy loss over ground-truth caption tokens w1,…,wTw_1, \dots, w_T:

    LXE=−∑t=1Tlog⁡P(wt∣w1,…,wt−1,I)\mathcal{L}_{\text{XE}} = -\sum_{t=1}^T \log P(w_t \mid w_1, \dots, w_{t-1}, I)

    where II is the visual encoder output.

    1. Self-Critical Sequence Training (SCST): After a predefined number of supervised epochs, training transitions to policy gradient reinforcement learning using CIDEr as the reward. The objective is:

    LRL=−(r(ws)−r(w^))∑t=1Tlog⁡P(wts∣w1s,…,wt−1s,I)\mathcal{L}_{\text{RL}} = -(r(w^s) - r(\hat{w})) \sum_{t=1}^T \log P(w_t^s \mid w_1^s, \dots, w_{t-1}^s, I)

    where wsw^s is a sampled caption sequence, w^\hat{w} is the baseline caption obtained via greedy decoding at test mode, and r(⋅)r(\cdot) denotes the sentence-level CIDEr reward.

  4. Knowl 4 — Data-Constrained Captioning Performance on MS COCO and Conceptual Captions

    data/table

    VisualGPT was benchmarked on data-constrained regimes using randomly sampled subsets of 0.1%, 0.5%, and 1% of the training data from MS COCO (567, 2,835, and 5,670 image-caption pairs) and Conceptual Captions (3,300, 16,500, and 33,000 pairs). Scores are reported as the average over 4 runs across BLEU-1 (B1), BLEU-4 (B4), METEOR (M), ROUGE-L (R), and CIDEr (C).

    MS COCO Conceptual Captions
    Method PLM B1 B4 M R C B1 B4 M R C
    0.1% training data
    Transformer None 57.4 13.1 16.7 40.7 40.8 12.4 2.4 4.9 15.2 21.2
    M2\text{M}^2 Transformer None 56.9 13.1 16.9 40.6 40.9 13.1 2.8 4.8 15.5 23.5
    AoA Transformer None 56.6 13.5 15.9 40.7 38.4 11.4 2.4 4.6 14.7 20.9
    X-Transformer None 56.7 12.9 16.5 40.6 40.4 12.8 2.7 4.7 15.3 23.1
    OSCAR BERT 53.8 11.9 17.1 39.5 41.0 12.2 2.4 4.3 14.8 21.9
    Transformer GPT 56.8 15.3 17.0 41.2 42.9 13.2 2.5 5.0 15.1 21.9
    M2\text{M}^2 Transformer GPT 54.9 14.7 16.6 41.1 41.0 11.9 2.6 4.9 15.4 24.0
    AoA Transformer GPT 55.5 14.4 16.2 40.7 40.1 11.8 2.8 4.6 13.9 20.5
    VisualGPT GPT 58.2 16.4 18.5 41.9 45.1 13.9 3.2 5.6 16.7 27.7
    0.5% training data
    Transformer None 62.8 18.8 19.4 25.2 59.2 13.2 3.3 5.5 16.3 29.6
    M2\text{M}^2 Transformer None 63.3 19.4 19.8 45.6 61.3 14.5 3.6 6.0 17.1 32.0
    AoA Transformer None 63.5 20.2 19.4 45.8 63.9 13.8 3.3 5.6 17.9 31.8
    X-Transformer None 62.9 19.0 19.6 45.7 62.0 14.2 3.5 5.8 17.3 32.1
    OSCAR BERT 59.2 18.0 21.0 45.3 60.2 14.4 3.7 6.1 17.2 33.5
    Transformer GPT 65.1 21.8 20.6 46.6 69.5 16.2 3.8 6.5 18.3 35.6
    M2\text{M}^2 Transformer GPT 64.7 21.8 20.7 47.1 68.5 13.9 3.6 6.0 17.2 34.1
    AoA Transformer GPT 64.2 21.2 20.5 46.5 67.2 14.8 3.6 6.2 17.6 34.1
    VisualGPT GPT 66.2 22.1 21.1 47.3 70.3 15.9 4.2 6.7 18.5 37.2
    1.0% training data
    Transformer None 66.0 21.9 21.1 47.3 71.9 13.9 3.7 6.3 18.1 37.9
    M2\text{M}^2 Transformer None 67.1 23.4 21.3 48.3 73.0 16.0 4.1 6.8 18.9 39.8
    AoA Transformer None 67.6 23.6 21.5 48.4 75.5 14.9 4.1 6.5 18.6 39.0
    X-Transformer None 67.0 23.6 21.2 48.1 47.1 15.6 4.0 6.6 18.7 39.5
    OSCAR BERT 67.2 23.3 22.5 49.1 78.4 16.1 4.2 6.7 18.9 40.6
    Transformer GPT 68.5 25.1 22.1 49.0 80.5 17.8 4.2 6.7 19.0 40.2
    M2\text{M}^2 Transformer GPT 68.2 25.0 22.4 49.2 80.4 15.4 3.9 6.5 17.9 39.1
    AoA Transformer GPT 68.5 24.6 22.0 48.6 78.4 15.4 3.9 6.5 17.9 38.5
    VisualGPT GPT 69.5 25.6 22.6 49.6 80.9 16.3 4.3 6.9 19.3 40.9

    VisualGPT outperforms all randomly initialized and GPT-initialized baselines across all metric configurations. On MS COCO, VisualGPT achieves CIDEr improvements over the best baseline of +4.1 points at 0.1%, +6.4 points at 0.5%, and +2.5 points at 1.0%. On Conceptual Captions, VisualGPT delivers CIDEr improvements of +4.2 points at 0.1%, +3.5 points at 0.5%, and +0.3 points at 1.0%.

  5. Knowl 5 — Medical Report Generation on the IU X-Ray Dataset

    empirical result

    VisualGPT was evaluated on the domain-specific IU X-ray benchmark, which consists of 7,470 radiography images and 3,955 human-written clinical reports (with a standard split of 5,226 training images / 2,770 reports and 2,770 test reports).

    Model B-1 B-2 B-3 B-4 R M C
    Att2in 22.4 12.9 8.9 6.8 30.8 - 29.7
    CoAtt 45.5 28.8 20.5 15.4 36.9 - 27.7
    HRGR 43.8 29.8 20.8 15.1 32.2 - 34.3
    CMAS-RL 46.4 30.1 21.0 15.4 37.1 - 27.5
    Chen et al. 47.0 30.4 21.9 16.5 37.1 18.7 -
    VisualGPT 48.0 31.3 22.2 15.9 37.4 20.5 49.7

    VisualGPT establishes a new state of the art on IU X-ray with a CIDEr score of 49.7, outperforming HRGR (34.3 CIDEr) by 15.4 points, while also attaining the top BLEU-1 (48.0), BLEU-2 (31.3), BLEU-3 (22.2), ROUGE-L (37.4), and METEOR (20.5) scores.

  6. Knowl 6 — Comparison with Semi-Supervised and Unsupervised Image Captioning

    empirical result

    VisualGPT was compared against semi-supervised and unsupervised image captioning methods using the experimental split protocol of Kim et al., where only 1% of MS COCO images (1,133 images with all corresponding captions) are available as paired training data:

    Model B-1 B-4 M R C
    Kim et al. 58.1 13.4 15.9 - 36.0
    Kim et al. + unpaired 63.0 18.7 20.7 - 55.2
    Gu et al. (unsupervised) 46.2 5.4 13.2 - 17.7
    Feng et al. (unsupervised) 58.9 18.6 17.9 - 54.9
    VisualGPT 67.1 24.3 21.9 48.6 75.8

    Using exclusively the 1% supervised paired images without any supplementary unpaired images or captions, VisualGPT achieves a CIDEr score of 75.8. This improves over the semi-supervised baseline of Kim et al. by 39.8 CIDEr points on the same 1% paired data, and exceeds the Kim et al. + unpaired variant (which incorporates the remaining 99% of MS COCO as unpaired images and text) by 20.6 CIDEr points. VisualGPT also surpasses unsupervised methods trained on tens of millions of unpaired data points.

  7. Knowl 7 — Ablation of SRAU Gating Asymmetry and Sparsity Threshold

    empirical result

    Ablation studies evaluated the specific architectural and parameter choices of the Self-Resurrecting Activation Unit (SRAU):

    1. Effect of Gate Normalization: SRAU was compared against a normalized, symmetric variant where gate activations are divided by their sum:

    B~vis[i,j]=Bvis[i,j]Bvis[i,j]+Blan[i,j],B~lan[i,j]=Blan[i,j]Bvis[i,j]+Blan[i,j]\tilde{B}^{\text{vis}}[i, j] = \frac{B^{\text{vis}}[i, j]}{B^{\text{vis}}[i, j] + B^{\text{lan}}[i, j]}, \quad \tilde{B}^{\text{lan}}[i, j] = \frac{B^{\text{lan}}[i, j]}{B^{\text{vis}}[i, j] + B^{\text{lan}}[i, j]}

    Normalized SRAU reduces CIDEr performance relative to standard SRAU by 2.7 (at 0.1%), 1.0 (at 0.5%), and 0.3 (at 1.0%) on MS COCO, and by 2.2 (at 0.1%), 1.3 (at 0.5%), and 0.6 (at 1.0%) on Conceptual Captions. This confirms that the asymmetric gradient path of SRAU is essential for enabling dead gates to self-resurrect.

    1. Threshold Parameter τ\tau: Setting τ=0\tau = 0 corresponds to Ordinary Complementary Sigmoid Gates (OCG), which output near-zero non-zero gradients that can overwrite pretrained weights. Across 0.1%, 0.5%, and 5% data splits on MS COCO, setting τ=0.2\tau = 0.2 yielded higher CIDEr scores (45.1, 70.3, and 103.0) than τ=0\tau = 0 (44.3, 69.0, and 99.5) and τ=0.1\tau = 0.1 (42.7, 67.9, and 100.5).

    2. Unbounded Activation Functions: Substituting sigmoid gating with Leaky ReLU or GELU caused training to diverge and crash due to unbounded activation ranges.

  8. Knowl 8 — Human Evaluation of Caption Accuracy and Object Hallucination

    empirical result

    Two user studies evaluated caption quality and object faithfulness on 250 randomly sampled test images from models trained on low-data splits of MS COCO:

    1. Human Preference Ranking: Five independent annotators per image selected the most accurate descriptive caption among VisualGPT and baseline models (plain Transformer, M2\text{M}^2 Transformer, AoA Transformer, each with 3 decoder layers). VisualGPT received the highest proportion of votes across all splits:

      • 0.1% training data: VisualGPT 39.2% vs. M2\text{M}^2 Transformer 30.9%, Transformer 18.4%, AoA Transformer 11.5%.
      • 0.5% training data: VisualGPT 39.1% vs. M2\text{M}^2 Transformer 22.8%, AoA Transformer 20.9%, Transformer 17.2%.
      • 1.0% training data: VisualGPT 37.4% vs. AoA Transformer 25.0%, M2\text{M}^2 Transformer 20.8%, Transformer 16.8%. All differences were statistically significant (p<0.05p < 0.05, Pearson's Chi-square test).
    2. Hallucination and Omission Rates (1% Split):

      • Object Omission: VisualGPT achieved a "No Omission" rate of 0.66 (719 No / 367 Yes), outperforming M2\text{M}^2 Transformer (0.59), plain Transformer (0.58), and AoA Transformer (0.58), with ground-truth captions scoring 0.93.
      • Object Hallucination: VisualGPT achieved a "No Hallucination" rate of 0.67 (720 No / 360 Yes), outperforming M2\text{M}^2 Transformer (0.62), AoA Transformer (0.61), and plain Transformer (0.60), with ground-truth captions scoring 0.96.
  9. Knowl 9 — Modality Gating Dynamics Across Decoder Layers and Word Categories

    empirical result

    Analysis of the visual gate BvisB^{\text{vis}} and linguistic gate BlanB^{\text{lan}} in VisualGPT reveals distinct token-level and layer-level specialization:

    1. Token-Level Specialization: The mean value of the visual gate vector assigned to individual generated tokens reflects semantic grounding. Highly visual nouns and verbs (e.g., "bench", "wooden", "sitting", "clock", "toilet", "desk", "snowy surface") receive high visual attention scores BvisB^{\text{vis}}, whereas syntactic determiners, prepositions, and connectives (e.g., "a", "the", "to", "of", "on") receive minimal visual attention and high linguistic gate weights BlanB^{\text{lan}}.
    2. Hierarchical Layer Specialization: Across the 12 decoder layers (layers 0 to 11):
      • Layers 0 to 2 are heavily dominated by linguistic gate weights BlanB^{\text{lan}} with minimal visual input BvisB^{\text{vis}}, preserving low-level syntactic structures.
      • Layers 3 to 9 progressively increase visual attention weights BvisB^{\text{vis}}, performing cross-modal feature fusion.
      • Layers 10 and 11 achieve a balanced distribution between visual and linguistic gates.
  10. Knowl 10 — Diminishing Adaptation Advantage with Abundant In-Domain Data

    limitation

    The performance margin of VisualGPT over standard captioning models trained from scratch diminishes as the quantity of in-domain paired training data grows large. This effect is more pronounced on datasets with constrained, repetitive vocabularies (such as MS COCO) than on datasets with diverse vocabularies (such as Conceptual Captions). Pretrained unimodal language representations provide substantial value in low-data regimes where training pairs fail to cover the vocabulary space, but their relative utility decreases when paired multimodal training corpora are abundant.

Coverage note — None was omitted; all contributed models, mechanisms, mathematical formulations, quantitative benchmark results, ablation analyses, human studies, and stated limitations are fully covered.

References

  1. 1.Harsh Agrawal, Karan Desai, Yufei Wang, Xinlei Chen, Rishabh Jain, Mark Johnson, Dhruv Batra, Devi Parikh, Stefan Lee, and Peter Anderson. nocaps: novel object captioning at scale. In ICCV, 2019. 1, 2
  2. 2.Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. Bottom-up and top-down attention for image captioning and visual question answering. In CVPR, 2018. 2
  3. 3.Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. Layer normalization. arXiv 1607.06450, 2016. 3
  4. 4.Satanjeev Banerjee and Alon Lavie. Meteor: An automatic metric for mt evaluation with improved correlation with human judgments. In Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization, 2005. 4
  5. 5.Yoshua Bengio, Rejean Ducharme, Pascal Vincent, and Christian Jauvin. A neural probabilistic language model. Journal of machine learning research, 3(Feb), 2003. 2
  6. 6.Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020. 1, 2
  7. 7.Pawel Budzianowski and Ivan Vulic. Hello, it’s GPT-2 - how can I help you? towards the use of pretrained language models for task-oriented dialogue systems. In Alexandra Birch, Andrew M. Finch, Hiroaki Hayashi, Ioannis Konstas, Thang Luong, Graham Neubig, Yusuke Oda, and Katsuhito Sudoh, editors, EMNLP-IJCNLP. Association for Computational Linguistics, 2019. 1
  8. 8.Long Chen, Hanwang Zhang, Jun Xiao, Liqiang Nie, Jian Shao, Wei Liu, and Tat-Seng Chua. Sca-cnn: Spatial and channel-wise attention in convolutional networks for image captioning. In CVPR, 2017. 2
  9. 9.Shizhe Chen, Qin Jin, Peng Wang, and Qi Wu. Say as you wish: Fine-grained control of image caption generation with abstract scene graphs. In CVPR, 2020. 2
  10. 10.Zhihong Chen, Yan Song, Tsung-Hui Chang, and Xiang Wan. Generating radiology reports via memory-driven transformer. In EMNLP, 2020. 6
  11. 11.Cesc Chunseong Park, Byeongchang Kim, and Gunhee Kim. Attend to you: Personalized image captioning with context sequence memory networks. In CVPR, 2017. 2
  12. 12.Marcella Cornia, Lorenzo Baraldi, and Rita Cucchiara. Show, control and tell: A framework for generating controllable and grounded captions. In CVPR, 2019. 2
  13. 13.Marcella Cornia, Matteo Stefanini, Lorenzo Baraldi, and Rita Cucchiara. Meshed-memory transformer for image captioning. In CVPR, 2020. 1, 2, 3, 5, 6, 7
  14. 14.Bo Dai, Sanja Fidler, Raquel Urtasun, and Dahua Lin. Towards diverse and natural image descriptions via a conditional gan. In ICCV, 2017. 2
  15. 15.Dina Demner-Fushman, Marc D Kohli, Marc B Rosenman, Sonya E Shooshan, Laritza Rodriguez, Sameer Antani, George R Thoma, and Clement J McDonald. Preparing a collection of radiology examinations for distribution and retrieval. Journal of the American Medical Informatics Association, 23(2), 2016. 1, 2, 4, 5
  16. 16.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. In NAACL, 2019. 1, 2, 6
  17. 17.Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon. Unified language model pre-training for natural language understanding and generation. In NeurIPS, 2019. 1
  18. 18.Obeida ElJundi., Mohamad Dhaybi., Kotaiba Mokadam., Hazem Hajj., and Daniel Asmar. Resources and end-to-end neural network models for arabic image captioning. In Proceedings of the 15th International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications - Volume 5: VISAPP,, pages 233–241. INSTICC, SciTePress, 2020. 1
  19. 19.Ali Farhadi, Mohsen Hejrati, Mohammad Amin Sadeghi, Peter Young, Cyrus Rashtchian, Julia Hockenmaier, and David Forsyth. Every picture tells a story: Generating sentences from images. In ECCV. Springer, 2010. 2
  20. 20.Yang Feng, Lin Ma, Wei Liu, and Jiebo Luo. Unsupervised image captioning. In CVPR, 2019. 2, 6
  21. 21.Sergey Golovanov, Rauf Kurbanov, Sergey Nikolenko, Kyryl Truskovskyi, Alexander Tselousov, and Thomas Wolf. Large-scale transfer learning for natural language generation. In ACL, 2019. 1
  22. 22.Jiuxiang Gu, Shafiq Joty, Jianfei Cai, and Gang Wang. Unpaired image captioning by language pivoting. In ECCV, 2018. 6
  23. 23.Lisa Anne Hendricks, Subhashini Venugopalan, Marcus Rohrbach, Raymond Mooney, Kate Saenko, and Trevor Darrell. Deep compositional captioning: Describing novel object categories without paired training data. In CVPR, 2016. 1, 2
  24. 24.Lun Huang, Wenmin Wang, Jie Chen, and Xiao-Yong Wei. Attention on attention for image captioning. In ICCV, 2019. 1, 2, 3, 5, 7
  25. 25.Baoyu Jing, Zeya Wang, and Eric Xing. Show, describe and conclude: On exploiting the structure information of chest x-ray reports. In ACL, 2019. 6
  26. 26.Baoyu Jing, Pengtao Xie, and Eric Xing. On the automatic generation of medical imaging reports. In ACL, 2018. 6
  27. 27.Justin Johnson, Andrej Karpathy, and Li Fei-Fei. Densecap: Fully convolutional localization networks for dense captioning. In CVPR, 2016. 2
  28. 28.Andrej Karpathy and Li Fei-Fei. Deep visual-semantic alignments for generating image descriptions. In CVPR, 2015. 1, 4
  29. 29.Dong-Jin Kim, Jinsoo Choi, Tae-Hyun Oh, and In So Kweon. Image captioning with very scarce supervised data: Adversarial semi-supervised learning approach. In Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan, editors, EMNLP-IJCNLP. Association for Computational Linguistics, 2019. 2
  30. 30.Dong-Jin Kim, Jinsoo Choi, Tae-Hyun Oh, and In So Kweon. Image captioning with very scarce supervised data: Adversarial semi-supervised learning approach. In EMNLP, Hong Kong, China, Nov. 2019. Association for Computational Linguistics. 6
  31. 31.Girish Kulkarni, Visruth Premraj, Vicente Ordonez, Sagnik Dhar, Siming Li, Yejin Choi, Alexander C Berg, and Tamara L Berg. Babytalk: Understanding and generating simple image descriptions. TPAMI, 35(12), 2013. 1, 2
  32. 32.Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. Albert: A lite bert for self-supervised learning of language representations. In ICLR, 2019. 2
  33. 33.Kuang-Huei Lee, Xi Chen, Gang Hua, Houdong Hu, and Xiaodong He. Stacked cross attention for image-text matching. In ECCV, 2018. 2
  34. 34.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. 2019. 1
  35. 35.Christy Y Li, Xiaodan Liang, Zhiting Hu, and Eric P Xing. Hybrid retrieval-generation reinforced agent for medical image report generation. In NeurIPS, 2018. 6
  36. 36.Guang Li, Linchao Zhu, Ping Liu, and Yi Yang. Entangled transformer for image captioning. In ICCV, 2019. 2
  37. 37.Siming Li, Girish Kulkarni, Tamara Berg, Alexander Berg, and Yejin Choi. Composing simple image descriptions using web-scale n-grams. In CoNLL, 2011. 2
  38. 38.Xiangyang Li and Shuqiang Jiang. Know more say less: Image captioning based on scene graphs. IEEE Transactions on Multimedia, 21(8), 2019. 2
  39. 39.Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, et al. Oscar: Object-semantics aligned pre-training for vision-language tasks. In ECCV. Springer, 2020. 2, 5, 6
  40. 40.Yuan Li, Xiaodan Liang, Zhiting Hu, and Eric P Xing. Hybrid retrieval-generation reinforced agent for medical image report generation. In NeurIPS. 2018. 1
  41. 41.Yehao Li, Ting Yao, Yingwei Pan, Hongyang Chao, and Tao Mei. Pointing novel objects in image captioning. In CVPR, 2019. 2
  42. 42.Chin-Yew Lin and Eduard Hovy. Manual and automatic evaluation of summaries. In Proceedings of the ACL-02 Workshop on Automatic Summarization-Volume 4, 2002. 5
  43. 43.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft COCO: Common objects in context. In ECCV. Springer, 2014. 1, 4
  44. 44.Nelson F Liu, Matt Gardner, Yonatan Belinkov, Matthew E Peters, and Noah A Smith. Linguistic knowledge and transferability of contextual representations. In NAACL, 2019. 3
  45. 45.Siqi Liu, Zhenhai Zhu, Ning Ye, Sergio Guadarrama, and Kevin Murphy. Improved image captioning via policy gradient optimization of SPIDEr. In ICCV, 2017. 2
  46. 46.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. RoBERTa: A robustly optimized BERT pretraining approach. arXiv Preprint, arXiv 1907.11692, 2019. 1
  47. 47.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized bert pretraining approach. 2019. 2
  48. 48.Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In Hanna M. Wallach, Hugo Larochelle, Alina Beygelzimer, Florence d’Alche-Buc, Emily B. Fox, and Roman Garnett, editors, NeurIPS, 2019. 2
  49. 49.Jiasen Lu, Caiming Xiong, Devi Parikh, and Richard Socher. Knowing when to look: Adaptive attention via a visual sentinel for image captioning. In CVPR, 2017. 2
  50. 50.Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh. Neural baby talk. In CVPR, 2018. 2
  51. 51.Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. Pointer sentinel mixture models. 2017. 2
  52. 52.Toma´s Mikolov, Stefan Kombrink, Luk a´s Burget, Jan Cernock y, and Sanjeev Khudanpur. Extensions of recurrent neural network language model. In ICASSP. IEEE, 2011. 2
  53. 53.Yingwei Pan, Ting Yao, Yehao Li, and Tao Mei. X-linear attention networks for image captioning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10971–10980, 2020. 5
  54. 54.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. Bleu: a method for automatic evaluation of machine translation. In ACL, 2002. 4
  55. 55.Karl Pearson. X. on the criterion that a given system of deviations from the probable in the case of a correlated system of variables is such that it can be reasonably supposed to have arisen from random sampling. The London, Edinburgh, and Dublin Philosophical Magazine and Journal of Science, 50(302), 1900. 8
  56. 56.Matthew E Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. Deep contextualized word representations. In NAACL-HLT, 2018. 2
  57. 57.Di Qi, Lin Su, Jia Song, Edward Cui, Taroon Bharti, and Arun Sacheti. Imagebert: Cross-modal pre-training with large-scale weak-supervised image-text data. arXiv preprint arXiv:2001.07966, 2020. 2
  58. 58.Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improving language understanding by generative pre-training. 2018. 1, 2
  59. 59.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners. OpenAI blog, 1(8), 2019. 1, 2, 3
  60. 60.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21:1–67, 2020. 1
  61. 61.Steven J Rennie, Etienne Marcheret, Youssef Mroueh, Jerret Ross, and Vaibhava Goel. Self-critical sequence training for image captioning. In CVPR, 2017. 2, 4, 6
  62. 62.Anna Rohrbach, Lisa Anne Hendricks, Kaylee Burns, Trevor Darrell, and Kate Saenko. Object hallucination in image captioning. In EMNLP, 2018. 8
  63. 63.Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In ACL, 2018. 1, 4
  64. 64.Rakshith Shetty, Marcus Rohrbach, Lisa Anne Hendricks, Mario Fritz, and Bernt Schiele. Speaking the same language: Matching machine to human captions by adversarial training. In ICCV, 2017. 2
  65. 65.Richard Socher and Li Fei-Fei. Connecting modalities: Semi-supervised segmentation and annotation of images using unaligned text corpora. In CVPR. IEEE, 2010. 2
  66. 66.Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. Vl-bert: Pre-training of generic visual-linguistic representations. In ICLR, 2020. 2
  67. 67.Hao Tan and Mohit Bansal. LXMERT: learning cross-modality encoder representations from transformers. In Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan, editors, EMNLP-IJCNLP. ACL, 2019. 2
  68. 68.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NeurIPS, 2017. 3, 5, 7
  69. 69.Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh. Cider: Consensus-based image description evaluation. In CVPR, 2015. 5
  70. 70.Subhashini Venugopalan, Lisa Anne Hendricks, Marcus Rohrbach, Raymond Mooney, Trevor Darrell, and Kate Saenko. Captioning images with diverse objects. In CVPR, 2017. 2
  71. 71.Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. Show and tell: Lessons learned from the 2015 MSCOCO image captioning challenge. TPAMI, 39(4), 2016. 2
  72. 72.Qingzhong Wang and Antoni B. Chan. Cnn+cnn: Convolutional decoders for image captioning. arXiv 1805.09019, 2018. 2
  73. 73.Yufei Wang, Zhe Lin, Xiaohui Shen, Scott Cohen, and Garrison W Cottrell. Skeleton key: Image captioning by skeleton-attribute decomposition. In CVPR, 2017. 2
  74. 74.Yike Wu, Shiwan Zhao, Jia Chen, Ying Zhang, Xiaojie Yuan, and Zhong Su. Improving captioning for low-resource languages by cycle consistency. In 2019 IEEE International Conference on Multimedia and Expo (ICME), pages 362–367, 2019. 1
  75. 75.Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio. Show, attend and tell: Neural image caption generation with visual attention. In ICML, 2015. 1, 2
  76. 76.Xu Yang, Kaihua Tang, Hanwang Zhang, and Jianfei Cai. Auto-encoding scene graphs for image captioning. In CVPR, 2019. 2
  77. 77.Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. Xlnet: Generalized autoregressive pretraining for language understanding. In NeurIPS, 2019. 1
  78. 78.Zhilin Yang, Ye Yuan, Yuexin Wu, William W Cohen, and Russ R Salakhutdinov. Review networks for caption generation. In D. Lee, M. Sugiyama, U. Luxburg, I. Guyon, and R. Garnett, editors, NeurIPS, volume 29, 2016. 2
  79. 79.Benjamin Z Yao, Xiong Yang, Liang Lin, Mun Wai Lee, and Song-Chun Zhu. I2t: Image parsing to text description. Proceedings of the IEEE, 98(8), 2010. 2
  80. 80.Ting Yao, Yingwei Pan, Yehao Li, and Tao Mei. Exploring visual relationship for image captioning. In ECCV, 2018. 2
  81. 81.Ting Yao, Yingwei Pan, Yehao Li, and Tao Mei. Hierarchy parsing for image captioning. In ICCV, 2019. 2
  82. 82.Pengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang, Lei Zhang, Lijuan Wang, Yejin Choi, and Jianfeng Gao. Vinvl: Revisiting visual representations in vision-language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5579–5588, 2021. 2
  83. 83.Mingchen Zhuge, Dehong Gao, Deng-Ping Fan, Linbo Jin, Ben Chen, Haoming Zhou, Minghui Qiu, and Ling Shao. Kaleido-bert: Vision-language pre-training on fashion domain. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12647–12657, 2021. 2

Citation

MLA
Chen, J., et al. “VisualGPT: Data-efficient Adaptation of Pretrained Language Models for Image Captioning”. arXiv, 2021, http://arxiv.org/abs/2102.10407v5.
APA
Chen, J., Guo, H., Yi, K., Li, B., & Elhoseiny, M. (2021). VisualGPT: Data-efficient Adaptation of Pretrained Language Models for Image Captioning. arXiv. http://arxiv.org/abs/2102.10407v5
Chicago
Chen, J., H. Guo, K. Yi, B. Li, and M. Elhoseiny. 2021. “VisualGPT: Data-efficient Adaptation of Pretrained Language Models for Image Captioning”. arXiv. http://arxiv.org/abs/2102.10407v5.
Harvard
Chen, J. et al. (2021) “VisualGPT: Data-efficient Adaptation of Pretrained Language Models for Image Captioning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2102.10407v5.
Vancouver
1. Chen J, Guo H, Yi K, Li B, Elhoseiny M (2021) VisualGPT: Data-efficient Adaptation of Pretrained Language Models for Image Captioning. arXiv

BibTeX

@article{chen2021visualgpt,
  title = {VisualGPT: Data-efficient Adaptation of Pretrained Language Models for Image Captioning},
  author = {Chen, Jun and Guo, Han and Yi, Kai and Li, Boyang and Elhoseiny, Mohamed},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2102.10407v5},
  eprint = {2102.10407}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE