VisualGPT: Data-efficient Adaptation of Pretrained Language Models for Image Captioning
Jun ChenHan GuoKai YiBoyang LiMohamed Elhoseiny
Proposes a data-efficient adaptation framework for image captioning that balances pretrained language model knowledge with visual features via a self-resurrecting activation unit, outperforming baselines under extreme low-data regimes.
Deploying machine learning models for image captioning typically requires hundreds of thousands of paired image and text examples. Manually curating and annotating such datasets is costly and labor-intensive, while web-scraped alternatives often introduce low-quality or inaccurate data. In specialized fields, such as medical diagnostic reporting and low-resource languages, gathering large collections of paired data is often impossible. Consequently, building automated captioning systems that learn effectively from minimal labeled data has become a critical operational objective.
The article demonstrates and evaluates VisualGPT, a framework designed to efficiently generate image captions using only tiny fractions of domain training data. The main objective is to adapt large, unimodal pretrained language models to visual captioning tasks without overwriting their rich linguistic knowledge during fine-tuning.
To achieve this, the authors combine a randomly initialized visual encoder with a caption decoder initialized from a pretrained language model. They introduce a self-resurrecting attention mechanism that selectively gates visual and textual inputs. This mechanism produces sparse activations to shield learned language structures from disruption while retaining the ability to reactivate dormant gating pathways during training. The authors evaluated the model across two standard benchmark image datasets—using sample splits as small as 0.1%, 0.5%, and 1% of the training data—as well as a specialized medical dataset comprising chest radiographs and clinical reports.
The primary findings show that VisualGPT significantly outperforms competing architectures in low-data regimes. When trained on only 0.1% to 1.0% of standard benchmark data, the system exceeded the performance of baseline models by up to 10.0% CIDEr score on one benchmark and 17.9% on another. In the medical domain, the model established a new state of the art on the chest radiograph dataset without requiring large-scale multimodal pretraining. Furthermore, human evaluations confirmed that human raters consistently preferred VisualGPT captions over alternatives by approximately 37% to 39% of total votes, while reporting significantly reduced rates of object hallucination and omission.
These results demonstrate that organizations can deploy high-performing image description models without incurring the massive financial and temporal costs of large-scale dataset curation. By effectively transferring unimodal language knowledge into cross-modal tasks, technical teams can bypass expensive multimodal pretraining steps and mitigate the risk of deploying models in niche or data-scarce domains.
Organizations operating in data-constrained domains should pilot this selective gating approach to adapt existing text-based models rather than collecting massive paired datasets from scratch. Development teams should use the provided open-source implementation to evaluate performance on their internal image workflows before making costly investments in manual annotation.
Readers should note that the performance advantage of this method diminishes as in-domain training data scales to full dataset capacity, where abundant paired examples reduce the relative value of transferred language priors. Confidence in these results is high for low-resource and data-restricted settings, though performance remains bounded by the vocabulary diversity of the underlying language model and the quality of initial object extraction.
- Paper: Knowing When to Look: Adaptive Attention via a Visual Sentinel for Image Captioning, Jiasen Lu et al. (2016). This work introduces the foundational visual sentinel mechanism for gating visual versus linguistic generation in image captioning, which directly precedes and informs VisualGPT's balance between linguistic priors and visual inputs.
- Paper: Show, Attend and Tell: Neural Image Caption Generation with Visual Attention, Kelvin Xu et al. (2015). This seminal paper established the encoder-decoder attention architecture for neural image captioning that VisualGPT builds upon and modifies with self-resurrecting attention.
- Paper: Self-Critical Sequence Training for Image Captioning, Steven J. Rennie et al. (2016). This foundational paper presents the sequence-level CIDEr optimization framework commonly used to train and evaluate modern captioning models like VisualGPT.
- Paper: Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning, Piyush Sharma et al. (2018). This work establishes the Conceptual Captions dataset, which provides the primary benchmark and data-efficient training regime evaluated in VisualGPT.
- Paper: CLIP-Adapter: Better Vision-Language Models with Feature Adapters, Peng Gao et al. (2021). This paper demonstrates parameter-efficient feature adapters for vision-language models, providing key context for lightweight adaptation of frozen representations.
- Paper: VisualBERT: A Simple and Performant Baseline for Vision and Language, Liunian Harold Li et al. (2019). This work establishes early transformer-based multimodal pre-training architectures combining visual features and language tokens.
- Paper: Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks, Xiujun Li et al. (2020). This study introduces anchor-based vision-language alignment techniques for image captioning and understanding tasks.
- Paper: MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models, Deyao Zhu et al. (2024). MiniGPT-4 advances the paradigm of connecting pre-trained visual encoders directly to large language models via lightweight adaptation layers.
- Paper: LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day, Chunyuan Li et al. (2023). LLaVA-Med builds upon rapid domain adaptation of large vision-language models for biomedical imaging, extending the clinical generation capabilities explored by VisualGPT on medical reports.
- Paper: Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models, Siddharth Karamcheti et al. (2024). This work provides an extensive empirical analysis of design choices when visually conditioning language models, including parameter freezing and alignment strategies.
- Paper: Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond, Jinze Bai et al. (2023). Qwen-VL scales cross-attention vision-to-language adaptation adapters to handle complex multimodal understanding, localization, and detailed captioning.
- Paper: Vision-Language Models for Vision Tasks: A Survey, Jingyi Zhang et al. (2023). This comprehensive survey categorizes the architectures and parameter-efficient adaptation methods used to transfer vision-language models to downstream recognition and captioning tasks.
- Paper: Multi-Modal Hallucination Control by Visual Information Grounding, Alessandro Favero et al. (2024). This paper analyzes and addresses the problem of language-prior dominance over visual input during multimodal generation, directly continuing VisualGPT's focus on balancing visual tokens with prior linguistic knowledge.
- Paper: Efficient Multimodal Fusion via Interactive Prompting, Yaowei Li et al. (2023). This paper extends parameter-efficient multimodal fusion through interactive cross-modal prompting between frozen unimodal vision and language models.
