Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts

Yan ZengXinsong ZhangHang Li

article2022ICML395 citations

Proposes X-VLM, a vision-language pre-training framework that learns multi-grained cross-modal alignments across objects, regions, and full images without relying on object detectors, achieving state-of-the-art results on visual grounding, retrieval, and reasoning tasks.

Listen

Existing vision-language models struggle to balance broad scene understanding with detailed object comprehension. Traditional methods typically rely on rigid object detectors that miss complex relationships among multiple objects, or they encode only entire images, losing critical fine-grained details needed for precise visual reasoning. This article addresses these trade-offs by introducing and evaluating X-VLM, an approach designed to learn multi-grained visual-linguistic alignments across objects, regions, and full images simultaneously.

The researchers developed a modular framework combining an image encoder, a text encoder, and a cross-modal encoder. Rather than relying on separate object detectors, the system is trained directly to locate visual concepts in images based on text descriptions while concurrently matching text to visual concepts at multiple levels of granularity. The evaluation was conducted across moderate dataset scales of 4 million and 16 million images, benchmarking the model across standard industry tasks such as image-text retrieval, visual question answering, natural language visual reasoning, visual grounding, and image captioning.

The findings demonstrate clear performance and efficiency advantages. When trained on just 4 million images, the proposed model achieved an 80.4% top-1 text retrieval score on MSCOCO, outperforming larger models like VinVL and ALIGN that were trained on billions of data points. In visual reasoning, it surpassed previous state-of-the-art benchmarks on both VQA and NLVR2 while operating roughly ten times faster during inference than detector-based models. On visual grounding benchmarks, it exceeded specialized models by up to 4.5%, directly predicting target regions rather than ranking external region proposals. Ablation studies confirmed that removing either regional concepts or bounding box prediction significantly degraded overall performance.

These results show that multi-granularity alignment delivers substantial operational benefits, including reduced computational overhead, smaller model parameter footprints (216 million parameters), and faster inference times. These characteristics lower deployment costs and energy consumption while maintaining superior accuracy. The source suggests scaling the pre-training data further and exploring different visual backbones, noting that the model's high grounding precision makes it particularly well-suited for fine-grained assistive visual applications. However, organizations should account for data quality constraints, as performance depends on dense annotations during pre-training and downstream validation across different domains.

arXiv: 2111.08276
Cover for Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts

Abstract

Most existing methods in vision language pre-training rely on object-centric features extracted through object detection and make fine-grained alignments between the extracted features and texts. It is challenging for these methods to learn relations among multiple objects. To this end, we propose a new method called X-VLM 1 to perform ‘multi-grained vision language pre-training.’ The key to learning multi-grained alignments is to locate visual concepts in the image given the associated texts, and in the meantime align the texts with the visual concepts, where the alignments are in multi-granularity. Experimental results show that X-VLM effectively leverages the learned multi-grained alignments to many downstream vision language tasks and consistently outperforms state-of-the-art methods.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Overview
  • 3.2. Vision Encoding
  • 3.3. Cross-Modal Modeling
  • 4. Experiment
  • 4.1. Pre-training Datasets
  • 4.2. Implementation Details
  • 4.3. Downstream Tasks
  • 4.4. Results on Image-Text Retrieval
  • 4.5. Results on Visual Reasoning
  • 4.6. Results on Visual Grounding
  • 4.7. Results on Image Captioning
  • 4.8. Ablation Study
  • 5. Conclusion and Discussion
  • Acknowledgements
  • References
  • A. Appendix
  • A.1. Statistics of Object and Region Annotations
  • A.2. Implementation Details of Downstream Tasks
  • A.3. Zero-Shot Image-Text Retrieval Results
  • A.4. Case Study

Citation

MLA
Zeng, Y., et al. “Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts”. International Conference on Machine Learning, vol. 162, 2022, pp. 25994–6009, https://proceedings.mlr.press/v162/zeng22c.html.
APA
Zeng, Y., Zhang, X., & Li, H. (2022). Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts. International Conference on Machine Learning, 162, 25994–26009. https://proceedings.mlr.press/v162/zeng22c.html
Chicago
Zeng, Y., X. Zhang, and H. Li. 2022. “Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts”. International Conference on Machine Learning 162: 25994–26009. https://proceedings.mlr.press/v162/zeng22c.html.
Harvard
Zeng, Y., Zhang, X. and Li, H. (2022) “Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts”, International Conference on Machine Learning. PMLR, pp. 25994–26009. Available at: https://proceedings.mlr.press/v162/zeng22c.html.
Vancouver
1. Zeng Y, Zhang X, Li H (2022) Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts. In: International Conference on Machine Learning. PMLR, pp 25994–26009

BibTeX

@InProceedings{pmlr-v162-zeng22c,
  title = 	 {Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts},
  author =       {Zeng, Yan and Zhang, Xinsong and Li, Hang},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {25994--26009},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/zeng22c/zeng22c.pdf},
  url = 	 {https://proceedings.mlr.press/v162/zeng22c.html},
  abstract = 	 {Most existing methods in vision language pre-training rely on object-centric features extracted through object detection and make fine-grained alignments between the extracted features and texts. It is challenging for these methods to learn relations among multiple objects. To this end, we propose a new method called X-VLM to perform ‘multi-grained vision language pre-training.’ The key to learning multi-grained alignments is to locate visual concepts in the image given the associated texts, and in the meantime align the texts with the visual concepts, where the alignments are in multi-granularity. Experimental results show that X-VLM effectively leverages the learned multi-grained alignments to many downstream vision language tasks and consistently outperforms state-of-the-art methods.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/