VCoder: Versatile Vision Encoders for Multimodal Large Language Models

Jitesh JainJianwei YangHumphrey Shi

article2024CVPR60 citations

Presents VCoder, a framework that integrates auxiliary perception modalities like segmentation and depth maps into multimodal large language models to resolve chronic errors in object identification, counting, and spatial hallucination.

Listen

Modern multimodal artificial intelligence systems perform exceptionally well at complex tasks such as visual question-answering and image captioning. However, they frequently fail at fundamental visual perception tasks, including identifying background objects, estimating depth order, and accurately counting items in cluttered scenes. This occurs largely because standard vision models focus predominantly on salient foreground objects described in caption datasets, leaving systems prone to significant hallucinations and basic perceptual errors.

The article demonstrates that augmenting Multimodal Large Language Models with dedicated perception inputs substantially improves their object identification, counting, and spatial ordering capabilities. Specifically, the article evaluates a new framework that feeds auxiliary perception signals directly into the language model architecture without disrupting its underlying reasoning capabilities.

To achieve this, the researchers developed an adapter framework called Versatile Vision Encoders, or VCoder, integrated into the open-source LLaVA-1.5 model. They constructed the COCO Segmentation Text dataset, which uses 280,000 training images paired with segmentation and depth maps from specialized off-the-shelf vision models to generate question-and-answer pairs covering semantic, instance, and panoptic object identification. The framework encodes these extra perception maps into distinct tokens, training only lightweight projection layers while freezing the base language and image components. The article also introduced standardized evaluation metrics: Count Score, Hallucination Score, and Depth Score.

The experimental findings show that VCoder significantly outperforms leading open-source models and proprietary systems like GPT-4V across all perception metrics. On panoptic object identification, VCoder achieved a Count Score of approximately 86% to 87% and reduced the Hallucination Score to around 11% to 12%, compared to GPT-4V, which achieved a Count Score of only 38.4% and an 83.0% Hallucination Score. Existing open-source models performed even lower, frequently recording Count Scores below 40%. Furthermore, for object depth ordering, VCoder reduced the Depth Score error from over 166 down to roughly 63 to 66.

These results indicate that relying purely on standard visual encoders introduces substantial operational risks for applications requiring precise inventory counting, spatial awareness, or environment tracking. Feeding dedicated perceptual representations directly into language models provides a computationally efficient path to eliminate object hallucinations without requiring expensive full-model retraining.

Organizations developing or deploying multimodal vision-language systems should incorporate dedicated perception adapters like VCoder rather than relying solely on raw image inputs for perception-critical tasks. Future technical efforts should focus on expanding the training dataset to broader, open-vocabulary object categories beyond standard closed sets and refining scoring metrics to handle vocabulary synonyms more flexibly without manual mapping.

arXiv: 2312.14233
Cover for VCoder: Versatile Vision Encoders for Multimodal Large Language Models

Abstract

Humans possess the remarkable skill of Visual Perception, the ability to see and understand the seen, helping them make sense of the visual world and, in turn, reason. Multimodal Large Language Models (MLLM) have recently achieved impressive performance on vision-language tasks ranging from visual question-answering and image captioning to visual reasoning and image generation. However, when prompted to identify or count (perceive) the entities in a given image, existing MLLM systems fail. Working towards developing an accurate MLLM system for perception and reasoning, we propose using Versatile vision encoders (VCoder) as perception eyes for Multimodal LLMs. We feed the VCoder with perception modalities such as segmentation or depth maps, improving the MLLM's perception abilities. Secondly, we leverage the images from COCO and outputs from off-the-shelf vision perception models to create our COCO Segmentation Text (COST) dataset for training and evaluating MLLMs on the object perception task. Thirdly, we introduce metrics to assess the object perception abilities in MLLMs on our COST dataset. Lastly, we provide extensive experimental evidence proving the VCoder's improved object-level perception skills over existing Multimodal LLMs, including GPT-4V. We open-source our dataset, code, and models to promote research.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Visual Perception
  • 2.2. Visual Understanding with LLMs
  • 2.3. Perception Hallucination in MLLMs
  • 3. Object Identification with MLLMs
  • 3.1. COST to Identify Objects with MLLMs
  • 3.2. VCoder for Multimodal LLMs
  • 3.3. Evaluating MLLMs for Object Identification
  • 4. Experiments
  • 4.1. Implementation Details
  • 4.2. Main Results
  • 5. Object Order Perception with MLLMs
  • 6. Limitations
  • 7. Conclusion
  • References

Knowls

  1. Knowl 1 — VCoder adapter for perception-conditioned MLLMs

    model/method

    VCoder is an adapter that adds visual perception modalities to a multimodal large language model (MLLM). In the paper’s implementation, the base model is LLaVA-1.5, which contains an RGB-image encoder (ImCoder), a two-layer MLP projector, and a language model. A segmentation map or depth map is processed by a modality-specific CLIP ViT encoder and a trainable two-layer MLP projector; these components form the VCoder.

    The RGB image, perception-control inputs, and user question are converted into embeddings with modality-specific tokens such as <img>, <seg>, <depth>, and <query>. The resulting embeddings are concatenated and passed to the language model. For object identification, the model uses a segmentation map as the additional control input; for object-order perception, it uses both a segmentation map and a depth map. During training, the original LLaVA-1.5 parameters are frozen and only the VCoder MLP projectors are optimized, preserving the base model’s existing reasoning behavior while adding object-perception information. The architecture diagram on page 5 shows this frozen-parameter adapter design.

  2. Knowl 2 — COST dataset for object identification and counting

    data/table

    COCO Segmentation Text (COST) is a question-answer dataset constructed to train and evaluate MLLMs on object-level perception. The dataset uses COCO images, segmentation outputs from OneFormer, question templates generated by GPT-4, and textual answers containing object names and counts. It covers three segmentation-derived tasks: semantic object identification, instance object identification, and panoptic object identification.

    For each image, the authors extract object classes and counts from OneFormer’s segmentation output and serialize them as an answer of the form “The objects present in the image are: [count] [object], ...”. A count is written only when it is greater than one. GPT-4 was prompted to produce 20 questions for each of the semantic, instance, and panoptic question buckets. The training split contains 280,000 images from COCO’s train2017, test2017, and unlabeled2017 splits, while the validation split contains the 5,000 images in val2017. The resulting data are intended to expose MLLMs to both salient foreground objects and less-salient background objects, including their counts.

  3. Knowl 3 — Count and hallucination scores for object perception

    equation

    The paper evaluates object identification by extracting object nouns and their counts from a ground-truth answer and an MLLM prediction. Let G={(oiG,ciG)}i=1NG=\{(o_i^G,c_i^G)\}_{i=1}^{N} be the ground-truth dictionary, where oiGo_i^G is a ground-truth object noun and ciG∈N>0c_i^G\in\mathbb{N}_{>0} is its count. Let P={(ojP,cjP)}j=1MP=\{(o_j^P,c_j^P)\}_{j=1}^{M} be the corresponding prediction dictionary. Let I(o,D)I(o,D) equal 11 if object noun oo is a key in dictionary DD, and 00 otherwise.

    The count score (CS) measures how accurately the predicted counts match ground-truth counts, averaged over ground-truth object nouns:

    CS=100N∑i=1N{min⁡(ciG,ciP)max⁡(ciG,ciP),I(oiG,P)=1,0,I(oiG,P)=0.\mathrm{CS}=\frac{100}{N}\sum_{i=1}^{N} \begin{cases} \dfrac{\min(c_i^G,c_i^P)}{\max(c_i^G,c_i^P)}, & I(o_i^G,P)=1,\\ 0, & I(o_i^G,P)=0. \end{cases}

    Here ciPc_i^P denotes the predicted count for the matched noun oiGo_i^G. Higher CS is better.

    The hallucination score (HS) measures extra or incorrectly counted predicted objects, averaged over predicted object nouns:

    HS=100M∑j=1M{1−min⁡(cjP,cjG)max⁡(cjP,cjG),I(ojP,G)=1,1,I(ojP,G)=0.\mathrm{HS}=\frac{100}{M}\sum_{j=1}^{M} \begin{cases} 1-\dfrac{\min(c_j^P,c_j^G)}{\max(c_j^P,c_j^G)}, & I(o_j^P,G)=1,\\ 1, & I(o_j^P,G)=0. \end{cases}

    Here cjGc_j^G denotes the ground-truth count for the matched noun ojPo_j^P. Lower HS is better. Before matching, the evaluation manually maps synonymous categories to a common COCO category, such as mapping “man,” “woman,” “boy,” “girl,” and “child” to “person.”

  4. Knowl 4 — Object-identification results on COST

    data/table

    The page-7 comparison table evaluates semantic, instance, and panoptic object identification on the COST validation set. CS is reported as a percentage and should be maximized; HS is a percentage and should be minimized. The comparison includes off-the-shelf MLLMs, models trained on COST, an RGB-control baseline, and VCoder models.

    Method Semantic Instance Panoptic
    CS HS CS HS CS HS
    GPT-4V – – – – 38.4 83.0
    MiniGPT-4 LLaMA-2-7b 6.2 92.2 5.6 97.7 6.2 94.9
    InstructBLIP Vicuna-7b 14.2 85.8 25.3 91.9 17.5 91.2
    LLaVA-1.5-7b 30.6 60.1 50.3 75.9 38.7 67.3
    LLaVA-1.5-13b 25.0 69.3 49.9 75.0 35.8 68.6
    CogVLM-17b 33.4 67.5 43.5 86.2 40.6 75.9
    COST IT LLaVA-1.5-7b 78.7 22.1 67.5 30.3 71.9 28.2
    Soft-Prompted LLaVA-1.5-7b 36.2 56.7 18.4 72.2 26.8 63.0
    ImCoder LLaVA-1.5-7b 78.9 22.7 64.0 29.4 70.8 27.9
    VCoder LLaVA-1.5-7b 88.6 10.4 71.1 26.9 86.0 12.8
    VCoder LLaVA-1.5-13b 89.0 10.0 73.3 25.0 87.2 11.6

    VCoder LLaVA-1.5-13b is the strongest model in every reported task, reaching CS values of 89.0, 73.3, and 87.2 and HS values of 10.0, 25.0, and 11.6 for semantic, instance, and panoptic identification, respectively. The VCoder models outperform both the original MLLMs and the COST-trained baselines. The RGB-image control baseline, ImCoder, performs below the segmentation-control VCoder, supporting the paper’s claim that segmentation provides more useful object-level control information than an additional RGB input. GPT-4V is evaluated only for panoptic identification and is substantially below both VCoder variants.

  5. Knowl 5 — Depth-conditioned object-order perception

    model/method

    COST is extended to object-order perception by adding depth information. The authors obtain monocular depth maps for COCO images with the publicly available DINOv2 DPT model and use OneFormer panoptic masks to aggregate depth values over each object. Objects are then ordered from foreground to background and converted into answers of the form “The depth order for objects present in the image is: [object 1], [object 2], ...”.

    When multiple objects belong to the same class, later instances receive suffixes such as person-2 so that their relative positions remain distinguishable. The resulting VCoder-DS LLaVA-1.5 receives both <depth> and <seg> control inputs. Its training mixture contains the COST object-identification and object-order data, plus approximately 200,000 randomly selected LLaVA-1.5 instruction-tuning image-conversation pairs with corresponding OneFormer segmentation maps.

  6. Knowl 6 — Depth score and object-order results

    data/table

    For object-order perception, the paper defines a depth score (DS) by comparing the ground-truth and predicted positions of objects in the foreground-to-background ordering. The score uses the absolute difference between the ground-truth rank and predicted rank for each object; lower DS indicates more accurate ordering.

    The page-8 results table reports the following COST validation performance:

    Method Depth Score
    LLaVA-1.5-7b 166.1
    LLaVA-1.5-13b 227.2
    VCoder-DS LLaVA-1.5-7b 65.9
    VCoder-DS LLaVA-1.5-13b 63.3

    Adding segmentation and depth control inputs through VCoder reduces the depth score from 166.1 to 65.9 for the 7-billion-parameter model and from 227.2 to 63.3 for the 13-billion-parameter model. Thus, the VCoder-DS models substantially improve foreground-to-background object ordering over their corresponding LLaVA-1.5 baselines.

  7. Knowl 7 — Training and evaluation configuration

    experimental setup

    The object-identification VCoder models use CLIP-ViT-L-336px encoders for the RGB image and segmentation control input, with modality-specific two-layer MLP projectors. Visual inputs are resized to 336×336336\times336 pixels, producing 576 visual tokens. The segmentation maps used for training and inference are generated by OneFormer with a DiNAT-L backbone trained on COCO.

    The VCoder-adapted LLaVA-1.5 models are trained for two epochs on COST with batch size 256 and learning rate 10−310^{-3}; the original LLaVA-1.5 parameters remain frozen. Semantic, instance, and panoptic identification examples are sampled uniformly. Training takes approximately 8 hours for the 7-billion-parameter model and 14 hours for the 13-billion-parameter model on eight A100 GPUs.

    All models are evaluated on the COST validation split with separately sampled questions for semantic, instance, and panoptic identification. The evaluation prompt requests a paragraph beginning with “The objects present in the image are” and requires counts to be written in words when greater than one. The depth-conditioned model is trained for one epoch with the same general hyperparameter settings, using DINOv2 ViT-L/14 DPT depth maps trained on NYUv2 together with OneFormer segmentation maps.

  8. Knowl 8 — Evidence that segmentation control and COST training are complementary

    empirical result

    The experiments show that training on COST alone improves object perception but does not match the performance of VCoder. COST IT LLaVA-1.5-7b reaches semantic, instance, and panoptic CS values of 78.7, 67.5, and 71.9, whereas the segmentation-controlled VCoder LLaVA-1.5-7b reaches 88.6, 71.1, and 86.0. The corresponding HS values decrease from 22.1, 30.3, and 28.2 to 10.4, 26.9, and 12.8.

    The RGB-control ImCoder baseline is also weaker than segmentation-controlled VCoder: its panoptic CS is 70.8 with HS 27.9, compared with 86.0 and 12.8 for the 7-billion-parameter VCoder. The authors interpret these results as evidence that both perception-focused instruction data and explicit segmentation representations are important. Off-the-shelf models perform relatively better on instance identification than on semantic or panoptic identification, consistent with stronger handling of salient objects than background objects.

  9. Knowl 9 — VCoder supports multiple perception modalities

    model/method

    VCoder is designed as a modality-extensible interface rather than a segmentation-only module. Each control modality has its own encoder/projector pathway and its own special input token; segmentation and depth are represented by <seg> and <depth>. The projected modality embeddings are concatenated with the RGB-image and text embeddings before entering the language model.

    In the paper, segmentation control is used for object identification, while segmentation plus depth control is used for object-order prediction. The same design is stated to support additional dense or structured perception inputs, such as keypoint maps, by adding a modality-specific encoder/projector and token. This allows the control modality to be selected according to the perception task.

  10. Knowl 10 — Limitations of COST and its evaluation metrics

    limitation

    COST inherits the category limitations of OneFormer, whose segmentation model is trained with a closed-set COCO vocabulary. Consequently, the dataset covers only a limited set of object categories and does not yet provide a broad real-world benchmark with many classes and varying levels of category granularity.

    The count score, hallucination score, and depth score rely on one-to-one word matching between model responses and reference answers. This requires manually defining mappings between synonyms, such as mapping different human descriptors to “person.” The paper identifies replacing these manually specified synonym mappings with a more robust matching procedure as an unresolved limitation.

Coverage note — No substantial contributed material was omitted; background, related work, references, acknowledgements, and the conclusion’s restatement of the contributions were excluded.

References

  1. 1.Aishwarya Agrawal, Jiasen Lu, Stanislaw Antol, Margaret Mitchell, C. Lawrence Zitnick, Dhruv Batra, and Devi Parikh. Vqa: Visual question answering, 2015. 2
  2. 2.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karen Simonyan. Flamingo: a visual language model for few-shot learning. In NeurIPS, 2022. 1, 3
  3. 3.Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, Alessandro Cappelli, Ruxandra Cojocaru, Merouane Debbah, Etienne Goffinet, Daniel Heslow, Julien Launay, Quentin Malartic, Badreddine Noune, Baptiste Pannier, and Guilherme Penedo. Falcon-40B: an open large language model with state-of-the-art performance. arXiv, 2023. 1
  4. 4.Anas Awadalla, Irena Gao, Josh Gardner, Jack Hessel, Yusuf Hanafy, Wanrong Zhu, Kalyani Marathe, Yonatan Bitton, Samir Gadre, Shiori Sagawa, Jenia Jitsev, Simon Kornblith, Pang Wei Koh, Gabriel Ilharco, Mitchell Wortsman, and Ludwig Schmidt. Openflamingo: An open-source framework for training large autoregressive vision-language models. arXiv, 2023. 1, 3
  5. 5.Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, Binyuan Hui, Luo Ji, Mei Li, Junyang Lin, Runji Lin, Dayiheng Liu, Gao Liu, Chengqiang Lu, Keming Lu, Jianxin Ma, Rui Men, Xingzhang Ren, Xuancheng Ren, Chuanqi Tan, Sinan Tan, Jianhong Tu, Peng Wang, Shijie Wang, Wei Wang, Shengguang Wu, Benfeng Xu, Jin Xu, An Yang, Hao Yang, Jian Yang, Shusheng Yang, Yang Yao, Bowen Yu, Hongyi Yuan, Zheng Yuan, Jianwei Zhang, Xingxuan Zhang, Yichang Zhang, Zhenru Zhang, Chang Zhou, Jingren Zhou, Xiaohuan Zhou, and Tianhang Zhu. Qwen technical report. arXiv, 2023. 3
  6. 6.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In ECCV, 2020. 3
  7. 7.Soravit Changpinyo, Piyush Sharma, Nan Ding, and Radu Soricut. Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. In CVPR, 2021. 2
  8. 8.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L. Yuille. Semantic image segmentation with deep convolutional nets and fully connected crfs. In ICLR, 2015. 3
  9. 9.Bowen Cheng, Maxwell D Collins, Yukun Zhu, Ting Liu, Thomas S Huang, Hartwig Adam, and Liang-Chieh Chen. Panoptic-deeplab: A simple, strong, and fast baseline for bottom-up panoptic segmentation. In CVPR, 2020. 3
  10. 10.Bowen Cheng, Ishan Misra, Alexander G. Schwing, Alexander Kirillov, and Rohit Girdhar. Masked-attention mask transformer for universal image segmentation. In CVPR, 2022. 3, 4
  11. 11.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, 2023. 1, 3
  12. 12.OpenCompass Contributors. Opencompass: A universal evaluation platform for foundation models. https://github.com/open-compass/opencompass, 2023. 3
  13. 13.Timothy F Cootes, Gareth J Edwards, and Christopher J Taylor. Active appearance models. IEEE Transactions on pattern analysis and machine intelligence, 23(6):681–685, 2001. 3
  14. 14.Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023. 1, 2, 3, 7
  15. 15.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2021. 3, 5, 6
  16. 16.David Eigen, Christian Puhrsch, and Rob Fergus. Depth map prediction from a single image using a multi-scale deep network. In Advances in neural information processing systems, pages 2366–2374, 2014. 3
  17. 17.Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Zhenyu Qiu, Wei Lin, Jinrui Yang, Xiawu Zheng, Ke Li, Xing Sun, and Rongrong Ji. Mme: A comprehensive evaluation benchmark for multimodal large language models. arXiv, 2023. 3
  18. 18.Rohit Girdhar, Mannat Singh, Nikhila Ravi, Laurens van der Maaten, Armand Joulin, and Ishan Misra. Omnivore: A Single Model for Many Visual Modalities. In CVPR, 2022. 3
  19. 19.Ali Hassani and Humphrey Shi. Dilated neighborhood attention transformer. arXiv:2209.15001, 2022. 7
  20. 20.Ali Hassani, Steven Walton, Jiachen Li, Shen Li, and Humphrey Shi. Neighborhood attention transformer. In CVPR, 2023. 7
  21. 21.Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick. Mask r-cnn. In ICCV, 2017. 3
  22. 22.Shaohan Huang, Li Dong, Wenhui Wang, Yaru Hao, Saksham Singhal, Shuming Ma, Tengchao Lv, Lei Cui, Owais Khan Mohammed, Qiang Liu, Kriti Aggarwal, Zewen Chi, Johan Bjorck, Vishrav Chaudhary, Subhojit Som, Xia Song, and Furu Wei. Language is not all you need: Aligning perception with language models. ArXiv, abs/2302.14045, 2023. 1, 3
  23. 23.Drew A Hudson and Christopher D Manning. Gqa: A new dataset for real-world visual reasoning and compositional question answering. CVPR, 2019. 3
  24. 24.Jitesh Jain, Anukriti Singh, Nikita Orlov, Zilong Huang, Jiachen Li, Steven Walton, and Humphrey Shi. Semask: Semantically masking transformer backbones for effective semantic segmentation. arXiv, 2021. 3
  25. 25.Jitesh Jain, Jiachen Li, MangTik Chiu, Ali Hassani, Nikita Orlov, and Humphrey Shi. OneFormer: One Transformer to Rule Universal Image Segmentation. In CVPR, 2023. 2, 3, 4, 5, 6, 7, 8
  26. 26.Yang Jin, Kun Xu, Kun Xu, Liwei Chen, Chao Liao, Jianchao Tan, Yadong Mu, et al. Unified language-vision pre-training in llm with dynamic discrete visual tokenization. arXiv, 2023. 1
  27. 27.Jing Yu Koh, Daniel Fried, and Ruslan Salakhutdinov. Generating images with multimodal language models. NeurIPS, 2023. 1
  28. 28.Deanna Kuhn. The Skills of Argument. Cambridge University Press, 1991. 1
  29. 29.Iro Laina, Christian Rupprecht, Vasileios Belagiannis, Federico Tombari, and Nassir Navab. Deeper depth prediction with fully convolutional residual networks. In 2016 Fourth International Conference on 3D Vision (3DV), pages 239–248. IEEE, 2016. 3
  30. 30.Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE. 3
  31. 31.Bo Li, Yuanhan Zhang, Liangyu Chen, Jinghao Wang, Fanyi Pu, Jingkang Yang, Chunyuan Li, and Ziwei Liu. Mimic-it: Multi-modal in-context instruction tuning. 2023. 3
  32. 32.Bo Li, Yuanhan Zhang, Liangyu Chen, Jinghao Wang, Jingkang Yang, and Ziwei Liu. Otter: A multi-modal model with in-context instruction tuning. arXiv preprint arXiv:2305.03726, 2023. 3
  33. 33.Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In ICML, 2022. 2
  34. 34.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In ICML, 2023. 2, 3
  35. 35.Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji-Rong Wen. Evaluating object hallucination in large vision-language models. In EMNLP, 2023. 3, 6
  36. 36.Tsung-Yi Lin, Michael Maire, Serge Belongie, Lubomir Bourdev, Ross Girshick, James Hays, Pietro Perona, Deva Ramanan, C. Lawrence Zitnick, and Piotr Dollar. Microsoft coco: Common objects in context. In ECCV, 2014. 2, 4, 5, 6, 7, 8
  37. 37.Fuxiao Liu, Tianrui Guan, Zongxia Li, Lichang Chen, Yaser Yacoob, Dinesh Manocha, and Tianyi Zhou. Hallusionbench: You see what you think? or you think what you see? an image-context reasoning benchmark challenging for gpt-4v (ision), llava-1.5, and other multi-modality models. arXiv preprint arXiv:2310.14566, 2023. 3
  38. 38.Fuxiao Liu, Kevin Lin, Linjie Li, Jianfeng Wang, Yaser Yacoob, and Lijuan Wang. Aligning large multi-modal model with robust instruction tuning. arXiv preprint arXiv:2306.14565, 2023. 3
  39. 39.Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning, 2023. 1, 2, 3, 5, 6, 7, 8
  40. 40.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. In NeurIPS, 2023. 1, 3
  41. 41.Jonathan Long, Evan Shelhamer, and Trevor Darrell. Fully convolutional networks for semantic segmentation. In CVPR, 2015. 3
  42. 42.Holy Lovenia, Wenliang Dai, Samuel Cahyawijaya, Ziwei Ji, and Pascale Fung. Negative object presence evaluation (nope) to measure object hallucination in vision-language models. arXiv, 2023. 6
  43. 43.Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh. Neural baby talk. In CVPR, 2018. 6
  44. 44.H. Moravec. Mind children: The future of robot and human intelligence. Harvard University Press, 1988. 4
  45. 45.Chong Mou, Xintao Wang, Liangbin Xie, Yanze Wu, Jian Zhang, Zhongang Qi, Ying Shan, and Xiaohu Qie. T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models. arXiv, 2023. 3, 5
  46. 46.Pushmeet Kohli Nathan Silberman, Derek Hoiem and Rob Fergus. Indoor segmentation and support inference from rgbd images. In ECCV, 2012. 7
  47. 47.OpenAI. Chatgpt. https://chat.openai.com/, 2022. 1
  48. 48.OpenAI. Gpt-4 technical report, 2023. 1, 2, 3, 4, 5, 7, 8
  49. 49.Maxime Oquab, Timothee Darcet, Theo Moutakanni, Huy V. Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Russell Howes, Po-Yao Huang, Hu Xu, Vasu Sharma, Shang-Wen Li, Wojciech Galuba, Mike Rabbat, Mido Assran, Nicolas Ballas, Gabriel Synnaeve, Ishan Misra, Herve Jegou, Julien Mairal, Patrick Labatut, Armand Joulin, and Piotr Bojanowski. Dinov2: Learning robust visual features without supervision, 2023. 2, 4, 5, 7
  50. 50.Vicente Ordonez, Girish Kulkarni, and Tamara Berg. Im2text: Describing images using 1 million captioned photographs. In NeurIPS, 2011. 2
  51. 51.Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao, Shaohan Huang, Shuming Ma, and Furu Wei. Kosmos-2: Grounding multimodal large language models to the world. ArXiv, abs/2306, 2023. 1, 3
  52. 52.Alec Radford, Jong Wook Kim, Chris Hallacy, A. Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In ICML, 2021. 2, 3
  53. 53.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. arXiv, 2021. 3, 5, 6
  54. 54.Rene Ranftl, Alexey Bochkovskiy, and Vladlen Koltun. Vision transformers for dense prediction. In ICCV, 2021. 2, 3, 4, 5, 7
  55. 55.Anna Rohrbach, Lisa Anne Hendricks, Kaylee Burns, Trevor Darrell, and Kate Saenko. Object hallucination in image captioning. In EMNLP, 2018. 3, 6
  56. 56.Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. Laion-400m: Open dataset of clip-filtered 400 million image-text pairs. In NeurIPS Workshops 2021, 2021. 2
  57. 57.Quan Sun, Qiying Yu, Yufeng Cui, Fan Zhang, Xiaosong Zhang, Yueze Wang, Hongcheng Gao, Jingjing Liu, Tiejun Huang, and Xinlong Wang. Generative pretraining in multi-modality. arXiv, 2023. 1
  58. 58.Alexander Toshev and Christian Szegedy. Deeppose: Human pose estimation via deep neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1653–1660, 2014. 3
  59. 59.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothee Lacroix, Baptiste Roziere, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. Llama: Open and efficient foundation language models. arXiv, 2023. 3
  60. 60.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. Llama 2: Open foundation and fine-tuned chat models. arXiv, 2023. 1, 3
  61. 61.Maria Tsimpoukelli, Jacob Menick, Serkan Cabi, S. M. Ali Eslami, Oriol Vinyals, and Felix Hill. Multimodal few-shot learning with frozen language models. In NeurIPS, 2021. 3
  62. 62.Z. Tu, Xiangrong Chen, Alan Yuille, and Song Zhu. Image parsing: Unifying segmentation, detection, and recognition. In IJCV, 2005. 3
  63. 63.Paul Viola and Michael Jones. Rapid object detection using a boosted cascade of simple features. In CVPR, 2001. 3
  64. 64.Weihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong, Ji Qi, Yan Wang, Junhui Ji, Zhuoyi Yang, Lei Zhao, Xixuan Song, Jiazheng Xu, Bin Xu, Juanzi Li, Yuxiao Dong, Ming Ding, and Jie Tang. Cogvlm: Visual expert for pretrained language models. arXiv, 2023. 7
  65. 65.Enze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar, Jose M. Alvarez, and Ping Luo. Segformer: Simple and efficient design for semantic segmentation with transformers. In NeurIPS, 2021. 3
  66. 66.Xingqian Xu, Jiayi Guo, Zhangyang Wang, Gao Huang, Irfan Essa, and Humphrey Shi. Prompt-free diffusion: Taking”text” out of text-to-image diffusion models. arXiv preprint arXiv:2305.16223, 2023. 3
  67. 67.Xingqian Xu, Zhangyang Wang, Eric Zhang, Kai Wang, and Humphrey Shi. Versatile diffusion: Text, images and variations all in one diffusion model. In ICCV, 2023. 3
  68. 68.Qinghao Ye, Haiyang Xu, Guohai Xu, Jiabo Ye, Ming Yan, Yiyang Zhou, Junyang Wang, Anwen Hu, Pengcheng Shi, Yaya Shi, Chaoya Jiang, Chenliang Li, Yuanhong Xu, Hehong Chen, Junfeng Tian, Qian Qi, Ji Zhang, and Fei Huang. mplug-owl: Modularization empowers large language models with multimodality, 2023. 1, 3
  69. 69.Yan Zeng, Hanbo Zhang, Jiani Zheng, Jiangnan Xia, Guoqiang Wei, Yang Wei, Yuchen Zhang, and Tao Kong. What matters in training a gpt4-style language model with multimodal inputs? arXiv preprint arXiv:2307.02469, 2023. 3
  70. 70.Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. In ICCV, 2023. 3, 5
  71. 71.Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, Longyue Wang, Anh Tuan Luu, Wei Bi, Freda Shi, and Shuming Shi. Siren’s song in the ai ocean: A survey on hallucination in large language models. arXiv, 2023. 3
  72. 72.Kaizhi Zheng, Xuehai He, and Xin Eric Wang. Minigpt-5: Interleaved vision-and-language generation via generative tokens. arXiv, 2023. 1
  73. 73.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric. P Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging llm-as-a-judge with mt-bench and chatbot arena, 2023. 6
  74. 74.Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv, 2023. 1, 3, 7
  75. 75.Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable detr: Deformable transformers for end-to-end object detection. arXiv, 2020. 3

Citation

MLA
Jain, J., et al. “VCoder: Versatile Vision Encoders for Multimodal Large Language Models”. arXiv, 2023, http://arxiv.org/abs/2312.14233v1.
APA
Jain, J., Yang, J., & Shi, H. (2023). VCoder: Versatile Vision Encoders for Multimodal Large Language Models. arXiv. http://arxiv.org/abs/2312.14233v1
Chicago
Jain, J., J. Yang, and H. Shi. 2023. “VCoder: Versatile Vision Encoders for Multimodal Large Language Models”. arXiv. http://arxiv.org/abs/2312.14233v1.
Harvard
Jain, J., Yang, J. and Shi, H. (2023) “VCoder: Versatile Vision Encoders for Multimodal Large Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2312.14233v1.
Vancouver
1. Jain J, Yang J, Shi H (2023) VCoder: Versatile Vision Encoders for Multimodal Large Language Models. arXiv

BibTeX

@article{jain2023vcoder,
  title = {VCoder: Versatile Vision Encoders for Multimodal Large Language Models},
  author = {Jain, Jitesh and Yang, Jianwei and Shi, Humphrey},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2312.14233v1},
  eprint = {2312.14233}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE