Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models

Siddharth KaramchetiSuraj NairAshwin BalakrishnaPercy LiangThomas KollarDorsa Sadigh

article2024ICML296 citations

Investigates key design decisions in vision-language models across visual encoders and training strategies, delivering a standardized evaluation framework and open-source models that outperform LLaVA-1.5 and InstructBLIP.

Listen

Visually-conditioned language models, which generate natural language responses from visual and textual inputs, are expanding rapidly across applications such as robotics, visual chat, and scene understanding. However, existing development practices rely on untested architectural and training conventions, while relying on subjective, model-based evaluation methods that obscure what drives actual performance.

The article systematically evaluates the primary design axes of these multimodal models—including optimization strategies, visual representations, language model selection, and dataset scaling—to establish an empirical foundation for efficient and effective model development.

To conduct this investigation, the authors developed an optimized, modular training framework alongside a standardized evaluation suite of twelve established objective benchmarks spanning visual question answering, object localization, and challenge sets probing spatial reasoning and hallucination. Using controlled single-variable comparisons at the 7-billion and 13-billion parameter scales, the study analyzed the isolated impact of each core design choice.

The analysis produced several key findings: First, eliminating the common multi-stage training pipeline in favor of direct, single-stage training improves aggregate performance while reducing training compute costs by 20% to 25%. Second, keeping the visual backbone frozen is critical; finetuning it during training severely degrades performance, particularly on localization tasks. Third, fusing complementary visual representations—specifically combining high-level contrastive features from SigLIP with low-level spatial features from DINOv2—yields significant 5% to 10% gains on localization and spatial reasoning benchmarks. Fourth, base language models perform comparably to instruction-tuned models while exhibiting less verbosity and lower hallucination rates, though including language-only safety data during training is essential to prevent harmful and biased outputs. Finally, models benefit significantly from training for two full epochs rather than one, and scaling dataset diversity improves performance more effectively than raw volume.

These findings demonstrate that organizations can achieve superior multimodal performance with substantially lower compute budgets and simplified engineering workflows. Conventional practices, such as complex multi-stage alignment and full model finetuning, add unnecessary cost and operational risk while degrading downstream accuracy. By consolidating these design principles, the authors introduced the PRISM model family, which consistently outperforms leading open-source models such as LLaVa v1.5 and InstructBLIP across all twelve benchmark tasks.

Technical leaders and engineering teams should adopt single-stage training workflows, freeze visual backbones, deploy fused visual representations, and allocate sufficient training time (two epochs) on diverse data mixtures. Teams utilizing base language models must retain language-only safety co-training data to mitigate toxic or biased generations. Further work should explore architectural methods to prevent visual representation collapse during full finetuning and investigate optimal pretraining mixtures for downstream multimodal integration.

Confidence in these findings is high for standard autoregressive architectures at the 7-billion to 13-billion parameter scale. However, decision-makers should exercise caution when extrapolating these conclusions to alternative architectures (such as resampler-based models), substantially larger scales (such as 70-billion-plus parameters), or extended conversational contexts that exceed standard single-turn objective benchmarks.

Cover for Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models

Abstract

Visually-conditioned language models (VLMs) have seen growing adoption in applications such as visual dialogue, scene understanding, and robotic task planning; adoption that has fueled a wealth of new models such as LLaVa, InstructBLIP, and PaLI-3. Despite the volume of new releases, key design decisions around image preprocessing, architecture, and optimization are under-explored, making it challenging to understand what factors account for model performance −- a challenge further complicated by the lack of objective, consistent evaluations. To address these gaps, we first compile a suite of standardized evaluations spanning visual question answering, object localization, and challenge sets that probe properties such as hallucination; evaluations that provide fine-grained insight VLM capabilities. Second, we rigorously investigate VLMs along key design axes, including pretrained visual representations and training from base vs. instruct-tuned language models, amongst others. We couple our analysis with three resource contributions: (1) a unified framework for evaluating VLMs, (2) optimized, flexible training code, and (3) checkpoints for all models, including a family of VLMs at the 7-13B scale that strictly outperform InstructBLIP and LLaVa v1.5, the state-of-the-art in open VLMs.

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 3 Evaluation Suite
  • 4 Experiments – Investigating Design Axes
  • 4.1 Optimization Procedure
  • 4.2 Image Processing & Visual Representations
  • 4.3 Integrating Language Models
  • 4.4 Scaling Properties: Training Time & Data
  • 5 Prism – Distilling Key Insights
  • 6 Limitations & Future Work
  • 7 Conclusion
  • References
  • A Training Visually-Conditioned Language Models
  • A.1 Pretraining Dataset Composition
  • A.2 Implementation – Architecture Components & Optimization
  • A.3 Training Hyperparameters
  • B Evaluation Protocol
  • B.1 Evaluation Procedures
  • B.2 Comparing Model Performance – Significance Testing
  • B.3 Exhaustive Results

Knowls

  1. Knowl 1 — Prismatic Visually-Conditioned Language Model Architecture and Training Objective

    model/method

    The Prismatic visually-conditioned language model (VLM) architecture connects a pretrained visual backbone to an autoregressive language model using a lightweight projector.

    Let ximg∈RH×W×Cx_{\text{img}} \in \mathbb{R}^{H \times W \times C} denote an input image and let uprompt=(u1,…,uK)u_{\text{prompt}} = (u_1, \dots, u_K) denote a sequence of text prompt tokens. The model comprises three primary components:

    1. Visual Representation Backbone (VωV_\omega): A Vision Transformer (ViT) with parameters ω\omega processes ximgx_{\text{img}} and extracts patch features from its penultimate layer:

    pimg=Vω(ximg)∈RL×hvisionp_{\text{img}} = V_\omega(x_{\text{img}}) \in \mathbb{R}^{L \times h_{\text{vision}}}

    where LL is the number of visual patch tokens and hvisionh_{\text{vision}} is the hidden dimension of the visual backbone.

    1. Vision-Language Projector (FψF_\psi): A 2-layer Multi-Layer Perceptron (MLP) with GELU activation functions and parameters ψ\psi projects each visual patch representation independently into the language model's embedding space:

    eimg=Fψ(pimg)∈RL×htexte_{\text{img}} = F_\psi(p_{\text{img}}) \in \mathbb{R}^{L \times h_{\text{text}}}

    where htexth_{\text{text}} is the embedding dimension of the language model.

    1. Language Model (LMθ\text{LM}_\theta): An autoregressive language model with parameters θ\theta. The projected visual embeddings eimge_{\text{img}} are prepended to the prompt embeddings eprompt=embed(uprompt)∈RK×htexte_{\text{prompt}} = \text{embed}(u_{\text{prompt}}) \in \mathbb{R}^{K \times h_{\text{text}}} via sequence-wise concatenation:

    E=[eimg;eprompt]∈R(L+K)×htextE = [e_{\text{img}}; e_{\text{prompt}}] \in \mathbb{R}^{(L + K) \times h_{\text{text}}}

    The language model autoregressively generates output tokens ugen=LMθ(E)u_{\text{gen}} = \text{LM}_\theta(E).

    Given a ground-truth target sequence u^gen=(u^1,…,u^T)\hat{u}_{\text{gen}} = (\hat{u}_1, \dots, \hat{u}_T), training minimizes the negative log-likelihood loss via gradient descent:

    L(ω,ψ,θ)=−∑t=1Tlog⁡pθ(u^t∣ximg,uprompt,u^<t)\mathcal{L}(\omega, \psi, \theta) = -\sum_{t=1}^T \log p_\theta(\hat{u}_t \mid x_{\text{img}}, u_{\text{prompt}}, \hat{u}_{<t})

  2. Knowl 2 — PRISM VLM Recipe and Model Family

    model/method

    The PRISM family of VLMs integrates the best-performing design choices identified across optimization, vision encoders, image processing, language modeling, and scaling axes into a unified training recipe:

    1. Optimization: Single-stage training directly on multimodal instruction data without an isolated projector pretraining stage, while keeping the visual encoder weights frozen.
    2. Visual Representation: A fused visual backbone combining DINOv2 ViT-L/14 (self-supervised) and SigLIP ViT-SO/14 (vision-language contrastive). Patch features from both backbones are concatenated along the channel dimension (hvision=1024+1152=2176h_{\text{vision}} = 1024 + 1152 = 2176) before passing into a 2-layer GELU MLP projector.
    3. Image Preprocessing: Naive image resizing (warping arbitrary aspect ratios directly to fixed square resolutions of 384×384384 \times 384 pixels) rather than cropping or letterbox padding.
    4. Language Model Backbone: A base language model (Llama-2 7B or 13B) trained using the prompt template <s> In: {Input} Out: {Response} </s> without system prompts.
    5. Safety Alignment: Co-training on 40K language-only conversations from ShareGPT alongside multimodal instruction data to establish conversational safeguards against toxic and biased generations.
    6. Data & Scaling: Training for 2 epochs on an expanded multimodal instruction mixture that incorporates the base LLaVa v1.5 mixture, LRV-Instruct, and LVIS-Instruct-4V.

    Under identical data and compute budgets (PRISM Controlled), this recipe outperforms LLaVa v1.5 across visual question answering, localization, and reasoning benchmarks while training in less than 9 hours on 8 ×\times A100 GPUs at the 7B scale.

  3. Knowl 3 — Elimination of Multi-Stage Projector Pretraining in VLM Optimization

    empirical result

    Standard VLM training pipelines (such as LLaVa v1.5) employ a two-stage training scheme: Stage 1 aligns the vision-language projector in isolation on image captioning data (558K samples) while freezing both the visual backbone and language model, and Stage 2 finetunes both the projector and the language model on multimodal instruction data (665K samples).

    Eliminating Stage 1 and training the randomly initialized projector and language model directly on the multimodal instruction dataset in a single stage produces models that match or statistically significantly outperform the two-stage baseline across a 12-benchmark evaluation suite (p=0.00558p = 0.00558, one-sided Fisher's tt-test).

    Single-stage training provides two practical advantages:

    1. It reduces total training compute and training wall-clock time by 20%–25% (e.g., reducing training time on 8 ×\times A100 GPUs from 13.09 hours to 8.80 hours for 7B parameter models, and from 23.32 hours to 15.75 hours for 13B parameter models).
    2. It eliminates the requirement for dedicated image-caption pretraining datasets.
  4. Knowl 4 — Performance Degradation from Full Finetuning of Visual Backbones

    empirical result

    Finetuning the visual encoder weights alongside the projector and language model during VLM training leads to significant performance degradation compared to keeping the visual encoder frozen (p=0.00381p = 0.00381, one-sided Fisher's tt-test).

    The performance drop is especially severe on tasks requiring fine-grained spatial reasoning and object localization:

    • On RefCOCO bounding-box grounding (0.5 IoU), performance drops from 64.08% (frozen ViT, single-stage) to 42.56% (finetuned ViT, single-stage) and 19.24% (finetuned ViT, multi-stage).
    • On RefCOCO+, performance drops from 58.19% (frozen) to 37.89% (single-stage finetuned) and 17.48% (multi-stage finetuned).
    • On RefCOCOg, performance drops from 58.03% (frozen) to 41.05% (single-stage finetuned) and 23.12% (multi-stage finetuned).
    • On OCID-Ref robotic clutter localization (0.25 IoU), performance drops from 44.58% (frozen) to 33.42% (single-stage finetuned) and 16.35% (multi-stage finetuned).

    This degradation indicates representation collapse in the visual encoder when optimized solely through autoregressive next-token language loss without explicit visual preservation objectives.

  5. Knowl 5 — Fused Visual Backbones Combining Contrastive and Self-Supervised Features

    empirical result

    Comparing pretrained visual representations across Vision Transformers (ViT-Large variants at 224px) reveals that vision-language contrastive models (CLIP and SigLIP) significantly outperform models trained via self-supervised learning (DINOv2) or classification (ImageNet-21K/1K) for VLM performance (p=7.11×10−8p = 7.11 \times 10^{-8}).

    However, fusing patch features from DINOv2 (which captures low-level spatial geometry) with SigLIP (which captures high-level semantics) by concatenating their patch embeddings along the channel dimension before the projector MLP yields significant gains over using SigLIP alone (p=0.00164p = 0.00164):

    • Localization benchmarks improve by 5%–12%: RefCOCO increases from 61.38% to 73.86%, RefCOCO+ from 55.76% to 67.29%, RefCOCOg from 56.84% to 67.85%, and OCID-Ref from 41.49% to 52.82%.
    • Hallucination probing on POPE improves from 86.52% to 88.30%.
    • General visual question answering on VQAv2 improves from 78.81% to 79.18% and GQA from 63.60% to 64.33%.

    In contrast, fusing DINOv2 with CLIP yields mixed results (p=0.37313p = 0.37313) and causes a sharp drop on TextVQA (from 49.66% down to 15.67%), making DINOv2 + SigLIP the strongest vision backbone configuration.

  6. Knowl 6 — Impact of Image Preprocessing and Input Resolution Scaling

    empirical result

    Experiments evaluating three image preprocessing strategies across visual backbones demonstrate that naive resizing performs best overall:

    1. Resize & Crop: Resizing and taking a center crop removes visual information and causes severe degradation on full-scene reasoning and localization (e.g., RefCOCO drops from 64.08% to 54.31% for CLIP 336px).
    2. Letterbox Padding: Padding non-square images with constant borders preserves the aspect ratio but introduces up to 40%+ uninformative 'dead pixels'.
    3. Naive Resizing: Non-uniformly stretching/squeezing the image directly to the target square resolution warps the aspect ratio but avoids dead pixels and cropping. It outperforms letterbox padding for CLIP (e.g., GQA increases from 62.57% to 63.48%; TextVQA increases from 44.45% to 49.66%) and performs comparably for SigLIP (p=0.0176p = 0.0176 aggregate difference across representations).

    Scaling input image resolution from 224px to higher resolutions (336px for CLIP, 384px for SigLIP) yields statistically significant performance improvements across all evaluation categories (p=6.05×10−4p = 6.05 \times 10^{-4}), though quadrupling patch count increases downstream language model attention computational complexity.

  7. Knowl 7 — Base vs. Instruct-Tuned Language Models and Language-Only Safety Co-training

    empirical result

    A head-to-head comparison of VLMs trained with base language models (Llama-2 7B/13B) versus instruction-tuned language models (Vicuna v1.5 7B/13B) shows no statistically significant difference in objective quantitative benchmark accuracy (p=0.34854p = 0.34854). However, instruct-tuned models generate more verbose responses and exhibit higher rates of object hallucination.

    Using base language models with higher NLP benchmark scores (e.g., Mistral 7B vs. Llama-2 7B) does not produce statistically significant gains in aggregate VLM capability (p=0.03097p = 0.03097), though Mistral slightly improves localization accuracy.

    When using base language models, co-training on 40K language-only instruction samples from ShareGPT does not degrade visual performance (p=0.13655p = 0.13655), but is essential for safety: without this language-only co-training data, base LMs produce toxic, biased, and racist outputs when given adversarial prompts with malicious intent, whereas co-trained models retain safety guardrails.

  8. Knowl 8 — VLM Scaling Dynamics: Training Epochs and Data Diversity

    empirical result

    Analysis of training epochs and dataset composition reveals key scaling behaviors for visually-conditioned language models:

    1. Epoch Scaling: Training for a single epoch (the default in prior work such as LLaVa and PaLI) leads to underfitting. Training for 2 epochs yields statistically significant improvements (p=0.00496p = 0.00496), especially on structured output tasks (e.g., RefCOCO accuracy rises from 64.08% at 1 epoch to 71.23% at 2 epochs, RefCOCO+ rises from 58.19% to 65.40%), before performance plateaus at 3 epochs (71.79% on RefCOCO).
    2. Dataset Diversity Scaling: Augmenting the base instruction mixture (665K samples) with additional instruction data yields statistically significant performance gains (p=0.01459p = 0.01459). Incorporating LRV-Instruct (which emphasizes visual diversity including charts, diagrams, and printings) delivers larger gains across benchmarks than incorporating LVIS-Instruct-4V (which focuses on dense synthetic captions), demonstrating that visual diversity is more critical than raw sample volume for VLM pretraining data mixtures.
  9. Knowl 9 — Standardized 12-Benchmark VLM Evaluation Suite and Significance Testing Protocol

    experimental setup

    To evaluate VLM capabilities with objective metrics, an evaluation suite is constructed comprising 12 standardized benchmarks across three core capability areas:

    1. Open-Ended Visual Question Answering:

      • VQAv2 (general visual reasoning)
      • GQA (compositional and spatial reasoning, test-dev split)
      • VizWiz (visual reasoning and unanswerable question detection)
      • TextVQA (reasoning over scene text without OCR input tokens) Metric: Standard VQA accuracy via greedy decoding.
    2. Object Localization (Grounding):

      • RefCOCO (referring expressions with spatial anchors)
      • RefCOCO+ (appearance-based referring expressions)
      • RefCOCOg (long referring descriptions)
      • OCID-Ref (out-of-distribution cluttered robotic scenes) Metric: Intersection over Union (IoU) accuracy (extIoU≥0.5 ext{IoU} \ge 0.5 for RefCOCO/+/g; extIoU≥0.25 ext{IoU} \ge 0.25 for OCID-Ref).
    3. Closed-Set Challenge Sets:

      • Visual Spatial Reasoning (VSR, zero-shot test split, 2-choice True/False spatial relation verification)
      • TallyQA (16-choice counting questions, numbers 0–15)
      • POPE (2-choice Yes/No object hallucination probing)
      • AI2 Diagrams (AI2D, 4-choice scientific diagram and chart question answering) Metric: Multiple-choice accuracy.

    Statistical Significance Testing: To compare two model configurations across heterogeneous metrics, normalized ZZ-scores are computed for each benchmark using the population mean and standard deviation across all evaluated models. A global average ZZ-score is computed across all 12 benchmarks, and a one-sided Fisher's tt-test is performed on paired differences at significance threshold p<0.01p < 0.01.

  10. Knowl 10 — Benchmark Performance Comparison of PRISM vs. Baseline Open VLMs

    data/table

    The table below reproduces evaluation results across 12 benchmarks comparing official baseline models (LLaVa v1.5, InstructBLIP) against PRISM model variants at 7B and 13B parameter scales. PRISM (Controlled) models use identical training data and budget to LLaVa v1.5, while full PRISM models incorporate fused DINOv2 + SigLIP features, 2-epoch training, and expanded data mixtures (LVIS-Instruct-4V and LRV-Instruct).

    Model VQAv2 GQA VizWiz TextVQA RefCOCO RefCOCO+ RefCOCOg OCIDRef VSR POPE TallyQA AI2D
    LLaVa v1.5 7B 76.54 61.58 54.24 46.13 55.12 49.47 50.92 35.07 51.47 86.57 62.06 54.10
    InstructBLIP 7B 76.12 48.41 32.02 33.54 N/A N/A N/A N/A 58.92 84.30 15.51 32.90
    Prism-CLIP 7B (Controlled) 77.87 63.65 56.10 50.31 66.42 60.14 60.56 44.12 66.61 86.83 60.86 55.46
    Prism-SigLIP 7B (Controlled) 79.12 63.98 58.99 55.79 64.74 58.58 60.56 43.63 65.14 87.07 64.54 55.48
    Prism-DINOSigLIP 7B (Controlled) 79.05 64.16 59.82 51.78 73.62 67.85 66.34 50.56 66.28 88.28 65.07 55.51
    Prism-DINOSigLIP 7B 80.97 65.27 52.82 55.64 77.78 73.08 71.04 54.12 59.57 88.12 66.70 55.65
    LLaVa v1.5 13B 78.13 63.17 56.66 48.99 66.75 61.36 60.85 45.56 69.07 87.10 64.83 57.13
    InstructBLIP 13B 59.46 42.92 30.65 27.90 N/A N/A N/A N/A 63.91 84.49 49.73 35.60
    Prism-CLIP 13B (Controlled) 78.83 64.10 57.09 52.22 70.92 65.95 65.03 47.32 65.96 86.96 65.71 56.64
    Prism-DINOSigLIP 13B (Controlled) 80.07 65.14 56.61 54.10 76.64 71.41 70.87 53.60 71.85 88.50 66.09 57.72
    Prism-DINOSigLIP 13B 81.66 66.13 58.01 57.08 79.39 75.55 72.73 54.62 72.18 88.07 70.41 57.96

    Prism-DINOSigLIP uniformly outperforms baseline models across both 7B and 13B parameter scales, achieving large advantages on object localization tasks (e.g., RefCOCO 77.78% vs. 55.12% for LLaVa v1.5 7B) and general VQA.

  11. Knowl 11 — Architectural Scope and Conversational Evaluation Limitations

    limitation

    The conclusions and architecture in this work are subject to two main limitations:

    1. Architectural Scope: The study exclusively investigates the patch-as-token prefix-concatenation VLM paradigm. It does not evaluate architectures with learned visual token downsampling (such as the Perceiver Resamplers used in Flamingo and IDEFICS) or cross-attention mechanisms. Because patch tokens are preserved individually, high input resolutions scale the number of visual tokens quadratically in language model attention computation (O(N2)O(N^2)). Furthermore, empirical validation is limited to 7B and 13B parameter scales and does not address whether these design conclusions hold at scales ≥70B\ge 70\text{B}.
    2. Evaluation Scope: Evaluations focus exclusively on single-turn, standardized benchmarks with deterministic greedy decoding and closed-form metrics. This protocol does not assess open-ended, multi-turn conversational capabilities, dialogue coherence, or long-context reasoning across extended human-agent visual interactions.

Coverage note — None was omitted; all key design investigations (optimization stages, visual backbones and ensembling, image preprocessing, LM selection and safety co-training, scaling data and epochs), the PRISM recipe, benchmark evaluation suite, and tabulated results are fully covered.

References

  1. 1.Acharya, M., Kafle, K., and Kanan, C. TallyQA: Answering complex counting questions. In Association for the Advancement of Artificial Intelligence (AAAI), 2018.
  2. 2.AI, A. Fuyu-8b: A multimodal architecture for AI agents, 2023.
  3. 3.Alabdulmohsin, I. M., Zhai, X., Kolesnikov, A., and Beyer, L. Getting ViT in shape: Scaling laws for compute-optimal model design. arXiv preprint arXiv:2305.13035, 2023.
  4. 4.Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., Ring, R., Rutherford, E., Cabi, S., Han, T., Gong, Z., Samangooei, S., Monteiro, M., Menick, J., Borgeaud, S., Brock, A., Nematzadeh, A., Sharifzadeh, S., Binkowski, M., Barreira, R., Vinyals, O., Zisserman, A., and Simonyan, K. Flamingo: a visual language model for few-shot learning. In Advances in Neural Information Processing Systems (NeurIPS), 2022.
  5. 5.Bai, J., Bai, S., Yang, S., Wang, S., Tan, S., Wang, P., Lin, J., Zhou, C., and Zhou, J. Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond. arXiv preprint arXiv:2308.12966, 2023.
  6. 6.Biderman, S., Schoelkopf, H., Anthony, Q. G., Bradley, H., O’Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., et al. Pythia: A suite for analyzing large language models across training and scaling. In International Conference on Machine Learning (ICML), 2023.
  7. 7.Bigham, J. P., Jayant, C., Ji, H., Little, G., Miller, A., Miller, R. C., Miller, R., Tatarowicz, A., White, B., White, S., and Yeh, T. VizWiz: nearly real-time answers to visual questions. In User Interface Software and Technology (UIST), pp. 333–342, 2010.
  8. 8.Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Choromanski, K., Ding, T., Driess, D., Finn, C., Florence, P. R., Fu, C., Arenas, M. G., Gopalakrishnan, K., Han, K., Hausman, K., Herzog, A., Hsu, J., Ichter, B., Irpan, A., Joshi, N. J., Julian, R. C., Kalashnikov, D., Kuang, Y., Leal, I., Levine, S., Michalewski, H., Mordatch, I., Pertsch, K., Rao, K., Reymann, K., Ryoo, M. S., Salazar, G., Sanketi, P. R., Sermanet, P., Singh, J., Singh, A., Soricut, R., Tran, H., Vanhoucke, V., Vuong, Q. H., Wahid, A., Welker, S., Wohlhart, P., Xiao, T., Yu, T., and Zitkovich, B. RT-2: Vision-language-action models transfer web knowledge to robotic control. arXiv preprint arXiv:2307.15818, 2023.
  9. 9.Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020.
  10. 10.Chen, K., Zhang, Z., Zeng, W., Zhang, R., Zhu, F., and Zhao, R. Shikra: Unleashing multimodal llms referential dialogue magic. arXiv preprint arXiv:2306.15195, 2023a.
  11. 11.Chen, M., Tworek, J., Jun, H., Yuan, Q., de Oliveira Pinto, H. P., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., Ryder, N., Pavlov, M., Power, A., Kaiser, L., Bavarian, M., Winter, C., Tillet, P., Such, F. P., Cummings, D., Plappert, M., Chantzis, F., Barnes, E., Herbert-Voss, A., Guss, W. H., Nichol, A., Paino, A., Tezak, N., Tang, J., Babuschkin, I., Balaji, S., Jain, S., Saunders, W., Hesse, C., Carr, A. N., Leike, J., Achiam, J., Misra, V., Morikawa, E., Radford, A., Knight, M., Brundage, M., Murati, M., Mayer, K., Welinder, P., McGrew, B., Amodei, D., McCandlish, S., Sutskever, I., and Zaremba, W. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021.
  12. 12.Chen, X., Wang, X., Beyer, L., Kolesnikov, A., Wu, J., Voigtlaender, P., Mustafa, B., Goodman, S., Alabdulmohsin, I. M., Padlewski, P., Salz, D. M., Xiong, X., Vlasic, D., Pavetic, F., Rong, K., Yu, T., Keysers, D., Zhai, X.-Q., and Soricut, R. PaLI-3 vision language models: Smaller, faster, stronger. arXiv preprint arXiv:2310.09199, 2023b.
  13. 13.Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., Webson, A., Gu, S. S., Dai, Z., Suzgun, M., Chen, X., Chowdhery, A., Valter, D., Narang, S., Mishra, G., Yu, A. W., Zhao, V., Huang, Y., Dai, A. M., Yu, H., Petrov, S., hsin Chi, E. H., Dean, J., Devlin, J., Roberts, A., Zhou, D., Le, Q. V., and Wei, J. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416, 2022.
  14. 14.Dai, W., Li, J., Li, D., Tiong, A. M. H., Zhao, J., Wang, W., Li, B. A., Fung, P., and Hoi, S. C. H. InstructBLIP: Towards general-purpose vision-language models with instruction tuning. arXiv preprint arXiv:2305.06500, 2023.
  15. 15.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations (ICLR), 2021.
  16. 16.Driess, D., Xia, F., Sajjadi, M. S. M., Lynch, C., Chowdhery, A., Ichter, B., Wahid, A., Tompson, J., Vuong, Q. H., Yu, T., Huang, W., Chebotar, Y., Sermanet, P., Duckworth, D., Levine, S., Vanhoucke, V., Hausman, K., Toussaint, M., Greff, K., Zeng, A., Mordatch, I., and Florence, P. R. Palm-e: An embodied multimodal language model. In International Conference on Machine Learning (ICML), 2023.
  17. 17.Gao, P., Han, J., Zhang, R., Lin, Z., Geng, S., Zhou, A., Zhang, W., Lu, P., He, C., Yue, X., Li, H., and Qiao, Y. J. Llama-adapter v2: Parameter-efficient visual instruction model. arXiv preprint arXiv:2304.15010, 2023.
  18. 18.Gong, T., Lyu, C., Zhang, S., Wang, Y., Zheng, M., Zhao, Q., Liu, K., Zhang, W., Luo, P., and Chen, K. Multimodal-GPT: A vision and language model for dialogue with humans. ArXiv, 0, 2023.
  19. 19.Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D. Making the V in VQA matter: Elevating the role of image understanding in visual question answering. In Computer Vision and Pattern Recognition (CVPR), 2017.
  20. 20.Gupta, A., Dollar, P., and Girshick, R. B. LVIS: A dataset for large vocabulary instance segmentation. In Computer Vision and Pattern Recognition (CVPR), 2019.
  21. 21.Hendrycks, D. and Gimpel, K. Gaussian error linear units (gelus). arXiv preprint arXiv:1606.08415, 2016.
  22. 22.Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. Measuring massive multitask language understanding. In International Conference on Learning Representations (ICLR), 2021.
  23. 23.Hudson, D. A. and Manning, C. D. GQA: A new dataset for real-world visual reasoning and compositional question answering. In Computer Vision and Pattern Recognition (CVPR), 2019.
  24. 24.Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de Las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.-A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E. Mistral 7b. arXiv preprint arXiv:2310.06825, 2023.
  25. 25.Karamcheti, S., Orr, L., Bolton, J., Zhang, T., Goel, K., Narayan, A., Bommasani, R., Narayanan, D., Hashimoto, T., Jurafsky, D., Manning, C. D., Potts, C., Re, C., and Liang, P. Mistral - a journey towards reproducible language model training, 2021.
  26. 26.Karamcheti, S., Nair, S., Chen, A. S., Kollar, T., Finn, C., Sadigh, D., and Liang, P. Language-driven representation learning for robotics. In Robotics: Science and Systems (RSS), 2023.
  27. 27.Kazemzadeh, S., Ordonez, V., Matten, M., and Berg, T. ReferItGame: Referring to objects in photographs of natural scenes. In Empirical Methods in Natural Language Processing (EMNLP), pp. 787–798, 2014.
  28. 28.Kembhavi, A., Salvato, M., Kolve, E., Seo, M., Hajishirzi, H., and Farhadi, A. A diagram is worth a dozen images. In European Conference on Computer Vision (ECCV), 2016.
  29. 29.Kerr, J., Kim, C. M., Goldberg, K., Kanazawa, A., and Tancik, M. LERF: Language embedded radiance fields. In International Conference on Computer Vision (ICCV), 2023.
  30. 30.Kobayashi, S., Matsumoto, E., and Sitzmann, V. Decomposing nerf for editing via feature field distillation. arXiv preprint arXiv:2205.15585, 2022.
  31. 31.Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidi, Y., Li, L.-J., Shamma, D. A., Bernstein, M. S., and Li, F.-F. Visual genome: Connecting language and vision using crowdsourced dense image annotations. International Journal of Computer Vision, 123:32–73, 2017.
  32. 32.Laurenson, H., Saulnier, L., Tronchon, L., Bekman, S., Singh, A., Lozhkov, A., Wang, T., Karamcheti, S., Rush, A. M., Kiela, D., Cord, M., and Sanh, V. OBELICS: An open web-scale filtered dataset of interleaved image-text documents. In Neural Information Processing Systems Track on Datasets and Benchmarks (NeurIPS Datasets and Benchmarks), 2023.
  33. 33.Li, B., Zhang, P., Yang, J., Zhang, Y., Pu, F., and Liu, Z. Otterhd: A high-resolution multi-modality model. arXiv preprint arXiv:2311.04219, 2023a.
  34. 34.Li, J., Li, D., Xiong, C., and Hoi, S. C. H. BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In International Conference on Machine Learning (ICML), 2022.
  35. 35.Li, J., Li, D., Savarese, S., and Hoi, S. C. H. BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International Conference on Machine Learning (ICML), 2023b.
  36. 36.Li, K., He, Y., Wang, Y., Li, Y., Wang, W., Luo, P., Wang, Y., Wang, L., and Qiao, Y. Videochat: Chat-centric video understanding. ArXiv, 0, 2023c.
  37. 37.Li, X. L. and Liang, P. Prefix-tuning: Optimizing continuous prompts for generation. In Association for Computational Linguistics (ACL), 2021.
  38. 38.Li, Y., Du, Y., Zhou, K., Wang, J., Zhao, W. X., and rong Wen, J. Evaluating object hallucination in large vision-language models. In Empirical Methods in Natural Language Processing (EMNLP), 2023d.
  39. 39.Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollar, P., and Zitnick, C. L. Microsoft COCO: Common objects in context. In European Conference on Computer Vision (ECCV), pp. 740–755, 2014.
  40. 40.Liu, F., Emerson, G. E. T., and Collier, N. Visual spatial reasoning. Transactions of the Association for Computational Linguistics (TACL), 11:635–651, 2022.
  41. 41.Liu, F., Lin, K., Li, L., Wang, J., Yacoob, Y., and Wang, L. Mitigating hallucination in large multi-modal models via robust instruction tuning. arXiv preprint arXiv:2306.14565, 2023a.
  42. 42.Liu, H., Li, C., Li, Y., and Lee, Y. J. Improved baselines with visual instruction tuning. arXiv preprint arXiv:2310.03744, 2023b.
  43. 43.Liu, H., Li, C., Wu, Q., and Lee, Y. J. Visual instruction tuning. In Advances in Neural Information Processing Systems (NeurIPS), 2023c.
  44. 44.Liu, X., Zhu, Y., Lan, Y., Yang, C., and Qiao, Y. Query-relevant images jailbreak large multi-modal models. arXiv preprint arXiv:2311.17600, 2023d.
  45. 45.Liu, Y., Duan, H., Zhang, Y., Li, B., Zhang, S., Zhao, W., Yuan, Y., Wang, J., He, C., Liu, Z., Chen, K., and Lin, D. MMBench: Is your multi-modal model an all-around player? arXiv preprint arXiv:2307.06281, 2023e.
  46. 46.Marino, K., Rastegari, M., Farhadi, A., and Mottaghi, R. OK-VQA: A visual question answering benchmark requiring external knowledge. In Computer Vision and Pattern Recognition (CVPR), 2019.
  47. 47.Mishra, A., Shekhar, S., Singh, A. K., and Chakraborty, A. OCR-VQA: Visual question answering by reading text in images. In International Conference on Document Analysis and Recognition (ICDAR), 2019.
  48. 48.OpenAI, Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., Avila, R., Babuschkin, I., Balaji, S., Balcom, V., Baltescu, P., Bao, H., Bavarian, M., Belgum, J., Bello, I., Berdine, J., Bernadett-Shapiro, G., Berner, C., Bogdonoff, L., Boiko, O., Boyd, M., Brakman, A.-L., Brockman, G., Brooks, T., Brundage, M., Button, K., Cai, T., Campbell, R., Cann, A., Carey, B., Carlson, C., Carmichael, R., Chan, B., Chang, C., Chantzis, F., Chen, D., Chen, S., Chen, R., Chen, J., Chen, M., Chess, B., Cho, C., Chu, C., Chung, H. W., Cummings, D., Currier, J., Dai, Y., Decareaux, C., Degry, T., Deutsch, N., Deville, D., Dhar, A., Dohan, D., Dowling, S., Dunning, S., Ecoffet, A., Eleti, A., Eloundou, T., Farhi, D., Fedus, L., Felix, N., Fishman, S. P., Forte, J., Fulford, I., Gao, L., Georges, E., Gibson, C., Goel, V., Gogineni, T., Goh, G., Gontijo-Lopes, R., Gordon, J., Grafstein, M., Gray, S., Greene, R., Gross, J., Gu, S. S., Guo, Y., Hallacy, C., Han, J., Harris, J., He, Y., Heaton, M., Heidecke, J., Hesse, C., Hickey, A., Hickey, W., Hoeschele, P., Houghton, B., Hsu, K., Hu, S., Hu, X., Huizinga, J., Jain, S., Jain, S., Jang, J., Jiang, A., Jiang, R., Jin, H., Jin, D., Jomoto, S., Jonn, B., Jun, H., Kaftan, T., Kaiser, L., Kamali, A., Kanitscheider, I., Keskar, N. S., Khan, T., Kilpatrick, L., Kim, J. W., Kim, C., Kim, Y., Kirchner, H., Kiros, J. R., Knight, M., Kokotajlo, D., Kondraciuk, L., Kondrich, A., Konstantinidis, A., Kosic, K., Krueger, G., Kuo, V., Lampe, M., Lan, I., Lee, T., Leike, J., Leung, J., Levy, D., Li, C. M., Lim, R., Lin, M., Lin, S., Litwin, M., Lopez, T., Lowe, R., Lue, P., Makanju, A. A., Malfacini, K., Manning, S., Markov, T., Markovski, Y., Martin, B., Mayer, K., Mayne, A., McGrew, B., McKinney, S. M., McLeavey, C., McMillan, P., McNeil, J., Medina, D., Mehta, A., Menick, J., Metz, L., Mishchenko, A., Mishkin, P., Monaco, V., Morikawa, E., Mossing, D. P., Mu, T., Murati, M., Murk, O., Mely, D., Nair, A., Nakano, R., Nayak, R., Neelakantan, A., Ngo, R., Noh, H., Long, O., O’Keefe, C., Pachocki, J. W., Paino, A., Palermo, J., Pantuliano, A., Parascandolo, G., Parish, J., Parparita, E., Passos, A., Pavlov, M., Peng, A., Perelman, A., de Avila Belbute Peres, F., Petrov, M., de Oliveira Pinto, H. P., Pokorny, M., Pokrass, M., Pong, V. H., Powell, T., Power, A., Power, B., Proehl, E., Puri, R., Radford, A., Rae, J., Ramesh, A., Raymond, C., Real, F., Rimbach, K., Ross, C., Rotsted, B., Roussez, H., Ryder, N., Saltarelli, M. D., Sanders, T., Santurkar, S., Sastry, G., Schmidt, H., Schnurr, D., Schulman, J., Selsam, D., Sheppard, K., Sherbakov, T., Shieh, J., Shoker, S., Shyam, P., Sidor, S., Sigler, E., Simens, M., Sitkin, J., Slama, K., Sohl, I., Sokolowsky, B. D., Song, Y., Staudacher, N., Such, F. P., Summers, N., Sutskever, I., Tang, J., Tezak, N. A., Thompson, M., Tillet, P., Tootoonchian, A., Tseng, E., Tuggle, P., Turley, N., Tworek, J., Uribe, J. F. C., Vallone, A., Vijayvergiya, A., Voss, C., Wainwright, C., Wang, J. J., Wang, A., Wang, B., Ward, J., Wei, J., Weinmann, C., Welihinda, A., Welinder, P., Weng, J., Weng, L., Wiethoff, M., Willner, D., Winter, C., Wolrich, S., Wong, H., Workman, L., Wu, S., Wu, J., Wu, M., Xiao, K., Xu, T., Yoo, S., Yu, K., Yuan, Q., Zaremba, W., Zellers, R., Zhang, C., Zhang, M., Zhao, S., Zheng, T., Zhuang, J., Zhuk, W., and Zoph, B. GPT-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  49. 49.Oquab, M., Darcet, T., Moutakanni, T., Vo, H. Q., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., Assran, M., Ballas, N., Galuba, W., Howes, R., Huang, P.-Y. B., Li, S.-W., Misra, I., Rabbat, M. G., Sharma, V., Synnaeve, G., Xu, H., Jegou, H., Mairal, J., Labatut, P., Joulin, A., and Bojanowski, P. DINOv2: Learning robust visual features without supervision. Transactions of Machine Learning Research (TMLR), 2023.
  50. 50.Ordonez, V., Kulkarni, G., and Berg, T. L. Im2Text: Describing images using 1 million captioned photographs. In Advances in Neural Information Processing Systems (NeurIPS), 2011.
  51. 51.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L. E., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R. J. Training language models to follow instructions with human feedback. arXiv, 2022.
  52. 52.Qi, X., Huang, K., Panda, A., Wang, M., and Mittal, P. Visual adversarial examples jailbreak aligned large language models. arXiv preprint arXiv:2306.13213, 2023.
  53. 53.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning (ICML), volume 139, pp. 8748–8763, 2021.
  54. 54.Rasley, J., Rajbhandari, S., Ruwase, O., and He, Y. DeepSpeed: System optimizations enable training deep learning models with over 100 billion parameters. In International Conference on Knowledge Discovery and Data Mining (KDD), 2020.
  55. 55.Schuhmann, C., Vencu, R., Beaumont, R., Kaczmarczyk, R., Mullis, C., Katta, A., Coombes, T., Jitsev, J., and Komatsuzaki, A. LAION-400M: Open dataset of CLIP-filtered 400 million image-text pairs. arXiv preprint arXiv:2111.02114, 2021.
  56. 56.Schwenk, D., Khandelwal, A., Clark, C., Marino, K., and Mottaghi, R. A-OKVQA: A benchmark for visual question answering using world knowledge. arXiv preprint arXiv:2206.01718, 2022.
  57. 57.ShareGPT. ShareGPT, 2023.
  58. 58.Sharma, P., Ding, N., Goodman, S., and Soricut, R. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Association for Computational Linguistics (ACL), 2018.
  59. 59.Sidorov, O., Hu, R., Rohrbach, M., and Singh, A. TextCaps: a dataset for image captioning with reading comprehension. In European Conference on Computer Vision (ECCV), 2020.
  60. 60.Singh, A., Natarajan, V., Shah, M., Jiang, Y., Chen, X., Batra, D., Parikh, D., and Rohrbach, M. Towards VQA models that can read. In Computer Vision and Pattern Recognition (CVPR), 2019.
  61. 61.Steiner, A., Kolesnikov, A., Zhai, X., Wightman, R., Uszkoreit, J., and Beyer, L. How to train your ViT? data, augmentation, and regularization in vision transformers. Transactions of Machine Learning Research (TMLR), 2021.
  62. 62.Subramanian, S., Narasimhan, M. G., Khangaonkar, K., Yang, K., Nagrani, A., Schmid, C., Zeng, A., Darrell, T., and Klein, D. Modular visual question answering via code generation. In Association for Computational Linguistics (ACL), 2023.
  63. 63.Sur’is, D., Menon, S., and Vondrick, C. ViperGPT: Visual inference via Python execution for reasoning. In International Conference on Computer Vision (ICCV), 2023.
  64. 64.Tan, H. H. and Bansal, M. LXMERT: Learning cross-modality encoder representations from transformers. In Empirical Methods in Natural Language Processing (EMNLP), 2019.
  65. 65.Touvron, H., Martin, L., Stone, K. R., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D. M., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A. S., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I. M., Korenev, A. V., Koura, P. S., Lachaux, M.-A., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T. Llama 2: Open foundations and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
  66. 66.Wang, J., Meng, L., Weng, Z., He, B., Wu, Z., and Jiang, Y.-G. To see is to believe: Prompting GPT-4V for better visual instruction tuning. arXiv preprint arXiv:2311.07574, 2023.
  67. 67.Wang, K.-J., Liu, Y.-H., Su, H.-T., Wang, J.-W., Wang, Y.-S., Hsu, W. H., and Chen, W.-C. OCID-Ref: A 3d robotic dataset with embodied language for clutter scene grounding. In Association for Computational Linguistics (ACL), 2021.
  68. 68.Wightman, R. Pytorch image models. https://github.com/rwightman/pytorch-image-models, 2019.
  69. 69.Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., and Brew, J. HuggingFace’s transformers: State-of-the-art natural language processing. arXiv preprint arXiv:1910.03771, 2019.
  70. 70.Ye, Q., Xu, H., Xu, G., Ye, J., Yan, M., Zhou, Y., Wang, J., Hu, A., Shi, P., Shi, Y., Li, C., Xu, Y., Chen, H., Tian, J., Qi, Q., Zhang, J., and Huang, F. mPLUG-Owl: Modularization empowers large language models with multimodality. arXiv preprint arXiv:2304.14178, 2023.
  71. 71.Yu, L., Poirson, P., Yang, S., Berg, A. C., and Berg, T. L. Modeling context in referring expressions. In European Conference on Computer Vision (ECCV), 2016.
  72. 72.Yu, W., Yang, Z., Li, L., Wang, J., Lin, K., Liu, Z., Wang, X., and Wang, L. MM-Vet: Evaluating large multimodal models for integrated capabilities. arXiv preprint arXiv:2308.02490, 2023.
  73. 73.Zellers, R., Holtzman, A., Rashkin, H., Bisk, Y., Farhadi, A., Roesner, F., and Choi, Y. Defending against neural fake news. In Advances in Neural Information Processing Systems (NeurIPS), pp. 9054–9065, 2019.
  74. 74.Zhai, X., Mustafa, B., Kolesnikov, A., and Beyer, L. Sigmoid loss for language image pre-training. In International Conference on Computer Vision (ICCV), 2023.
  75. 75.Zhao, Y., Gu, A., Varma, R., Luo, L., chin Huang, C., Xu, M., Wright, L., Shojanazeri, H., Ott, M., Shleifer, S., Desmaison, A., Balioglu, C., Nguyen, B., Chauhan, G., Hao, Y., and Li, S. PyTorch FSDP: Experiences on scaling fully sharded data parallel. In Very Large Data Bases (VLDB), 2023.
  76. 76.Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., and Stoica, I. Judging LLM-as-a-judge with MT-bench and chatbot arena. arXiv preprint arXiv:2306.05686, 2023.

Citation

MLA
Karamcheti, S., et al. “Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models”. arXiv, 2024, http://arxiv.org/abs/2402.07865v2.
APA
Karamcheti, S., Nair, S., Balakrishna, A., Liang, P., Kollar, T., & Sadigh, D. (2024). Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models. arXiv. http://arxiv.org/abs/2402.07865v2
Chicago
Karamcheti, S., S. Nair, A. Balakrishna, P. Liang, T. Kollar, and D. Sadigh. 2024. “Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models”. arXiv. http://arxiv.org/abs/2402.07865v2.
Harvard
Karamcheti, S. et al. (2024) “Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.07865v2.
Vancouver
1. Karamcheti S, Nair S, Balakrishna A, Liang P, Kollar T, Sadigh D (2024) Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models. arXiv

BibTeX

@article{karamcheti2024prismatic,
  title = {Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models},
  author = {Karamcheti, Siddharth and Nair, Suraj and Balakrishna, Ashwin and Liang, Percy and Kollar, Thomas and Sadigh, Dorsa},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.07865v2},
  eprint = {2402.07865}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/