Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models
Siddharth KaramchetiSuraj NairAshwin BalakrishnaPercy LiangThomas KollarDorsa Sadigh
Investigates key design decisions in vision-language models across visual encoders and training strategies, delivering a standardized evaluation framework and open-source models that outperform LLaVA-1.5 and InstructBLIP.
Visually-conditioned language models, which generate natural language responses from visual and textual inputs, are expanding rapidly across applications such as robotics, visual chat, and scene understanding. However, existing development practices rely on untested architectural and training conventions, while relying on subjective, model-based evaluation methods that obscure what drives actual performance.
The article systematically evaluates the primary design axes of these multimodal models—including optimization strategies, visual representations, language model selection, and dataset scaling—to establish an empirical foundation for efficient and effective model development.
To conduct this investigation, the authors developed an optimized, modular training framework alongside a standardized evaluation suite of twelve established objective benchmarks spanning visual question answering, object localization, and challenge sets probing spatial reasoning and hallucination. Using controlled single-variable comparisons at the 7-billion and 13-billion parameter scales, the study analyzed the isolated impact of each core design choice.
The analysis produced several key findings: First, eliminating the common multi-stage training pipeline in favor of direct, single-stage training improves aggregate performance while reducing training compute costs by 20% to 25%. Second, keeping the visual backbone frozen is critical; finetuning it during training severely degrades performance, particularly on localization tasks. Third, fusing complementary visual representations—specifically combining high-level contrastive features from SigLIP with low-level spatial features from DINOv2—yields significant 5% to 10% gains on localization and spatial reasoning benchmarks. Fourth, base language models perform comparably to instruction-tuned models while exhibiting less verbosity and lower hallucination rates, though including language-only safety data during training is essential to prevent harmful and biased outputs. Finally, models benefit significantly from training for two full epochs rather than one, and scaling dataset diversity improves performance more effectively than raw volume.
These findings demonstrate that organizations can achieve superior multimodal performance with substantially lower compute budgets and simplified engineering workflows. Conventional practices, such as complex multi-stage alignment and full model finetuning, add unnecessary cost and operational risk while degrading downstream accuracy. By consolidating these design principles, the authors introduced the PRISM model family, which consistently outperforms leading open-source models such as LLaVa v1.5 and InstructBLIP across all twelve benchmark tasks.
Technical leaders and engineering teams should adopt single-stage training workflows, freeze visual backbones, deploy fused visual representations, and allocate sufficient training time (two epochs) on diverse data mixtures. Teams utilizing base language models must retain language-only safety co-training data to mitigate toxic or biased generations. Further work should explore architectural methods to prevent visual representation collapse during full finetuning and investigate optimal pretraining mixtures for downstream multimodal integration.
Confidence in these findings is high for standard autoregressive architectures at the 7-billion to 13-billion parameter scale. However, decision-makers should exercise caution when extrapolating these conclusions to alternative architectures (such as resampler-based models), substantially larger scales (such as 70-billion-plus parameters), or extended conversational contexts that exceed standard single-turn objective benchmarks.
- Paper: Improved Baselines with Visual Instruction Tuning, Haotian Liu et al. (2024). It establishes the LLaVA-1.5 baseline and design choices that Prismatic VLMs directly investigates, benchmarks against, and aims to outperform.
- Paper: InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning, Wenliang Dai et al. (2023). It introduces the InstructBLIP architecture and instruction-tuning framework, which serves as one of the primary open VLM baselines dissected and evaluated in Prismatic VLMs.
- Paper: Visual Instruction Tuning, Haotian Liu et al. (2023). It introduces the foundational visual instruction tuning paradigm connecting vision encoders to LLMs that underpins the VLM design space analyzed by the source.
- Paper: DINOv2: Learning Robust Visual Features without Supervision, Maxime Oquab et al. (2023). It provides the self-supervised DINOv2 vision representations that the source systematically tests and compares against contrastive encoders in its visual backbone ablations.
- Paper: Evaluating Object Hallucination in Large Vision-Language Models, Yifan Li et al. (2023). It introduces the POPE hallucination benchmark used by the source paper as a core component of its standardized VLM evaluation suite.
- Paper: MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models, Chaoyou Fu et al. (2023). It defines a comprehensive, multi-task benchmark for multimodal language models that informs the standardized evaluation methodologies adopted in the source study.
- Paper: MMBench: Is Your Multi-modal Model an All-around Player?, Yuanzhan Liu et al. (2023). It establishes fine-grained perception and reasoning benchmarks that form part of the objective evaluation landscape synthesized in Prismatic VLMs.
- Paper: BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models, Junnan Li et al. (2023). It details the modular integration of frozen vision backbones and LLMs via learnable connectors, representing a key architectural axis explored in the source.
- Paper: OpenVLA: An Open-Source Vision-Language-Action Model, Moo Jin Kim et al. (2024). It directly builds upon the fused vision representation and VLM design principles identified in Prismatic VLMs to train open-source vision-language-action policies for robotics.
- Paper: CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models, Qingqing Zhao et al. (2025). It extends VLA architectures built on prismatic-style VLM backbones by integrating visual chain-of-thought intermediate reasoning for complex robotic manipulation.
- Paper: LLaVA-OneVision: Easy Visual Task Transfer, Bo Li et al. (2024). It advances the open VLM architecture space by applying unified representations across single-image, multi-image, and video transfer scenarios.
- Paper: Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution, Peng Wang et al. (2024). It extends open VLM design by implementing dynamic-resolution perception and multidimensional positional embeddings across images and video.
- Paper: Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling, Zhe Chen et al. (2024). It investigates advanced model, data, and test-time scaling recipes to extend the performance boundaries of modular open-source multimodal systems.
- Paper: InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models, Jinguo Zhu et al. (2025). It explores native multimodal pretraining recipes and preference optimization, offering a next-generation progression beyond modular visual instruction tuning.
- Paper: Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis, Chaoyou Fu et al. (2025). It expands multimodal evaluation beyond static images into comprehensive temporal and dynamic video reasoning benchmarks for large vision-language models.
- Paper: Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation, Zhiheng Liu et al. (2026). It challenges the visual encoder design paradigm explored in Prismatic VLMs by testing end-to-end patch pixel embeddings for multimodal understanding and generation.
