Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Peter TongEllis BrownPenghao WuSanghyun WooAdithya IyerSai Charitha AkulaShusheng YangJihan YangManoj MiddepoguZiteng Wang

article2024NeurIPS765 citations

Presents a fully open, vision-centric suite of multimodal language models backed by systematic evaluations of over twenty vision encoders, a token-reducing spatial aggregator, and complete open-source training recipes and benchmarks.

Listen

Recent progress in multimodal large language models—systems that combine language models with visual processing—has been largely driven by scaling up text-based language backbones. However, the design of the visual perception components remains underexplored, creating models that often rely on language cues as shortcuts rather than developing genuine visual grounding. This limitation leads to deficiencies in real-world visual perception tasks such as reading charts, recognizing fine spatial layouts, and assessing 3D depth. The article systematically evaluates the vision design space for multimodal language models across visual representations, connector architectures, instruction tuning recipes, and training data mixtures. Its main objective is to establish an open, vision-centric framework and introduce Cambrian-1, a family of state-of-the-art multimodal models designed to bridge the gap between visual representation learning and multimodal reasoning.

To conduct this evaluation, the researchers tested over 20 distinct vision encoders across language-supervised, self-supervised, and specialized vision models using a standardized two-stage instruction-tuning recipe. They critically examined existing multimodal benchmarks, uncovering that several standard evaluations show less than a 5% difference between having vision enabled versus disabled. To address this benchmarking gap, the team developed CV-Bench, a curated 2,638-example benchmark probing fundamental 2D spatial relationships and 3D depth awareness. The researchers also engineered the Spatial Vision Aggregator, a dynamic connector that integrates high-resolution feature maps from multiple vision backbones while compressing the visual token count to 576 tokens. Finally, they curated a 9.78-million sample instruction dataset, termed Cambrian-10M, refining it through source balancing and targeted data generation into an optimized 7-million sample mix.

Key findings demonstrate that combining diverse visual backbones significantly improves multimodal performance. While language-supervised models like CLIP provide strong baselines, integrating self-supervised encoders like DINOv2 and high-resolution convolutional architectures substantially enhances vision-centric and document understanding capabilities. Second, the Spatial Vision Aggregator delivers superior accuracy across benchmarks while reducing the visual token footprint to roughly one-fifth of competing architectures such as LLaVA-NeXT. Third, instruction data curation and category balancing proved critical; curating the raw dataset into the balanced 7-million sample mix yielded higher overall accuracy than training on the full uncurated dataset. Fourth, the researchers discovered that heavy exposure to short-answer visual question datasets causes models to lose conversational fluency—a problem mitigated by inserting explicit formatting system prompts during training.

These findings indicate that visual grounding can be dramatically improved without exponentially increasing computational inference costs, as effective token aggregation delivers high performance at lower token counts. This is particularly relevant for organizations seeking to deploy efficient, high-accuracy document analysis, robotics, and image understanding systems. For future work, development teams should adopt multi-encoder vision setups, utilize spatial aggregation mechanisms, and prioritize balanced instruction mixtures over sheer data volume. The primary limitation noted is that the current model uses a fixed token compression rather than dynamic native-resolution handling for extreme aspect ratios or ultra-high resolutions. Nonetheless, the empirical results provide high confidence that vision-centric architectures and open evaluation recipes substantially narrow the gap between open-source models and leading proprietary systems.

No sufficiently relevant recommendations were found.

Cover for Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Abstract

We introduce Cambrian-1, a family of multimodal LLMs (MLLMs) designed with a vision-centric approach. While stronger language models can enhance multimodal capabilities, the design choices for vision components are often insufficiently explored and disconnected from visual representation learning research. This gap hinders accurate sensory grounding in real-world scenarios. Our study uses LLMs and visual instruction tuning as an interface to evaluate various visual representations, offering new insights into different models and architectures -- self-supervised, strongly supervised, or combinations thereof -- based on experiments with over 20 vision encoders. We critically examine existing MLLM benchmarks, address the difficulties involved in consolidating and interpreting results from various tasks, and introduce a new vision-centric benchmark, CV-Bench. To further improve visual grounding, we propose the Spatial Vision Aggregator (SVA), a dynamic and spatially-aware connector that integrates high-resolution vision features with LLMs while reducing the number of tokens. Additionally, we discuss the curation of high-quality visual instruction-tuning data from publicly available sources, emphasizing the importance of data source balancing and distribution ratio. Collectively, Cambrian-1 not only achieves state-of-the-art performance but also serves as a comprehensive, open cookbook for instruction-tuned MLLMs. We provide model weights, code, supporting tools, datasets, and detailed instruction-tuning and evaluation recipes. We hope our release will inspire and accelerate advancements in multimodal systems and visual representation learning.

Table of Contents

  • 1 Introduction
  • 2 Multimodal LLMs: Preliminaries and Related Work
  • 3 Evaluating Visual Representations through MLLMs
  • 3.1 Analyzing the Benchmarks
  • 3.2 Cambrian Vision-Centric Benchmark (CV-Bench)
  • 3.3 Instruction Tuning Recipes
  • 3.4 MLLMs as a Visual Representation Evaluator
  • 3.5 Combining Multiple Vision Encoders
  • 4 Spatial Vision Aggregator (SVA): A New Connector Design
  • 5 Instruction Tuning Data for Training MLLMs
  • 5.1 Data Collection
  • 5.2 Data Curation
  • 5.3 Alleviating the “Answer Machine Phenomenon” via System Prompts
  • 6 State of the Art Performance
  • 7 Discussion
  • References
  • A Training, Infrastructure, and Implementation
  • B Analyzing the Benchmarks
  • C Cambrian Vision-Centric Benchmark (CV-Bench)
  • D Vision Models in MLLMs
  • D.1 Details of Vision Models
  • D.2 Full Results of Different Vision Backbones
  • D.3 Model Ensemble
  • E Data
  • E.1 Catalog of Visual Instruction Data
  • E.2 Additional System Prompts used in Cambrian Data
  • E.3 Data Engine
  • E.4 Full results on data curation experiment
  • E.5 737K and 5M Mixes
  • E.6 Test Image Leakage in Visual Instruction Training Data
  • E.7 Broader Impacts
  • F Implementation Details
  • G Evaluation Details
  • G.1 System Prompts Used in Evaluation
  • G.2 Ablation Study on Fuzzy Matching Vs LLM Judgement
  • H Potential Misuse & Mitigation Strategies

Knowls

  1. Knowl 1 — Spatial Vision Aggregator Connector Architecture

    model/method

    The Spatial Vision Aggregator (SVA) is a multimodal connector designed to aggregate visual features from NN vision encoders of varying spatial resolutions and architectures into a compact set of visual tokens for a Large Language Model (LLM), while preserving spatial geometry.

    Let Fk∈RmkL×mkL×CF_k \in \mathbb{R}^{m_k L \times m_k L \times C} denote the feature map of the kk-th vision encoder, where L×LL \times L represents the spatial dimensions of the query grid, mk∈Z+m_k \in \mathbb{Z}^+ is an integer resolution multiplier, and CC is the hidden channel dimension. A learnable latent token x∈RCx \in \mathbb{R}^C is repeated L×LL \times L times to initialize a 2D query grid X∈RL2×CX \in \mathbb{R}^{L^2 \times C}.

    To impose spatial inductive bias, the query token at coordinate (i,j)(i, j) (0≤i,j<L0 \le i, j < L) cross-attends only to the corresponding sub-region slice across all NN encoders: Fk[mk⋅i:mk⋅(i+1), mk⋅j:mk⋅(j+1)]∈Rmk2×CF_k[m_k \cdot i : m_k \cdot (i+1), \, m_k \cdot j : m_k \cdot (j+1)] \in \mathbb{R}^{m_k^2 \times C}

    The updated query vector qi,j∗∈R1×Cq^*_{i,j} \in \mathbb{R}^{1 \times C} is computed as: qi,j∗=softmax(qi,j[ki,j,1,ki,j,2,…,ki,j,N]⊤C)[vi,j,1,vi,j,2,…,vi,j,N]q^*_{i,j} = \text{softmax}\left(\frac{q_{i,j} [k_{i,j,1}, k_{i,j,2}, \dots, k_{i,j,N}]^\top}{\sqrt{C}}\right) [v_{i,j,1}, v_{i,j,2}, \dots, v_{i,j,N}] where:

    • qi,j=WQxi,j∈R1×Cq_{i,j} = W_Q x_{i,j} \in \mathbb{R}^{1 \times C}, with learnable query projection matrix WQ∈RC×CW_Q \in \mathbb{R}^{C \times C}. qi,jq_{i,j} is concatenated with a global feature obtained by global average pooling over the encoder feature maps.
    • ki,j,k=WKkFk[mk⋅i:mk⋅(i+1), mk⋅j:mk⋅(j+1)]∈Rmk2×Ck_{i,j,k} = W_K^k F_k[m_k \cdot i : m_k \cdot (i+1), \, m_k \cdot j : m_k \cdot (j+1)] \in \mathbb{R}^{m_k^2 \times C}, with encoder-specific key projection WKk∈RC×CW_K^k \in \mathbb{R}^{C \times C}.
    • vi,j,k=WVkFk[mk⋅i:mk⋅(i+1), mk⋅j:mk⋅(j+1)]∈Rmk2×Cv_{i,j,k} = W_V^k F_k[m_k \cdot i : m_k \cdot (i+1), \, m_k \cdot j : m_k \cdot (j+1)] \in \mathbb{R}^{m_k^2 \times C}, with encoder-specific value projection WVk∈RC×CW_V^k \in \mathbb{R}^{C \times C}.

    When mk>1m_k > 1, learnable mk×mkm_k \times m_k positional encodings are added to the vision features. SVA defines capacity through DD (number of stacked cross-attention layers) and GG (number of distinct query groups aggregating in parallel). To mitigate information loss, SVA embeds multi-layer cross-attention layers (D=1,G=1D=1, G=1) at fixed layer strides throughout the LLM transformer blocks, allowing intermediate layers repeated access to uncompressed visual representations.

  2. Knowl 2 — Cambrian Vision-Centric Benchmark (CV-Bench)

    experimental setup

    The Cambrian Vision-Centric Benchmark (CV-Bench) evaluates 2D and 3D visual perception in Multimodal Large Language Models (MLLMs). It comprises 2,638 manually inspected, multi-choice visual question answering (VQA) examples generated from ADE20K, COCO, and Omni3D.

    CV-Bench evaluates four core tasks:

    1. 2D Spatial Relationship (650 samples; ADE20K, COCO): Identifies the relative position (left-right or top-bottom) of a target object relative to an anchor object, using bounding box visual prompts to remove instance ambiguity.
    2. 2D Object Counting (788 samples; ADE20K, COCO): Determines the count of a target category, including zero-count existence checks, with distractor choices constructed near the true count.
    3. 3D Depth Order (600 samples; Omni3D): Identifies which of two distinct objects is closer to the camera. Object A is closer than Object B if the farthest 3D bounding-box vertex of A is closer to the camera than the nearest 3D vertex of B by a specified Euclidean offset.
    4. 3D Relative Distance (600 samples; Omni3D): Evaluates three distinct objects (an anchor and objects A and B) to determine which is closer to the anchor. Object A is closer than B if the maximum Euclidean distance from A's vertices to the anchor is shorter than the minimum Euclidean distance from B's vertices to the anchor by a specified offset.

    The benchmark calculates accuracy by weighting 2D and 3D sub-tasks equally: Accuracy2D=AccuracyCOCO+AccuracyADE20K2\text{Accuracy}_{2\text{D}} = \frac{\text{Accuracy}_{\text{COCO}} + \text{Accuracy}_{\text{ADE20K}}}{2} Overall Accuracy=Accuracy2D+Accuracy3D2\text{Overall Accuracy} = \frac{\text{Accuracy}_{2\text{D}} + \text{Accuracy}_{3\text{D}}}{2}

  3. Knowl 3 — Visual Instruction Tuning Data Curation and Balancing in Cambrian-7M

    model/method

    To address severe class and source imbalances in public visual instruction datasets, Cambrian-10M (~9.78M raw instruction samples) is curated into Cambrian-7M through per-source data capping and ratio balancing across seven capability domains.

    1. Per-Source Data Thresholding (tt): An upper bound tt is enforced on the number of samples taken from any single data source to suppress the long tail of large, repetitive synthetic datasets (such as CLEVR and DVQA). Sweeps over t∈{150k,250k,350k,450k}t \in \{150\text{k}, 250\text{k}, 350\text{k}, 450\text{k}\} reveal an elbow effect where t∈[250k,350k]t \in [250\text{k}, 350\text{k}] maximizes average downstream accuracy across benchmarks (54.31%54.31\% at t=250kt=250\text{k} vs. 53.74%53.74\% at t=150kt=150\text{k}).

    2. Category Distribution Mixture: Data is balanced according to seven functional categories:

    • General Conversation & VQA: ~33.3%
    • OCR & Chart Data: ~27.6%
    • Language-Only Instruction Data: ~23.8%
    • Object Counting Data: ~8.5%
    • Mathematics: ~3.2%
    • Science: ~2.9%
    • Code Generation: ~0.8%

    Curation experiments demonstrated that while scaling OCR data directly improves OCR and chart benchmarks, an excessively high OCR proportion degrades general VQA and vision-centric spatial performance. Instruction-tuning on the curated Cambrian-7M achieves 55.9%55.9\% overall average benchmark accuracy, surpassing the uncurated Cambrian-10M (54.8%54.8\%) and baseline LLaVA-665K (40.7%40.7\%).

  4. Knowl 4 — Benchmark Performance Comparison of the Cambrian-1 Family

    data/table

    The Cambrian-1 family of MLLMs combines four vision encoders (OpenAI CLIP ViT-L/14@336, SigLIP ViT-SO400M/14@384, OpenCLIP ConvNeXt-XXL@1024, and DINOv2 ViT-L/14@518) via the Spatial Vision Aggregator (SVA), producing 576 visual tokens. Models are pre-trained on 2.5M adapter data and fine-tuned on Cambrian-7M across three LLM backbones (LLaMA-3-Instruct-8B, Vicuna-1.5-13B, and Hermes-2-Yi-34B).

    Model LLM Backbone # Vis Tokens General Avg Knowledge Avg OCR Chart Avg Vision-Centric Avg
    Mini-Gemini-HD-8B LLaMA-3-Ins-8B 2880 72.7 55.7 62.9 51.5
    LLaVA-NeXT-8B LLaMA-3-Ins-8B 2880 72.5 55.6 63.9 56.6
    Cambrian-1-8B LLaMA-3-Ins-8B 576 75.9 61.3 71.3 65.0
    Mini-Gemini-HD-13B Vicuna-1.5-13B 2880 68.6 54.1 60.8 49.4
    LLaVA-NeXT-13B Vicuna-1.5-13B 2880 70.0 53.7 62.9 55.9
    Cambrian-1-13B Vicuna-1.5-13B 576 75.7 60.2 71.3 62.2
    Mini-Gemini-HD-34B Hermes2-Yi-34B 2880 80.6 62.4 68.1 63.8
    LLaVA-NeXT-34B Hermes2-Yi-34B 2880 79.3 62.5 67.7 64.0
    Cambrian-1-34B Hermes2-Yi-34B 576 81.4 67.0 71.9 68.5
    GPT-4V Proprietary Unknown 75.8 65.2 77.4 62.4

    Despite using 576 visual tokens (one-fifth of the 2880 tokens used by Mini-Gemini-HD and LLaVA-NeXT), Cambrian-1 outperforms competing open-source models across all model sizes, with significant advantages on OCR & Chart (+7.4%+7.4\% over LLaVA-NeXT-8B) and Vision-Centric benchmarks (+8.4%+8.4\% over LLaVA-NeXT-8B).

  5. Knowl 5 — Multi-Vision Encoder Ensembling and Dynamic Cross-Attention Allocation

    empirical result

    Combining distinct vision encoder representations (language-supervised ViTs, self-supervised ViTs, and high-resolution ConvNets) improves overall MLLM performance over any single visual backbone, particularly on vision-centric and high-resolution perception tasks.

    When ensembling OpenAI CLIP ViT-L/14@336, SigLIP ViT-SO400M/14@384, OpenCLIP ConvNeXt-XXL@1024, and DINOv2 ViT-L/14@518 inside the Spatial Vision Aggregator (SVA) on Cambrian-1-8B, the attention score distribution dynamically adapts across tasks:

    • General Natural Images (GQA): Attention weights are distributed evenly across backbones (SigLIP: 29.7%29.7\%, ConvNeXt: 27.7%27.7\%, DINOv2: 24.1%24.1\%, CLIP: 18.5%18.5\%).
    • High-Resolution Text and Documents (DocVQA): ConvNeXt attention increases to 44.5%44.5\% and SigLIP to 31.1%31.1\%, while DINOv2 drops to 11.0%11.0\% and CLIP to 13.4%13.4\%, showing ConvNeXt's utility for dense 2D feature extraction.
    • Scientific Diagrams (ScienceQA): SigLIP receives the largest weight (35.2%35.2\%), followed by ConvNeXt (30.9%30.9\%), DINOv2 (17.6%17.6\%), and CLIP (16.3%16.3\%).

    Adding self-supervised DINOv2 to language-supervised backbones yields consistent gains on spatial benchmarks (CV-Bench 2D/3D and RealWorldQA) and improves OCR performance over single-encoder baselines.

  6. Knowl 6 — Impact of Connector Pre-Training and Vision Encoder Unfreezing

    empirical result

    Evaluating instruction-tuning recipes across 23 vision backbones with a Vicuna-1.5-7B LLM on a 737K data mix reveals two key findings regarding training stages and backbone freezing:

    1. Two-Stage vs. One-Stage Training: Skipping connector pre-training reduces performance across all benchmark categories. Training the connector first on image caption adapter data before joint LLM fine-tuning provides significant gains that scale with adapter dataset volume (0M→0.5M→1.2M0\text{M} \to 0.5\text{M} \to 1.2\text{M}). For example, SigLIP SO400M average benchmark score rises from 47.57%47.57\% (0M) to 50.41%50.41\% (0.5M) and 53.91%53.91\% (1.2M).

    2. Freezing vs. Unfreezing Vision Encoders: Unfreezing the vision backbone during the instruction fine-tuning stage yields widespread performance improvements across general, OCR & chart, and vision-centric categories:

    • Language-supervised encoders benefit broadly, particularly on OCR and chart tasks (e.g., OpenCLIP ConvNeXt-L@512 OCR & Chart score improves from 28.00%28.00\% to 35.40%35.40\%).
    • Self-supervised encoders gain significantly on vision-centric benchmarks (e.g., DINOv2 ViT-L/14@336 MMVP score improves from 21.33%21.33\% to 26.00%26.00\%).

    Unfreezing increases training computation time by approximately 50%–55%50\%\text{--}55\%.

  7. Knowl 7 — Closing the SSL vs. Language-Supervised Encoder Gap via Scaled Tuning

    empirical result

    Language-supervised vision encoders (e.g., CLIP, SigLIP) natively outperform self-supervised visual encoders (e.g., DINOv2, MAE, I-JEPA) on multimodal benchmarks, especially in OCR and chart understanding, due to text-rich web-scale pre-training data. However, this performance gap can be bridged by scaling multimodal instruction tuning data and unfreezing the visual encoder.

    In controlled experiments comparing OpenAI CLIP ViT-L/14@336 and DINOv2 ViT-L/14@336:

    • At 0.7M instruction tuning data with frozen encoders, CLIP strongly outperforms DINOv2 on General (62.0%62.0\% vs. 56.9%56.9\%) and OCR & Chart (36.8%36.8\% vs. 16.0%16.0\%) averages.
    • When instruction tuning data is scaled to 5.0M samples with an unfrozen vision encoder, DINOv2's overall average reaches 47.40%47.40\%, surpassing the 0.7M frozen CLIP baseline on General (61.62%61.62\% vs. 61.96%61.96\%) and Vision-Centric tasks (60.98%60.98\% on CV-Bench 2D and 34.67%34.67\% on MMVP), while narrowing the knowledge gap.
  8. Knowl 8 — Alleviating the Answer Machine Phenomenon via Response-Formatting System Prompts

    model/method

    When multimodal LLMs are fine-tuned on instruction datasets composed predominantly of short-answer VQA benchmarks, they exhibit catastrophic forgetting of general conversational fluency. They default to terse single-word or letter outputs even for open-ended queries—a failure mode termed the 'answer machine phenomenon'.

    To eliminate this behavior without reducing benchmark scores, specific response-formatting system prompts are prepended to training examples based on target response format:

    • Short phrase/word: "Answer the question using a single word or phrase."
    • Multiple choice: "Answer with the option's letter from the given choices directly."
    • Direct value: "Give the short answer directly."
    • Reasoning/Chain-of-thought: "First show your reasoning process and then give the final answer."

    Training with these explicit format prompts teaches the LLM backbone to associate short answers strictly with format-constrained prompts. During standard chat without format constraints, the model produces detailed, conversational responses and step-by-step reasoning while maintaining high benchmark accuracy.

  9. Knowl 9 — Targeted Internet Data Collection Engine for Scientific Visual Instruction Tuning

    algorithm

    The targeted internet data collection engine automatically curates high-quality scientific visual question answering data from structured web resources (e.g., Wikipedia) to expand scarce knowledge-centric multimodal data.

    Input: Set of target knowledge domains and subfields SS (e.g., Physics -> Electromagnetism)
    Input: Search API, HTML Parser, LLM generator MtopicM_\text{topic} (GPT-4), VQA generator MqaM_\text{qa} (GPT-3.5)
    Output: Dataset DVQAD_\text{VQA} of (image, question, answer) tuples
    Initialize DVQA=∅D_\text{VQA} = \emptyset
    for each subfield s∈Ss \in S do
        Topics $T_s = M_\text{topic}("\text{Generate domain topics for }" + s)
        for each topic τ∈Ts\tau \in T_s do
            URLs $U_\tau = \text{SearchEngineAPI}(\tau, \text{max\_results}=10)
            for each URL u∈Uτu \in U_\tau do
                Tuples P=HTMLParser(u) // extracts (img_url, caption, context_text)P = \text{HTMLParser}(u) \text{ // extracts (img\_url, caption, context\_text)}
                for each (img_url,caption,context)∈P(img\_url, caption, context) \in P do
                    if \text{WordCount}(context) < 50 then
                        continue
                    end if
                    image = \text{DownloadAndConvertToPNG}(img\_url)
                    $(Q, A) = M_\text{qa}(\text{ConstructPrompt}(caption, context))
                    DVQA=DVQA∪{(image,Q,A)}D_\text{VQA} = D_\text{VQA} \cup \{(image, Q, A)\}
                end for
            end for
        end for
    end for
    return DVQAD_\text{VQA}

    Processing 30 fields and 3,660 topics yielded 36,600 parsed web pages and generated 161k scientific VQA data points (a 400% increase over existing combined open scientific VQA sources). Difference hashing (dHash) confirmed a 0.06% image overlap (32 images total) with common evaluation test sets, demonstrating negligible test set leakage.

  10. Knowl 10 — Blind Evaluation Analysis and Capability Clustering of MLLM Benchmarks

    empirical result

    Auditing commonly used MLLM evaluation benchmarks by comparing model performance with visual inputs enabled versus disabled (text-only prompt) revealed significant language-shortcut biases across standard evaluation suites:

    • Language Shortcut Benchmarks: SQA-I, MMMU, MathVista, and AI2D show less than a 5%5\% score reduction when visual input is completely removed, indicating that these benchmarks predominantly evaluate the base LLM rather than multimodal perception.
    • Language Prior Bias: TextVQA and GQA achieve scores nearly 40%40\% higher under blind (vision-disabled) evaluation than random guessing, showing heavy exploitation of language priors.
    • Vision-Dependent Benchmarks: MMVP and MME Perception performance with disabled visual inputs drops below random guessing, demonstrating strong reliance on visual feature grounding.

    Principal Component Analysis (PCA) on cross-model performance across 23 vision backbones clusters MLLM benchmarks into four distinct capability groups:

    1. General (e.g., MME, MMB, SEED-I, GQA)
    2. Knowledge (e.g., MMMU, MathVista, SQA-I, AI2D)
    3. Chart & OCR (e.g., ChartQA, OCRBench, TextVQA, DocVQA)
    4. Vision-Centric (e.g., MMVP, RealWorldQA, CV-Bench)

Coverage note — Omitted minor infrastructure-specific engineering workarounds for TPU TorchXLA distributed training and extended catalog listings of raw benchmark-specific system prompts from the appendices.

References

  1. 1.M. Acharya, K. Kafle, and C. Kanan. “TallyQA: Answering complex counting questions”. In: AAAI. 2019.
  2. 2.A. Agrawal et al. “Don’t just assume; look and answer: Overcoming priors for visual question answering”. In: CVPR. 2018.
  3. 3.A. Ahmadyan et al. “Objectron: A Large Scale Dataset of Object-Centric Videos in the Wild with Pose Annotations”. In: CVPR (2021).
  4. 4.AI@Meta. “Llama 3 Model Card”. In: (2024).
  5. 5.H. A. Alawwad et al. “Enhancing Textbook Question Answering Task with Large Language Models and Retrieval Augmented Generation”. In: arXiv preprint arXiv:2402.05128 (2024).
  6. 6.J.-B. Alayrac et al. “Flamingo: a visual language model for few-shot learning”. In: NeurIPS. 2022.
  7. 7.T. Aquinas. Quaestiones Disputatae de Veritate. q.2 a.3 arg.19, 1259.
  8. 8.Aristotle. Metaphysics. Ed. by T. by W. D. Ross. The Internet Classics Archive, 350BCE.
  9. 9.M. Assran et al. “Self-supervised learning from images with a joint-embedding predictive architecture”. In: CVPR. 2023.
  10. 10.J. Bai et al. “Qwen Technical Report”. In: arXiv preprint arXiv:2309.16609 (2023).
  11. 11.J. Bai et al. “Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond”. In: (2023).
  12. 12.M. E. Banani et al. “Probing the 3D Awareness of Visual Foundation Models”. In: arXiv preprint arXiv:2404.08636 (2024).
  13. 13.G. Baruch et al. “ARKitScenes - A Diverse Real-World Dataset for 3D Indoor Scene Understanding Using Mobile RGB-D Data”. In: NeurIPS Datasets and Benchmarks Track (Round 1). 2021.
  14. 14.J. Belouadi, A. Lauscher, and S. Eger. “Automatikz: Text-guided synthesis of scientific vector graphics with tikz”. In: ICLR. 2024.
  15. 15.R. Birkl, D. Wofk, and M. Müller. “Midas v3. 1–a model zoo for robust monocular relative depth estimation”. In: arXiv preprint arXiv:2307.14460 (2023).
  16. 16.A. F. Biten et al. “Latr: Layout-aware transformer for scene-text vqa”. In: CVPR. 2022.
  17. 17.A. F. Biten et al. “Scene text visual question answering”. In: ICCV. 2019.
  18. 18.G. Brazil et al. “Omni3d: A large benchmark and model for 3d object detection in the wild”. In: CVPR. 2023.
  19. 19.J. Buchner. imagehash (fork). https://github.com/JohannesBuchner/imagehash. 2021.
  20. 20.H. Caesar et al. “nuscenes: A multimodal dataset for autonomous driving”. In: CVPR. 2020.
  21. 21.J. Cha et al. “Honeybee: Locality-enhanced projector for multimodal llm”. In: CVPR. 2024.
  22. 22.S. Cha et al. “Visually Dehallucinative Instruction Generation: Know What You Don’t Know”. In: arXiv preprint arXiv:2402.09717 (2024).
  23. 23.D. J. Chalmers. “Does Thought Require Sensory Grounding? From Pure Thinkers to Large Language Models”. In: Proceedings and Addresses of the American Philosophical Association 97 (2023), pp. 22–45.
  24. 24.Y. Chang et al. “A survey on evaluation of large language models”. In: ACM Transactions on Intelligent Systems and Technology 15.3 (2024), pp. 1–45.
  25. 25.G. H. Chen et al. “ALLaVA: Harnessing GPT4V-synthesized Data for A Lite Vision-Language Model”. In: arXiv preprint arXiv:2402.11684 (2024).
  26. 26.L. Chen et al. “Are We on the Right Way for Evaluating Large Vision-Language Models?” In: arXiv preprint arXiv:2403.20330 (2024).
  27. 27.L. Chen et al. “Sharegpt4v: Improving large multi-modal models with better captions”. In: arXiv preprint arXiv:2311.12793 (2023).
  28. 28.X. Chen et al. “Pali: A jointly-scaled multilingual language-image model”. In: ICLR. 2023.
  29. 29.X. Chen, S. Xie, and K. He. “An empirical study of training self-supervised vision transformers”. In: ICCV. 2021.
  30. 30.Z. Chen et al. “How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites”. In: arXiv preprint arXiv:2404.16821 (2024).
  31. 31.Z. Chen et al. “Finqa: A dataset of numerical reasoning over financial data”. In: EMNLP. 2021.
  32. 32.Z. Cheng et al. “HiTab: A hierarchical table dataset for question answering and natural language generation”. In: ACL. 2022.
  33. 33.M. Cherti et al. “Reproducible scaling laws for contrastive language-image learning”. In: CVPR. 2023.
  34. 34.W.-L. Chiang et al. “Chatbot arena: An open platform for evaluating llms by human preference”. In: arXiv preprint arXiv:2403.04132 (2024).
  35. 35.X. Chu et al. “Mobilevlm v2: Faster and stronger baseline for vision language model”. In: arXiv preprint arXiv:2402.03766 (2024).
  36. 36.M. Conover et al. Free Dolly: Introducing the World’s First Truly Open Instruction-Tuned LLM. 2023. U R L: https://www.databricks.com/blog/2023/04/12/dolly-first-open-commercially-viable-instruction-tuned-llm (visited on 06/30/2023).
  37. 37.W. Dai et al. “Instructblip: Towards general-purpose vision-language models with instruction tuning”. In: NeurIPS. 2024.
  38. 38.H. Dong et al. “Rlhf workflow: From reward modeling to online rlhf”. In: arXiv preprint arXiv:2405.07863 (2024).
  39. 39.A. Dosovitskiy et al. “An image is worth 16x16 words: Transformers for image recognition at scale”. In: ICLR. 2021.
  40. 40.A. Fang et al. “Data filtering networks”. In: ICLR. 2024.
  41. 41.X. Fu et al. “BLINK: Multimodal Large Language Models Can See but Not Perceive”. In: arXiv preprint arXiv:2404.12390 (2024).
  42. 42.S. Y. Gadre et al. “Datacomp: In search of the next generation of multimodal datasets”. In: vol. 36. 2024.
  43. 43.J. Gao et al. “G-llava: Solving geometric problem with multi-modal large language model”. In: arXiv preprint arXiv:2312.11370 (2023).
  44. 44.P. Gao et al. “Llama-adapter v2: Parameter-efficient visual instruction model”. In: arXiv preprint arXiv:2304.15010 (2023).
  45. 45.P. Gao et al. “SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models”. In: arXiv preprint arXiv:2402.05935 (2024).
  46. 46.Y. Ge et al. “Planting a seed of vision in large language model”. In: arXiv preprint arXiv:2307.08041 (2023).
  47. 47.A. Geiger, P. Lenz, and R. Urtasun. “Are we ready for Autonomous Driving? The KITTI Vision Benchmark Suite”. In: CVPR. 2012.
  48. 48.R. Geirhos et al. “Shortcut learning in deep neural networks”. In: Nature Machine Intelligence (2020).
  49. 49.R. Girshick et al. “Rich feature hierarchies for accurate object detection and semantic segmentation”. In: CVPR. 2014.
  50. 50.Google. Gemini. 2023.
  51. 51.Y. Goyal et al. “Making the v in vqa matter: Elevating the role of image understanding in visual question answering”. In: CVPR. 2017.
  52. 52.D. Gurari et al. “Vizwiz grand challenge: Answering visual questions from blind people”. In: CVPR. 2018.
  53. 53.K. He et al. “Masked autoencoders are scalable vision learners”. In: CVPR. 2022.
  54. 54.X. He et al. “PathVQA: 30000+ Questions for Medical Visual Question Answering”. In: CoRR abs/2003.10286 (2020).
  55. 55.T. Hiippala et al. “AI2D-RST: A multimodal corpus of 1000 primary school science diagrams”. In: Language Resources and Evaluation 55 (2021), pp. 661–688.
  56. 56.J. Hoffmann et al. “Training compute-optimal large language models”. In: NeurIPS (2023).
  57. 57.Y.-C. Hsiao, F. Zubach, M. Wang, et al. “Screenqa: Large-scale question-answer pairs over mobile app screenshots”. In: arXiv preprint arXiv:2209.08199 (2022).
  58. 58.D. A. Hudson and C. D. Manning. “GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering”. In: CVPR. 2019.
  59. 59.A. Jaegle et al. “Perceiver: General perception with iterative attention”. In: ICML. 2021.
  60. 60.J. Johnson et al. “Clevr: A diagnostic dataset for compositional language and elementary visual reasoning”. In: CVPR. 2017.
  61. 61.N. Jouppi et al. “Tpu v4: An optically reconfigurable supercomputer for machine learning with hardware support for embeddings”. In: Proceedings of the 50th Annual International Symposium on Computer Architecture. 2023.
  62. 62.K. Kafle et al. “Dvqa: Understanding data visualizations via question answering”. In: CVPR. 2018.
  63. 63.S. Kantharaj et al. “Chart-to-text: A large-scale benchmark for chart summarization”. In: ACL. 2022.
  64. 64.S. Karamcheti et al. “Prismatic vlms: Investigating the design space of visually-conditioned language models”. In: arXiv preprint arXiv:2402.07865 (2024).
  65. 65.M. Kazemi et al. “Geomverse: A systematic evaluation of large models for geometric reasoning”. In: 2023.
  66. 66.A. Kembhavi et al. “A diagram is worth a dozen images”. In: ECCV. 2016.
  67. 67.D. Kiela et al. “The hateful memes challenge: Detecting hate speech in multimodal memes”. In: NeurIPS. 2020.
  68. 68.G. Kim et al. “Donut: Document understanding transformer without ocr”. In: ECCV. 2022.
  69. 69.A. Kirillov et al. “Segment anything”. In: ICCV. 2023.
  70. 70.R. Krishna et al. “Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations”. In: IJCV (2016).
  71. 71.LAION. laion/gpt4v-dataset. 2023.
  72. 72.H. Laurençon, L. Tronchon, and V. Sanh. “Unlocking the conversion of Web Screenshots into HTML Code with the WebSight Dataset”. In: arXiv preprint arXiv:2403.09029 (2024).
  73. 73.H. Laurençon et al. “What matters when building vision-language models?” In: arXiv preprint arXiv:2405.02246 (2024).
  74. 74.A. C. Li et al. “Internet Explorer: Targeted Representation Learning on the Open Web”. In: ICML. 2023.
  75. 75.A. C. Li et al. “Your diffusion model is secretly a zero-shot classifier”. In: ICCV. 2023.
  76. 76.B. Li et al. LLaVA-NeXT: Stronger LLMs Supercharge Multimodal Capabilities in the Wild. 2024.
  77. 77.L. Li et al. “Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models”. In: arXiv preprint arXiv:2403.00231 (2024).
  78. 78.Y. Li et al. “Mini-gemini: Mining the potential of multi-modality vision language models”. In: arXiv preprint arXiv:2403.18814 (2024).
  79. 79.W. Lian et al. OpenOrca: An Open Dataset of GPT Augmented FLAN Reasoning Traces. https://https://huggingface.co/Open-Orca/OpenOrca. 2023.
  80. 80.T.-Y. Lin et al. “Microsoft coco: Common objects in context”. In: ECCV. 2014.
  81. 81.H. Liu et al. “Improved baselines with visual instruction tuning”. In: arXiv preprint arXiv:2310.03744 (2023).
  82. 82.H. Liu et al. LLaVA-NeXT: Improved reasoning, OCR, and world knowledge. 2024.
  83. 83.H. Liu et al. “Visual Instruction Tuning”. In: NeurIPS. 2023.
  84. 84.Y. Liu et al. “Mmbench: Is your multi-modal model an all-around player?” In: arXiv preprint arXiv:2307.06281 (2023).
  85. 85.Y. Liu et al. “On the hidden mystery of ocr in large multimodal models”. In: arXiv preprint arXiv:2305.07895 (2023).
  86. 86.Z. Liu and K. He. “A Decade’s Battle on Dataset Bias: Are We There Yet?” In: arXiv preprint arXiv:2403.08632 (2024).
  87. 87.Z. Liu et al. “A convnet for the 2020s”. In: CVPR. 2022.
  88. 88.H. Lu et al. “DeepSeek-VL: towards real-world vision-language understanding”. In: arXiv preprint arXiv:2403.05525 (2024).
  89. 89.P. Lu et al. “Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning”. In: ICLR. 2023.
  90. 90.P. Lu et al. “Iconqa: A new benchmark for abstract diagram understanding and visual language reasoning”. In: NeurIPS. 2021.
  91. 91.P. Lu et al. “Inter-GPS: Interpretable geometry problem solving with formal language and symbolic reasoning”. In: ACL. 2021.
  92. 92.P. Lu et al. “Learn to explain: Multimodal reasoning via thought chains for science question answering”. In: NeurIPS. 2022.
  93. 93.P. Lu et al. “Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts”. In: ICLR (2023).
  94. 94.Z. Luo et al. “Wizardcoder: Empowering code large language models with evol-instruct”. In: ICLR. 2024.
  95. 95.A. Majumdar et al. “OpenEQA: Embodied Question Answering in the Era of Foundation Models”. In: 2nd Workshop on Mobile Manipulation and Embodied Intelligence at ICRA 2024. 2024.
  96. 96.K. Marino et al. “OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge”. In: CVPR. 2019.
  97. 97.A. Masry et al. “Chartqa: A benchmark for question answering about charts with visual and logical reasoning”. In: ACL. 2022.
  98. 98.M. Mathew, D. Karatzas, and C. Jawahar. “Docvqa: A dataset for vqa on document images”. In: WACV. 2021.
  99. 99.B. McKinzie et al. “Mm1: Methods, analysis & insights from multimodal llm pre-training”. In: arXiv preprint arXiv:2403.09611 (2024).
  100. 100.A. Mitra et al. Orca-Math: Unlocking the potential of SLMs in Grade School Math. 2024. arXiv: 2402.14830 [cs.CL].
  101. 101.OpenAI. ChatGPT. 2022.
  102. 102.OpenAI. gpt4o. 2024.
  103. 103.M. Oquab et al. “Dinov2: Learning robust visual features without supervision”. In: TMLR (2023).
  104. 104.L. Ouyang et al. “Training language models to follow instructions with human feedback”. In: NeurIPS. 2022.
  105. 105.A. Parker. In the blink of an eye: how vision sparked the big bang of evolution. 2003.
  106. 106.P. Pasupat and P. Liang. “Compositional semantic parsing on semi-structured tables”. In: ACL. 2015.
  107. 107.J. Piaget, M. Cook, et al. The origins of intelligence in children. Vol. 8. 5. International Universities Press New York, 1952.
  108. 108.J. Pont-Tuset et al. “Connecting Vision and Language with Localized Narratives”. In: ECCV. 2020.
  109. 109.A. Radford et al. “Learning transferable visual models from natural language supervision”. In: ICML. 2021.
  110. 110.R. Rafailov et al. “Direct preference optimization: Your language model is secretly a reward model”. In: NeurIPS. 2024.
  111. 111.M. Roberts et al. “Hypersim: A Photorealistic Synthetic Dataset for Holistic Indoor Scene Understanding”. In: ICCV. 2021.
  112. 112.R. Rombach et al. “High-Resolution Image Synthesis With Latent Diffusion Models”. In: CVPR. 2022.
  113. 113.O. Russakovsky et al. “Imagenet large scale visual recognition challenge”. In: IJCV (2015).
  114. 114.O. Sanseviero. LLM Evals and Benchmarking. 2022.
  115. 115.C. Schuhmann et al. “Laion-5b: An open large-scale dataset for training next generation image-text models”. In: NeurIPS. 2022.
  116. 116.D. Schwenk et al. “A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge”. In: ECCV. 2022.
  117. 117.M. Shridhar et al. “ALFWorld: Aligning Text and Embodied Environments for Interactive Learning”. In: ICLR. 2021.
  118. 118.C. Si et al. “Design2Code: How Far Are We From Automating Front-End Engineering?” In: arXiv preprint arXiv:2403.03163 (2024).
  119. 119.O. Sidorov et al. TextCaps: a Dataset for Image Captioning with Reading Comprehension. 2020. arXiv: 2003.12462 [cs.CV].
  120. 120.A. Singh et al. “Towards vqa models that can read”. In: CVPR. 2019.
  121. 121.S. Song, S. P. Lichtenberg, and J. Xiao. “Sun rgb-d: A rgb-d scene understanding benchmark suite”. In: CVPR. 2015.
  122. 122.Q. Sun et al. “Eva-clip: Improved training techniques for clip at scale”. In: arXiv preprint arXiv:2303.15389 (2023).
  123. 123.R. Tanaka, K. Nishida, and S. Yoshida. “VisualMRC: Machine Reading Comprehension on Document Images”. In: AAAI. 2021.
  124. 124.B. J. Tang, A. Boggust, and A. Satyanarayan. “Vistext: A benchmark for semantically rich chart captioning”. In: arXiv preprint arXiv:2307.05356 (2023).
  125. 125.S. Tong, E. Jones, and J. Steinhardt. “Mass-producing failures of multimodal systems with language models”. In: NeurIPS. 2024.
  126. 126.S. Tong et al. “Eyes wide shut? exploring the visual shortcomings of multimodal llms”. In: CVPR. 2024.
  127. 127.H. Touvron et al. “LLaMA 2: Open foundation and fine-tuned chat models”. In: (2023).
  128. 128.H. Touvron et al. “LLaMA: Open and efficient foundation language models”. In: arXiv preprint arXiv:2302.13971 (2023).
  129. 129.H. Tu et al. “How many unicorns are in this image? a safety evaluation benchmark for vision llms”. In: arXiv preprint arXiv:2311.16101 (2023).
  130. 130.K. Vishniakov, Z. Shen, and Z. Liu. “ConvNet vs Transformer, Supervised vs CLIP: Beyond ImageNet Accuracy”. In: ICML. 2024.
  131. 131.J. Wang et al. “To see is to believe: Prompting gpt-4v for better visual instruction tuning”. In: arXiv preprint arXiv:2311.07574 (2023).
  132. 132.K. Wang et al. “Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset”. In: arXiv preprint arXiv:2402.14804 (2024).
  133. 133.J. Wei et al. “Chain-of-thought prompting elicits reasoning in large language models”. In: NeurIPS. 2022.
  134. 134.C. Wendler. wendlerc/RenderedText. 2023.
  135. 135.H. Wu et al. “Q-instruct: Improving low-level visual abilities for multi-modality foundation models”. In: arXiv preprint arXiv:2311.06783 (2023).
  136. 136.P. Wu and S. Xie. “V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs”. In: CVPR. 2024.
  137. 137.xAI. grok. 2024.
  138. 138.H. Xu et al. “Demystifying clip data”. In: ICLR. 2024.
  139. 139.A. Young et al. “Yi: Open foundation models by 01. ai”. In: arXiv preprint arXiv:2403.04652 (2024).
  140. 140.L. Yu et al. Modeling Context in Referring Expressions. 2016. arXiv: 1608.00272 [cs.CV].
  141. 141.T. Yu et al. “Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback”. In: arXiv preprint arXiv:2312.00849 (2023).
  142. 142.X. Yue et al. “Mammoth: Building math generalist models through hybrid instruction tuning”. In: ICLR. 2024.
  143. 143.X. Yue et al. “Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi”. In: CVPR. 2024.
  144. 144.M. Yuksekgonul et al. “When and why vision-language models behave like bags-of-words, and what to do about it?” In: ICLR. 2022.
  145. 145.X. Zhai et al. “Sigmoid loss for language image pre-training”. In: ICCV. 2023.
  146. 146.Y. Zhai et al. “Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning”. In: arXiv preprint arXiv:2405.10292 (2024).
  147. 147.Y. Zhai et al. “Investigating the catastrophic forgetting in multimodal large language models”. In: CPAL. 2024.
  148. 148.C. Zhang et al. “Raven: A dataset for relational and analogical visual reasoning”. In: CVPR. 2019.
  149. 149.Y. Zhang et al. “Llavar: Enhanced visual instruction tuning for text-rich image understanding”. In: arXiv preprint arXiv:2306.17107 (2023).
  150. 150.Y. Zhao et al. “Pytorch fsdp: experiences on scaling fully sharded data parallel”. In: arXiv preprint arXiv:2304.11277 (2023).
  151. 151.L. Zheng et al. “Judging llm-as-a-judge with mt-bench and chatbot arena”. In: NeurIPS. 2024.
  152. 152.T. Zheng et al. “OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement”. In: arXiv preprint arXiv:2402.14658 (2024).
  153. 153.V. Zhong, C. Xiong, and R. Socher. “Seq2sql: Generating structured queries from natural language using reinforcement learning”. In: 2017.
  154. 154.B. Zhou et al. “Semantic understanding of scenes through the ade20k dataset”. In: IJCV (2019).
  155. 155.K. Zhou et al. “Don’t Make Your LLM an Evaluation Benchmark Cheater”. In: arXiv preprint arXiv:2311.01964 (2023).
  156. 156.B. Zhu et al. Starling-7b: Improving llm helpfulness & harmlessness with rlaif. 2023.
  157. 157.D. Zhu et al. “Minigpt-4: Enhancing vision-language understanding with advanced large language models”. In: arXiv preprint arXiv:2304.10592 (2023).
  158. 158.F. Zhu et al. “TAT-QA: A question answering benchmark on a hybrid of tabular and textual content in finance”. In: ACL. 2021.
  159. 159.Y. Zhu et al. “Visual7w: Grounded question answering in images”. In: CVPR. 2016.

Citation

MLA
Tong, S., et al. “Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs”. arXiv, 2024, http://arxiv.org/abs/2406.16860v2.
APA
Tong, S., Brown, E., Wu, P., Woo, S., Middepogu, M., Akula, S. C., Yang, J., Yang, S., Iyer, A., Pan, X., Wang, Z., Fergus, R., LeCun, Y., & Xie, S. (2024). Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs. arXiv. http://arxiv.org/abs/2406.16860v2
Chicago
Tong, S., E. Brown, P. Wu, et al. 2024. “Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs”. arXiv. http://arxiv.org/abs/2406.16860v2.
Harvard
Tong, S. et al. (2024) “Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2406.16860v2.
Vancouver
1. Tong S, Brown E, Wu P, et al (2024) Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs. arXiv

BibTeX

@article{tong2024cambrian,
  title = {Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs},
  author = {Tong, Shengbang and Brown, Ellis and Wu, Penghao and Woo, Sanghyun and Middepogu, Manoj and Akula, Sai Charitha and Yang, Jihan and Yang, Shusheng and Iyer, Adithya and Pan, Xichen and Wang, Ziteng and Fergus, Rob and LeCun, Yann and Xie, Saining},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2406.16860v2},
  eprint = {2406.16860}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors