InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Jinguo ZhuWeiyun WangZhe ChenZhaoyang LiuShenglong YeLixin GuYuchen DuanHao TianWeijie SuJie Shao

article2025arXiv1,820 citations

Presents InternVL3, a native multimodal model trained jointly on visual and textual corpora that matches leading proprietary systems on the MMMU benchmark by combining variable visual position encoding with advanced test-time scaling.

Listen

Modern multimodal artificial intelligence systems, which process both text and images, typically adapt pre-existing, text-only large language models through complex, multi-stage training pipelines. This post-hoc adaptation often introduces significant cross-modality alignment challenges, requires delicate parameter-freezing schedules, and risks degrading core language competencies. As demand grows for unified models capable of handling multidisciplinary reasoning, complex documents, and real-world visual environments, developing more integrated, resource-efficient training paradigms has become essential.

The article introduces and evaluates InternVL3, an open-source family of multimodal large language models ranging from 1 billion to 78 billion parameters. The main objective is to demonstrate that a native pre-training paradigm—which jointly optimizes linguistic and visual representations in a single initial stage—delivers state-of-the-art multimodal performance while preserving and enhancing pure-text capabilities.

To achieve this, the authors implemented a unified pre-training architecture that updates all model parameters simultaneously using an optimal 1:3 ratio of pure text (50 billion tokens) to multimodal data (150 billion tokens). The model incorporates Variable Visual Position Encoding to support longer contexts, alongside advanced post-training involving supervised fine-tuning across 21.7 million samples and Mixed Preference Optimization. The evaluation relied on comprehensive benchmarking across standard multimodal test suites, language understanding suites, and real-world vision tasks, supported by an optimized distributed training framework that delivered training speedups of 50% to 200%.

The key findings highlight substantial performance advantages across several domains. First, the largest variant, InternVL3-78B, established a new open-source state of the art with a 72.2 score on the multidisciplinary MMMU benchmark, matching or exceeding leading proprietary systems like ChatGPT-4o and Claude 3.5 Sonnet. Second, the series demonstrated strong mathematical and logical reasoning, scoring 79.0 on MathVista, which rose to 80.5 when applying test-time scaling with a process reward critic model. Third, the models achieved superior results on text-rich and practical tasks, including a score of 906 on OCRBench and an 88.7% accuracy on graphical user interface grounding. Finally, purely textual evaluation confirmed that joint multimodal pre-training improved text and coding capabilities compared to baseline chat models derived from the same base text architecture.

These findings imply that joint multimodal pre-training eliminates the need for cumbersome, multi-stage alignment pipelines, lowering integration overhead without sacrificing language quality. For organizations deploying computer vision and language solutions, InternVL3 demonstrates that open-source architectures can rival costly closed-source alternatives across document processing, user interface automation, and spatial analysis.

Based on these results, decision-makers should consider adopting or piloting the publicly released open-source weights and training data for enterprise workflows requiring multimodal understanding. Further work should explore applying the model to dynamic, interactive user interface agents and specialized spatial-temporal environments.

Regarding limitations, the article notes that larger model variants exhibited plateaus in visual grounding performance due to a relative reduction in grounding-specific pre-training data. Additionally, closed-source models such as Gemini 2.5 Pro still retain narrow performance edges in specific areas like video-context comprehension and visual hallucination reduction. Stakeholders should maintain moderate caution when deploying the system in critical tasks where fine-grained object localization is paramount.

Cover for InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Abstract

We introduce InternVL3, a significant advancement in the InternVL series featuring a native multimodal pre-training paradigm. Rather than adapting a text-only large language model (LLM) into a multimodal large language model (MLLM) that supports visual inputs, InternVL3 jointly acquires multimodal and linguistic capabilities from both diverse multimodal data and pure-text corpora during a single pre-training stage. This unified training paradigm effectively addresses the complexities and alignment challenges commonly encountered in conventional post-hoc training pipelines for MLLMs. To further improve performance and scalability, InternVL3 incorporates variable visual position encoding (V2PE) to support extended multimodal contexts, employs advanced post-training techniques such as supervised fine-tuning (SFT) and mixed preference optimization (MPO), and adopts test-time scaling strategies alongside an optimized training infrastructure. Extensive empirical evaluations demonstrate that InternVL3 delivers superior performance across a wide range of multi-modal tasks. In particular, InternVL3-78B achieves a score of 72.2 on the MMMU benchmark, setting a new state-of-the-art among open-source MLLMs. Its capabilities remain highly competitive with leading proprietary models, including ChatGPT-4o, Claude 3.5 Sonnet, and Gemini 2.5 Pro, while also maintaining strong pure-language proficiency. In pursuit of open-science principles, we will publicly release both the training data and model weights to foster further research and development in next-generation MLLMs.

Table of Contents

  • 1 Introduction
  • 2 InternVL3
  • 2.1 Model Architecture
  • 2.2 Native Multimodal Pre-Training
  • 2.3 Post-Training
  • 2.4 Test-Time Scaling
  • 2.5 Infrastructure
  • 3 Experiments
  • 3.1 Overall Comparison to Other Advanced MLLMs
  • 3.2 Multimodal Reasoning and Mathematics
  • 3.3 OCR, Chart, and Document Understanding
  • 3.4 Multi-Image Understanding
  • 3.5 Real-World Comprehension
  • 3.6 Comprehensive Multimodal Evaluation
  • 3.7 Multimodal Hallucination Evaluation
  • 3.8 Visual Grounding
  • 3.9 Multimodal Multilingual Understanding
  • 3.10 Video Understanding
  • 3.11 GUI Grounding
  • 3.12 Spatial Reasoning
  • 3.13 Evaluation on Language Capability
  • 3.14 Ablation Study
  • 4 Conclusion
  • References

Knowls

  1. Knowl 1 — Native Multimodal Pre-Training Paradigm in InternVL3

    model/method

    Unlike conventional multimodal large language model (MLLM) training pipelines that start with a pre-trained, text-only LLM and adapt it via multi-stage alignment or parameter freezing, InternVL3 uses a native multimodal pre-training strategy. In this paradigm, all model parameters across the vision transformer (InternViT), the 2-layer MLP projector, and the language model (initialized directly from pre-trained base LLMs without instruction tuning) are updated jointly from the start on a mixture of large-scale text-only corpora and diverse multimodal datasets (image-text, video-text, and interleaved sequences).

    To balance modalities, the pre-training tokens follow a 1:3 ratio of pure-language data to multimodal data over a 200-billion-token budget (50 billion language tokens and 150 billion multimodal tokens). Gradients are computed strictly over textual tokens, using visual tokens solely as conditioning context. The training objective is:

    Ltext-only(θ)=−∑i=2,xi∈TextLwi⋅log⁡pθ(xi∣x1,…,xi−1)\mathcal{L}_{\text{text-only}}(\theta) = -\sum_{i=2, x_i \in \text{Text}}^L w_i \cdot \log p_\theta(x_i \mid x_1, \dots, x_{i-1})

    where θ\theta denotes model parameters, x=(x1,…,xL)x = (x_1, \dots, x_L) is a sequence of length LL, and wiw_i is a square-averaging token weight defined as:

    wi=1l0.5w_i = \frac{1}{l^{0.5}}

    where ll is the number of text tokens in the sample on which loss is computed. Square averaging prevents gradient bias toward excessively long or short sequences.

  2. Knowl 2 — Variable Visual Position Encoding (V2PE)

    model/method

    To support long multimodal contexts without exceeding the positional index window of the underlying language model, InternVL3 incorporates Variable Visual Position Encoding (V2PE). While standard MLLMs assign position indices pip_i that increment uniformly by 11 for all tokens regardless of modality, V2PE applies a modality-specific recursive update:

    pi=pi−1+{1,if xi is a textual token,δ,if xi is a visual tokenp_i = p_{i-1} + \begin{cases} 1, & \text{if } x_i \text{ is a textual token}, \\ \delta, & \text{if } x_i \text{ is a visual token} \end{cases}

    with p1=0p_1 = 0 and δ≤1\delta \le 1.

    Key properties of V2PE:

    • Within a single image, the step size δ\delta is kept constant to preserve intra-image 2D spatial relationships.
    • During training, δ\delta is randomly sampled for each image from the discrete set of fractional values:

    δ∈Δ={1,12,14,18,116,132,164,1128,1256}\delta \in \Delta = \left\{ 1, \frac{1}{2}, \frac{1}{4}, \frac{1}{8}, \frac{1}{16}, \frac{1}{32}, \frac{1}{64}, \frac{1}{128}, \frac{1}{256} \right\}

    • During inference, δ\delta can be dynamically selected based on input context length. Setting δ=1\delta = 1 reverts the mechanism to standard 1D sequential positional encoding.
  3. Knowl 3 — Mixed Preference Optimization (MPO) Framework

    model/method

    To resolve the exposure bias and distribution shift between ground-truth training tokens and autoregressively generated tokens during Chain-of-Thought (CoT) reasoning, InternVL3 applies Mixed Preference Optimization (MPO) in its second post-training phase following supervised fine-tuning (SFT). The MPO objective combines preference loss Lp\mathcal{L}_p, quality loss Lq\mathcal{L}_q, and language model generation loss Lg\mathcal{L}_g:

    LMPO=wpLp+wqLq+wgLg\mathcal{L}_{\text{MPO}} = w_p \mathcal{L}_p + w_q \mathcal{L}_q + w_g \mathcal{L}_g

    where wp,wq,wg∈R+w_p, w_q, w_g \in \mathbb{R}^+ weight each component.

    1. Preference Loss (Direct Preference Optimization, DPO): Captures relative ranking between chosen response ycy_c and rejected response yry_r given prompt xx:

    Lp=−log⁡σ(βlog⁡πθ(yc∣x)π0(yc∣x)−βlog⁡πθ(yr∣x)π0(yr∣x))\mathcal{L}_p = -\log \sigma \left( \beta \log \frac{\pi_\theta(y_c \mid x)}{\pi_0(y_c \mid x)} - \beta \log \frac{\pi_\theta(y_r \mid x)}{\pi_0(y_r \mid x)} \right)

    where σ\sigma is the sigmoid function, β\beta is the KL penalty coefficient, πθ\pi_\theta is the policy model, and π0\pi_0 is the reference model initialized from the SFT stage.

    1. Quality Loss (Binary Classifier Optimization, BCO): Enforces absolute quality discrimination for chosen and rejected responses independently:

    Lq=Lq++Lq−\mathcal{L}_q = \mathcal{L}_q^+ + \mathcal{L}_q^- Lq+=−log⁡σ(βlog⁡πθ(yc∣x)π0(yc∣x)−δreward)\mathcal{L}_q^+ = -\log \sigma \left( \beta \log \frac{\pi_\theta(y_c \mid x)}{\pi_0(y_c \mid x)} - \delta_{\text{reward}} \right) Lq−=−log⁡σ(−(βlog⁡πθ(yr∣x)π0(yr∣x)−δreward))\mathcal{L}_q^- = -\log \sigma \left( -\left( \beta \log \frac{\pi_\theta(y_r \mid x)}{\pi_0(y_r \mid x)} - \delta_{\text{reward}} \right) \right)

    where δreward\delta_{\text{reward}} is a reward shift parameter tracking the running moving average of previous rewards to stabilize training.

    1. Generation Loss (LM Loss): Standard cross-entropy loss applied exclusively to chosen responses ycy_c to maintain fluent generative ability.
  4. Knowl 4 — Visual Process Reward Model (VisualPRM) and Best-of-N Test-Time Scaling

    model/method

    InternVL3 integrates test-time scaling for complex multimodal reasoning and mathematics using a Best-of-NN (BoN) sampling strategy evaluated by VisualPRM-8B as the step-level verifier.

    VisualPRM formulates step-level process supervision as a multi-turn chat task. Given an image II, a question qq, and a candidate step-by-step solution s={s0,s1,…,sn}s = \{s_0, s_1, \dots, s_n\}:

    1. The image II, question qq, and initial step s0s_0 are fed into turn 1.
    2. In each subsequent turn ii, step sis_i is presented, and the model predicts step correctness ci∈{+,−}c_i \in \{+, -\}:

    ci∼M(yi∣I,q,s≤i)c_i \sim \mathcal{M}(y_i \mid I, q, s_{\le i})

    1. During inference, the quality score for step ii is the predicted probability of emitting the token "++":

    Score(si)=P(ci=′+′∣I,q,s≤i)\text{Score}(s_i) = P(c_i = '+' \mid I, q, s_{\le i})

    1. The overall solution score is the arithmetic mean of all step scores: 1n+1∑i=0nScore(si)\frac{1}{n+1} \sum_{i=0}^n \text{Score}(s_i). The highest-scoring candidate among NN rollouts (e.g., N=8N=8) is selected as the final answer.
  5. Knowl 5 — InternVL3 Model Configurations and Specifications

    data/table

    The InternVL3 family follows the ViT-MLP-LLM architecture. Visual inputs are encoded by InternViT variants, processed by a pixel unshuffle operation reducing 448×448448 \times 448 pixel tiles by a factor of 4 into 256 visual tokens, and mapped into the LLM embedding space via a randomly initialized 2-layer MLP.

    Model Name Total Parameters Vision Encoder Language Model OpenCompass Academic
    InternVL3-1B 0.9B InternViT-300M-448px-V2.5 Qwen2.5-0.5B 57.4
    InternVL3-2B 1.9B InternViT-300M-448px-V2.5 Qwen2.5-1.5B 63.9
    InternVL3-8B 8.1B InternViT-300M-448px-V2.5 Qwen2.5-7B 73.3
    InternVL3-9B 9.2B InternViT-300M-448px-V2.5 InternLM3-8B 72.4
    InternVL3-14B 15.1B InternViT-300M-448px-V2.5 Qwen2.5-14B 75.5
    InternVL3-38B 38.4B InternViT-6B-448px-V2.5 Qwen2.5-32B 77.3
    InternVL3-78B 78.4B InternViT-6B-448px-V2.5 Qwen2.5-72B 79.5

    All language models are initialized strictly from base models rather than instruction-tuned variants.

  6. Knowl 6 — InternEVO Distributed Training Infrastructure for MLLMs

    model/method

    To train dense vision-language models up to 78.4B parameters across thousands of GPUs, the InternEVO framework was extended with multimodal-specific distributed optimizations:

    1. Decoupled Module Sharding: The ViT, MLP, and LLM modules use decoupled ZeRO sharding and communication strategies, allowing computation and communication to overlap across disparate module architectures.
    2. Dynamic Workload Balancing: Addresses the load imbalance caused by fluctuating ratios of visual and textual tokens per batch by dynamically distributing computational workloads across modules.
    3. Hybrid Parallelism for Long Contexts: Incorporates combinations of data parallelism, tensor parallelism, pipeline parallelism, head parallelism, and sequence parallelism to support context lengths up to 32K tokens.
    4. Efficiency Gains: Yields a training speedup of 50% to 200% over the InternVL2.5 infrastructure under identical compute budgets.
  7. Knowl 7 — Multimodal Reasoning and Mathematics Benchmark Performance

    empirical result

    InternVL3 achieves state-of-the-art results among open-source MLLMs on multimodal reasoning and mathematical benchmarks, with additional gains when paired with Best-of-8 (Bo8) test-time scaling via VisualPRM.

    Key benchmark results:

    • MMMU: InternVL3-78B scores 72.2 (vs. InternVL2.5-78B at 70.0, Qwen2.5-VL-72B at 68.2, Claude-3.5 Sonnet at 66.4, ChatGPT-4o at 72.9, and Gemini-2.5 Pro at 74.7).
    • MathVista: InternVL3-8B achieves 71.6 (75.2 w/ Bo8); InternVL3-78B achieves 79.0 (80.5 w/ Bo8), compared to 74.8 for Qwen2.5-VL-72B and 71.6 for ChatGPT-4o-latest.
    • MathVision: InternVL3-78B scores 43.1 (vs. 39.3 for Qwen2.5-VL-72B and 31.2 for GPT-4o-20241120).
    • MathVerse (Vision-Only split): InternVL3-38B scores 48.2 (54.2 w/ Bo8, a +6.0 boost); InternVL3-78B scores 51.0 (54.2 w/ Bo8, a +3.2 boost).
    • Overall Reasoning Mean: Across MMMU, MathVista, MathVision, MathVerse, DynaMath, WeMath, and LogicVista, InternVL3-78B reaches an overall average of 54.6 (56.5 w/ Bo8), outperforming Qwen2.5-VL-72B (52.8) and GPT-4o-20241120 (47.9).
  8. Knowl 8 — OCR, Chart, and Document Understanding Performance

    empirical result

    Across nine text-centric and document benchmarks, InternVL3 demonstrates strong visual-textual perception:

    • OCRBench: InternVL3-8B reaches 880; InternVL3-78B reaches 906, outperforming Qwen2.5-VL-72B (885), GPT-4o (894), and Gemini-2.5 Pro (862).
    • DocVQA: InternVL3-8B achieves 92.7%; InternVL3-78B achieves 95.4% (vs. Claude-3.5 Sonnet at 95.2% and Qwen2.5-VL-72B at 96.4%).
    • ChartQA: InternVL3-8B achieves 86.6%; InternVL3-78B reaches 89.7% (vs. Qwen2.5-VL-72B at 89.5% and Claude-3.5 Sonnet at 90.8%).
    • AI2D: InternVL3-78B achieves 89.7% / 96.0% (with / without mark), outperforming Claude-3.5 Sonnet (81.2% / 94.7%) and GPT-4o (84.6% / 94.2%).
    • VCR (Visual Caption Restoration): InternVL3-78B achieves 96.0 / 98.6 (EM / Jaccard).
    • CharXiv: InternVL3-78B scores 46.0 / 85.1 (RQ / DQ).
  9. Knowl 9 — Multi-Image, Video, and Real-World Comprehension Performance

    empirical result

    InternVL3 establishes strong cross-image, video, and real-world evaluation results:

    1. Multi-Image Understanding:
      • On BLINK (val), InternVL3-78B scores 66.3 (vs. GPT-4o-20240513 at 68.0 and Qwen2.5-VL-72B at 64.4).
      • On MMT-Bench (val), InternVL3-78B scores 73.2 (vs. InternVL2.5-78B at 70.8 and GPT-4o at 65.4).
      • On Mantis-Eval, InternVL3-2B reaches 65.9 (+11.1 over InternVL2.5-2B), and InternVL3-78B scores 79.3.
    2. Real-World Benchmarks:
      • On RealWorldQA, InternVL3-78B scores 78.0 (vs. GPT-4o at 75.4 and Qwen2.5-VL-72B at 75.7).
      • On MME-RealWorld (EN), InternVL3-78B achieves 65.4, substantially outperforming GPT-4o (45.2) and Claude-3.5 Sonnet (51.6).
      • On R-Bench, InternVL3-78B achieves 77.4.
    3. Video Understanding:
      • On Video-MME (with subtitles), InternVL3-78B reaches 75.7% (72.7% without subtitles).
      • On MVBench, InternVL3-78B achieves 78.7 (vs. Qwen2.5-VL-72B at 70.4 and Gemini-1.5-Pro at 81.3).
      • On LongVideoBench, InternVL3-78B achieves 65.7%.
  10. Knowl 10 — GUI Grounding and 3D Spatial Intelligence (VSI-Bench)

    empirical result

    InternVL3 demonstrates enhanced grounding on user interfaces and 3D spatial environments:

    1. GUI Grounding:
      • On ScreenSpot, InternVL3-78B achieves 88.7% accuracy, outperforming Qwen2.5-VL-72B (87.1%), UI-TARS-72B (88.4%), and Gemini 2.0 (84.0%). InternVL3-38B achieves 85.6% vs. GPT-4o (18.1%).
      • On ScreenSpot-V2, InternVL3-8B achieves 81.4%, InternVL3-38B reaches 88.3%, and InternVL3-78B reaches 90.9% (surpassing UI-TARS-72B at 90.3%).
    2. 3D Spatial Reasoning (VSI-Bench):
      • InternVL3-8B scores 42.1 overall, leading all open-source MLLMs.
      • InternVL3-38B and InternVL3-78B score 48.9 and 48.4 overall, respectively, outperforming GPT-4o (34.0), Gemini-1.5 Flash, and Gemini-1.5 Pro (45.4).
      • Sub-task performance for InternVL3-78B on VSI-Bench includes 71.2 on object counting, 53.7 on absolute distance estimation, 55.9 on relative distance estimation, and 54.5 on appearance order prediction.
  11. Knowl 11 — Language Capability Preservation in Native Multimodal Training

    empirical result

    Rather than suffering from catastrophic forgetting or linguistic degradation, InternVL3 models exhibit higher language capabilities than their corresponding text-only chat counterparts (Qwen2.5 Chat models) initialized from identical base weights.

    Key comparisons across standard NLP benchmarks:

    • MMLU: InternVL3-8B scores 77.3 vs. Qwen2.5-7B Chat (74.2); InternVL3-78B scores 86.9 vs. Qwen2.5-72B Chat (84.4).
    • CMMLU / C-Eval: InternVL3-78B scores 89.9 / 89.5 vs. Qwen2.5-72B Chat (87.4 / 88.1).
    • GSM8K: InternVL3-8B scores 83.1 vs. Qwen2.5-7B Chat (80.1); InternVL3-78B scores 90.5 vs. Qwen2.5-72B Chat (88.2).
    • Overall NLP Score: InternVL3-1B achieves 42.4 (vs. 33.5 for Qwen2.5-0.5B Chat); InternVL3-8B achieves 72.9 (vs. 69.4 for Qwen2.5-7B Chat); InternVL3-78B achieves 80.5 (vs. 78.9 for Qwen2.5-72B Chat).

    This preservation and improvement are attributed to the ~25% high-quality pure-language pre-training corpus, full parameter joint optimization, and pure-text data included in SFT.

  12. Knowl 12 — Ablation on Positional Increment Step $\delta$ and MPO Post-Training

    empirical result

    Ablation experiments quantify the isolated contributions of V2PE step sizes and MPO fine-tuning:

    1. Positional Increment δ\delta in V2PE (InternVL3-8B Pre-trained): Evaluating across 12 standard benchmarks (TextVQA, VizWiz, ChartQA, DocVQA, AI2D, InfoVQA, GQA, SQA-I, POPE, Tiny LVLM, MMMU val, SEED v1 image):

      • Standard fixed encoding (without V2PE / δ=1\delta = 1): 75.2 overall score.
      • V2PE with δ=1/256\delta = 1/256: 75.0 overall.
      • V2PE with δ=1/64\delta = 1/64: 75.3 overall.
      • V2PE with δ=1/16\delta = 1/16: 75.6 overall.
      • V2PE with δ=1/4\delta = 1/4: 75.9 overall (optimal, with DocVQA at 91.0, AI2D at 81.8, InfoVQA at 71.7, GQA at 61.2).
      • V2PE with δ=1/1\delta = 1/1: 75.7 overall.
    2. Impact of Mixed Preference Optimization (MPO): Using ~300K preference rollout samples generated by SFT variants (a strict subset of SFT data), MPO consistently boosts reasoning across all scales:

      • InternVL3-1B: 24.6 →\rightarrow 25.1 (+0.5)
      • InternVL3-2B: 30.7 →\rightarrow 32.4 (+1.7)
      • InternVL3-8B: 41.4 →\rightarrow 44.3 (+2.9)
      • InternVL3-9B: 41.6 →\rightarrow 43.1 (+1.5)
      • InternVL3-14B: 46.2 →\rightarrow 49.9 (+3.7)
      • InternVL3-38B: 48.3 →\rightarrow 52.8 (+4.5)
      • InternVL3-78B: 50.5 →\rightarrow 54.6 (+4.1)

Coverage note — None was omitted; all primary contributions—native multimodal pre-training, V2PE, MPO, VisualPRM test-time scaling, model configurations, InternEVO infrastructure, cross-benchmark evaluation results, and ablation studies—are fully represented.

References

  1. 1.Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  2. 2.Anonymous. CG-bench: Clue-grounded question answering benchmark for long video understanding. In Submitted to The Thirteenth International Conference on Learning Representations, 2024. under review.
  3. 3.Anthropic. The claude 3 model family: Opus, sonnet, haiku. https://www.anthropic.com, 2024.
  4. 4.Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al. Program synthesis with large language models. arXiv preprint arXiv:2108.07732, 2021.
  5. 5.Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-vl: A frontier large vision-language model with versatile abilities. arXiv preprint arXiv:2308.12966, 2023.
  6. 6.Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, et al. Qwen2. 5-vl technical report. arXiv preprint arXiv:2502.13923, 2025.
  7. 7.Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, et al. Qwen2.5-vl technical report. arXiv preprint arXiv:2502.13923, 2025.
  8. 8.Loubna Ben Allal, Anton Lozhkov, Guilherme Penedo, Thomas Wolf, and Leandro von Werra. Smollm-corpus, 2024.
  9. 9.Ali Furkan Biten, Ruben Tito, Andres Mafla, Lluis Gomez, Marçal Rusinol, Ernest Valveny, CV Jawahar, and Dimosthenis Karatzas. Scene text visual question answering. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4291–4301, 2019.
  10. 10.Jie Cao and Jing Xiao. An augmented benchmark dataset for geometric question answering through dual parallel text encoding. In Proceedings of the 29th International Conference on Computational Linguistics, pages 1511–1520, 2022.
  11. 11.Shuaichen Chang, David Palzer, Jialin Li, Eric Fosler-Lussier, and Ningchuan Xiao. Mapqa: A dataset for question answering on choropleth maps. arXiv preprint arXiv:2211.08545, 2022.
  12. 12.Keqin Chen, Zhao Zhang, Weili Zeng, Richong Zhang, Feng Zhu, and Rui Zhao. Shikra: Unleashing multimodal llm’s referential dialogue magic. arXiv preprint arXiv:2306.15195, 2023.
  13. 13.Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Jiaqi Wang, Yu Qiao, Dahua Lin, et al. Are we on the right way for evaluating large vision-language models? arXiv preprint arXiv:2403.20330, 2024.
  14. 14.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021.
  15. 15.Qiaoling Chen, Diandian Gu, Guoteng Wang, Xun Chen, YingTong Xiong, Ting Huang, Qinghao Hu, Xin Jin, Yonggang Wen, Tianwei Zhang, et al. Internevo: Efficient long-sequence large language model training via hybrid parallelism and redundant sharding. arXiv preprint arXiv:2401.09149, 2024.
  16. 16.Qiguang Chen, Libo Qin, Jin Zhang, Zhi Chen, Xiao Xu, and Wanxiang Che. M3cot: A novel benchmark for multi-domain multi-step multi-modal chain-of-thought. arXiv preprint arXiv:2405.16473, 2024.
  17. 17.Wenhu Chen, Ming Yin, Max Ku, Pan Lu, Yixin Wan, Xueguang Ma, Jianyu Xu, Xinyi Wang, and Tony Xia. Theoremqa: A theorem-driven question answering dataset. In Houda Bouamor, Juan Pino, and Kalika Bali, editors, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023, pages 7889–7901. Association for Computational Linguistics, 2023.
  18. 18.Zhe Chen, Weiyun Wang, Yue Cao, Yangzhou Liu, Zhangwei Gao, Erfei Cui, Jinguo Zhu, Shenglong Ye, Hao Tian, Zhaoyang Liu, et al. Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling. arXiv preprint arXiv:2412.05271, 2024.
  19. 19.Zhe Chen, Weiyun Wang, Hao Tian, Shenglong Ye, Zhangwei Gao, Erfei Cui, Wenwen Tong, Kongzhi Hu, Jiapeng Luo, Zheng Ma, et al. How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites. arXiv preprint arXiv:2404.16821, 2024.
  20. 20.Zhe Chen, Weiyun Wang, Hao Tian, Shenglong Ye, Zhangwei Gao, Erfei Cui, Wenwen Tong, Kongzhi Hu, Jiapeng Luo, Zheng Ma, et al. How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites. arXiv preprint arXiv:2404.16821, 2024.
  21. 21.Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 24185–24198, 2024.
  22. 22.Kanzhi Cheng, Qiushi Sun, Yougang Chu, Fangzhi Xu, Yantao Li, Jianbing Zhang, and Zhiyong Wu. Seeclick: Harnessing gui grounding for advanced visual gui agents. arXiv preprint arXiv:2401.10935, 2024.
  23. 23.Zesen Cheng, Sicong Leng, Hang Zhang, Yifei Xin, Xin Li, Guanzheng Chen, Yongxin Zhu, Wenqi Zhang, Ziyang Luo, Deli Zhao, et al. Videollama 2: Advancing spatial-temporal modeling and audio understanding in video-llms. arXiv preprint arXiv:2406.07476, 2024.
  24. 24.Christopher Clark and Matt Gardner. Simple and effective multi-paragraph reading comprehension. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, pages 845–855, 2018.
  25. 25.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.
  26. 26.OpenCompass Contributors. Opencompass: A universal evaluation platform for foundation models. https://github.com/open-compass/opencompass, 2023.
  27. 27.X.AI Corp. Grok-1.5 vision preview: Connecting the digital and physical worlds with our first multimodal model. https://x.ai/blog/grok-1.5v, 2024.
  28. 28.Wenliang Dai, Nayeon Lee, Boxin Wang, Zhuolin Yang, Zihan Liu, Jon Barker, Tuomas Rintamaki, Mohammad Shoeybi, Bryan Catanzaro, and Wei Ping. Nvlm: Open frontier-class multimodal llms. arXiv preprint arXiv:2409.11402, 2024.
  29. 29.Google Deepmind. Gemini 2.0 is now available to everyone. https://blog.google/technology/google-deepmind/gemini-model-updates-february-2025/, 202.
  30. 30.Google Deepmind. Introducing gemini 2.0: our new ai model for the agentic era. https://blog.google/technology/google-deepmind/google-gemini-ai-update-december-2024/, 2024.
  31. 31.Matt Deitke, Christopher Clark, Sangho Lee, Rohun Tripathi, Yue Yang, Jae Sung Park, Mohammadreza Salehi, Niklas Muennighoff, Kyle Lo, Luca Soldaini, et al. Molmo and pixmo: Open weights and open data for state-of-the-art multimodal models. arXiv preprint arXiv:2409.17146, 2024.
  32. 32.Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Bin Wang, Linke Ouyang, Songyang Zhang, Haodong Duan, Wenwei Zhang, Yining Li, et al. Internlm-xcomposer2-4khd: A pioneering large vision-language model handling resolutions from 336 pixels to 4k hd. arXiv preprint arXiv:2404.06512, 2024.
  33. 33.Haodong Duan, Junming Yang, Yuxuan Qiao, Xinyu Fang, Lin Chen, Yuan Liu, Xiaoyi Dong, Yuhang Zang, Pan Zhang, Jiaqi Wang, et al. Vlmevalkit: An open-source toolkit for evaluating large multi-modality models. In Proceedings of the 32nd ACM International Conference on Multimedia, pages 11198–11201, 2024.
  34. 34.Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024.
  35. 35.Xinyu Fang, Kangrui Mao, Haodong Duan, Xiangyu Zhao, Yining Li, Dahua Lin, and Kai Chen. Mmbench-video: A long-form multi-shot benchmark for holistic video understanding. arXiv preprint arXiv:2406.14515, 2024.
  36. 36.Li Fei-Fei, Rob Fergus, and Pietro Perona. Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories. In Conference on Computer Vision and Pattern Recognition Workshop, pages 178–178, 2004.
  37. 37.Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Zhenyu Qiu, Wei Lin, Jinrui Yang, Xiawu Zheng, et al. Mme: A comprehensive evaluation benchmark for multimodal large language models. arXiv preprint arXiv:2306.13394, 2023.
  38. 38.Chaoyou Fu, Yuhan Dai, Yondong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, et al. Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis. arXiv preprint arXiv:2405.21075, 2024.
  39. 39.Xingyu Fu, Yushi Hu, Bangzheng Li, Yu Feng, Haoyu Wang, Xudong Lin, Dan Roth, Noah A Smith, Wei-Chiu Ma, and Ranjay Krishna. Blink: Multimodal large language models can see but not perceive. arXiv preprint arXiv:2404.12390, 2024.
  40. 40.Jiahui Gao, Renjie Pi, Jipeng Zhang, Jiacheng Ye, Wanjun Zhong, Yufei Wang, Lanqing Hong, Jianhua Han, Hang Xu, Zhenguo Li, et al. G-llava: Solving geometric problem with multi-modal large language model. arXiv preprint arXiv:2312.11370, 2023.
  41. 41.Zhangwei Gao, Zhe Chen, Erfei Cui, Yiming Ren, Weiyun Wang, Jinguo Zhu, Hao Tian, Shenglong Ye, Junjun He, Xizhou Zhu, et al. Mini-internvl: A flexible-transfer pocket multimodal model with 5% parameters and 90% performance. arXiv preprint arXiv:2410.16261, 2024.
  42. 42.Junqi Ge, Ziyi Chen, Jintao Lin, Jinguo Zhu, Xihui Liu, Jifeng Dai, and Xizhou Zhu. V2pe: Improving multi-modal long-context capability of vision-language models with variable visual position encoding. arXiv preprint arXiv:2412.09616, 2024.
  43. 43.Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6904–6913, 2017.
  44. 44.Shuhao Gu, Jialing Zhang, Siyuan Zhou, Kevin Yu, Zhaohu Xing, Liangdong Wang, Zhou Cao, Jintao Jia, Zhuoyi Zhang, Yixuan Wang, et al. Infinity-mm: Scaling multimodal performance with large-scale and high-quality instruction data. arXiv preprint arXiv:2410.18558, 2024.
  45. 45.Tianrui Guan, Fuxiao Liu, Xiyang Wu, Ruiqi Xian, Zongxia Li, Xiaoyu Liu, Xijun Wang, Lichang Chen, Furong Huang, Yaser Yacoob, et al. Hallusionbench: An advanced diagnostic suite for entangled language hallucination & visual illusion in large vision-language models. arXiv preprint arXiv:2310.14566, 2023.
  46. 46.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. In The International Conference on Learning Representations, 2020.
  47. 47.Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring mathematical problem solving with the MATH dataset. In Joaquin Vanschoren and Sai-Kit Yeung, editors, Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1, NeurIPS Datasets and Benchmarks 2021, December 2021, virtual, 2021.
  48. 48.Yuzhen Huang, Yuzhuo Bai, Zhihao Zhu, Junlei Zhang, Jinghan Zhang, Tangjun Su, Junteng Liu, Chuancheng Lv, Yikai Zhang, Yao Fu, et al. C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models. Advances in Neural Information Processing Systems, 36, 2024.
  49. 49.Zheng Huang, Kai Chen, Jianhua He, Xiang Bai, Dimosthenis Karatzas, Shijian Lu, and CV Jawahar. Icdar2019 competition on scanned receipt ocr and information extraction. In 2019 International Conference on Document Analysis and Recognition (ICDAR), pages 1516–1520. IEEE, 2019.
  50. 50.Drew A Hudson and Christopher D Manning. Gqa: A new dataset for real-world visual reasoning and compositional question answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6700–6709, 2019.
  51. 51.Dongfu Jiang, Xuan He, Huaye Zeng, Cong Wei, Max Ku, Qian Liu, and Wenhu Chen. Mantis: Interleaved multi-image instruction tuning. arXiv preprint arXiv:2405.01483, 2024.
  52. 52.Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. arXiv preprint arXiv:1705.03551, 2017.
  53. 53.Seungjae Jung, Gunsoo Han, Daniel Wontae Nam, and Kyoung-Woon On. Binary classifier optimization for large language model alignment. arXiv preprint arXiv:2404.04656, 2024.
  54. 54.Kushal Kafle, Brian Price, Scott Cohen, and Christopher Kanan. Dvqa: Understanding data visualizations via question answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5648–5656, 2018.
  55. 55.Mehran Kazemi, Hamidreza Alvari, Ankit Anand, Jialin Wu, Xi Chen, and Radu Soricut. Geomverse: A systematic evaluation of large models for geometric reasoning. arXiv preprint arXiv:2312.12241, 2023.
  56. 56.Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara Berg. Referitgame: Referring to objects in photographs of natural scenes. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing, pages 787–798, 2014.
  57. 57.Aniruddha Kembhavi, Mike Salvato, Eric Kolve, Minjoon Seo, Hannaneh Hajishirzi, and Ali Farhadi. A diagram is worth a dozen images. In European Conference on Computer Vision, pages 235–251, 2016.
  58. 58.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. Natural questions: a benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:453–466, 2019.
  59. 59.Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. Race: Large-scale reading comprehension dataset from examinations. arXiv preprint arXiv:1704.04683, 2017.
  60. 60.Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Yanwei Li, Ziwei Liu, and Chunyuan Li. Llava-onevision: Easy visual task transfer. arXiv preprint arXiv:2408.03326, 2024.
  61. 61.Bohao Li, Yuying Ge, Yi Chen, Yixiao Ge, Ruimao Zhang, and Ying Shan. Seed-bench-2-plus: Benchmarking multimodal large language models with text-rich visual comprehension. arXiv preprint arXiv:2404.16790, 2024.
  62. 62.Chunyi Li, Jianbo Zhang, Zicheng Zhang, Haoning Wu, Yuan Tian, Wei Sun, Guo Lu, Xiaohong Liu, Xiongkuo Min, Weisi Lin, et al. R-bench: Are your large multimodal model robust to real-world corruptions? arXiv preprint arXiv:2410.05474, 2024.
  63. 63.Haonan Li, Yixuan Zhang, Fajri Koto, Yifei Yang, Hai Zhao, Yeyun Gong, Nan Duan, and Timothy Baldwin. Cmmlu: Measuring massive multitask language understanding in chinese. arXiv preprint arXiv:2306.09212, 2023.
  64. 64.KunChang Li, Yinan He, Yi Wang, Yizhuo Li, Wenhai Wang, Ping Luo, Yali Wang, Limin Wang, and Yu Qiao. Videochat: Chat-centric video understanding. arXiv preprint arXiv:2305.06355, 2023.
  65. 65.Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Yi Liu, Zun Wang, Jilan Xu, Guo Chen, Ping Luo, et al. Mvbench: A comprehensive multi-modal video understanding benchmark. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22195–22206, 2024.
  66. 66.Yanghao Li, Chao-Yuan Wu, Haoqi Fan, Karttikeya Mangalam, Bo Xiong, Jitendra Malik, and Christoph Feichtenhofer. Mvitv2: Improved multiscale vision transformers for classification and detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4804–4814, 2022.
  67. 67.Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji-Rong Wen. Evaluating object hallucination in large vision-language models. In The Conference on Empirical Methods in Natural Language Processing, pages 292–305, 2023.
  68. 68.Zhang Li, Biao Yang, Qiang Liu, Zhiyin Ma, Shuo Zhang, Jingxu Yang, Yabo Sun, Yuliang Liu, and Xiang Bai. Monkey: Image resolution and text label are important things for large multi-modal models. arXiv preprint arXiv:2311.06607, 2023.
  69. 69.Zhiqi Li, Guo Chen, Shilong Liu, Shihao Wang, Vibashan VS, Yishen Ji, Shiyi Lan, Hao Zhang, Yilin Zhao, Subhashree Radhakrishnan, et al. Eagle 2: Building post-training data strategies from scratch for frontier vision-language models. arXiv preprint arXiv:2501.14818, 2025.
  70. 70.Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. Let’s verify step by step. In The Twelfth International Conference on Learning Representations, 2023.
  71. 71.Ji Lin, Hongxu Yin, Wei Ping, Pavlo Molchanov, Mohammad Shoeybi, and Song Han. Vila: On pre-training for visual language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 26689–26699, 2024.
  72. 72.Adam Dahlgren Lindström and Savitha Sam Abraham. Clevr-math: A dataset for compositional language, visual and mathematical reasoning. arXiv preprint arXiv:2208.05358, 2022.
  73. 73.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. Advances in Neural Information Processing Systems, 36, 2023.
  74. 74.Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jianwei Yang, Hang Su, et al. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. In European Conference on Computer Vision, pages 38–55. Springer, 2025.
  75. 75.Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, et al. Mmbench: Is your multi-modal model an all-around player? arXiv preprint arXiv:2307.06281, 2023.
  76. 76.Yuliang Liu, Zhang Li, Hongliang Li, Wenwen Yu, Mingxin Huang, Dezhi Peng, Mingyu Liu, Mingrui Chen, Chunyuan Li, Lianwen Jin, et al. On the hidden mystery of ocr in large multimodal models. arXiv preprint arXiv:2305.07895, 2023.
  77. 77.Zihan Liu, Yang Chen, Mohammad Shoeybi, Bryan Catanzaro, and Wei Ping. Acemath: Advancing frontier math reasoning with post-training and reward modeling. arXiv preprint, 2024.
  78. 78.Zuyan Liu, Yuhao Dong, Ziwei Liu, Winston Hu, Jiwen Lu, and Yongming Rao. Oryx mllm: On-demand spatial-temporal understanding at arbitrary resolution. arXiv preprint arXiv:2409.12961, 2024.
  79. 79.Dakuan Lu, Xiaoyu Tan, Rui Xu, Tianchu Yao, Chao Qu, Wei Chu, Yinghui Xu, and Yuan Qi. Scp-116k: A high-quality problem-solution dataset and a generalized pipeline for automated extraction in the higher education science domain, 2025.
  80. 80.Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao. Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts. arXiv preprint arXiv:2310.02255, 2023.
  81. 81.Pan Lu, Ran Gong, Shibiao Jiang, Liang Qiu, Siyuan Huang, Xiaodan Liang, and Song-Chun Zhu. Inter-gps: Interpretable geometry problem solving with formal language and symbolic reasoning. arXiv preprint arXiv:2105.04165, 2021.
  82. 82.Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. Learn to explain: Multimodal reasoning via thought chains for science question answering. Advances in Neural Information Processing Systems, 35:2507–2521, 2022.
  83. 83.Pan Lu, Liang Qiu, Jiaqi Chen, Tony Xia, Yizhou Zhao, Wei Zhang, Zhou Yu, Xiaodan Liang, and Song-Chun Zhu. Iconqa: A new benchmark for abstract diagram understanding and visual language reasoning. arXiv preprint arXiv:2110.13214, 2021.
  84. 84.Shiyin Lu, Yang Li, Qing-Guo Chen, Zhao Xu, Weihua Luo, Kaifu Zhang, and Han-Jia Ye. Ovis: Structural embedding alignment for multimodal large language model. arXiv preprint arXiv:2405.20797, 2024.
  85. 85.Xudong Lu, Yinghao Chen, Cheng Chen, Hui Tan, Boheng Chen, Yina Xie, Rui Hu, Guanxin Tan, Renshou Wu, Yan Hu, et al. Bluelm-v-3b: Algorithm and system co-design for multimodal large language models on mobile devices. arXiv preprint arXiv:2411.10640, 2024.
  86. 86.Yujie Lu, Dongfu Jiang, Wenhu Chen, William Yang Wang, Yejin Choi, and Bill Yuchen Lin. Wildvision: Evaluating vision-language models in the wild with human preferences. arXiv preprint arXiv:2406.11069, 2024.
  87. 87.Liangchen Luo, Yinxiao Liu, Rosanne Liu, Samrat Phatale, Harsh Lara, Yunxuan Li, Lei Shu, Yun Zhu, Lei Meng, Jiao Sun, et al. Improve mathematical reasoning in language models by automated process supervision. arXiv preprint arXiv:2406.06592, 2, 2024.
  88. 88.Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy. Generation and comprehension of unambiguous object descriptions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11–20, 2016.
  89. 89.Andrés Marafioti, Orr Zohar, Miquel Farré, Merve Noyan, Elie Bakouch, Pedro Cuenca, Cyril Zakka, Loubna Ben Allal, Anton Lozhkov, Nouamane Tazi, et al. Smolvlm: Redefining small and efficient multimodal models. arXiv preprint arXiv:2504.05299, 2025.
  90. 90.Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi. Ok-vqa: A visual question answering benchmark requiring external knowledge. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3195–3204, 2019.
  91. 91.Ahmed Masry, Xuan Long Do, Jia Qing Tan, Shafiq Joty, and Enamul Hoque. Chartqa: A benchmark for question answering about charts with visual and logical reasoning. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, pages 2263–2279, 2022.
  92. 92.Minesh Mathew, Viraj Bagal, Rubèn Tito, Dimosthenis Karatzas, Ernest Valveny, and CV Jawahar. Infographicvqa. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 1697–1706, 2022.
  93. 93.Minesh Mathew, Dimosthenis Karatzas, and CV Jawahar. Docvqa: A dataset for vqa on document images. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 2200–2209, 2021.
  94. 94.Nat McAleese, Rai Michael Pokorny, Juan Felipe Ceron Uribe, Evgenia Nitishinskaya, Maja Trebacz, and Jan Leike. Llm critics help catch llm bugs. arXiv preprint arXiv:2407.00215, 2024.
  95. 95.Fanqing Meng, Jin Wang, Chuanhao Li, Quanfeng Lu, Hao Tian, Jiaqi Liao, Xizhou Zhu, Jifeng Dai, Yu Qiao, Ping Luo, et al. Mmiu: Multimodal multi-image understanding for evaluating large vision-language models. arXiv preprint arXiv:2408.02718, 2024.
  96. 96.Anand Mishra, Shashank Shekhar, Ajeet Kumar Singh, and Anirban Chakraborty. Ocr-vqa: Visual question answering by reading text in images. In International Conference on Document Analysis and Recognition, pages 947–952, 2019.
  97. 97.OpenAI. Gpt-4v(ision) system card. https://cdn.openai.com/papers/GPTV_System_Card.pdf, 2023.
  98. 98.OpenAI. Gpt-4o system card. https://openai.com/index/gpt-4o-system-card/, 2025.
  99. 99.Runqi Qiao, Qiuna Tan, Guanting Dong, Minhui Wu, Chong Sun, Xiaoshuai Song, Zhuoma GongQue, Shanglin Lei, Zhe Wei, Miaoxuan Zhang, et al. We-math: Does your large multimodal model achieve human-like mathematical reasoning? arXiv preprint arXiv:2407.01284, 2024.
  100. 100.Yujia Qin, Yining Ye, Junjie Fang, Haoming Wang, Shihao Liang, Shizuo Tian, Junda Zhang, Jiahao Li, Yunxin Li, Shijue Huang, et al. Ui-tars: Pioneering automated gui interaction with native agents. arXiv preprint arXiv:2501.12326, 2025.
  101. 101.Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36, 2024.
  102. 102.Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy Lillicrap, Jean-baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530, 2024.
  103. 103.Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. Winogrande: An adversarial winograd schema challenge at scale. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 8732–8740, 2020.
  104. 104.Minjoon Seo, Hannaneh Hajishirzi, Ali Farhadi, Oren Etzioni, and Clint Malcolm. Solving geometry problems: Combining text and diagram interpretation. In Proceedings of the 2015 conference on empirical methods in natural language processing, pages 1466–1476, 2015.
  105. 105.Min Shi, Fuxiao Liu, Shihao Wang, Shijia Liao, Subhashree Radhakrishnan, De-An Huang, Hongxu Yin, Karan Sapra, Yaser Yacoob, Humphrey Shi, et al. Eagle: Exploring the design space for multimodal llms with mixture of encoders. arXiv preprint arXiv:2408.15998, 2024.
  106. 106.Wenhao Shi, Zhiqiang Hu, Yi Bin, Junhua Liu, Yang Yang, See-Kiong Ng, Lidong Bing, and Roy Ka-Wei Lee. Math-llava: Bootstrapping mathematical reasoning for multimodal large language models. arXiv preprint arXiv:2406.17294, 2024.
  107. 107.Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach. Towards vqa models that can read. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8317–8326, 2019.
  108. 108.Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar. Scaling llm test-time compute optimally can be more effective than scaling model parameters. arXiv preprint arXiv:2408.03314, 2024.
  109. 109.Hai-Long Sun, Da-Wei Zhou, Yang Li, Shiyin Lu, Chao Yi, Qing-Guo Chen, Zhao Xu, Weihua Luo, Kaifu Zhang, De-Chuan Zhan, et al. Parrot: Multilingual visual instruction tuning. arXiv preprint arXiv:2406.02539, 2024.
  110. 110.Kai Sun, Dian Yu, Dong Yu, and Claire Cardie. Investigating prior knowledge for challenging chinese machine reading comprehension. Transactions of the Association for Computational Linguistics, 8:141–155, 2020.
  111. 111.Zhiqing Sun, Sheng Shen, Shengcao Cao, Haotian Liu, Chunyuan Li, Yikang Shen, Chuang Gan, Liang-Yan Gui, Yu-Xiong Wang, Yiming Yang, et al. Aligning large multimodal models with factually augmented rlhf. arXiv preprint arXiv:2309.14525, 2023.
  112. 112.Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V Le, Ed H Chi, Denny Zhou, et al. Challenging big-bench tasks and whether chain-of-thought can solve them. arXiv preprint arXiv:2210.09261, 2022.
  113. 113.Jingqun Tang, Qi Liu, Yongjie Ye, Jinghui Lu, Shu Wei, Chunhui Lin, Wanqing Li, Mohamad Fitri Faiz Bin Mahmood, Hao Feng, Zhen Zhao, et al. Mtvqa: Benchmarking multilingual text-centric visual question answering. arXiv preprint arXiv:2405.11985, 2024.
  114. 114.Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023.
  115. 115.Qwen Team. Qvq: To see the world with wisdom, December 2024.
  116. 116.Shengbang Tong, Ellis Brown, Penghao Wu, Sanghyun Woo, Manoj Middepogu, Sai Charitha Akula, Jihan Yang, Shusheng Yang, Adithya Iyer, Xichen Pan, et al. Cambrian-1: A fully open, vision-centric exploration of multimodal llms. arXiv preprint arXiv:2406.16860, 2024.
  117. 117.v DeepMind. Gemini 2.5 pro. https://deepmind.google/technologies/gemini/pro/, 2025.
  118. 118.Fei Wang, Xingyu Fu, James Y Huang, Zekun Li, Qin Liu, Xiaogeng Liu, Mingyu Derek Ma, Nan Xu, Wenxuan Zhou, Kai Zhang, et al. Muirbench: A comprehensive benchmark for robust multi-image understanding. arXiv preprint arXiv:2406.09411, 2024.
  119. 119.Ke Wang, Junting Pan, Weikang Shi, Zimu Lu, Mingjie Zhan, and Hongsheng Li. Measuring multimodal mathematical reasoning with math-vision dataset. arXiv preprint arXiv:2402.14804, 2024.
  120. 120.Peiyi Wang, Lei Li, Zhihong Shao, RX Xu, Damai Dai, Yifei Li, Deli Chen, Yu Wu, and Zhifang Sui. Math-shepherd: Verify and reinforce llms step-by-step without human annotations. arXiv preprint arXiv:2312.08935, 2023.
  121. 121.Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, et al. Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution. arXiv preprint arXiv:2409.12191, 2024.
  122. 122.Peng Wang, Shijie Wang, Junyang Lin, Shuai Bai, Xiaohuan Zhou, Jingren Zhou, Xinggang Wang, and Chang Zhou. One-peace: Exploring one general representation model toward unlimited modalities. arXiv:2305.11172, 2023.
  123. 123.Weihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong, Ji Qi, Yan Wang, Junhui Ji, Zhuoyi Yang, Lei Zhao, Xixuan Song, et al. Cogvlm: Visual expert for pretrained language models. arXiv preprint arXiv:2311.03079, 2023.
  124. 124.Weiyun Wang, Zhe Chen, Wenhai Wang, Yue Cao, Yangzhou Liu, Zhangwei Gao, Jinguo Zhu, Xizhou Zhu, Lewei Lu, Yu Qiao, and Jifeng Dai. Enhancing the reasoning ability of multimodal large language models via mixed preference optimization. arXiv preprint arXiv:2411.10442, 2024.
  125. 125.Weiyun Wang, Zhangwei Gao, Lianjie Chen, Zhe Chen, Jinguo Zhu, Xiangyu Zhao, Yangzhou Liu, Yue Cao, Shenglong Ye, Xizhou Zhu, et al. Visualprm: An effective process reward model for multimodal reasoning. arXiv preprint arXiv:2503.10291, 2025.
  126. 126.Weiyun Wang, Yiming Ren, Haowen Luo, Tiantong Li, Chenxiang Yan, Zhe Chen, Wenhai Wang, Qingyun Li, Lewei Lu, Xizhou Zhu, et al. The all-seeing project v2: Towards general relation comprehension of the open world. arXiv preprint arXiv:2402.19474, 2024.
  127. 127.Weiyun Wang, Min Shi, Qingyun Li, Wenhai Wang, Zhenhang Huang, Linjie Xing, Zhe Chen, Hao Li, Xizhou Zhu, Zhiguo Cao, et al. The all-seeing project: Towards panoptic visual recognition and understanding of the open world. In The International Conference on Learning Representations, 2024.
  128. 128.Zirui Wang, Mengzhou Xia, Luxi He, Howard Chen, Yitao Liu, Richard Zhu, Kaiqu Liang, Xindi Wu, Haotian Liu, Sadhika Malladi, et al. Charxiv: Charting gaps in realistic chart understanding in multimodal llms. arXiv preprint arXiv:2406.18521, 2024.
  129. 129.Haoning Wu, Dongxu Li, Bei Chen, and Junnan Li. Longvideobench: A benchmark for long-context interleaved video-language understanding. arXiv preprint arXiv:2407.15754, 2024.
  130. 130.Zhiyong Wu, Zhenyu Wu, Fangzhi Xu, Yian Wang, Qiushi Sun, Chengyou Jia, Kanzhi Cheng, Zichen Ding, Liheng Chen, Paul Pu Liang, et al. Os-atlas: A foundation action model for generalist gui agents. arXiv preprint arXiv:2410.23218, 2024.
  131. 131.Yijia Xiao, Edward Sun, Tianyu Liu, and Wei Wang. Logicvista: Multimodal llm logical reasoning benchmark in visual contexts. arXiv preprint arXiv:2407.04973, 2024.
  132. 132.Yiheng Xu, Zekun Wang, Junli Wang, Dunjie Lu, Tianbao Xie, Amrita Saha, Doyen Sahoo, Tao Yu, and Caiming Xiong. Aguvis: Unified pure vision agents for autonomous gui interaction. 2024.
  133. 133.B. Yan, Yi Jiang, Jiannan Wu, D. Wang, Ping Luo, Zehuan Yuan, and Huchuan Lu. Universal instance perception as object discovery and retrieval. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023.
  134. 134.Jihan Yang, Shusheng Yang, Anjali Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie. Thinking in Space: How Multimodal Large Language Models See, Remember and Recall Spaces. arXiv preprint arXiv:2412.14171, 2024.
  135. 135.Yuan Yao, Tianyu Yu, Ao Zhang, Chongyi Wang, Junbo Cui, Hongji Zhu, Tianchi Cai, Haoyu Li, Weilin Zhao, Zhihui He, et al. Minicpm-v: A gpt-4v level mllm on your phone. arXiv preprint arXiv:2408.01800, 2024.
  136. 136.Qinghao Ye, Haiyang Xu, Jiabo Ye, Ming Yan, Haowei Liu, Qi Qian, Ji Zhang, Fei Huang, and Jingren Zhou. mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration. arXiv preprint arXiv:2311.04257, 2023.
  137. 137.Kaining Ying, Fanqing Meng, Jin Wang, Zhiqian Li, Han Lin, Yue Yang, Hao Zhang, Wenbo Zhang, Yuqi Lin, Shuo Liu, Jiayi Lei, Quanfeng Lu, Runjian Chen, Peng Xu, Renrui Zhang, Haozhe Zhang, Peng Gao, Yali Wang, Yu Qiao, Ping Luo, Kaipeng Zhang, and Wenqi Shao. Mmt-bench: A comprehensive multimodal benchmark for evaluating large vision-language models towards multitask agi. arXiv preprint arXiv:2404.16006, 2024.
  138. 138.Weihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Zicheng Liu, Xinchao Wang, and Lijuan Wang. Mm-vet: Evaluating large multimodal models for integrated capabilities. arXiv preprint arXiv:2308.02490, 2023.
  139. 139.Weihao Yu, Zhengyuan Yang, Linfeng Ren, Linjie Li, Jianfeng Wang, Kevin Lin, Chung-Ching Lin, Zicheng Liu, Lijuan Wang, and Xinchao Wang. Mm-vet v2: A challenging benchmark to evaluate large multimodal models for integrated capabilities. arXiv preprint arXiv:2408.00765, 2024.
  140. 140.Ya-Qi Yu, Minghui Liao, Jiwen Zhang, and Jihao Wu. Texthawk2: A large vision-language model excels in bilingual ocr and grounding with 16x fewer tokens. arXiv preprint arXiv:2410.05261, 2024.
  141. 141.Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, et al. Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi. arXiv preprint arXiv:2311.16502, 2023.
  142. 142.Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a machine really finish your sentence? In Proceedings of the Annual Meeting of the Association for Computational Linguistics, pages 4791–4800, 2019.
  143. 143.Haotian Zhang, Mingfei Gao, Zhe Gan, Philipp Dufter, Nina Wenzel, Forrest Huang, Dhruti Shah, Xianzhi Du, Bowen Zhang, Yanghao Li, et al. Mm1.5: Methods, analysis & insights from multimodal llm fine-tuning. arXiv preprint arXiv:2409.20566, 2024.
  144. 144.Haotian Zhang, Haoxuan You, Philipp Dufter, Bowen Zhang, Chen Chen, Hong-You Chen, Tsu-Jui Fu, William Yang Wang, Shih-Fu Chang, Zhe Gan, et al. Ferret-v2: An improved baseline for referring and grounding with large language models. arXiv preprint arXiv:2404.07973, 2024.
  145. 145.Peiyuan Zhang, Kaichen Zhang, Bo Li, Guangtao Zeng, Jingkang Yang, Yuanhan Zhang, Ziyue Wang, Haoran Tan, Chunyuan Li, and Ziwei Liu. Long context transfer from language to vision. arXiv preprint arXiv:2406.16852, 2024.
  146. 146.Renrui Zhang, Dongzhi Jiang, Yichi Zhang, Haokun Lin, Ziyu Guo, Pengshuo Qiu, Aojun Zhou, Pan Lu, Kai-Wei Chang, Peng Gao, et al. Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? arXiv preprint arXiv:2403.14624, 2024.
  147. 147.Renrui Zhang, Xinyu Wei, Dongzhi Jiang, Yichi Zhang, Ziyu Guo, Chengzhuo Tong, Jiaming Liu, Aojun Zhou, Bin Wei, Shanghang Zhang, et al. Mavis: Mathematical visual instruction tuning. arXiv preprint arXiv:2407.08739, 2024.
  148. 148.Tianyu Zhang, Suyuchen Wang, Lu Li, Ge Zhang, Perouz Taslakian, Sai Rajeswar, Jie Fu, Bang Liu, and Yoshua Bengio. Vcr: Visual caption restoration. arXiv preprint arXiv:2406.06462, 2024.
  149. 149.Xiaotian Zhang, Chunyang Li, Yi Zong, Zhengyu Ying, Liang He, and Xipeng Qiu. Evaluating the performance of large language models on gaokao benchmark. arXiv preprint arXiv:2305.12474, 2023.
  150. 150.Y Zhang, B Li, H Liu, Y Lee, L Gui, D Fu, J Feng, Z Liu, and C Li. Llava-next: A strong zero-shot video understanding model. 2024.
  151. 151.Yi-Fan Zhang, Huanyu Zhang, Haochen Tian, Chaoyou Fu, Shuangqing Zhang, Junfei Wu, Feng Li, Kun Wang, Qingsong Wen, Zhang Zhang, et al. Mme-realworld: Could your multimodal llm challenge high-resolution real-world scenarios that are difficult for humans? arXiv preprint arXiv:2408.13257, 2024.
  152. 152.Zhenru Zhang, Chujie Zheng, Yangzhen Wu, Beichen Zhang, Runji Lin, Bowen Yu, Dayiheng Liu, Jingren Zhou, and Junyang Lin. The lessons of developing process reward models in mathematical reasoning. arXiv preprint arXiv:2501.07301, 2025.
  153. 153.Bingchen Zhao, Yongshuo Zong, Letian Zhang, and Timothy Hospedales. Benchmarking multi-image understanding in vision and language models: Perception, knowledge, reasoning, and multi-hop reasoning. arXiv preprint arXiv:2406.12742, 2024.
  154. 154.Junjie Zhou, Yan Shu, Bo Zhao, Boya Wu, Shitao Xiao, Xi Yang, Yongping Xiong, Bo Zhang, Tiejun Huang, and Zheng Liu. Mlvu: A comprehensive benchmark for multi-task long video understanding. arXiv preprint arXiv:2406.04264, 2024.
  155. 155.Chengke Zou, Xingang Guo, Rui Yang, Junyu Zhang, Bin Hu, and Huan Zhang. Dynamath: A dynamic visual benchmark for evaluating mathematical reasoning robustness of vision language models. arXiv preprint arXiv:2411.00836, 2024.

Citation

MLA
Zhu, J., et al. “InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models”. arXiv, 2025, http://arxiv.org/abs/2504.10479v3.
APA
Zhu, J., Wang, W., Chen, Z., Liu, Z., Ye, S., Gu, L., Tian, H., Duan, Y., Su, W., Shao, J., Gao, Z., Cui, E., Wang, X., Cao, Y., Liu, Y., Wei, X., Zhang, H., Wang, H., Xu, W., … Wang, W. (2025). InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models. arXiv. http://arxiv.org/abs/2504.10479v3
Chicago
Zhu, J., W. Wang, Z. Chen, et al. 2025. “InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models”. arXiv. http://arxiv.org/abs/2504.10479v3.
Harvard
Zhu, J. et al. (2025) “InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2504.10479v3.
Vancouver
1. Zhu J, Wang W, Chen Z, et al (2025) InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models. arXiv

BibTeX

@article{zhu2025internvl3,
  title = {InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models},
  author = {Zhu, Jinguo and Wang, Weiyun and Chen, Zhe and Liu, Zhaoyang and Ye, Shenglong and Gu, Lixin and Tian, Hao and Duan, Yuchen and Su, Weijie and Shao, Jie and Gao, Zhangwei and Cui, Erfei and Wang, Xuehui and Cao, Yue and Liu, Yangzhou and Wei, Xingguang and Zhang, Hongjie and Wang, Haomin and Xu, Weiye and Li, Hao and Wang, Jiahao and Deng, Nianchen and Li, Songze and He, Yinan and Jiang, Tan and Luo, Jiapeng and Wang, Yi and He, Conghui and Shi, Botian and Zhang, Xingcheng and Shao, Wenqi and He, Junjun and Xiong, Yingtong and Qu, Wenwen and Sun, Peng and Jiao, Penglong and Lv, Han and Wu, Lijun and Zhang, Kaipeng and Deng, Huipeng and Ge, Jiaye and Chen, Kai and Wang, Limin and Dou, Min and Lu, Lewei and Zhu, Xizhou and Lu, Tong and Lin, Dahua and Qiao, Yu and Dai, Jifeng and Wang, Wenhai},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2504.10479v3},
  eprint = {2504.10479}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/