InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Jinguo ZhuWeiyun WangZhe ChenZhaoyang LiuShenglong YeLixin GuYuchen DuanHao TianWeijie SuJie Shao
Presents InternVL3, a native multimodal model trained jointly on visual and textual corpora that matches leading proprietary systems on the MMMU benchmark by combining variable visual position encoding with advanced test-time scaling.
Modern multimodal artificial intelligence systems, which process both text and images, typically adapt pre-existing, text-only large language models through complex, multi-stage training pipelines. This post-hoc adaptation often introduces significant cross-modality alignment challenges, requires delicate parameter-freezing schedules, and risks degrading core language competencies. As demand grows for unified models capable of handling multidisciplinary reasoning, complex documents, and real-world visual environments, developing more integrated, resource-efficient training paradigms has become essential.
The article introduces and evaluates InternVL3, an open-source family of multimodal large language models ranging from 1 billion to 78 billion parameters. The main objective is to demonstrate that a native pre-training paradigm—which jointly optimizes linguistic and visual representations in a single initial stage—delivers state-of-the-art multimodal performance while preserving and enhancing pure-text capabilities.
To achieve this, the authors implemented a unified pre-training architecture that updates all model parameters simultaneously using an optimal 1:3 ratio of pure text (50 billion tokens) to multimodal data (150 billion tokens). The model incorporates Variable Visual Position Encoding to support longer contexts, alongside advanced post-training involving supervised fine-tuning across 21.7 million samples and Mixed Preference Optimization. The evaluation relied on comprehensive benchmarking across standard multimodal test suites, language understanding suites, and real-world vision tasks, supported by an optimized distributed training framework that delivered training speedups of 50% to 200%.
The key findings highlight substantial performance advantages across several domains. First, the largest variant, InternVL3-78B, established a new open-source state of the art with a 72.2 score on the multidisciplinary MMMU benchmark, matching or exceeding leading proprietary systems like ChatGPT-4o and Claude 3.5 Sonnet. Second, the series demonstrated strong mathematical and logical reasoning, scoring 79.0 on MathVista, which rose to 80.5 when applying test-time scaling with a process reward critic model. Third, the models achieved superior results on text-rich and practical tasks, including a score of 906 on OCRBench and an 88.7% accuracy on graphical user interface grounding. Finally, purely textual evaluation confirmed that joint multimodal pre-training improved text and coding capabilities compared to baseline chat models derived from the same base text architecture.
These findings imply that joint multimodal pre-training eliminates the need for cumbersome, multi-stage alignment pipelines, lowering integration overhead without sacrificing language quality. For organizations deploying computer vision and language solutions, InternVL3 demonstrates that open-source architectures can rival costly closed-source alternatives across document processing, user interface automation, and spatial analysis.
Based on these results, decision-makers should consider adopting or piloting the publicly released open-source weights and training data for enterprise workflows requiring multimodal understanding. Further work should explore applying the model to dynamic, interactive user interface agents and specialized spatial-temporal environments.
Regarding limitations, the article notes that larger model variants exhibited plateaus in visual grounding performance due to a relative reduction in grounding-specific pre-training data. Additionally, closed-source models such as Gemini 2.5 Pro still retain narrow performance edges in specific areas like video-context comprehension and visual hallucination reduction. Stakeholders should maintain moderate caution when deploying the system in critical tasks where fine-grained object localization is paramount.
- Paper: Intern VL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks, Zhe Chen et al. (2024). Introduces the foundational InternVL architecture and InternViT encoder, establishing the core lineage and design motivations that InternVL3 directly advances.
- Paper: MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI, Xiang Yue et al. (2023). Presents the expert-level multimodal understanding benchmark MMMU, which serves as the primary evaluation metric and state-of-the-art claim in InternVL3.
- Paper: Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution, Peng Wang et al. (2024). Develops dynamic resolution and multidimensional rotary position embeddings for vision-language models, foundational concepts that InternVL3 builds upon with its variable visual position encoding (V2PE).
- Paper: LLaVA-OneVision: Easy Visual Task Transfer, Bo Li et al. (2024). Demonstrates open-source unified cross-scenario transfer across images and video, representing the immediate prior paradigm of post-hoc multimodal alignment that InternVL3 aims to surpass with native pre-training.
- Paper: Improved Baselines with Visual Instruction Tuning, Haotian Liu et al. (2024). Details essential visual instruction tuning and data formatting baselines that underpin modern supervised fine-tuning recipes for large multimodal models.
- Paper: MMBench: Is Your Multi-modal Model an All-around Player?, Yuanzhan Liu et al. (2023). Establishes standard comprehensive evaluation benchmarks and circular evaluation protocols widely adopted in assessing advanced open-source vision-language models.
- Paper: Qwen3-VL Technical Report, Shuai Bai et al. (2025). Builds upon native multimodal pre-training recipes and position encoding techniques to scale dense and mixture-of-experts vision-language models up to 235B parameters with extended context lengths.
- Paper: Innovator-VL: A Multimodal Large Language Model for Scientific Discovery, Zichen Wen et al. (2026). Extends general-purpose open-source multimodal recipes to specialized scientific discovery and reasoning domains through tailored post-training and mid-training pipelines.
- Paper: Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces, Jihan Yang et al. (2025). Evaluates how high-capacity multimodal models like InternVL3 perceive, recall, and reason about spatial environments in complex video sequences.
- Paper: Qwen3.5-Omni Technical Report, Qwen Team (2026). Broadens native multimodal architectures into unified omnimodal systems that integrate streaming audio, speech synthesis, and agentic execution.
- Paper: Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation, Zhiheng Liu et al. (2026). Explores an alternative next-generation multimodal paradigm that unifies understanding and generation directly on raw pixel embeddings without discrete vision encoders.
