Qwen3-VL Technical Report

Shuai BaiYuxuan CaiRui-Zhe ChenKe-qin ChenXiong-Hui ChenZesen ChengLiang-Hao DengWei DingRongyao FangChang Gao

article2025arXiv2,274 citations

Presents the Qwen3-VL model family, delivering native 256K-token interleaved multimodal context processing across dense and mixture-of-experts architectures through upgraded spatial-temporal positional embeddings, multi-level visual feature fusion, and explicit video timestamp alignment.

Listen

Modern vision–language models must balance advanced multimodal reasoning across text, images, and video without degrading their core language processing abilities. The article introduces and evaluates the Qwen3-VL model family to demonstrate how targeted architectural upgrades, comprehensive data curation, and multi-stage training can deliver state-of-the-art multimodal performance while preserving or improving pure-text capabilities.

The authors develop dense and mixture-of-experts model variants ranging from 2 billion to 235 billion parameters, supporting native context lengths up to 256,000 tokens. To resolve common multi-modal bottlenecks, the design incorporates interleaved multi-dimensional rotary positional embeddings for balanced spatial-temporal modeling, DeepStack cross-layer visual token injection to preserve fine-grained representations, and explicit textual timestamps to replace sparse positional IDs in long video contexts. The training framework utilizes a four-stage pretraining pipeline scaled across massive compute infrastructure, followed by supervised fine-tuning, strong-to-weak knowledge distillation, and reinforcement learning tailored for both standard direct execution and extended chain-of-thought reasoning.

Empirical evaluations show that the flagship 235B parameter mixture-of-experts model consistently achieves top-tier or state-of-the-art results across diverse benchmarks, including complex multimodal reasoning, optical character recognition across 39 languages, 2D and 3D spatial grounding, and agentic interface navigation. In long-context retrieval evaluations, the model maintained 100% accuracy on 30-minute video sequences and 99.5% accuracy when extrapolated to 2 hours of input. Notably, the integration of vision capabilities does not compromise textual proficiency; the flagship model outperforms comparable text-only language models on complex mathematical and coding benchmarks like AIME-25 and LiveCodeBench.

These findings indicate that multimodal models can serve as single, unified engines for enterprise workflows involving lengthy technical documentation, autonomous graphical user interface interaction, and fine-grained visual search without requiring separate specialized text and vision models. Furthermore, the results reveal that augmenting models with external tools often delivers greater perceptual accuracy improvements than simply increasing model parameter size, highlighting an efficient path to improve operational performance.

Organizations adopting these models should select between standard and reasoning-oriented variants based on specific latency, computational budget, and problem complexity constraints. Future operational integration should focus on interactive agent workflows, real-time control, and unified generation architectures. However, decision-makers should note that long-video benchmark comparisons faced constraints due to API limits across proprietary competitor baselines, requiring careful domain-specific validation before large-scale production deployment.

Cover for Qwen3-VL Technical Report

Abstract

We introduce Qwen3-VL, the most capable vision-language model in the Qwen series to date, achieving superior performance across a broad range of multimodal benchmarks. It natively supports interleaved contexts of up to 256K tokens, seamlessly integrating text, images, and video. The model family includes both dense (2B/4B/8B/32B) and mixture-of-experts (30B-A3B/235B-A22B) variants to accommodate diverse latency-quality trade-offs. Qwen3-VL delivers three core pillars: (i) markedly stronger pure-text understanding, surpassing comparable text-only backbones in several cases; (ii) robust long-context comprehension with a native 256K-token window for both text and interleaved multimodal inputs, enabling faithful retention, retrieval, and cross-referencing across long documents and videos; and (iii) advanced multimodal reasoning across single-image, multi-image, and video tasks, demonstrating leading performance on comprehensive evaluations such as MMMU and visual-math benchmarks (e.g., MathVista and MathVision). Architecturally, we introduce three key upgrades: (i) an enhanced interleaved-MRoPE for stronger spatial-temporal modeling across images and video; (ii) DeepStack integration, which effectively leverages multi-level ViT features to tighten vision-language alignment; and (iii) text-based time alignment for video, evolving from T-RoPE to explicit textual timestamp alignment for more precise temporal grounding. Under comparable token budgets and latency constraints, Qwen3-VL achieves superior performance in both dense and Mixture-of-Experts (MoE) architectures. We envision Qwen3-VL serving as a foundational engine for image-grounded reasoning, agentic decision-making, and multimodal code intelligence in real-world workflows.

Table of Contents

  • 1 Introduction
  • 2 Model Architecture
  • 2.1 Interleaved MRoPE
  • 2.2 DeepStack
  • 2.3 Video Timestamp
  • 3 Pre-Training
  • 3.1 Training Recipe
  • 3.2 Pre-Training Data
  • 3.2.1 Image Caption and Interleaved Text-Image Data
  • 3.2.2 Knowledge
  • 3.2.3 OCR, Document Parsing and Long Document Understanding
  • 3.2.4 Grounding and Counting
  • 3.2.5 Spatial Understanding and 3D Recognition
  • 3.2.6 Code
  • 3.2.7 Video
  • 3.2.8 Science, Technology, Engineering, and Mathematics (STEM)
  • 3.2.9 Agent
  • 4 Post-Training
  • 4.1 Training Recipe
  • 4.2 Cold Start Data
  • 4.2.1 SFT Data
  • 4.2.2 Long-CoT Cold Start Data
  • 4.3 Strong-to-Weak Distillation
  • 4.4 Reinforcement Learning
  • 4.4.1 Reasoning Reinforcement Learning
  • 4.4.2 General Reinforcement Learning
  • 4.5 Thinking with Images
  • 4.6 Infrastructure
  • 5 Evaluation
  • 5.1 General Visual Question Answering
  • 5.2 Multimodal Reasoning
  • 5.3 Alignment and Subjective Tasks
  • 5.4 Text Recognition and Document Understanding
  • 5.5 2D and 3D Grounding
  • 5.6 Fine-grained Perception
  • 5.7 Multi-Image Understanding
  • 5.8 Embodied and Spatial Understanding
  • 5.9 Video Understanding
  • 5.10 Agent
  • 5.11 Text-Centric Tasks
  • 5.12 Ablation Study
  • 5.12.1 Vision Encoder
  • 5.12.2 DeepStack
  • 5.12.3 Needle-in-a-Haystack
  • 6 Conclusion
  • 7 Contributions and Acknowledgments
  • References
  • A Benchmarks
  • B Evaluation Prompts
  • B.1 STEM & Puzzle
  • B.2 GeneralVQA
  • B.3 Alignment
  • B.4 Document-Understanding
  • B.5 2D/3D Grounding
  • B.6 Embodied/Spatial Understanding
  • B.7 Multi-Image
  • B.8 Video Understanding
  • B.9 Perception with Tool
  • B.10 Coding
  • B.11 Agent

Knowls

  1. Knowl 1 — Qwen3-VL Architecture and Model Variants

    model/method

    Qwen3-VL is a family of vision–language foundation models instantiated in four dense configurations (2B, 4B, 8B, and 32B parameters) and two Mixture-of-Experts (MoE) configurations (Qwen3-VL-30B-A3B and Qwen3-VL-235B-A22B, where the flagship 235B model activates 22B parameters per token). The architecture comprises three integrated components:

    1. Vision Encoder: Built upon the SigLIP-2 architecture initialized from pre-trained checkpoints and continuously trained with dynamic visual resolutions. The 2B and 4B models use SigLIP2-Large (300M parameters), while models with 8B or more parameters utilize SigLIP2-SO (400M parameters). Dynamic input resolutions are handled using 2D Rotary Position Embeddings (2D-RoPE) combined with size-interpolated absolute position embeddings.
    2. MLP Vision-Language Merger: A two-layer Multilayer Perceptron (MLP) compresses 2×22 \times 2 neighboring spatial visual feature patches from the vision encoder into a single visual token matching the hidden dimension of the language model backbone.
    3. Large Language Model (LLM): An autoregressive decoder backbone based on the Qwen3 series supporting native context windows up to 262,144 (256K) tokens for interleaved sequences of text, images, and video.
  2. Knowl 2 — Interleaved Multimodal Rotary Position Embedding

    model/method

    Multimodal Rotary Position Embedding (MRoPE) decomposes positional coordinates for multimodal sequences into temporal (tt), horizontal (hh), and vertical (ww) axes. Partitioning the embedding channels into contiguous chunks dedicated to tt, hh, and ww assigns distinct, non-overlapping frequency bands to each axis, creating an imbalanced frequency spectrum that degrades long-context and long-video modeling.

    Interleaved MRoPE addresses this spectral bias by uniformly interleaving the tt, hh, and ww positional components across all frequency bands of the embedding channels. Distributing each spatial and temporal coordinate axis evenly across both low- and high-frequency rotary channels eliminates dimensional spectral imbalance and enhances long-range temporal modeling in video inputs.

  3. Knowl 3 — DeepStack Multi-Level Vision-to-Language Feature Injection

    model/method

    DeepStack in Qwen3-VL routes intermediate visual representations from the Vision Transformer (ViT) directly into the early layers of the large language model (LLM) decoder to preserve fine-grained visual details without increasing visual token context length.

    Visual feature maps are extracted from three distinct intermediate depth levels of the vision encoder (capturing low-, mid-, and high-level visual features). Each extracted feature map is projected by a dedicated lightweight vision-language merger MLP to match the LLM hidden dimension. The resulting projected tokens are added directly via residual connections to the hidden states of the first three layers of the LLM decoder. This cross-layer injection enriches visual feature fidelity while leaving the sequence length seen by self-attention unchanged.

  4. Knowl 4 — Textual Token-Based Video Timestamp Alignment

    model/method

    In place of synchronizing video temporal positioning by mapping absolute timestamps directly to numeric position IDs—which yields excessively large, sparse position indices over long videos and requires uniform frame rate sampling during training—Qwen3-VL encodes video time explicitly via textual timestamp tokens.

    Each temporal video patch group is prefixed with a formatted textual string representing the current timecode. Timestamps are formatted during training in both pure seconds (e.g., <3.0 seconds>) and hours:minutes:seconds notation (e.g., HH:MM:SS). This textual encoding allows the language model to parse and reason over explicit temporal markers directly within the token sequence, improving performance on time-grounded video tasks such as dense video captioning, temporal activity localization, and multi-turn video question answering.

  5. Knowl 5 — Four-Stage Progressive Pre-Training Pipeline

    model/method

    Qwen3-VL is pre-trained across four progressive stages that scale trainable components, token budgets, and sequence context windows:

    Stage Objective Trainable Parameters Token Budget Sequence Length
    S0 Vision–Language Alignment Merger only 67B 8,192
    S1 Multimodal Pre-Training All parameters ∼\sim1T 8,192
    S2 Long-Context Pre-Training All parameters ∼\sim1T 32,768
    S3 Ultra-Long-Context Adaptation All parameters 100B 262,144
    1. Stage 0 (Vision--Language Alignment): Keeps the vision encoder and LLM frozen while training only the MLP merger on 67B tokens of image captions, OCR data, and visual knowledge at sequence length 8,192.
    2. Stage 1 (Multimodal Pre-Training): Unfreezes all parameters across encoder, merger, and LLM for end-to-end joint training on ∼\sim1T tokens (composed of text-only and multimodal interleaved data) at sequence length 8,192. A square-root-normalized per-token loss is applied to balance text and visual optimization objectives.
    3. Stage 2 (Long-Context Pre-Training): Scales the sequence length to 32,768 across ∼\sim1T tokens, increasing the proportions of long-form text, multi-turn agent interactions, and extended video sequences.
    4. Stage 3 (Ultra-Long-Context Adaptation): Pushes context length to 262,144 (256K) tokens across a curated 100B-token dataset focused on multi-page long documents, full books, and long-duration videos.
  6. Knowl 6 — Post-Training Pipeline: SFT, Distillation, and Reinforcement Learning

    model/method

    The post-training methodology for Qwen3-VL consists of three stages:

    1. Supervised Fine-Tuning (SFT): Conducted first at 32K context length and extended to 256K context length on ∼\sim1.2 million samples (one-third text-only, two-thirds multimodal). Data is partitioned into standard formats for non-thinking models and Chain-of-Thought (CoT) trajectories for thinking models. Visual math problems undergo multimodal necessity filtering, discarding any sample that a text-only baseline (Qwen3-30B-nothink) solves correctly without the image.
    2. Strong-to-Weak Distillation: Distills capabilities from stronger teacher models into smaller student models using text-only data via off-policy response distillation followed by on-policy Kullback-Leibler (KL) divergence logit alignment.
    3. Reinforcement Learning (RL): Bifurcated into:
      • Reasoning RL: Uses the Soft Adaptive Policy Optimization (SAPO) policy gradient algorithm on verifiably checkable domains (mathematics, code, visual puzzles, and grounding), sampling 16 rollouts per prompt and filtering out queries with pass rates above 90% or equal to 0%.
      • General RL: Multi-task RL optimizing instruction following and preference alignment using rule-based rewards for structured constraints and model-based judges (Qwen2.5-VL-72B-Instruct and Qwen3) for open-ended queries.
  7. Knowl 7 — Two-Stage Tool-Integrated Reinforcement Learning for Visual Agents

    model/method

    To train Qwen3-VL in active "thinking with images" workflows, the model uses a two-stage agentic reinforcement learning paradigm:

    1. Stage 1 (Bootstrap & Cold-Start): A seed dataset of 10,000 two-turn visual grounding examples is constructed. Supervised fine-tuning is performed on Qwen2.5-VL-32B to emulate agent trajectories (think→act→analyze feedback→answer\text{think} \rightarrow \text{act} \rightarrow \text{analyze feedback} \rightarrow \text{answer}), followed by tool-integrated RL.
    2. Stage 2 (Distillation & Scaling): The stage 1 agent is used to generate 120,000 multi-turn agent interaction trajectories across diverse visual tasks. Qwen3-VL undergoes SFT on these trajectories and is subsequently optimized with tool-integrated RL.

    During reinforcement learning, optimization is guided by three complementary rewards:

    • Answer Accuracy Reward: Evaluated by Qwen3-32B verifying whether the final predicted answer is correct.
    • Multi-Turn Reasoning Reward: Evaluated by Qwen2.5-VL-72B checking whether tool feedback was accurately interpreted during step-by-step reasoning.
    • Tool-Calling Reward: Penalizes divergence between the actual number of tool calls and an offline expert target estimated by Qwen2.5-VL-72B based on problem complexity, preventing the policy from degenerating into single-call shortcuts.
  8. Knowl 8 — Normalized Coordinate System for Visual Grounding

    definition

    In Qwen3-VL, 2D visual grounding and spatial coordinates are mapped to a continuous normalized coordinate system scaled to the fixed integer range [0,1000][0, 1000] along both horizontal and vertical spatial dimensions, irrespective of the original image's native resolution or aspect ratio.

    • A 2D bounding box is represented as [x1,y1,x2,y2]∈[0,1000]4[x_1, y_1, x_2, y_2] \in [0, 1000]^4, where (x1,y1)(x_1, y_1) denotes the top-left corner and (x2,y2)(x_2, y_2) denotes the bottom-right corner.
    • A 2D point reference is represented as [x,y]∈[0,1000]2[x, y] \in [0, 1000]^2.

    This normalization ensures scale and aspect ratio invariance across heterogeneous inputs, simplifying post-processing across downstream object detection, counting, and GUI agent operations.

  9. Knowl 9 — Long-Context Needle-in-a-Haystack Video Retrieval Performance

    empirical result

    The long-context retrieval fidelity of Qwen3-VL-235B-A22B-Instruct was evaluated on a video Needle-in-a-Haystack benchmark, where a target evidence frame was inserted at varying depth percentages (0% to 100%) across video sequences of increasing length sampled at 1 frame per second (FPS) with dynamic resolution.

    • Within the native pre-trained context window of 256K tokens (corresponding to up to 30 minutes of video duration), the model achieved an exact 100% retrieval and question-answering accuracy across all inserted positions and depths.
    • When extrapolated to sequences of up to 1,048,576 tokens (1M tokens, corresponding to approximately 120 minutes / 2 hours of video) using YaRN-based positional interpolation, the model maintained an accuracy of 99.5%.
  10. Knowl 10 — Empirical Multimodal Benchmark Evaluation of Qwen3-VL-235B-A22B

    empirical result

    Qwen3-VL-235B-A22B was evaluated across multimodal reasoning, general visual question answering, alignment, document understanding, grounding, and video benchmarks in both thinking and instruct (non-thinking) modes:

    Benchmark Qwen3-VL-235B-A22B Gemini 2.5 Pro GPT-5 Claude Opus 4.1
    Thinking Instruct Thinking Budget-128 High Minimal Thinking Non-thinking
    MMMU 80.6 78.7 81.7 80.9 84.2 74.4 78.4 77.2
    MMMU-Pro 69.3 68.1 68.8 71.2 78.4 62.7 64.8 60.7
    MathVistamini_{\text{mini}} 85.8 84.9 82.7 77.7 81.3 50.9 75.5 74.5
    MathVision 74.6 66.5 73.3 66.0 70.9 45.8 64.3 57.7
    MathVersemini_{\text{mini}} 85.0 72.5 82.9 65.9 84.1 43.0 70.6 68.1
    DynaMath 82.8 79.4 80.0 78.5 85.4 74.0 75.1 72.0
    HallusionBench 66.7 63.2 63.7 60.9 65.7 53.7 60.4 55.1
    MIA-Bench 92.7 91.3 92.3 91.3 92.4 92.6 91.2 90.0
    DocVQAtest_{\text{test}} 96.5 97.1 92.6 94.0 91.5 89.6 92.5 89.2
    OCRBench 875 920 866 872 810 787 764 750
    MMLongBenchDoc 56.2 57.0 55.6 51.2 51.5 42.4 54.5 48.1
    RefCOCO-avg 92.1 91.9 74.6 – 66.8 – – –
    MuirBench 80.1 73.0 77.2 74.0 77.5 66.5 – –
    MVBench 75.2 76.5 69.9 65.8 75.3 64.6 61.4 59.0

    Key findings include:

    1. On STEM and visual math reasoning benchmarks, Qwen3-VL-235B-A22B-Thinking achieves top scores on MathVistamini_{\text{mini}} (85.8), MathVision (74.6), and MathVersemini_{\text{mini}} (85.0), outperforming both Gemini 2.5 Pro Thinking and Claude Opus 4.1.
    2. In document literacy and long-document understanding, Qwen3-VL-235B-A22B-Instruct achieves 97.1 on DocVQAtest_{\text{test}}, 920 on OCRBench, and 57.0 on MMLongBenchDoc.
    3. On multi-image understanding, Qwen3-VL-235B-A22B-Thinking reaches 80.1 on MuirBench.
  11. Knowl 11 — Ablation of DeepStack and Qwen3-ViT Modules

    empirical result

    Ablation experiments evaluated the isolated impact of DeepStack feature injection and Qwen3-ViT vision encoder pre-training:

    1. DeepStack Multi-Level Injection: Evaluated on an internal 15B-A2B language backbone pre-trained on 200B tokens across 11 validation sets without post-training:
      • Average zero-shot score increased from 74.7 (baseline without DeepStack) to 76.0 (with DeepStack).
      • Major improvements were observed in visual text and chart comprehension: InfoVQA improved from 71.9 to 74.2 (+2.3), ChartQA from 81.5 to 83.3 (+1.8), DocVQA from 89.5 to 91.1 (+1.6), and OCRBench from 81.0 to 83.6 (+2.6).
    2. Qwen3-ViT Vision Encoder: Evaluated against the standard SigLIP-2 checkpoint:
      • In zero-shot CLIP pre-training evaluations, Qwen3-ViT increased OmniBench performance from 36.9 to 45.5 (+8.6) while maintaining comparable ImageNet-1K zero-shot accuracy (84.6 vs. 84.2).
      • When integrated with a 1.7B Qwen3 LLM and trained on 1.5T tokens, Qwen3-ViT outperformed the SigLIP-2 backbone across downstream VLM benchmarks: OCRBench (78.7 vs. 77.2), AI2D (76.2 vs. 74.1), RealWorldQA (66.1 vs. 58.7), InfoVQA (67.0 vs. 65.3), and OmniBench (53.0 vs. 50.1).
  12. Knowl 12 — Preservation of Pure-Text Reasoning Capabilities in Vision-Language Training

    empirical result

    Vision–language training often causes degradation in pure-text linguistic and symbolic reasoning proficiency. Qwen3-VL preserves and improves text capabilities relative to corresponding text-only base LLMs through balanced square-root loss reweighting, text-only distillation, and curated reasoning mixtures:

    Benchmark Qwen3-VL-235B-A22B Qwen3-235B-A22B Qwen3-VL-235B-A22B Qwen3-235B-A22B
    Instruct Instruct-2507 Thinking Thinking-2507
    MMLU-Pro 81.8 83.0 83.8 84.4
    MMLU-Redux 92.2 93.1 93.7 93.8
    GPQA 74.3 77.5 77.1 81.1
    AIME-25 74.7 70.3 89.7 92.3
    HMMT-25 57.4 55.4 77.4 83.9
    LiveBench 74.8 75.4 79.6 78.4
    LiveCodeBench v6 54.3 51.8 70.1 74.1
    IFEval 87.8 88.7 88.2 87.8
    WritingBench 85.5 85.2 86.7 88.3

    In the Instruct mode, Qwen3-VL-235B-A22B outperforms its text-only LLM counterpart on mathematical and coding benchmarks (AIME-25: 74.7 vs. 70.3; HMMT-25: 57.4 vs. 55.4; LiveCodeBench v6: 54.3 vs. 51.8), demonstrating that multimodal alignment does not impair base symbolic reasoning.

Coverage note — Appendix B benchmark prompt templates and granular performance rows for smaller model variants (2B, 4B, 8B, 30B, 32B) across every individual sub-benchmark were omitted as standalone knowls, with representative performance trends and core architectures retained.

References

  1. 1.Pravesh Agrawal, Szymon Antoniak, Emma Bou Hanna, Baptiste Bout, Devendra Chaplot, Jessica Chudnovsky, Diogo Costa, Baudouin De Monicault, Saurabh Garg, Theophile Gervet, et al. Pixtral 12b. arXiv preprint arXiv:2410.07073, 2024.
  2. 2.AIME. Aime problems and solutions, 2025. URL https://artofproblemsolving.com/wiki/index.php/AIMEProblemsandSolutions.
  3. 3.Anthropic. Claude opus 4.1, 2025. URL https://www.anthropic.com/news/claude-opus-4-1.
  4. 4.Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin. Qwen2.5-vl technical report, 2025.
  5. 5.Gilad Baruch, Zhuoyuan Chen, Afshin Dehghan, Tal Dimry, Yuri Feigin, Peter Fu, Thomas Gebauer, Brandon Joffe, Daniel Kurz, Arik Schwartz, et al. Arkitscenes: A diverse real-world dataset for 3d indoor scene understanding using mobile rgb-d data. arXiv preprint arXiv:2111.08897, 2021.
  6. 6.Garrick Brazil, Abhinav Kumar, Julian Straub, Nikhila Ravi, Justin Johnson, and Georgia Gkioxari. Omni3d: A large benchmark and model for 3d object detection in the wild. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 13154–13164, 2023.
  7. 7.Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Jiaqi Wang, Yu Qiao, Dahua Lin, et al. Are we on the right way for evaluating large vision-language models? arXiv:2403.20330, 2024a.
  8. 8.Shimin Chen, Xiaohan Lan, Yitian Yuan, Zequn Jie, and Lin Ma. Timemarker: A versatile video-llm for long and short video understanding with superior temporal localization ability. arXiv preprint arXiv:2411.18211, 2024b.
  9. 9.Yitong Chen, Lingchen Meng, Wujian Peng, Zuxuan Wu, and Yu-Gang Jiang. Comp: Continual multi-modal pre-training for vision foundation models. arXiv preprint arXiv:2503.18931, 2025.
  10. 10.Kanzhi Cheng, Qiushi Sun, Yougang Chu, Fangzhi Xu, Yantao Li, Jianbing Zhang, and Zhiyong Wu. Seeclick: Harnessing gui grounding for advanced visual gui agents. arXiv preprint arXiv:2401.10935, 2024.
  11. 11.Xianfu Cheng, Wei Zhang, Shiwei Zhang, Jian Yang, Xiangyuan Guan, Xianjie Wu, Xiang Li, Ge Zhang, Jiaheng Liu, Yuying Mai, et al. Simplevqa: Multimodal factuality evaluation for multimodal large language models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 4637–4646, 2025.
  12. 12.Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, et al. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261, 2025.
  13. 13.Matt Deitke, Christopher Clark, Sangho Lee, Rohun Tripathi, Yue Yang, Jae Sung Park, Mohammadreza Salehi, Niklas Muennighoff, Kyle Lo, Luca Soldaini, et al. Molmo and pixmo: Open weights and open data for state-of-the-art multimodal models. arXiv preprint arXiv:2409.17146, 2024.
  14. 14.Shizhe Diao, Yu Yang, Yonggan Fu, Xin Dong, Dan Su, Markus Kliegl, Zijia Chen, Peter Belcak, Yoshi Suhara, Hongxu Yin, et al. Climb: Clustering-based iterative data mixture bootstrapping for language model pre-training. arXiv preprint arXiv:2504.13161, 2025.
  15. 15.Matthijs Douze, Alexandr Guzhva, Chengqi Deng, Jeff Johnson, Gergely Szilvasy, Pierre-Emmanuel Mazaré, Maria Lomeli, Lucas Hosseini, and Hervé Jégou. The faiss library. 2024.
  16. 16.Mengfei Du, Binhao Wu, Zejun Li, Xuanjing Huang, and Zhongyu Wei. Embspatial-bench: Benchmarking spatial understanding for embodied tasks with large vision-language models. arXiv preprint arXiv:2406.05756, 2024.
  17. 17.Chengqi Duan, Kaiyue Sun, Rongyao Fang, Manyuan Zhang, Yan Feng, Ying Luo, Yufang Liu, Ke Wang, Peng Pei, Xunliang Cai, et al. Codeplot-cot: Mathematical visual reasoning by thinking with code-driven images. arXiv preprint arXiv:2510.11718, 2025.
  18. 18.Chaoyou Fu, Yuhan Dai, Yondong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, et al. Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis. arXiv:2405.21075, 2024a.
  19. 19.Ling Fu, Biao Yang, Zhebin Kuang, Jiajun Song, Yuzhe Li, Linghao Zhu, Qidi Luo, Xinyu Wang, Hao Lu, Mingxin Huang, Zhang Li, Guozhi Tang, Bin Shan, Chunhui Lin, Qi Liu, Binghong Wu, Hao Feng, Hao Liu, Can Huang, Jingqun Tang, Wei Chen, Lianwen Jin, Yuliang Liu, and Xiang Bai. Ocrbench v2: An improved benchmark for evaluating large multimodal models on visual text localization and reasoning, 2024b. URL https://arxiv.org/abs/2501.00321.
  20. 20.Xingyu Fu, Yushi Hu, Bangzheng Li, Yu Feng, Haoyu Wang, Xudong Lin, Dan Roth, Noah A Smith, Wei-Chiu Ma, and Ranjay Krishna. Blink: Multimodal large language models can see but not perceive. In European Conference on Computer Vision, pp. 148–166. Springer, 2024c.
  21. 21.Chang Gao, Chujie Zheng, Xiong-Hui Chen, Kai Dang, Shixuan Liu, Bowen Yu, An Yang, Shuai Bai, Jingren Zhou, and Junyang Lin. Soft adaptive policy optimization. arXiv preprint arXiv:2511.20347, 2025.
  22. 22.Jiyang Gao, Chen Sun, Zhenheng Yang, and Ram Nevatia. Tall: Temporal activity localization via language query. In Proceedings of the IEEE international conference on computer vision, pp. 5267–5275, 2017.
  23. 23.Aryo Pradipta Gema, Joshua Ong Jun Leang, Giwon Hong, Alessio Devoto, Alberto Carlo Maria Mancino, Rohit Saxena, Xuanli He, Yu Zhao, Xiaotang Du, Mohammad Reza Ghasemi Madani, Claire Barale, Robert McHardy, Joshua Harris, Jean Kaddour, Emile van Krieken, and Pasquale Minervini. Are we done with mmlu? CoRR, abs/2406.04127, 2024. doi: 10.48550/ARXIV.2406.04127. URL https://doi.org/10.48550/arXiv.2406.04127.
  24. 24.Tianrui Guan, Fuxiao Liu, Xiyang Wu, Ruiqi Xian, Zongxia Li, Xiaoyu Liu, Xijun Wang, Lichang Chen, Furong Huang, Yaser Yacoob, Dinesh Manocha, and Tianyi Zhou. Hallusionbench: An advanced diagnostic suite for entangled language hallucination & visual illusion in large vision-language models, 2023.
  25. 25.Yun He, Di Jin, Chaoqi Wang, Chloe Bi, Karishma Mandyam, Hejia Zhang, Chen Zhu, Ning Li, Tengyu Xu, Hongjiang Lv, Shruti Bhosale, Chenguang Zhu, Karthik Abinav Sankararaman, Eryk Helenowski, Melanie Kambadur, Aditya Tayade, Hao Ma, Han Fang, and Sinong Wang. Multi-if: Benchmarking llms on multi-turn and multilingual instructions following. CoRR, abs/2410.15553, 2024. doi: 10.48550/ARXIV.2410.15553. URL https://doi.org/10.48550/arXiv.2410.15553.
  26. 26.HMMT. Hmmt 2025. https://www.hmmt.org, 2025.
  27. 27.Kairui Hu, Penghao Wu, Fanyi Pu, Wang Xiao, Yuanhan Zhang, Xiang Yue, Bo Li, and Ziwei Liu. Video-mmmu: Evaluating knowledge acquisition from multi-discipline professional videos. arXiv preprint arXiv:2501.13826, 2025.
  28. 28.Jie Huang, Xuejing Liu, Sibo Song, Ruibing Hou, Hong Chang, Junyang Lin, and Shuai Bai. Revisiting multimodal positional encoding in vision-language models, 2025.
  29. 29.Naman Jain, King Han, Alex Gu, Wen-Ding Li, Fanjia Yan, Tianjun Zhang, Sida Wang, Armando Solar-Lezama, Koushik Sen, and Ion Stoica. Livecodebench: Holistic and contamination free evaluation of large language models for code. CoRR, abs/2403.07974, 2024. doi: 10.48550/ARXIV.2403.07974. URL https://doi.org/10.48550/arXiv.2403.07974.
  30. 30.Bowen Jin, Hansi Zeng, Zhenrui Yue, Jinsung Yoon, Sercan Arik, Dong Wang, Hamed Zamani, and Jiawei Han. Search-r1: Training llms to reason and leverage search engines with reinforcement learning. arXiv preprint arXiv:2503.09516, 2025.
  31. 31.Jeff Johnson, Matthijs Douze, and Hervé Jégou. Billion-scale similarity search with GPUs. IEEE Transactions on Big Data, 7(3):535–547, 2019.
  32. 32.Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara Berg. Referitgame: Referring to objects in photographs of natural scenes. In EMNLP, 2014.
  33. 33.Aniruddha Kembhavi, Michael Salvato, Eric Kolve, Minjoon Seo, Hannaneh Hajishirzi, and Ali Farhadi. A diagram is worth a dozen images. ArXiv, abs/1603.07396, 2016.
  34. 34.Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper Uijlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci, Alexander Kolesnikov, et al. The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale. International journal of computer vision, pp. 1956–1981, 2020.
  35. 35.Xin Lai, Junyi Li, Wei Li, Tao Liu, Tianjian Li, and Hengshuang Zhao. Mini-o3: Scaling up reasoning patterns and interaction turns for visual search. arXiv preprint arXiv:2509.07969, 2025.
  36. 36.Hugo Laurençon, Lucile Saulnier, Léo Tronchon, Stas Bekman, Amanpreet Singh, Anton Lozhkov, Thomas Wang, Siddharth Karamcheti, Alexander Rush, Douwe Kiela, et al. Obelics: An open web-scale filtered dataset of interleaved image-text documents. Advances in Neural Information Processing Systems, 36: 71683–71702, 2023.
  37. 37.Jinke Li, Jiarui Yu, Chenxing Wei, Hande Dong, Qiang Lin, Liangjing Yang, Zhicai Wang, and Yanbin Hao. Unisvg: A unified dataset for vector graphic understanding and generation with multimodal large language models. In Proceedings of the 33rd ACM International Conference on Multimedia, pp. 13156–13163, 2025a.
  38. 38.Kaixin Li, Yuchen Tian, Qisheng Hu, Ziyang Luo, Zhiyong Huang, and Jing Ma. Mmcode: Benchmarking multimodal large language models for code generation with visually rich programming problems. In Findings of the Association for Computational Linguistics: EMNLP 2024, pp. 736–783, 2024a.
  39. 39.Kaixin Li, Ziyang Meng, Hongzhan Lin, Ziyang Luo, Yuchen Tian, Jing Ma, Zhiyong Huang, and Tat-Seng Chua. Screenspot-pro: Gui grounding for professional high-resolution computer use, 2025b. URL https://likaixin2000.github.io/papers/ScreenSpot_Pro.pdf. Preprint.
  40. 40.Kaixin Li et al. Iconstack, 2025c. URL https://huggingface.co/datasets/likaixin/IconStack-48M-Rendered-Train.
  41. 41.Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Yi Liu, Zun Wang, Jilan Xu, Guo Chen, Ping Luo, et al. Mvbench: A comprehensive multi-modal video understanding benchmark. In CVPR, 2024b.
  42. 42.Liunian Harold Li, Pengchuan Zhang, Haotian Zhang, Jianwei Yang, Chunyuan Li, Yiwu Zhong, Lijuan Wang, Lu Yuan, Lei Zhang, Jenq-Neng Hwang, et al. Grounded language-image pre-training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10965–10975, 2022.
  43. 43.Qingyun Li, Zhe Chen, Weiyun Wang, Wenhai Wang, Shenglong Ye, Zhenjiang Jin, Guanzhou Chen, Yinan He, Zhangwei Gao, Erfei Cui, et al. Omnicorpus: An unified multimodal corpus of 10 billion-level images interleaved with text. arXiv preprint arXiv:2406.08418, 2024c.
  44. 44.Tianle Li, Wei-Lin Chiang, Evan Frick, Lisa Dunlap, Tianhao Wu, Banghua Zhu, Joseph E. Gonzalez, and Ion Stoica. From crowdsourced data to high-quality benchmarks: Arena-hard and benchbuilder pipeline. CoRR, abs/2406.11939, 2024d. doi: 10.48550/ARXIV.2406.11939. URL https://doi.org/10.48550/arXiv.2406.11939.
  45. 45.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In ECCV, 2014.
  46. 46.Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chun yue Li, Jianwei Yang, Hang Su, Jun-Juan Zhu, and Lei Zhang. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. arXiv:2303.05499, 2023a.
  47. 47.Yuan Liu, Haodong Duan, Bo Li Yuanhan Zhang, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, Kai Chen, and Dahua Lin. Mmbench: Is your multi-modal model an all-around player? arXiv:2307.06281, 2023b.
  48. 48.Yuliang Liu, Zhang Li, Mingxin Huang, Biao Yang, Wenwen Yu, Chunyuan Li, Xu-Cheng Yin, Cheng-Lin Liu, Lianwen Jin, and Xiang Bai. Ocrbench: on the hidden mystery of ocr in large multimodal models. Science China Information Sciences, 67(12), December 2024. ISSN 1869-1919. doi: 10.1007/s11432-024-4235-6. URL http://dx.doi.org/10.1007/s11432-024-4235-6.
  49. 49.Dunjie Lu, Yiheng Xu, Junli Wang, Haoyuan Wu, Xinyuan Wang, Zekun Wang, Junlin Yang, Hongjin Su, Jixuan Chen, Junda Chen, Yuchen Mao, Jingren Zhou, Junyang Lin, Binyuan Hui, and Tao Yu. Videoagenttrek: Computer use pretraining from unlabeled videos, 2025. URL https://arxiv.org/abs/2510.19488.
  50. 50.Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao. Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts. arXiv preprint arXiv:2310.02255, 2023.
  51. 51.Yubo Ma, Yuhang Zang, Liangyu Chen, Meiqi Chen, Yizhu Jiao, Xinze Li, Xinyuan Lu, Ziyu Liu, Yan Ma, Xiaoyi Dong, et al. Mmlongbench-doc: Benchmarking long-context document understanding with visualizations. Advances in Neural Information Processing Systems, 37:95963–96010, 2024.
  52. 52.Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy. Generation and comprehension of unambiguous object descriptions. In CVPR, 2016.
  53. 53.Ahmed Masry, Do Xuan Long, Jia Qing Tan, Shafiq Joty, and Enamul Hoque. Chartqa: A benchmark for question answering about charts with visual and logical reasoning. arXiv:2203.10244, 2022.
  54. 54.Minesh Mathew, Viraj Bagal, Rubèn Pérez Tito, Dimosthenis Karatzas, Ernest Valveny, and C.V. Jawahar. Infographicvqa. 2022 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pp. 2582–2591, 2021a.
  55. 55.Minesh Mathew, Dimosthenis Karatzas, and CV Jawahar. Docvqa: A dataset for vqa on document images. In WACV, 2021b.
  56. 56.Lingchen Meng, Jianwei Yang, Rui Tian, Xiyang Dai, Zuxuan Wu, Jianfeng Gao, and Yu-Gang Jiang. Deepstack: Deeply stacking visual tokens is surprisingly simple and effective for lmms. In Advances in Neural Information Processing Systems, volume 37, pp. 23464–23487, 2024.
  57. 57.OpenAI. Gpt-5 system card, 2025. URL https://cdn.openai.com/gpt-5-system-card.pdf.
  58. 58.Linke Ouyang, Yuan Qu, Hongbin Zhou, Jiawei Zhu, Rui Zhang, Qunshu Lin, Bin Wang, Zhiyuan Zhao, Man Jiang, Xiaomeng Zhao, Jin Shi, Fan Wu, Pei Chu, Minghao Liu, Zhenxiang Li, Chao Xu, Bo Zhang, Botian Shi, Zhongying Tu, and Conghui He. Omnidocbench: Benchmarking diverse pdf document parsing with comprehensive annotations, 2024. URL https://arxiv.org/abs/2412.07626.
  59. 59.Samuel J. Paech. Eq-bench: An emotional intelligence benchmark for large language models. CoRR, abs/2312.06281, 2023. doi: 10.48550/ARXIV.2312.06281. URL https://doi.org/10.48550/arXiv.2312.06281.
  60. 60.Roni Paiss, Ariel Ephrat, Omer Tov, Shiran Zada, Inbar Mosseri, Michal Irani, and Tali Dekel. Teaching clip to count to ten. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 3170–3180, 2023.
  61. 61.Shishir G. Patil, Huanzhi Mao, Charlie Cheng-Jie Ji, Fanjia Yan, Vishnu Suresh, Ion Stoica, and Joseph E. Gonzalez. The berkeley function calling leaderboard (bfcl): From tool use to agentic evaluation of large language models. In Advances in Neural Information Processing Systems, 2024.
  62. 62.Yusu Qian, Hanrong Ye, Jean-Philippe Fauconnier, Peter Grasch, Yinfei Yang, and Zhe Gan. Mia-bench: Towards better instruction following evaluation of multimodal llms. arXiv preprint arXiv:2407.01509, 2024.
  63. 63.Runqi Qiao, Qiuna Tan, Guanting Dong, Minhui Wu, Chong Sun, Xiaoshuai Song, Zhuoma GongQue, Shanglin Lei, Zhe Wei, Miaoxuan Zhang, et al. We-math: Does your large multimodal model achieve human-like mathematical reasoning? arXiv preprint arXiv:2407.01284, 2024.
  64. 64.Pooyan Rahmanzadehgervi, Logan Bolton, Mohammad Reza Taesiri, and Anh Totti Nguyen. Vision language models are blind: Failing to translate detailed visual features into words, 2025. URL https://arxiv.org/abs/2407.06581.
  65. 65.Christopher Rawles, Sarah Clinckemaillie, Yifan Chang, Jonathan Waltz, Gabrielle Lau, Marybeth Fair, Alice Li, William Bishop, Wei Li, Folawiyo Campbell-Ajala, et al. Androidworld: A dynamic benchmarking environment for autonomous agents. arXiv:2405.14573, 2024.
  66. 66.David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R. Bowman. GPQA: A graduate-level google-proof q&a benchmark. CoRR, abs/2311.12022, 2023. doi: 10.48550/ARXIV.2311.12022. URL https://doi.org/10.48550/arXiv.2311.12022.
  67. 67.Jonathan Roberts, Mohammad Reza Taesiri, Ansh Sharma, Akash Gupta, Samuel Roberts, Ioana Croitoru, Simion-Vlad Bogolin, Jialu Tang, Florian Langer, et al. Zerobench: An impossible visual benchmark for contemporary large multimodal models, 2025. URL https://arxiv.org/abs/2502.09696.
  68. 68.Mike Roberts, Jason Ramapuram, Anurag Ranjan, Atulit Kumar, Miguel Angel Bautista, Nathan Paczan, Russ Webb, and Joshua M Susskind. Hypersim: A photorealistic synthetic dataset for holistic indoor scene understanding. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 10912–10922, 2021.
  69. 69.Angelika Romanou, Negar Foroutan, Anna Sotnikova, Zeming Chen, Sree Harsha Nelaturu, Shivalika Singh, Rishabh Maheshwary, Micol Altomare, Mohamed A. Haggag, Imanol Schlag, et al. INCLUDE: evaluating multilingual language understanding with regional knowledge. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net, 2025.
  70. 70.Shuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng, Gang Yu, Xiangyu Zhang, Jing Li, and Jian Sun. Objects365: A large-scale, high-quality dataset for object detection. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 8430–8439, 2019.
  71. 71.Chenglei Si, Yanzhe Zhang, Ryan Li, Zhengyuan Yang, Ruibo Liu, and Diyi Yang. Design2code: Benchmarking multimodal code generation for automated front-end engineering. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pp. 3956–3974, 2025.
  72. 72.Chan Hee Song, Valts Blukis, Jonathan Tremblay, Stephen Tyree, Yu Su, and Stan Birchfield. Robospatial: Teaching spatial understanding to 2d and 3d vision-language models for robotics. In Proceedings of the Computer Vision and Pattern Recognition Conference, pp. 15768–15780, 2025a.
  73. 73.Shuran Song, Samuel P Lichtenberg, and Jianxiong Xiao. Sun rgb-d: A rgb-d scene understanding benchmark suite. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 567–576, 2015.
  74. 74.Yueqi Song, Tianyue Ou, Yibo Kong, Zecheng Li, Graham Neubig, and Xiang Yue. Visualpuzzles: Decoupling multimodal reasoning evaluation from domain knowledge. arXiv preprint arXiv:2504.10342, 2025b. URL https://arxiv.org/abs/2504.10342.
  75. 75.Gemini Robotics Team, Saminda Abeyruwan, Joshua Ainslie, Jean-Baptiste Alayrac, Montserrat Gonzalez Arenas, Travis Armstrong, Ashwin Balakrishna, Robert Baruch, Maria Bauza, Michiel Blokzijl, et al. Gemini robotics: Bringing ai into the physical world. arXiv preprint arXiv:2503.20020, 2025.
  76. 76.M-A-P Team. Supergpqa: Scaling LLM evaluation across 285 graduate disciplines. CoRR, abs/2502.14739, 2025. doi: 10.48550/ARXIV.2502.14739. URL https://doi.org/10.48550/arXiv.2502.14739.
  77. 77.Michael Tschannen, Alexey Gritsenko, Xiao Wang, Muhammad Ferjad Naeem, Ibrahim Alabdulmohsin, Nikhil Parthasarathy, Talfan Evans, Lucas Beyer, Ye Xia, Basil Mustafa, et al. Siglip 2: Multilingual vision-language encoders with improved semantic understanding, localization, and dense features. arXiv preprint arXiv:2502.14786, 2025.
  78. 78.Fei Wang, Xingyu Fu, James Y Huang, Zekun Li, Qin Liu, Xiaogeng Liu, Mingyu Derek Ma, Nan Xu, Wenxuan Zhou, Kai Zhang, et al. Muirbench: A comprehensive benchmark for robust multi-image understanding. arXiv preprint arXiv:2406.09411, 2024a.
  79. 79.Ke Wang, Junting Pan, Weikang Shi, Zimu Lu, Houxing Ren, Aojun Zhou, Mingjie Zhan, and Hongsheng Li. Measuring multimodal mathematical reasoning with math-vision dataset. Advances in Neural Information Processing Systems, 37:95095–95169, 2024b.
  80. 80.Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Junyang Lin. Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution. arXiv:2409.12191, 2024c.
  81. 81.Weihan Wang, Zehai He, Wenyi Hong, Yean Cheng, Xiaohan Zhang, Ji Qi, Xiaotao Gu, Shiyu Huang, Bin Xu, Yuxiao Dong, et al. Lvbench: An extreme long video understanding benchmark. arXiv preprint arXiv:2406.08035, 2024d.
  82. 82.Wenbin Wang, Liang Ding, Minyan Zeng, Xiabin Zhou, Li Shen, Yong Luo, and Dacheng Tao. Divide, conquer and combine: A training-free framework for high-resolution image perception in multimodal large language models. arXiv preprint, 2024e. URL https://arxiv.org/abs/2408.15556.
  83. 83.Xinyuan Wang, Bowen Wang, Dunjie Lu, Junlin Yang, Tianbao Xie, Junli Wang, Jiaqi Deng, Xiaole Guo, Yiheng Xu, Chen Henry Wu, et al. Opencua: Open foundations for computer-use agents. arXiv preprint arXiv:2508.09123, 2025a.
  84. 84.Yiming Wang, Pei Zhang, Jialong Tang, Haoran Wei, Baosong Yang, Rui Wang, Chenshu Sun, Feitong Sun, Jiran Zhang, Junxuan Wu, Qiqian Cang, Yichang Zhang, Fei Huang, Junyang Lin, et al. Polymath: Evaluating mathematical reasoning in multilingual contexts. CoRR, abs/2504.18428, 2025b. doi: 10.48550/ARXIV.2504.18428. URL https://doi.org/10.48550/arXiv.2504.18428.
  85. 85.Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jiang, Tianle Li, et al. MMLU-Pro: A more robust and challenging multi-task language understanding benchmark. CoRR, abs/2406.01574, 2024f.
  86. 86.Zhexu Wang, Yiping Liu, Yejie Wang, Wenyang He, Bofei Gao, Muxi Diao, Yanxu Chen, Kelin Fu, Flood Sung, Zhilin Yang, Tianyu Liu, and Weiran Xu. Ojbench: A competition level code benchmark for large language models. CoRR, abs/2506.16395, 2025c. doi: 10.48550/ARXIV.2506.16395. URL https://doi.org/10.48550/arXiv.2506.16395.
  87. 87.Zirui Wang, Mengzhou Xia, Luxi He, Howard Chen, Yitao Liu, Richard Zhu, Kaiqu Liang, Xindi Wu, Haotian Liu, Sadhika Malladi, Alexis Chevalier, Sanjeev Arora, and Danqi Chen. Charxiv: Charting gaps in realistic chart understanding in multimodal llms. arXiv preprint arXiv:2406.18521, 2024g.
  88. 88.Alexander Wettig, Kyle Lo, Sewon Min, Hannaneh Hajishirzi, Danqi Chen, and Luca Soldaini. Organize the web: Constructing domains enhances pre-training data curation. arXiv preprint arXiv:2502.10341, 2025.
  89. 89.Colin White, Samuel Dooley, Manley Roberts, Arka Pal, Benjamin Feuer, Siddhartha Jain, Ravid Shwartz-Ziv, Neel Jain, et al. Livebench: A challenging, contamination-free LLM benchmark. CoRR, abs/2406.19314, 2024. doi: 10.48550/ARXIV.2406.19314. URL https://doi.org/10.48550/arXiv.2406.19314.
  90. 90.Jinming Wu, Zihao Deng, Wei Li, Yiding Liu, Bo You, Bo Li, Zejun Ma, and Ziwei Liu. Mmsearch-r1: Incentivizing lmms to search. arXiv preprint arXiv:2506.20670, 2025a.
  91. 91.Penghao Wu and Saining Xie. V*: Guided visual search as a core mechanism in multimodal llms. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 13084–13094, June 2024.
  92. 92.Yuning Wu, Jiahao Mei, Ming Yan, Chenliang Li, Shaopeng Lai, Yuran Ren, Zijia Wang, Ji Zhang, Mengyue Wu, Qin Jin, and Fei Huang. Writingbench: A comprehensive benchmark for generative writing. CoRR, abs/2503.05244, 2025b. doi: 10.48550/ARXIV.2503.05244. URL https://doi.org/10.48550/arXiv.2503.05244.
  93. 93.xAI. Realworldqa: A benchmark for real-world spatial understanding. https://huggingface.co/datasets/xai-org/RealworldQA, 2024. Accessed: 2025-04-26.
  94. 94.Yijia Xiao, Edward Sun, Tianyu Liu, and Wei Wang. Logicvista: Multimodal llm logical reasoning benchmark in visual contexts. arXiv preprint arXiv:2407.04973, 2024.
  95. 95.Tianbao Xie, Jiaqi Deng, Xiaochuan Li, Junlin Yang, Haoyuan Wu, Jixuan Chen, Wenjing Hu, Xinyuan Wang, Yuhui Xu, Zekun Wang, Yiheng Xu, Junli Wang, Doyen Sahoo, Tao Yu, and Caiming Xiong. Scaling computer-use grounding via user interface decomposition and synthesis, 2025a. URL https://arxiv.org/abs/2505.13227.
  96. 96.Tianbao Xie, Mengqi Yuan, Danyang Zhang, Xinzhuang Xiong, Zhennan Shen, Zilong Zhou, Xinyuan Wang, Yanxu Chen, Jiaqi Deng, Junda Chen, Bowen Wang, Haoyuan Wu, Jixuan Chen, Junli Wang, Dunjie Lu, Hao Hu, and Tao Yu. Introducing osworld-verified. xlang.ai, July 2025b. URL https://xlang.ai/blog/osworld-verified.
  97. 97.Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, et al. Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments. Advances in Neural Information Processing Systems, 37:52040–52094, 2025c.
  98. 98.Weiye Xu, Jiahao Wang, Weiyun Wang, Zhe Chen, Wengang Zhou, Aijun Yang, Lewei Lu, Houqiang Li, Xiaohua Wang, Xizhou Zhu, et al. Visulogic: A benchmark for evaluating visual reasoning in multi-modal large language models, 2025. URL https://arxiv.org/abs/2504.15279.
  99. 99.An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, et al. Qwen3 technical report, 2025a.
  100. 100.Cheng Yang, Chufan Shi, Yaxin Liu, Bo Shui, Junjie Wang, Mohan Jing, Linran Xu, Xinyu Zhu, Siheng Li, Yuxiang Zhang, et al. Chartmimic: Evaluating lmm’s cross-modal reasoning capability via chart-to-code generation. arXiv preprint arXiv:2406.09961, 2024a.
  101. 101.Jihan Yang, Shusheng Yang, Anjali W Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie. Thinking in space: How multimodal large language models see, remember, and recall spaces. In Proceedings of the Computer Vision and Pattern Recognition Conference, pp. 10632–10643, 2025b.
  102. 102.Zhibo Yang, Jun Tang, Zhaohai Li, Pengfei Wang, Jianqiang Wan, Humen Zhong, Xuejing Liu, Mingkun Yang, Peng Wang, Shuai Bai, LianWen Jin, and Junyang Lin. Cc-ocr: A comprehensive and challenging ocr benchmark for evaluating large multimodal models in literacy, 2024b. URL https://arxiv.org/abs/2412.02210.
  103. 103.Jiabo Ye, Xi Zhang, Haiyang Xu, Haowei Liu, Junyang Wang, Zhaoqing Zhu, Ziwei Zheng, et al. Mobile-agent-v3: Fundamental agents for gui automation. arXiv preprint arXiv:2508.15144, 2025.
  104. 104.Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, et al. Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 9556–9567, 2024a.
  105. 105.Xiang Yue, Tianyu Zheng, Yuansheng Ni, Yubo Wang, Kai Zhang, Shengbang Tong, Yuxuan Sun, Botao Yu, Ge Zhang, Huan Sun, et al. Mmmu-pro: A more robust multi-discipline multimodal understanding benchmark. arXiv preprint arXiv:2409.02813, 2024b.
  106. 106.Renrui Zhang, Dongzhi Jiang, Yichi Zhang, Haokun Lin, Ziyu Guo, Pengshuo Qiu, Aojun Zhou, Pan Lu, Kai-Wei Chang, Yu Qiao, et al. Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pp. 169–186. Springer, 2024.
  107. 107.Yilun Zhao, Lujing Xie, Haowei Zhang, Guo Gan, Yitao Long, Zhiyuan Hu, Tongyan Hu, Weiyuan Chen, Chuhan Li, Junyang Song, Zhijian Xu, Chengye Wang, et al. Mmvu: Measuring expert-level multi-discipline video understanding, 2025. URL https://arxiv.org/abs/2501.12380.
  108. 108.Ziwei Zheng, Michael Yang, Jack Hong, Chenxiao Zhao, Guohai Xu, Le Yang, Chao Shen, and Xing Yu. Deepeyes: Incentivizing" thinking with images" via reinforcement learning. arXiv preprint arXiv:2505.14362, 2025.
  109. 109.Enshen Zhou, Jingkun An, Cheng Chi, Yi Han, Shanyu Rong, Chi Zhang, Pengwei Wang, Zhongyuan Wang, Tiejun Huang, Lu Sheng, et al. Roborefer: Towards spatial referring with reasoning in vision-language models for robotics. arXiv preprint arXiv:2506.04308, 2025.
  110. 110.Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha Brahma, Sujoy Basu, Yi Luan, Denny Zhou, and Le Hou. Instruction-following evaluation for large language models. CoRR, abs/2311.07911, 2023. doi: 10.48550/ARXIV.2311.07911. URL https://doi.org/10.48550/arXiv.2311.07911.
  111. 111.Junjie Zhou, Yan Shu, Bo Zhao, Boya Wu, Shitao Xiao, Xi Yang, Yongping Xiong, Bo Zhang, Tiejun Huang, and Zheng Liu. Mlvu: A comprehensive benchmark for multi-task long video understanding. arXiv preprint arXiv:2406.04264, 2024.
  112. 112.Wanrong Zhu, Jack Hessel, Anas Awadalla, Samir Yitzhak Gadre, Jesse Dodge, Alex Fang, Youngjae Yu, Ludwig Schmidt, William Yang Wang, and Yejin Choi. Multimodal c4: An open, billion-scale corpus of images interleaved with text. Advances in Neural Information Processing Systems, 36:8958–8974, 2023.
  113. 113.Chengke Zou, Xingang Guo, Rui Yang, Junyu Zhang, Bin Hu, and Huan Zhang. Dynamath: A dynamic visual benchmark for evaluating mathematical reasoning robustness of vision language models. arXiv preprint arXiv:2411.00836, 2024.

Citation

MLA
Bai, S., et al. “Qwen3-VL Technical Report”. arXiv, 2025, http://arxiv.org/abs/2511.21631v2.
APA
Bai, S., Cai, Y., Chen, R., Chen, K., Chen, X., Cheng, Z., Deng, L., Ding, W., Gao, C., Ge, C., Ge, W., Guo, Z., Huang, Q., Huang, J., Huang, F., Hui, B., Jiang, S., Li, Z., Li, M., … Zhu, K. (2025). Qwen3-VL Technical Report. arXiv. http://arxiv.org/abs/2511.21631v2
Chicago
Bai, S., Y. Cai, R. Chen, et al. 2025. “Qwen3-VL Technical Report”. arXiv. http://arxiv.org/abs/2511.21631v2.
Harvard
Bai, S. et al. (2025) “Qwen3-VL Technical Report”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2511.21631v2.
Vancouver
1. Bai S, Cai Y, Chen R, et al (2025) Qwen3-VL Technical Report. arXiv

BibTeX

@article{bai2025qwen3,
  title = {Qwen3-VL Technical Report},
  author = {Bai, Shuai and Cai, Yuxuan and Chen, Ruizhe and Chen, Keqin and Chen, Xionghui and Cheng, Zesen and Deng, Lianghao and Ding, Wei and Gao, Chang and Ge, Chunjiang and Ge, Wenbin and Guo, Zhifang and Huang, Qidong and Huang, Jie and Huang, Fei and Hui, Binyuan and Jiang, Shutong and Li, Zhaohai and Li, Mingsheng and Li, Mei and Li, Kaixin and Lin, Zicheng and Lin, Junyang and Liu, Xuejing and Liu, Jiawei and Liu, Chenglong and Liu, Yang and Liu, Dayiheng and Liu, Shixuan and Lu, Dunjie and Luo, Ruilin and Lv, Chenxu and Men, Rui and Meng, Lingchen and Ren, Xuancheng and Ren, Xingzhang and Song, Sibo and Sun, Yuchong and Tang, Jun and Tu, Jianhong and Wan, Jianqiang and Wang, Peng and Wang, Pengfei and Wang, Qiuyue and Wang, Yuxuan and Xie, Tianbao and Xu, Yiheng and Xu, Haiyang and Xu, Jin and Yang, Zhibo and Yang, Mingkun and Yang, Jianxin and Yang, An and Yu, Bowen and Zhang, Fei and Zhang, Hang and Zhang, Xi and Zheng, Bo and Zhong, Humen and Zhou, Jingren and Zhou, Fan and Zhou, Jing and Zhu, Yuanzhi and Zhu, Ke},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2511.21631v2},
  eprint = {2511.21631}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors