SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Dongyang LiuRenrui ZhangLongtian QiuSiyuan HuangWeifeng LinShitian ZhaoShijie GengZiyi LinPeng JinKaipeng Zhang

article2024ICML148 citations

Presents an efficient multimodal architecture and a unified one-stage training framework across multiple base language models ranging from 1.1B to 8×7B parameters, demonstrating how systematic scaling of both training data and model capacity boosts multimodal performance across diverse visual-language tasks.

Listen

Recent advancements in artificial intelligence have accelerated the development of multimodal systems capable of interpreting both visual imagery and natural language. However, many current open-source models remain constrained by narrow training datasets that falter in specialized domains—such as reading text-dense documents, interpreting charts, and solving visual mathematics—and by limited model size options that are either too large for mobile devices or too small for complex reasoning. The article addresses these challenges by developing SPHINX-X, an adaptable family of multimodal large language models designed to scale training data coverage and parameter sizes efficiently while streamlining model architecture.

To achieve this, the authors restructured their prior framework into a more efficient, single-stage training pipeline. They reduced visual processing overhead by pairing two complementary vision encoders, eliminated redundant processing of empty image padding with learnable skip tokens, and trained the entire system at once. The training utilized a broad collection of multimodal data, including millions of conversational samples, visual detection and pose tasks, an in-house document dataset of three million text-dense pages, and detailed visual marking annotations. The authors applied this unified training across four base language models ranging from a compact 1.1-billion parameter model to a high-capacity mixture-of-experts architecture activating subsets of an eight-by-seven-billion parameter base.

Benchmarking revealed several key outcomes. First, scaling both model size and dataset diversity produced consistent performance improvements across diverse tasks. The expanded 13-billion parameter version systematically outperformed its predecessor across standard visual and text benchmarks. Second, the mixture-of-experts configuration achieved leading results among open-source models on mathematical and scientific reasoning, matching or exceeding proprietary models such as GPT-4V on specialized visual perception and user interface localization tests. Third, despite being trained solely on still images, the models outperformed dedicated video architectures on video question-answering benchmarks when evaluated on sampled video frames. Finally, the compact 1.1-billion parameter model preserved viable multimodal capabilities, demonstrating feasibility for edge and mobile environments.

These findings indicate that consolidating complex, multi-stage training into a single comprehensive workflow reduces engineering friction and computational waste without sacrificing capability. For organizations evaluating artificial intelligence integration, the results show that generalist models trained on diverse domain data can effectively handle specialized tasks such as optical character recognition, document layout analysis, and interface grounding without requiring separate, dedicated models for each use case.

Decision-makers can evaluate a tiered adoption strategy, deploying compact models for on-device applications and larger mixture-of-experts models for complex enterprise reasoning. However, users should note key limitations: the models exhibited weaker results on multidisciplinary academic tests due to gaps in college-level subject training data, and video tracking across time remains limited without dedicated temporal fine-tuning. Future efforts should focus on integrating multi-discipline educational datasets and optimizing expert pruning to enhance deployment efficiency.

No sufficiently relevant recommendations were found.

Cover for SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Abstract

We propose SPHINX-X, an extensive Multi-modality Large Language Model (MLLM) series developed upon SPHINX. To improve the architecture and training efficiency, we modify the SPHINX framework by removing redundant visual encoders, bypassing fully-padded sub-images with skip tokens, and simplifying multi-stage training into a one-stage all-in-one paradigm. To fully unleash the potential of MLLMs, we assemble a comprehensive multi-domain and multi-modal dataset covering publicly available resources in language, vision, and vision-language tasks. We further enrich this collection with our curated OCR intensive and Set-of-Mark datasets, extending the diversity and generality. By training over different base LLMs including TinyLlama-1.1B, InternLM2-7B, LLaMA2-13B, and Mixtral-8×7B, we obtain a spectrum of MLLMs that vary in parameter size and multilingual capabilities. Comprehensive benchmarking reveals a strong correlation between the multi-modal performance with the data and parameter scales. Code and models are released at https://github.com/Alpha-VLLM/LLaMA2-Accessory.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. A Revisit of SPHINX
  • 3.2. SPHINX-X
  • Bypassing Fully-padded Sub-images with Skip Tokens.
  • 3.3. Training Data of SPHINX-X
  • 3.4. SPHINX-X with Different LLMs
  • 4. Experiment
  • 4.1. Experimental Settings
  • 4.2. Performance Evaluation
  • 4.3. SPHINX-MoE on other MLLM Benchmarks
  • 4.4. Performance of SPHINX-Plus on Video Analysis
  • 4.5. Demonstrations of SPHINX-X
  • 5. Conclusion
  • Acknowledgements
  • Impact Statement
  • References
  • A. Appendix
  • A.1. Analysis of Routing Mechanisms in SPHINX-MoE
  • A.1.1. INFERENCE WITH DIFFERENT NUMBERS OF ACTIVATING EXPERTS
  • A.1.2. EXPERTS' USAGE DISTRIBUTION ON DIFFERENT DOMAINS AND DIFFERENT MODALITIES
  • A.1.3. PRUNE SOME OF THE EXPERTS WHEN INFERENCE
  • A.2. Video Analysis on MVBench
  • A.3. Additional details on the training dataset

Knowls

  1. Knowl 1 — SPHINX-X retains two complementary visual encoders

    model/method

    SPHINX-X simplifies the original SPHINX visual encoder mixture by retaining DINOv2 and CLIP-ConvNeXt and removing CLIP-ViT and the Q-former. The retained encoders combine different pretraining approaches (self-supervised and weakly supervised) and architectures (vision transformer and convolutional network); the authors call this pair the Mixture of Visual experts (MoV). The change is intended to preserve complementary visual representations while reducing the computation spent encoding images and their high-resolution sub-images. In the reported training setup, the visual encoders are frozen while the other model modules are optimized.

  2. Knowl 2 — Learnable skip tokens bypass fully padded image crops

    model/method

    For high-resolution inputs, SPHINX-X uses a downsampled global image together with local sub-images. In the standard 448 × 448 setup, an image is divided into four 224 × 224 sub-images after scaling and zero-padding; images with large aspect ratios can produce crops containing only padding. SPHINX-X replaces each such fully padded crop with a learnable skip token rather than encoding its zero pixels. The token preserves the crop's relative position in the sequence presented to the language model, while avoiding unnecessary visual encoding and reducing the language-model input sequence length.

  3. Knowl 3 — One-stage all-in-one multimodal instruction tuning

    model/method

    SPHINX-X replaces SPHINX's two-stage training and weight-mixing procedure with a single training stage over the collected datasets. Data from different tasks is converted into a unified multi-turn conversational format and used together, rather than being assigned to manually chosen stage-specific combinations. Training updates the language model and intermediate projection layers while keeping the two MoV visual encoders fixed. The stated motivation is to simplify large-scale multi-task training while maintaining performance.

  4. Knowl 4 — Broad training corpus and targeted OCR and Set-of-Marks data

    data/table

    SPHINX-X combines language, visual-task, and vision-language instruction data with two targeted multimodal collections. The reported sample counts are: language data—1.8M multi-turn-dialog examples, 0.6M math examples, and 80k coding examples; visual instruction data—4.9M detection, 0.3M human-pose, 1M classification, and 1M grounding examples; vision-language instruction data—0.7M VQA, 0.5M captioning, and 0.4M visual-instruction examples; targeted data—3M OCR examples from PaperText, 1.0M text-layout and spotting examples, and Set-of-Marks (SoM) examples consisting of 5k natural-image, 1k website/mobile/desktop-agent, 2k OCR-related, 1k document-image, and 1k multipanel examples. The corpus covers conversation, reasoning, coding, detection, pose, grounding, general and text-oriented VQA, captions, and visual instruction following; the full collection is used in the one-stage training procedure.

    PaperText was created from large-scale PDFs collected from Common Crawl and arXiv. PyMuPDF was used to render pages and extract text with bounding boxes; Unicode checks and text-splitting/merging procedures were applied before converting the approximately 3M text-dense pages into question-answer examples. For SoM data, annotations such as boxes and masks were used to place points, boxes, polygons, or identifiers on images from different domains. GPT-4V was prompted with these marked images to produce global captions, region descriptions, and object-relation analyses. During SPHINX-X training, the model receives the unmarked image; the marks are represented in the conversation as language and coordinates.

  5. Knowl 5 — Four base language models define the SPHINX-X family

    model/method

    SPHINX-X is trained with four base language models to cover different parameter scales and capabilities: SPHINX-Tiny uses TinyLlama-1.1B for a compact model; SPHINX-Intern2 uses InternLM2-7B, chosen for bilingual Chinese-English language capability; SPHINX-Plus uses LLaMA2-13B and the expanded dataset, and is initialized from the original SPHINX model; SPHINX-MoE uses Mixtral-8×7B, whose transformer layers have eight feed-forward experts with two experts activated per token during training. SPHINX-Plus-2K is a higher-resolution SPHINX-Plus variant: it raises image resolution from 448 × 448 to 672 × 672 and divides the input into a 3 × 3 set of sub-images.

  6. Knowl 6 — Training and optimization configuration

    experimental setup

    All SPHINX-X models use the one-stage training procedure, with every module except the visual encoders optimized. The learning rate is 5 × 10⁻⁶ for SPHINX-MoE and 2 × 10⁻⁵ for the other models; it warms up linearly during the first 0.01 epoch and then decays to zero with a cosine schedule. Optimization uses AdamW with weight decay 0 and betas (0.9, 0.95), with an effective batch size of 256. Training combines ZeRO2-style data parallelism and Megatron-style model parallelism; the model-parallel size is 8 for SPHINX-MoE, 2 for SPHINX-Plus, and 1 for the other variants. SPHINX-Plus is initialized from SPHINX, whereas the other variants use their base language models and original visual encoders with randomly initialized linear projection layers.

  7. Knowl 7 — The larger SPHINX-Plus data mixture improves many—but not all—benchmarks

    empirical result

    SPHINX-Plus and SPHINX both use a LLaMA2-13B base, allowing a comparison of the original system with the expanded-data, one-stage SPHINX-Plus training setup. On representative multimodal benchmarks, SPHINX-Plus scores higher on MM-Bench (71.0 versus 67.1), SEED-Bench (74.8 versus 71.6), MM-Vet (47.9 versus 36.6), MathVista (36.8 versus 27.5), InfiMM-Eval (39.5 versus 30.7), and QBench (68.6 versus 65.8). The gains are not universal: SPHINX scores higher on POPE (90.8 versus 89.1), MME perception (1560.2 versus 1457.7), MME cognition (310.0 versus 283.6), LLaVA-Bench (74.3 versus 71.7), and CCbench (27.9 versus 25.6). The comparison is consistent with benefits from the broader data and training changes, but does not isolate data scale as the sole cause because the training procedure also changes.

  8. Knowl 8 — Larger language-model variants generally score higher, with benchmark-dependent exceptions

    empirical result

    Across the SPHINX-X variants, larger base language models usually achieve stronger results on the reported multimodal benchmarks, although rankings vary by task and the comparison also involves different pretrained language models. For TinyLlama-1.1B, InternLM2-7B, LLaMA2-13B SPHINX-Plus, and Mixtral-8×7B SPHINX-MoE, respectively, POPE scores are 82.2, 86.9, 89.1, and 89.6; MM-Bench scores are 56.6, 57.9, 71.0, and 71.3; and MathVista scores are 26.4, 35.5, 36.8, and 42.7. Other metrics do not rise monotonically: SEED-Bench scores are 17.1, 68.8, 74.8, and 73.0, while MM-Vet scores are 23.8, 36.5, 47.9, and 40.9. Thus, the results support a broad association between model scale and multimodal performance rather than a uniform improvement on every benchmark.

  9. Knowl 9 — SPHINX-MoE is strong on selected visual reasoning benchmarks but weaker on multidisciplinary exams

    empirical result

    On MathVerse, SPHINX-MoE scores 15.6 overall, tying the reported open-source LLaVA-NeXT score of 15.6; its scores across the text-dominant, text-lite, text-only, vision-intensive, vision-dominant, and vision-only subsets are 22.2, 16.4, 18.3, 14.8, 12.6, and 9.1. On SciVerse, SPHINX-MoE scores 37.3 overall, above the reported open-source LLaVA-NeXT score of 34.9; its text-only, knowledge-lite, knowledge-rich, professional-knowledge, vision-dominant, and vision-only scores are 41.1, 38.9, 38.8, 41.3, 36.3, and 31.4. SPHINX-MoE also scores 49.3 on MMVP, compared with 38.7 for GPT-4V, and its AesBench aesthetic perception and expression scores are 72.93 and 73.32, compared with GPT-4V's 72.08 and 70.16. However, its MMMU validation score is 31.1 versus GPT-4V's 56.8, and its CMMMU validation/test scores are 29.3/29.6 versus 42.5/43.7 for GPT-4V. The authors attribute the weakness on multidisciplinary tasks to insufficient multimodal multidisciplinary training data.

  10. Knowl 10 — An image-only SPHINX-Plus model performs well on video QA but misses temporal reasoning

    limitation

    SPHINX-Plus was trained as an image-based model without video training. For the reported Video-Bench evaluation, videos were evenly sampled and the middle frame was used as the model's representative image. Under this setup, SPHINX-Plus obtains an average score of 45.1, compared with 39.0 for the original SPHINX and 38.3 for the highest-scoring video-specialized model listed in that comparison. Its results are especially strong on some visual and prior-knowledge questions, including 68.5 on MSVD-QA and 53.1 on ActivityNet-QA. Performance is weak on MOT, where it scores 11.1; the authors note that this kind of task requires modeling temporal relationships, which the image-only model does not represent.

Coverage note — The appendix's detailed MoE expert-routing and expert-pruning analyses are omitted because they are diagnostic follow-up studies rather than central components of the SPHINX-X method or its main scaling claims.

References

  1. 1.Agrawal, A., Lu, J., Antol, S., Mitchell, M., Zitnick, C. L., Parikh, D., and Batra, D. Vqa: Visual question answering. International Journal of Computer Vision, 123:4 – 31, 2015.
  2. 2.Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al. Flamingo: a visual language model for few-shot learning. Advances in Neural Information Processing Systems, 35:23716–23736, 2022.
  3. 3.Bai, J., Bai, S., Yang, S., Wang, S., Tan, S., Wang, P., Lin, J., Zhou, C., and Zhou, J. Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond. arXiv preprint arXiv:2308.12966, 2023.
  4. 4.Bao, H., Wang, W., Dong, L., Liu, Q., Mohammed, O. K., Aggarwal, K., Som, S., Piao, S., and Wei, F. Vlmo: Unified vision-language pre-training with mixture-of-modality-experts. In Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., and Oh, A. (eds.), Advances in Neural Information Processing, 2022.
  5. 5.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. In Advances in neural information processing systems, pp. 1877–1901, 2020.
  6. 6.Cai, R., Song, Z., Guan, D., Chen, Z., Luo, X., Yi, C., and Kot, A. Benchlmm: Benchmarking cross-style visual capability of large multimodal models. ArXiv, abs/2312.02896, 2023a.
  7. 7.Cai, R., Song, Z., Guan, D., Chen, Z., Luo, X., Yi, C., and Kot, A. Benchlmm: Benchmarking cross-style visual capability of large multimodal models. arXiv preprint arXiv:2312.02896, 2023b.
  8. 8.Cao, Y., Xu, X., Sun, C., Huang, X., and Shen, W. Towards generic anomaly detection and understanding: Large-scale visual-linguistic model (gpt-4v) takes the lead. arXiv preprint arXiv:2311.02782, 2023.
  9. 9.Chen, K., Zhang, Z., Zeng, W., Zhang, R., Zhu, F., and Zhao, R. Shikra: Unleashing multimodal llm’s referential dialogue magic. arXiv preprint arXiv:2306.15195, 2023a.
  10. 10.Chen, L., Li, J., wen Dong, X., Zhang, P., He, C., Wang, J., Zhao, F., and Lin, D. Sharegpt4v: Improving large multi-modal models with better captions. ArXiv, abs/2311.12793, 2023b. URL https://api.semanticscholar.org/CorpusID:265308687.
  11. 11.Chen, W., Wang, H., Chen, J., Zhang, Y., Wang, H., LI, S., Zhou, X., and Wang, W. Y. Tabfact: A large-scale dataset for table-based fact verification. ArXiv, abs/1909.02164, 2019.
  12. 12.Cheng, K., Sun, Q., Chu, Y., Xu, F., Li, Y., Zhang, J., and Wu, Z. Seeclick: Harnessing gui grounding for advanced visual gui agents. arXiv preprint arXiv:2401.10935, 2024.
  13. 13.Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality. https://lmsys.org/blog/2023-03-30-vicuna/, March 2023.
  14. 14.Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311, 2022.
  15. 15.Contributors, O. Opencompass: A universal evaluation platform for foundation models. https://github.com/open-compass/opencompass, 2023.
  16. 16.Dai, W., Li, J., Li, D., Tiong, A. M. H., Zhao, J., Wang, W., Li, B., Fung, P., and Hoi, S. Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023.
  17. 17.Ding, N., Chen, Y., Xu, B., Qin, Y., Zheng, Z., Hu, S., Liu, Z., Sun, M., and Zhou, B. Enhancing chat language models by scaling high-quality instructional conversations. arXiv preprint arXiv:2305.14233, 2023.
  18. 18.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N. An image is worth 16x16 words: Transformers for image recognition at scale. ArXiv, abs/2010.11929, 2020.
  19. 19.Fedus, W., Zoph, B., and Shazeer, N. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. The Journal of Machine Learning Research, 23(1):5232–5270, 2022.
  20. 20.Fu, C., Chen, P., Shen, Y., Qin, Y., Zhang, M., Lin, X., Qiu, Z., Lin, W., Yang, J., Zheng, X., et al. Mme: A comprehensive evaluation benchmark for multimodal large language models. arXiv preprint arXiv:2306.13394, 2023a.
  21. 21.Fu, C., Chen, P., Shen, Y., Qin, Y., Zhang, M., Lin, X., Yang, J., Zheng, X., Li, K., Sun, X., Wu, Y., and Ji, R. Mme: A comprehensive evaluation benchmark for multimodal large language models. arXiv preprint arXiv:2306.13394, 2023b.
  22. 22.Fu, C., Zhang, R., Lin, H., Wang, Z., Gao, T., Luo, Y., Huang, Y., Zhang, Z., Qiu, L., Ye, G., et al. A challenger to gpt-4v? early explorations of gemini in visual expertise. arXiv preprint arXiv:2312.12436, 2023c.
  23. 23.Gao, P., Han, J., Zhang, R., Lin, Z., Geng, S., Zhou, A., Zhang, W., Lu, P., He, C., Yue, X., Li, H., and Qiao, Y. Llama-adapter v2: Parameter-efficient visual instruction model. arXiv preprint arXiv:2304.15010, 2023.
  24. 24.Ge, Z., Xinrun, D., Bei, C., Yiming, L., Tongxu, L., Tianyu, Z., Kang, Z., Yuyang, C., Chunpu, X., Shuyue, G., Haoran, Z., Xingwei, Q., Junjie, W., Ruibin, Y., Yizhi, L., Zekun, W., Yudong, L., Yu-Hsuan, T., Fengji, Z., Chenghua, L., Wenhao, H., Wenhu, C., and Jie, F. Cmmmu: A chinese massive multi-discipline multimodal understanding benchmark. arXiv preprint arXiv:2401.20847, 2024.
  25. 25.Gemini Team, G. Gemini: a family of highly capable multi-modal models. arXiv preprint arXiv:2312.11805, 2023.
  26. 26.Geng, H., Wei, S., Deng, C., Shen, B., Wang, H., and Guibas, L. Sage: Bridging semantic and actionable parts for generalizable articulated-object manipulation under language instructions. arXiv preprint arXiv:2312.01307, 2023.
  27. 27.Ghosal, D., Chia, Y. K., Majumder, N., and Poria, S. Flacuna: Unleashing the problem solving power of vicuna using flan fine-tuning, 2023.
  28. 28.Guan, T., Liu, F., Li, X. W. R. X. Z., Wang, X. L. X., Yacoob, L. C. F. H. Y., and Zhou, D. M. T. Hallusionbench: An advanced diagnostic suite for entangled language hallucination & visual illusion in large vision-language models. arXiv e-prints, pp. arXiv–2310, 2023.
  29. 29.Guo, Z., Zhang, R., Zhu, X., Tang, Y., Ma, X., Han, J., Chen, K., Gao, P., Li, X., Li, H., et al. Point-bind & point-llm: Aligning point cloud with multi-modality for 3d understanding, generation, and instruction following. arXiv preprint arXiv:2309.00615, 2023.
  30. 30.Guo, Z., Zhang, R., Chen, H., Gao, J., Gao, P., Li, H., and Heng, P.-A. Sciverse. arXiv preprint, 2024. URL https://sciverse-cuhk.github.io/.
  31. 31.Gupta, A., Dollar, P., and Girshick, R. B. Lvis: A dataset for ´ large vocabulary instance segmentation. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 5351–5359, 2019.
  32. 32.Gurari, D., Li, Q., Stangl, A., Guo, A., Lin, C., Grauman, K., Luo, J., and Bigham, J. P. Vizwiz grand challenge: Answering visual questions from blind people. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3608–3617, 2018.
  33. 33.Han, J., Zhang, R., Shao, W., Gao, P., Xu, P., Xiao, H., Zhang, K., Liu, C., Wen, S., Guo, Z., et al. Imagebind-llm: Multi-modality instruction tuning. arXiv preprint arXiv:2309.03905, 2023a.
  34. 34.Han, X., You, Q., Liu, Y., Chen, W., Zheng, H., Mrini, K., Lin, X., Wang, Y., Zhai, B., Yuan, J., Wang, H., and Yang, H. Infimm-eval: Complex open-ended reasoning evaluation for multi-modal large language models. 2023b.
  35. 35.He, H., Yao, W., Ma, K., Yu, W., Dai, Y., Zhang, H., Lan, Z., and Yu, D. Webvoyager: Building an end-to-end web agent with large multimodal models. arXiv preprint arXiv:2401.13919, 2024.
  36. 36.Hu, A., Shi, Y., Xu, H., Ye, J., Ye, Q., Yan, M., Li, C., Qian, Q., Zhang, J., and Huang, F. mplug-paperowl: Scientific diagram analysis with the multimodal large language model. arXiv preprint arXiv:2311.18248, 2023.
  37. 37.Huang, Y., Yuan, Q., Sheng, X., Yang, Z., Wu, H., Chen, P., Yang, Y., Li, L., and Lin, W. Aesbench: An expert benchmark for multimodal large language models on image aesthetics perception. arXiv preprint arXiv:2401.08276, 2024.
  38. 38.Hudson, D. A. and Manning, C. D. Gqa: A new dataset for real-world visual reasoning and compositional question answering. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 6693–6702, 2019.
  39. 39.Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al. Mistral 7b. arXiv preprint arXiv:2310.06825, 2023.
  40. 40.Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., Casas, D. d. l., Hanna, E. B., Bressand, F., et al. Mixtral of experts. arXiv preprint arXiv:2401.04088, 2024a.
  41. 41.Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., de Las Casas, D., Hanna, E. B., Bressand, F., Lengyel, G., Bour, G., Lample, G., Lavaud, L. R., Saulnier, L., Lachaux, M., Stock, P., Subramanian, S., Yang, S., Antoniak, S., Scao, T. L., Gervet, T., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E. Mixtral of experts. Arxiv 2401.04088, 2024b.
  42. 42.Jin, P., Takanobu, R., Zhang, C., Cao, X., and Yuan, L. Chat-univi: Unified visual representation empowers large language models with image and video understanding. arXiv preprint arXiv:2311.08046, 2023.
  43. 43.Johnson, J., Hariharan, B., van der Maaten, L., Fei-Fei, L., Zitnick, C. L., and Girshick, R. B. Clevr: A diagnostic dataset for compositional language and elementary visual reasoning. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 1988–1997, 2016.
  44. 44.Kafle, K., Cohen, S. D., Price, B. L., and Kanan, C. Dvqa: Understanding data visualizations via question answering. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 5648–5656, 2018. URL https://api.semanticscholar.org/CorpusID:4445015.
  45. 45.Kazemzadeh, S., Ordonez, V., andre Matten, M., and Berg, T. L. Referitgame: Referring to objects in photographs of natural scenes. In Conference on Empirical Methods in Natural Language Processing, 2014.
  46. 46.Kembhavi, A., Salvato, M., Kolve, E., Seo, M., Hajishirzi, H., and Farhadi, A. A diagram is worth a dozen images. ArXiv, abs/1603.07396, 2016. URL https://api.semanticscholar.org/CorpusID:2682274.
  47. 47.Kim, G., Hong, T., Yim, M., Park, J., Yim, J., Hwang, W., Yun, S., Han, D., and Park, S. Donut: Document understanding transformer without ocr. arXiv preprint arXiv:2111.15664, 7:15, 2021.
  48. 48.Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L.-J., Shamma, D. A., et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations. International journal of computer vision, 123:32–73, 2017.
  49. 49.Kuznetsova, A., Rom, H., Alldrin, N. G., Uijlings, J. R. R., Krasin, I., Pont-Tuset, J., Kamali, S., Popov, S., Malloci, M., Kolesnikov, A., Duerig, T., and Ferrari, V. The open images dataset v4. International Journal of Computer Vision, 128:1956 – 1981, 2018.
  50. 50.Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., and Chen, Z. Gshard: Scaling giant models with conditional computation and automatic sharding. arXiv preprint arXiv:2006.16668, 2020.
  51. 51.Li, B., Wang, R., Wang, G., Ge, Y., Ge, Y., and Shan, Y. Seed-bench: Benchmarking multimodal llms with generative comprehension. ArXiv, abs/2307.16125, 2023a.
  52. 52.Li, B., Zhang, K., Zhang, H., Guo, D., Zhang, R., Li, F., Zhang, Y., Liu, Z., and Li, C. Llava-next: Stronger llms supercharge multimodal capabilities in the wild, May 2024. URL https://llava-vl.github.io/blog/2024-05-10-llava-next-stronger-llms/.
  53. 53.Li, J., Li, D., Xiong, C., and Hoi, S. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In International Conference on Machine Learning, pp. 12888–12900. PMLR, 2022.
  54. 54.Li, J., Li, D., Savarese, S., and Hoi, S. BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., and Scarlett, J. (eds.), Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pp. 19730–19742. PMLR, 23–29 Jul 2023b. URL https://proceedings.mlr.press/v202/li23q.html.
  55. 55.Li, J., Li, D., Savarese, S., and Hoi, S. C. H. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International Conference on Machine Learning, 2023c. URL https://api.semanticscholar.org/CorpusID:256390509.
  56. 56.Li, K., He, Y., Wang, Y., Li, Y., Wang, W., Luo, P., Wang, Y., Wang, L., and Qiao, Y. Videochat: Chat-centric video understanding. arXiv preprint arXiv:2305.06355, 2023d.
  57. 57.Li, K., Wang, Y., He, Y., Li, Y., Wang, Y., Liu, Y., Wang, Z., Xu, J., Chen, G., Luo, P., et al. Mvbench: A comprehensive multi-modal video understanding benchmark. arXiv preprint arXiv:2311.17005, 2023e.
  58. 58.Li, Y., Du, Y., Zhou, K., Wang, J., Zhao, W. X., and Wen, J.-R. Evaluating object hallucination in large vision-language models. arXiv preprint arXiv:2305.10355, 2023f.
  59. 59.Lian, W., Goodson, B., Pentland, E., Cook, A., Vong, C., and ”Teknium”. Openorca: An open dataset of gpt augmented flan reasoning traces. https://https://huggingface.co/Open-Orca/OpenOrca, 2023.
  60. 60.Lin, T.-Y., Maire, M., Belongie, S. J., Hays, J., Perona, P., Ramanan, D., Dollar, P., and Zitnick, C. L. Microsoft ´ coco: Common objects in context. In European Conference on Computer Vision, 2014.
  61. 61.Lin, Z., Liu, C., Zhang, R., Gao, P., Qiu, L., Xiao, H., Qiu, H., Lin, C., Shao, W., Chen, K., et al. Sphinx: The joint mixing of weights, tasks, and visual embeddings for multi-modal large language models. arXiv preprint arXiv:2311.07575, 2023.
  62. 62.Liu, H., Li, C., Li, Y., and Lee, Y. J. Improved baselines with visual instruction tuning, 2023a.
  63. 63.Liu, H., Li, C., Wu, Q., and Lee, Y. J. Visual instruction tuning. In NeurIPS, 2023b.
  64. 64.Liu, S., Zeng, Z., Ren, T., Li, F., Zhang, H., Yang, J., yue Li, C., Yang, J., Su, H., Zhu, J.-J., and Zhang, L. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. ArXiv, abs/2303.05499, 2023c.
  65. 65.Liu, Y., Duan, H., Zhang, Y., Li, B., Zhang, S., Zhao, W., Yuan, Y., Wang, J., He, C., Liu, Z., et al. Mmbench: Is your multi-modal model an all-around player? arXiv preprint arXiv:2307.06281, 2023d.
  66. 66.Liu, Z., Mao, H., Wu, C.-Y., Feichtenhofer, C., Darrell, T., and Xie, S. A convnet for the 2020s. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 11976–11986, 2022.
  67. 67.Lu, P., Gong, R., Jiang, S., Qiu, L., Huang, S., Liang, X., and Zhu, S.-C. Inter-gps: Interpretable geometry problem solving with formal language and symbolic reasoning. In Annual Meeting of the Association for Computational Linguistics, 2021a. URL https://api.semanticscholar.org/CorpusID:234337054.
  68. 68.Lu, P., Qiu, L., Chen, J., Xia, T., Zhao, Y., Zhang, W., Yu, Z., Liang, X., and Zhu, S.-C. Iconqa: A new benchmark for abstract diagram understanding and visual language reasoning. ArXiv, abs/2110.13214, 2021b.
  69. 69.Lu, P., Mishra, S., Xia, T., Qiu, L., Chang, K.-W., Zhu, S.-C., Tafjord, O., Clark, P., and Kalyan, A. Learn to explain: Multimodal reasoning via thought chains for science question answering. ArXiv, abs/2209.09513, 2022.
  70. 70.Lu, P., Bansal, H., Xia, T., Liu, J., yue Li, C., Hajishirzi, H., Cheng, H., Chang, K.-W., Galley, M., and Gao, J. Mathvista: Evaluating math reasoning in visual contexts with gpt-4v, bard, and other large multimodal models. ArXiv, abs/2310.02255, 2023.
  71. 71.Luo, Z., Xu, C., Zhao, P., Sun, Q., Geng, X., Hu, W., Tao, C., Ma, J., Lin, Q., and Jiang, D. Wizardcoder: Empowering code large language models with evol-instruct. arXiv preprint arXiv:2306.08568, 2023.
  72. 72.Lv, T., Huang, Y., Chen, J., Cui, L., Ma, S., Chang, Y., Huang, S., Wang, W., Dong, L., Luo, W., et al. Kosmos-2.5: A multimodal literate model. arXiv preprint arXiv:2309.11419, 2023.
  73. 73.Maaz, M., Rasheed, H., Khan, S., and Khan, F. S. Video-chatgpt: Towards detailed video understanding via large vision and language models. arXiv preprint arXiv:2306.05424, 2023.
  74. 74.Mao, J., Huang, J., Toshev, A., Camburu, O.-M., Yuille, A. L., and Murphy, K. P. Generation and comprehension of unambiguous object descriptions. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 11–20, 2015.
  75. 75.Marino, K., Rastegari, M., Farhadi, A., and Mottaghi, R. Ok-vqa: A visual question answering benchmark requiring external knowledge. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 3190–3199, 2019.
  76. 76.Masry, A., Long, D., Tan, J. Q., Joty, S., and Hoque, E. ChartQA: A benchmark for question answering about charts with visual and logical reasoning. In Findings of the Association for Computational Linguistics: ACL 2022, pp. 2263–2279, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.findings-acl.177. URL https://aclanthology.org/2022.findings-acl.177.
  77. 77.Mathew, M., Bagal, V., Tito, R. P., Karatzas, D., Valveny, E., and Jawahar, C. Infographicvqa. 2022 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pp. 2582–2591, 2021a.
  78. 78.Mathew, M., Karatzas, D., and Jawahar, C. Docvqa: A dataset for vqa on document images. In Proceedings of the IEEE/CVF winter conference on applications of computer vision, pp. 2200–2209, 2021b.
  79. 79.Meng, F., Shao, W., Lu, Q., Gao, P., Zhang, K., Qiao, Y., and Luo, P. Chartassisstant: A universal chart multimodal language model via chart-to-table pre-training and multi-task instruction tuning. arXiv preprint arXiv:2401.02384, 2024.
  80. 80.Microsoft. Phi-2, 2023. URL https://huggingface.co/microsoft/phi-2.
  81. 81.Mishra, A., Shekhar, S., Singh, A. K., and Chakraborty, A. Ocr-vqa: Visual question answering by reading text in images. 2019 International Conference on Document Analysis and Recognition (ICDAR), pp. 947–952, 2019.
  82. 82.Ning, M., Zhu, B., Xie, Y., Lin, B., Cui, J., Yuan, L., Chen, D., and Yuan, L. Video-bench: A comprehensive benchmark and toolkit for evaluating video-based large language models. arXiv preprint arXiv:2311.16103, 2023.
  83. 83.OpenAI. GPT-4V(ision) system card, 2023. URL https://openai.com/research/gpt-4v-system-card.
  84. 84.Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., et al. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193, 2023.
  85. 85.Ordonez, V., Kulkarni, G., and Berg, T. Im2text: Describing images using 1 million captioned photographs. In Shawe-Taylor, J., Zemel, R., Bartlett, P., Pereira, F., and Weinberger, K. (eds.), Advances in Neural Information Processing Systems, volume 24. Curran Associates, Inc., 2011. URL https://proceedings.neurips.cc/paper_files/paper/2011/file/5dd9db5e033da9c6fb5ba83c7a7ebea9-Paper.pdf.
  86. 86.Pasupat, P. and Liang, P. Compositional semantic parsing on semi-structured tables. In Annual Meeting of the Association for Computational Linguistics, 2015. URL https://api.semanticscholar.org/CorpusID:9027681.
  87. 87.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  88. 88.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, 2021. URL https://api.semanticscholar.org/CorpusID:231591445.
  89. 89.Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y. Zero: Memory optimizations toward training trillion parameter models. In SC20: International Conference for High Performance Computing, Networking, Storage and Analysis, pp. 1–16. IEEE, 2020.
  90. 90.Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M. S., Berg, A. C., and Fei-Fei, L. Imagenet large scale visual recognition challenge. International Journal of Computer Vision, 115:211 – 252, 2014.
  91. 91.Schuhmann, C., Vencu, R., Beaumont, R., Kaczmarczyk, R., Mullis, C., Katta, A., Coombes, T., Jitsev, J., and Komatsuzaki, A. Laion-400m: Open dataset of clip-filtered 400 million image-text pairs. arXiv preprint arXiv:2111.02114, 2021.
  92. 92.Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., et al. Laion-5b: An open large-scale dataset for training next generation image-text models. Advances in Neural Information Processing Systems, 35:25278–25294, 2022.
  93. 93.Shao, S., Li, Z., Zhang, T., Peng, C., Yu, G., Zhang, X., Li, J., and Sun, J. Objects365: A large-scale, high-quality dataset for object detection. 2019 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 8429–8438, 2019.
  94. 94.Shao, W., Hu, Y., Gao, P., Lei, M., Zhang, K., Meng, F., Xu, P., Huang, S., Li, H., Qiao, Y., et al. Tiny lvlm-ehub: Early multimodal experiments with bard. arXiv preprint arXiv:2308.03729, 2023.
  95. 95.Sharma, P., Ding, N., Goodman, S., and Soricut, R. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 2556–2565, 2018.
  96. 96.Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. arXiv preprint arXiv:1701.06538, 2017.
  97. 97.Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B. Megatron-lm: Training multi-billion parameter language models using model parallelism. arXiv preprint arXiv:1909.08053, 2019.
  98. 98.Singh, A., Natarajan, V., Shah, M., Jiang, Y., Chen, X., Batra, D., Parikh, D., and Rohrbach, M. Towards vqa models that can read. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 8309–8318, 2019.
  99. 99.Stanislawek, T., Grali’nski, F., Wr’oblewska, A., Lipi’nski, D., Kaliska, A., Rosalska, P., Topolski, B., and Biecek, P. Kleister: Key information extraction datasets involving long documents with complex layouts. In IEEE International Conference on Document Analysis and Recognition, 2021.
  100. 100.Su, Y., Lan, T., Li, H., Xu, J., Wang, Y., and Cai, D. Pandagpt: One model to instruction-follow them all. arXiv preprint arXiv:2305.16355, 2023.
  101. 101.Svetlichnaya, S. Deepform, 2020. URL https://wandb.ai/stacey/deepform_v1/reports/DeepForm-Understand-/Structured-Documents-at-Scale.
  102. 102.Tanaka, R., Nishida, K., and Yoshida, S. Visualmrc: Machine reading comprehension on document images. ArXiv, abs/2101.11272, 2021.
  103. 103.Team, I. Internlm: A multilingual language model with progressively enhanced capabilities, 2023.
  104. 104.Tong, S., Liu, Z., Zhai, Y., Ma, Y., LeCun, Y., and Xie, S. Eyes wide shut? exploring the visual shortcomings of multimodal llms. Arxiv 2401.06209, 2024.
  105. 105.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Roziere, B., Goyal, N., Hambro, E., Azhar, F., et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023a.
  106. 106.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023b.
  107. 107.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. In Advances in neural information processing systems, 2017.
  108. 108.Wang, J., Meng, L., Weng, Z., He, B., Wu, Z., and Jiang, Y.-G. To see is to believe: Prompting gpt-4v for better visual instruction tuning. ArXiv, abs/2311.07574, 2023a. URL https://api.semanticscholar.org/CorpusID:265150580.
  109. 109.Wang, W., Chen, Z., Chen, X., Wu, J., Zhu, X., Zeng, G., Luo, P., Lu, T., Zhou, J., Qiao, Y., et al. Visionllm: Large language model is also an open-ended decoder for vision-centric tasks. arXiv preprint arXiv:2305.11175, 2023b.
  110. 110.Wen, L., Fu, D., Li, X., Cai, X., Ma, T., Cai, P., Dou, M., Shi, B., He, L., and Qiao, Y. Dilu: A knowledge-driven approach to autonomous driving with large language models. arXiv preprint arXiv:2309.16292, 2023.
  111. 111.Workshop, B., Scao, T. L., Fan, A., Akiki, C., Pavlick, E., Ilic, S., Hesslow, D., Castagn ´ e, R., Luccioni, A. S., Yvon, ´ F., et al. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100, 2022.
  112. 112.Yang, A., Xiao, B., Wang, B., Zhang, B., Bian, C., Yin, C., Lv, C., Pan, D., Wang, D., Yan, D., et al. Baichuan 2: Open large-scale language models. arXiv preprint arXiv:2309.10305, 2023a.
  113. 113.Yang, J., Zeng, A., Zhang, R., and Zhang, L. Unipose: Detecting any keypoints. ArXiv, abs/2310.08530, 2023b.
  114. 114.Yang, J., Zhang, H., Li, F., Zou, X., Li, C., and Gao, J. Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v. arXiv preprint arXiv:2310.11441, 2023c.
  115. 115.Yang, S., Liu, J., Zhang, R., Pan, M., Guo, Z., Li, X., Chen, Z., Gao, P., Guo, Y., and Zhang, S. Lidar-llm: Exploring the potential of large language models for 3d lidar understanding. arXiv preprint arXiv:2312.14074, 2023d.
  116. 116.Yang, Z., Li, L., Lin, K., Wang, J., Lin, C.-C., Liu, Z., and Wang, L. The dawn of lmms: Preliminary explorations with gpt-4v (ision). arXiv preprint arXiv:2309.17421, 9 (1):1, 2023e.
  117. 117.Yang, Z., Liu, J., Han, Y., Chen, X., Huang, Z., Fu, B., and Yu, G. Appagent: Multimodal agents as smartphone users. arXiv preprint arXiv:2312.13771, 2023f.
  118. 118.Ye, J., Hu, A., Xu, H., Ye, Q., Yan, M., Dan, Y., Zhao, C., Xu, G., Li, C., Tian, J., et al. mplug-docowl: Modularized multimodal large language model for document understanding. arXiv preprint arXiv:2307.02499, 2023a.
  119. 119.Ye, Q., Xu, H., Xu, G., Ye, J., Yan, M., Zhou, Y., Wang, J., Hu, A., Shi, P., Shi, Y., Jiang, C., Li, C., Xu, Y., Chen, H., Tian, J., Qian, Q., Zhang, J., and Huang, F. mplug-owl: Modularization empowers large language models with multimodality, 2023b.
  120. 120.Ye, Q., Xu, H., Ye, J., Yan, M., Hu, A., Liu, H., Qian, Q., Zhang, J., Huang, F., and Zhou, J. mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration, 2023c.
  121. 121.Yim, M., Kim, Y., Cho, H.-C., and Park, S. Synthtiger: Synthetic text image generator towards better text recognition models. In International Conference on Document Analysis and Recognition, pp. 109–124. Springer, 2021.
  122. 122.Yu, L., Jiang, W., Shi, H., Yu, J., Liu, Z., Zhang, Y., Kwok, J. T., Li, Z., Weller, A., and Liu, W. Metamath: Bootstrap your own mathematical questions for large language models. arXiv preprint arXiv:2309.12284, 2023a.
  123. 123.Yu, W., Yang, Z., Li, L., Wang, J., Lin, K., Liu, Z., Wang, X., and Wang, L. Mm-vet: Evaluating large multimodal models for integrated capabilities. ArXiv, abs/2308.02490, 2023b.
  124. 124.Yue, X., Ni, Y., Zhang, K., Zheng, T., Liu, R., Zhang, G., Stevens, S., Jiang, D., Ren, W., Sun, Y., Wei, C., Yu, B., Yuan, R., Sun, R., Yin, M., Zheng, B., Yang, Z., Liu, Y., Huang, W., Sun, H., Su, Y., and Chen, W. Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi. arXiv preprint arXiv:2311.16502, 2023a.
  125. 125.Yue, X., Qu, X., Zhang, G., Fu, Y., Huang, W., Sun, H., Su, Y., and Chen, W. Mammoth: Building math generalist models through hybrid instruction tuning. arXiv preprint arXiv:2309.05653, 2023b.
  126. 126.Zhang, H., Li, X., and Bing, L. Video-llama: An instruction-tuned audio-visual language model for video understanding. arXiv preprint arXiv:2306.02858, 2023a.
  127. 127.Zhang, P., Zeng, G., Wang, T., and Lu, W. Tinyllama: An open-source small language model. ArXiv, abs/2401.02385, 2024a. URL https://api.semanticscholar.org/CorpusID:266755802.
  128. 128.Zhang, P., Zeng, G., Wang, T., and Lu, W. Tinyllama: An open-source small language model. arXiv preprint arXiv:2401.02385, 2024b.
  129. 129.Zhang, R., Guo, Z., Zhang, W., Li, K., Miao, X., Cui, B., Qiao, Y., Gao, P., and Li, H. Pointclip: Point cloud understanding by clip. In CVPR 2022, 2022a.
  130. 130.Zhang, R., Hu, X., Li, B., Huang, S., Deng, H., Li, H., Qiao, Y., and Gao, P. Prompt, generate, then cache: Cascade of foundation models makes strong few-shot learners. CVPR 2023, 2023b.
  131. 131.Zhang, R., Han, J., Zhou, A., Hu, X., Yan, S., Lu, P., Li, H., Gao, P., and Qiao, Y. LLaMA-adapter: Efficient fine-tuning of large language models with zero-initialized attention. In The Twelfth International Conference on Learning Representations, 2024c. URL https://openreview.net/forum?id=d4UiXAHN2W.
  132. 132.Zhang, R., Jiang, D., Zhang, Y., Lin, H., Guo, Z., Qiu, P., Zhou, A., Lu, P., Chang, K.-W., Gao, P., et al. Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? arXiv preprint arXiv:2403.14624, 2024d.
  133. 133.Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068, 2022b.
  134. 134.Zhang, Y., Zhang, R., Gu, J., Zhou, Y., Lipka, N., Yang, D., and Sun, T. Llavar: Enhanced visual instruction tuning for text-rich image understanding. ArXiv, abs/2306.17107, 2023c. URL https://api.semanticscholar.org/CorpusID:259287523.
  135. 135.Zhu, D., Chen, J., Shen, X., Li, X., and Elhoseiny, M. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592, 2023a.
  136. 136.Zhu, X., Zhang, R., He, B., Guo, Z., Zeng, Z., Qin, Z., Zhang, S., and Gao, P. Pointclip v2: Prompting clip and gpt for powerful 3d open-world learning. ICCV 2023, 2023b.

Citation

MLA
Liu, D., et al. “SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models”. arXiv, 2024, http://arxiv.org/abs/2402.05935v3.
APA
Liu, D., Zhang, R., Qiu, L., Huang, S., Lin, W., Zhao, S., Geng, S., Lin, Z., Jin, P., Zhang, K., Shao, W., Xu, C., He, C., He, J., Shao, H., Lu, P., Li, H., Qiao, Y., & Gao, P. (2024). SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models. arXiv. http://arxiv.org/abs/2402.05935v3
Chicago
Liu, D., R. Zhang, L. Qiu, et al. 2024. “SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models”. arXiv. http://arxiv.org/abs/2402.05935v3.
Harvard
Liu, D. et al. (2024) “SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.05935v3.
Vancouver
1. Liu D, Zhang R, Qiu L, et al (2024) SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models. arXiv

BibTeX

@article{liu2024sphinx,
  title = {SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models},
  author = {Liu, Dongyang and Zhang, Renrui and Qiu, Longtian and Huang, Siyuan and Lin, Weifeng and Zhao, Shitian and Geng, Shijie and Lin, Ziyi and Jin, Peng and Zhang, Kaipeng and Shao, Wenqi and Xu, Chao and He, Conghui and He, Junjun and Shao, Hao and Lu, Pan and Li, Hongsheng and Qiao, Yu and Gao, Peng},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.05935v3},
  eprint = {2402.05935}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/