SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models
Dongyang LiuRenrui ZhangLongtian QiuSiyuan HuangWeifeng LinShitian ZhaoShijie GengZiyi LinPeng JinKaipeng Zhang
Presents an efficient multimodal architecture and a unified one-stage training framework across multiple base language models ranging from 1.1B to 8×7B parameters, demonstrating how systematic scaling of both training data and model capacity boosts multimodal performance across diverse visual-language tasks.
Recent advancements in artificial intelligence have accelerated the development of multimodal systems capable of interpreting both visual imagery and natural language. However, many current open-source models remain constrained by narrow training datasets that falter in specialized domains—such as reading text-dense documents, interpreting charts, and solving visual mathematics—and by limited model size options that are either too large for mobile devices or too small for complex reasoning. The article addresses these challenges by developing SPHINX-X, an adaptable family of multimodal large language models designed to scale training data coverage and parameter sizes efficiently while streamlining model architecture.
To achieve this, the authors restructured their prior framework into a more efficient, single-stage training pipeline. They reduced visual processing overhead by pairing two complementary vision encoders, eliminated redundant processing of empty image padding with learnable skip tokens, and trained the entire system at once. The training utilized a broad collection of multimodal data, including millions of conversational samples, visual detection and pose tasks, an in-house document dataset of three million text-dense pages, and detailed visual marking annotations. The authors applied this unified training across four base language models ranging from a compact 1.1-billion parameter model to a high-capacity mixture-of-experts architecture activating subsets of an eight-by-seven-billion parameter base.
Benchmarking revealed several key outcomes. First, scaling both model size and dataset diversity produced consistent performance improvements across diverse tasks. The expanded 13-billion parameter version systematically outperformed its predecessor across standard visual and text benchmarks. Second, the mixture-of-experts configuration achieved leading results among open-source models on mathematical and scientific reasoning, matching or exceeding proprietary models such as GPT-4V on specialized visual perception and user interface localization tests. Third, despite being trained solely on still images, the models outperformed dedicated video architectures on video question-answering benchmarks when evaluated on sampled video frames. Finally, the compact 1.1-billion parameter model preserved viable multimodal capabilities, demonstrating feasibility for edge and mobile environments.
These findings indicate that consolidating complex, multi-stage training into a single comprehensive workflow reduces engineering friction and computational waste without sacrificing capability. For organizations evaluating artificial intelligence integration, the results show that generalist models trained on diverse domain data can effectively handle specialized tasks such as optical character recognition, document layout analysis, and interface grounding without requiring separate, dedicated models for each use case.
Decision-makers can evaluate a tiered adoption strategy, deploying compact models for on-device applications and larger mixture-of-experts models for complex enterprise reasoning. However, users should note key limitations: the models exhibited weaker results on multidisciplinary academic tests due to gaps in college-level subject training data, and video tracking across time remains limited without dedicated temporal fine-tuning. Future efforts should focus on integrating multi-discipline educational datasets and optimizing expert pruning to enhance deployment efficiency.
- Paper: Llama 2: Open Foundation and Fine-Tuned Chat Models, Hugo Touvron et al. (2023). SPHINX-X trains variants on LLaMA 2, so reading its base-model report first clarifies the language-model foundation the series adapts.
- Paper: Improved Baselines with Visual Instruction Tuning, Haotian Liu et al. (2024). This LLaVA study establishes visual instruction-tuning design choices that help frame SPHINX-X’s changes to multimodal architecture and training.
No sufficiently relevant recommendations were found.
