Language Is Not All You Need: Aligning Perception with Language Models
Shaohan HuangLi DongWenhui WangYaru HaoSaksham SinghalShuming MaTengchao LvLei CuiOwais Khan MohammedBarun Patra
Presents Kosmos-1, a multimodal language model trained from scratch on interleaved text and image data that aligns visual perception with language generation to perform zero-shot multimodal dialogue, OCR-free reading, and nonverbal reasoning without task-specific fine-tuning.
Large language models excel at processing pure text, but their inability to natively perceive other modalities such as vision limits their ability to ground knowledge in the physical world and perform visually driven tasks. Bridging the gap between language processing and visual perception is critical for expanding artificial intelligence into high-value domains, including document intelligence, robotics, and graphical user interface interaction.
The article demonstrates and evaluates KOSMOS-1, a multimodal large language model designed to natively perceive multimodal inputs, learn in context, and follow instructions across language, vision, and cross-modal tasks without task-specific gradient fine-tuning.
The researchers built a 1.6-billion-parameter model using a Transformer-based causal language model backbone integrated with a visual encoder. They trained the system from scratch on web-scale datasets comprising approximately 360 billion tokens across standard text corpora, image-caption pairs, and roughly 71 million interleaved image-text documents. To align instruction-following behavior across modalities, the team applied language-only instruction tuning using curated datasets. The model was evaluated across various zero-shot and few-shot benchmarks covering language understanding, image captioning, visual question answering, optical character recognition (OCR)-free reading, web page comprehension, zero-shot image classification, and abstract nonverbal reasoning.
The evaluation produced several significant findings. First, KOSMOS-1 demonstrated strong perception capabilities, achieving zero-shot image captioning CIDEr scores of 84.7 on COCO and 67.1 on Flickr30k, outperforming larger baseline models such as Flamingo-3B and Flamingo-9B. Second, the model exhibited robust cross-modal transfer: instruction tuning performed solely on language data boosted visual question answering accuracy by 4.3 percentage points on VQAv2, and multimodal pretraining improved visual commonsense reasoning over language-only models by 14.7 percentage points on object color prediction. Third, when provided with verbal category descriptions in context, zero-shot fine-grained image classification accuracy surged from 61.7% to 90.0%. Fourth, the model achieved OCR-free language understanding directly from images, attaining 67.1% accuracy on rendered sentiment text (improving to 72.9% via multimodal chain-of-thought prompting) without external OCR tools. Finally, on a newly introduced 50-example Raven Progressive Matrices benchmark evaluating nonverbal reasoning, the model scored 22% to 26% accuracy, exceeding the 17% random baseline.
These findings indicate that treating language models as general-purpose interfaces capable of directly ingesting visual tokens enables native multimodal interaction without sacrificing core language capabilities. Multimodal pretraining allows models to acquire richer commonsense knowledge about physical objects that text alone cannot provide, while structural comprehension unlocks practical capabilities like extracting information directly from rendered web pages or images.
Organizations developing or deploying multimodal systems should explore aligning perception directly within language model decoders rather than relying on complex pipelines of disjoint OCR and vision tools. For future system enhancements, the article recommends scaling the model size, incorporating additional sensory modalities such as speech, and expanding the architecture to support multimodal generation tasks such as instruction-guided image creation.
While the model demonstrates promising zero-shot capabilities, its nonverbal reasoning score on Raven's matrices remains substantially below human adult levels, and evaluation was performed on a relatively compact 1.6-billion-parameter architecture at a standard 224x224 image resolution. Stakeholders should view these results with high confidence regarding the feasibility of cross-modal transfer and unified architectures, but exercise caution before deploying such models in highly complex visual reasoning workflows without further scaling and validation.
- Paper: SimVLM: Simple Visual Language Model Pretraining with Weak Supervision, Zirui Wang et al. (2021). SimVLM’s single generative objective over image and text offers a useful precursor to KOSMOS-1’s causal language-model approach to multimodal inputs.
- Paper: PaLI: A Jointly-Scaled Multilingual Language-Image Model, Xi Chen et al. (2022). PaLI establishes the jointly scaled vision-language model design that helps contextualize KOSMOS-1’s integration of a visual encoder with a language model.
- Paper: LXMERT: Learning Cross-Modality Encoder Representations from Transformers, Hao Tan et al. (2019). LXMERT provides an earlier account of pretraining shared vision-language representations for tasks such as visual question answering, clarifying the cross-modal learning problem KOSMOS-1 addresses.
- Paper: Kosmos-2: Grounding Multimodal Large Language Models to the World, Zhiliang Peng et al. (2023). KOSMOS-2 carries KOSMOS-1’s multimodal language-model approach forward by grounding generated text in image regions and enabling spatially precise interaction.
