Language Is Not All You Need: Aligning Perception with Language Models

Shaohan HuangLi DongWenhui WangYaru HaoSaksham SinghalShuming MaTengchao LvLei CuiOwais Khan MohammedBarun Patra

article2023NeurIPS733 citations

Presents Kosmos-1, a multimodal language model trained from scratch on interleaved text and image data that aligns visual perception with language generation to perform zero-shot multimodal dialogue, OCR-free reading, and nonverbal reasoning without task-specific fine-tuning.

Listen

Large language models excel at processing pure text, but their inability to natively perceive other modalities such as vision limits their ability to ground knowledge in the physical world and perform visually driven tasks. Bridging the gap between language processing and visual perception is critical for expanding artificial intelligence into high-value domains, including document intelligence, robotics, and graphical user interface interaction.

The article demonstrates and evaluates KOSMOS-1, a multimodal large language model designed to natively perceive multimodal inputs, learn in context, and follow instructions across language, vision, and cross-modal tasks without task-specific gradient fine-tuning.

The researchers built a 1.6-billion-parameter model using a Transformer-based causal language model backbone integrated with a visual encoder. They trained the system from scratch on web-scale datasets comprising approximately 360 billion tokens across standard text corpora, image-caption pairs, and roughly 71 million interleaved image-text documents. To align instruction-following behavior across modalities, the team applied language-only instruction tuning using curated datasets. The model was evaluated across various zero-shot and few-shot benchmarks covering language understanding, image captioning, visual question answering, optical character recognition (OCR)-free reading, web page comprehension, zero-shot image classification, and abstract nonverbal reasoning.

The evaluation produced several significant findings. First, KOSMOS-1 demonstrated strong perception capabilities, achieving zero-shot image captioning CIDEr scores of 84.7 on COCO and 67.1 on Flickr30k, outperforming larger baseline models such as Flamingo-3B and Flamingo-9B. Second, the model exhibited robust cross-modal transfer: instruction tuning performed solely on language data boosted visual question answering accuracy by 4.3 percentage points on VQAv2, and multimodal pretraining improved visual commonsense reasoning over language-only models by 14.7 percentage points on object color prediction. Third, when provided with verbal category descriptions in context, zero-shot fine-grained image classification accuracy surged from 61.7% to 90.0%. Fourth, the model achieved OCR-free language understanding directly from images, attaining 67.1% accuracy on rendered sentiment text (improving to 72.9% via multimodal chain-of-thought prompting) without external OCR tools. Finally, on a newly introduced 50-example Raven Progressive Matrices benchmark evaluating nonverbal reasoning, the model scored 22% to 26% accuracy, exceeding the 17% random baseline.

These findings indicate that treating language models as general-purpose interfaces capable of directly ingesting visual tokens enables native multimodal interaction without sacrificing core language capabilities. Multimodal pretraining allows models to acquire richer commonsense knowledge about physical objects that text alone cannot provide, while structural comprehension unlocks practical capabilities like extracting information directly from rendered web pages or images.

Organizations developing or deploying multimodal systems should explore aligning perception directly within language model decoders rather than relying on complex pipelines of disjoint OCR and vision tools. For future system enhancements, the article recommends scaling the model size, incorporating additional sensory modalities such as speech, and expanding the architecture to support multimodal generation tasks such as instruction-guided image creation.

While the model demonstrates promising zero-shot capabilities, its nonverbal reasoning score on Raven's matrices remains substantially below human adult levels, and evaluation was performed on a relatively compact 1.6-billion-parameter architecture at a standard 224x224 image resolution. Stakeholders should view these results with high confidence regarding the feasibility of cross-modal transfer and unified architectures, but exercise caution before deploying such models in highly complex visual reasoning workflows without further scaling and validation.

Cover for Language Is Not All You Need: Aligning Perception with Language Models

Abstract

A big convergence of language, multimodal perception, action, and world modeling is a key step toward artificial general intelligence. In this work, we introduce Kosmos-1, a Multimodal Large Language Model (MLLM) that can perceive general modalities, learn in context (i.e., few-shot), and follow instructions (i.e., zero-shot). Specifically, we train Kosmos-1 from scratch on web-scale multimodal corpora, including arbitrarily interleaved text and images, image-caption pairs, and text data. We evaluate various settings, including zero-shot, few-shot, and multimodal chain-of-thought prompting, on a wide range of tasks without any gradient updates or finetuning. Experimental results show that Kosmos-1 achieves impressive performance on (i) language understanding, generation, and even OCR-free NLP (directly fed with document images), (ii) perception-language tasks, including multimodal dialogue, image captioning, visual question answering, and (iii) vision tasks, such as image recognition with descriptions (specifying classification via text instructions). We also show that MLLMs can benefit from cross-modal transfer, i.e., transfer knowledge from language to multimodal, and from multimodal to language. In addition, we introduce a dataset of Raven IQ test, which diagnoses the nonverbal reasoning capability of MLLMs.

Table of Contents

  • 1 Introduction: From LLMs to MLLMs
  • 2 Kosmos-1: A Multimodal Large Language Model
  • 2.1 Input Representation
  • 2.2 Multimodal Large Language Models (MLLMs)
  • 2.3 Training Objective
  • 3 Model Training
  • 3.1 Multimodal Training Data
  • 3.2 Training Setup
  • 3.3 Language-Only Instruction Tuning
  • 4 Evaluation
  • 4.1 Perception-Language Tasks
  • 4.1.1 Evaluation Setup
  • 4.1.2 Results
  • 4.2 IQ Test: Nonverbal Reasoning
  • 4.2.1 Evaluation Setup
  • 4.2.2 Results
  • 4.3 OCR-Free Language Understanding
  • 4.3.1 Evaluation Setup
  • 4.3.2 Results
  • 4.4 Web Page Question Answering
  • 4.4.1 Evaluation Setup
  • 4.4.2 Results
  • 4.5 Multimodal Chain-of-Thought Prompting
  • 4.5.1 Evaluation Setup
  • 4.5.2 Results
  • 4.6 Zero-Shot Image Classification
  • 4.6.1 Evaluation Setup
  • 4.6.2 Results
  • 4.7 Zero-Shot Image Classification with Descriptions
  • 4.7.1 Evaluation Setup
  • 4.7.2 Results
  • 4.8 Language Tasks
  • 4.8.1 Evaluation Setup
  • 4.8.2 Results
  • 4.9 Cross-modal Transfer
  • 4.9.1 Transfer from Language to Multimodal: Language-Only Instruction Tuning
  • 4.9.2 Transfer from Multimodal to Language: Visual Commonsense Reasoning
  • 5 Conclusion
  • References
  • A Hyperparameters
  • A.1 Training
  • A.2 Language-Only Instruction Tuning
  • B Datasets
  • B.1 Pretraning
  • B.1.1 Text Corpora
  • B.1.2 Image-Caption Pairs
  • B.1.3 Interleaved Data
  • B.2 Data Format
  • C Evaluation
  • C.1 Input Format Used for Perception-Language Tasks
  • C.2 Language Tasks
  • C.3 WebSRC Task Examples

Knowls

  1. Knowl 1 — KOSMOS-1 represents multimodal inputs in a causal language-model sequence

    model/method

    KOSMOS-1 uses a Transformer decoder as a shared interface for text and perception. Text tokens are mapped to embeddings, while images are encoded by a pretrained CLIP ViT-L/14 vision encoder and reduced to fewer embeddings with an attentive Resampler. The resulting vectors are flattened into one sequence, with special tokens such as <s>, </s>, <image>, and </image> marking sequence and image boundaries; text and image content can be interleaved. A left-to-right causal decoder processes this sequence and predicts subsequent text tokens. The decoder uses the MAGNETO Transformer variant and XPOS relative position encoding.

  2. Knowl 2 — Multimodal next-token prediction trains KOSMOS-1 across paired and interleaved data

    model/method

    KOSMOS-1 is trained with an autoregressive next-token objective on monomodal text, image-caption pairs, and documents containing interleaved text and images. The objective is to maximize the likelihood of the next token given preceding context. Although image embeddings condition the model, only discrete tokens, such as text tokens, contribute to the training loss. Text data trains language capabilities, while paired and interleaved image-text data align visual perception with language; interleaved documents also provide multimodal language-modeling examples.

  3. Knowl 3 — KOSMOS-1 combines web text, image-caption pairs, and filtered interleaved web pages

    data/table

    The training corpus combines three data forms to support language modeling and image-language alignment. Text sources include The Pile and Common Crawl snapshots, including CC-2020-50 and CC-2021-04, as well as CC-Stories and RealNews; some Pile sources, including GitHub, arXiv, Stack Exchange, and PubMed Central, were excluded. Image-caption data comes from English LAION-2B, LAION-400M, COYO-700M, and Conceptual Captions. For interleaved data, filtering an initial collection of about 2 billion Common Crawl pages yielded about 71 million English web pages with meaningful text and images. Pages without interspersed images were removed, as were small or single-colored images and incoherent or spam-like text. Documents were limited to five images, and half of the documents containing only one image were randomly discarded.

  4. Knowl 4 — KOSMOS-1 uses a 1.3-billion-parameter decoder and is trained for 300,000 steps

    experimental setup

    The KOSMOS-1 decoder has 24 layers, hidden size 2,048, feed-forward intermediate size 8,192, and 32 attention heads, for about 1.3 billion decoder parameters and about 1.6 billion parameters in the full model. The image encoder is CLIP ViT-L/14 with 1,024-dimensional features; its parameters were frozen except for the final layer, and training images were resized to 224×224224\times224. Training lasted 300,000 steps, or approximately 360 billion tokens, with a batch of 1.2 million tokens apportioned as 0.5 million from text, 0.5 million from image-caption pairs, and 0.2 million from interleaved data. The model used AdamW with β=(0.9,0.98)\beta=(0.9,0.98), weight decay 0.01, dropout 0.1, a peak learning rate of 2×10−42\times10^{-4}, 375 warmup steps, and linear decay thereafter. Magneto initialization was used for optimization stability.

  5. Knowl 5 — Language-only instruction tuning transfers instruction-following gains to visual tasks

    empirical result

    KOSMOS-1 received language-only instruction tuning using Unnatural Instructions (68,478 core instruction-input-output examples) and 54,000 randomly selected FLANv2 examples, mixed with the training corpora. Tuning lasted 10,000 steps with learning rate 2×10−52\times10^{-5}; the loss was applied to the output, not the instruction or input. In the reported ablation, tuning improved Flickr30k CIDEr by 1.9 points, VQAv2 accuracy by 4.3 points, and VizWiz accuracy by 1.3 points. The COCO CIDEr entries went in the opposite direction: 84.7 with tuning versus 87.6 without it, despite the paper's accompanying general claim of improvements. These results indicate transfer of instruction-following behavior from language training to some vision-language evaluations, rather than uniform gains on every reported metric.

  6. Knowl 6 — KOSMOS-1 performs zero-shot and few-shot image captioning and visual question answering

    empirical result

    KOSMOS-1 was evaluated at 224×224224\times224 image resolution on COCO and Flickr30k captioning and on VQAv2 and VizWiz visual question answering. Captioning used beam search (beam size 5) and CIDEr/SPICE metrics; question answering used greedy decoding and VQA accuracy. KOSMOS-1 achieved 84.7 CIDEr and 16.8 SPICE on COCO, and 67.1 CIDEr and 14.5 SPICE on Flickr30k in zero-shot captioning. Its Flickr30k CIDEr exceeded Flamingo-3B's 60.6 and Flamingo-9B's 61.5; those Flamingo comparisons used two text-only demonstrations, whereas KOSMOS-1 used no examples. Zero-shot VQA accuracy was 51.0 on VQAv2 and 29.2 on VizWiz; Flamingo-3B/9B scored 49.2/51.8 and 28.9/28.8, respectively, with two text-only examples in their prompts.

    In few-shot evaluation, KOSMOS-1 captioning CIDEr scores for k=2,4,8k=2,4,8 were 99.6, 101.7, and 96.7 on COCO and 70.0, 75.3, and 68.0 on Flickr30k. Few-shot VQA accuracies for k=2,4,8k=2,4,8 were 51.4, 51.8, and 51.4 on VQAv2 and 31.4, 35.3, and 39.0 on VizWiz. Thus, more examples did not produce a monotonic improvement in every setting, while VizWiz VQA improved at each tested shot count.

  7. Knowl 7 — KOSMOS-1 transfers visual commonsense knowledge to text-only reasoning

    empirical result

    To test transfer from multimodal training to language-only reasoning, KOSMOS-1 and a language-model baseline trained on the same text data were evaluated using text inputs only. RELATIVESIZE asks binary comparisons about the relative sizes of object pairs; MEMORYCOLOR and COLORTERMS ask for object colors from 11 labels. KOSMOS-1 obtained 94.2%, 76.1%, and 73.1% accuracy on these datasets, compared with 92.7%, 61.4%, and 63.4% for the language-only baseline. The respective gains were 1.5, 14.7, and 9.7 percentage points. The paper interprets this pattern as evidence that visual knowledge acquired during multimodal training can support language-only commonsense judgments.

  8. Knowl 8 — KOSMOS-1 shows above-chance zero-shot performance on a 50-item Raven matrices set

    experimental setup

    The authors assembled 50 Raven Progressive Matrices examples from websites. Each problem presents three, four, or eight matrix images and six candidate completions, with one correct candidate. For each candidate, the images are flattened into the prompt and paired with textual task cues; the model is asked whether the candidate is correct. The candidate with the highest probability of producing “Yes” is selected. Accuracy was 17% for random choice, 22% for KOSMOS-1, and 26% for KOSMOS-1 without language-only instruction tuning. The results are above the random baseline, but the dataset is small, and the authors note a substantial gap from average adult performance.

  9. Knowl 9 — KOSMOS-1 reads and classifies text in images without an external OCR system

    empirical result

    KOSMOS-1 was evaluated without external OCR on rendered sentiment text (Rendered SST-2) and multimodal hate-speech memes (HatefulMemes). It achieved 67.1% accuracy on the Rendered SST-2 test set and 63.9% ROC AUC on the HatefulMemes validation set. On Rendered SST-2, CLIP ViT-L/14 scored 64.0% accuracy. On HatefulMemes, CLIP ViT-L/14 scored 63.3% ROC AUC and Flamingo-9B scored 57.0%; Flamingo's prompt explicitly included OCR text, while KOSMOS-1 used no such external text. The results support the model's ability to recognize and use text rendered within images.

  10. Knowl 10 — Multimodal chain-of-thought prompting improves OCR-free sentiment classification

    model/method

    For multimodal chain-of-thought prompting, KOSMOS-1 first receives an image and the instruction “Introduce this picture in detail:” to generate a textual rationale describing its contents. The generated rationale is then included before a sentiment question, and the model predicts “positive” or “negative.” On Rendered SST-2, this two-stage prompting achieved 72.9% accuracy, compared with 67.1% for standard prompting, a 5.8-point increase. The authors attribute the gain to generating intermediate text that helps the model recognize image text and infer sentence sentiment.

  11. Knowl 11 — Natural-language descriptions improve zero-shot visual category classification

    empirical result

    KOSMOS-1 was tested on two forms of zero-shot image classification using natural-language category names or descriptions. On ImageNet, it generated category names with top-1 accuracy of 4.0% without constraints and 38.1% when decoding was restricted to the 1,000 ImageNet labels; the corresponding GIT scores were 1.9% and 33.5%. In a separate three-group binary bird-classification set, each category had a verbal description and 20 images. Classification accuracy was 90.0% when the two competing categories were described in the prompt, compared with 61.7% when the prompt supplied category names without descriptions. The description-based setup lets users specify distinctions in language rather than relying only on a fixed label list.

  12. Knowl 12 — KOSMOS-1 is competitive with a matched language model on language-only benchmarks

    empirical result

    The authors compared KOSMOS-1 with a language-only model trained on the same text corpora and training setup, without instruction tuning in either model. They evaluated eight language tasks under zero-shot, one-shot, and four-shot prompting; the entries below are the reported task scores, with averages across all eight tasks.

    Task LLM 0-shot KOSMOS-1 0-shot LLM 1-shot KOSMOS-1 1-shot LLM 4-shot KOSMOS-1 4-shot
    StoryCloze 72.9 72.1 72.9 72.2 73.1 72.3
    HellaSwag 50.4 50.0 50.2 50.0 50.4 50.3
    Winograd 71.6 69.8 71.2 68.4 70.9 69.8
    Winogrande 56.7 54.8 56.7 54.5 57.0 55.7
    PIQA 73.2 72.9 73.0 72.5 72.6 72.3
    BoolQ 56.4 56.4 55.1 57.2 58.7 59.2
    CB 39.3 44.6 41.1 48.2 42.9 53.6
    COPA 68.0 63.0 69.0 64.0 69.0 64.0
    Average 61.1 60.5 61.2 60.9 61.8 62.2

    The language-only model had a higher average in zero-shot and one-shot evaluation, while KOSMOS-1 had the higher average with four demonstrations. KOSMOS-1's especially higher CB results contribute to its four-shot average advantage; it did not outperform the baseline on every task.

  13. Knowl 13 — Image layout helps KOSMOS-1 answer questions about web pages

    empirical result

    On WebSRC, a web-page question-answering benchmark requiring both semantic and structural understanding, KOSMOS-1 was compared with a language-only model trained on the same text data and setup. Both received extracted page text; KOSMOS-1 additionally received the page image. Exact match (EM) and F1 were 15.8 and 31.3 for KOSMOS-1, versus 7.6 and 17.9 for the language model. KOSMOS-1 evaluated without extracted page text scored 3.8 EM and 10.6 F1. Thus, in this setup, page images improved performance over text alone, while extracted text still contributed substantially to KOSMOS-1's results.

Coverage note — Future plans to scale the model further and add speech capability are omitted because they are proposed directions, not completed contributions.

References

  1. 1.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karen Simonyan. Flamingo: a visual language model for few-shot learning. In Advances in Neural Information Processing Systems, 2022.
  2. 2.Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. Spice: Semantic propositional image caption evaluation. In ECCV, pages 382–398, 2016.
  3. 3.Armen Aghajanyan, Bernie Huang, Candace Ross, Vladimir Karpukhin, Hu Xu, Naman Goyal, Dmytro Okhonko, Mandar Joshi, Gargi Ghosh, Mike Lewis, and Luke Zettlemoyer. CM3: A causal masked multimodal model of the Internet. ArXiv, abs/2201.07520, 2022.
  4. 4.Elia Bruni, Gemma Boleda, Marco Baroni, and Nam Khanh Tran. Distributional semantics in technicolor. In ACL, 2012.
  5. 5.Hessam Bagherinezhad, Hannaneh Hajishirzi, Yejin Choi, and Ali Farhadi. Are elephants bigger than butterflies? reasoning about sizes of objects. ArXiv, abs/1602.00753, 2016.
  6. 6.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc., 2020.
  7. 7.Minwoo Byeon, Beomhee Park, Haecheon Kim, Sungjun Lee, Woonhyuk Baek, and Saehoon Kim. Coyo-700m: Image-text pair dataset, 2022.
  8. 8.Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. Piqa: Reasoning about physical commonsense in natural language. In Thirty-Fourth AAAI Conference on Artificial Intelligence, 2020.
  9. 9.Zewen Chi, Li Dong, Shaohan Huang, Damai Dai, Shuming Ma, Barun Patra, Saksham Singhal, Payal Bajaj, Xia Song, Xian-Ling Mao, Heyan Huang, and Furu Wei. On the representation collapse of sparse mixture of experts. In Advances in Neural Information Processing Systems, 2022.
  10. 10.Patricia A Carpenter, Marcel A Just, and Peter Shell. What one intelligence test measures: a theoretical account of the processing in the raven progressive matrices test. Psychological review, 97(3):404, 1990.
  11. 11.Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. BoolQ: Exploring the surprising difficulty of natural yes/no questions. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 2924–2936, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics.
  12. 12.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek B Rao, Parker Barnes, Yi Tay, Noam M. Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Benton C. Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier García, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Oliveira Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Díaz, Orhan Firat, Michele Catasta, Jason Wei, Kathleen S. Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. PaLM: Scaling language modeling with pathways. ArXiv, abs/2204.02311, 2022.
  13. 13.Soravit Changpinyo, Piyush Sharma, Nan Ding, and Radu Soricut. Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3558–3568, 2021.
  14. 14.Xingyu Chen, Zihan Zhao, Lu Chen, JiaBao Ji, Danyang Zhang, Ao Luo, Yuxuan Xiong, and Kai Yu. WebSRC: A dataset for web-based structural reading comprehension. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 4173–4185, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics.
  15. 15.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR 2009), 20-25 June 2009, Miami, Florida, USA, pages 248–255. IEEE Computer Society, 2009.
  16. 16.Marie-Catherine de Marneffe, Mandy Simons, and Judith Tonhauser. The Commitment-Bank: Investigating projection in naturally occurring discourse. Proceedings of Sinn und Bedeutung, 23(2):107–124, Jul. 2019.
  17. 17.Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al. The pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027, 2020.
  18. 18.Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In CVPR, pages 6325–6334, 2017.
  19. 19.Danna Gurari, Qing Li, Abigale J Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P Bigham. Vizwiz grand challenge: Answering visual questions from blind people. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3608–3617, 2018.
  20. 20.Dan Hendrycks and Kevin Gimpel. Gaussian error linear units (GELUs). arXiv preprint arXiv:1606.08415, 2016.
  21. 21.Yaru Hao, Haoyu Song, Li Dong, Shaohan Huang, Zewen Chi, Wenhui Wang, Shuming Ma, and Furu Wei. Language models are general-purpose interfaces. ArXiv, abs/2206.06336, 2022.
  22. 22.Or Honovich, Thomas Scialom, Omer Levy, and Timo Schick. Unnatural instructions: Tuning language models with (almost) no human labor, 2022.
  23. 23.John and Jean Raven. Raven Progressive Matrices, pages 223–237. Springer US, Boston, MA, 2003.
  24. 24.Andrej Karpathy and Li Fei-Fei. Deep visual-semantic alignments for generating image descriptions. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39(4):664–676, 2017.
  25. 25.Douwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami, Amanpreet Singh, Pratik Ringshia, and Davide Testuggine. The hateful memes challenge: Detecting hate speech in multimodal memes. In Advances in Neural Information Processing Systems, volume 33, pages 2611–2624, 2020.
  26. 26.Taku Kudo and John Richardson. SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. In EMNLP, pages 66–71, 2018.
  27. 27.Jing Yu Koh, Ruslan Salakhutdinov, and Daniel Fried. Grounding language models to images for multimodal generation. arXiv preprint arXiv:2301.13823, 2023.
  28. 28.Hector Levesque, Ernest Davis, and Leora Morgenstern. The winograd schema challenge. In Thirteenth International Conference on the Principles of Knowledge Representation and Reasoning, 2012.
  29. 29.Hector J. Levesque, Ernest Davis, and Leora Morgenstern. The winograd schema challenge. In Principles of Knowledge Representation and Reasoning, 2012.
  30. 30.Shayne Longpre, Le Hou, Tu Vu, Albert Webson, Hyung Won Chung, Yi Tay, Denny Zhou, Quoc V Le, Barret Zoph, Jason Wei, et al. The flan collection: Designing data and methods for effective instruction tuning. arXiv preprint arXiv:2301.13688, 2023.
  31. 31.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. ArXiv, abs/2301.12597, 2023.
  32. 32.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In ECCV, pages 740–755, 2014.
  33. 33.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. RoBERTa: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019.
  34. 34.Nasrin Mostafazadeh, Michael Roth, Annie Louis, Nathanael Chambers, and James Allen. Lsdsem 2017 shared task: The story cloze test. In Proceedings of the 2nd Workshop on Linking Models of Lexical, Sentential and Discourse-level Semantics, pages 46–51, 2017.
  35. 35.Shuming Ma, Hongyu Wang, Shaohan Huang, Wenhui Wang, Zewen Chi, Li Dong, Alon Benhaim, Barun Patra, Vishrav Chaudhary, Xia Song, and Furu Wei. TorchScale: Transformers at scale. CoRR, abs/2211.13184, 2022.
  36. 36.Tobias Norlund, Lovisa Hagström, and Richard Johansson. Transferring knowledge from vision to language: How to achieve it and how to measure it? ArXiv, abs/2109.11321, 2021.
  37. 37.Melissa Roemmele, Cosmin Adrian Bejan, and Andrew S. Gordon. Choice of plausible alternatives: An evaluation of commonsense causal reasoning. In AAAI Spring Symposium, 2011.
  38. 38.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PMLR, 2021.
  39. 39.Jack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier, and Timothy P. Lillicrap. Compressive transformers for long-range sequence modelling. In ICLR, 2020.
  40. 40.Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. WinoGrande: An adversarial winograd schema challenge at scale. In AAAI, pages 8732–8740, 2020.
  41. 41.Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al. Laion-5b: An open large-scale dataset for training next generation image-text models. arXiv preprint arXiv:2210.08402, 2022.
  42. 42.Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, ACL 2018, Melbourne, Australia, July 15-20, 2018, Volume 1: Long Papers, pages 2556–2565. Association for Computational Linguistics, 2018.
  43. 43.Yutao Sun, Li Dong, Barun Patra, Shuming Ma, Shaohan Huang, Alon Benhaim, Vishrav Chaudhary, Xia Song, and Furu Wei. A length-extrapolatable transformer. arXiv preprint arXiv:2212.10554, 2022.
  44. 44.Shaden Smith, Mostofa Patwary, Brandon Norick, Patrick LeGresley, Samyam Rajbhandari, Jared Casper, Zhun Liu, Shrimai Prabhumoye, George Zerveas, Vijay Korthikanti, Elton Zhang, Rewon Child, Reza Yazdani Aminabadi, Julie Bernauer, Xia Song, Mohammad Shoeybi, Yuxiong He, Michael Houston, Saurabh Tiwary, and Bryan Catanzaro. Using DeepSpeed and Megatron to train Megatron-Turing NLG 530B, a large-scale generative language model, 2022.
  45. 45.Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. Megatron-lm: Training multi-billion parameter language models using model parallelism. arXiv preprint arXiv:1909.08053, 2019.
  46. 46.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, pages 1631–1642, Seattle, Washington, USA, October 2013. Association for Computational Linguistics.
  47. 47.Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. Laion-400m: Open dataset of clip-filtered 400 million image-text pairs. arXiv preprint arXiv:2111.02114, 2021.
  48. 48.Maria Tsimpoukelli, Jacob Menick, Serkan Cabi, S. M. Ali Eslami, Oriol Vinyals, and Felix Hill. Multimodal few-shot learning with frozen language models. In Neural Information Processing Systems, 2021.
  49. 49.Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh. Cider: Consensus-based image description evaluation. In CVPR, pages 4566–4575, 2015.
  50. 50.Wenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck, Zhiliang Peng, Qiang Liu, Kriti Aggarwal, Owais Mohammed, Saksham Singhal, Subhojit Som, and Furu Wei. Image as a foreign language: BEiT pretraining for all vision and vision-language tasks. ArXiv, abs/2208.10442, 2022.
  51. 51.Catherine Wah, Steve Branson, Peter Welinder, Pietro Perona, and Serge J. Belongie. The caltech-ucsd birds-200-2011 dataset. 2011.
  52. 52.Chengyi Wang, Sanyuan Chen, Yu Wu, Zi-Hua Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, Lei He, Sheng Zhao, and Furu Wei. Neural codec language models are zero-shot text to speech synthesizers. ArXiv, abs/2301.02111, 2023.
  53. 53.Weizhi Wang, Li Dong, Hao Cheng, Haoyu Song, Xiaodong Liu, Xifeng Yan, Jianfeng Gao, and Furu Wei. Visually-augmented language modeling. In International Conference on Learning Representations, 2023.
  54. 54.Hongyu Wang, Shuming Ma, Li Dong, Shaohan Huang, Dongdong Zhang, and Furu Wei. DeepNet: Scaling Transformers to 1,000 layers. CoRR, abs/2203.00555, 2022.
  55. 55.Hongyu Wang, Shuming Ma, Shaohan Huang, Li Dong, Wenhui Wang, Zhiliang Peng, Yu Wu, Payal Bajaj, Saksham Singhal, Alon Benhaim, Barun Patra, Zhun Liu, Vishrav Chaudhary, Xia Song, and Furu Wei. Foundation transformers. CoRR, abs/2210.06423, 2022.
  56. 56.Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. SuperGLUE: A stickier benchmark for general-purpose language understanding systems. arXiv preprint arXiv:1905.00537, 2019.
  57. 57.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. Chain of thought prompting elicits reasoning in large language models. arXiv preprint arXiv:2201.11903, 2022.
  58. 58.Jianfeng Wang, Zhengyuan Yang, Xiaowei Hu, Linjie Li, Kevin Lin, Zhe Gan, Zicheng Liu, Ce Liu, and Lijuan Wang. GIT: A generative image-to-text transformer for vision and language. CoRR, abs/2205.14100, 2022.
  59. 59.Michihiro Yasunaga, Armen Aghajanyan, Weijia Shi, Rich James, Jure Leskovec, Percy Liang, Mike Lewis, Luke Zettlemoyer, and Wen tau Yih. Retrieval-augmented multimodal language modeling. ArXiv, abs/2211.12561, 2022.
  60. 60.Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. TACL, 2:67–78, 2014.
  61. 61.Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019.

Citation

MLA
Huang, S., et al. “Language Is Not All You Need: Aligning Perception with Language Models”. arXiv, 2023, http://arxiv.org/abs/2302.14045v2.
APA
Huang, S., Dong, L., Wang, W., Hao, Y., Singhal, S., Ma, S., Lv, T., Cui, L., Mohammed, O. K., Patra, B., Liu, Q., Aggarwal, K., Chi, Z., Bjorck, J., Chaudhary, V., Som, S., Song, X., & Wei, F. (2023). Language Is Not All You Need: Aligning Perception with Language Models. arXiv. http://arxiv.org/abs/2302.14045v2
Chicago
Huang, S., L. Dong, W. Wang, et al. 2023. “Language Is Not All You Need: Aligning Perception with Language Models”. arXiv. http://arxiv.org/abs/2302.14045v2.
Harvard
Huang, S. et al. (2023) “Language Is Not All You Need: Aligning Perception with Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2302.14045v2.
Vancouver
1. Huang S, Dong L, Wang W, et al (2023) Language Is Not All You Need: Aligning Perception with Language Models. arXiv

BibTeX

@article{huang2023language,
  title = {Language Is Not All You Need: Aligning Perception with Language Models},
  author = {Huang, Shaohan and Dong, Li and Wang, Wenhui and Hao, Yaru and Singhal, Saksham and Ma, Shuming and Lv, Tengchao and Cui, Lei and Mohammed, Owais Khan and Patra, Barun and Liu, Qiang and Aggarwal, Kriti and Chi, Zewen and Bjorck, Johan and Chaudhary, Vishrav and Som, Subhojit and Song, Xia and Wei, Furu},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2302.14045v2},
  eprint = {2302.14045}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission