Grounding Language Models to Images for Multimodal Inputs and Outputs

Jing Yu KohRuslan SalakhutdinovDaniel Fried

article2023ICML166 citations

Presents FROMAGe, an efficient approach that equips frozen text-only language models to process arbitrarily interleaved image-text inputs and generate text interspersed with retrieved images by training only lightweight linear translation layers.

Listen

State-of-the-art large language models demonstrate impressive conversational and reasoning skills, but because they are trained strictly on text, they lack visual grounding in the physical world. Adapting these models to understand and generate multimodal content typically requires massive computational infrastructure and web-scale datasets of interleaved image-text documents. This high resource barrier limits rapid experimentation and prevents organizations from easily deploying systems capable of fluid visual and textual dialogue.

The article demonstrates an efficient method called FROMAGe (Frozen Retrieval Over Multimodal Data for Autoregressive Generation). Its primary objective is to visually ground a frozen, pretrained text-only language model to process interleaved image-and-text inputs and generate text interleaved with retrieved images.

To achieve this, the authors connected a frozen 6.7-billion-parameter text model to a frozen visual model using lightweight, trainable linear mapping layers and a dedicated retrieval token. The system was trained on standard paired image-caption data (approximately 3.1 million examples) using a dual objective: generating captions from visual prefixes and learning contrastive embeddings for cross-modal retrieval. Because 97 percent of the model parameters remain frozen, the training required only 5.5 million parameter updates and was completed in 24 hours on a single graphics processing unit, contrasting sharply with competing frameworks that require hundreds to thousands of processors over multiple weeks.

The experimental findings show that the model outperforms established baselines when contextual depth increases. On contextual image retrieval using the Visual Storytelling dataset, the model improved top-one retrieval accuracy by roughly 77 percent relative to a standard vision-language baseline when provided with full image-text story context. While baseline models degraded by roughly 50 percent on long, temporally dependent text descriptions, this approach leveraged additional context to improve retrieval accuracy. Furthermore, human evaluations confirmed that conditioning on interleaved images and captions produced significantly more coherent narratives and relevant descriptions than single-modality inputs. In zero-shot visual dialogue benchmarks, the model outperformed several existing systems in answering dialogue questions and surpassed the baseline by 17.5 percent in conversation-based image retrieval.

These results demonstrate that organizations can successfully bridge text and vision without retraining expensive foundation models from scratch. Keeping the core language model frozen preserves its pre-existing reasoning, in-context learning, and world knowledge while significantly reducing training costs, memory overhead, and implementation timelines. Using image retrieval instead of open-ended image generation also provides a distinct governance advantage: organizations can strictly curate and filter the pool of candidate images to reduce safety and compliance risks.

Decision-makers considering multimodal conversational interfaces can adopt this modular framework to upgrade existing language backbones at low computational cost. Before deploying such systems to production, teams should implement logit adjustments or structured prompting to ensure the model triggers image retrievals reliably, and they must curate retrieval image pools to prevent biased or inappropriate visuals. Future work should focus on fine-tuning with multimodal dialogue instructions to improve spontaneous retrieval generation and exploring the integration of novel image synthesis.

The findings are subject to certain boundaries. Because visual outputs are generated through retrieval rather than synthesis, the model cannot generate novel or out-of-distribution imagery, such as fantastical scenes. The system also inherits the standard risks of underlying language models, including occasional factual errors or degenerative text repetition. Despite these constraints, the confidence in the core approach is high, as consistent performance gains were verified across multiple benchmarks, ablation studies, and scaling evaluations.

  • Paper: NExT-GPT: Any-to-Any Multimodal LLM, Shengqiong Wu et al. (2024). NExT-GPT extends the interleaved multimodal input-output idea beyond images to audio and video, using lightweight adapters to connect frozen foundation models.
Cover for Grounding Language Models to Images for Multimodal Inputs and Outputs

Abstract

We propose an efficient method to ground pre-trained text-only language models to the visual domain, enabling them to process arbitrarily interleaved image-and-text data, and generate text interleaved with retrieved images. Our method leverages the abilities of language models learnt from large scale text-only pretraining, such as in-context learning and free-form text generation. We keep the language model frozen, and fine-tune input and output linear layers to enable cross-modality interactions. This allows our model to process arbitrarily interleaved image-and-text inputs, and generate free-form text interleaved with retrieved images. We achieve strong zero-shot performance on grounded tasks such as contextual image retrieval and multimodal dialogue, and showcase compelling interactive abilities. Our approach works with any off-the-shelf language model and paves the way towards an effective, general solution for leveraging pretrained language models in visually grounded settings.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Model Architecture
  • 3.2. Translating Between Image-and-Text
  • 3.3. Training Setup
  • 3.4. Data and Implementation Details
  • 4. Experiments
  • 4.1. Contextual Retrieval from Multimodal Inputs
  • 4.2. Visual Dialogue
  • 4.3. Qualitative Results
  • 5. Analysis
  • 5.1. Ablation Experiments
  • 5.2. The Effect of Context
  • 5.3. In-context Learning and Text Generation
  • 6. Future Work
  • 7. Conclusion
  • Acknowledgements
  • References
  • A. Qualitative Examples
  • A.1. Comparison Against CM3
  • B. Further Analysis
  • B.1. Details on Freezing Ablation
  • B.2. Joint Retrieval + Captioning Training
  • B.3. Image-Text Concatenation for Captioning
  • B.4. Image-Text Concatenation for Retrieval
  • B.5. Scaling Properties
  • B.6. Text Generation Results
  • C. Human Evaluation Procedure
  • D. Current Limitations and Broader Impacts

Knowls

  1. Knowl 1 — FROMAGe bridges frozen language and vision models with learned linear projections

    model/method

    FROMAGe grounds a pretrained autoregressive text-only language model and a pretrained visual encoder while keeping both backbones frozen. For an image yy, the visual encoder returns vϕ(y)∈Rmv_\phi(y)\in\mathbb{R}^m. A learned matrix Wc∈Rm×kdW_c\in\mathbb{R}^{m\times kd} maps this vector to kk vectors of dimension dd, after reshaping the result to k×dk\times d; these vectors act as soft visual-prefix tokens in the language model’s input embedding space. Thus text tokens and mapped image representations can be interleaved in the same input sequence.

    For image retrieval, FROMAGe appends a learned special token, [RET][\mathrm{RET}], to a caption. The language model’s final-layer hidden representation at that token, hθ(x)∈Rph_\theta(x)\in\mathbb{R}^p, is mapped by Wt∈Rp×qW_t\in\mathbb{R}^{p\times q} into a retrieval space. The visual representation is mapped into the same space by Wi∈Rm×qW_i\in\mathbb{R}^{m\times q}. At inference, generating [RET][\mathrm{RET}] triggers retrieval of an image from a candidate image set using the projected text representation. The language model and visual encoder stay fixed; the projections and the [RET][\mathrm{RET}] embedding are learned. This design supports text generation interleaved with retrieved images as well as inputs containing interleaved images and text.

  2. Knowl 2 — Captioning and bidirectional contrastive retrieval are trained jointly

    equation

    FROMAGe trains on paired captions xix_i and images yiy_i, for i=1,…,Ni=1,\ldots,N. For captioning, let sits_{it} be token tt of caption xix_i, and let si,<ts_{i,<t} denote its preceding tokens. If pθp_\theta is the frozen language model and vϕ(yi)TWcv_\phi(y_i)^T W_c is the mapped image prefix, the batch captioning loss is

    Lc=−1N∑i=1N∑t=1Tilog⁡pθ ⁣(sit∣vϕ(yi)TWc,si,<t),L_c=-\frac{1}{N}\sum_{i=1}^{N}\sum_{t=1}^{T_i}\log p_\theta\!\left(s_{it}\mid v_\phi(y_i)^T W_c,s_{i,<t}\right),

    where TiT_i is the number of tokens in caption ii.

    For retrieval, the caption ends with [RET][\mathrm{RET}]. Define projected vectors ui=hθ(xi)TWt∈Rqu_i=h_\theta(x_i)^T W_t\in\mathbb{R}^q and zi=vϕ(yi)TWi∈Rqz_i=v_\phi(y_i)^T W_i\in\mathbb{R}^q, and let sim⁡(xi,yj)\operatorname{sim}(x_i,y_j) be their cosine similarity, uiTzj/(∥ui∥∥zj∥)u_i^Tz_j/(\lVert u_i\rVert\lVert z_j\rVert). The text-to-image and image-to-text InfoNCE losses treat paired items as positives and the other items in the batch as negatives:

    Lt2i=−1N∑i=1Nlog⁡exp⁡(sim⁡(xi,yi)/τ)∑j=1Nexp⁡(sim⁡(xi,yj)/τ),Li2t=−1N∑i=1Nlog⁡exp⁡(sim⁡(xi,yi)/τ)∑j=1Nexp⁡(sim⁡(xj,yi)/τ),L_{t2i}=-\frac{1}{N}\sum_{i=1}^{N}\log\frac{\exp(\operatorname{sim}(x_i,y_i)/\tau)}{\sum_{j=1}^{N}\exp(\operatorname{sim}(x_i,y_j)/\tau)},\qquad L_{i2t}=-\frac{1}{N}\sum_{i=1}^{N}\log\frac{\exp(\operatorname{sim}(x_i,y_i)/\tau)}{\sum_{j=1}^{N}\exp(\operatorname{sim}(x_j,y_i)/\tau)},

    where τ\tau is a learned temperature. The combined objective is L=λcLc+λr(Lt2i+Li2t)L=\lambda_cL_c+\lambda_r(L_{t2i}+L_{i2t}), with λc\lambda_c and λr\lambda_r the captioning and retrieval weights. Gradients update only WcW_c, WtW_t, WiW_i, and the [RET][\mathrm{RET}] embedding.

  3. Knowl 3 — Zero-shot contextual retrieval improves as FROMAGe receives interleaved story context

    empirical result

    On Visual Storytelling (VIST), each test example is a temporally ordered sequence of five image-caption pairs. FROMAGe is evaluated on retrieving the final image from its caption, from the five captions, or from the preceding five captions and four images. Recall@k results are shown below; the dagger marks retrieval over images not previously shown in the input story. Values are reported as in the paper.

    Model Input R@1 R@5 R@10
    CLIP ViT-L/14 1 caption 11.9 25.5 32.2
    FROMAGe 1 caption 9.0 21.1 28.7
    CLIP ViT-L/14 5 captions 5.9 19.5 28.0
    FROMAGe 5 captions 10.4 23.8 31.7
    BLIP†^{\dagger} 5 captions 6.2 16.8 23.4
    CLIP ViT-L/14†^{\dagger} 5 captions 8.8 22.3 29.8
    FROMAGe†^{\dagger} 5 captions 11.6 24.7 32.8
    CLIP ViT-L/14 5 captions, 4 images 2.4 21.3 34.0
    FROMAGe†^{\dagger} 5 captions, 4 images 15.6 36.5 45.8

    CLIP performs better on the single-caption setting, but its recall drops with five captions and drops further when averaged image and text context is supplied. FROMAGe improves with longer and multimodal context; with five captions and four images, its R@1 is 15.6, compared with 8.8 for CLIP given five captions and unseen retrieval targets.

  4. Knowl 4 — FROMAGe supports both candidate-answer selection and text-to-image retrieval in Visual Dialog

    empirical result

    The zero-shot Visual Dialog evaluation measures image-and-text-to-text (IT2T) answer selection and text-to-image (T2I) retrieval. For IT2T, FROMAGe scores each candidate question-and-answer sequence by perplexity and ranks candidates from lowest to highest perplexity. The table reports model size, finetuning-data size, IT2T metrics, and T2I Recall@k; dashes indicate unreported values, and “Incapable” indicates that the model cannot perform the task.

    Model Trainable params Finetuning data NDCG MRR IT2T R@1 IT2T R@5 IT2T R@10 T2I R@1 T2I R@5 T2I R@10
    ViLBERT 114M 3.1M 11.6 6.9 2.6 7.2 11.3 – – –
    CLIP ViT-L/14 300M 400M 10.9 8.5 3.1 8.7 15.9 17.7 38.9 50.2
    Flamingo 10.2B 1.8B 52.0 – – – – Incapable Incapable Incapable
    ESPER 4M 0.5M 22.3 25.7 14.6 – – Incapable Incapable Incapable
    FROMAGe 5.5M 3.1M 16.5 22.0 17.6 20.1 25.1 20.8 44.9 56.0

    FROMAGe’s IT2T R@1 of 17.6 exceeds ESPER’s 14.6, while its NDCG and MRR are below ESPER’s. Its T2I R@1 of 20.8 is 17.5% higher relative to CLIP’s 17.7; ESPER and Flamingo cannot perform the image-retrieval task.

  5. Knowl 5 — Multimodal context improves retrieval and supports story-like generation

    empirical result

    In VIST context ablations, retrieval improves as preceding image-caption pairs are added. The reported R@1 rises from 9.0 with one caption to 12.8 with two captions and one image, and to 15.6 with five captions and four images. The comparison indicates that adding an image can help more than adding text descriptions alone. On Visual Dialog image retrieval, performance also increases as more dialogue rounds are provided; with the full dialogue, FROMAGe’s R@1 is reported as 17.5% higher relative to CLIP.

    Human comparisons of generated VIST text found that conditioning on the full preceding image-and-caption context produced more coherent stories than conditioning on only the final image or only the text descriptions. Full multimodal context was judged more relevant to the images than text-only context, but less relevant than a single-image input, which more often elicited factual, caption-like descriptions. The results support the distinction between image-focused captioning and context-conditioned story generation.

  6. Knowl 6 — Freezing the language model preserves downstream generalization

    empirical result

    An ablation using OPT 1.3B compared FROMAGe with a frozen language-model backbone against a version in which that backbone was finetuned. Although finetuning reduced training and validation loss on Conceptual Captions, it harmed downstream performance: VIST contextual-retrieval R@1 fell from 12.8 to 6.2, and Visual Dialog IT2T R@1 fell from 14.6 to 1.0. The authors interpret this as evidence that freezing the language model helps retain its pretrained in-context learning and zero-shot generalization abilities.

  7. Knowl 7 — A learned retrieval token substantially improves image retrieval

    empirical result

    FROMAGe adds a learned [RET][\mathrm{RET}] token to the language-model vocabulary and appends it to training captions, so its hidden state can serve as the caption representation for image retrieval. In the VIST setting with full multimodal context, the version without this dedicated token obtains R@1 of 11.3, while the version with it obtains R@1 of 15.6, a reported relative gain of 38.1%. The token also allows autoregressive generation to signal when an image should be retrieved and inserted into the output.

  8. Knowl 8 — Random concatenation trains captioning to handle sequences of images

    model/method

    During captioning training, FROMAGe randomly concatenates distinct image-caption examples sequentially with probability 0.5. This exposes the model to multiple images in one input and encourages it to associate each caption with the correct image. On VIST, enabling this captioning-time concatenation raises R@1 from 11.6 to 15.6 when the input contains five captions and four images; the authors report similar performance on VisDial with or without this augmentation. Applying concatenation to retrieval training, where the model must retrieve a separate image for each of two captions, did not help these downstream tasks: VIST R@1 decreased from 15.6 to 14.4.

  9. Knowl 9 — FROMAGe inherits limitations of its language model and retrieval-based image output

    limitation

    FROMAGe retrieves images from a fixed candidate set rather than synthesizing novel images, so its image outputs are limited by the contents of that set and may be poor for prompts unlikely to occur in natural-image data. It also does not reliably generate [RET][\mathrm{RET}] during inference and tends to favor ordinary text tokens; the authors report that multiplying [RET][\mathrm{RET}] logits by 1.3–1.5, using in-context examples, or explicitly asking for images can help. Because the system uses a frozen language-model backbone, it can inherit failures such as repetitive or incoherent text, failure to follow instructions, misinformation, and toxic or socially biased content. Retrieved images can also reflect biases in the training and retrieval data.

Coverage note — The supplementary model-size scaling results and the MS-COCO/VQAv2 auxiliary evaluations are omitted because they are secondary to the core method and contextual-retrieval and dialogue findings captured here.

References

  1. 1.Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G. S., Davis, A., Dean, J., Devin, M., et al. Tensorflow: Large-scale machine learning on heterogeneous distributed systems. arXiv preprint arXiv:1603.04467, 2016.
  2. 2.Aghajanyan, A., Huang, B., Ross, C., Karpukhin, V., Xu, H., Goyal, N., Okhonko, D., Joshi, M., Ghosh, G., Lewis, M., et al. Cm3: A causal masked multimodal model of the internet. arXiv preprint arXiv:2201.07520, 2022.
  3. 3.Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al. Flamingo: a visual language model for few-shot learning. NeurIPS, 2022.
  4. 4.Banerjee, S. and Lavie, A. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, 2005.
  5. 5.Bender, E. M., Gebru, T., McMillan-Major, A., and Shmitchell, S. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pp. 610–623, 2021.
  6. 6.Birhane, A., Prabhu, V. U., and Kahembwe, E. Multimodal datasets: misogyny, pornography, and malignant stereotypes. arXiv preprint arXiv:2110.01963, 2021.
  7. 7.Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258, 2021.
  8. 8.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. NeurIPS, 2020.
  9. 9.Chan, S. C., Santoro, A., Lampinen, A. K., Wang, J. X., Singh, A., Richemond, P. H., McClelland, J., and Hill, F. Data distributional properties drive emergent few-shot learning in transformers. NeurIPS, 2022.
  10. 10.Chopra, S., Hadsell, R., and LeCun, Y. Learning a similarity metric discriminatively, with application to face verification. In CVPR, 2005.
  11. 11.Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311, 2022.
  12. 12.Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., et al. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416, 2022.
  13. 13.Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q. V., and Salakhutdinov, R. Transformer-xl: Attentive language models beyond a fixed-length context. ACL, 2019.
  14. 14.Das, A., Kottur, S., Gupta, K., Singh, A., Yadav, D., Moura, J. M., Parikh, D., and Batra, D. Visual dialog. In CVPR, 2017.
  15. 15.Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L. Llm. int8 (): 8-bit matrix multiplication for transformers at scale. NeurIPS, 2022.
  16. 16.Ding, M., Zheng, W., Hong, W., and Tang, J. Cogview2: Faster and better text-to-image generation via hierarchical transformers. arXiv preprint arXiv:2204.14217, 2022.
  17. 17.Eichenberg, C., Black, S., Weinbach, S., Parcalabescu, L., and Frank, A. Magma–multimodal augmentation of generative models through adapter-based finetuning. EMNLP, 2022.
  18. 18.Esser, P., Rombach, R., and Ommer, B. Taming transformers for high-resolution image synthesis. In CVPR, 2021.
  19. 19.Gehman, S., Gururangan, S., Sap, M., Choi, Y., and Smith, N. A. Realtoxicityprompts: Evaluating neural toxic degeneration in language models. EMNLP, 2020.
  20. 20.Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D. Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In CVPR, 2017.
  21. 21.Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al. Training compute-optimal large language models. NeurIPS, 2022.
  22. 22.Holtzman, A., Buys, J., Du, L., Forbes, M., and Choi, Y. The curious case of neural text degeneration. ICLR, 2020.
  23. 23.Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S. Parameter-efficient transfer learning for nlp. In ICML, 2019.
  24. 24.Huang, T.-H., Ferraro, F., Mostafazadeh, N., Misra, I., Agrawal, A., Devlin, J., Girshick, R., He, X., Kohli, P., Batra, D., et al. Visual storytelling. In NAACL-HLT, 2016.
  25. 25.Jia, C., Yang, Y., Xia, Y., Chen, Y.-T., Parekh, Z., Pham, H., Le, Q., Sung, Y.-H., Li, Z., and Duerig, T. Scaling up visual and vision-language representation learning with noisy text supervision. In ICLR, 2021.
  26. 26.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. ICLR, 2015.
  27. 27.Lester, B., Al-Rfou, R., and Constant, N. The power of scale for parameter-efficient prompt tuning. EMNLP, 2021.
  28. 28.Levesque, H., Davis, E., and Morgenstern, L. The winograd schema challenge. In Thirteenth international conference on the principles of knowledge representation and reasoning, 2012.
  29. 29.Li, J., Li, D., Xiong, C., and Hoi, S. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In ICML, 2022a.
  30. 30.Li, X. L. and Liang, P. Prefix-tuning: Optimizing continuous prompts for generation. ACL, 2021.
  31. 31.Li, X. L., Holtzman, A., Fried, D., Liang, P., Eisner, J., Hashimoto, T., Zettlemoyer, L., and Lewis, M. Contrastive decoding: Open-ended text generation as optimization. arXiv preprint arXiv:2210.15097, 2022b.
  32. 32.Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollar, P., and Zitnick, C. L. Microsoft coco: Common objects in context. In ECCV, 2014.
  33. 33.Lu, J., Batra, D., Parikh, D., and Lee, S. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. NeurIPS, 2019.
  34. 34.Lu, K., Grover, A., Abbeel, P., and Mordatch, I. Pretrained transformers as universal computation engines. AAAI, 2022.
  35. 35.Merullo, J., Castricato, L., Eickhoff, C., and Pavlick, E. Linearly mapping from image to text space. arXiv preprint arXiv:2209.15162, 2022.
  36. 36.Oord, A. v. d., Li, Y., and Vinyals, O. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2018.
  37. 37.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. Training language models to follow instructions with human feedback. arXiv preprint arXiv:2203.02155, 2022.
  38. 38.Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J. Bleu: a method for automatic evaluation of machine translation. In ACL, 2002.
  39. 39.Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al. Pytorch: An imperative style, high-performance deep learning library. NeurIPS, 2019.
  40. 40.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  41. 41.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. Learning transferable visual models from natural language supervision. In ICLR, 2021.
  42. 42.Rae, J. W., Borgeaud, S., Cai, T., Millican, K., Hoffmann, J., Song, F., Aslanides, J., Henderson, S., Ring, R., Young, S., et al. Scaling language models: Methods, analysis & insights from training gopher. arXiv preprint arXiv:2112.11446, 2021.
  43. 43.Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I. Zero-shot text-to-image generation. In ICML, 2021.
  44. 44.Reed, S., Akata, Z., Yan, X., Logeswaran, L., Schiele, B., and Lee, H. Generative adversarial text to image synthesis. In ICML, 2016.
  45. 45.Schuhmann, C., Vencu, R., Beaumont, R., Kaczmarczyk, R., Mullis, C., Katta, A., Coombes, T., Jitsev, J., and Komatsuzaki, A. Laion-400m: Open dataset of clip-filtered 400 million image-text pairs. arXiv preprint arXiv:2111.02114, 2021.
  46. 46.Sennrich, R., Haddow, B., and Birch, A. Neural machine translation of rare words with subword units. ACL, 2015.
  47. 47.Sharma, P., Ding, N., Goodman, S., and Soricut, R. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. ACL, 2018.
  48. 48.Smith, S., Patwary, M., Norick, B., LeGresley, P., Rajbhandari, S., Casper, J., Liu, Z., Prabhumoye, S., Zerveas, G., Korthikanti, V., et al. Using deepspeed and megatron to train megatron-turing nlg 530b, a large-scale generative language model. arXiv preprint arXiv:2201.11990, 2022.
  49. 49.Tan, B., Yang, Z., Al-Shedivat, M., Xing, E. P., and Hu, Z. Progressive generation of long text with pretrained language models. NAACL, 2021.
  50. 50.Tay, Y., Wei, J., Chung, H. W., Tran, V. Q., So, D. R., Shakeri, S., Garcia, X., Zheng, H. S., Rao, J., Chowdhery, A., et al. Transcending scaling laws with 0.1% extra compute. arXiv preprint arXiv:2210.11399, 2022.
  51. 51.Tsimpoukelli, M., Menick, J. L., Cabi, S., Eslami, S., Vinyals, O., and Hill, F. Multimodal few-shot learning with frozen language models. NeurIPS, 2021.
  52. 52.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. NeurIPS, 2017.
  53. 53.Wang, P., Yang, A., Men, R., Lin, J., Bai, S., Li, Z., Ma, J., Zhou, C., Zhou, J., and Yang, H. Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework. ICML, 2022.
  54. 54.Wei, J., Bosma, M., Zhao, V. Y., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V. Finetuned language models are zero-shot learners. ICLR, 2021.
  55. 55.Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., et al. Emergent abilities of large language models. TMLR, 2022.
  56. 56.Yang, K., Peng, N., Tian, Y., and Klein, D. Re3: Generating longer stories with recursive reprompting and revision. EMNLP, 2022.
  57. 57.Yu, J., Li, X., Koh, J. Y., Zhang, H., Pang, R., Qin, J., Ku, A., Xu, Y., Baldridge, J., and Wu, Y. Vector-quantized image modeling with improved vqgan. ICLR, 2021.
  58. 58.Yu, J., Xu, Y., Koh, J. Y., Luong, T., Baid, G., Wang, Z., Vasudevan, V., Ku, A., Yang, Y., Ayan, B. K., et al. Scaling autoregressive models for content-rich text-to-image generation. TMLR, 2022a.
  59. 59.Yu, Y., Chung, J., Yun, H., Hessel, J., Park, J., Lu, X., Ammanabrolu, P., Zellers, R., Bras, R. L., Kim, G., et al. Multimodal knowledge alignment with reinforcement learning. arXiv preprint arXiv:2205.12630, 2022b.
  60. 60.Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068, 2022.

Citation

MLA
Koh, J. Y., et al. “Grounding Language Models to Images for Multimodal Inputs and Outputs”. International Conference on Machine Learning, vol. 202, 2023, pp. 17283–300, https://proceedings.mlr.press/v202/koh23a.html.
APA
Koh, J. Y., Salakhutdinov, R., & Fried, D. (2023). Grounding Language Models to Images for Multimodal Inputs and Outputs. International Conference on Machine Learning, 202, 17283–17300. https://proceedings.mlr.press/v202/koh23a.html
Chicago
Koh, J. Y., R. Salakhutdinov, and D. Fried. 2023. “Grounding Language Models to Images for Multimodal Inputs and Outputs”. International Conference on Machine Learning 202: 17283–300. https://proceedings.mlr.press/v202/koh23a.html.
Harvard
Koh, J.Y., Salakhutdinov, R. and Fried, D. (2023) “Grounding Language Models to Images for Multimodal Inputs and Outputs”, International Conference on Machine Learning. PMLR, pp. 17283–17300. Available at: https://proceedings.mlr.press/v202/koh23a.html.
Vancouver
1. Koh JY, Salakhutdinov R, Fried D (2023) Grounding Language Models to Images for Multimodal Inputs and Outputs. In: International Conference on Machine Learning. PMLR, pp 17283–17300

BibTeX

@InProceedings{pmlr-v202-koh23a,
  title = 	 {Grounding Language Models to Images for Multimodal Inputs and Outputs},
  author =       {Koh, Jing Yu and Salakhutdinov, Ruslan and Fried, Daniel},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {17283--17300},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/koh23a/koh23a.pdf},
  url = 	 {https://proceedings.mlr.press/v202/koh23a.html},
  abstract = 	 {We propose an efficient method to ground pretrained text-only language models to the visual domain, enabling them to process arbitrarily interleaved image-and-text data, and generate text interleaved with retrieved images. Our method leverages the abilities of language models learnt from large scale text-only pretraining, such as in-context learning and free-form text generation. We keep the language model frozen, and finetune input and output linear layers to enable cross-modality interactions. This allows our model to process arbitrarily interleaved image-and-text inputs, and generate free-form text interleaved with retrieved images. We achieve strong zero-shot performance on grounded tasks such as contextual image retrieval and multimodal dialogue, and showcase compelling interactive abilities. Our approach works with any off-the-shelf language model and paves the way towards an effective, general solution for leveraging pretrained language models in visually grounded settings.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/