Grounding Language Models to Images for Multimodal Inputs and Outputs
Jing Yu KohRuslan SalakhutdinovDaniel Fried
Presents FROMAGe, an efficient approach that equips frozen text-only language models to process arbitrarily interleaved image-text inputs and generate text interspersed with retrieved images by training only lightweight linear translation layers.
State-of-the-art large language models demonstrate impressive conversational and reasoning skills, but because they are trained strictly on text, they lack visual grounding in the physical world. Adapting these models to understand and generate multimodal content typically requires massive computational infrastructure and web-scale datasets of interleaved image-text documents. This high resource barrier limits rapid experimentation and prevents organizations from easily deploying systems capable of fluid visual and textual dialogue.
The article demonstrates an efficient method called FROMAGe (Frozen Retrieval Over Multimodal Data for Autoregressive Generation). Its primary objective is to visually ground a frozen, pretrained text-only language model to process interleaved image-and-text inputs and generate text interleaved with retrieved images.
To achieve this, the authors connected a frozen 6.7-billion-parameter text model to a frozen visual model using lightweight, trainable linear mapping layers and a dedicated retrieval token. The system was trained on standard paired image-caption data (approximately 3.1 million examples) using a dual objective: generating captions from visual prefixes and learning contrastive embeddings for cross-modal retrieval. Because 97 percent of the model parameters remain frozen, the training required only 5.5 million parameter updates and was completed in 24 hours on a single graphics processing unit, contrasting sharply with competing frameworks that require hundreds to thousands of processors over multiple weeks.
The experimental findings show that the model outperforms established baselines when contextual depth increases. On contextual image retrieval using the Visual Storytelling dataset, the model improved top-one retrieval accuracy by roughly 77 percent relative to a standard vision-language baseline when provided with full image-text story context. While baseline models degraded by roughly 50 percent on long, temporally dependent text descriptions, this approach leveraged additional context to improve retrieval accuracy. Furthermore, human evaluations confirmed that conditioning on interleaved images and captions produced significantly more coherent narratives and relevant descriptions than single-modality inputs. In zero-shot visual dialogue benchmarks, the model outperformed several existing systems in answering dialogue questions and surpassed the baseline by 17.5 percent in conversation-based image retrieval.
These results demonstrate that organizations can successfully bridge text and vision without retraining expensive foundation models from scratch. Keeping the core language model frozen preserves its pre-existing reasoning, in-context learning, and world knowledge while significantly reducing training costs, memory overhead, and implementation timelines. Using image retrieval instead of open-ended image generation also provides a distinct governance advantage: organizations can strictly curate and filter the pool of candidate images to reduce safety and compliance risks.
Decision-makers considering multimodal conversational interfaces can adopt this modular framework to upgrade existing language backbones at low computational cost. Before deploying such systems to production, teams should implement logit adjustments or structured prompting to ensure the model triggers image retrievals reliably, and they must curate retrieval image pools to prevent biased or inappropriate visuals. Future work should focus on fine-tuning with multimodal dialogue instructions to improve spontaneous retrieval generation and exploring the integration of novel image synthesis.
The findings are subject to certain boundaries. Because visual outputs are generated through retrieval rather than synthesis, the model cannot generate novel or out-of-distribution imagery, such as fantastical scenes. The system also inherits the standard risks of underlying language models, including occasional factual errors or degenerative text repetition. Despite these constraints, the confidence in the core approach is high, as consistent performance gains were verified across multiple benchmarks, ablation studies, and scaling evaluations.
- Paper: BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models, Junnan Li et al. (2023). BLIP-2 shows how frozen image encoders and language models can be connected through a lightweight interface, a key design premise behind grounding pretrained language models to visual inputs and outputs.
- Paper: NExT-GPT: Any-to-Any Multimodal LLM, Shengqiong Wu et al. (2024). NExT-GPT extends the interleaved multimodal input-output idea beyond images to audio and video, using lightweight adapters to connect frozen foundation models.
