Built independently by an author, for readers. Read the story and support ChapterPal

keyword

multimodal prompts

Multimodal prompts are structured inputs provided to artificial intelligence models that combine two or more distinct data modalities, such as text, images, audio, video, or physical sensor readings, to guide understanding, reasoning, or content generation. Unlike traditional unimodal prompts that rely solely on a single medium like plain text, multimodal prompts allow models to interpret information whose meaning is distributed across different representations simultaneously. In machine learning architectures, these diverse inputs are encoded and aligned through mechanisms such as shared embedding spaces or cross-attention layers, enabling tasks that range from visual question answering and conditioned image synthesis to embodied robotic decision-making. By integrating complementary symbolic and perceptual signals into a unified query or instruction, multimodal prompts provide richer contextual grounding and allow for more precise control over complex model behaviors and outputs.

2 items

IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Hu Ye, Jun Zhang, Siyi Liu, Xiao Han, Wei Yang

OrganizationsTencent

Why you should read this

Introduces a lightweight, decoupled cross-attention adapter that equips pretrained text-to-image diffusion models with image prompting capabilities while maintaining full compatibility with text prompts and structural controls without retraining the base model.

Recent years have witnessed the strong power of large text-to-image diffusion models for the impressive generative capability to create high-fidelity images. However, it is very tricky to generate desired images using only text prompt as it often involves complex prompt engineering. An alternative to text prompt is image prompt, as the saying goes: "an image is worth a thousand words". Although existing methods of direct fine-tuning from pretrained models are effective, they require large computing resources and are not compatible with other base models, text prompt, and structural controls. In this paper, we present IP-Adapter, an effective and lightweight adapter to achieve image prompt capability for the pretrained text-to-image diffusion models. The key design of our IP-Adapter is decoupled cross-attention mechanism that separates cross-attention layers for text features and image features. Despite the simplicity of our method, an IP-Adapter with only 22M parameters can achieve comparable or even better performance to a fully fine-tuned image prompt model. As we freeze the pretrained diffusion model, the proposed IP-Adapter can be generalized not only to other custom models fine-tuned from the same base model, but also to controllable generation using existing controllable tools. With the benefit of the decoupled cross-attention strategy, the image prompt can also work well with the text prompt to achieve multimodal image generation. The project page is available at \url{this https URL}.

Added

2026-09-24

PaLM-E: An Embodied Multimodal Language Model

PaLM-E: An Embodied Multimodal Language Model

Danny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, Wenlong Huang, Yevgen Chebotar, Pierre Sermanet, Daniel Duckworth, Sergey Levine, Vincent Vanhoucke, Karol Hausman, Marc Toussaint, Klaus Greff, Andy Zeng, Igor Mordatch, Pete Florence

OrganizationsGoogleTechnische Universität Berlin

Why you should read this

Integrates real-world continuous sensor data into a language model, allowing the AI to reason about and act upon the physical world.

Large language models excel at a wide range of complex tasks. However, enabling general inference in the real world, e.g., for robotics problems, raises the challenge of grounding. We propose embodied language models to directly incorporate real-world continuous sensor modalities into language models and thereby establish the link between words and percepts. Input to our embodied language model are multi-modal sentences that interleave visual, continuous state estimation, and textual input encodings. We train these encodings end-to-end, in conjunction with a pre-trained large language model, for multiple embodied tasks including sequential robotic manipulation planning, visual question answering, and captioning. Our evaluations show that PaLM-E, a single large embodied multimodal model, can address a variety of embodied reasoning tasks, from a variety of observation modalities, on multiple embodiments, and further, exhibits positive transfer: the model benefits from diverse joint training across internet-scale language, vision, and visual-language domains. Our largest model, PaLM-E-562B with 562B parameters, in addition to being trained on robotics tasks, is a visual-language generalist with state-of-the-art performance on OK-VQA, and retains generalist language capabilities with increasing scale.

Added

2026-01-28