An image prompt adapter is a lightweight neural network module that enables pretrained text-to-image diffusion models to use visual inputs as conditioning prompts alongside or instead of text, without retraining or altering the underlying base model. By extracting visual features from a reference image and injecting them into the generation process through dedicated attention mechanisms, it allows users to guide attributes such as subject identity, composition, and artistic style. Because the core generative model remains frozen, the adapter provides a computationally efficient and modular way to achieve multimodal generation while preserving full compatibility with existing text prompts, customized checkpoints, and structural control tools.