Built independently by an author, for readers. Read the story and support ChapterPal

keyword

task-aware vision-language condition

A task-aware vision-language condition is a multimodal guiding input in visual generation systems that combines structural visual inputs, descriptive text prompts, and explicit task instructions to control image synthesis. In this framework, the visual condition provides spatial or geometric guidance, such as edge maps, depth layouts, or pose skeletons, while the natural language prompt directs the semantic content, context, and stylistic appearance of the generated output. The task-aware component explicitly specifies the particular control or transformation task being performed, enabling a single unified model to dynamically adapt its internal processing across diverse condition-to-image tasks without requiring separate models for each control modality.

1 item

UniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild

UniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild

Can Qin, Shu Zhang, Ning Yu, Yihao Feng, Xinyi Yang, Yingbo Zhou, Huan Wang, Juan Carlos Niebles, Caiming Xiong, Silvio Savarese, Stefano Ermon, Yun Fu, Ran Xu

OrganizationsNortheastern UniversitySalesforceStanford University

Why you should read this

Presents UniControl, a unified diffusion framework that consolidates multiple condition-to-image generation tasks into a single compact model using a mixture-of-experts adapter and a task-aware HyperNet to achieve zero-shot adaptation to unseen visual conditions.

Achieving machine autonomy and human control often represent divergent objectives in the design of interactive AI systems. Visual generative foundation models such as Stable Diffusion show promise in navigating these goals, especially when prompted with arbitrary languages. However, they often fall short in generating images with spatial, structural, or geometric controls. The integration of such controls, which can accommodate various visual conditions in a single unified model, remains an unaddressed challenge. In response, we introduce UniControl , a new generative foundation model that consolidates a wide array of controllable condition-to-image (C2I) tasks within a singular framework, while still allowing for arbitrary language prompts. UniControl enables pixel-level-precise image generation, where visual conditions primarily influence the generated structures and language prompts guide the style and context. To equip UniControl with the capacity to handle diverse visual conditions, we augment pretrained text-to-image diffusion models and introduce a task-aware HyperNet to modulate the diffusion models, enabling the adaptation to different C2I tasks simultaneously. Trained on nine unique C2I tasks, UniControl demonstrates impressive zero-shot generation abilities with unseen visual conditions. Experimental results show that UniControl often surpasses the performance of single-task-controlled methods of comparable model sizes. This control versatility positions UniControl as a significant advancement in the realm of controllable visual generation.

Added

2026-09-26