UniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild

Can QinShu ZhangNing YuYihao FengXinyi YangYingbo ZhouHuan WangJuan Carlos NieblesCaiming XiongSilvio Savarese

article2023NeurIPS246 citations

Presents UniControl, a unified diffusion framework that consolidates multiple condition-to-image generation tasks into a single compact model using a mixture-of-experts adapter and a task-aware HyperNet to achieve zero-shot adaptation to unseen visual conditions.

Listen

Visual generative models such as text-to-image diffusion systems have advanced rapidly, yet text prompts alone struggle to provide fine-grained geometric, spatial, and structural control. Existing solutions, notably ControlNet, enable visual conditioning (such as edge maps or depth maps) but require training and hosting a dedicated model for every specific control modality. This single-task approach creates significant compute, memory, and operational bottlenecks when deploying multi-modal control systems at scale.

The article introduces UniControl, a unified generative foundation framework designed to consolidate diverse condition-to-image tasks into a single diffusion model. The primary objective is to demonstrate that a single unified architecture can simultaneously process arbitrary natural language prompts alongside multiple distinct visual control conditions without sacrificing precision or increasing model size proportionally.

To achieve this, the authors built the MultiGen-20M dataset, containing over 20 million image-text-condition triplets across nine distinct tasks grouped into five categories: edges, region maps, human pose skeletons, 3D geometric maps, and image editing. The framework augments a pretrained Stable Diffusion model using two lightweight modules: a Mixture-of-Experts adapter (allocating roughly 70,000 parameters per task to extract low-level features) and a 12-million-parameter task-aware HyperNet that dynamically modulates the network layers based on natural language task instructions. In total, UniControl comprises about 1.44 billion parameters, compared to approximately 4.32 billion parameters required by an equivalent ensemble of nine separate task-specific models.

Experimental results show that UniControl consistently matches or outperforms specialized single-task baselines across multiple benchmarks. In quantitative evaluations measuring image quality and realism via Fréchet Inception Distance, UniControl achieved superior average scores (24.0) relative to existing single-task methods such as ControlNet (26.7) and GLIGEN. Comprehensive user studies involving over 7,000 evaluations confirmed that human raters preferred UniControl outputs over ControlNet baselines across all evaluated tasks. Crucially, the model demonstrates strong zero-shot generalization capabilities, successfully combining hybrid visual conditions (such as depth and skeleton maps simultaneously) and executing unseen visual tasks like colorization, deblurring, inpainting, and sketch-to-image conversion without task-specific training.

These findings indicate that generative vision models benefit substantially from multi-task pretraining by exploiting shared representations across visual domains. For organizations deploying generative AI, UniControl offers a roughly 67% reduction in parameters for a multi-task setup relative to maintaining nine separate ControlNet models, significantly lowering infrastructure overhead, memory footprints, and serving latency while improving output fidelity.

Stakeholders and engineering teams looking to adopt controllable visual generation should consider unifying their condition-specific pipelines into multi-task architectures rather than scaling independent models. Before deploying in safety-critical or high-fidelity production environments, organizations should conduct targeted pilot tests, curate domain-specific training data, and integrate guardrail filtering, as the underlying model can still inherit dataset biases that occasionally cause distorted human anatomy or blurred facial features. Overall, the provided evidence offers high confidence in the architectural efficiency and versatility of the unified approach.

Cover for UniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild

Abstract

Achieving machine autonomy and human control often represent divergent objectives in the design of interactive AI systems. Visual generative foundation models such as Stable Diffusion show promise in navigating these goals, especially when prompted with arbitrary languages. However, they often fall short in generating images with spatial, structural, or geometric controls. The integration of such controls, which can accommodate various visual conditions in a single unified model, remains an unaddressed challenge. In response, we introduce UniControl , a new generative foundation model that consolidates a wide array of controllable condition-to-image (C2I) tasks within a singular framework, while still allowing for arbitrary language prompts. UniControl enables pixel-level-precise image generation, where visual conditions primarily influence the generated structures and language prompts guide the style and context. To equip UniControl with the capacity to handle diverse visual conditions, we augment pretrained text-to-image diffusion models and introduce a task-aware HyperNet to modulate the diffusion models, enabling the adaptation to different C2I tasks simultaneously. Trained on nine unique C2I tasks, UniControl demonstrates impressive zero-shot generation abilities with unseen visual conditions. Experimental results show that UniControl often surpasses the performance of single-task-controlled methods of comparable model sizes. This control versatility positions UniControl as a significant advancement in the realm of controllable visual generation.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 UniControl
  • 3.1 Training Setup
  • 3.2 Model Design
  • 3.3 Task Generalization Ability
  • 4 Experiments
  • 4.1 Experiment Setup
  • 4.2 Visual Comparison
  • 4.3 Quantitative Evaluation
  • 4.4 Zero-shot Generalization
  • 5 Conclusion and Discussion
  • References
  • A Details of Implementation
  • A.1 MOE-Style Adapter
  • A.2 Task-aware HyperNet
  • A.3 Data Collection
  • B Numerical Analysis of Task-Aware Modulated ControlNet
  • C Zero-shot-task Results and Analysis
  • D Details of User Study
  • E Failure Cases
  • F Additional Results

Citation

MLA
Qin, C., et al. “UniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild”. Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 42961–92, https://proceedings.neurips.cc/paper_files/paper/2023/file/862f45ccecb2275851bc8acebb8b4d65-Paper-Conference.pdf.
APA
Qin, C., Zhang, S., Yu, N., Feng, Y., Yang, X., Zhou, Y., Wang, H., Niebles, J. C., Xiong, C., Savarese, S., Ermon, S., Fu, Y., & Xu, R. (2023). UniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild. Advances in Neural Information Processing Systems, 36, 42961–42992. https://proceedings.neurips.cc/paper_files/paper/2023/file/862f45ccecb2275851bc8acebb8b4d65-Paper-Conference.pdf
Chicago
Qin, C., S. Zhang, N. Yu, et al. 2023. “UniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild”. Advances in Neural Information Processing Systems 36: 42961–92. https://proceedings.neurips.cc/paper_files/paper/2023/file/862f45ccecb2275851bc8acebb8b4d65-Paper-Conference.pdf.
Harvard
Qin, C. et al. (2023) “UniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 42961–42992. Available at: https://proceedings.neurips.cc/paper_files/paper/2023/file/862f45ccecb2275851bc8acebb8b4d65-Paper-Conference.pdf.
Vancouver
1. Qin C, Zhang S, Yu N, et al (2023) UniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 42961–42992

BibTeX

@inproceedings{qin2023unicontrol,
  title = {UniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild},
  author = {Qin, Can and Zhang, Shu and Yu, Ning and Feng, Yihao and Yang, Xinyi and Zhou, Yingbo and Wang, Huan and Niebles, Juan Carlos and Xiong, Caiming and Savarese, Silvio and Ermon, Stefano and Fu, Yun and Xu, Ran},
  year = {2023},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {36},
  pages = {42961-42992},
  url = {https://proceedings.neurips.cc/paper_files/paper/2023/file/862f45ccecb2275851bc8acebb8b4d65-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors