Built independently by an author, for readers. Read the story and support ChapterPal

keyword

pre-trained stable diffusion

Pre-trained Stable Diffusion refers to an instance of the Stable Diffusion latent text-to-image generative model whose neural network parameters have already been optimized on massive datasets of paired images and textual captions. Operating within a compressed latent representation space, this pre-trained model captures extensive visual and semantic priors, allowing it to generate high-resolution images from natural language prompts without requiring initial training from scratch. In artificial intelligence research and downstream applications, it commonly serves as a foundational generative backbone that can be deployed directly for image generation, fine-tuned on specialized domains, or augmented with additional conditioning mechanisms to support tasks such as spatial layout control, multi-instance editing, and three-dimensional structure synthesis.

2 items

MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis

MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis

Dewei Zhou, You Li, Fan Ma, Xiaoting Zhang, Yi Yang

OrganizationsHuaweiZhejiang University

Why you should read this

Proposes a divide-and-conquer framework for Stable Diffusion that enables precise multi-instance text-to-image synthesis by shading individual instances separately through dedicated attention mechanisms before aggregating them into a unified layout.

We present a Multi-Instance Generation (MIG) task, simultaneously generating multiple instances with diverse controls in one image. Given a set of predefined coordinates and their corresponding descriptions, the task is to ensure that generated instances are accurately at the designated locations and that all instances' attributes adhere to their corresponding description. This broadens the scope of current research on Single-instance generation, elevating it to a more versatile and practical dimension. Inspired by the idea of divide and conquer, we introduce an innovative approach named Multi-Instance Generation Controller (MIGC) to address the challenges of the MIG task. Initially, we break down the MIG task into several subtasks, each involving the shading of a single instance. To ensure precise shading for each instance, we introduce an instance enhancement attention mechanism. Lastly, we aggregate all the shaded instances to provide the necessary information for accurately generating multiple instances in stable diffusion (SD). To evaluate how well generation models perform on the MIG task, we provide a COCO-MIG benchmark along with an evaluation pipeline. Extensive experiments were conducted on the proposed COCO-MIG benchmark, as well as on various commonly used benchmarks. The evaluation results illustrate the exceptional control capabilities of our model in terms of quantity, position, attribute, and interaction. Code and demos will be released at https://migcproject.github.io/.

Added

2026-09-26

Wonder3D: Single Image to 3D Using Cross-Domain Diffusion

Wonder3D: Single Image to 3D Using Cross-Domain Diffusion

Xiaoxiao Long, Yuan-Chen Guo, Cheng Lin, Yuan Liu, Zhiyang Dou, Lingjie Liu, Yuexin Ma, Song-Hai Zhang, Marc Habermann, Christian Theobalt, Wenping Wang

OrganizationsMax Planck Institute for InformaticsShanghaiTech UniversityTexas A&M UniversityTsinghua UniversityUniversity of Hong KongUniversity of PennsylvaniaVAST

Why you should read this

Proposes a cross-domain diffusion framework that jointly generates consistent multi-view normal maps and color images to extract detailed, high-fidelity 3D meshes from a single image in just two to three minutes.

In this work, we introduce Wonder3D, a novel method for efficiently generating high-fidelity textured meshes from single-view images. Recent methods based on Score Distillation Sampling (SDS) have shown the potential to recover 3D geometry from 2D diffusion priors, but they typically suffer from time-consuming per-shape optimization and inconsistent geometry. In contrast, certain works directly produce 3D information via fast network inferences, but their results are often of low quality and lack geometric details. To holistically improve the quality, consistency, and efficiency of single-view reconstruction tasks, we propose a cross-domain diffusion model that generates multi-view normal maps and the corresponding color images. To ensure the consistency of generation, we employ a multi-view cross-domain attention mechanism that facilitates information exchange across views and modalities. Lastly, we introduce a geometry-aware normal fusion algorithm that extracts high-quality surfaces from the multi-view 2D representations in only 2 ∼ 3 minutes. Our extensive evaluations demonstrate that our method achieves high-quality reconstruction results, robust generalization, and good efficiency compared to prior works.

Added

2026-09-26