DemoFusion: Democratising High-Resolution Image Generation With No $

Ruoyi DuDongliang ChangTimothy M. HospedalesYi-Zhe SongZhanyu Ma

article2024CVPR127 citations

Proposes DemoFusion, a training-free framework that enables open-source latent diffusion models like SDXL to generate high-resolution images up to 4096×4096 on consumer-grade GPUs by combining progressive upscaling, skip residuals, and dilated sampling.

Listen

Generating high-resolution visuals using Generative Artificial Intelligence currently requires immense capital investments in high-end hardware, extensive datasets, and energy. Because training costs escalate rapidly with image resolution, state-of-the-art image synthesis is increasingly centralized within well-funded corporations and placed behind commercial paywalls. Open-source models remain constrained to standard native resolutions, creating a significant barrier for individual creators and academic researchers seeking higher-quality outputs.

The article demonstrates that existing open-source latent diffusion models already possess untapped prior knowledge capable of producing high-resolution outputs without any additional training or fine-tuning. It presents a novel framework called DemoFusion, which scales image synthesis to four times, sixteen times, or higher resolutions while running entirely on a single consumer-grade graphics processing unit.

To overcome the structural distortions and repetitive artifacts common in patch-based image generation, the framework modifies the inference procedure through three combined mechanisms: progressive upscaling, skip residuals, and dilated sampling. The evaluation tested the system against existing baselines, including standard diffusion models, super-resolution enhancement, and concurrent training-free methods, across 1,000 text prompts from a standard open dataset.

The findings show that DemoFusion achieves the strongest overall performance across standard image quality, diversity, and text-alignment benchmarks. Crop-based evaluations confirm that it generates substantially richer, authentic local details compared to super-resolution upscaling, which merely smooths low-resolution inputs without creating new fine features. Furthermore, it successfully preserves global semantic coherence, avoiding the duplicated limbs and distorted geometries observed in other patch-based and kernel-dilated generation techniques.

These results demonstrate that organizations and researchers can produce ultra-high-resolution imagery without expensive infrastructure investments, retraining costs, or reliance on proprietary cloud services. The primary operational trade-off is runtime; generating higher-resolution images takes longer because of the progressive passes. However, the system generates fast, low-resolution intermediate previews within seconds, enabling users to rapidly refine prompts before committing to full high-resolution rendering.

Decision-makers should consider adopting this plug-and-play approach to reduce computational budgets in high-resolution image workflows. Future work should focus on training specialized base models tailored for progressive patch fusion to further mitigate edge-case limitations, such as occasional background object repetitions or local artifacts in sharp close-up compositions.

  • Paper: High-Resolution Image Synthesis with Latent Diffusion Models, Robin Rombach et al. (2022). Introduces Latent Diffusion Models (LDMs), the foundational architecture and generative prior that DemoFusion directly adapts for training-free high-resolution image synthesis.
  • Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). Establishes the core formulation of Denoising Diffusion Probabilistic Models (DDPM) upon which all modern latent and progressive diffusion frameworks are built.
  • Paper: Cascaded Diffusion Models for High Fidelity Image Generation, Jonathan Ho et al. (2021). Demonstrates cascaded and multi-stage upscaling architectures for diffusion models, establishing the progressive generation paradigm adapted by DemoFusion's inference pipeline.
  • Paper: Denoising Diffusion Implicit Models, Jiaming Song et al. (2021). Presents Denoising Diffusion Implicit Models (DDIM), which provide the deterministic reverse sampling trajectory essential for DemoFusion's skip residuals and progressive upscaling.
  • Paper: Image Super-Resolution via Iterative Refinement, Chitwan Saharia et al. (2021). Pioneers iterative refinement super-resolution using diffusion models, providing the core baseline context for why standard super-resolution smooths rather than generates authentic local details.
  • Paper: Super-resolution from a single image, Daniel Glasner et al. (2009). Provides the foundational visual computing insight regarding internal patch redundancy across scales that motivates patch-based upscaling and dilated sampling strategies.
  • Paper: Taming Transformers for High-Resolution Image Synthesis, Patrick Esser et al. (2020). Introduces high-resolution two-stage image generation via discrete latent representations, establishing the underlying perceptual compression principles used in modern diffusion models.
Cover for DemoFusion: Democratising High-Resolution Image Generation With No $

Abstract

High-resolution image generation with Generative Artificial Intelligence (GenAI) has immense potential but, due to the enormous capital investment required for training, it is increasingly centralised to a few large corporations, and hidden behind paywalls. This paper aims to democratise high-resolution GenAI by advancing the frontier of high-resolution generation while remaining accessible to a broad audience. We demonstrate that existing Latent Diffusion Models (LDMs) possess untapped potential for higher-resolution image generation. Our novel DemoFusion framework seamlessly extends open-source GenAI models, employing Progressive Upscaling, Skip Residual, and Dilated Sampling mechanisms to achieve higher-resolution image generation. The progressive nature of DemoFusion requires more passes, but the intermediate results can serve as “previews”, facilitating rapid prompt iteration.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Methodology
  • 3.1. Preliminaries
  • 3.2. Progressive Upscaling
  • 3.3. Skip Residual
  • 3.4. Dilated Sampling
  • 4. Experiments
  • 4.1. Comparison
  • 4.2. Ablation Study
  • 5. Limitations and Opportunities
  • 6. Conclusion
  • Acknowledgement
  • References

Knowls

  1. Knowl 1 — DemoFusion High-Resolution Image Generation Framework

    model/method

    DemoFusion is a tuning-free, plug-and-play framework that extends pre-trained Latent Diffusion Models (LDMs, such as Stable Diffusion XL) to synthesize high-resolution images (e.g., 4×4\times, 16×16\times, or higher area scaling, reaching 2048×20482048\times 2048 or 4096×40964096\times 4096 pixels) on consumer hardware without additional training.

    Directly prompting LDMs at high resolutions or generating non-overlapping independent patches results in structural collapse, while standard patch-fusion methods (such as MultiDiffusion) produce repetitive patterns due to a lack of global semantic context. DemoFusion resolves this by exploiting the prior knowledge of cropped scenes already present in base LDMs through three coordinated inference mechanisms: Progressive Upscaling across scale stages, Skip Residual guidance across diffusion steps, and Dilated Sampling in the latent space.

  2. Knowl 2 — Progressive Upscaling in Latent Diffusion Space

    equation

    To generate an image scaled by an area factor KK using an LDM with base latent dimensions c×h×wc \times h \times w, DemoFusion sets the side-length scaling factor to S=KS = \sqrt{K}, aiming for a target latent dimension c×H×Wc \times H \times W where H=ShH = S h and W=SwW = S w. Rather than generating the target resolution in a single pass, the process is divided into SS progressive phases s∈{1,2,…,S}s \in \{1, 2, \dots, S\}.

    The full progressive generation distribution is formulated as:

    pθ(z0S∣zT1)=pθ(z01∣zT1)∏s=2S(q(zT′s∣z0′s)pθ(z0s∣zT′s))p_\theta(\mathbf{z}^S_0 | \mathbf{z}^1_T) = p_\theta(\mathbf{z}^1_0 | \mathbf{z}^1_T) \prod_{s=2}^S \left( q(\mathbf{z}'^s_T | \mathbf{z}'^s_0) p_\theta(\mathbf{z}^s_0 | \mathbf{z}'^s_T) \right)

    where:

    • z01∈Rc×h×w\mathbf{z}^1_0 \in \mathbb{R}^{c \times h \times w} is synthesized during the initial phase s=1s=1 from Gaussian noise zT1∼N(0,I)\mathbf{z}^1_T \sim \mathcal{N}(\mathbf{0}, \mathbf{I}) using the standard reverse diffusion process pθ(z01∣zT1)=∏t=1Tpθ(zt−11∣zt1)p_\theta(\mathbf{z}^1_0 | \mathbf{z}^1_T) = \prod_{t=1}^T p_\theta(\mathbf{z}^1_{t-1} | \mathbf{z}^1_t).
    • For each subsequent phase s∈{2,…,S}s \in \{2, \dots, S\}, the cleaned latent from the preceding phase z0s−1\mathbf{z}^{s-1}_0 is upscaled via spatial interpolation (e.g., bicubic interpolation) z0′s=inter(z0s−1)∈Rc×sh×sw\mathbf{z}'^s_0 = \text{inter}(\mathbf{z}^{s-1}_0) \in \mathbb{R}^{c \times s h \times s w}.
    • q(zT′s∣z0′s)=∏t=1Tq(zt′s∣zt−1′s)q(\mathbf{z}'^s_T | \mathbf{z}'^s_0) = \prod_{t=1}^T q(\mathbf{z}'^s_t | \mathbf{z}'^s_{t-1}) is the forward diffusion process adding Gaussian noise to step TT.
    • pθ(z0s∣zT′s)=∏t=1Tpθ(zt−1s∣zts)p_\theta(\mathbf{z}^s_0 | \mathbf{z}'^s_T) = \prod_{t=1}^T p_\theta(\mathbf{z}^s_{t-1} | \mathbf{z}^s_t) is the patch-fused reverse denoising process at scale ss producing the refined latent z0s\mathbf{z}^s_0.
  3. Knowl 3 — Skip Residual Guidance Formulation

    equation

    To preserve global semantic coherence from the lower-resolution representation z0′s=inter(z0s−1)\mathbf{z}'^s_0 = \text{inter}(\mathbf{z}^{s-1}_0) and avoid information loss or upsampling noise artifacts during phase ss, DemoFusion modifies the reverse step pθ(zt−1s∣zts)p_\theta(\mathbf{z}^s_{t-1} | \mathbf{z}^s_t) to condition on a fused latent z^ts\hat{\mathbf{z}}^s_t:

    z^ts=c1×zt′s+(1−c1)×zts\hat{\mathbf{z}}^s_t = c_1 \times \mathbf{z}'^s_t + (1 - c_1) \times \mathbf{z}^s_t

    where:

    • zt′s\mathbf{z}'^s_t is the forward-diffused (noise-inversed) latent obtained from z0′s\mathbf{z}'^s_0 at time-step t∈[1,T]t \in [1, T].
    • zts\mathbf{z}^s_t is the current intermediate denoising latent at step tt.
    • c1c_1 is a time-dependent scaled cosine decay factor:

    c1=(1+cos⁡(T−tTπ)2)α1c_1 = \left( \frac{1 + \cos\left(\frac{T - t}{T} \pi\right)}{2} \right)^{\alpha_1}

    with scaling hyperparameter α1\alpha_1. In earlier denoising steps (t≈Tt \approx T), c1≈1c_1 \approx 1, steering the layout according to the global lower-resolution structure. In later steps (t→0t \to 0), c1→0c_1 \to 0, allowing local denoising paths to optimize fine details without constraint.

  4. Knowl 4 — Dilated Latent Sampling and Gaussian Filtering

    equation

    To provide diffusion paths with global receptive context across the full image canvas at scale ss, DemoFusion applies shifted dilated sampling to the latent representation zt∈Rc×H×W\mathbf{z}_t \in \mathbb{R}^{c \times H \times W}.

    To prevent graininess resulting from non-overlapping global paths, a Gaussian filter G(⋅)G(\cdot) with kernel size 4s−34s - 3 is applied to zt\mathbf{z}_t prior to sampling:

    Ztglobal=[z0,t,…,zm,t,…,zM,t]=Sglobal(G(zt))\mathcal{Z}_t^{global} = [\mathbf{z}_{0,t}, \dots, \mathbf{z}_{m,t}, \dots, \mathbf{z}_{M,t}] = \mathcal{S}_{global}(G(\mathbf{z}_t))

    where each zm,t∈Rc×h×w\mathbf{z}_{m,t} \in \mathbb{R}^{c \times h \times w}, the dilation stride is set to ss, and M=s2M = s^2. The standard deviation σ(t)\sigma(t) of G(⋅)G(\cdot) decays according to:

    σ(t)=c3×(σ1−σ2)+σ2,c3=(1+cos⁡(T−tTπ)2)α3\sigma(t) = c_3 \times (\sigma_1 - \sigma_2) + \sigma_2, \quad c_3 = \left( \frac{1 + \cos\left(\frac{T - t}{T} \pi\right)}{2} \right)^{\alpha_3}

    where σ1\sigma_1 and σ2\sigma_2 are the initial and final standard deviations, and α3\alpha_3 is a scaling exponent.

    After independently predicting the denoised latents zm,t−1\mathbf{z}_{m,t-1} via pθ(zm,t−1∣zm,t)p_\theta(\mathbf{z}_{m,t-1} | \mathbf{z}_{m,t}), the reconstructed global latent Rglobal(Zt−1global)\mathcal{R}_{global}(\mathcal{Z}_{t-1}^{global}) is blended with the reconstructed local overlapping patch latent Rlocal(Zt−1local)\mathcal{R}_{local}(\mathcal{Z}_{t-1}^{local}) (from stride-based crop sampling Slocal\mathcal{S}_{local}):

    zt−1=c2×Rglobal(Zt−1global)+(1−c2)×Rlocal(Zt−1local)\mathbf{z}_{t-1} = c_2 \times \mathcal{R}_{global}(\mathcal{Z}_{t-1}^{global}) + (1 - c_2) \times \mathcal{R}_{local}(\mathcal{Z}_{t-1}^{local})

    with dynamic weight:

    c2=(1+cos⁡(T−tTπ)2)α2c_2 = \left( \frac{1 + \cos\left(\frac{T - t}{T} \pi\right)}{2} \right)^{\alpha_2}

    where α2\alpha_2 is a scaling hyperparameter. This assigns dominance to global structure early in the reverse process (t≈Tt \approx T) and shifts focus to local texture refinement late in the process (t→0t \to 0).

  5. Knowl 5 — Quantitative Benchmark Evaluation on High-Resolution Synthesis

    data/table

    Quantitative performance comparison evaluated on 1K1\text{K} randomly sampled captions from LAION-5B across three generation resolutions using SDXL as the underlying generative model on an RTX 3090 GPU.

    Metrics include:

    • Global metrics: FID ↓\downarrow, Inception Score (IS) ↑\uparrow, and CLIP Score ↑\uparrow (evaluated by resizing full images to 299×299299 \times 299).
    • Crop-based local detail metrics: FIDcrop↓\text{FID}_{crop} \downarrow and IScrop↑\text{IS}_{crop} \uparrow (evaluated by cropping 1×1\times resolution local patches and resizing to 299×299299 \times 299).
    Method FID ↓\downarrow IS ↑\uparrow FIDcrop↓\text{FID}_{crop} \downarrow IScrop↑\text{IS}_{crop} \uparrow CLIP ↑\uparrow Time
    2048 ×\times 2048
    SDXL Direct Inference 79.66 13.47 73.91 17.38 28.12 1 min
    MultiDiffusion 75.93 14.56 70.93 17.85 28.97 3 min
    SDXL + BSRGAN 66.41 16.22 67.42 21.11 29.61 1 min
    SCALECRAFTER 69.91 15.72 68.36 19.44 29.51 1 min
    DemoFusion (Ours) 65.73 16.41 64.81 21.40 29.68 3 min
    2048 ×\times 4096
    SDXL Direct Inference 97.08 14.12 96.41 18.01 27.29 3 min
    MultiDiffusion 89.38 14.17 82.78 18.87 28.66 6 min
    SDXL + BSRGAN 68.70 16.29 75.03 21.76 29.01 1 min
    SCALECRAFTER 80.16 15.29 83.08 19.56 28.87 6 min
    DemoFusion (Ours) 73.15 16.37 71.35 23.55 29.05 11 min
    4096 ×\times 4096
    SDXL Direct Inference 105.65 14.01 98.59 19.47 25.64 8 min
    MultiDiffusion 97.98 13.84 79.45 19.73 28.62 15 min
    SDXL + BSRGAN 66.44 16.21 77.20 22.42 29.63 1 min
    SCALECRAFTER 87.50 15.20 84.36 20.32 29.04 19 min
    DemoFusion (Ours) 74.11 16.11 70.34 24.28 29.57 25 min

    While SDXL+BSRGAN attains competitive full-image FID/IS because BSRGAN adheres strictly to low-resolution input geometry, DemoFusion significantly outperforms all baselines on FIDcrop\text{FID}_{crop} and IScrop\text{IS}_{crop} (e.g., reaching FIDcrop=70.34\text{FID}_{crop} = 70.34 and IScrop=24.28\text{IS}_{crop} = 24.28 at 4096×40964096\times 4096), demonstrating superior synthesis of authentic high-resolution local details.

  6. Knowl 6 — Complementary Interactions of DemoFusion Core Components

    empirical result

    Ablation over all combinations of Progressive Upscaling (PU), Skip Residual (SR), and Dilated Sampling (DS) at 9×9\times resolution (3072×30723072\times 3072) reveals strong mutual dependencies:

    1. Baseline failure: Removing all three components reduces the system to naive patch-based diffusion, producing extreme repetition and loss of semantic coherence.
    2. Skip Residuals (SR): Introduces low-resolution structural guidance, eliminating gross structural duplications and maintaining scene-level semantic layout.
    3. Dilated Sampling (DS): Establishes global receptive paths to align local denoising directions towards the global optimum. However, when used alone, independent global paths introduce grainy textures and amplify interpolation artifacts.
    4. Progressive Upscaling (PU): Mitigates artificial upsampling noise by scaling in gradual intermediate steps rather than single-step large-scale upsampling.

    Together, SR suppresses graininess from DS, and PU removes the large-scale interpolation noise amplified by DS, making all three techniques jointly necessary to achieve coherent global structure and sharp local details.

  7. Knowl 7 — Generalization Across Base Latent Diffusion Models

    empirical result

    DemoFusion functions as a model-agnostic, tuning-free inference wrapper across different base LDMs:

    • Stable Diffusion 1.5 (native resolution 512×512512\times 512): DemoFusion extends generation up to 9×9\times resolution (1536×15361536\times 1536).
    • Stable Diffusion 2.1 (native resolution 768×768768\times 768): DemoFusion extends generation up to 9×9\times resolution (2304×23042304\times 2304).
    • SDXL (native resolution 1024×10241024\times 1024): DemoFusion extends generation to 4×4\times (2048×20482048\times 2048), 9×9\times (3072×30723072\times 3072), 16×16\times (4096×40964096\times 4096), and higher.

    While DemoFusion successfully generates coherent high-resolution outputs across all models without retraining, visual fidelity and detail quality scale proportionally with the representation power and cropping priors of the underlying base model, producing the best results with SDXL.

  8. Knowl 8 — Computational and Structural Limitations of DemoFusion

    limitation

    DemoFusion is subject to several practical and structural limitations:

    1. Inference Latency: Because patch-based generation requires overlapping local passes and progressive multi-phase upscaling, inference time scales significantly with resolution (e.g., 25 minutes for 4096×40964096\times 4096 on a single RTX 3090, compared to 1 minute for single-pass SDXL or SDXL+BSRGAN).
    2. Reliance on Training Priors: The framework relies entirely on the base LDM's inherent exposure to cropped images during pre-training. When synthesizing sharp close-up images, the model can occasionally hallucinate irrational local anatomy or structures.
    3. Background Repetition: Small, repetitive object patterns can still occasionally appear in sparse or uniform background regions.
    4. Model Correlation: As a tuning-free approach, generative visual quality is fundamentally bounded by the base generative capability of the underlying LDM.

Coverage note — None was omitted; all key algorithmic formulations (Progressive Upscaling, Skip Residuals, Dilated Sampling with Gaussian filtering), full quantitative experimental benchmarks, component ablations, cross-model evaluations, and stated limitations are fully covered.

References

  1. 1.Stability AI. Stable diffusion: A latent text-to-image diffusion model. https://stability.ai/blog/stable-diffusion-public-release, 2022. 1, 2
  2. 2.Omer Bar-Tal, Lior Yariv, Yaron Lipman, and Tali Dekel. Multidiffusion: Fusing diffusion paths for controlled image generation. In ICML, 2023. 2, 3, 5, 7
  3. 3.Tim Brooks, Aleksander Holynski, and Alexei A Efros. Instructpix2pix: Learning to follow image editing instructions. In CVPR, 2023. 3
  4. 4.Lucy Chai, Michael Gharbi, Eli Shechtman, Phillip Isola, and Richard Zhang. Any-resolution training for high-resolution image synthesis. In ECCV, 2022. 7
  5. 5.Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. In NeurIPS, 2021. 3
  6. 6.Xiao Han, Yukang Cao, Kai Han, Xiatian Zhu, Jiankang Deng, Yi-Zhe Song, Tao Xiang, and Kwan-Yee K Wong. Headsculpt: Crafting 3d head avatars with text. In NeurIPS, 2023. 3
  7. 7.Yingqing He, Shaoshu Yang, Haoxin Chen, Xiaodong Cun, Menghan Xia, Yong Zhang, Xintao Wang, Ran He, Qifeng Chen, and Ying Shan. Scalecrafter: Tuning-free higher-resolution visual generation with diffusion models. arXiv preprint arXiv:2310.07702, 2023. 3, 5, 7
  8. 8.Yingqing He, Tianyu Yang, Yong Zhang, Ying Shan, and Qifeng Chen. Latent video diffusion models for high-fidelity long video generation. arXiv preprint arXiv:2211.13221, 2023. 3
  9. 9.Amir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman, Yael Pritch, and Daniel Cohen-or. Prompt-to-prompt image editing with cross-attention control. In ICLR, 2022. 3, 4
  10. 10.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. NeurIPS, 2017. 7
  11. 11.Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P Kingma, Ben Poole, Mohammad Norouzi, David J Fleet, et al. Imagen video: High definition video generation with diffusion models. arXiv preprint arXiv:2210.02303, 2022. 3
  12. 12.Jonathan Ho, Chitwan Saharia, William Chan, David J Fleet, Mohammad Norouzi, and Tim Salimans. Cascaded diffusion models for high fidelity image generation. The Journal of Machine Learning Research, 2022. 3
  13. 13.Emiel Hoogeboom, Jonathan Heek, and Tim Salimans. simple diffusion: End-to-end diffusion for high resolution images. arXiv preprint arXiv:2301.11093, 2023. 3
  14. 14.Ajay Jain, Amber Xie, and Pieter Abbeel. Vectorfusion: Text-to-svg by abstracting pixel-based diffusion models. In CVPR, 2023. 3
  15. 15.Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen. Progressive growing of gans for improved quality, stability, and variation. In ICLR, 2018. 4
  16. 16.Yuseung Lee, Kunho Kim, Hyunjin Kim, and Minhyuk Sung. Syncdiffusion: Coherent montage via synchronized joint diffusions. arXiv preprint arXiv:2306.05178, 2023. 3
  17. 17.Tingting Liao, Hongwei Yi, Yuliang Xiu, Jiaxaing Tang, Yangyi Huang, Justus Thies, and Michael J Black. Tada! text to animatable digital avatars. arXiv preprint arXiv:2308.10899, 2023. 3
  18. 18.Chen-Hsuan Lin, Jun Gao, Luming Tang, Towaki Takikawa, Xiaohui Zeng, Xun Huang, Karsten Kreis, Sanja Fidler, Ming-Yu Liu, and Tsung-Yi Lin. Magic3d: High-resolution text-to-3d content creation. In CVPR, 2023. 3
  19. 19.MidJourney. Midjourney: An independent research lab. https://www.midjourney.com/, 2022. 1, 2
  20. 20.Ron Mokady, Amir Hertz, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or. Null-text inversion for editing real images using guided diffusion models. In CVPR, 2023. 3, 4
  21. 21.Chong Mou, Xintao Wang, Liangbin Xie, Jian Zhang, Zhongang Qi, Ying Shan, and Xiaohu Qie. T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models. arXiv preprint arXiv:2302.08453, 2023. 3
  22. 22.Kam Woh Ng, Xiatian Zhu, Yi-Zhe Song, and Tao Xiang. Dreamcreature: Crafting photorealistic virtual creatures from imagination. arXiv preprint arXiv:2311.15477, 2023. 3
  23. 23.OpenAI. Dall·e: Creating images from text. https://openai.com/blog/dall-e/, 2021. 1, 2
  24. 24.Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. Sdxl: Improving latent diffusion models for high-resolution image synthesis. arXiv preprint arXiv:2307.01952, 2023. 1, 2, 3, 5, 7
  25. 25.Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Mildenhall. Dreamfusion: Text-to-3d using 2d diffusion. In ICLR, 2022. 3
  26. 26.Zhiyu Qu, Tao Xiang, and Yi-Zhe Song. Sketchdreamer: Interactive text-augmented creative sketch ideation. In BMVC, 2023. 3
  27. 27.Zhiyu Qu, Lan Yang, Honggang Zhang, Tao Xiang, Kaiyue Pang, and Yi-Zhe Song. Wired perspectives: Multi-view wire art embraces generative ai. In CVPR, 2024. 3
  28. 28.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In ICML, 2021. 7
  29. 29.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In CVPR, 2022. 3
  30. 30.Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. In CVPR, 2023. 3
  31. 31.Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen. Improved techniques for training gans. NeurIPS, 2016. 7
  32. 32.Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al. Laion-5b: An open large-scale dataset for training next generation image-text models. NeurIPS, 2022. 7
  33. 33.Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In ICML, 2015. 3
  34. 34.Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. In ICLR, 2021. 3
  35. 35.Jianyi Wang, Zongsheng Yue, Shangchen Zhou, Kelvin CK Chan, and Chen Change Loy. Exploiting diffusion prior for real-world image super-resolution. arXiv preprint arXiv:2305.07015, 2023. 3
  36. 36.Tengfei Wang, Bo Zhang, Ting Zhang, Shuyang Gu, Jianmin Bao, Tadas Baltrusaitis, Jingjing Shen, Dong Chen, Fang Wen, Qifeng Chen, et al. Rodin: A generative model for sculpting 3d digital avatars using diffusion. In CVPR, 2023. 3
  37. 37.Jay Zhangjie Wu, Yixiao Ge, Xintao Wang, Stan Weixian Lei, Yuchao Gu, Yufei Shi, Wynne Hsu, Ying Shan, Xiaohu Qie, and Mike Zheng Shou. Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation. In CVPR, 2023. 3
  38. 38.Jiale Xu, Xintao Wang, Weihao Cheng, Yan-Pei Cao, Ying Shan, Xiaohu Qie, and Shenghua Gao. Dream3d: Zero-shot text-to-3d synthesis using 3d shape prior and text-to-image diffusion models. In CVPR, 2023. 3
  39. 39.Fisher Yu and Koltun Vladlen. Multi-scale context aggregation by dilated convolutions. In ICLR, 2016. 5
  40. 40.Kai Zhang, Jingyun Liang, Luc Van Gool, and Radu Timofte. Designing a practical degradation model for deep blind image super-resolution. In CVPR, 2021. 3, 5, 7
  41. 41.Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. In ICCV, 2023. 3
  42. 42.Qingping Zheng, Yuanfan Guo, Jiankang Deng, Jianhua Han, Ying Li, Songcen Xu, and Hang Xu. Any-size-diffusion: Toward efficient text-driven synthesis for any-size hd images. arXiv preprint arXiv:2308.16582, 2023. 2, 3

Citation

MLA
Du, R., et al. “DemoFusion: Democratising High-Resolution Image Generation With No $$$”. arXiv, 2023, http://arxiv.org/abs/2311.16973v2.
APA
Du, R., Chang, D., Hospedales, T., Song, Y.-Z., & Ma, Z. (2023). DemoFusion: Democratising High-Resolution Image Generation With No $$$. arXiv. http://arxiv.org/abs/2311.16973v2
Chicago
Du, R., D. Chang, T. Hospedales, Y.-Z. Song, and Z. Ma. 2023. “DemoFusion: Democratising High-Resolution Image Generation With No $$$”. arXiv. http://arxiv.org/abs/2311.16973v2.
Harvard
Du, R. et al. (2023) “DemoFusion: Democratising High-Resolution Image Generation With No $$$”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2311.16973v2.
Vancouver
1. Du R, Chang D, Hospedales T, Song Y-Z, Ma Z (2023) DemoFusion: Democratising High-Resolution Image Generation With No $$$. arXiv

BibTeX

@article{du2023demofusion,
  title = {DemoFusion: Democratising High-Resolution Image Generation With No $$$},
  author = {Du, Ruoyi and Chang, Dongliang and Hospedales, Timothy and Song, Yi-Zhe and Ma, Zhanyu},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2311.16973v2},
  eprint = {2311.16973}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE