DemoFusion: Democratising High-Resolution Image Generation With No $
Ruoyi DuDongliang ChangTimothy M. HospedalesYi-Zhe SongZhanyu Ma
Proposes DemoFusion, a training-free framework that enables open-source latent diffusion models like SDXL to generate high-resolution images up to 4096×4096 on consumer-grade GPUs by combining progressive upscaling, skip residuals, and dilated sampling.
Generating high-resolution visuals using Generative Artificial Intelligence currently requires immense capital investments in high-end hardware, extensive datasets, and energy. Because training costs escalate rapidly with image resolution, state-of-the-art image synthesis is increasingly centralized within well-funded corporations and placed behind commercial paywalls. Open-source models remain constrained to standard native resolutions, creating a significant barrier for individual creators and academic researchers seeking higher-quality outputs.
The article demonstrates that existing open-source latent diffusion models already possess untapped prior knowledge capable of producing high-resolution outputs without any additional training or fine-tuning. It presents a novel framework called DemoFusion, which scales image synthesis to four times, sixteen times, or higher resolutions while running entirely on a single consumer-grade graphics processing unit.
To overcome the structural distortions and repetitive artifacts common in patch-based image generation, the framework modifies the inference procedure through three combined mechanisms: progressive upscaling, skip residuals, and dilated sampling. The evaluation tested the system against existing baselines, including standard diffusion models, super-resolution enhancement, and concurrent training-free methods, across 1,000 text prompts from a standard open dataset.
The findings show that DemoFusion achieves the strongest overall performance across standard image quality, diversity, and text-alignment benchmarks. Crop-based evaluations confirm that it generates substantially richer, authentic local details compared to super-resolution upscaling, which merely smooths low-resolution inputs without creating new fine features. Furthermore, it successfully preserves global semantic coherence, avoiding the duplicated limbs and distorted geometries observed in other patch-based and kernel-dilated generation techniques.
These results demonstrate that organizations and researchers can produce ultra-high-resolution imagery without expensive infrastructure investments, retraining costs, or reliance on proprietary cloud services. The primary operational trade-off is runtime; generating higher-resolution images takes longer because of the progressive passes. However, the system generates fast, low-resolution intermediate previews within seconds, enabling users to rapidly refine prompts before committing to full high-resolution rendering.
Decision-makers should consider adopting this plug-and-play approach to reduce computational budgets in high-resolution image workflows. Future work should focus on training specialized base models tailored for progressive patch fusion to further mitigate edge-case limitations, such as occasional background object repetitions or local artifacts in sharp close-up compositions.
- Paper: High-Resolution Image Synthesis with Latent Diffusion Models, Robin Rombach et al. (2022). Introduces Latent Diffusion Models (LDMs), the foundational architecture and generative prior that DemoFusion directly adapts for training-free high-resolution image synthesis.
- Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). Establishes the core formulation of Denoising Diffusion Probabilistic Models (DDPM) upon which all modern latent and progressive diffusion frameworks are built.
- Paper: Cascaded Diffusion Models for High Fidelity Image Generation, Jonathan Ho et al. (2021). Demonstrates cascaded and multi-stage upscaling architectures for diffusion models, establishing the progressive generation paradigm adapted by DemoFusion's inference pipeline.
- Paper: Denoising Diffusion Implicit Models, Jiaming Song et al. (2021). Presents Denoising Diffusion Implicit Models (DDIM), which provide the deterministic reverse sampling trajectory essential for DemoFusion's skip residuals and progressive upscaling.
- Paper: Image Super-Resolution via Iterative Refinement, Chitwan Saharia et al. (2021). Pioneers iterative refinement super-resolution using diffusion models, providing the core baseline context for why standard super-resolution smooths rather than generates authentic local details.
- Paper: Super-resolution from a single image, Daniel Glasner et al. (2009). Provides the foundational visual computing insight regarding internal patch redundancy across scales that motivates patch-based upscaling and dilated sampling strategies.
- Paper: Taming Transformers for High-Resolution Image Synthesis, Patrick Esser et al. (2020). Introduces high-resolution two-stage image generation via discrete latent representations, establishing the underlying perceptual compression principles used in modern diffusion models.
- Paper: Scaling Rectified Flow Transformers for High-Resolution Image Synthesis, Patrick Esser et al. (2024). Scales high-resolution generation via rectified flow transformers and multimodal attention (MM-DiT), offering a trained architectural alternative to training-free inference upscaling.
- Paper: SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis, Dustin Podell et al. (2024). Improves native latent diffusion at larger resolutions through multi-stage refinement and conditioning, providing a powerful base model suitable for progressive patch fusion methods like DemoFusion.
- Paper: FiT: Flexible Vision Transformer for Diffusion Model, Zeyu Lu et al. (2024). Extends diffusion transformers to arbitrary aspect ratios and unrestricted resolutions using dynamic visual token sequences, addressing structural constraints in high-resolution image synthesis.
- Paper: DiffEditor: Boosting Accuracy and Flexibility on Diffusion-Based Image Editing, Chong Mou et al. (2024). Applies diffusion sampling control and localized gradient guidance to precise image manipulation and editing tasks.
- Paper: Fast ODE-based Sampling for Diffusion Models in Around 5 Steps, Zhenyu Zhou et al. (2024). Develops few-step ODE solvers for diffusion sampling trajectories, directly addressing the inference latency and computational trade-offs highlighted in progressive multi-pass generation.
- Paper: UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics, Xi Chen et al. (2025). Unifies high-resolution multi-image generation and instruction editing into a single large-scale diffusion transformer trained across progressive resolutions.
- Paper: Masked Autoencoders Are Effective Tokenizers for Diffusion Models, Hao Chen et al. (2025). Analyzes and redesigns the latent tokenizer spaces of diffusion models, directly impacting reconstruction fidelity and generative quality in high-resolution synthesis.
