MatFuse: Controllable Material Generation with Diffusion Models
Giuseppe VecchioRenato SortinoSimone PalazzoConcetto Spampinato
Proposes MatFuse, a unified latent diffusion framework that combines multimodal conditioning inputs like sketches, palettes, and text with a multi-encoder VQ-GAN architecture to enable highly controllable generation and volumetric inpainting of SVBRDF material maps.
The demand for high-quality, photorealistic 3D materials is expanding rapidly across video games, industrial simulation, and architectural visualization. However, creating these materials manually remains an expensive and time-consuming process that requires specialized technical artistry. While generative artificial intelligence has shown promise in automating visual asset creation, existing tools provide limited creative control, suffer from unstable training, or restrict generation to narrow, pre-defined material categories. The article introduces and evaluates MatFuse, a generative diffusion framework designed to synthesize and edit realistic 3D material maps using flexible, multimodal user inputs.
To address control and quality limitations, the article developed an approach combining a multi-encoder compression network with a latent diffusion model. The system was trained on a dataset of approximately 160,000 material samples and evaluated across 3.2 million renders. By separating material properties—including surface color, roughness, specularity, and geometric normals—into independent latent representations, the framework accepts multiple simultaneous prompts, such as text descriptions, reference images, color palettes, and structural sketches. The article also evaluated a targeted editing technique termed volumetric inpainting, which enables users to reconstruct specific missing maps or edit localized regions without altering the rest of the material.
The findings confirm that MatFuse delivers high-fidelity results with strong fidelity to user guidance. On image quality benchmarks, MatFuse achieved a CLIP-IQA score of 0.431, nearly matching the upper-bound ground-truth score of 0.471. It also achieved a Fréchet Inception Distance of 158.53, substantially outperforming baseline diffusion models at 231.64 and competing generative tools like TileGen at 184.81. In a user preference study with 100 participants evaluating 25 material pairs across five categories, MatFuse received 1,078 votes compared to 949 for TileGen, demonstrating a statistically significant user preference for its realism and visual quality. Ablation experiments verified that using dedicated encoders for each property map, combined with a rendering-consistency loss, reduced map reconstruction errors by more than half compared to standard single-encoder architectures.
These results indicate that multimodal diffusion models can significantly reduce production timelines and operational costs in 3D content creation workflows by giving artists intuitive, granular control over digital surfaces. However, practical deployment faces hardware constraints: synthesizing materials requires substantial graphics processing memory (around 18 to 24 gigabytes), which limits output resolution and the capture of ultra-fine surface details. Furthermore, the current implementation does not produce seamlessly tileable patterns for large digital surfaces. Teams seeking to adopt this technology should conduct pilot implementations in lower-resolution prototyping pipelines while directing future development toward patch-based architectures for higher resolutions, automated texture estimation from real photographs, and seamless tiling capabilities.
- Paper: High-Resolution Image Synthesis with Latent Diffusion Models, Robin Rombach et al. (2022). MatFuse directly builds upon the Latent Diffusion Model architecture to compress high-dimensional visual maps and execute generative diffusion within a computationally tractable latent space.
- Paper: MatSynth: A Modern PBR Materials Dataset, Giuseppe Vecchio et al. (2024). MatSynth establishes the modern benchmark and large-scale dataset framework for physically based rendering material acquisition and generation that informs PBR diffusion workflows like MatFuse.
- Paper: Adding Conditional Control to Text-to-Image Diffusion Models, Lvmin Zhang et al. (2023). ControlNet introduced the foundational multi-condition encoding mechanism for spatial and structural guidance in diffusion models that MatFuse adapts for multimodal material map generation.
- Paper: Classifier-Free Diffusion Guidance, Jonathan Ho et al. (2022). This paper establishes classifier-free guidance, the core conditioning mechanism used by multimodal diffusion frameworks such as MatFuse to balance input prompt fidelity with generative realism.
- Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). This seminal work establishes the foundational denoising diffusion probabilistic formulation and reverse-step generative training objectives underlying MatFuse.
- Paper: T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion Models, Chong Mou et al. (2023). T2I-Adapter introduces lightweight, multi-input structural and color conditioning adapters for diffusion models, which informs MatFuse's multi-encoder architecture for multimodal prompt fusion.
- Paper: SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations, Chenlin Meng et al. (2022). SDEdit provides the theoretical and algorithmic basis for noise-guided inpainting and user-directed stroke editing adapted in MatFuse's localized volumetric editing pipeline.
- Paper: Paint-it: Text-to-Texture Synthesis via Deep Convolutional Texture Map Optimization and Physically-Based Rendering, Kim Youwang et al. (2024). Paint-it builds on text-driven PBR texture generation principles like those in MatFuse by applying deep convolutional neural parameterization and differentiable rendering to synthesize PBR maps directly onto untextured 3D meshes.
- Paper: DemoFusion: Democratising High-Resolution Image Generation With No $, Ruoyi Du et al. (2024). DemoFusion extends latent diffusion generation by offering a training-free, patch-based progressive upscaling framework that addresses the high-resolution hardware bottlenecks highlighted in MatFuse.
- Paper: Masked Autoencoders Are Effective Tokenizers for Diffusion Models, Hao Chen et al. (2025). This work explores advanced masked autoencoders as tokenizers for diffusion latent spaces, providing an avenue to improve the multi-encoder latent representations used to encode property maps in MatFuse.
