MatFuse: Controllable Material Generation with Diffusion Models

Giuseppe VecchioRenato SortinoSimone PalazzoConcetto Spampinato

article2024CVPR53 citations

Proposes MatFuse, a unified latent diffusion framework that combines multimodal conditioning inputs like sketches, palettes, and text with a multi-encoder VQ-GAN architecture to enable highly controllable generation and volumetric inpainting of SVBRDF material maps.

Listen

The demand for high-quality, photorealistic 3D materials is expanding rapidly across video games, industrial simulation, and architectural visualization. However, creating these materials manually remains an expensive and time-consuming process that requires specialized technical artistry. While generative artificial intelligence has shown promise in automating visual asset creation, existing tools provide limited creative control, suffer from unstable training, or restrict generation to narrow, pre-defined material categories. The article introduces and evaluates MatFuse, a generative diffusion framework designed to synthesize and edit realistic 3D material maps using flexible, multimodal user inputs.

To address control and quality limitations, the article developed an approach combining a multi-encoder compression network with a latent diffusion model. The system was trained on a dataset of approximately 160,000 material samples and evaluated across 3.2 million renders. By separating material properties—including surface color, roughness, specularity, and geometric normals—into independent latent representations, the framework accepts multiple simultaneous prompts, such as text descriptions, reference images, color palettes, and structural sketches. The article also evaluated a targeted editing technique termed volumetric inpainting, which enables users to reconstruct specific missing maps or edit localized regions without altering the rest of the material.

The findings confirm that MatFuse delivers high-fidelity results with strong fidelity to user guidance. On image quality benchmarks, MatFuse achieved a CLIP-IQA score of 0.431, nearly matching the upper-bound ground-truth score of 0.471. It also achieved a Fréchet Inception Distance of 158.53, substantially outperforming baseline diffusion models at 231.64 and competing generative tools like TileGen at 184.81. In a user preference study with 100 participants evaluating 25 material pairs across five categories, MatFuse received 1,078 votes compared to 949 for TileGen, demonstrating a statistically significant user preference for its realism and visual quality. Ablation experiments verified that using dedicated encoders for each property map, combined with a rendering-consistency loss, reduced map reconstruction errors by more than half compared to standard single-encoder architectures.

These results indicate that multimodal diffusion models can significantly reduce production timelines and operational costs in 3D content creation workflows by giving artists intuitive, granular control over digital surfaces. However, practical deployment faces hardware constraints: synthesizing materials requires substantial graphics processing memory (around 18 to 24 gigabytes), which limits output resolution and the capture of ultra-fine surface details. Furthermore, the current implementation does not produce seamlessly tileable patterns for large digital surfaces. Teams seeking to adopt this technology should conduct pilot implementations in lower-resolution prototyping pipelines while directing future development toward patch-based architectures for higher resolutions, automated texture estimation from real photographs, and seamless tiling capabilities.

arXiv: 2308.11408
Cover for MatFuse: Controllable Material Generation with Diffusion Models

Abstract

Creating high-quality materials in computer graphics is a challenging and time-consuming task, which requires great expertise. To simplify this process, we introduce MatFuse, a unified approach that harnesses the generative power of diffusion models for creation and editing of 3D materials. Our method integrates multiple sources of conditioning, including color palettes, sketches, text, and pictures, enhancing creative possibilities and granting fine-grained control over material synthesis. Additionally, MatFuse enables map-level material editing capabilities through latent manipulation by means of a multi-encoder compression model which learns a disentangled latent representation for each map. We demonstrate the effectiveness of MatFuse under multiple conditioning settings and explore the potential of material editing. Finally, we assess the quality of the generated materials both quantitatively in terms of CLIP-IQA and FID scores and qualitatively by conducting a user study. Source code for training MatFuse and supplemental materials are publicly available at https://gvecchio.com/matfuse.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. MatFuse Architecture
  • 3.1. Latent Diffusion Model
  • 3.2. Conditioning Mechanisms
  • 3.3. Material Editing via Volumetric Inpainting
  • 4. Experimental Results
  • 4.1. Datasets
  • 4.2. Training Procedure
  • 4.3. Generation Results
  • 4.3.1 Single-conditional Generation
  • 4.3.2 Multi-conditional Generation
  • 4.4. Quantitative Evaluation
  • 4.5. Qualitative Comparison and User Study
  • 4.6. Material Editing Results
  • 4.7. Ablation Study
  • 5. Limitations and Future Work
  • 6. Conclusion
  • 7. Acknowledgments
  • References

Knowls

  1. Knowl 1 — Unified latent-diffusion model for SVBRDF generation

    model/method

    MatFuse is a unified generative model for synthesizing a material as a coordinated set of SVBRDF reflectance maps: diffuse, normal, roughness, and specular. It contains two stages: a learned compression model maps the aligned material maps from pixel space into a compact latent space, and a latent diffusion model generates latent material representations from Gaussian noise. A decoder converts the generated latent representation back into the four material maps. The diffusion model accepts no condition, one condition, or several conditions simultaneously, including text, reference images, color palettes, and sketches. The same latent representation also supports map-level and spatial material editing.

  2. Knowl 2 — Map-specific multi-encoder compression with rendering-aware training

    model/method

    MatFuse replaces a single material encoder with one encoder and one quantization codebook per input map. For a material consisting of NN maps x={M1,…,MN}x=\{M_1,\ldots,M_N\}, where each Mi∈RH×W×3M_i\in\mathbb{R}^{H\times W\times 3} has pixel dimensions H×WH\times W, encoder EiE_i produces a map-specific latent tensor that is quantized by QiQ_i:

    zi=Qi(Ei(Mi)),zi∈Rh×w×ci.z_i=Q_i(E_i(M_i)),\qquad z_i\in\mathbb{R}^{h\times w\times c_i}.

    The quantized tensors are concatenated along their channel dimension and decoded jointly:

    z=concat⁡(z1,…,zN)∈Rh×w×c,c=∑i=1Nci,x^=D(z).z=\operatorname{concat}(z_1,\ldots,z_N)\in\mathbb{R}^{h\times w\times c},\qquad c=\sum_{i=1}^{N}c_i,\qquad \hat{x}=D(z).

    Here, h×wh\times w is the latent spatial resolution, cic_i is the number of latent channels allocated to map ii, cc is the total latent channel count, DD is the shared decoder, and x^\hat{x} is the reconstructed map set. Separate encoders and codebooks encourage map-specific, approximately disentangled latent features, while concatenation lets the decoder model cross-map consistency. Compression training combines pixel-space L2L_2 reconstruction loss, LPIPS perceptual loss, a patch-based adversarial loss, and codebook commitment loss. MatFuse additionally minimizes a rendering loss equal to the mean-squared error between renders of the ground-truth maps and reconstructed maps, encouraging the separately reconstructed maps to produce a coherent material appearance.

  3. Knowl 3 — Latent diffusion objective and sampling process

    equation

    After training the multi-encoder compression model, MatFuse trains a U-Net diffusion model with self-attention to model the distribution of the quantized material latents rather than pixel-space maps. Let z0z_0 be a clean encoded material latent, tt a randomly sampled diffusion timestep, ztz_t the noisy latent at that timestep, and ϵt\epsilon_t the Gaussian noise added by the forward process. For an optional condition yy, let τ(y)\tau(y) denote its learned embedding; when no condition is supplied, this input is omitted. The diffusion network ϵθ\epsilon_\theta is trained with

    Ldiff=Et,z0,ϵ[∥ϵt−ϵθ(zt,t,τ(y))∥22].\mathcal{L}_{\mathrm{diff}}=\mathbb{E}_{t,z_0,\epsilon}\left[\left\|\epsilon_t-\epsilon_\theta\bigl(z_t,t,\tau(y)\bigr)\right\|_2^2\right].

    At generation time, a latent sampled from a standard normal distribution is iteratively denoised by the trained network and then decoded into material maps. In the reported implementation, denoising uses the DDIM schedule with T=50T=50 sampling steps.

  4. Knowl 4 — Composable global and local conditioning

    model/method

    MatFuse separates conditioning into global and local mechanisms. Text prompts and reference images are embedded with a pretrained CLIP encoder and injected at every diffusion denoising step through multi-head cross-attention between the latent features and the condition embedding. A palette condition is obtained by counting colors in an input image, clustering colors whose CIE76 distance is below a chosen threshold, retaining the five most prevalent colors, and projecting those colors into a one-dimensional embedding with a multilayer perceptron trained jointly with the diffusion model.

    Sketches provide local structural control. A sketch is projected to the latent spatial resolution by a small convolutional network trained jointly with the diffusion U-Net, and the resulting feature map is concatenated with the noisy latent. The paper uses both hand-drawn sketches and Canny edges extracted from a material render under diffuse lighting. During classifier-free conditioning training, each individual condition is randomly dropped with probability 50%50\%, while all conditions are dropped together with probability 10%10\%. This enables the same model to generate unconditionally, under a single modality, or under combinations such as a sketch plus text, image, or palette.

  5. Knowl 5 — Volumetric inpainting for map-level material editing

    model/method

    The map-specific latent channels learned by the multi-encoder compression model allow MatFuse to edit selected material properties without regenerating every map. Let zz be the encoded latent of a supplied material, zt+1z_{t+1} the current noisier latent at the adjacent denoising step, and mm a binary mask with the same spatial and channel dimensions as the latent. At denoising step tt, MatFuse forms the edited latent input as

    z~t=z⊙m+zt+1⊙(1−m),\tilde{z}_t=z\odot m+z_{t+1}\odot(1-m),

    where ⊙\odot denotes element-wise multiplication. Masking spatial positions edits selected regions; masking channels associated with one map edits or regenerates that material property; fully masking a map generates the complete missing map. Spatial and channel masking can also be combined, allowing selected regions of only selected maps to be regenerated, optionally under text, image, palette, or sketch conditioning. This joint spatial-and-channel masking is termed volumetric inpainting.

  6. Knowl 6 — Training data and optimization configuration

    experimental setup

    MatFuse is trained using the SVBRDF collection of approximately 20,000 blended materials, represented by diffuse, normal, specular, and roughness maps, together with 320 materials from PolyHaven. Each PolyHaven material is augmented with 1 crop at 4K resolution, 4 at 2K, 16 at 1K, 64 at 512 pixels, and 256 at 256 pixels, producing 431 crops per material and 140,508 crops overall. All crops are resized to 256×256256\times256 pixels. The combined material collection contains approximately 160,000 materials or material crops, and each crop is rendered under five environment maps with four rotations (0∘0^\circ, 90∘90^\circ, 180∘180^\circ, and 270∘270^\circ), producing 20 renders per crop and approximately 3.2 million renders.

    The compression model uses Adam with batch size 4, learning rate 10−410^{-4}, and 4,000,000 iterations. Its adversarial loss is enabled after 300,000 iterations, and the encoders use downsampling factor f=8f=8. The diffusion model uses AdamW with batch size 20 for 500,000 iterations, learning rate 10−410^{-4}, and a linear warm-up beginning at 10−610^{-6}. The diffusion noise schedule uses linearly varying β\beta values, and inference uses 50 DDIM steps.

  7. Knowl 7 — Qualitative control over material appearance and structure

    empirical result

    Qualitative experiments show that MatFuse transfers the intended semantics of global conditions and the intended geometry of local conditions into the generated material maps. Text prompts control broad appearance; for example, adding an adjective such as “shiny” produces a visibly less rough material. Image prompts transfer visual characteristics such as highlights and surface appearance, while palette conditioning reproduces the selected colors primarily in the diffuse map. Sketch conditioning transfers the drawn pattern into the material geometry through the normal map and works with both clean hand-drawn sketches and noisy Canny-derived sketches. When a sketch is combined with a global condition, the generated material preserves the sketch-defined spatial structure while also reflecting the text, image, or palette semantics.

  8. Knowl 8 — Quantitative generation quality

    data/table

    MatFuse was evaluated by rendering 2,000 unconditional samples at 512×512512\times512 resolution. CLIP-IQA was computed using the contrastive pair “high-quality/low-quality,” and FID was computed against the dataset renders. Higher CLIP-IQA and lower FID indicate better results. The comparison includes ground-truth renders, an unconditional latent-diffusion baseline, and TileGen.

    Could not parse LaTeX table

    MatFuse improves substantially over the unconditional LDM baseline in both metrics, reaches a CLIP-IQA score close to the ground-truth upper bound and nearly equal to TileGen, and obtains the best FID. Its lower FID than TileGen indicates greater similarity to the ground-truth render distribution under this evaluation, whereas the paper attributes TileGen’s weaker FID partly to its more restricted sample diversity.

  9. Knowl 9 — Human preference over a class-conditioned material generator

    empirical result

    A pairwise user study compared MatFuse with TileGen using 100 computer-science MS/PhD students. Each participant evaluated 25 randomly selected material pairs from a pool of 100 samples spanning leather, wood, marble, stone, and ceramic, judging realism and rendering quality. MatFuse was globally conditioned with the material class name, whereas TileGen was trained specifically on the compared categories. Across the study, MatFuse received 1,078 votes, TileGen received 949, and participants expressed no preference in 473 cases. A chi-square test reported a statistically significant preference for MatFuse over TileGen, with χ2=16.41\chi^2=16.41 and p<0.05p<0.05.

  10. Knowl 10 — Ablation of map-specific encoding and rendering loss

    data/table

    The compression ablation measures RMSE between reconstructed and ground-truth maps. The Base models use a single encoder with the indicated codebook size, while Multi Enc. uses MatFuse’s separate encoder and latent representation for each material map. The multi-encoder model produces lower reconstruction error for every map and for the rendered material than all single-encoder variants. A second ablation compares the baseline perceptual reconstruction loss with the same loss augmented by the rendering loss LrenderL_{\mathrm{render}}; adding the rendering loss improves the rendered result and reduces the errors for diffuse, roughness, and specular maps.

    Could not parse LaTeX table
    Could not parse LaTeX table

    The results support both architectural choices: separate map encoders improve map reconstruction and provide structured latent control, while the rendering loss improves cross-map appearance consistency.

Coverage note — The paper’s resource limitations, lack of tileability, and proposed future extensions such as patch-based high-resolution generation and ControlNet-style local conditioning were omitted because they are stated caveats or future directions rather than demonstrated contributions.

References

  1. 1.Miika Aittala, Timo Aila, and Jaakko Lehtinen. Reflectance modeling by neural texture synthesis. ACM Transactions on Graphics (ToG), 35(4):1–13, 2016.
  2. 2.Martin Arjovsky, Soumith Chintala, and Leon Bottou. Wasserstein generative adversarial networks. In International conference on machine learning, pages 214–223. PMLR, 2017.
  3. 3.Sai Bi, Zexiang Xu, Kalyan Sunkavalli, David Kriegman, and Ravi Ramamoorthi. Deep 3d capture: Geometry and reflectance from sparse multi-view images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5960–5969, 2020.
  4. 4.Andrew Brock, Jeff Donahue, and Karen Simonyan. Large scale gan training for high fidelity natural image synthesis. arXiv preprint arXiv:1809.11096, 2018.
  5. 5.John Canny. A computational approach to edge detection. IEEE Transactions on Pattern Analysis and Machine Intelligence, PAMI-8(6):679–698, 1986.
  6. 6.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition, pages 248–255, 2009.
  7. 7.Valentin Deschaintre, Miika Aittala, Fredo Durand, George Drettakis, and Adrien Bousseau. Single-image svbrdf capture with a rendering-aware deep network. ACM Transactions on Graphics (ToG), 37(4):1–15, 2018.
  8. 8.Valentin Deschaintre, Miika Aittala, Fredo Durand, George Drettakis, and Adrien Bousseau. Flexible SVBRDF capture with a multi-image deep network. In Computer Graphics Forum, pages 1–13. Wiley Online Library, 2019.
  9. 9.Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. Advances in Neural Information Processing Systems, 34:8780–8794, 2021.
  10. 10.Alexey Dosovitskiy and Thomas Brox. Generating images with perceptual similarity metrics based on deep networks. Advances in neural information processing systems, 29, 2016.
  11. 11.Patrick Esser, Robin Rombach, and Bjorn Ommer. Taming transformers for high-resolution image synthesis. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12873–12883, 2021.
  12. 12.Duan Gao, Xiao Li, Yue Dong, Pieter Peers, Kun Xu, and Xin Tong. Deep inverse rendering for high-resolution svbrdf estimation from an arbitrary number of images. ACM Trans. Graph., 38(4):134–1, 2019.
  13. 13.Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets. In Advances in Neural Information Processing Systems. Curran Associates, Inc., 2014.
  14. 14.Darya Guarnera, Giuseppe Claudio Guarnera, Abhijeet Ghosh, Cornelia Denk, and Mashhuda Glencross. Brdf representation and acquisition. In Computer Graphics Forum, pages 625–650. Wiley Online Library, 2016.
  15. 15.Pascal Guehl, Remi Allegre, J-M Dischler, Bedrich Benes, and Eric Galin. Semi-procedural textures using point process texture basis functions. In Computer Graphics Forum, pages 159–171. Wiley Online Library, 2020.
  16. 16.Ishaan Gulrajani, Faruk Ahmed, Martin Arjovsky, Vincent Dumoulin, and Aaron C Courville. Improved training of wasserstein gans. Advances in neural information processing systems, 30, 2017.
  17. 17.Jie Guo, Shuichang Lai, Chengzhi Tao, Yuelong Cai, Lei Wang, Yanwen Guo, and Ling-Qi Yan. Highlight-aware two-stream network for single-image svbrdf acquisition. ACM Transactions on Graphics (TOG), 40(4):1–14, 2021.
  18. 18.Yu Guo, Cameron Smith, Milos Hašan, Kalyan Sunkavalli, and Shuang Zhao. MaterialGAN: reflectance capture using a generative svbrdf model. arXiv preprint arXiv:2010.00114, 2020.
  19. 19.Zhen He, Jie Guo, Yan Zhang, Qinghao Tu, Mufan Chen, Yanwen Guo, Pengyu Wang, and Wei Dai. Text2Mat: Generating Materials from Text. In Pacific Graphics Short Papers and Posters. The Eurographics Association, 2023.
  20. 20.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In Advances in Neural Information Processing Systems, 2017.
  21. 21.Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598, 2022.
  22. 22.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 33:6840–6851, 2020.
  23. 23.Yiwei Hu, Milos Hašan, Paul Guerrero, Holly Rushmeier, and Valentin Deschaintre. Controlling material appearance by examples. In Computer Graphics Forum, pages 117–128. Wiley Online Library, 2022.
  24. 24.Lianghua Huang, Di Chen, Yu Liu, Yujun Shen, Deli Zhao, and Jingren Zhou. Composer: Creative and controllable image synthesis with composable conditions. arXiv preprint arXiv:2302.09778, 2023.
  25. 25.Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A Efros. Image-to-image translation with conditional adversarial networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1125–1134, 2017.
  26. 26.Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen. Progressive growing of gans for improved quality, stability, and variation. arXiv preprint arXiv:1710.10196, 2017.
  27. 27.Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Analyzing and improving the image quality of stylegan. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8110–8119, 2020.
  28. 28.Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Analyzing and improving the image quality of stylegan. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8110–8119, 2020.
  29. 29.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  30. 30.Brenden M Lake, Tomer D Ullman, Joshua B Tenenbaum, and Samuel J Gershman. Building machines that learn and think like people. Behavioral and brain sciences, 40:e253, 2017.
  31. 31.Xiao Li, Yue Dong, Pieter Peers, and Xin Tong. Modeling surface appearance from a single photograph using self-augmented convolutional neural networks. ACM Transactions on Graphics (ToG), 36(4):1–11, 2017.
  32. 32.Zhengqin Li, Kalyan Sunkavalli, and Manmohan Chandraker. Materials for masses: Svbrdf acquisition with a single mobile phone image. In Proceedings of the European Conference on Computer Vision (ECCV), pages 72–87, 2018.
  33. 33.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017.
  34. 34.Rosalie Martin, Arthur Roullier, Romain Rouffet, Adrien Kaiser, and Tamy Boubekeur. Materia: Single image high-resolution material capture in the wild. In Computer Graphics Forum, pages 163–177. Wiley Online Library, 2022.
  35. 35.Lars Mescheder. On the convergence properties of gan training. arXiv preprint arXiv:1801.04406, 1:16, 2018.
  36. 36.Luke Metz, Ben Poole, David Pfau, and Jascha Sohl-Dickstein. Unrolled generative adversarial networks. arXiv preprint arXiv:1611.02163, 2016.
  37. 37.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PMLR, 2021.
  38. 38.Alexander Raistrick, Lahav Lipson, Zeyu Ma, Lingjie Mei, Mingzhe Wang, Yiming Zuo, Karhan Kayan, Hongyu Wen, Beining Han, Yihan Wang, Alejandro Newell, Hei Law, Ankit Goyal, Kaiyu Yang, and Jia Deng. Infinite photorealistic worlds using procedural generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12630–12641, 2023.
  39. 39.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10684–10695, 2022.
  40. 40.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical image computing and computer-assisted intervention, pages 234–241. Springer, 2015.
  41. 41.Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen. Improved techniques for training gans. Advances in neural information processing systems, 29, 2016.
  42. 42.Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning, pages 2256–2265. PMLR, 2015.
  43. 43.Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502, 2020.
  44. 44.Zhihang Song, Zimin He, Xingyu Li, Qiming Ma, Ruibo Ming, Zhiqi Mao, Huaxin Pei, Lihui Peng, Jianming Hu, Danya Yao, and Yi Zhang. Synthetic datasets for autonomous driving: A survey, 2023.
  45. 45.Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott E Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. Going deeper with convolutions. corr abs/1409.4842 (2014), 2014.
  46. 46.Aaron Van Den Oord, Oriol Vinyals, et al. Neural discrete representation learning. Advances in neural information processing systems, 30, 2017.
  47. 47.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  48. 48.Giuseppe Vecchio, Simone Palazzo, and Concetto Spampinato. Surfacenet: Adversarial svbrdf estimation from a single image. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 12840–12848, 2021.
  49. 49.Giuseppe Vecchio, Simone Palazzo, Dario C Guastella, Ignacio Carlucho, Stefano V Albrecht, Giovanni Muscato, and Concetto Spampinato. Midgard: A simulation platform for autonomous navigation in unstructured environments. arXiv preprint arXiv:2205.08389, 2022.
  50. 50.Jianyi Wang, Kelvin CK Chan, and Chen Change Loy. Exploring clip for assessing the look and feel of images. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 2555–2563, 2023.
  51. 51.Kai Zhang, Jingyun Liang, Luc Van Gool, and Radu Timofte. Designing a practical degradation model for deep blind image super-resolution. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4791–4800, 2021.
  52. 52.Lvmin Zhang and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. arXiv preprint arXiv:2302.05543, 2023.
  53. 53.Xilong Zhou and Nima Khademi Kalantari. Adversarial single-image svbrdf estimation with hybrid training. In Computer Graphics Forum, pages 315–325. Wiley Online Library, 2021.
  54. 54.Xilong Zhou, Milos Hasan, Valentin Deschaintre, Paul Guerrero, Kalyan Sunkavalli, and Nima Khademi Kalantari. TileGen: Tileable, controllable material generation and capture. In SIGGRAPH Asia 2022 Conference Papers, pages 1–9, 2022.

Citation

MLA
Vecchio, G., et al. “MatFuse: Controllable Material Generation with Diffusion Models”. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp. 4429–38, https://doi.org/10.1109/CVPR52733.2024.00424.
APA
Vecchio, G., Sortino, R., Palazzo, S., & Spampinato, C. (2024). MatFuse: Controllable Material Generation with Diffusion Models. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 4429–4438. https://doi.org/10.1109/CVPR52733.2024.00424
Chicago
Vecchio, G., R. Sortino, S. Palazzo, and C. Spampinato. 2024. “MatFuse: Controllable Material Generation with Diffusion Models”. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 4429–38. https://doi.org/10.1109/CVPR52733.2024.00424.
Harvard
Vecchio, G. et al. (2024) “MatFuse: Controllable Material Generation with Diffusion Models”, 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 4429–4438. Available at: https://doi.org/10.1109/CVPR52733.2024.00424.
Vancouver
1. Vecchio G, Sortino R, Palazzo S, Spampinato C (2024) MatFuse: Controllable Material Generation with Diffusion Models. In: 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 4429–4438

BibTeX

@inproceedings{Vecchio_2024, title={MatFuse: Controllable Material Generation with Diffusion Models}, url={http://dx.doi.org/10.1109/CVPR52733.2024.00424}, DOI={10.1109/cvpr52733.2024.00424}, booktitle={2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Vecchio, Giuseppe and Sortino, Renato and Palazzo, Simone and Spampinato, Concetto}, year={2024}, month=June, pages={4429–4438} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE