Paint-it: Text-to-Texture Synthesis via Deep Convolutional Texture Map Optimization and Physically-Based Rendering

Kim YouwangTae-Hyun OhGerard Pons-Moll

article2024CVPR103 citations

Presents a text-driven 3D texturing framework that re-parameterizes physically-based rendering maps with deep convolutional neural kernels to filter out noisy score-distillation gradients, synthesizing high-quality, relightable texture maps within fifteen minutes.

Listen

Generating realistic and diverse 3D digital assets is essential for industries like gaming, film, and virtual reality, but manual texture authoring remains labor-intensive and costly. While recent artificial intelligence approaches attempt to generate 3D assets automatically from simple text descriptions, existing methods often yield low visual fidelity, fail to model complex surface reflections, or produce formats that are difficult to integrate into standard production pipelines.

The article demonstrates and evaluates Paint-it, an automated text-driven system that synthesizes high-fidelity, physically-based rendering texture maps for untextured 3D meshes without requiring paired text-and-3D training datasets.

To overcome the visual artifacts and noise typical of generative 2D-to-3D optimization, the approach introduces a deep convolutional neural parameterization of texture maps. Instead of directly optimizing individual pixel values, the framework optimizes the parameters of a convolutional network using feedback from a pre-trained text-to-image diffusion model. The system renders the textured mesh differentiably across multiple views, estimating diffuse color, surface roughness, metalness, and normal maps to evaluate photorealism and adherence to the text prompt across diverse object, human, and animal meshes.

The findings show that neural parameterization inherently schedules optimization from low-frequency structural shapes to high-frequency details, effectively filtering out noisy gradients that degrade standard pixel-based optimization. In quantitative evaluations on benchmark datasets, Paint-it achieved superior image realism scores (a Fréchet Inception Distance of 34.46) compared to existing state-of-the-art methods (which ranged from 37.89 to 58.79). In human perceptual studies involving 30 evaluators, it was the only method to achieve a realism score above 4 on a 5-point scale (4.37 versus 2.71 to 3.34 for competitors), while generating complete texture sets within 15 to 30 minutes.

These results indicate that studios and developers can rapidly produce production-ready, editable 3D textures that seamlessly integrate into standard commercial graphics engines. Because the system disentangles material properties like roughness and lighting reflections from surface geometry, assets can be dynamically relit and modified without remeshing or introducing visual seams, significantly reducing production turnaround time.

Organizations exploring automated 3D asset creation should consider incorporating convolutional re-parameterization workflows to streamline asset texturing pipelines. Before full-scale industrial deployment, teams should investigate optimization speed enhancements—such as integrating faster consistency models or building feed-forward networks trained on generated texture sets—to reduce per-asset synthesis times below the current 15-to-30-minute threshold. Confidence in the qualitative and perceptual performance is high across general object categories, though production planning must account for per-instance compute time during large-scale batch generation.

arXiv: 2312.11360
Cover for Paint-it: Text-to-Texture Synthesis via Deep Convolutional Texture Map Optimization and Physically-Based Rendering

Abstract

We present Paint-it, a text-driven high-fidelity texture map synthesis method for 3D meshes via neural re-parameterized texture optimization. Paint-it synthesizes texture maps from a text description by synthesis-through-optimization, exploiting the Score-Distillation Sampling (SDS). We observe that directly applying SDS yields undesirable texture quality due to its noisy gradients. We reveal the importance of texture parameterization when using SDS. Specifically, we propose Deep Convolutional Physically-Based Rendering (DC-PBR) parameterization, which re-parameterizes the physically-based rendering (PBR) texture maps with randomly initialized convolution-based neural kernels, instead of a standard pixel-based parameterization. We show that DC-PBR inherently schedules the optimization curriculum according to texture frequency and naturally filters out the noisy signals from SDS. In experiments, Paint-it obtains remarkable quality PBR texture maps within 15 min., given only a text description. We demonstrate the generalizability and practicality of Paint-it by synthesizing high-quality texture maps for large-scale mesh datasets and showing test-time applications such as relighting and material control using a popular graphics engine. Project page: https://kim-youwang.github.io/paint-it.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Paint-it: Text-Driven PBR Texture Synthesis via Neural Re-parameterized Optimization
  • 3.1. Preliminary: Score-Distillation Sampling
  • 3.2. Goal of Paint-it
  • 3.3. DC-PBR: Deep Convolutional PBR Texture Map Re-parameterization
  • 3.4. Text-driven DC-PBR Optimization
  • 4. Analysis: Effect of the Deep Convolutional Re-parameterization for PBR Texture Maps
  • 4.1. Analysis of Fitting Behavior
  • 4.2. Analysis of Optimization with the SDS Loss
  • 5. Experiments
  • 5.1. Qualitative Results
  • 5.2. Comparison with Competing Methods
  • 5.3. Ablation studies
  • 6. Discussion, Limitation, and Conclusion
  • References

Knowls

  1. Knowl 1 — Deep Convolutional Physically-Based Rendering Parameterization

    model/method

    Instead of directly optimizing the per-pixel values of 2D texture maps, the Deep Convolutional Physically-Based Rendering (DC-PBR) parameterization represents multi-channel physically-based rendering (PBR) texture maps via a neural surrogate Tθ\mathcal{T}_\theta.

    Let Tθ\mathcal{T}_\theta be a convolutional U-Net with skip connections parameterized by weights θ\theta. A fixed Gaussian noise tensor z∼N(0,I)∈RH×W×3\mathbf{z} \sim \mathcal{N}(0, \mathbf{I}) \in \mathbb{R}^{H \times W \times 3} serves as the static input, where HH and WW denote the height and width of the target texture maps. The network outputs the disentangled PBR texture maps:

    [Kθd,Kθrm,Kθn]=Tθ(z)[\mathbf{K}^d_\theta, \mathbf{K}^{rm}_\theta, \mathbf{K}^n_\theta] = \mathcal{T}_\theta(\mathbf{z})

    where:

    • Kθd∈RH×W×3\mathbf{K}^d_\theta \in \mathbb{R}^{H \times W \times 3} is the diffuse (albedo) texture map,
    • Kθrm∈RH×W×2\mathbf{K}^{rm}_\theta \in \mathbb{R}^{H \times W \times 2} represents the roughness and metalness texture maps, and
    • Kθn∈RH×W×3\mathbf{K}^n_\theta \in \mathbb{R}^{H \times W \times 3} represents the detail surface normal map in tangent space.

    The combined output tensor has shape H×W×(3+2+3)H \times W \times (3 + 2 + 3). Re-parameterizing texture maps through randomly initialized convolutional neural kernels provides an implicit prior that couples spatial frequencies across the entire texture space during optimization.

  2. Knowl 2 — Cook-Torrance Physically-Based Shading for DC-PBR Textured Meshes

    equation

    Given a 3D surface point p\mathbf{p} on a mesh textured by local parameters indexed from DC-PBR maps [Kθd,Kθrm,Kθn][\mathbf{K}^d_\theta, \mathbf{K}^{rm}_\theta, \mathbf{K}^n_\theta] via UV coordinates—specifically diffuse color kθd∈R3\mathbf{k}^d_\theta \in \mathbb{R}^3, roughness kθr∈Rk^r_\theta \in \mathbb{R}, metalness kθm∈Rk^m_\theta \in \mathbb{R}, and perturbed tangent-space normal direction kθn∈R3\mathbf{k}^n_\theta \in \mathbb{R}^3—the specularity kθs∈R3\mathbf{k}^s_\theta \in \mathbb{R}^3 is defined as:

    kθs=0.04⋅(1−kθm)+kθm⋅kθd\mathbf{k}^s_\theta = 0.04 \cdot (1 - k^m_\theta) + k^m_\theta \cdot \mathbf{k}^d_\theta

    The total outgoing radiance Lθ(p,ω)L_\theta(\mathbf{p}, \boldsymbol{\omega}) viewed from direction ω\boldsymbol{\omega} under incident environment lighting LiL_i integrated over the hemisphere Ω\Omega around the perturbed surface normal nθ\mathbf{n}_\theta is decomposed into a diffuse component Lθd(p)L^d_\theta(\mathbf{p}) and a specular component Lθs(p,ω)L^s_\theta(\mathbf{p}, \boldsymbol{\omega}) using the Cook-Torrance microfacet specular shading model:

    Lθ(p,ω)=Lθd(p)+Lθs(p,ω)L_\theta(\mathbf{p}, \boldsymbol{\omega}) = L^d_\theta(\mathbf{p}) + L^s_\theta(\mathbf{p}, \boldsymbol{\omega})

    Lθd(p)=kθd(1−kθm)∫ΩLi(p,ωi)(ωi⋅nθ) dωiL^d_\theta(\mathbf{p}) = \mathbf{k}^d_\theta (1 - k^m_\theta) \int_{\Omega} L_i(\mathbf{p}, \boldsymbol{\omega}_i) (\boldsymbol{\omega}_i \cdot \mathbf{n}_\theta) \, d\boldsymbol{\omega}_i

    Lθs(p,ω)=∫ΩDθFθGθ4(ω⋅nθ)(ωi⋅nθ)Li(p,ωi)(ωi⋅nθ) dωiL^s_\theta(\mathbf{p}, \boldsymbol{\omega}) = \int_{\Omega} \frac{D_\theta F_\theta G_\theta}{4 (\boldsymbol{\omega} \cdot \mathbf{n}_\theta)(\boldsymbol{\omega}_i \cdot \mathbf{n}_\theta)} L_i(\mathbf{p}, \boldsymbol{\omega}_i) (\boldsymbol{\omega}_i \cdot \mathbf{n}_\theta) \, d\boldsymbol{\omega}_i

    where ωi\boldsymbol{\omega}_i is the incident lighting direction, DθD_\theta is the microfacet distribution function, FθF_\theta is the Fresnel reflection term (dependent on kθs\mathbf{k}^s_\theta), and GθG_\theta is the geometric attenuation function, with DθD_\theta and GθG_\theta depending on roughness kθrk^r_\theta.

  3. Knowl 3 — Text-Driven DC-PBR Texture Optimization via Score-Distillation Sampling

    model/method

    Paint-it synthesizes text-conforming PBR texture maps by optimizing the neural network parameters θ\theta of the DC-PBR generator Tθ\mathcal{T}_\theta using Score-Distillation Sampling (SDS) guidance from a pre-trained 2D text-to-image diffusion model ϵ^ϕ\hat{\boldsymbol{\epsilon}}_\phi.

    Let RM(Kθd,Kθrm,Kθn)\mathcal{R}^M(\mathbf{K}^d_\theta, \mathbf{K}^{rm}_\theta, \mathbf{K}^n_\theta) denote a differentiable mesh rasterizer rendering multi-view images Iθ\mathbf{I}_\theta of mesh MM with PBR textures [Kθd,Kθrm,Kθn]=Tθ(z)[\mathbf{K}^d_\theta, \mathbf{K}^{rm}_\theta, \mathbf{K}^n_\theta] = \mathcal{T}_\theta(\mathbf{z}) under random camera viewpoints and environment lighting. The optimization objective is:

    θ∗=arg⁡min⁡θEt,ϵ[∥ϵ^ϕ(RtM(Kθd,Kθrm,Kθn);y,t)−ϵ∥22]\theta^* = \arg\min_\theta \mathbb{E}_{t, \boldsymbol{\epsilon}} \left[ \left\| \hat{\boldsymbol{\epsilon}}_\phi\left(\mathcal{R}^M_t(\mathbf{K}^d_\theta, \mathbf{K}^{rm}_\theta, \mathbf{K}^n_\theta); y, t\right) - \boldsymbol{\epsilon} \right\|_2^2 \right]

    where yy is the input text prompt, t∼U(0,1)t \sim \mathcal{U}(0, 1) is the diffusion noise timestep, ϵ∼N(0,I)\boldsymbol{\epsilon} \sim \mathcal{N}(0, \mathbf{I}) is the added perturbation noise, and RtM\mathcal{R}^M_t is the noisy rendered image at timestep tt.

    The parameter gradient used to iteratively update the neural weights θ\theta is computed as:

    ∇θLSDS(ϕ,Iθ)=Et,ϵ[m(t)(ϵ^ϕ(Iθ,t;y,t)−ϵ)∂Iθ∂θ]\nabla_\theta \mathcal{L}_{\text{SDS}}(\phi, \mathbf{I}_\theta) = \mathbb{E}_{t, \boldsymbol{\epsilon}} \left[ m(t) \left( \hat{\boldsymbol{\epsilon}}_\phi(\mathbf{I}_{\theta, t}; y, t) - \boldsymbol{\epsilon} \right) \frac{\partial \mathbf{I}_\theta}{\partial \theta} \right]

    where m(t)m(t) is a timestep-conditioned weighting scalar. The resulting high-quality PBR texture maps are obtained as [Kd∗,Krm∗,Kn∗]=Tθ∗(z)[\mathbf{K}^{d*}, \mathbf{K}^{rm*}, \mathbf{K}^{n*}] = \mathcal{T}_{\theta^*}(\mathbf{z}).

  4. Knowl 4 — Frequency Scheduling and Noise Filtering in DC-PBR Optimization

    empirical result

    Spatial Fourier energy analysis across five non-overlapping frequency bands (from lowest to highest frequencies) demonstrates why neural re-parameterization succeeds over direct pixel optimization when guided by Score-Distillation Sampling (SDS):

    • Pixel-Based Optimization: Under direct pixel parameterization with SDS loss (even when regularized with total variation loss LTV\mathcal{L}_{\text{TV}}), all frequency bands increase in energy simultaneously from initial iterations. The optimizer fits high-frequency noisy gradients induced by the random perturbation noise ϵ\boldsymbol{\epsilon}, timesteps tt, and camera poses, yielding severely degraded and jittered texture maps.
    • DC-PBR Optimization: The convolutional U-Net parameterization Tθ\mathcal{T}_\theta enforces an implicit spectral bias. It fits low-frequency components (overall structure, base colors, skin tones) rapidly in early iterations and defers fitting higher-frequency bands to later iterations.

    This inherent frequency-scheduled learning naturally filters out high-frequency noisy signals from the SDS loss without requiring manual filtering heuristics, producing a coarse-to-fine texture synthesis process.

  5. Knowl 5 — Quantitative Texture Quality and Realism Benchmark on Objaverse

    data/table

    Paint-it was evaluated against four text-driven 3D texturing baselines on a subset of meshes from the Objaverse dataset using identical text prompts: Latent-Paint, Fantasia3D, Text2Tex, and TEXTure. Texture fidelity was evaluated by Fréchet Inception Distance (FID, lower is better) on multi-view rendered images, and perceived realism was evaluated via a user study (mean score from 30 respondents on a 1–5 scale where 4 corresponds to realistic, higher is better).

    Method Latent-Paint Fantasia3D Text2Tex TEXTure Paint-it (Ours)
    FID (↓\downarrow) 41.11 58.79 37.89 38.40 34.46
    User score (↑\uparrow) 3.22 2.71 3.34 3.04 4.37

    Paint-it achieves the lowest FID score (34.46) and the highest realism rating (4.37), surpassing the realistic threshold (score of 4). Competing inpainting approaches (Text2Tex, TEXTure) and latent optimization (Latent-Paint) suffer from baked-in lighting and seam artifacts due to diffuse-only representations, while Fantasia3D produces non-smooth materials and geometric seams due to coordinate-based MLP estimation and surface re-meshing.

  6. Knowl 6 — PBR Material Disentanglement and Pseudo-Stereoscopic Detail Generation

    empirical result

    Disentangling synthesized textures into diffuse (Kθd\mathbf{K}^d_\theta), roughness/metalness (Kθrm\mathbf{K}^{rm}_\theta), and tangent-space normal maps (Kθn\mathbf{K}^n_\theta) provides distinct visual advantages over diffuse-only texturing:

    • Physical Material Separation: DC-PBR accurately differentiates matte materials (e.g., mushrooms, clothing) from metallic and highly specular materials (e.g., teapots, crowns) based on text prompts, preventing baked-in specular reflections.
    • Pseudo-Stereoscopic Relief: The synthesized detail normal map Kθn\mathbf{K}^n_\theta creates the visual appearance of relief details (such as gems and embossed patterns on flat surface geometry) via perturbed surface normals nθ\mathbf{n}_\theta interacting with light directions.
    • Engine Compatibility: Because texture maps are optimized directly onto the input mesh's existing UV layout without modifying topology or re-meshing, the synthesized assets directly support test-time relighting with High-Dynamic Range (HDR) environment maps and manual material property adjustments in graphics software like Blender.
  7. Knowl 7 — Per-Mesh Optimization Time Requirement

    limitation

    Because Paint-it optimizes neural U-Net parameters θ\theta on a per-instance basis via iterative Score-Distillation Sampling and multi-view differentiable rasterization, generating the full set of PBR texture maps requires approximately 15 to 30 minutes per 3D mesh. The method does not offer instantaneous feed-forward texture inference.

Coverage note — None was omitted; all key contributions—the DC-PBR re-parameterization, Cook-Torrance BRDF differentiable rendering pipeline, SDS optimization formulation, Fourier frequency scheduling analysis, empirical benchmarks, material disentanglement behaviors, and runtime limitations—are fully represented.

References

  1. 1.https://ami.postech.ac.kr/members. 8
  2. 2.http://virtualhumans.mpi-inf.mpg.de/people.html. 8
  3. 3.https://renderpeople.com/, 2023. 6, 7
  4. 4.Connelly Barnes, Eli Shechtman, Dan B Goldman, and Adam Finkelstein. The generalized PatchMatch correspondence algorithm. In European Conference on Computer Vision (ECCV), 2010. 6
  5. 5.Tianshi Cao, Karsten Kreis, Sanja Fidler, Nicholas Sharp, and KangXue Yin. Texfusion: Synthesizing 3d textures with text-guided image diffusion models. In IEEE International Conference on Computer Vision (ICCV), 2023. 2, 3
  6. 6.Yukang Cao, Yan-Pei Cao, Kai Han, Ying Shan, and Kwan-Yee K. Wong. Dreamavatar: Text-and-shape guided 3d human avatar generation via diffusion models. arXiv preprint, arxiv:2304.00916, 2023. 3
  7. 7.Eric R. Chan, Connor Z. Lin, Matthew A. Chan, Koki Nagano, Boxiao Pan, Shalini De Mello, Orazio Gallo, Leonidas Guibas, Jonathan Tremblay, Sameh Khamis, Tero Karras, and Gordon Wetzstein. Efficient geometry-aware 3D generative adversarial networks. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 1
  8. 8.Dave Zhenyu Chen, Yawar Siddiqui, Hsin-Ying Lee, Sergey Tulyakov, and Matthias Nießner. Text2tex: Text-driven texture synthesis via diffusion models. In IEEE International Conference on Computer Vision (ICCV), 2023. 2, 3, 6, 7, 8
  9. 9.Rui Chen, Yongwei Chen, Ningxin Jiao, and Kui Jia. Fantasia3d: Disentangling geometry and appearance for high-quality text-to-3d content creation. In IEEE International Conference on Computer Vision (ICCV), 2023. 3, 6, 7
  10. 10.Xu Chen, Tianjian Jiang, Jie Song, Jinlong Yang, Michael Black, Andreas Geiger, and Otmar Hilliges. gdna: Towards generative detailed neural avatars. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 1
  11. 11.Yongwei Chen, Rui Chen, Jiabao Lei, Yabin Zhang, and Kui Jia. Tango: Text-driven photorealistic and robust 3d stylization via lighting decomposition. In Advances in Neural Information Processing Systems (NeurIPS), 2022. 3
  12. 12.R. L. Cook and K. E. Torrance. A reflectance model for computer graphics. ACM Transactions on Graphics (SIGGRAPH), 1(1), 1982. 4
  13. 13.Antonio Criminisi, Patrick Perez, and Kentaro Toyama. Object removal by exemplar-based inpainting. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2003. 6
  14. 14.Matt Deitke, Dustin Schwenk, Jordi Salvador, Luca Weihs, Oscar Michel, Eli VanderBilt, Ludwig Schmidt, Kiana Ehsani, Aniruddha Kembhavi, and Ali Farhadi. Objaverse: A universe of annotated 3d objects. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 6, 7, 8
  15. 15.Zijian Dong, Xu Chen, Michael J.Black Jinlong Yang, Otmar Hilliges, and Andreas Geiger. AG3D: Learning to generate 3D avatars from 2D image collections. In IEEE International Conference on Computer Vision (ICCV), 2023. 1
  16. 16.Leon Gatys, Alexander S Ecker, and Matthias Bethge. Texture synthesis using convolutional neural networks. In Advances in Neural Information Processing Systems (NeurIPS), 2015. 6
  17. 17.Reinhard Heckel and Mahdi Soltanolkotabi. Denoising and regularization via exploiting the structural bias of convolutional generators. In International Conference on Learning Representations (ICLR), 2020. 6
  18. 18.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In Advances in Neural Information Processing Systems (NeurIPS), 2017. 8
  19. 19.Fangzhou Hong, Mingyuan Zhang, Liang Pan, Zhongang Cai, Lei Yang, and Ziwei Liu. Avatarclip: Zero-shot text-driven generation and animation of 3d avatars. ACM Transactions on Graphics (SIGGRAPH), 41(4):1–19, 2022. 3
  20. 20.Yukun Huang, Jianan Wang, Yukai Shi, Xianbiao Qi, Zheng-Jun Zha, and Lei Zhang. Dreamtime: An improved optimization strategy for text-to-3d content creation. arXiv preprint, arxiv:2306.12422, 2023.
  21. 21.Yukun Huang, Jianan Wang, Ailing Zeng, He Cao, Xianbiao Qi, Yukai Shi, Zheng-Jun Zha, and Lei Zhang. Dreamwaltz: Make a scene with complex 3d animatable avatars. arXiv preprint, arxiv:2305.12529, 2023. 1, 3
  22. 22.Ajay Jain, Ben Mildenhall, Jonathan T. Barron, Pieter Abbeel, and Ben Poole. Zero-shot text-guided object generation with dream fields. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  23. 23.Ruixiang Jiang, Can Wang, Jingbo Zhang, Menglei Chai, Mingming He, Dongdong Chen, and Jing Liao. Avatarcraft: Transforming text into neural human avatars with parameterized shape and pose control. In IEEE International Conference on Computer Vision (ICCV), 2023. 1, 3
  24. 24.Samuli Laine, Janne Hellsten, Tero Karras, Yeongho Seol, Jaakko Lehtinen, and Timo Aila. Modular primitives for high-performance differentiable rendering. ACM Transactions on Graphics (SIGGRAPH), 39(6), 2020. 5
  25. 25.Chen-Hsuan Lin, Jun Gao, Luming Tang, Towaki Takikawa, Xiaohui Zeng, Xun Huang, Karsten Kreis, Sanja Fidler, Ming-Yu Liu, and Tsung-Yi Lin. Magic3d: High-resolution text-to-3d content creation. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 1, 3
  26. 26.Qianli Ma, Jinlong Yang, Anurag Ranjan, Sergi Pujades, Gerard Pons-Moll, Siyu Tang, and Michael Black. Learning to dress 3d people in generative clothing. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2020. 1
  27. 27.Gal Metzer, Elad Richardson, Or Patashnik, Raja Giryes, and Daniel Cohen-Or. Latent-nerf for shape-guided generation of 3d shapes and textures. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 2, 3, 6, 7
  28. 28.Tomer Michaeli and Michal Irani. Nonparametric blind super-resolution. In IEEE International Conference on Computer Vision (ICCV), 2013. 6
  29. 29.Oscar Michel, Roi Bar-On, Richard Liu, Sagie Benaim, and Rana Hanocka. Text2mesh: Text-driven neural stylization for meshes. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 3
  30. 30.Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. In European Conference on Computer Vision (ECCV), 2020. 1, 3
  31. 31.Jeong Joon Park, Peter Florence, Julian Straub, Richard Newcombe, and Steven Lovegrove. Deepsdf: Learning continuous signed distance functions for shape representation. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019. 1
  32. 32.Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed A. A. Osman, Dimitrios Tzionas, and Michael J. Black. Expressive body capture: 3D hands, face, and body from a single image. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019. 1
  33. 33.Ben Poole, Ajay Jain, Jonathan T. Barron, and Ben Mildenhall. Dreamfusion: Text-to-3d using 2d diffusion. In International Conference on Learning Representations (ICLR), 2022. 1, 2, 3
  34. 34.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning (ICML), 2021. 3
  35. 35.Elad Richardson, Gal Metzer, Yuval Alaluf, Raja Giryes, and Daniel Cohen-Or. Texture: Text-guided texturing of 3d shapes. ACM Transactions on Graphics (SIGGRAPH), 2023. 2, 3, 6, 7, 8
  36. 36.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 1, 3
  37. 37.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L. Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, Seyedeh Sara Mahdavi, Raphael Gontijo Lopes, Tim Salimans, Jonathan Ho, David Fleet, and Mohammad Norouzi. Photorealistic text-to-image diffusion models with deep language understanding. In Advances in Neural Information Processing Systems (NeurIPS), 2022. 1, 3
  38. 38.Zenglin Shi, Pascal Mettes, Subhransu Maji, and Cees G M Snoek. On measuring and controlling the spectral bias of the deep image prior. International Journal of Computer Vision, 2022. 5
  39. 39.Yawar Siddiqui, Justus Thies, Fangchang Ma, Qi Shan, Matthias Nießner, and Angela Dai. Texturify: Generating textures on 3d shape surfaces. In European Conference on Computer Vision (ECCV), 2022. 3
  40. 40.Yang Song, Prafulla Dhariwal, Mark Chen, and Ilya Sutskever. Consistency models. In International Conference on Machine Learning (ICML), 2023. 8
  41. 41.Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky. Deep image prior. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018. 5, 6
  42. 42.Haochen Wang, Xiaodan Du, Jiahao Li, Raymond A. Yeh, and Greg Shakhnarovich. Score jacobian chaining: Lifting pre-trained 2d diffusion models for 3d generation. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 3
  43. 43.Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He. Non-local neural networks. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018. 6
  44. 44.Zhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao, Chongxuan Li, Hang Su, and Jun Zhu. Prolificdreamer: High-fidelity and diverse text-to-3d generation with variational score distillation. In Advances in Neural Information Processing Systems (NeurIPS), 2023. 3
  45. 45.Jiajun Wu, Chengkai Zhang, Tianfan Xue, William T Freeman, and Joshua B Tenenbaum. Learning a probabilistic latent space of object shapes via 3d generative-adversarial modeling. In Advances in Neural Information Processing Systems (NeurIPS), 2016. 1
  46. 46.Kim Youwang, Kim Ji-Yeon, and Tae-Hyun Oh. CLIP-Actor: Text-driven recommendation and stylization for animating human meshes. In European Conference on Computer Vision (ECCV), 2022. 3
  47. 47.Silvia Zuffi, Angjoo Kanazawa, David Jacobs, and Michael J. Black. 3D menagerie: Modeling the 3D shape and pose of animals. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017. 6, 7

Citation

MLA
Youwang, K., et al. “Paint-it: Text-to-Texture Synthesis via Deep Convolutional Texture Map Optimization and Physically-Based Rendering”. arXiv, 2023, http://arxiv.org/abs/2312.11360v2.
APA
Youwang, K., Oh, T.-H., & Pons-Moll, G. (2023). Paint-it: Text-to-Texture Synthesis via Deep Convolutional Texture Map Optimization and Physically-Based Rendering. arXiv. http://arxiv.org/abs/2312.11360v2
Chicago
Youwang, K., T.-H. Oh, and G. Pons-Moll. 2023. “Paint-it: Text-to-Texture Synthesis via Deep Convolutional Texture Map Optimization and Physically-Based Rendering”. arXiv. http://arxiv.org/abs/2312.11360v2.
Harvard
Youwang, K., Oh, T.-H. and Pons-Moll, G. (2023) “Paint-it: Text-to-Texture Synthesis via Deep Convolutional Texture Map Optimization and Physically-Based Rendering”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2312.11360v2.
Vancouver
1. Youwang K, Oh T-H, Pons-Moll G (2023) Paint-it: Text-to-Texture Synthesis via Deep Convolutional Texture Map Optimization and Physically-Based Rendering. arXiv

BibTeX

@article{youwang2023paint,
  title = {Paint-it: Text-to-Texture Synthesis via Deep Convolutional Texture Map Optimization and Physically-Based Rendering},
  author = {Youwang, Kim and Oh, Tae-Hyun and Pons-Moll, Gerard},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2312.11360v2},
  eprint = {2312.11360}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE