Paint-it: Text-to-Texture Synthesis via Deep Convolutional Texture Map Optimization and Physically-Based Rendering
Kim YouwangTae-Hyun OhGerard Pons-Moll
Presents a text-driven 3D texturing framework that re-parameterizes physically-based rendering maps with deep convolutional neural kernels to filter out noisy score-distillation gradients, synthesizing high-quality, relightable texture maps within fifteen minutes.
Generating realistic and diverse 3D digital assets is essential for industries like gaming, film, and virtual reality, but manual texture authoring remains labor-intensive and costly. While recent artificial intelligence approaches attempt to generate 3D assets automatically from simple text descriptions, existing methods often yield low visual fidelity, fail to model complex surface reflections, or produce formats that are difficult to integrate into standard production pipelines.
The article demonstrates and evaluates Paint-it, an automated text-driven system that synthesizes high-fidelity, physically-based rendering texture maps for untextured 3D meshes without requiring paired text-and-3D training datasets.
To overcome the visual artifacts and noise typical of generative 2D-to-3D optimization, the approach introduces a deep convolutional neural parameterization of texture maps. Instead of directly optimizing individual pixel values, the framework optimizes the parameters of a convolutional network using feedback from a pre-trained text-to-image diffusion model. The system renders the textured mesh differentiably across multiple views, estimating diffuse color, surface roughness, metalness, and normal maps to evaluate photorealism and adherence to the text prompt across diverse object, human, and animal meshes.
The findings show that neural parameterization inherently schedules optimization from low-frequency structural shapes to high-frequency details, effectively filtering out noisy gradients that degrade standard pixel-based optimization. In quantitative evaluations on benchmark datasets, Paint-it achieved superior image realism scores (a Fréchet Inception Distance of 34.46) compared to existing state-of-the-art methods (which ranged from 37.89 to 58.79). In human perceptual studies involving 30 evaluators, it was the only method to achieve a realism score above 4 on a 5-point scale (4.37 versus 2.71 to 3.34 for competitors), while generating complete texture sets within 15 to 30 minutes.
These results indicate that studios and developers can rapidly produce production-ready, editable 3D textures that seamlessly integrate into standard commercial graphics engines. Because the system disentangles material properties like roughness and lighting reflections from surface geometry, assets can be dynamically relit and modified without remeshing or introducing visual seams, significantly reducing production turnaround time.
Organizations exploring automated 3D asset creation should consider incorporating convolutional re-parameterization workflows to streamline asset texturing pipelines. Before full-scale industrial deployment, teams should investigate optimization speed enhancements—such as integrating faster consistency models or building feed-forward networks trained on generated texture sets—to reduce per-asset synthesis times below the current 15-to-30-minute threshold. Confidence in the qualitative and perceptual performance is high across general object categories, though production planning must account for per-instance compute time during large-scale batch generation.
- Paper: DreamFusion: Text-to-3D using 2D Diffusion, Ben Poole et al. (2023). Introduces Score Distillation Sampling (SDS) to optimize 3D representations using 2D text-to-image diffusion models, establishing the foundational distillation paradigm used by Paint-it.
- Paper: ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score Distillation, Zhengyi Wang et al. (2023). Presents Variational Score Distillation (VSD) for high-fidelity text-to-3D synthesis, providing direct mathematical grounding and baseline context for diffusion guidance in 3D optimization.
- Paper: Magic3D: High-Resolution Text-to-3D Content Creation, Chen-Hsuan Lin et al. (2022). Demonstrates coarse-to-fine 3D mesh texturing and differentiable rendering guided by 2D latent diffusion, serving as an important predecessor to texture map optimization pipelines.
- Paper: Texture Synthesis Using Convolutional Neural Networks, Leon A. Gatys et al. (2015). Establishes the foundational principles of using deep convolutional neural network representations and feature statistics for texture synthesis and optimization.
- Paper: Adding Conditional Control to Text-to-Image Diffusion Models, Lvmin Zhang et al. (2023). Introduces ControlNet for adding structural and spatial conditioning to text-to-image diffusion models, informing multi-view and geometry-aligned generative guidance.
- Paper: Text-to-3D using Gaussian Splatting, Zilong Chen et al. (2024). Explores explicit 3D Gaussian Splatting representations with score distillation, offering an alternative explicit 3D optimization strategy beyond convolutional texture mapping on meshes.
- Paper: Magic123: One Image to High-Quality 3D Object Generation Using Both 2D and 3D Diffusion Priors, Guocheng Qian et al. (2024). Combines 2D and 3D diffusion priors for single-image-to-3D mesh generation, extending the principles of multi-view diffusion supervision to image-conditioned 3D synthesis.
- Paper: XCube: Large-Scale 3D Generative Modeling using Sparse Voxel Hierarchies, Xuanchi Ren et al. (2024). Applies hierarchical sparse voxel structures and cascaded latent diffusion to scale 3D generative modeling beyond single-mesh texturing to large-scale scenes.
