IRON: Inverse Rendering by Optimizing Neural SDFs and Materials from Photometric Images
Kai ZhangFujun LuanZhengqi LiNoah Snavely
Proposes a hybrid inverse rendering framework that couples volumetric neural signed distance fields with edge-aware physics-based surface rendering to reconstruct high-fidelity textured meshes directly compatible with standard graphics engines.
Digitizing real-world objects into high-quality 3D digital assets is a critical capability for augmented reality, virtual reality, and modern computer graphics. While traditional mesh-based methods struggle with topological changes during optimization and existing neural techniques often fuse lighting and material properties in ways incompatible with standard editing software, creating realistic, editable 3D content remains technically challenging.
The article demonstrates a novel inverse rendering pipeline called IRON that reconstructs high-fidelity 3D geometry and spatially varying material properties directly from multi-view flashlight photographs, producing assets that convert cleanly into standard triangle meshes and texture maps.
The authors evaluated the framework on synthetic datasets of nine complex objects and real-world captures of five physical subjects photographed under a co-located camera and flashlight setup. The technical approach applies a two-stage hybrid optimization: it first optimizes a neural signed distance field and diffuse albedo using volumetric radiance to establish correct global shape and topology, and then refines geometry and separates material parameters via an edge-aware physics-based surface rendering algorithm that samples subpixel depth discontinuities.
The evaluation produced several key findings. First, on synthetic benchmarks, the proposed pipeline achieved a geometric surface reconstruction error (Chamfer L1 distance of 0.0014) that is approximately 70% lower than mesh-based differentiable baselines (0.0048) and nearly eight times lower than volumetric neural methods (0.0111). Second, novel viewpoint and relighting quality improved substantially, yielding an SSIM score of 0.9747 compared to 0.9358 and 0.8252 for baseline techniques. Third, ablation testing showed that incorporating edge-aware gradient sampling prevents geometric distortion and failure along object silhouettes, while standard surface rendering models stall without edge handling. Finally, on real-world captures, the pipeline reconstructed sharp, detailed textures without the severe blurring found in purely volumetric approaches.
These findings indicate that integrating neural distance fields with edge-aware physics-based surface rendering eliminates the core trade-off between optimization flexibility and downstream utility. Production pipelines can deploy these reconstructed models directly into industry-standard graphics tools for raytracing, relighting, and material editing without requiring manual topological cleanup or remeshing.
Organizations seeking to streamline 3D asset generation should consider adopting hybrid neural inverse rendering for scanning physical inventory, provided the capture environment allows controlled co-located lighting. Future technical development should focus on extending the pipeline to handle ambient illumination, multi-bounce light reflections in concave regions, and non-opaque materials such as glass or translucent plastics.
- Paper: NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view Reconstruction, Peng Wang et al. (2021). NeuS establishes the foundational method of learning neural implicit signed distance functions via volume rendering for multi-view surface reconstruction, which IRON directly adopts in its first optimization stage.
- Paper: Volume Rendering of Neural Implicit Surfaces, Lior Yariv et al. (2021). VolSDF provides the theoretical and practical framework for converting signed distance fields into volume densities for differentiable rendering, forming a core precursor to IRON's hybrid SDF pipeline.
- Paper: NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis, Ben Mildenhall et al. (2020). NeRF introduces the coordinate-based volumetric radiance field optimization from 2D images that underlies the first stage of IRON's inverse rendering framework.
- Paper: DeepSDF: Learning Continuous Signed Distance Functions for Shape Representation, Jeong Joon Park et al. (2019). DeepSDF introduces continuous neural implicit representations of 3D geometry via signed distance functions, which IRON utilizes and optimizes for scene geometry.
- Paper: Ref-NeRF: Structured View-Dependent Appearance for Neural Radiance Fields, Dor Verbin et al. (2022). Ref-NeRF establishes techniques for structured view-dependent appearance and surface normal regularization in neural fields, directly informing IRON's material and lighting disentanglement.
- Paper: Magic3D: High-Resolution Text-to-3D Content Creation, Chen-Hsuan Lin et al. (2022). Magic3D builds on the coarse-to-fine paradigm of extracting and refining textured meshes via differentiable rendering to synthesize high-resolution 3D assets from text prompts.
- Paper: 2D Gaussian Splatting for Geometrically Accurate Radiance Fields, Binbin Huang et al. (2024). 2D Gaussian Splatting advances geometrically accurate surface reconstruction and differentiable rendering using explicit surface-aligned primitives as an alternative to implicit SDF optimization.
- Paper: LRM: Large Reconstruction Model for Single Image to 3D, Yicong Hong et al. (2024). LRM extends the challenge of 3D geometry and appearance reconstruction from multi-view per-scene optimization to a generalizable single-image feed-forward model.
- Paper: Structured 3D Latents for Scalable and Versatile 3D Generation, Jianfeng Xiang et al. (2025). TRELLIS explores a unified structured latent framework capable of decoding high-fidelity 3D assets into multiple representations including meshes and radiance fields.
