Neural Fields Meet Explicit Geometric Representations for Inverse Rendering of Urban Scenes
Zian WangTianchang ShenJun GaoShengyu HuangJacob MunkbergJon HasselgrenZan GojcicWenzheng ChenSanja Fidler
Presents a hybrid inverse rendering framework that couples neural fields for primary ray representation with explicit reconstructed meshes for efficient secondary ray tracing, enabling photorealistic relighting, shadow casting, and virtual object insertion in large-scale urban environments.
Digital twins of large outdoor environments are increasingly vital for simulation, augmented reality, and virtual content creation. While recent neural field techniques excel at creating photorealistic 3D views, they typically bake lighting and shadows directly into the reconstructed geometry, preventing users from altering illumination or seamlessly inserting virtual objects. Conversely, traditional 3D mesh rendering methods allow lighting modifications but fail to scale effectively to complex, expansive outdoor environments.
The article addresses this gap by introducing a hybrid inverse rendering framework named FEGR. Its core objective is to jointly extract precise 3D geometry, spatially varying material properties, and high dynamic range lighting from ordinary sets of captured images, enabling flexible relighting and virtual modifications in large urban settings.
FEGR combines continuous neural fields with explicit 3D mesh structures. The pipeline uses neural fields to capture fine surface geometry, colors, and material attributes, while extracting an explicit 3D mesh to trace complex secondary light bounces and cast shadows using hardware-accelerated physics-based ray tracing. The framework was evaluated on multi-illumination outdoor benchmark datasets as well as real-world autonomous driving data captured under single lighting conditions with moving cameras and LiDAR depth assistance.
The key findings demonstrate that FEGR consistently outperforms leading baseline methods across both visual quality and lighting reconstruction accuracy. On the multi-illumination outdoor benchmark, the framework achieved significantly higher image fidelity, reducing reconstruction error compared to earlier approaches. Ablation tests showed that physically modeling shadows and exposure compensation contributed performance gains of up to 1.5 decibels in visual quality. In real-world driving captures, FEGR successfully disentangled true surface colors from shadows where competing methods failed. Furthermore, in a formal user study evaluating virtual object insertion, human evaluators preferred FEGR's photorealism and shadow accuracy over state-of-the-art baselines by margins of roughly 69% to 86%.
These findings indicate that hybrid rendering offers a viable path to creating editable, high-fidelity digital replicas of real-world outdoor environments. By accurately separating environmental lighting from surface materials, the approach reduces the need for expensive synthetic asset modeling and expands the usability of captured video in autonomous vehicle simulation and spatial computing. Unlike previous techniques limited to single objects or fixed lighting, this formulation handles large-scale, complex scenes captured under single or multiple illumination passes.
For practical implementation, organizations looking to build relightable digital twins should consider hybrid deferred pipelines that blend neural fields with explicit geometry rather than relying solely on volumetric rendering. However, decision-makers should recognize current limitations: the framework relies on static scene assumptions and handcrafted semantic regularizations to resolve lighting ambiguities in single-capture datasets. Further development is needed to incorporate dynamic object handling and data-driven priors before deploying the system in rapidly changing, non-static environments.
- Paper: IRON: Inverse Rendering by Optimizing Neural SDFs and Materials from Photometric Images, Kai Zhang et al. (2022). Presents a foundational inverse rendering framework that couples neural signed distance fields and materials for surface extraction and physics-based rendering.
- Paper: Ref-NeRF: Structured View-Dependent Appearance for Neural Radiance Fields, Dor Verbin et al. (2022). Introduces structured view-dependent appearance and reflection modeling in neural fields, directly informing material and specular decomposition techniques.
- Paper: NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view Reconstruction, Peng Wang et al. (2021). Formulates the neural implicit surface volume rendering framework that enables recovering high-fidelity explicit geometric meshes from posed multi-view images.
- Paper: Volume Rendering of Neural Implicit Surfaces, Lior Yariv et al. (2021). Develops the foundational theory for transforming learnable signed distance functions into volumetric density fields for robust geometry reconstruction.
- Paper: NeRF in the Wild: Neural Radiance Fields for Unconstrained Photo Collections, Ricardo Martin-Brualla et al. (2021). Establishes techniques for handling unconstrained outdoor photo collections by disentangling appearance variations and transient elements from static scene geometry.
- Paper: Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance Fields, Jonathan T. Barron et al. (2022). Introduces space contraction and proposal sampling for unbounded 360-degree scenes, critical for scaling neural field rendering to outdoor urban environments.
- Paper: Rendering synthetic objects into real scenes: bridging traditional and image-based graphics with global illumination and high dynamic range photography, P. Debevec (1998). Establishes the seminal differential rendering formulation for seamlessly inserting synthetic objects and casting shadows into real captured environments.
- Paper: NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis, Ben Mildenhall et al. (2020). Introduces neural radiance fields and differentiable volume rendering upon which neural scene representation and novel view synthesis are built.
- Paper: VastGaussian: Vast 3D Gaussians for Large Scene Reconstruction, Jiaqi Lin et al. (2024). Scales explicit rendering primitives to vast outdoor scenes while decoupling appearance variations and illumination shifts across large environments.
- Paper: 2D Gaussian Splatting for Geometrically Accurate Radiance Fields, Binbin Huang et al. (2024). Advances explicit primitive rendering by employing oriented 2D surfaces to dramatically improve reconstructed 3D surface geometry fidelity.
- Paper: DL3DV-10K: A Large-Scale Scene Dataset for Deep Learning-based 3D Vision, Lu Ling et al. (2024). Introduces a massive benchmark spanning complex real-world outdoor lighting and material phenomena to rigorously evaluate large-scale 3D novel view synthesis.
