Neural Fields Meet Explicit Geometric Representations for Inverse Rendering of Urban Scenes

Zian WangTianchang ShenJun GaoShengyu HuangJacob MunkbergJon HasselgrenZan GojcicWenzheng ChenSanja Fidler

article2023CVPR125 citations

Presents a hybrid inverse rendering framework that couples neural fields for primary ray representation with explicit reconstructed meshes for efficient secondary ray tracing, enabling photorealistic relighting, shadow casting, and virtual object insertion in large-scale urban environments.

Listen

Digital twins of large outdoor environments are increasingly vital for simulation, augmented reality, and virtual content creation. While recent neural field techniques excel at creating photorealistic 3D views, they typically bake lighting and shadows directly into the reconstructed geometry, preventing users from altering illumination or seamlessly inserting virtual objects. Conversely, traditional 3D mesh rendering methods allow lighting modifications but fail to scale effectively to complex, expansive outdoor environments.

The article addresses this gap by introducing a hybrid inverse rendering framework named FEGR. Its core objective is to jointly extract precise 3D geometry, spatially varying material properties, and high dynamic range lighting from ordinary sets of captured images, enabling flexible relighting and virtual modifications in large urban settings.

FEGR combines continuous neural fields with explicit 3D mesh structures. The pipeline uses neural fields to capture fine surface geometry, colors, and material attributes, while extracting an explicit 3D mesh to trace complex secondary light bounces and cast shadows using hardware-accelerated physics-based ray tracing. The framework was evaluated on multi-illumination outdoor benchmark datasets as well as real-world autonomous driving data captured under single lighting conditions with moving cameras and LiDAR depth assistance.

The key findings demonstrate that FEGR consistently outperforms leading baseline methods across both visual quality and lighting reconstruction accuracy. On the multi-illumination outdoor benchmark, the framework achieved significantly higher image fidelity, reducing reconstruction error compared to earlier approaches. Ablation tests showed that physically modeling shadows and exposure compensation contributed performance gains of up to 1.5 decibels in visual quality. In real-world driving captures, FEGR successfully disentangled true surface colors from shadows where competing methods failed. Furthermore, in a formal user study evaluating virtual object insertion, human evaluators preferred FEGR's photorealism and shadow accuracy over state-of-the-art baselines by margins of roughly 69% to 86%.

These findings indicate that hybrid rendering offers a viable path to creating editable, high-fidelity digital replicas of real-world outdoor environments. By accurately separating environmental lighting from surface materials, the approach reduces the need for expensive synthetic asset modeling and expands the usability of captured video in autonomous vehicle simulation and spatial computing. Unlike previous techniques limited to single objects or fixed lighting, this formulation handles large-scale, complex scenes captured under single or multiple illumination passes.

For practical implementation, organizations looking to build relightable digital twins should consider hybrid deferred pipelines that blend neural fields with explicit geometry rather than relying solely on volumetric rendering. However, decision-makers should recognize current limitations: the framework relies on static scene assumptions and handcrafted semantic regularizations to resolve lighting ambiguities in single-capture datasets. Further development is needed to incorporate dynamic object handling and data-driven priors before deploying the system in rapidly changing, non-static environments.

arXiv: 2304.03266
Cover for Neural Fields Meet Explicit Geometric Representations for Inverse Rendering of Urban Scenes

Abstract

Reconstruction and intrinsic decomposition of scenes from captured imagery would enable many applications such as relighting and virtual object insertion. Recent NeRF based methods achieve impressive fidelity of 3D reconstruction, but bake the lighting and shadows into the radiance field, while mesh-based methods that facilitate intrinsic decomposition through differentiable rendering have not yet scaled to the complexity and scale of outdoor scenes. We present a novel inverse rendering framework for large urban scenes capable of jointly reconstructing the scene geometry, spatially-varying materials, and HDR lighting from a set of posed RGB images with optional depth. Specifically, we use a neural field to account for the primary rays, and use an explicit mesh (reconstructed from the underlying neural field) for modeling secondary rays that produce higher-order lighting effects such as cast shadows. By faithfully disentangling complex geometry and materials from lighting effects, our method enables photorealistic relighting with specular and shadow effects on several outdoor datasets. Moreover, it supports physics-based scene manipulations such as virtual object insertion with ray-traced shadow casting.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Neural Intrinsic Scene Representation
  • 3.2. Hybrid Deferred Rendering
  • 3.3. Optimizing the Neural Scene Representation
  • 4. Experiments
  • 4.1. Datasets and evaluation setting
  • 4.2. Evaluation of Inverse Rendering
  • 4.3. Application to virtual object insertion
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — FEGR hybrid inverse-rendering pipeline

    model/method

    FEGR reconstructs an urban scene from posed RGB images, optionally supplemented with LiDAR depth, by jointly estimating its geometry, spatially varying materials, and HDR illumination. The method combines two representations: a neural intrinsic field renders primary camera rays at high spatial resolution, while an explicit mesh extracted from the learned geometry handles secondary rays and their visibility. The resulting differentiable hybrid renderer supports novel-view synthesis, relighting under new environment maps, and physics-based virtual object insertion with cast shadows. A single intrinsic scene representation is shared across all captures; when the input contains multiple illumination conditions, each condition receives its own HDR environment map.

  2. Knowl 2 — Neural intrinsic scene and lighting representation

    model/method

    FEGR represents the intrinsic scene properties with a neural field FϕF_\phi that maps a 3D location x∈R3\mathbf{x}\in\mathbb{R}^3 to a signed-distance value s∈Rs\in\mathbb{R}, surface normal n∈R3\mathbf{n}\in\mathbb{R}^3, base color kd∈R3\mathbf{k}_d\in\mathbb{R}^3, and two PBR material parameters ks∈R2\mathbf{k}_s\in\mathbb{R}^2. The two components of ks\mathbf{k}_s are roughness and metallicity under the Disney material model. In implementation, separate MLPs predict s=fSDF(x;θSDF)s=f_{\mathrm{SDF}}(\mathbf{x};\theta_{\mathrm{SDF}}), n=fnorm(x;θnorm)\mathbf{n}=f_{\mathrm{norm}}(\mathbf{x};\theta_{\mathrm{norm}}), and (kd,ks)=fmat(x;θmat)(\mathbf{k}_d,\mathbf{k}_s)=f_{\mathrm{mat}}(\mathbf{x};\theta_{\mathrm{mat}}); all three use multi-resolution hash positional encoding.

    The outdoor illumination is represented by an HDR sky dome at infinity. An MLP fenv(d;θenv)f_{\mathrm{env}}(\mathbf{d};\theta_{\mathrm{env}}) maps a 2D environment-direction coordinate d\mathbf{d} to RGB HDR intensity e∈R3\mathbf{e}\in\mathbb{R}^3. For MM illumination conditions, FEGR retains one shared intrinsic field and learns MM separate HDR sky maps.

  3. Knowl 3 — Neural volumetric rendering of the primary-ray G-buffer

    equation

    For an image of height hh and width ww, FEGR renders a G-buffer G∈Rh×w×8G\in\mathbb{R}^{h\times w\times 8} containing a normal map N∈Rh×w×3N\in\mathbb{R}^{h\times w\times 3}, base-color map Kd∈Rh×w×3K_d\in\mathbb{R}^{h\times w\times 3}, material map M∈Rh×w×2M\in\mathbb{R}^{h\times w\times 2}, and radial-depth map D∈Rh×wD\in\mathbb{R}^{h\times w}. For a camera ray with origin o∈R3\mathbf{o}\in\mathbb{R}^3, unit direction u∈R3\mathbf{u}\in\mathbb{R}^3, and point r(t)=o+tu\mathbf{r}(t)=\mathbf{o}+t\mathbf{u}, the rendered base color is

    Kd(r)=∫tntfT(t) ρ(r(t)) kd(r(t)) dt,T(t)=exp⁡(−∫tntρ(r(q)) dq),K_d(\mathbf{r})=\int_{t_n}^{t_f}T(t)\,\rho(\mathbf{r}(t))\,\mathbf{k}_d(\mathbf{r}(t))\,dt, \qquad T(t)=\exp\left(-\int_{t_n}^{t}\rho(\mathbf{r}(q))\,dq\right),

    where tnt_n and tft_f are the near and far ray bounds, T(t)T(t) is accumulated transmittance, and ρ\rho is the opaque volume density. Density is obtained from the learned SDF fSDFf_{\mathrm{SDF}} as

    ρ(r(t))=max⁡(−ddtΦκ ⁣(fSDF(r(t)))Φκ ⁣(fSDF(r(t))),0),Φκ(z)=Sigmoid⁡(κz),\rho(\mathbf{r}(t))=\max\left( \frac{-\frac{d}{dt}\Phi_\kappa\!\left(f_{\mathrm{SDF}}(\mathbf{r}(t))\right)} {\Phi_\kappa\!\left(f_{\mathrm{SDF}}(\mathbf{r}(t))\right)},0\right), \qquad \Phi_\kappa(z)=\operatorname{Sigmoid}(\kappa z),

    with learnable scalar sharpness parameter κ\kappa. Normals and material parameters are alpha-composited analogously. The radial depth used to locate a surface point is

    D(r)=∫tntfT(t) ρ(r(t)) t dt.D(\mathbf{r})=\int_{t_n}^{t_f}T(t)\,\rho(\mathbf{r}(t))\,t\,dt.

    This neural volume-rendering stage supplies the primary-ray appearance and the per-pixel intrinsic quantities required by the later shading stage.

  4. Knowl 4 — Explicit-mesh secondary-ray shading with visibility

    model/method

    FEGR applies a physically based shading pass to each G-buffer pixel. The pixel's normal, base color, and material values are combined with its rendered depth to obtain a 3D surface point x\mathbf{x}. Marching cubes extracts an explicit surface mesh SS from the optimized signed-distance field. The outgoing radiance in direction ωo\boldsymbol{\omega}_o is evaluated with the non-emissive rendering equation

    Lo(x,ωo)=∫Ωfr(x,ωo,ωi) Li(x,ωi) ∣n⋅ωi∣ dωi,L_o(\mathbf{x},\boldsymbol{\omega}_o)=\int_{\Omega} f_r(\mathbf{x},\boldsymbol{\omega}_o,\boldsymbol{\omega}_i)\,L_i(\mathbf{x},\boldsymbol{\omega}_i)\,\left|\mathbf{n}\cdot\boldsymbol{\omega}_i\right|\,d\boldsymbol{\omega}_i,

    where Ω\Omega is the incident hemisphere, frf_r is the simplified Disney BRDF, n\mathbf{n} is the surface normal, and LiL_i is incoming radiance. For an incident direction ωi\boldsymbol{\omega}_i, mesh ray tracing gives binary visibility vi(x,ωi,S)v_i(\mathbf{x},\boldsymbol{\omega}_i,S): it is 00 when the ray is blocked by SS and 11 otherwise. Incoming radiance is therefore

    Li(x,ωi)=vi(x,ωi,S) fenv(ωi;θenv),L_i(\mathbf{x},\boldsymbol{\omega}_i)=v_i(\mathbf{x},\boldsymbol{\omega}_i,S)\,f_{\mathrm{env}}(\boldsymbol{\omega}_i;\theta_{\mathrm{env}}),

    so cast shadows and other higher-order effects arise from explicit visibility rather than being baked into diffuse color. The implementation traces 512 secondary rays per shading point, importance-sampling both the BSDF and the HDR environment map and combining the samples with multiple importance sampling. OptiX accelerates the ray–mesh queries. During optimization, the mesh is regenerated every 20 iterations as the SDF changes; after optimization, the environment network can be evaluated once per texel of an exported HDR environment image for more efficient importance sampling.

  5. Knowl 5 — End-to-end objective and geometric supervision

    model/method

    FEGR optimizes the neural scene field and HDR lighting jointly using observed RGB images and, when available, LiDAR ranges. The total objective is

    L=Lrender+λdepthLdepth+λradLrad+λnormLnorm+λshadeLshade+λregLreg,\mathcal{L}=\mathcal{L}_{\mathrm{render}}+\lambda_{\mathrm{depth}}\mathcal{L}_{\mathrm{depth}}+\lambda_{\mathrm{rad}}\mathcal{L}_{\mathrm{rad}}+\lambda_{\mathrm{norm}}\mathcal{L}_{\mathrm{norm}}+\lambda_{\mathrm{shade}}\mathcal{L}_{\mathrm{shade}}+\lambda_{\mathrm{reg}}\mathcal{L}_{\mathrm{reg}},

    where each λ\lambda is a balancing weight. For a batch of camera rays R\mathcal{R}, the primary image loss is an RGB L1 error,

    Lrender=1∣R∣∑r∈R∣Crender(r)−Cgt(r)∣,\mathcal{L}_{\mathrm{render}}=\frac{1}{|\mathcal{R}|}\sum_{\mathbf{r}\in\mathcal{R}}\left|\mathbf{C}_{\mathrm{render}}(\mathbf{r})-\mathbf{C}_{\mathrm{gt}}(\mathbf{r})\right|,

    with rendered and ground-truth RGB values Crender\mathbf{C}_{\mathrm{render}} and Cgt\mathbf{C}_{\mathrm{gt}}. An auxiliary radiance MLP frad(x,u;θrad)f_{\mathrm{rad}}(\mathbf{x},\mathbf{u};\theta_{\mathrm{rad}}) is volumetrically rendered to produce Crad(r)\mathbf{C}_{\mathrm{rad}}(\mathbf{r}) and supplies geometry supervision through

    Lrad=1∣R∣∑r∈R∣Crad(r)−Cgt(r)∣.\mathcal{L}_{\mathrm{rad}}=\frac{1}{|\mathcal{R}|}\sum_{\mathbf{r}\in\mathcal{R}}\left|\mathbf{C}_{\mathrm{rad}}(\mathbf{r})-\mathbf{C}_{\mathrm{gt}}(\mathbf{r})\right|.

    The auxiliary radiance field is discarded after optimization. For LiDAR rays Rd\mathcal{R}_d, with measured radial distances Dgt(r)D_{\mathrm{gt}}(\mathbf{r}), depth supervision is

    Ldepth=1∣Rd∣∑r∈Rd∣D(r)−Dgt(r)∣.\mathcal{L}_{\mathrm{depth}}=\frac{1}{|\mathcal{R}_d|}\sum_{\mathbf{r}\in\mathcal{R}_d}\left|D(\mathbf{r})-D_{\mathrm{gt}}(\mathbf{r})\right|.

    To preserve both detailed rendered normals and smooth SDF geometry, FEGR compares the volumetrically rendered normal nx\mathbf{n}_{\mathbf{x}} with the normalized SDF gradient n~x=−∇xfSDF(x)/∥∇xfSDF(x)∥\widetilde{\mathbf{n}}_{\mathbf{x}}=-\nabla_{\mathbf{x}}f_{\mathrm{SDF}}(\mathbf{x})/\|\nabla_{\mathbf{x}}f_{\mathrm{SDF}}(\mathbf{x})\| using

    Lnorm=1∣R∣∑r∈Rcos⁡−1(∣n~x⋅nx∣).\mathcal{L}_{\mathrm{norm}}=\frac{1}{|\mathcal{R}|}\sum_{\mathbf{r}\in\mathcal{R}}\cos^{-1}\left(\left|\widetilde{\mathbf{n}}_{\mathbf{x}}\cdot\mathbf{n}_{\mathbf{x}}\right|\right).

    Additional regularizers Lreg\mathcal{L}_{\mathrm{reg}} constrain the ill-posed intrinsic decomposition, while Lshade\mathcal{L}_{\mathrm{shade}} prevents shadows from being absorbed into diffuse albedo.

  6. Knowl 6 — Semantic shading prior and two-stage optimization

    model/method

    To discourage the optimizer from explaining cast shadows by darkening the diffuse albedo, FEGR introduces one learnable RGB albedo ksemb∈R3\mathbf{k}_{\mathrm{sem}}^b\in\mathbb{R}^3 for each of BB semantic classes. Let Rb\mathcal{R}_b be the camera rays assigned to class bb by an off-the-shelf semantic segmentation network, let sdiffuse(r)s_{\mathrm{diffuse}}(\mathbf{r}) be the deferred diffuse shading, and let Cgt(r)\mathbf{C}_{\mathrm{gt}}(\mathbf{r}) be the observed RGB value. The semantic shading penalty is

    Lshade=1B∑b=1B1∣Rb∣∑r∈Rb∣Cdiffuseb(r)−Cgt(r)∣,Cdiffuseb(r)=ksembsdiffuse(r).\mathcal{L}_{\mathrm{shade}}=\frac{1}{B}\sum_{b=1}^{B}\frac{1}{|\mathcal{R}_b|}\sum_{\mathbf{r}\in\mathcal{R}_b}\left|\mathbf{C}_{\mathrm{diffuse}}^b(\mathbf{r})-\mathbf{C}_{\mathrm{gt}}(\mathbf{r})\right|, \qquad \mathbf{C}_{\mathrm{diffuse}}^b(\mathbf{r})=\mathbf{k}_{\mathrm{sem}}^b s_{\mathrm{diffuse}}(\mathbf{r}).

    Because one albedo is shared by all pixels in a semantic class, this restricted albedo model cannot reproduce individual shadow patterns easily; the optimization is consequently encouraged to explain those patterns through the HDR environment and geometry. FEGR first initializes geometry by optimizing with the auxiliary radiance supervision alone, then optimizes the scene intrinsics and lighting using the full objective.

  7. Knowl 7 — Multi-illumination outdoor relighting performance and ablations

    data/table

    On the NeRF-OSR outdoor dataset, FEGR was optimized on 13, 12, and 11 illumination sessions for three scenes, then evaluated by relighting each scene with environment maps from five held-out sessions. Dynamic objects, sky, and vegetation were excluded using semantic masks. The following results compare average PSNR, where higher is better, and MSE, where lower is better. The ablations isolate the explicit-mesh-only representation, removal of secondary-ray shadows, and removal of per-channel exposure compensation.

    Site 1 Site 2 Site 3
    Method PSNR MSE PSNR MSE PSNR MSE
    NeRF-OSR 19.34 0.012 16.35 0.027 15.66 0.029
    Ours 21.53 0.007 17.00 0.023 17.57 0.018
    Ours (mesh only) 18.94 0.013 16.50 0.025 16.86 0.021
    Ours (w/o shadow) 20.62 0.009 16.17 0.028 16.15 0.024
    Ours (w/o exposure) 20.70 0.009 16.70 0.025 16.09 0.025

    The complete FEGR model has the best PSNR and MSE on all three sites, outperforming NeRF-OSR in every listed metric. Using only the reconstructed mesh degrades performance substantially, showing that high-resolution neural-field rendering is important for primary-ray appearance. Removing shadow ray tracing or exposure compensation also lowers performance, demonstrating that explicit visibility and photometric correction contribute independently to relighting quality.

  8. Knowl 8 — Single-illumination inverse rendering on autonomous-driving scenes

    empirical result

    FEGR was evaluated on two autonomous-driving scene collections captured under only one illumination condition: a 20-second Waymo Open Dataset clip recorded by five pinhole cameras and a 64-beam LiDAR at 10 Hz, and an in-house RoadData collection recorded by eight 3848×21683848\times2168 cameras and a 128-beam LiDAR. For RoadData, only the front-facing camera with a 120∘120^\circ field of view was used; LiDAR supplied depth supervision for both collections. The scenes span environments up to approximately 200 m×200 m200\,\mathrm{m}\times200\,\mathrm{m}, include complex geometry, occlusions, spatially varying materials, unknown potentially high-intensity sunlight, and motion blur or HDR artifacts from fast capture motion of approximately 10 m/s10\,\mathrm{m/s} on the Waymo sequence.

    Compared qualitatively with Nvdiffrecmc, FEGR reconstructs cleaner base colors, more accurate geometry, and higher-resolution environment lighting while separating cast shadows from diffuse albedo. It produces photo-realistic re-rendered views and relighting results despite the single-illumination capture. The comparison is especially relevant because Nvdiffrecmc was designed for outside-looking-in, 360-degree view coverage, whereas these driving captures provide restricted inside-looking-out views; relying heavily on sparse or incorrect LiDAR depth in that setting can produce artifacts on unobserved surfaces such as windows.

  9. Knowl 9 — Virtual object insertion with ray-traced shadows

    empirical result

    FEGR uses its recovered HDR sky illumination and explicit scene mesh to insert a virtual car into the autonomous-driving scenes. The recovered dominant sunlight direction produces shadows whose locations and boundaries agree with the surrounding real scene, while the estimated materials and environment support realistic reflections.

    A user study compared FEGR with the methods of Hold-Geoffroy et al. and Wang et al. Each comparison showed participants two augmented images containing the same inserted car, presented in random order; participants judged which result was more photorealistic based on cast-shadow and reflection quality. For each baseline comparison, 9 users evaluated 29 examples, and the majority vote determined the preference for each example. FEGR was preferred in 86.2% of comparisons against Hold-Geoffroy et al. and in 68.9% against Wang et al., indicating a strong user preference for its lighting and insertion results.

  10. Knowl 10 — Limitations of FEGR

    limitation

    FEGR remains dependent on manually designed priors and regularization terms because inverse rendering is highly ill-posed, particularly when only a single illumination condition is available. The method therefore does not guarantee a unique or physically correct decomposition in all cases. It is also limited to static scenes, as its neural-field representation does not model scene dynamics; extending it with dynamic neural representations is identified as future work.

Coverage note — No substantial contributed material was deliberately omitted; implementation details from the appendix were not expanded because they are supplementary rather than load-bearing contributions.

References

  1. 1.Harry Barrow, J Tenenbaum, A Hanson, and E Riseman. Recovering intrinsic scene characteristics. Comput. Vis. Syst, 2:3–26, 1978. 1, 2
  2. 2.Sean Bell, Kavita Bala, and Noah Snavely. Intrinsic images in the wild. ACM Transactions on Graphics (TOG), 33(4):159, 2014. 2
  3. 3.Sai Bi, Zexiang Xu, Pratul Srinivasan, Ben Mildenhall, Kalyan Sunkavalli, Miloš Hašan, Yannick Hold-Geoffroy, David Kriegman, and Ravi Ramamoorthi. Neural reflectance fields for appearance acquisition. arXiv preprint arXiv:2008.03824, 2020. 2
  4. 4.Sai Bi, Zexiang Xu, Kalyan Sunkavalli, Miloš Hašan, Yannick Hold-Geoffroy, David Kriegman, and Ravi Ramamoorthi. Deep reflectance volumes: Relightable reconstructions from multi-view photometric images. In ECCV, pages 294–311. Springer, 2020. 2, 3
  5. 5.Boming Zhao and Bangbang Yang, Zhenyang Li, Zuoyue Li, Guofeng Zhang, Jiashu Zhao, Dawei Yin, Zhaopeng Cui, and Hujun Bao. Factorized and controllable neural re-rendering of outdoor scene for photo extrapolation. In Proceedings of the 30th ACM International Conference on Multimedia, 2022. 2
  6. 6.Mark Boss, Raphael Braun, Varun Jampani, Jonathan T. Barron, Ce Liu, and Hendrik P.A. Lensch. Nerd: Neural reflectance decomposition from image collections. In ICCV, 2021. 2, 3
  7. 7.Mark Boss, Andreas Engelhardt, Abhishek Kar, Yuanzhen Li, Deqing Sun, Jonathan T. Barron, Hendrik P.A. Lensch, and Varun Jampani. SAMURAI: Shape And Material from Unconstrained Real-world Arbitrary Image collections. In Advances in Neural Information Processing Systems (NeurIPS), 2022. 2
  8. 8.Mark Boss, Varun Jampani, Raphael Braun, Ce Liu, Jonathan T. Barron, and Hendrik P.A. Lensch. Neural-pil: Neural pre-integrated lighting for reflectance decomposition. In Advances in Neural Information Processing Systems (NeurIPS), 2021. 2
  9. 9.Mark Boss, Varun Jampani, Kihwan Kim, Hendrik P.A. Lensch, and Jan Kautz. Two-shot spatially-varying brdf and shape estimation. In CVPR, 2020. 2
  10. 10.Adrien Bousseau, Sylvain Paris, and Frédo Durand. User-assisted intrinsic images. In ACM Transactions on Graphics (TOG), volume 28, page 130. ACM, 2009. 2
  11. 11.Brent Burley. Physically-based shading at disney. 2012. 4
  12. 12.Wenzheng Chen, Huan Ling, Jun Gao, Edward Smith, Jaako Lehtinen, Alec Jacobson, and Sanja Fidler. Learning to predict 3d objects with an interpolation-based differentiable renderer. In NeurIPS, 2019. 3
  13. 13.Wenzheng Chen, Joey Litalien, Jun Gao, Zian Wang, Clement Fuji Tsang, Sameh Khalis, Or Litany, and Sanja Fidler. DIB-R++: Learning to predict lighting and material with a hybrid differentiable renderer. In NeurIPS, 2021. 2, 3, 5
  14. 14.Sara Fridovich-Keil, Alex Yu, Matthew Tancik, Qinhong Chen, Benjamin Recht, and Angjoo Kanazawa. Plenoxels: Radiance fields without neural networks. In CVPR, 2022. 1
  15. 15.Hang Gao, Ruilong Li, Shubham Tulsiani, Bryan Russell, and Angjoo Kanazawa. Monocular dynamic view synthesis: A reality check. In NeurIPS, 2022. 8
  16. 16.Mathieu Garon, Kalyan Sunkavalli, Sunil Hadap, Nathan Carr, and Jean-François Lalonde. Fast spatially-varying indoor lighting estimation. In CVPR, pages 6908–6917, 2019. 3
  17. 17.Roger Grosse, Micah K Johnson, Edward H Adelson, and William T Freeman. Ground truth dataset and baseline evaluations for intrinsic image algorithms. In ICCV, pages 2335–2342. IEEE, 2009. 2
  18. 18.Jon Hasselgren, Nikolai Hofmann, and Jacob Munkberg. Shape, Light, and Material Decomposition from Images using Monte Carlo Rendering and Denoising. arXiv:2206.03380, 2022. 2, 3, 5, 6, 7
  19. 19.Yannick Hold-Geoffroy, Akshaya Athawale, and Jean-François Lalonde. Deep sky modeling for single image outdoor lighting estimation. In CVPR, pages 6927–6935, 2019. 3, 6, 8
  20. 20.Yannick Hold-Geoffroy, Kalyan Sunkavalli, Sunil Hadap, Emiliano Gambaretto, and Jean-François Lalonde. Deep outdoor illumination estimation. In CVPR, pages 7312–7321, 2017. 3
  21. 21.James T Kajiya. The rendering equation. In Proceedings of the 13th annual conference on Computer graphics and interactive techniques, pages 143–150, 1986. 4
  22. 22.Balazs Kovacs, Sean Bell, Noah Snavely, and Kavita Bala. Shading annotations in the wild. In CVPR, pages 6998–7007, 2017. 2
  23. 23.Zhengfei Kuang, Kyle Olszewski, Menglei Chai, Zeng Huang, Panos Achlioptas, and Sergey Tulyakov. NeROIC: Neural object capture and rendering from online image collections. Computing Research Repository (CoRR), abs/2201.02533, 2022. 2
  24. 24.Edwin H Land and John J McCann. Lightness and retinex theory. Josa, 61(1):1–11, 1971. 2
  25. 25.Chloe LeGendre, Wan-Chun Ma, Graham Fyffe, John Flynn, Laurent Charbonnel, Jay Busch, and Paul Debevec. Deeplight: Learning illumination for unconstrained mobile mixed reality. In CVPR, pages 5918–5928, 2019. 3
  26. 26.Quewei Li, Jie Guo, Yang Fei, Feichao Li, and Yanwen Guo. Neulighting: Neural lighting for free viewpoint outdoor scene relighting with unconstrained photo collections. In Soon Ki Jung, Jehee Lee, and Adam W. Bargteil, editors, SIGGRAPH Asia 2022 Conference Papers, SA 2022, Daegu, Republic of Korea, December 6-9, 2022, pages 13:1–13:9. ACM, 2022. 3
  27. 27.Zhengqin Li, Mohammad Shafiei, Ravi Ramamoorthi, Kalyan Sunkavalli, and Manmohan Chandraker. Inverse rendering for complex indoor scenes: Shape, spatially-varying lighting and svbrdf from a single image. In CVPR, pages 2475–2484, 2020. 2
  28. 28.Zhengqi Li and Noah Snavely. Cgintrinsics: Better intrinsic image decomposition through physically-based rendering. In ECCV, pages 371–387, 2018. 2
  29. 29.Zhengqin Li, Zexiang Xu, Ravi Ramamoorthi, Kalyan Sunkavalli, and Manmohan Chandraker. Learning to reconstruct shape and spatially-varying reflectance from a single image. ACM Transactions on Graphics (TOG), 37(6):1–11, 2018. 3
  30. 30.Zhengqin Li, Ting-Wei Yu, Shen Sang, Sarah Wang, Sai Bi, Zexiang Xu, Hong-Xing Yu, Kalyan Sunkavalli, Miloš Hašan, Ravi Ramamoorthi, et al. Openrooms: An end-to-end open framework for photorealistic indoor scene datasets. arXiv preprint arXiv:2007.12868, 2020. 2
  31. 31.Andrew Liu, Shiry Ginosar, Tinghui Zhou, Alexei A. Efros, and Noah Snavely. Learning to factorize and relight a city. In ECCV, 2020. 3
  32. 32.William E. Lorensen and Harvey E. Cline. Marching cubes: A high resolution 3d surface construction algorithm. SIGGRAPH Comput. Graph., 21(4):163–169, aug 1987. 4
  33. 33.Ricardo Martin-Brualla, Noha Radwan, Mehdi S. M. Sajjadi, Jonathan T. Barron, Alexey Dosovitskiy, and Daniel Duckworth. NeRF in the Wild: Neural Radiance Fields for Unconstrained Photo Collections. In arXiv, 2020. 3
  34. 34.Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. arXiv preprint arXiv:2003.08934, 2020. 1, 2, 3
  35. 35.Thomas Müller, Alex Evans, Christoph Schied, and Alexander Keller. Instant neural graphics primitives with a multiresolution hash encoding. ACM Transactions on Graphics (TOG), 41(4):102:1–102:15, July 2022. 1, 4
  36. 36.Jacob Munkberg, Jon Hasselgren, Tianchang Shen, Jun Gao, Wenzheng Chen, Alex Evans, Thomas Mueller, and Sanja Fidler. Extracting Triangular 3D Models, Materials, and Lighting From Images. arXiv:2111.12503, 2021. 2, 3, 5
  37. 37.Merlin Nimier-David, Zhao Dong, Wenzel Jakob, and Anton Kaplanyan. Material and lighting reconstruction for complex indoor scenes with texture-space differentiable rendering. 2021. 3
  38. 38.Merlin Nimier-David, Delio Vicini, Tizian Zeltner, and Wenzel Jakob. Mitsuba 2: A retargetable forward and inverse renderer. ACM Transactions on Graphics (TOG), 38(6), Dec. 2019. 2
  39. 39.Steven G. Parker, James Bigler, Andreas Dietrich, Heiko Friedrich, Jared Hoberock, David Luebke, David McAllister, Morgan McGuire, Keith Morley, Austin Robison, and Martin Stich. Optix: A general purpose ray tracing engine. ACM Trans. Graph., 29(4), jul 2010. 2, 5
  40. 40.Viktor Rudnev, Mohamed Elgharib, William Smith, Lingjie Liu, Vladislav Golyanik, and Christian Theobalt. Nerf for outdoor scene relighting. In ECCV, 2022. 2, 3, 4, 6, 7
  41. 41.Soumyadip Sengupta, Jinwei Gu, Kihwan Kim, Guilin Liu, David W. Jacobs, and Jan Kautz. Neural inverse rendering of an indoor scene from a single image. In ICCV, 2019. 2
  42. 42.Pratul P. Srinivasan, Boyang Deng, Xiuming Zhang, Matthew Tancik, Ben Mildenhall, and Jonathan T. Barron. Nerv: Neural reflectance and visibility fields for relighting and view synthesis. In CVPR, 2021. 2, 3
  43. 43.Cheng Sun, Min Sun, and Hwann-Tzong Chen. Direct voxel grid optimization: Super-fast convergence for radiance fields reconstruction. In CVPR, 2022. 1
  44. 44.Pei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aurelien Chouard, Vijaysai Patnaik, Paul Tsui, James Guo, Yin Zhou, Yuning Chai, Benjamin Caine, et al. Scalability in perception for autonomous driving: Waymo open dataset. In CVPR, 2020. 6
  45. 45.Matthew Tancik, Vincent Casser, Xinchen Yan, Sabeek Pradhan, Ben Mildenhall, Pratul Srinivasan, Jonathan T. Barron, and Henrik Kretzschmar. Block-NeRF: Scalable large scene neural view synthesis. arXiv, 2022. 1
  46. 46.Jiajun Tang, Yongjie Zhu, Haoyu Wang, Jun-Hoong Chan, Si Li, and Boxin Shi. Estimating spatially-varying lighting in urban scenes with disentangled representation. In ECCV, 2022. 3
  47. 47.Andrew Tao, Karan Sapra, and Bryan Catanzaro. Hierarchical multi-scale attention for semantic segmentation. arXiv preprint arXiv:2005.10821, 2020. 6
  48. 48.Haithem Turki, Deva Ramanan, and Mahadev Satyanarayanan. Mega-nerf: Scalable construction of large-scale nerfs for virtual fly-throughs. In CVPR, pages 12922–12931, June 2022. 1
  49. 49.Eric Veach and Leonidas J Guibas. Optimally combining sampling techniques for monte carlo rendering. In Proceedings of the 22nd annual conference on Computer graphics and interactive techniques, pages 419–428, 1995. 5
  50. 50.Dor Verbin, Peter Hedman, Ben Mildenhall, Todd Zickler, Jonathan T. Barron, and Pratul P. Srinivasan. Ref-NeRF: Structured view-dependent appearance for neural radiance fields. CVPR, 2022. 3
  51. 51.Peng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt, Taku Komura, and Wenping Wang. Neus: Learning neural implicit surfaces by volume rendering for multi-view reconstruction. NeurIPS, 2021. 4
  52. 52.Zian Wang, Wenzheng Chen, David Acuna, Jan Kautz, and Sanja Fidler. Neural light field estimation for street scenes with differentiable virtual object insertion. In ECCV, 2022. 2, 3, 6, 8
  53. 53.Zian Wang, Jonah Philion, Sanja Fidler, and Jan Kautz. Learning indoor inverse rendering with 3d spatially-varying lighting. In ICCV, 2021. 2
  54. 54.Felix Wimbauer, Shangzhe Wu, and Christian Rupprecht. De-rendering 3d objects in the wild. In CVPR, 2022. 2
  55. 55.Yuanbo Xiangli, Linning Xu, Xingang Pan, Nanxuan Zhao, Anyi Rao, Christian Theobalt, Bo Dai, and Dahua Lin. Bungeenerf: Progressive neural radiance field for extreme multi-scale scene rendering. In ECCV, 2022. 1
  56. 56.Yao Yao, Jingyang Zhang, Jingbo Liu, Yihang Qu, Tian Fang, David McKinnon, Yanghai Tsin, and Long Quan. Neilf: Neural incident light field for physically-based material estimation. In European Conference on Computer Vision (ECCV), 2022. 3
  57. 57.Ye Yu and William AP Smith. Inverserendernet: Learning single image inverse rendering. In CVPR, 2019. 2
  58. 58.Jinsong Zhang, Kalyan Sunkavalli, Yannick Hold-Geoffroy, Sunil Hadap, Jonathan Eisenman, and Jean-François Lalonde. All-weather deep outdoor lighting estimation. In CVPR, pages 10158–10166, 2019. 3
  59. 59.Jason Y. Zhang, Gengshan Yang, Shubham Tulsiani, and Deva Ramanan. NeRS: Neural reflectance surfaces for sparse-view 3d reconstruction in the wild. In Conference on Neural Information Processing Systems, 2021. 3
  60. 60.Kai Zhang, Fujun Luan, Zhengqi Li, and Noah Snavely. Iron: Inverse rendering by optimizing neural sdfs and materials from photometric images. In CVPR, 2022. 2, 3
  61. 61.Kai Zhang, Fujun Luan, Qianqian Wang, Kavita Bala, and Noah Snavely. PhySG: Inverse rendering with spherical gaussians for physics-based material editing and relighting. In CVPR, 2021. 2, 3, 5
  62. 62.Kai Zhang, Gernot Riegler, Noah Snavely, and Vladlen Koltun. Nerf++: Analyzing and improving neural radiance fields, 2020. 4
  63. 63.Xiuming Zhang, Pratul P Srinivasan, Boyang Deng, Paul Debevec, William T Freeman, and Jonathan T Barron. Nerfactor: Neural factorization of shape and reflectance under an unknown illumination. ACM Transactions on Graphics (TOG), 40(6):1–18, 2021. 2, 3, 6
  64. 64.Yuanqing Zhang, Jiaming Sun, Xingyi He, Huan Fu, Rongfei Jia, and Xiaowei Zhou. Modeling indirect illumination for inverse rendering. In CVPR, 2022. 2, 3
  65. 65.Qi Zhao, Ping Tan, Qiang Dai, Li Shen, Enhua Wu, and Stephen Lin. A closed-form solution to retinex with nonlocal texture constraints. 34(7):1437–1444, 2012. 2
  66. 66.Yongjie Zhu, Yinda Zhang, Si Li, and Boxin Shi. Spatially-varying outdoor lighting estimation from intrinsics. In CVPR, 2021. 3

Citation

MLA
Wang, Z., et al. “Neural Fields Meet Explicit Geometric Representation for Inverse Rendering of Urban Scenes”. arXiv, 2023, http://arxiv.org/abs/2304.03266v1.
APA
Wang, Z., Shen, T., Gao, J., Huang, S., Munkberg, J., Hasselgren, J., Gojcic, Z., Chen, W., & Fidler, S. (2023). Neural Fields meet Explicit Geometric Representation for Inverse Rendering of Urban Scenes. arXiv. http://arxiv.org/abs/2304.03266v1
Chicago
Wang, Z., T. Shen, J. Gao, et al. 2023. “Neural Fields Meet Explicit Geometric Representation for Inverse Rendering of Urban Scenes”. arXiv. http://arxiv.org/abs/2304.03266v1.
Harvard
Wang, Z. et al. (2023) “Neural Fields meet Explicit Geometric Representation for Inverse Rendering of Urban Scenes”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2304.03266v1.
Vancouver
1. Wang Z, Shen T, Gao J, Huang S, Munkberg J, Hasselgren J, Gojcic Z, Chen W, Fidler S (2023) Neural Fields meet Explicit Geometric Representation for Inverse Rendering of Urban Scenes. arXiv

BibTeX

@article{wang2023neural,
  title = {Neural Fields meet Explicit Geometric Representation for Inverse Rendering of Urban Scenes},
  author = {Wang, Zian and Shen, Tianchang and Gao, Jun and Huang, Shengyu and Munkberg, Jacob and Hasselgren, Jon and Gojcic, Zan and Chen, Wenzheng and Fidler, Sanja},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2304.03266v1},
  eprint = {2304.03266}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE