Renderable Neural Radiance Map for Visual Navigation

Obin KwonJeongho ParkSonghwai Oh

article2023CVPR80 citationsHighlight

Proposes a 2D grid-based neural radiance map that embeds visual scene features into latent codes to enable real-time camera tracking, image-based localization, and image-goal movement across unseen environments without scene-specific retraining.

Listen

Autonomous visual navigation in unfamiliar indoor environments remains a significant challenge for robotics. Conventional approaches rely on simple geometric maps that lack rich visual information, while newer neural rendering methods are computationally intensive and require pre-training on specific target environments. This limits their practical deployment in real-time robotics where agents must quickly navigate unseen spaces under noisy real-world conditions.

The article introduces and evaluates the Renderable Neural Radiance Map, a novel grid-based spatial memory framework designed to capture 3D environmental appearance efficiently. The primary objective is to demonstrate that embedding visual latent codes into a 2D map enables real-time localization, novel view image synthesis, and robust image-goal navigation in unfamiliar environments without requiring per-scene optimization.

The authors designed a modular framework using a pre-trained encoder to convert color and depth camera images into latent vectors that are registered on a 2D grid, paired with a decoder capable of rendering 3D views. Evaluation was conducted through simulations using the Habitat platform across 72 indoor scenes from the Gibson dataset and standardized test paths. The system's performance was measured against leading reinforcement learning and modular baseline methods across camera tracking, image-based localization, and image-goal navigation tasks under sensor and movement noise.

The evaluation produced several key findings. First, the proposed framework achieved a 65.7% success rate in complex curved navigation scenarios, outperforming the existing state of the art by 18.6 percentage points. Second, the system operates with high computational efficiency, running mapping at 91.9 Hz and image localization at 56.8 Hz, while camera tracking operates at 5.0 Hz. Third, the localization framework achieved a 99% recall rate within a 50-centimeter error threshold in recorded maps and retained 97.4% accuracy even when one-third of the scene's objects changed. Finally, the framework demonstrated zero-shot generalization by successfully navigating and localizing in unseen environments without environment-specific fine-tuning.

These results demonstrate that structured spatial representations embedded with visual features provide a more sample-efficient and robust foundation for navigation than computationally demanding reinforcement learning policies. The high processing speeds and resilience to environmental alterations reduce operational risks and hardware overhead, making neural radiance concepts practical for real-time mobile robotics and automated facility inspections.

Organizations developing autonomous mobile systems should consider adopting grid-based neural radiance representations for vision-guided search and inspection tasks. Next steps should focus on deploying the framework onto physical robotic hardware to assess real-world sensor dynamics and developing graph-based optimization mechanisms to correct accumulated drift over extended operational timelines.

The findings are supported by comprehensive comparative simulations across standard industry benchmarks. However, confidence should be tempered by the fact that evaluations were conducted in simulated indoor environments with 3-degree-of-freedom movement. Additionally, the system faces limitations when correcting past mapping errors once features are registered, and performance degrades in environments with highly repetitive or visually ambiguous scenes.

arXiv: 2303.00304
Cover for Renderable Neural Radiance Map for Visual Navigation

Abstract

We propose a novel type of map for visual navigation, a renderable neural radiance map (RNR-Map), which is designed to contain the overall visual information of a 3D environment. The RNR-Map has a grid form and consists of latent codes at each pixel. These latent codes are embedded from image observations, and can be converted to the neural radiance field which enables image rendering given a camera pose. The recorded latent codes implicitly contain visual information about the environment, which makes the RNR-Map visually descriptive. This visual information in RNR-Map can be a useful guideline for visual localization and navigation. We develop localization and navigation frameworks that can effectively utilize the RNR-Map. We evaluate the proposed frameworks on camera tracking, visual localization, and image-goal navigation. Experimental results show that the RNR-Map-based localization framework can find the target location based on a single query image with fast speed and competitive accuracy compared to other baselines. Also, this localization framework is robust to environmental changes, and even finds the most visually similar places when a query image from a different environment is given. The proposed navigation framework outperforms the existing image-goal navigation methods in difficult scenarios, under odometry and actuation noises. The navigation framework shows 65.7% success rate in curved scenarios of the NRNS [21] dataset, which is an improvement of 18.6% over the current state-of-the-art. Project page: https://rllab-snu.github.io/projects/RNR-Map/

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. RNR-Map
  • 4. Localization
  • 4.1. Image-Based Localization
  • 4.2. Camera Tracking
  • 5. Navigation
  • 5.1. Mapping Module
  • 5.2. Localization Module
  • 5.3. Navigation Module
  • 5.4. Implementation Details
  • 6. Experiments
  • 6.1. Localization
  • 6.2. Image-Goal Navigation
  • 6.2.1 Baselines
  • 6.2.2 Task setup
  • 6.2.3 Image-goal navigation Results
  • 7. Conclusion
  • References

Knowls

  1. Knowl 1 — Renderable neural radiance map construction

    model/method

    An RNR-Map is a 2-D spatial grid m∈RU×V×Dm\in\mathbb{R}^{U\times V\times D} whose cell m(u,v)∈RDm(u,v)\in\mathbb{R}^{D} is a latent visual code for a spatial region, rather than an occupancy or semantic label. A robot observes an RGB-D image It∈RH×W×4I_t\in\mathbb{R}^{H\times W\times 4} at a 3-DoF pose pt=(xt,yt,at)p_t=(x_t,y_t,a_t) and uses a pretrained encoder to produce pixel features Ct∈RH×W×DC_t\in\mathbb{R}^{H\times W\times D}. Each feature is assigned to the corresponding map cell through inverse camera projection and spatial quantization.

    For pixel coordinates (h,w)(h,w), depth dh,wd_{h,w}, camera intrinsic matrix KK, extrinsic parameters [R∣t][R\mid t], and map resolution ss, the corresponding world position qh,w=(qx,qy,qz)⊤q_{h,w}=(q_x,q_y,q_z)^\top and grid coordinates (u,v)(u,v) are

    qh,w=dh,wR−1K−1[hw1]−t,u=round⁡(qxs),v=round⁡(qys).q_{h,w}=d_{h,w}R^{-1}K^{-1}\begin{bmatrix}h\\w\\1\end{bmatrix}-t, \qquad u=\operatorname{round}\left(\frac{q_x}{s}\right), \qquad v=\operatorname{round}\left(\frac{q_y}{s}\right).

    For the set Xu,vX_{u,v} of encoded pixel features assigned to cell (u,v)(u,v), the local map code and observation-count mask are

    Xu,v={ch,w∈Ct: (u,v)=(round⁡(qx(h,w)/s),round⁡(qy(h,w)/s))},X_{u,v}=\{c_{h,w}\in C_t:\ (u,v)=\left(\operatorname{round}(q_x(h,w)/s),\operatorname{round}(q_y(h,w)/s)\right)\}, mtl(u,v)=1ntl(u,v)∑ci∈Xu,vci,ntl(u,v)=∣Xu,v∣.m_t^l(u,v)=\frac{1}{n_t^l(u,v)}\sum_{c_i\in X_{u,v}}c_i, \qquad n_t^l(u,v)=|X_{u,v}|.

    The global map is built incrementally without rendering. If a newly observed local cell and an existing global cell have counts ntl(u,v)n_t^l(u,v) and nt−1(u,v)n_{t-1}(u,v), respectively, their codes are fused by count-weighted averaging:

    mt(u,v)=mtl(u,v)ntl(u,v)+mt−1(u,v)nt−1(u,v)ntl(u,v)+nt−1(u,v),nt(u,v)=ntl(u,v)+nt−1(u,v).m_t(u,v)=\frac{m_t^l(u,v)n_t^l(u,v)+m_{t-1}(u,v)n_{t-1}(u,v)}{n_t^l(u,v)+n_{t-1}(u,v)}, \qquad n_t(u,v)=n_t^l(u,v)+n_{t-1}(u,v).

    The update is applied to observed cells; unobserved cells retain zero observation count. Thus repeated observations of the same spatial region accumulate a visual latent representation using odometry and known camera calibration.

  2. Knowl 2 — Latent-code decoding and reconstruction training

    model/method

    The RNR-Map encoder and decoder are trained jointly so that latent codes preserve enough visual information to reconstruct RGB-D observations. Given a map mm and a camera pose pp, the decoder samples the latent codes along the camera ray associated with every output pixel, converts those samples into locally conditioned radiance fields using modulation-linear-layer conditioning, and volume-renders an RGB-D image Fdec(m,p;θdec)F_{\mathrm{dec}}(m,p;\theta_{\mathrm{dec}}). The reconstruction workflow illustrated on page 3 consists of registration of observations into a map followed by radiance-field decoding from a query pose.

    For a sequence of NN training images IiI_i with poses pip_i, the encoder parameters θenc\theta_{\mathrm{enc}} first construct maps (mig,nig)(m_i^g,n_i^g) incrementally. The training objective is

    L(θenc,θdec)=1N∑i=1N∥Ii−Fdec(mNg,pi;θdec)∥1,\mathcal{L}(\theta_{\mathrm{enc}},\theta_{\mathrm{dec}}) =\frac{1}{N}\sum_{i=1}^{N}\left\|I_i-F_{\mathrm{dec}}(m_N^g,p_i;\theta_{\mathrm{dec}})\right\|_1,

    where mNgm_N^g is the final map made from the NN observations and ∥⋅∥1\|\cdot\|_1 is the elementwise RGB-D L1L_1 error. Because the radiance field is conditioned on image-derived latent codes rather than optimized for one fixed scene, the pretrained encoder-decoder can embed and render observations in previously unseen environments without additional scene-specific optimization. It can also synthesize novel viewpoints within observed 3-D regions; rendering is limited by observed points and produces gray output in unobserved regions.

  3. Knowl 3 — Cross-correlation image-based localization

    model/method

    The image-based localization function FlocF_{\mathrm{loc}} searches an RNR-Map for the region most visually related to a target RGB-D image ItrgI_{\mathrm{trg}}. Since the target pose is unknown, the target image is registered at the arbitrary origin p0=(0,0,0)p_0=(0,0,0) to produce mtrg=Freg(Itrg,p0;θenc)m_{\mathrm{trg}}=F_{\mathrm{reg}}(I_{\mathrm{trg}},p_0;\theta_{\mathrm{enc}}). Three learned components are then used: FkF_k embeds the current RNR-Map mm, FqF_q embeds the target map mtrgm_{\mathrm{trg}}, and FEF_E converts cross-correlation responses into a spatial probability heatmap and an orientation prediction.

    The target embedding is rotated through RR discrete orientations. If Rot⁡r\operatorname{Rot}_r denotes rotation by the rr-th angle in {0∘,360∘/R,…,360∘(R−1)/R}\{0^\circ,360^\circ/R,\ldots,360^\circ(R-1)/R\}, the localization function is

    Floc(m,mtrg;θloc)=FE(Conv⁡(Fk(m),{Rot⁡r(Fq(mtrg))}r=1R)),F_{\mathrm{loc}}(m,m_{\mathrm{trg}};\theta_{\mathrm{loc}}) =F_E\left(\operatorname{Conv}\left(F_k(m),\{\operatorname{Rot}_r(F_q(m_{\mathrm{trg}}))\}_{r=1}^{R}\right)\right),

    where Conv⁡\operatorname{Conv} is cross-correlation, θloc\theta_{\mathrm{loc}} contains the parameters of FkF_k, FqF_q, and FEF_E, and the output heatmap E∈RU×VE\in\mathbb{R}^{U\times V} assigns each map cell a probability of being the target location. The orientation head discretizes the target heading into 18 bins.

    Training uses a Gaussian ground-truth heatmap EgtE_{\mathrm{gt}} centered at the true map cell and a cross-entropy orientation target. For predicted heatmap E^\hat E and orientation a^trg\hat a_{\mathrm{trg}}, the loss is

    Lloc=DKL(Egt,E^)+CE(atrg,a^trg),\mathcal{L}_{\mathrm{loc}}=D_{\mathrm{KL}}(E_{\mathrm{gt}},\hat E)+CE(a_{\mathrm{trg}},\hat a_{\mathrm{trg}}),

    where DKLD_{\mathrm{KL}} is KL divergence and CECE is cross-entropy. The page-5 localization diagram depicts the same process: target-map encoding, multi-angle cross-correlation with the RNR-Map, and prediction of a spatial heatmap and target angle.

  4. Knowl 4 — Differentiable camera tracking under odometry noise

    model/method

    RNR-Map camera tracking refines a noisy odometric pose by minimizing photometric error between the current RGB-D observation and an image rendered from the previously built map. Let p^t−1\hat p_{t-1} be the previously estimated pose, let Δpt\Delta p_t be the measured relative odometry increment, and let ItI_t be the current observation. The initial pose estimate is pˉt=p^t−1+Δpt\bar p_t=\hat p_{t-1}+\Delta p_t. Using the differentiability of the decoder, the corrected pose is

    p^t=Ftrack(mt−1,pˉt)=arg⁡min⁡δpt∥Fdec(mt−1,pˉt+δpt)−It∥1,\hat p_t=F_{\mathrm{track}}(m_{t-1},\bar p_t) =\arg\min_{\delta p_t}\left\|F_{\mathrm{dec}}(m_{t-1},\bar p_t+\delta p_t)-I_t\right\|_1,

    where mt−1m_{t-1} is the map before incorporating the current observation and δpt\delta p_t is a 3-DoF pose correction. Gradient-based optimization is initialized at pˉt\bar p_t and can use only a small subset of image pixels to reduce computation, allowing tracking during navigation.

  5. Knowl 5 — RNR-Map image-goal navigation system

    model/method

    The proposed image-goal navigation system combines a persistent target representation, incremental RNR-Map construction, learned localization, camera tracking, and occupancy-based motion planning. At the beginning of an episode, the target image ItrgI_{\mathrm{trg}} is encoded into a target map mtrgm_{\mathrm{trg}}. During navigation, each current RGB-D observation updates the current RNR-Map mtm_t and a separate occupancy map; the occupancy map records occupied, free, and unseen cells and is used for collision avoidance.

    The localization module applies FlocF_{\mathrm{loc}} to mtm_t and mtrgm_{\mathrm{trg}} to obtain a target-likelihood heatmap EE, while FtrackF_{\mathrm{track}} corrects the agent pose before the next map update. For exploration, a generalized Voronoi graph is constructed over the occupancy map. Each graph node receives a latent score from the heatmap EE and an exploration score proportional to the number of unseen occupancy pixels in its neighborhood; their sum determines visitation priority. The highest-priority node is selected as an exploration target.

    A point-navigation controller follows the shortest occupancy-map path to the selected node and heuristically avoids obstacles. A learned stopper FstopF_{\mathrm{stop}} takes mtrgm_{\mathrm{trg}} and mtm_t and predicts whether the target has been reached. Once the target is detected, keypoint matching followed by Perspective-n-Point with RANSAC estimates the relative target pose, and the point-navigation controller performs the final approach. The navigation-system diagram on page 5 represents these mapping, localization, exploration, point-navigation, and stopping modules as a closed loop.

  6. Knowl 6 — Image-goal navigation evaluation protocol

    experimental setup

    The navigation experiments use the Habitat simulator with the Gibson and NRNS image-goal navigation datasets. Gibson contains 72 houses for training and 14 houses for validation; the authors collected 200 random navigation trajectories per training scene to train the encoder, decoder, localization networks, and stopper. The NRNS benchmark has easy, medium, and hard difficulty levels with straight and curved path types, with 1,000 episodes per difficulty-path combination except for hard-straight, which has 806 episodes.

    The agent receives only current RGB-D observations and odometry. Experiments use a directional camera with 90∘90^\circ horizontal field of view, noisy odometry and actuation, and four discrete actions: move forward 0.250.25 m, turn left 10∘10^\circ, turn right 10∘10^\circ, or stop. Each episode lasts at most 500 steps and is successful when the stop action is taken within 1 m of the target. Performance is reported using success rate (SR) and success weighted by path length (SPL), where SPL rewards successful paths while penalizing unnecessary path length.

  7. Knowl 7 — Image-goal navigation performance

    data/table

    The main quantitative comparison evaluates image-goal navigation under the noisy protocol. Each table entry is SR/SPL in percent; within each path type, the four columns are easy, medium, hard, and overall. The table reproduced from page 8 compares the proposed RNR-Map system with behavior cloning, reinforcement learning, occupancy-map exploration, topological exploration, and last-mile variants.

    Could not parse LaTeX table

    RNR-Map obtains the best curved-path overall result among the listed methods, with 65.7 SR and 40.8 SPL, compared with 55.4 SR and 37.4 SPL for OVRL + SLING and 43.7 SR and 14.3 SPL for NRNS + SLING. It also exceeds the other listed methods on curved easy, medium, and hard success rates, demonstrating that the learned latent target signal is especially useful when the route is not straight.

  8. Knowl 8 — Localization accuracy and robustness to visual change

    empirical result

    The RNR-Map localization components were evaluated for camera tracking and image-based localization. Camera tracking achieved a reported pose error of 0.1080.108 m while operating at 55 Hz, which the authors considered adequate for real-time navigation. Image-based localization recovered previously observed images with 99% recall when an inlier was required to be within 5050 cm.

    The localization method was also tested when the map and query image did not match exactly. In an object-change condition where 33.9% of observed images had changed, it localized the query within 5050 cm in 97.4% of cases and within 20∘20^\circ orientation error in 97.5% of cases. In a novel-environment condition, where the query image came from a different scene, the selected location achieved 94.5% of the best possible visual similarity available in the current scene. The authors report failures primarily when several places look alike or when the query image contains little visual information.

  9. Knowl 9 — Runtime and component ablation

    data/table

    The runtime and ablation results quantify whether RNR-Map can operate online and which components matter for navigation. Runtime was measured on a desktop with an Intel i7-9700KF CPU at 3.60 GHz and an NVIDIA GeForce RTX 2080 Ti GPU. The page-8 runtime table reports both frequency and per-call latency.

    Could not parse LaTeX table

    The ablation values below are averages over straight and curved scenarios; each entry is SR/SPL in percent. Noise indicates whether odometry and actuation noise are enabled, while FlocF_{\mathrm{loc}} and FtrackF_{\mathrm{track}} indicate whether the corresponding modules are active.

    Could not parse LaTeX table

    The results show that removing the target-localization heatmap causes a large performance drop, from 66.9/42.3 overall with all modules and noise to 59.8/39.4 when tracking is also removed, and to 52.6/33.7 when localization is absent in the no-noise comparison. Activating camera tracking improves noisy navigation from 59.8 SR and 39.4 SPL to 66.9 SR and 42.3 SPL. All major operations, including rendering, have runtimes compatible with online use, although rendering-based tracking is the slowest component.

  10. Knowl 10 — Odometry-error correction limitation

    limitation

    Once observations have been embedded into grid cells, the global RNR-Map is difficult to correct if the past odometry estimates were wrong: the latent codes have already been registered at potentially incorrect positions. The paper does not implement loop closure to resolve this problem. It proposes a future extension in which local RNR-Maps are connected by a pose graph and their poses are optimized jointly so that the rendered observations become geometrically consistent.

Coverage note — No substantial contribution was omitted; detailed network layer specifications and supplementary qualitative or MP3D analyses were left out because they provide implementation detail or corroboration rather than additional load-bearing method or findings.

References

  1. 1.Michal Adamkiewicz, Timothy Chen, Adam Caccavale, Rachel Gardner, Preston Culbertson, Jeannette Bohg, and Mac Schwager. Vision-Only Robot Navigation in a Neural Radiance World. IEEE Robotics and Automation Letters, 7(2):4606–4613, 2022.
  2. 2.Ziad Al-Halah, Santhosh K. Ramakrishnan, and Kristen Grauman. Zero Experience Required: Plug & Play Modular Transfer Learning for Semantic Visual Navigation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  3. 3.Peter Anderson, Angel X. Chang, Devendra Singh Chaplot, Alexey Dosovitskiy, Saurabh Gupta, Vladlen Koltun, Jana Kosecka, Jitendra Malik, Roozbeh Mottaghi, Manolis Savva, and Amir Roshan Zamir. On Evaluation of Embodied Navigation Agents. arXiv preprint arXiv:1807.06757, 2018.
  4. 4.Ivan Anokhin, Kirill Demochkin, Taras Khakhulin, Gleb Sterkin, Victor Lempitsky, and Denis Korzhenkov. Image Generators with Conditionally-Independent Pixel Synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  5. 5.Valts Blukis, Taeyeop Lee, Jonathan Tremblay, Bowen Wen, In So Kweon, Kuk-Jin Yoon, Dieter Fox, and Stan Birchfield. Neural Fields for Robotic Object Manipulation from a Single Image. arXiv preprint arXiv:2210.12126, 2022.
  6. 6.Arunkumar Byravan, Jan Humplik, Leonard Hasenclever, Arthur Brussee, Francesco Nori, Tuomas Haarnoja, Ben Moran, Steven Bohez, Fereshteh Sadeghi, Bojan Vujatovic, et al. NeRF2Real: Sim2real Transfer of Vision-guided Bipedal Motion Skills using Neural Radiance Fields. arXiv preprint arXiv:2210.04932, 2022.
  7. 7.Vincent Cartillier, Zhile Ren, Neha Jain, Stefan Lee, Irfan Essa, and Dhruv Batra. Semantic mapnet: Building allocentric semanticmaps and representations from egocentric views. arXiv preprint arXiv:2010.01191, 2020.
  8. 8.Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Halber, Matthias Niessner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang. Matterport3D: Learning from RGB-D Data in Indoor Environments. International Conference on 3D Vision (3DV), 2017.
  9. 9.Devendra Singh Chaplot, Dhiraj Gandhi, Abhinav Gupta, and Ruslan Salakhutdinov. Object Goal Navigation using Goal-Oriented Semantic Exploration. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), 2020.
  10. 10.Devendra Singh Chaplot, Dhiraj Gandhi, Saurabh Gupta, Abhinav Gupta, and Ruslan Salakhutdinov. Learning To Explore Using Active Neural SLAM. In Proceedings of the International Conference on Learning Representations (ICLR), 2020.
  11. 11.Tao Chen, Saurabh Gupta, and Abhinav Gupta. Learning Exploration Policies for Navigation. In Proceedings of the International Conference on Learning Representations (ICLR), 2019.
  12. 12.Qiyu Dai, Yan Zhu, Yiran Geng, Ciyu Ruan, Jiazhao Zhang, and He Wang. GraspNeRF: Multiview-based 6-DoF Grasp Detection for Transparent and Specular Objects Using Generalizable NeRF. arXiv preprint arXiv:2210.06575, 2022.
  13. 13.Terrance DeVries, Miguel Angel Bautista, Nitish Srivastava, Graham W. Taylor, and Joshua M. Susskind. Unconstrained scene generation with locally conditioned radiance fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021.
  14. 14.Danny Driess, Zhiao Huang, Yunzhu Li, Russ Tedrake, and Marc Toussaint. Learning Multi-Object Dynamics with Compositional Neural Radiance Fields. In Proceedings of the Conference on Robot Learning (CoRL), 2022.
  15. 15.Danny Driess, Ingmar Schubert, Pete Florence, Yunzhu Li, and Marc Toussaint. Reinforcement learning with neural radiance fields. arXiv preprint arXiv:2206.01634, 2022.
  16. 16.Jiafei Duan, Samson Yu, Hui Li Tan, Hongyuan Zhu, and Cheston Tan. A Survey of Embodied AI: From Simulators to Research Tasks. IEEE Transactions on Emerging Topics in Computational Intelligence, 2022.
  17. 17.Alberto Elfes. Using occupancy grids for mobile robot perception and navigation. Computer, 22(6):46–57, 1989.
  18. 18.Martin A Fischler and Robert C Bolles. Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography. Communications of the ACM, 24(6):381–395, 1981.
  19. 19.Georgios Georgakis, Bernadette Bucher, Karl Schmeckpeper, Siddharth Singh, and Kostas Daniilidis. Learning to Map for Active Semantic Goal Navigation. In Proceedings of the International Conference on Learning Representations (ICLR), 2022.
  20. 20.Saurabh Gupta, James Davidson, Sergey Levine, Rahul Sukthankar, and Jitendra Malik. Cognitive Mapping and Planning for Visual Navigation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2017.
  21. 21.Meera Hahn, Devendra Chaplot, Shubham Tulsiani, Mustafa Mukadam, James Rehg, and Abhinav Gupta. No RL, No Simulation: Learning to Navigate without Navigating. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), 2021.
  22. 22.Joao F. Henriques and Andrea Vedaldi. MapNet: An Allocentric Spatial Memory for Mapping Environments. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2018.
  23. 23.Jeffrey Ichnowski*, Yahav Avigal*, Justin Kerr, and Ken Goldberg. Dex-NeRF: Using a neural radiance field to grasp transparent objects. In Conference on Robot Learning (CoRL), 2020.
  24. 24.Jonghoek Kim, Fumin Zhang, and Magnus Egerstedt. A provably complete exploration strategy by constructing voronoi diagrams. Autonomous Robots, 29(3):367–380, 2010.
  25. 25.Vincent Lepetit, Francesc Moreno-Noguer, and Pascal Fua. Epnp: An accurate O(n) solution to the PnP problem. International Journal of Computer Vision, 81(2):155–166, 2009.
  26. 26.Yunzhu Li, Shuang Li, Vincent Sitzmann, Pulkit Agrawal, and Antonio Torralba. 3D Neural Scene Representations for Visuomotor Control. In Conference on Robot Learning, 2021.
  27. 27.Chen-Hsuan Lin, Wei-Chiu Ma, Antonio Torralba, and Simon Lucey. BARF: Bundle-Adjusting Neural Radiance Fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021.
  28. 28.Dominic Maggio, Marcus Abate, Jingnan Shi, Courtney Mario, and Luca Carlone. Loc-NeRF: Monte Carlo Localization using Neural Radiance Fields. arXiv preprint arXiv:2209.09050, 2022.
  29. 29.Manolis Savva, Abhishek Kadian, Oleksandr Maksymets, Yili Zhao, Erik Wijmans, Bhavana Jain, Julian Straub, Jia Liu, Vladlen Koltun, Jitendra Malik, Devi Parikh, and Dhruv Batra. Habitat: A Platform for Embodied AI Research. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019.
  30. 30.Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis. In Proceedings of the European Conference on Computer Vision (ECCV), 2020.
  31. 31.Keiji Nagatani and Howie Choset. Toward robust sensor based exploration by constructing reduced generalized voronoi graph. In Proceedings 1999 IEEE/RSJ International Conference on Intelligent Robots and Systems. Human and Environment Friendly Robots with High Intelligence and Emotional Quotients (Cat. No. 99CH36289), volume 3, pages 1687–1692. IEEE, 1999.
  32. 32.Emilio Parisotto and Ruslan Salakhutdinov. Neural Map: Structured Memory for Deep Reinforcement Learning. In International Conference on Learning Representations, 2018.
  33. 33.Benjamin Planche, Xuejian Rong, Ziyan Wu, Srikrishna Karanam, Harald Kosch, YingLi Tian, Jan Ernst, and Hutter Andreas. Incremental Scene Synthesis. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), 2019.
  34. 34.Ziad Al-Halah Santhosh Kumar Ramakrishnan and Kristen Grauman. Occupancy Anticipation for Efficient Exploration and Navigation. In Proceedings of the European Conference on Computer Vision (ECCV), 2020.
  35. 35.Edgar Sucar, Shikun Liu, Joseph Ortiz, and Andrew Davison. iMAP: Implicit Mapping and Positioning in Real-Time. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021.
  36. 36.Saim Wani, Shivansh Patel, Unnat Jain, Angel X. Chang, and Manolis Savva. MultiON: Benchmarking Semantic Map Memory using Multi-Object Navigation. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), 2020.
  37. 37.Justin Wasserman, Karmesh Yadav, Girish Chowdhary, Abhinav Gupta, and Unnat Jain. Last-Mile Embodied Visual Navigation. In Proceedings of the Conference on Robot Learning (CoRL), 2022.
  38. 38.Erik Wijmans, Abhishek Kadian, Ari Morcos, Stefan Lee, Irfan Essa, Devi Parikh, Manolis Savva, and Dhruv Batra. DD-PPO: Learning near-perfect pointgoal navigators from 2.5 billion frames. In Proceedings of the International Conference on Learning Representations (ICLR), 2020.
  39. 39.Fei Xia, Amir R. Zamir, Zhi-Yang He, Alexander Sax, Jitendra Malik, and Silvio Savarese. Gibson Env: Real-World Perception for Embodied Agents. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2018.
  40. 40.Karmesh Yadav, Ram Ramrakhya, Arjun Majumdar, Vincent-Pierre Berges, Sachit Kuhar, Dhruv Batra, Alexei Baevski, and Oleksandr Maksymets. Offline Visual Representation Learning for Embodied Navigation. arXiv preprint arXiv:2204.13226, 2022.
  41. 41.Qiwen Zhang, David Whitney, Florian Shkurti, and Ioannis Rekleitis. Ear-based exploration on hybrid metric/topological maps. In IEEE/RSJ International Conference on Intelligent Robots and Systems, 2014.
  42. 42.Zihan Zhu, Songyou Peng, Viktor Larsson, Weiwei Xu, Hujun Bao, Zhaopeng Cui, Martin R. Oswald, and Marc Pollefeys. NICE-SLAM: Neural Implicit Scalable Encoding for SLAM. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  43. 43.Xinkai Zuo, Fan Yang, Yifan Liang, Zhou Gang, Fei Su, Haihong Zhu, and Lin Li. An Improved Autonomous Exploration Framework for Indoor Mobile Robotics Using Reduced Approximated Generalized Voronoi Graphs. ISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information Sciences, 1:351–359, 2020.

Citation

MLA
Kwon, O., et al. “Renderable Neural Radiance Map for Visual Navigation”. arXiv, 2023, http://arxiv.org/abs/2303.00304v4.
APA
Kwon, O., Park, J., & Oh, S. (2023). Renderable Neural Radiance Map for Visual Navigation. arXiv. http://arxiv.org/abs/2303.00304v4
Chicago
Kwon, O., J. Park, and S. Oh. 2023. “Renderable Neural Radiance Map for Visual Navigation”. arXiv. http://arxiv.org/abs/2303.00304v4.
Harvard
Kwon, O., Park, J. and Oh, S. (2023) “Renderable Neural Radiance Map for Visual Navigation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2303.00304v4.
Vancouver
1. Kwon O, Park J, Oh S (2023) Renderable Neural Radiance Map for Visual Navigation. arXiv

BibTeX

@article{kwon2023renderable,
  title = {Renderable Neural Radiance Map for Visual Navigation},
  author = {Kwon, Obin and Park, Jeongho and Oh, Songhwai},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2303.00304v4},
  eprint = {2303.00304}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE