Reinforcement Learning with Neural Radiance Fields

Danny DriessIngmar SchubertPete FlorenceYunzhu LiMarc Toussaint

article2022NeurIPS72 citations

Demonstrates that using Neural Radiance Fields as a self-supervision signal to train visual encoders yields 3D-aware latent state representations that substantially improve the sample efficiency and generalization of reinforcement learning agents on complex robotic manipulation tasks.

Listen

Autonomous robotic systems often struggle to learn complex manipulation tasks directly from raw visual inputs, especially when object shapes vary. Conventional reinforcement learning methods frequently rely either on simplified, hand-engineered state estimates that fail to generalize across diverse object geometries or on standard two-dimensional image encoders that lack three-dimensional spatial awareness.

The article demonstrates that supervising state representation learning with Neural Radiance Fields—a computer vision technique that reconstructs three-dimensional scenes from two-dimensional images—significantly improves the sample efficiency and overall success rate of reinforcement learning agents.

The authors implemented a two-stage framework, termed NeRF-RL. First, an encoder-decoder architecture was pretrained offline using multi-view camera images collected from random environment interactions. The encoder compresses multi-camera views into a compact latent representation, while a latent-conditioned Neural Radiance Field decoder reconstructs the scene from novel viewpoints, embedding three-dimensional inductive biases into the representation. Second, the encoder was frozen and used directly as the input state for standard reinforcement learning algorithms across three challenging robotic simulation environments: hanging mugs with varying shapes on hooks, pushing differently sized objects, and opening sliding doors with variable handle geometries.

The findings show that representations trained with Neural Radiance Field decoders consistently outperformed all alternative methods. Agents using object-compositional Neural Radiance Field supervision achieved the highest success rates and learned faster than models using standard two-dimensional convolutional autoencoders, contrastive learning techniques, and even hand-engineered expert keypoints. In the door-opening task, the method achieved near-perfect task completion, whereas baseline methods plateaued below a fifty percent success rate or exhibited severe training instability. The researchers also confirmed that higher visual reconstruction quality directly correlated with improved policy performance.

These results indicate that embedding three-dimensional structural understanding into representation learning resolves critical visual ambiguities, such as occlusions and geometry variations, without requiring real-time three-dimensional rendering during operation. While pretraining the encoder required up to two days of compute compared to half a day for simpler contrastive models, the operational inference time remained extremely low at approximately seven milliseconds per step, introducing no runtime latency for deployed systems.

Organizations developing automated robotic manipulation should consider adopting Neural Radiance Field supervision for vision-based learning pipelines when geometric variability is high. Further work is recommended to validate the approach on physical hardware and investigate methods for updating representations online during live operations rather than relying entirely on offline pretraining.

arXiv: 2206.01634
Cover for Reinforcement Learning with Neural Radiance Fields

Abstract

It is a long-standing problem to find effective representations for training reinforcement learning (RL) agents. This paper demonstrates that learning state representations with supervision from Neural Radiance Fields (NeRFs) can improve the performance of RL compared to other learned representations or even low-dimensional, hand-engineered state information. Specifically, we propose to train an encoder that maps multiple image observations to a latent space describing the objects in the scene. The decoder built from a latent-conditioned NeRF serves as the supervision signal to learn the latent space. An RL algorithm then operates on the learned latent space as its state representation. We call this NeRF-RL. Our experiments indicate that NeRF as supervision leads to a latent space better suited for the downstream RL tasks involving robotic object manipulations like hanging mugs on hooks, pushing objects, or opening doors. Video: https://dannydriess.github.io/nerf-rl

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Background
  • 3.1 Reinforcement Learning
  • 3.2 Neural Radiance Fields (NeRFs)
  • 4 Learning State Representations for RL with NeRF Supervision
  • 4.1 Using Latent-Conditioned NeRF for RL
  • 4.2 Overview: Auto-Encoder with Latent-Conditioned NeRF Decoder
  • 4.3 Latent-Conditioned NeRF Decoder Details
  • 4.4 Encoder Details
  • 5 Baselines / Alternative State Representations
  • 6 Experiments
  • 6.1 Environments
  • 6.2 Results
  • 7 Discussion
  • 8 Conclusion
  • Acknowledgments
  • References
  • Checklist

Knowls

  1. Knowl 1 — NeRF-RL Framework for State Representation Learning

    model/method

    NeRF-RL is a framework that trains a visual state representation encoder for reinforcement learning (RL) using supervision from a latent-conditioned Neural Radiance Field (NeRF) decoder.

    The framework operates in two stages:

    1. Offline Representation Pretraining: An encoder Ω\Omega and a latent-conditioned (or object-compositional) NeRF decoder DD are trained end-to-end on an offline dataset of multi-view image observations collected via random interactions with the environment. The encoder maps multi-view posed images (and optional object segmentation masks) to a latent state space z∈Z⊂Rkz \in \mathcal{Z} \subset \mathbb{R}^k. The NeRF decoder provides the self-supervision signal by reconstructing the multi-view scene via volumetric rendering.
    2. Downstream Policy Optimization: The pretrained encoder Ω\Omega is frozen. An RL algorithm operates directly on the extracted latent vector z=Ω(y)z = \Omega(y) as its state representation to train a policy π(a∣z)\pi(a \mid z) and value/Q-functions on downstream manipulation tasks without fine-tuning the vision backbone.
  2. Knowl 2 — Object-Compositional Latent-Conditioned NeRF Decoder

    model/method

    In the object-compositional variant of NeRF-RL, each object j∈{1,…,m}j \in \{1, \dots, m\} in a scene is assigned an individual latent code zj∈Rkz_j \in \mathbb{R}^k. A single shared neural network ff predicts the volume density σj(x)∈R≥0\sigma_j(x) \in \mathbb{R}_{\ge 0} and emitted RGB color cj(x)∈R3c_j(x) \in \mathbb{R}^3 for any 3D world coordinate x∈R3x \in \mathbb{R}^3 conditioned on the object code:

    (σj(x),cj(x))=f(x,zj)(\sigma_j(x), c_j(x)) = f(x, z_j)

    To represent the full scene in 3D, the individual object radiance fields are composed into a global model via 3D spatial composition:

    σ(x)=∑j=1mσj(x),c(x)=1σ(x)∑j=1mσj(x)cj(x)\sigma(x) = \sum_{j=1}^m \sigma_j(x), \qquad c(x) = \frac{1}{\sigma(x)} \sum_{j=1}^m \sigma_j(x) c_j(x)

    Using the composed field (σ,c)(\sigma, c), an image I=D(z1:m,K)I = D(z_{1:m}, K) is synthesized for a target camera matrix KK by numerically integrating color along each camera ray r(α)=r(0)+αdr(\alpha) = r(0) + \alpha d (with ray origin r(0)r(0), direction dd, and near/far bounds αn,αf\alpha_n, \alpha_f) via volume rendering:

    C(r)=∫αnαfT(r,α)σ(r(α))c(r(α)) dα,T(r,α)=exp⁡(−∫αnασ(r(u)) du)C(r) = \int_{\alpha_n}^{\alpha_f} T(r, \alpha) \sigma(r(\alpha)) c(r(\alpha)) \, d\alpha, \qquad T(r, \alpha) = \exp\left(-\int_{\alpha_n}^\alpha \sigma(r(u)) \, du\right)

  3. Knowl 3 — Multi-View Image Reconstruction Loss for NeRF-RL Auto-Encoder

    equation

    Given an observation y=(I1:V,K1:V,M1:V)y = (I^{1:V}, K^{1:V}, M^{1:V}) consisting of RGB images Ii∈R3×h×wI^i \in \mathbb{R}^{3 \times h \times w}, camera projection matrices Ki∈R3×4K^i \in \mathbb{R}^{3 \times 4} (intrinsics and extrinsics), and per-view binary object masks Mji∈{0,1}h×wM_j^i \in \{0, 1\}^{h \times w} for VV camera views and mm objects, the latent vector for object jj is produced by zj=Ω(I1:V,K1:V,Mj1:V)∈Rkz_j = \Omega(I^{1:V}, K^{1:V}, M_j^{1:V}) \in \mathbb{R}^k.

    The auto-encoder parameters are trained by minimizing the L2L_2 image reconstruction loss over all views i∈{1,…,V}i \in \{1, \dots, V\}:

    Li=∥Ii∘Mtoti−D(Ω(I1:V,K1:V,M1:m1:V),Ki)∥22\mathcal{L}^i = \left\| I^i \circ M_{\text{tot}}^i - D\left(\Omega\left(I^{1:V}, K^{1:V}, M_{1:m}^{1:V}\right), K^i\right) \right\|_2^2

    where ∘\circ denotes the Hadamard (element-wise) product, Mtoti=⋁j=1mMjiM_{\text{tot}}^i = \bigvee_{j=1}^m M_j^i is the non-background scene mask for view ii, and D(z1:m,Ki)D(z_{1:m}, K^i) is the latent-conditioned NeRF rendering from viewpoint KiK^i.

  4. Knowl 4 — Multi-View Encoder Architectures for NeRF-RL

    model/method

    Two alternative encoder architectures map multi-view observations to object latent vectors zjz_j:

    1. Per-Image 2D CNN Encoder ("Image Encoder"): For each object jj, the 2D CNN encoder computes:

    zj=ΩCNN(I1:V,K1:V,Mj1:V)=hMLP(1V∑i=1VgMLP(ECNN(Ii∘Mji),Ki))z_j = \Omega_{\text{CNN}}\left(I^{1:V}, K^{1:V}, M_j^{1:V}\right) = h_{\text{MLP}}\left(\frac{1}{V} \sum_{i=1}^V g_{\text{MLP}}\left(E_{\text{CNN}}\left(I^i \circ M_j^i\right), K^i\right)\right)

    where ECNNE_{\text{CNN}} is a ResNet-18 feature extractor applied to the masked image Ii∘MjiI^i \circ M_j^i, concatenated with flattened camera matrix KiK^i, mapped by gMLPg_{\text{MLP}}, averaged across all VV views, and finally processed by MLP hMLPh_{\text{MLP}}. For global scenes without per-object masks, m=1m=1 and Mj=1i=MtotiM_{j=1}^i = M_{\text{tot}}^i.

    1. Neural Field 3D CNN Encoder ("Field Encoder"): A continuous neural vector field ϕ(I1:V,Mj1:V):R3→RE\phi(I^{1:V}, M_j^{1:V}): \mathbb{R}^3 \to \mathbb{R}^E is constructed by querying 3D coordinates and projecting them into camera coordinates of each view to extract pixel-aligned 2D features. The field is queried on a discrete 3D spatial grid Xh∈RdX×hX×wXX_h \in \mathbb{R}^{d_X \times h_X \times w_X} covering the workspace, and mapped to the latent vector via 3D convolutions:

    zj=E3D CNN(ϕ(I1:V,Mj1:V)(Xh))z_j = E_{\text{3D CNN}}\left(\phi\left(I^{1:V}, M_j^{1:V}\right)(X_h)\right)

  5. Knowl 5 — Empirical Superiority of NeRF-RL Over Visual Baselines and Keypoints

    empirical result

    Across sparse-reward manipulation tasks evaluated under shape and initial state randomization:

    1. Sample Efficiency and Asymptotic Performance: Compositional NeRF-RL (NeRF-RL comp. + field and NeRF-RL comp. + image) consistently outperforms 2D CNN autoencoders (Conv. Autoencoder, c and Conv. Autoencoder, g), single-view contrastive learning (CURL), multi-view contrastive learning (Multi-CURL), and ground-truth expert 3D keypoints.
    2. Mug on Hook: Compositional NeRF-RL achieves ∼80%\sim 80\% success within 2×1062 \times 10^6 transitions. CURL initially improves up to 59%59\% but degrades and destabilizes; global CNN autoencoders also become unstable, and keypoints reach only ∼60%\sim 60\%.
    3. Planar Pushing: Compositional NeRF-RL variants reach ∼80–85%\sim 80\text{--}85\% success within 1.5×1061.5 \times 10^6 transitions. The global NeRF-RL variant achieves ∼65%\sim 65\%, outperforming all non-NeRF baselines (which remain below 50%50\%).
    4. Door Opening: NeRF-RL comp. + field reaches ≈100%\approx 100\% success within 4×1054 \times 10^5 transitions. CURL achieves ∼70%\sim 70\%, expert keypoints plateau at ∼65%\sim 65\%, and CNN autoencoders fail to exceed 50%50\%.
  6. Knowl 6 — Dominance of Decoder 3D Inductive Bias Over Encoder Inductive Bias

    empirical result

    Comparing combinations of 2D CNN vs. 3D Neural Field encoders with NeRF vs. 2D CNN decoders reveals that the 3D inductive bias injected by the latent-conditioned NeRF decoder is the primary driver of RL performance:

    • Using a compositional NeRF decoder yields high RL success rates and sample efficiency regardless of whether the encoder is a standard 2D CNN (comp. + image) or a 3D neural field (comp. + field).
    • Replacing the NeRF decoder with a standard 2D deconvolutional CNN decoder (Conv. Autoencoder, c and g) or training the 3D field encoder with an InfoNCE contrastive loss (Multi-CURL) leads to severe performance degradation or failure to learn.
    • Incorporating 3D inductive biases into the encoder alone is insufficient; the volumetric 3D reconstruction supervision from the NeRF decoder is necessary to produce latent representations that generalize across varied object geometries.
  7. Knowl 7 — Robotic Manipulation Benchmarks for Shape-Generalizing RL

    experimental setup

    Evaluations are performed using Proximal Policy Optimization (PPO) on four camera views with known calibration matrices across three sparse-reward manipulation environments where object geometry varies per episode:

    1. Mug on Hook: The agent applies 3D translation actions to a mug to hang it stably on a hook. Mug shapes, handle openings, and hook geometries are randomized across distributions. Sparse reward R=1R=1 is provided if the mug remains stably hung without falling when tested under a virtual physics drop simulation, else R=0R=0.
    2. Planar Pushing: A pusher moves in 2D to push yellow box objects into a left table zone and blue box objects into a right table zone. Object dimensions, colors, and initial positions are randomized. Sparse reward R=1R=1 is assigned if the box reaches its target zone, else R=0R=0.
    3. Door Opening: A 3-DoF end-effector pushes a sliding door open. Handle position, height, and dimensions are randomized per episode. Sparse reward R=1R=1 is assigned if the door opens past a target displacement threshold, else R=0R=0.
  8. Knowl 8 — State Representation Learning Frameworks Evaluated in NeRF-RL

    data/table

    The state representations compared in the paper differ in their encoder architecture, decoder structure, compositional masking, usage of NeRF, and representation learning loss:

    Method Encoder Decoder Compositional NeRF Loss
    NeRF-RL comp.+field (Ours) 3D CNN comp. 3D NeRF ✓ ✓ image reconstr.: L2L_2
    NeRF-RL comp.+image (Ours) 2D CNN comp. 3D NeRF ✓ ✓ image reconstr.: L2L_2
    NeRF-RL global+image (Ours) 2D CNN global 3D NeRF ✗ ✓ image reconstr.: L2L_2
    Conv. Autoencoder, c 2D CNN comp. 2D CNN ✓ ✗ image reconstr.: L2L_2
    Conv. Autoencoder, g 2D CNN 2D CNN ✗ ✗ image reconstr.: L2L_2
    CURL 2D CNN – ✗ ✗ contrast: InfoNCE
    Multi-CURL 3D CNN – ✓ ✗ contrast: InfoNCE
    Keypoints Chosen by expert knowledge and perfect extraction –

    These configurations isolate the benefits of NeRF volume rendering supervision relative to deconvolutional pixel reconstruction, contrastive instance discrimination, and direct low-dimensional keypoints.

  9. Knowl 9 — Effects of Viewpoint Count and Pretraining Data Volume on NeRF-RL

    empirical result

    Ablation experiments on the representation learning stage show:

    1. Correlation with Rendering Quality: There is a direct positive correlation between the NeRF multi-view reconstruction quality (PSNR/reconstruction accuracy) of the pretrained auto-encoder and downstream RL task success.
    2. Sensitivity to Viewpoint Count: In 2D planar pushing, reducing camera views from 4 to 2 or 1 causes minimal performance loss. However, for 3D tasks requiring fine geometric alignment (e.g., Mug on Hook), reducing views from 4 to 2 significantly degrades downstream RL performance due to unresolved occlusions and loss of 3D geometry accuracy.
    3. Pretraining Dataset Size: Decreasing the amount of offline random interaction data used to pretrain the encoder does not significantly impair downstream RL performance, indicating sample-efficient representation learning.
  10. Knowl 10 — Computational and Operational Limitations of NeRF-RL

    limitation

    The NeRF-RL framework exhibits several practical constraints:

    1. Pretraining Computational Demand: Volumetric ray rendering makes training the NeRF decoder computationally intensive, taking up to 2 days on an NVIDIA RTX 3090 (compared to ≈0.5\approx 0.5 days for contrastive methods like CURL). Consequently, using NeRF reconstruction as an online auxiliary loss during RL is computationally prohibitive without reduced loss update frequencies.
    2. Offline Data Coverage: Pretraining relies on random exploration to cover the relevant state and shape space. In environments where target states cannot be reached by random actions, offline pretraining without guided exploration may be inadequate.
    3. Multi-View and Segmentation Priors: The approach assumes calibrated multi-view camera setups (known extrinsics and intrinsics) and requires accurate 2D instance segmentation masks to enable the object-compositional variant.

Coverage note — None was omitted; all primary contributions, models (compositional and global NeRF decoders, 2D CNN and 3D field encoders), mathematical formulations, experimental setups, comparative baseline results, ablations, and stated limitations are fully represented.

References

  1. 1.C. Finn, X. Y. Tan, Y. Duan, T. Darrell, S. Levine, and P. Abbeel. Deep spatial autoencoders for visuomotor learning. In 2016 IEEE International Conference on Robotics and Automation (ICRA), pages 512–519. IEEE, 2016.
  2. 2.R. Jonschkowski, R. Hafner, J. Scholz, and M. Riedmiller. Pves: Position-velocity encoders for unsupervised learning of structured state representations. arXiv preprint arXiv:1705.09805, 2017.
  3. 3.D. Dwibedi, J. Tompson, C. Lynch, and P. Sermanet. Learning actionable representations from visual observations. In 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 1577–1584. IEEE, 2018.
  4. 4.T. D. Kulkarni, A. Gupta, C. Ionescu, S. Borgeaud, M. Reynolds, A. Zisserman, and V. Mnih. Unsupervised learning of object keypoints for perception and control. Advances in neural information processing systems, 32, 2019.
  5. 5.M. Laskin, A. Srinivas, and P. Abbeel. Curl: Contrastive unsupervised representations for reinforcement learning. In International Conference on Machine Learning, pages 5639–5650. PMLR, 2020.
  6. 6.L. Manuelli, Y. Li, P. Florence, and R. Tedrake. Keypoints into the future: Self-supervised correspondence in model-based reinforcement learning. arXiv preprint arXiv:2009.05085, 2020.
  7. 7.M. Vecerik, J.-B. Regli, O. Sushkov, D. Barker, R. Pevceviciute, T. Rothörl, C. Schuster, R. Hadsell, L. Agapito, and J. Scholz. S3k: Self-supervised semantic keypoints for robotic manipulation via multi-view consistency. arXiv preprint arXiv:2009.14711, 2020.
  8. 8.L. Manuelli, W. Gao, P. Florence, and R. Tedrake. kpam: Keypoint affordances for category-level robotic manipulation. arXiv preprint arXiv:1903.06684, 2019.
  9. 9.V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller. Playing atari with deep reinforcement learning. arXiv preprint arXiv:1312.5602, 2013.
  10. 10.V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al. Human-level control through deep reinforcement learning. nature, 518(7540):529–533, 2015.
  11. 11.Y. Li, S. Li, V. Sitzmann, P. Agrawal, and A. Torralba. 3d neural scene representations for visuomotor control. In Conference on Robot Learning, pages 112–123. PMLR, 2022.
  12. 12.S. Lange and M. Riedmiller. Deep auto-encoder neural networks in reinforcement learning. In The 2010 International Joint Conference on Neural Networks (IJCNN), 2010.
  13. 13.A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  14. 14.A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  15. 15.I. Akinola, J. Varley, and D. Kalashnikov. Learning precise 3d manipulation from multiple uncalibrated cameras. In 2020 IEEE International Conference on Robotics and Automation (ICRA), pages 4616–4622. IEEE, 2020.
  16. 16.B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. In European conference on computer vision, pages 405–421. Springer, 2020.
  17. 17.F. Dellaert and L. Yen-Chen. Neural volume rendering: Nerf and beyond, 2021.
  18. 18.S. A. Eslami, D. Jimenez Rezende, F. Besse, F. Viola, A. S. Morcos, M. Garnelo, A. Ruderman, A. A. Rusu, I. Danihelka, K. Gregor, et al. Neural scene representation and rendering. Science, 360(6394):1204–1210, 2018.
  19. 19.J. J. Park, P. Florence, J. Straub, R. Newcombe, and S. Lovegrove. Deepsdf: Learning continuous signed distance functions for shape representation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 165–174, 2019.
  20. 20.L. Mescheder, M. Oechsle, M. Niemeyer, S. Nowozin, and A. Geiger. Occupancy networks: Learning 3d reconstruction in function space. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4460–4470, 2019.
  21. 21.Z. Chen and H. Zhang. Learning implicit fields for generative shape modeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5939–5948, 2019.
  22. 22.V. Sitzmann, M. Zollhöfer, and G. Wetzstein. Scene representation networks: Continuous 3d-structure-aware neural scene representations. Advances in Neural Information Processing Systems, 32, 2019.
  23. 23.Y. Xie, T. Takikawa, S. Saito, O. Litany, S. Yan, N. Khan, F. Tombari, J. Tompkin, V. Sitzmann, and S. Sridhar. Neural fields in visual computing and beyond. arXiv preprint arXiv:2111.11426, 2021.
  24. 24.A. Tewari, J. Thies, B. Mildenhall, P. Srinivasan, E. Tretschk, Y. Wang, C. Lassner, V. Sitzmann, R. Martin-Brualla, S. Lombardi, et al. Advances in neural rendering. arXiv preprint arXiv:2111.05849, 2021.
  25. 25.J. T. Barron, B. Mildenhall, M. Tancik, P. Hedman, R. Martin-Brualla, and P. P. Srinivasan. Mip-nerf: A multiscale representation for anti-aliasing neural radiance fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5855–5864, 2021.
  26. 26.J. T. Barron, B. Mildenhall, D. Verbin, P. P. Srinivasan, and P. Hedman. Mip-nerf 360: Unbounded anti-aliased neural radiance fields. arXiv preprint arXiv:2111.12077, 2021.
  27. 27.M. Tancik, V. Casser, X. Yan, S. Pradhan, B. Mildenhall, P. P. Srinivasan, J. T. Barron, and H. Kretzschmar. Block-nerf: Scalable large scene neural view synthesis. arXiv preprint arXiv:2202.05263, 2022.
  28. 28.M. Boss, R. Braun, V. Jampani, J. T. Barron, C. Liu, and H. Lensch. NeRD: Neural reflectance decomposition from image collections. https://arxiv.org/abs/2012.03918, 2020.
  29. 29.P. Srinivasan, B. Deng, X. Zhang, M. Tancik, B. Mildenhall, and J. T. Barron. NeRV: Neural reflectance and visibility fields for relighting and view synthesis. https://arxiv.org/abs/2012.03927, 2020.
  30. 30.X. Zhang, P. P. Srinivasan, B. Deng, P. Debevec, W. T. Freeman, and J. T. Barron. Nerfactor: Neural factorization of shape and reflectance under an unknown illumination. https://arxiv.org/abs/2106.01970, 2021.
  31. 31.L. Liu, J. Gu, K. Z. Lin, T.-S. Chua, and C. Theobalt. Neural sparse voxel fields. In Advances in Neural Information Processing Systems (NeurIPS), volume 33, 2020.
  32. 32.D. Lindell, J. Martel, and G. Wetzstein. AutoInt: Automatic integration for fast neural volume rendering. https://arxiv.org/abs/2012.01714, 2020.
  33. 33.D. Rebain, W. Jiang, S. Yazdani, K. Li, K. M. Yi, and A. Tagliasacchi. DeRF: Decomposed radiance fields. https://arxiv.org/abs/2011.12490, 2020.
  34. 34.T. Neff, P. Stadlbauer, M. Parger, A. Kurz, J. H. Mueller, C. R. A. Chaitanya, A. S. Kaplanyan, and M. Steinberger. DONeRF: Towards Real-Time Rendering of Compact Neural Radiance Fields using Depth Oracle Networks. Computer Graphics Forum, 40(4), 2021. ISSN 1467-8659. doi: 10.1111/cgf.14340. URL https://doi.org/10.1111/cgf.14340.
  35. 35.S. J. Garbin, M. Kowalski, M. Johnson, J. Shotton, and J. Valentin. Fastnerf: High-fidelity neural rendering at 200fps. https://arxiv.org/abs/2103.10380, 2021.
  36. 36.C. Reiser, S. Peng, Y. Liao, and A. Geiger. Kilonerf: Speeding up neural radiance fields with thousands of tiny mlps, 2021.
  37. 37.A. Yu, R. Li, M. Tancik, H. Li, R. Ng, and A. Kanazawa. Plenoctrees for real-time rendering of neural radiance fields. In arXiv, 2021.
  38. 38.S. Lombardi, T. Simon, G. Schwartz, M. Zollhoefer, Y. Sheikh, and J. Saragih. Mixture of volumetric primitives for efficient neural rendering, 2021.
  39. 39.A. Yu, S. Fridovich-Keil, M. Tancik, Q. Chen, B. Recht, and A. Kanazawa. Plenoxels: Radiance fields without neural networks. arXiv preprint arXiv:2112.05131, 2021.
  40. 40.V. Sitzmann, S. Rezchikov, W. T. Freeman, J. B. Tenenbaum, and F. Durand. Light field networks: Neural scene representations with single-evaluation rendering. In arXiv, 2021.
  41. 41.T. Müller, A. Evans, C. Schied, and A. Keller. Instant neural graphics primitives with a multiresolution hash encoding. ACM Trans. Graph., 41(4):102:1–102:15, July 2022. doi: 10.1145/3528223.3530127. URL https://doi.org/10.1145/3528223.3530127.
  42. 42.R. Martin-Brualla, N. Radwan, M. S. Sajjadi, J. T. Barron, A. Dosovitskiy, and D. Duckworth. Nerf in the wild: Neural radiance fields for unconstrained photo collections. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7210–7219, 2021.
  43. 43.A. Yu, V. Ye, M. Tancik, and A. Kanazawa. pixelnerf: Neural radiance fields from one or few images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4578–4587, 2021.
  44. 44.Q. Wang, Z. Wang, K. Genova, P. P. Srinivasan, H. Zhou, J. T. Barron, R. Martin-Brualla, N. Snavely, and T. Funkhouser. Ibrnet: Learning multi-view image-based rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4690–4699, 2021.
  45. 45.K. Zhang, G. Riegler, N. Snavely, and V. Koltun. NERF++: Analyzing and improving neural radiance fields. https://arxiv.org/abs/2010.07492, 2020.
  46. 46.M. Niemeyer and A. Geiger. GIRAFFE: Representing scenes as compositional generative neural feature fields. https://arxiv.org/abs/2011.12100, 2020.
  47. 47.M. Guo, A. Fathi, J. Wu, and T. Funkhouser. Object-centric neural scene rendering. https://arxiv.org/abs/2012.08503, 2020.
  48. 48.W. Yuan, Z. Lv, T. Schmidt, and S. Lovegrove. Star: Self-supervised tracking and reconstruction of rigid objects in motion with neural rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13144–13152, 2021.
  49. 49.Z. Wang, T. Bagautdinov, S. Lombardi, T. Simon, J. Saragih, J. Hodgins, and M. Zollhöfer. Learning compositional radiance fields of dynamic human heads. https://arxiv.org/abs/2012.09955, 2020.
  50. 50.J. Ost, F. Mannan, N. Thuerey, J. Knodt, and F. Heide. Neural scene graphs for dynamic scenes. https://arxiv.org/abs/2011.10379, 2020.
  51. 51.H.-X. Yu, L. J. Guibas, and J. Wu. Unsupervised discovery of object radiance fields. arXiv preprint arXiv:2107.07905, 2021.
  52. 52.B. Yang, Y. Zhang, Y. Xu, Y. Li, H. Zhou, H. Bao, G. Zhang, and Z. Cui. Learning object-compositional neural radiance field for editable scene rendering. In International Conference on Computer Vision (ICCV), October 2021.
  53. 53.D. Driess, Z. Huang, Y. Li, R. Tedrake, and M. Toussaint. Learning multi-object dynamics with compositional neural radiance fields. arXiv preprint arXiv:2202.11855, 2022.
  54. 54.H. Zhang, R. Wang, J. Zhang, C. Li, G. Yang, P. Spincemaille, T. Nguyen, and Y. Wang. Nerd: Neural representation of distribution for medical image segmentation. arXiv preprint arXiv:2103.04020, 2021.
  55. 55.L. Yen-Chen, P. Florence, J. T. Barron, A. Rodriguez, P. Isola, and T.-Y. Lin. iNeRF: Inverting neural radiance fields for pose estimation. IROS, 2021.
  56. 56.M. Adamkiewicz, T. Chen, A. Caccavale, R. Gardner, P. Culbertson, J. Bohg, and M. Schwager. Vision-only robot navigation in a neural radiance world. IEEE Robotics and Automation Letters, 7(2):4606–4613, 2022.
  57. 57.J. Ichnowski, Y. Avigal, J. Kerr, and K. Goldberg. Dex-nerf: Using a neural radiance field to grasp transparent objects. arXiv preprint arXiv:2110.14217, 2021.
  58. 58.L. Yen-Chen, P. Florence, J. T. Barron, T.-Y. Lin, A. Rodriguez, and P. Isola. NeRF-Supervision: Learning dense object descriptors from neural radiance fields. In IEEE Conference on Robotics and Automation (ICRA), 2022.
  59. 59.K. Karunratanakul, J. Yang, Y. Zhang, M. J. Black, K. Muandet, and S. Tang. Grasping field: Learning implicit representations for human grasps. In 2020 International Conference on 3D Vision (3DV), pages 333–344. IEEE, 2020.
  60. 60.J.-S. Ha, D. Driess, and M. Toussaint. Learning neural implicit functions as object representations for robotic manipulation. arXiv preprint arXiv:2112.04812, 2021.
  61. 61.A. Simeonov, Y. Du, A. Tagliasacchi, J. B. Tenenbaum, A. Rodriguez, P. Agrawal, and V. Sitzmann. Neural descriptor fields: Se (3)-equivariant object representations for manipulation. arXiv preprint arXiv:2112.05124, 2021.
  62. 62.Y. Wi, P. Florence, A. Zeng, and N. Fazeli. Virdo: Visio-tactile implicit representations of deformable objects. arXiv preprint arXiv:2202.00868, 2022.
  63. 63.D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al. Mastering the game of go with deep neural networks and tree search. nature, 529(7587):484–489, 2016.
  64. 64.T. Chen, S. Kornblith, M. Norouzi, and G. Hinton. A simple framework for contrastive learning of visual representations. In International conference on machine learning, pages 1597–1607. PMLR, 2020.
  65. 65.K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9729–9738, 2020.
  66. 66.A. Srinivas, M. Laskin, and P. Abbeel. Curl: Contrastive unsupervised representations for reinforcement learning. arXiv preprint arXiv:2004.04136, 2020.
  67. 67.B. You, O. Arenz, Y. Chen, and J. Peters. Integrating contrastive learning with dynamic models for reinforcement learning from images. Neurocomputing, 2022.
  68. 68.P. Sermanet, C. Lynch, Y. Chebotar, J. Hsu, E. Jang, S. Schaal, S. Levine, and G. Brain. Time-contrastive networks: Self-supervised learning from video. In 2018 IEEE International Conference on Robotics and Automation (ICRA), pages 1134–1141. IEEE, 2018.
  69. 69.A. Kinose, M. Okada, R. Okumura, and T. Taniguchi. Multi-view dreaming: Multi-view world model with contrastive learning. arXiv preprint arXiv:2203.11024, 2022.
  70. 70.K. Chen, Y. Lee, and H. Soh. Multi-modal mutual information (mummi) training for robust self-supervised deep reinforcement learning. In 2021 IEEE International Conference on Robotics and Automation (ICRA), pages 4274–4280. IEEE, 2021.
  71. 71.S. Nair, A. Rajeswaran, V. Kumar, C. Finn, and A. Gupta. R3m: A universal visual representation for robot manipulation. arXiv preprint arXiv:2203.12601, 2022.
  72. 72.A. Stooke, K. Lee, P. Abbeel, and M. Laskin. Decoupling representation learning from reinforcement learning. In International Conference on Machine Learning, pages 9870–9879. PMLR, 2021.
  73. 73.S. Parisi, A. Rajeswaran, S. Purushwalkam, and A. Gupta. The unsurprising effectiveness of pre-trained vision models for control. arXiv preprint arXiv:2203.03580, 2022.
  74. 74.D. Yarats, A. Zhang, I. Kostrikov, B. Amos, J. Pineau, and R. Fergus. Improving sample efficiency in model-free reinforcement learning from images. arXiv preprint arXiv:1910.01741, 2019.
  75. 75.A. Zhang, R. McAllister, R. Calandra, Y. Gal, and S. Levine. Learning invariant representations for reinforcement learning without reconstruction. arXiv preprint arXiv:2006.10742, 2020.
  76. 76.K. Zakka, A. Zeng, P. Florence, J. Tompson, J. Bohg, and D. Dwibedi. Xirl: Cross-embodiment inverse reinforcement learning. In Conference on Robot Learning, pages 537–546. PMLR, 2022.
  77. 77.T. Xiao, I. Radosavovic, T. Darrell, and J. Malik. Masked visual pre-training for motor control. arXiv preprint arXiv:2203.06173, 2022.
  78. 78.Y. Seo, K. Lee, S. James, and P. Abbeel. Reinforcement learning with action-free pre-training from videos. arXiv preprint arXiv:2203.13880, 2022.
  79. 79.T. Lesort, N. Díaz-Rodríguez, J.-F. Goudou, and D. Filliat. State representation learning for control: An overview. Neural Networks, 108:379–392, 2018.
  80. 80.M. Niemeyer and A. Geiger. Giraffe: Representing scenes as compositional generative neural feature fields. In Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2021.
  81. 81.K. Stelzner, K. Kersting, and A. R. Kosiorek. Decomposing 3d scenes into objects via unsupervised volume segmentation. arXiv preprint arXiv:2104.01148, 2021.
  82. 82.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
  83. 83.S. Saito, Z. Huang, R. Natsume, S. Morishima, A. Kanazawa, and H. Li. Pifu: Pixel-aligned implicit function for high-resolution clothed human digitization. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2304–2314, 2019.
  84. 84.R. Hadsell, S. Chopra, and Y. LeCun. Dimensionality reduction by learning an invariant mapping. In 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’06), volume 2, pages 1735–1742. IEEE, 2006.
  85. 85.A. Van den Oord, Y. Li, and O. Vinyals. Representation learning with contrastive predictive coding. arXiv e-prints, pages arXiv–1807, 2018.
  86. 86.J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017.
  87. 87.D. Driess, J.-S. Ha, M. Toussaint, and R. Tedrake. Learning models as functionals of signed-distance fields for manipulation planning. In Conference on Robot Learning (CoRL), 2021.
  88. 88.K. He, G. Gkioxari, P. Dollár, and R. Girshick. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision, pages 2961–2969, 2017.
  89. 89.A. Raffin, A. Hill, M. Ernestus, A. Gleave, A. Kanervisto, and N. Dormann. Stable baselines3. https://github.com/DLR-RM/stable-baselines3, 2019.

Citation

MLA
Driess, D., et al. “Reinforcement Learning with Neural Radiance Fields”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 16931–45, https://proceedings.neurips.cc/paper_files/paper/2022/file/6c294f059e3d77d58dbb8fe48f21fe00-Paper-Conference.pdf.
APA
Driess, D., Schubert, I., Florence, P., Li, Y., & Toussaint, M. (2022). Reinforcement Learning with Neural Radiance Fields. Advances in Neural Information Processing Systems, 35, 16931–16945. https://proceedings.neurips.cc/paper_files/paper/2022/file/6c294f059e3d77d58dbb8fe48f21fe00-Paper-Conference.pdf
Chicago
Driess, D., I. Schubert, P. Florence, Y. Li, and M. Toussaint. 2022. “Reinforcement Learning with Neural Radiance Fields”. Advances in Neural Information Processing Systems 35: 16931–45. https://proceedings.neurips.cc/paper_files/paper/2022/file/6c294f059e3d77d58dbb8fe48f21fe00-Paper-Conference.pdf.
Harvard
Driess, D. et al. (2022) “Reinforcement Learning with Neural Radiance Fields”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 16931–16945. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/6c294f059e3d77d58dbb8fe48f21fe00-Paper-Conference.pdf.
Vancouver
1. Driess D, Schubert I, Florence P, Li Y, Toussaint M (2022) Reinforcement Learning with Neural Radiance Fields. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 16931–16945

BibTeX

@inproceedings{driess2022reinforcement,
  title = {Reinforcement Learning with Neural Radiance Fields},
  author = {Driess, Danny and Schubert, Ingmar and Florence, Pete and Li, Yunzhu and Toussaint, Marc},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {16931-16945},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/6c294f059e3d77d58dbb8fe48f21fe00-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission