Reinforcement Learning with Neural Radiance Fields
Danny DriessIngmar SchubertPete FlorenceYunzhu LiMarc Toussaint
Demonstrates that using Neural Radiance Fields as a self-supervision signal to train visual encoders yields 3D-aware latent state representations that substantially improve the sample efficiency and generalization of reinforcement learning agents on complex robotic manipulation tasks.
Autonomous robotic systems often struggle to learn complex manipulation tasks directly from raw visual inputs, especially when object shapes vary. Conventional reinforcement learning methods frequently rely either on simplified, hand-engineered state estimates that fail to generalize across diverse object geometries or on standard two-dimensional image encoders that lack three-dimensional spatial awareness.
The article demonstrates that supervising state representation learning with Neural Radiance Fields—a computer vision technique that reconstructs three-dimensional scenes from two-dimensional images—significantly improves the sample efficiency and overall success rate of reinforcement learning agents.
The authors implemented a two-stage framework, termed NeRF-RL. First, an encoder-decoder architecture was pretrained offline using multi-view camera images collected from random environment interactions. The encoder compresses multi-camera views into a compact latent representation, while a latent-conditioned Neural Radiance Field decoder reconstructs the scene from novel viewpoints, embedding three-dimensional inductive biases into the representation. Second, the encoder was frozen and used directly as the input state for standard reinforcement learning algorithms across three challenging robotic simulation environments: hanging mugs with varying shapes on hooks, pushing differently sized objects, and opening sliding doors with variable handle geometries.
The findings show that representations trained with Neural Radiance Field decoders consistently outperformed all alternative methods. Agents using object-compositional Neural Radiance Field supervision achieved the highest success rates and learned faster than models using standard two-dimensional convolutional autoencoders, contrastive learning techniques, and even hand-engineered expert keypoints. In the door-opening task, the method achieved near-perfect task completion, whereas baseline methods plateaued below a fifty percent success rate or exhibited severe training instability. The researchers also confirmed that higher visual reconstruction quality directly correlated with improved policy performance.
These results indicate that embedding three-dimensional structural understanding into representation learning resolves critical visual ambiguities, such as occlusions and geometry variations, without requiring real-time three-dimensional rendering during operation. While pretraining the encoder required up to two days of compute compared to half a day for simpler contrastive models, the operational inference time remained extremely low at approximately seven milliseconds per step, introducing no runtime latency for deployed systems.
Organizations developing automated robotic manipulation should consider adopting Neural Radiance Field supervision for vision-based learning pipelines when geometric variability is high. Further work is recommended to validate the approach on physical hardware and investigate methods for updating representations online during live operations rather than relying entirely on offline pretraining.
- Paper: NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis, Ben Mildenhall et al. (2020). Introduces Neural Radiance Fields (NeRF), the foundational 3D volumetric scene representation that the source relies on and integrates into reinforcement learning.
- Paper: Scene Representation Networks: Continuous 3D-Structure-Aware Neural Scene Representations, V. Sitzmann et al. (2019). Establishes continuous, structure-aware neural scene representations learned directly from 2D images, providing key groundwork for neural rendering in visual understanding.
- Paper: CURL: Contrastive Unsupervised Representations for Reinforcement Learning, Aravind Srinivas et al. (2020). Demonstrates self-supervised auxiliary representation learning for visual reinforcement learning, serving as a primary baseline and conceptual predecessor to the source.
- Paper: Dream to Control: Learning Behaviors by Latent Imagination, Danijar Hafner et al. (2019). Introduces visual representation learning and latent imagination for continuous control, motivating visual feature extraction for RL policies.
- Paper: End-to-End Training of Deep Visuomotor Policies, Sergey Levine et al. (2015). Provides foundational principles for learning robotic visuomotor policies directly from raw camera images.
- Paper: QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation, Dmitry Kalashnikov et al. (2018). Pioneers scalable vision-based deep reinforcement learning for robotic manipulation, establishing the core control challenges that the source addresses.
- Paper: D-NeRF: neural radiance fields for dynamic scenes, Albert Pumarola et al. (2021). Extends neural radiance fields to dynamic scenes, highlighting spatial-temporal representations relevant to moving robotic manipulation environments.
- Paper: Renderable Neural Radiance Map for Visual Navigation, Obin Kwon et al. (2023). Extends neural radiance field representations into 2D spatial memory maps for downstream robotic visual navigation without requiring per-scene optimization.
- Paper: On Pre-Training for Visuo-Motor Control: Revisiting a Learning-from-Scratch Baseline, Nicklas Hansen et al. (2023). Critically re-evaluates the efficacy and robustness of frozen pre-trained visual representations versus learning from scratch across robotic manipulation tasks.
- Paper: VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training, Yecheng Jason Ma et al. (2023). Builds upon self-supervised visual pretraining to provide both visual state representations and universal zero-shot reward functions for robotic control.
- Paper: RobustNeRF: Ignoring Distractors with Robust Losses, Sara Sabour et al. (2023). Improves neural radiance field reconstruction under dynamic distractors and occlusions commonly encountered in autonomous robotic environments.
- Paper: OpenNeRF: Open Set 3D Neural Scene Segmentation with Pixel-Wise Features and Rendered Novel Views, Francis Engelmann et al. (2024). Lifts dense visual-language semantics into 3D neural radiance fields to support open-vocabulary robotic scene understanding.
- Paper: CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models, Qingqing Zhao et al. (2025). Advances vision-based robotic manipulation by generating intermediate visual subgoals before executing control actions.
