RobustNeRF: Ignoring Distractors with Robust Losses
Sara SabourSuhani VoraDaniel DuckworthIvan KrasinDavid J. FleetAndrea Tagliasacchi
Presents an outlier-aware loss formulation for neural radiance fields that removes transient distractors and shadows from multi-view image captures without relying on semantic segmentation masks.
Creating accurate 3D digital representations from collections of 2D images is vital for applications in augmented and virtual reality, autonomous robotics, and digital mapping. Neural radiance fields (NeRF) have become the standard technique for this task, but they rely heavily on the assumption that scenes remain completely static during image capture. In real-world environments, transient distractors—such as moving people, passing vehicles, and shifting shadows—frequently violate this assumption. Conventional training methods struggle to distinguish between transient objects and normal light reflections, resulting in severe visual artifacts, cloudy geometry, and degraded 3D models.
The article demonstrates that treating transient distractors as statistical outliers during optimization enables clean 3D scene reconstruction without requiring manual data labeling or complex pre-processing. The authors introduce RobustNeRF, a straightforward method that integrates an outlier-filtering mechanism directly into existing neural rendering pipelines.
To evaluate this approach, the authors designed a robust training strategy based on trimmed least-squares and iteratively reweighted optimization. The technique identifies pixels with large prediction errors, applies spatial smoothing across neighboring pixels, and filters patches based on the principle that real-world distractors occupy continuous regions rather than isolated pixels. The evaluation benchmarked RobustNeRF against standard baselines and leading dynamic-scene models across both synthetic datasets and real-world scenes captured manually and autonomously by robotic arms.
The findings show that RobustNeRF consistently outperforms existing approaches in visual fidelity and computational efficiency. On real-world natural scenes, it exceeds standard baselines by 1.3 to 4.7 dB in Peak Signal-to-Noise Ratio (PSNR), successfully removing distractors and eliminating cloudy artifacts. Against specialized dynamic reconstruction models, it achieves up to a 12 dB PSNR improvement on scenes with numerous shifting objects while requiring substantially less memory—using 2.3 times less peak memory overall and 37 times less when normalized for batch size. Furthermore, sensitivity analyses demonstrate that the model maintains stable reconstruction quality above 31 dB even as the proportion of cluttered training images increases, whereas baseline performance steadily drops.
These results imply that high-quality 3D reconstructions can be reliably captured in unconstrained, real-world conditions without costly manual cleanup or strict environmental controls. Because the proposed optimization logic requires few hyperparameters and integrates directly into standard neural rendering frameworks, organizations can deploy it with minimal development overhead and lower compute costs compared to multi-component dynamic models.
Based on these results, technical teams working on 3D reconstruction and mapping should integrate trimmed, patch-based robust loss formulations into their existing pipelines when capturing uncontrolled environments. However, decision-makers should note a key performance trade-off: on datasets that are entirely free of distractors, the trimming mechanism slightly reduces statistical efficiency, leading to marginally lower reconstruction fidelity and longer training times. Further work is recommended to adapt the spatial filtering window for very small distractors and explore learned weight functions for fully automated scene capture.
- Paper: NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis, Ben Mildenhall et al. (2020). Introduces the foundational coordinate-based neural radiance field and volume rendering formulation that RobustNeRF directly builds upon and makes robust to transient distractors.
- Paper: NeRF in the Wild: Neural Radiance Fields for Unconstrained Photo Collections, Ricardo Martin-Brualla et al. (2021). Pioneers the modeling of unconstrained photo collections by decoupling transient components in neural radiance fields, providing the primary baseline and conceptual benchmark for RobustNeRF's outlier-loss approach.
- Paper: D2NeRF: Self-Supervised Decoupling of Dynamic and Static Objects from a Monocular Video, Tianhao Wu et al. (2022). Develops a multi-field strategy for decoupling static backgrounds and dynamic objects from monocular inputs, representing an alternative dynamic-scene paradigm evaluated against RobustNeRF's single-field robust formulation.
- Paper: Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance Fields, Jonathan T. Barron et al. (2022). Establishes techniques for unbounded 360-degree neural reconstruction and proposal sampling, representing the core rendering backbone architecture deployed in modern wild capture settings.
- Paper: Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance Fields, Jonathan T. Barron et al. (2021). Introduces cone-tracing and integrated positional encodings for multiscale anti-aliasing in neural radiance fields, fundamental to modern NeRF optimization pipelines.
- Paper: D-NeRF: neural radiance fields for dynamic scenes, Albert Pumarola et al. (2021). Formulates dynamic scene view synthesis via canonical fields and deformation networks, serving as standard context for dynamic neural scene representations.
- Paper: 3D Gaussian Splatting for Real-Time Radiance Field Rendering, Bernhard Kerbl et al. (2023). Transitions the static radiance field representation from volumetric MLPs to explicit 3D Gaussian primitives for real-time rendering and fast optimization.
- Paper: 4D Gaussian Splatting for Real-Time Dynamic Scene Rendering, Guanjun Wu et al. (2023). Extends explicit Gaussian radiance fields to dynamic environments using deformation networks, building on the broader goal of capturing moving scenes in real time.
- Paper: DL3DV-10K: A Large-Scale Scene Dataset for Deep Learning-based 3D Vision, Lu Ling et al. (2024). Provides a large-scale, diverse real-world benchmark to rigorously evaluate view-synthesis and reconstruction models against complex lighting, clutter, and unconstrained environments.
- Paper: 2D Gaussian Splatting for Geometrically Accurate Radiance Fields, Binbin Huang et al. (2024). Builds on radiance field modeling advances by introducing 2D planar surfels to improve geometric accuracy and surface reconstruction.
- Paper: VastGaussian: Vast 3D Gaussians for Large Scene Reconstruction, Jiaqi Lin et al. (2024). Scales radiance field reconstruction to vast, unconstrained outdoor spaces by combining spatial partitioning with appearance variation modeling.
