RobustNeRF: Ignoring Distractors with Robust Losses

Sara SabourSuhani VoraDaniel DuckworthIvan KrasinDavid J. FleetAndrea Tagliasacchi

article2023CVPR111 citationsCVPR 2023 Highlight

Presents an outlier-aware loss formulation for neural radiance fields that removes transient distractors and shadows from multi-view image captures without relying on semantic segmentation masks.

Listen

Creating accurate 3D digital representations from collections of 2D images is vital for applications in augmented and virtual reality, autonomous robotics, and digital mapping. Neural radiance fields (NeRF) have become the standard technique for this task, but they rely heavily on the assumption that scenes remain completely static during image capture. In real-world environments, transient distractors—such as moving people, passing vehicles, and shifting shadows—frequently violate this assumption. Conventional training methods struggle to distinguish between transient objects and normal light reflections, resulting in severe visual artifacts, cloudy geometry, and degraded 3D models.

The article demonstrates that treating transient distractors as statistical outliers during optimization enables clean 3D scene reconstruction without requiring manual data labeling or complex pre-processing. The authors introduce RobustNeRF, a straightforward method that integrates an outlier-filtering mechanism directly into existing neural rendering pipelines.

To evaluate this approach, the authors designed a robust training strategy based on trimmed least-squares and iteratively reweighted optimization. The technique identifies pixels with large prediction errors, applies spatial smoothing across neighboring pixels, and filters patches based on the principle that real-world distractors occupy continuous regions rather than isolated pixels. The evaluation benchmarked RobustNeRF against standard baselines and leading dynamic-scene models across both synthetic datasets and real-world scenes captured manually and autonomously by robotic arms.

The findings show that RobustNeRF consistently outperforms existing approaches in visual fidelity and computational efficiency. On real-world natural scenes, it exceeds standard baselines by 1.3 to 4.7 dB in Peak Signal-to-Noise Ratio (PSNR), successfully removing distractors and eliminating cloudy artifacts. Against specialized dynamic reconstruction models, it achieves up to a 12 dB PSNR improvement on scenes with numerous shifting objects while requiring substantially less memory—using 2.3 times less peak memory overall and 37 times less when normalized for batch size. Furthermore, sensitivity analyses demonstrate that the model maintains stable reconstruction quality above 31 dB even as the proportion of cluttered training images increases, whereas baseline performance steadily drops.

These results imply that high-quality 3D reconstructions can be reliably captured in unconstrained, real-world conditions without costly manual cleanup or strict environmental controls. Because the proposed optimization logic requires few hyperparameters and integrates directly into standard neural rendering frameworks, organizations can deploy it with minimal development overhead and lower compute costs compared to multi-component dynamic models.

Based on these results, technical teams working on 3D reconstruction and mapping should integrate trimmed, patch-based robust loss formulations into their existing pipelines when capturing uncontrolled environments. However, decision-makers should note a key performance trade-off: on datasets that are entirely free of distractors, the trimming mechanism slightly reduces statistical efficiency, leading to marginally lower reconstruction fidelity and longer training times. Further work is recommended to adapt the spatial filtering window for very small distractors and explore learned weight functions for fully automated scene capture.

arXiv: 2302.00833
Cover for RobustNeRF: Ignoring Distractors with Robust Losses

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Sensitivity to outliers
  • 3.2. Robustness to outliers
  • 3.3. Robustness via Trimmed Least Squares
  • 4. Experiments
  • 4.1. Baselines
  • 4.2. Datasets
  • 4.3. Evaluation
  • 5. Conclusions
  • References

Knowls

  1. Knowl 1 — RobustNeRF Trimmed Loss Weighting Mechanism

    model/method

    RobustNeRF computes binary inlier weights ω(r)∈{0,1}\omega(\mathbf{r}) \in \{0, 1\} for each pixel ray r\mathbf{r} at training iteration tt using a three-stage spatial regularization pipeline based on residuals ϵ(r)=∥C(r;θ(t−1))−Ci(r)∥2\epsilon(\mathbf{r}) = \|\mathbf{C}(\mathbf{r}; \theta^{(t-1)}) - \mathbf{C}_i(\mathbf{r})\|_2, where C(r;θ(t−1))\mathbf{C}(\mathbf{r}; \theta^{(t-1)}) is the NeRF color prediction from the previous iteration and Ci(r)\mathbf{C}_i(\mathbf{r}) is the ground-truth observed RGB color:

    1. Median Trimming: Compute initial inlier classifications ω~(r)\tilde{\omega}(\mathbf{r}) by selecting residuals below a percentile threshold (set to the 50th percentile / median): ω~(r)=I(ϵ(r)≤Tϵ),Tϵ=Medianr{ϵ(r)}\tilde{\omega}(\mathbf{r}) = \mathbb{I}(\epsilon(\mathbf{r}) \le T_\epsilon), \quad T_\epsilon = \text{Median}_{\mathbf{r}}\{\epsilon(\mathbf{r})\} where I(⋅)\mathbb{I}(\cdot) is the indicator function.

    2. Spatial Diffusion: Convolve the initial inlier map with a 3×33 \times 3 box filter B3×3\mathcal{B}_{3 \times 3} and apply a threshold to enforce spatial continuity of inliers: W(r)=I((ω~(r)⊛B3×3)≥T⊛),T⊛=0.5\mathcal{W}(\mathbf{r}) = \mathbb{I}((\tilde{\omega}(\mathbf{r}) \circledast \mathcal{B}_{3 \times 3}) \ge T_\circledast), \quad T_\circledast = 0.5

    3. Patch-Level Aggregation: Label an entire 8×88 \times 8 pixel patch R8(r)\mathcal{R}_8(\mathbf{r}) centered around ray r\mathbf{r} based on the mean spatial inlier rate over a larger 16×1616 \times 16 neighborhood R16(r)\mathcal{R}_{16}(\mathbf{r}): ω(R8(r))=I(Es∼R16(r)[W(s)]≥TR),TR=0.6\omega(\mathcal{R}_8(\mathbf{r})) = \mathbb{I}\left(\mathbb{E}_{\mathbf{s} \sim \mathcal{R}_{16}(\mathbf{r})}[\mathcal{W}(\mathbf{s})] \ge T_\mathcal{R}\right), \quad T_\mathcal{R} = 0.6

    The final weight mask is the union of the inlier masks across these filtering stages. This procedure suppresses transient distractors while preserving high-frequency static textures that initially present large residuals.

  2. Knowl 2 — Iteratively Reweighted Least Squares Objective for Distractor-Robust NeRF

    equation

    In RobustNeRF, training with distractors is formulated as an Iteratively Reweighted Least Squares (IRLS) optimization problem. At training iteration tt, the robust photometric loss for a camera ray r\mathbf{r} sampled from training image Ci\mathbf{C}_i is given by:

    Lrobustr,i(θ(t))=ω(ϵ(t−1)(r))⋅∥C(r;θ(t))−Ci(r)∥22\mathcal{L}_{\text{robust}}^{\mathbf{r}, i}(\theta^{(t)}) = \omega(\epsilon^{(t-1)}(\mathbf{r})) \cdot \|\mathbf{C}(\mathbf{r}; \theta^{(t)}) - \mathbf{C}_i(\mathbf{r})\|_2^2

    where C(r;θ(t))\mathbf{C}(\mathbf{r}; \theta^{(t)}) is the RGB color predicted along ray r\mathbf{r} by the neural radiance field parameterized by θ(t)\theta^{(t)}, Ci(r)\mathbf{C}_i(\mathbf{r}) is the ground-truth observed RGB color in image ii, and ϵ(t−1)(r)=∥C(r;θ(t−1))−Ci(r)∥2\epsilon^{(t-1)}(\mathbf{r}) = \|\mathbf{C}(\mathbf{r}; \theta^{(t-1)}) - \mathbf{C}_i(\mathbf{r})\|_2 is the error residual computed with network parameters frozen at iteration t−1t-1. The weighting function ω(⋅)∈{0,1}\omega(\cdot) \in \{0, 1\} dynamically identifies outlier/distractor rays and excludes them from parameter updates.

  3. Knowl 3 — Gradient Pathology of Standard Robust Estimators in NeRF Optimization

    theoretical result

    When optimizing a neural radiance field under a robust M-estimator loss κ(ϵ(θ))\kappa(\epsilon(\theta)) with residual ϵ(θ)=∥C(r;θ)−Ci(r)∥2\epsilon(\theta) = \|\mathbf{C}(\mathbf{r}; \theta) - \mathbf{C}_i(\mathbf{r})\|_2, the gradient with respect to network parameters θ\theta at iteration tt follows the chain rule:

    ∂κ(ϵ(θ))∂θ∣θ(t)=∂κ(ϵ)∂ϵ∣ϵ(θ(t))⋅∂ϵ(θ)∂θ∣θ(t)\left. \frac{\partial \kappa(\epsilon(\theta))}{\partial \theta} \right|_{\theta^{(t)}} = \left. \frac{\partial \kappa(\epsilon)}{\partial \epsilon} \right|_{\epsilon(\theta^{(t)})} \cdot \left. \frac{\partial \epsilon(\theta)}{\partial \theta} \right|_{\theta^{(t)}}

    Early and midway through NeRF training, unresolved high-frequency static details produce error residuals ϵ(θ(t))\epsilon(\theta^{(t)}) of comparable magnitude to those caused by transient scene distractors (outliers). Under strongly robust redescending kernels (such as Geman-McClure or Lorentzian), the derivative ∂κ(ϵ)∂ϵ→0\frac{\partial \kappa(\epsilon)}{\partial \epsilon} \to 0 for large residuals. Consequently, the gradient signal for high-frequency static details is severely downweighted alongside outliers, causing fine-grained textures to be learned extremely slowly or erased altogether.

  4. Knowl 4 — Quantitative Distractor Removal Benchmark on Real-World Scenes

    data/table

    Reconstruction quality on novel views of four real-world natural scenes containing distractors, evaluated against held-out distractor-free ground truth frames. Metrics reported are LPIPS (lower is better), SSIM (higher is better), and PSNR in dB (higher is better):

    Method Statue Android Crab BabyYoda
    LPIPS↓\downarrow SSIM↑\uparrow PSNR↑\uparrow LPIPS↓\downarrow SSIM↑\uparrow PSNR↑\uparrow LPIPS↓\downarrow SSIM↑\uparrow PSNR↑\uparrow LPIPS↓\downarrow SSIM↑\uparrow PSNR↑\uparrow
    mip-NeRF 360 (L2L_2) 0.36 0.66 19.09 0.40 0.65 19.35 0.27 0.77 25.73 0.31 0.75 22.97
    mip-NeRF 360 (L1L_1) 0.30 0.72 19.55 0.40 0.66 19.38 0.22 0.79 26.69 0.22 0.80 26.15
    mip-NeRF 360 (Charbonnier) 0.30 0.73 19.64 0.40 0.66 19.53 0.21 0.80 27.72 0.23 0.80 25.22
    D2NeRF\text{D}^2\text{NeRF} 0.48 0.49 19.09 0.43 0.57 20.61 0.42 0.68 21.18 0.44 0.65 17.32
    RobustNeRF 0.28 0.75 20.89 0.31 0.65 21.72 0.21 0.81 30.75 0.20 0.83 30.87
    mip-NeRF 360 (clean upper bound) 0.19 0.80 23.57 0.31 0.71 23.10 0.16 0.84 32.55 0.16 0.84 32.63

    RobustNeRF outperforms all baselines across all scenes, outperforming the standard mip-NeRF 360 baselines by 1.3 to 4.7 dB PSNR and D2NeRF\text{D}^2\text{NeRF} by up to 13.5 dB PSNR in the presence of extensive distractors (e.g., BabyYoda). Standard loss functions force mip-NeRF 360 to reconstruct distractors as semi-transparent cloudy floaters, while D2NeRF\text{D}^2\text{NeRF} struggles when the number of moving distractor objects is large (100 to 150 objects).

  5. Knowl 5 — Quantitative Evaluation on Synthetic Kubric Distractor Benchmark

    data/table

    Evaluation of novel view synthesis on the synthetic Kubric benchmark datasets introduced by D2NeRF\text{D}^2\text{NeRF}, containing 200 training frames with moving distractors and 100 distractor-free evaluation views. Metrics reported are LPIPS (lower is better), MS-SSIM (higher is better), and PSNR in dB (higher is better):

    Method Car Cars Bag Chairs Pillow
    LPIPS↓\downarrow MS-SSIM↑\uparrow PSNR↑\uparrow LPIPS↓\downarrow MS-SSIM↑\uparrow PSNR↑\uparrow LPIPS↓\downarrow MS-SSIM↑\uparrow PSNR↑\uparrow LPIPS↓\downarrow MS-SSIM↑\uparrow PSNR↑\uparrow LPIPS↓\downarrow MS-SSIM↑\uparrow PSNR↑\uparrow
    NeRF-W .218 .814 24.23 .243 .873 24.51 .139 .791 20.65 .150 .681 23.77 .088 .935 28.24
    NSFF .200 .806 24.90 .620 .376 10.29 .108 .892 25.62 .682 .284 12.82 .782 .343 4.55
    NeuralDiff .065 .952 31.89 .098 .921 25.93 .117 .910 29.02 .112 .722 24.42 .565 .652 20.09
    D2NeRF\text{D}^2\text{NeRF} .062 .975 34.27 .090 .953 26.27 .076 .979 34.14 .095 .707 24.63 .076 .979 36.58
    RobustNeRF .013 .988 37.73 .063 .957 26.31 .006 .995 41.82 .007 .992 41.23 .018 .990 38.95

    RobustNeRF achieves state-of-the-art results across all five Kubric scenes, outperforming dynamic scene decomposition methods (NeuralDiff, D2NeRF\text{D}^2\text{NeRF}, NSFF) and transient-embedding models (NeRF-W), achieving PSNR gains exceeding 7 dB on the Bag and Chairs scenes over D2NeRF\text{D}^2\text{NeRF}.

  6. Knowl 6 — Ablation Analysis of Loss Regularization Components

    data/table

    Ablation study evaluating the individual components of the RobustNeRF loss on the Crab scene against distractor-free ground truth views. Metrics reported are LPIPS, SSIM, PSNR (dB), and optimization updates required to attain PSNR ≥30\ge 30 dB:

    Loss Configuration LPIPS↓\downarrow SSIM↑\uparrow PSNR↑\uparrow Updates to PSNR=30
    mip-NeRF 360 (L2L_2) 0.31 0.75 22.97 –
    + robust trimming (Eq. 8) 0.39 0.60 18.21 –
    + spatial smoothing (Eq. 9) 0.22 0.81 30.01 250K
    + patch-level aggregation (Eq. 10) 0.21 0.81 30.75 70K
    Oracle (trained on clean data) 0.16 0.84 32.55 25K

    Blindly applying trimmed least squares (+ robust) eliminates distractors but prematurely removes fine high-frequency static textures, dropping PSNR to 18.21 dB. Incorporating 3×33 \times 3 spatial smoothing (+ smoothing) recovers fine texture detail, raising PSNR to 30.01 dB, but requires 250K iterations. Adding 8×88 \times 8 patch-level aggregation within 16×1616 \times 16 neighborhoods (+ patching) accelerates training convergence by ≈3.6×\approx 3.6\times (reaching 30 dB in 70K updates vs. 250K updates).

  7. Knowl 7 — Robustness to Clutter Density in Training Data

    empirical result

    When evaluating reconstruction quality on the BabyYoda dataset across varying proportions of cluttered training images (ranging from 0% to 100% of images containing distractors while keeping the total dataset size constant), standard mip-NeRF 360 with L2L_2 loss exhibits a monotonic decrease in test PSNR from 33 dB down to 25 dB as distractor prevalence increases. In contrast, RobustNeRF maintains stable reconstruction accuracy, consistently staying above 31 dB PSNR across all clutter fractions from 10% to 100%.

  8. Knowl 8 — Statistical Inefficiency on Clean Distractor-Free Captures

    limitation

    Because RobustNeRF employs a trimmed least-squares estimator that unconditionally discards a percentile of pixels with the largest residuals during initial training phases, it incurs statistical inefficiency on purely static, distractor-free capture datasets. When trained on perfectly clean scenes, RobustNeRF achieves slightly lower reconstruction quality than standard mip-NeRF 360 (e.g., 30.87 dB vs. 32.63 dB on distractor-free BabyYoda) and requires more optimization steps to converge to equivalent accuracy levels.

  9. Knowl 9 — Multi-View Benchmark for Distractor Robustness with Robotic Arm Ground Truth

    experimental setup

    The evaluation dataset comprises seven natural multi-view scenes spanning three capture environments (apartment, street, and robotics lab) and synthetic scenes generated via Kubric. In contrast to monocular dynamic video datasets that enforce temporal continuity, training images are captured as unordered photo sets with independent distractor movements (varying from 1 distractor in Statue up to 150 unique moving objects in BabyYoda). For robotics lab scenes (Crab and BabyYoda), exact pixel-aligned pairs of distractor-laden and distractor-free views are captured via an automated robotic arm, enabling ground-truth evaluation without distractor contamination.

Coverage note — None was omitted; all key contributions—including the robust IRLS formulation, spatial/patch trimmed kernel design, gradient failure analysis of classical robust losses, synthetic and real-world benchmarks, ablation experiments, and limitations—have been covered.

References

  1. 1.Michal Adamkiewicz, Timothy Chen, Adam Caccavale, Rachel Gardner, Preston Culbertson, Jeannette Bohg, and Mac Schwager. Vision-only robot navigation in a neural radiance world. IEEE Robotics and Automation Letters, 2022. 1, 2
  2. 2.Jonathan T. Barron. A general and adaptive robust loss function. Proc. CVPR, 2019. 3, 4
  3. 3.Jonathan T. Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-Brualla, and Pratul P. Srinivasan. Mip-nerf: A multiscale representation for anti-aliasing neural radiance fields. ICCV, 2021. 3
  4. 4.Jonathan T Barron, Ben Mildenhall, Dor Verbin, Pratul P Srinivasan, and Peter Hedman. Mip-nerf 360: Unbounded anti-aliased neural radiance fields. In Proc. CVPR, 2022. 2, 5, 8
  5. 5.Mark Boss, Andreas Engelhardt, Abhishek Kar, Yuanzhen Li, Deqing Sun, Jonathan T. Barron, Hendrik P.A. Lensch, and Varun Jampani. SAMURAI: Shape And Material from Unconstrained Real-world Arbitrary Image collections. In Proc. NeurIPS, 2022. 2
  6. 6.Zhiqin Chen, Thomas Funkhouser, Peter Hedman, and Andrea Tagliasacchi. Mobilenerf: Exploiting the polygon rasterization pipeline for efficient neural field rendering on mobile architectures. arXiv preprint arXiv:2208.00277, 2022. 1, 2
  7. 7.Dmitry Chetverikov, Dmitry Svirko, Dmitry Stepanov, and Pavel Krsek. The trimmed iterative closest point algorithm. In International Conference on Pattern Recognition, 2002. 5
  8. 8.Shin-Fang Chng, Sameera Ramasinghe, Jamie Sherrah, and Simon Lucey. Garf: Gaussian activated radiance fields for high fidelity reconstruction and pose estimation. arXiv eprints, 2022. 2
  9. 9.Yilun Du, Yinan Zhang, Hong-Xing Yu, Joshua B Tenenbaum, and Jiajun Wu. Neural radiance flow for 4d view synthesis and video processing. In Proc. ICCV. IEEE Computer Society, 2021. 2
  10. 10.Jiemin Fang, Taoran Yi, Xinggang Wang, Lingxi Xie, Xiaopeng Zhang, Wenyu Liu, Matthias Nießner, and Qi Tian. Fast dynamic radiance fields with time-aware neural voxels. arXiv preprint arXiv:2205.15285, 2022. 2
  11. 11.Chen Gao, Ayush Saraf, Johannes Kopf, and Jia-Bin Huang. Dynamic view synthesis from dynamic monocular video. In Proc. ICCV, 2021. 2
  12. 12.Klaus Greff, Francois Belletti, Lucas Beyer, Carl Doersch, Yilun Du, Daniel Duckworth, David J Fleet, Dan Gnanapragasam, Florian Golemo, Charles Herrmann, Thomas Kipf, Abhijit Kundu, Dmitry Lagun, Issam Laradji, Hsueh-Ti (Derek) Liu, Henning Meyer, Yishu Miao, Derek Nowrouzezahrai, Cengiz Oztireli, Etienne Pot, Noha Radwan, Daniel Rebain, Sara Sabour, Mehdi S. M. Sajjadi, Matan Sela, Vincent Sitzmann, Austin Stone, Deqing Sun, Suhani Vora, Ziyu Wang, Tianhao Wu, Kwang Moo Yi, Fangcheng Zhong, and Andrea Tagliasacchi. Kubric: a scalable dataset generator. In Proc. CVPR, 2022. 6
  13. 13.Peter Hedman, Pratul P. Srinivasan, Ben Mildenhall, Jonathan T. Barron, and Paul Debevec. Baking neural radiance fields for real-time view synthesis. Proc. ICCV, 2021. 2
  14. 14.Yoonwoo Jeong, Seokjun Ahn, Christopher Choy, Anima Anandkumar, Minsu Cho, and Jaesik Park. Self-calibrating neural radiance fields. In Proc. ICCV, 2021. 2
  15. 15.Animesh Karnewar, Tobias Ritschel, Oliver Wang, and Niloy Mitra. Relu fields: The little non-linearity that could. TOG (Proc. SIGGRAPH), 2022. 2
  16. 16.KN Kutulakos and SM Seitz. A theory of shape by space carving. IJCV, 2000. 3
  17. 17.Tianye Li, Mira Slavcheva, Michael Zollhoefer, Simon Green, Christoph Lassner, Changil Kim, Tanner Schmidt, Steven Lovegrove, Michael Goesele, Richard Newcombe, et al. Neural 3d video synthesis from multi-view video. In Proc. CVPR, 2022. 2
  18. 18.Zhengqi Li, Simon Niklaus, Noah Snavely, and Oliver Wang. Neural scene flow fields for space-time view synthesis of dynamic scenes. In Proc. CVPR, 2021. 2, 8
  19. 19.Chen-Hsuan Lin, Wei-Chiu Ma, Antonio Torralba, and Simon Lucey. Barf: Bundle-adjusting neural radiance fields. In Proc. CVPR, 2021. 2
  20. 20.Jia-Wei Liu, Yan-Pei Cao, Weijia Mao, Wenqiao Zhang, David Junhao Zhang, Jussi Keppo, Ying Shan, Xiaohu Qie, and Mike Zheng Shou. Devrf: Fast deformable voxel radiance fields for dynamic scenes. arXiv preprint arXiv:2205.15723, 2022. 2
  21. 21.Stephen Lombardi, Tomas Simon, Jason Saragih, Gabriel Schwartz, Andreas Lehrmann, and Yaser Sheikh. Neural volumes: Learning dynamic renderable volumes from images. ACM TOG, 2019. 1
  22. 22.Li Ma, Xiaoyu Li, Jing Liao, Qi Zhang, Xuan Wang, Jue Wang, and Pedro V. Sander. Deblur-nerf: Neural radiance fields from blurry images. In Proc. CVPR, 2022. 2
  23. 23.Ricardo Martin-Brualla, Noha Radwan, Mehdi SM Sajjadi, Jonathan T Barron, Alexey Dosovitskiy, and Daniel Duckworth. Nerf in the wild: Neural radiance fields for unconstrained photo collections. In Proc. CVPR, 2021. 2, 3, 4, 8
  24. 24.Ben Mildenhall, Peter Hedman, Ricardo Martin-Brualla, Pratul P. Srinivasan, and Jonathan T. Barron. NeRF in the dark: High dynamic range view synthesis from noisy raw images. In Proc. CVPR, 2021. 2
  25. 25.Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. In Proc. ECCV, 2020. 1, 2
  26. 26.Ben Mildenhall, Dor Verbin, Pratul P. Srinivasan, Peter Hedman, Ricardo Martin-Brualla, and Jonathan T. Barron. Multinerf: a code release for Mip-NeRF 360, Ref-NeRF, and RawNeRF, 2022. 5
  27. 27.Thomas Müller, Alex Evans, Christoph Schied, and Alexander Keller. Instant neural graphics primitives with a multiresolution hash encoding. TOG (Proc. SIGGRAPH), 2022. 2
  28. 28.Michael Niemeyer, Jonathan T. Barron, Ben Mildenhall, Mehdi S. M. Sajjadi, Andreas Geiger, and Noha Radwan. Regnerf: Regularizing neural radiance fields for view synthesis from sparse inputs. In Proc. CVPR, 2022. 2
  29. 29.Michael Oechsle, Songyou Peng, and Andreas Geiger. Unisurf: Unifying neural implicit surfaces and radiance fields for multi-view reconstruction. In Proc. ICCV, 2021. 2
  30. 30.Keunhong Park, Utkarsh Sinha, Peter Hedman, Jonathan T Barron, Sofien Bouaziz, Dan B Goldman, Ricardo Martin-Brualla, and Steven M Seitz. Hypernerf: a higher-dimensional representation for topologically varying neural radiance fields. ACM Transactions on Graphics (TOG), 2021. 2
  31. 31.Ben Poole, Ajay Jain, Jonathan T. Barron, and Ben Mildenhall. Dreamfusion: Text-to-3d using 2d diffusion. arXiv preprint 2209.14988, 2022. 2
  32. 32.Albert Pumarola, Enric Corona, Gerard Pons-Moll, and Francesc Moreno-Noguer. D-nerf: Neural radiance fields for dynamic scenes. In Proc. CVPR, 2021. 2
  33. 33.Daniel Rebain, Mark Matthews, Kwang Moo Yi, Dmitry Lagun, and Andrea Tagliasacchi. LOLNeRF: Learn from One Look. In Proc. CVPR, 2022. 2, 3
  34. 34.Jeremy Reizenstein, Roman Shapovalov, Philipp Henzler, Luca Sbordone, Patrick Labatut, and David Novotny. Common objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction. In Proc. ICCV, 2021. 1
  35. 35.Konstantinos Rematas, Andrew Liu, Pratul P. Srinivasan, Jonathan T. Barron, Andrea Tagliasacchi, Tom Funkhouser, and Vittorio Ferrari. Urban radiance fields. Proc. CVPR, 2022. 1, 2, 3
  36. 36.Antoni Rosinol, John J Leonard, and Luca Carlone. Nerf-slam: Real-time dense monocular slam with neural radiance fields. arXiv preprint arXiv:2210.13641, 2022. 1
  37. 37.Johannes L Schonberger and Jan-Michael Frahm. Structure-from-motion revisited. In Proc. CVPR, 2016. 1
  38. 38.Johannes Lutz Schönberger, Enliang Zheng, Marc Pollefeys, and Jan-Michael Frahm. Pixelwise view selection for unstructured multi-view stereo. In Proc. ECCV, 2016. 6
  39. 39.Vincent Sitzmann, Michael Zollhöfer, and Gordon Wetzstein. Scene representation networks: Continuous 3d-structure-aware neural scene representations. In Proc. NeurIPs, 2019. 1
  40. 40.Weiwei Sun, Wei Jiang, Andrea Tagliasacchi, Eduard Trulls, and Kwang Moo Yi. ACNe: Attentive Context Normalization for Robust Permutation-Equivariant Learning. Proc. CVPR, 2020. 4
  41. 41.Andrea Tagliasacchi and Hao Li. Modern techniques and applications for real-time non-rigid registration. In Proc. SIGGRAPH Asia (Technical Course Notes), 2016. 3, 4
  42. 42.Andrea Tagliasacchi and Ben Mildenhall. Volume Rendering Digest (for NeRF), 2022. 2
  43. 43.Matthew Tancik, Vincent Casser, Xinchen Yan, Sabeek Pradhan, Ben Mildenhall, Pratul P Srinivasan, Jonathan T Barron, and Henrik Kretzschmar. Block-nerf: Scalable large scene neural view synthesis. In Proc. CVPR, 2022. 2
  44. 44.Ayush Tewari, Justus Thies, Ben Mildenhall, Pratul Srinivasan, Edgar Tretschk, W Yifan, Christoph Lassner, Vincent Sitzmann, Ricardo Martin-Brualla, Stephen Lombardi, et al. Advances in neural rendering. In Computer Graphics Forum, 2022. 1
  45. 45.Edgar Tretschk, Ayush Tewari, Vladislav Golyanik, Michael Zollhöfer, Christoph Lassner, and Christian Theobalt. Non-rigid neural radiance fields: Reconstruction and novel view synthesis of a dynamic scene from monocular video. In Proc. ICCV, 2021. 2
  46. 46.Vadim Tschernezki, Diane Larlus, and Andrea Vedaldi. NeuralDiff: Segmenting 3D objects that move in egocentric videos. In Proc. 3DV, 2021. 8
  47. 47.Dor Verbin, Peter Hedman, Ben Mildenhall, Todd Zickler, Jonathan T. Barron, and Pratul P. Srinivasan. Ref-NeRF: Structured view-dependent appearance for neural radiance fields. Proc. CVPR, 2022. 2
  48. 48.Suhani Vora, Noha Radwan, Klaus Greff, Henning Meyer, Kyle Genova, Mehdi SM Sajjadi, Etienne Pot, Andrea Tagliasacchi, and Daniel Duckworth. Nesf: Neural semantic fields for generalizable semantic segmentation of 3d scenes. TMLR, 2021. 2
  49. 49.Chaoyang Wang, Ben Eckart, Simon Lucey, and Orazio Gallo. Neural trajectory fields for dynamic novel view synthesis. arXiv preprint arXiv:2105.05994, 2021. 2
  50. 50.Jiepeng Wang, Peng Wang, Xiaoxiao Long, Christian Theobalt, Taku Komura, Lingjie Liu, and Wenping Wang. Neuris: Neural reconstruction of indoor scenes using normal priors. arXiv preprint, 2022. 2
  51. 51.Peng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt, Taku Komura, and Wenping Wang. Neus: Learning neural implicit surfaces by volume rendering for multi-view reconstruction. Proc. NeurIPS, 2021. 2
  52. 52.Z. Wang, A.C. Bovik, H.R. Sheikh, and E.P. Simoncelli. Image quality assessment: From error visibility to structural similarity. IEEE Transactions on Image Processing, 2004. 5
  53. 53.Tianhao Wu, Fangcheng Zhong, Forrester Cole, Andrea Tagliasacchi, and Cengiz Oztireli. D2nerf: Self-supervised decoupling of dynamic and static objects from a monocular video. In Proc. NeurIPS, 2022. 2, 4, 6, 7, 8
  54. 54.Wenqi Xian, Jia-Bin Huang, Johannes Kopf, and Changil Kim. Space-time neural irradiance fields for free-viewpoint video. In Proc. CVPR, 2021. 2
  55. 55.Yiheng Xie, Towaki Takikawa, Shunsuke Saito, Or Litany, Shiqin Yan, Numair Khan, Federico Tombari, James Tompkin, Vincent Sitzmann, and Srinath Sridhar. Neural fields in visual computing and beyond. Comput. Graph. Forum, 2022. 1
  56. 56.Alex Yu, Vickie Ye, Matthew Tancik, and Angjoo Kanazawa. pixelnerf: Neural radiance fields from one or few images. In Proc. CVPR, 2021. 2
  57. 57.Zehao Yu, Songyou Peng, Michael Niemeyer, Torsten Sattler, and Andreas Geiger. Monosdf: Exploring monocular geometric cues for neural implicit surface reconstruction. In Proc. NeurIPS, 2022. 2
  58. 58.Jason Y Zhang, Deva Ramanan, and Shubham Tulsiani. Relpose: Predicting probabilistic relative rotation for single objects in the wild. In Proc. ECCV, 2022. 2
  59. 59.Kai Zhang, Gernot Riegler, Noah Snavely, and Vladlen Koltun. Nerf++: Analyzing and improving neural radiance fields. arXiv preprint arXiv:2010.07492, 2020. 2
  60. 60.Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proc. CVPR, 2018. 5
  61. 61.Zihan Zhu, Songyou Peng, Viktor Larsson, Weiwei Xu, Hujun Bao, Zhaopeng Cui, Martin R Oswald, and Marc Pollefeys. Nice-slam: Neural implicit scalable encoding for slam. In Proc. CVPR, 2022. 1

Citation

MLA
Sabour, S., et al. “RobustNeRF: Ignoring Distractors with Robust Losses”. arXiv, 2023, http://arxiv.org/abs/2302.00833v2.
APA
Sabour, S., Vora, S., Duckworth, D., Krasin, I., Fleet, D. J., & Tagliasacchi, A. (2023). RobustNeRF: Ignoring Distractors with Robust Losses. arXiv. http://arxiv.org/abs/2302.00833v2
Chicago
Sabour, S., S. Vora, D. Duckworth, I. Krasin, D. J. Fleet, and A. Tagliasacchi. 2023. “RobustNeRF: Ignoring Distractors with Robust Losses”. arXiv. http://arxiv.org/abs/2302.00833v2.
Harvard
Sabour, S. et al. (2023) “RobustNeRF: Ignoring Distractors with Robust Losses”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2302.00833v2.
Vancouver
1. Sabour S, Vora S, Duckworth D, Krasin I, Fleet DJ, Tagliasacchi A (2023) RobustNeRF: Ignoring Distractors with Robust Losses. arXiv

BibTeX

@article{sabour2023robustnerf,
  title = {RobustNeRF: Ignoring Distractors with Robust Losses},
  author = {Sabour, Sara and Vora, Suhani and Duckworth, Daniel and Krasin, Ivan and Fleet, David J. and Tagliasacchi, Andrea},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2302.00833v2},
  eprint = {2302.00833}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE