Panoptic Lifting for 3D Scene Understanding with Neural Fields

Yawar SiddiquiLorenzo PorziSamuel Rota BulòNorman MüllerMatthias NießnerAngela DaiPeter Kontschieder

article2023CVPR211 citations

Presents a neural field framework that reconstructs unified 3D panoptic scene representations directly from multi-view images and noisy 2D segmentations without requiring ground-truth 3D annotations or 3D bounding box detectors.

Listen

Building accurate three-dimensional models of physical environments that capture geometry, visual appearance, and object-level semantic identities is essential for emerging technologies such as virtual reality, autonomous vehicles, and robotic navigation. While standard two-dimensional computer vision tools can segment individual photographs into object categories and instances, they struggle to maintain consistency across multiple viewpoints. Two-dimensional models frequently produce conflicting category labels and lack the capacity to track persistent object identities from frame to frame, creating significant hurdles for downstream applications that require a coherent, full-scene understanding.

The article demonstrates a novel framework called Panoptic Lifting, which evaluates how noisy, machine-generated two-dimensional image segmentations can be lifted into a unified, view-consistent three-dimensional volumetric representation without requiring manual three-dimensional annotations. To achieve this, the authors combine an explicit volumetric neural field for appearance and density with lightweight neural networks for semantic classes and object instances. The model aligns inconsistent two-dimensional instance identifiers to persistent three-dimensional identities using an optimal linear assignment process. It incorporates robustness techniques including test-time data augmentations to refine confidence estimates, a segment consistency objective to prevent object fragmentation, bounded probability fields, and gradient blocking to keep label noise from corrupting scene geometry. The authors evaluated the approach across standard synthetic and real-world benchmark datasets—Hypersim, Replica, and ScanNet—as well as in-the-wild smartphone captures.

The analysis demonstrates substantial performance improvements over existing state-of-the-art methods. In scene-level panoptic quality, Panoptic Lifting outperformed competing neural field baselines by 8.4 percentage points on Hypersim, 13.8 percentage points on Replica, and 10.6 percentage points on ScanNet. It also improved semantic segmentation accuracy by approximately 6 to 18 percentage points over standard 2D and 3D baselines. Ablation experiments confirmed that each robustness mechanism was vital, as removing all proposed additions decreased segmentation quality by 8 percentage points and scene-level panoptic quality by 11 percentage points. Furthermore, the resulting representations successfully enabled downstream interactive applications, such as selective object deletion, duplication, and spatial transformation.

These findings indicate that high-quality three-dimensional panoptic understanding can be achieved directly from standard photographs and off-the-shelf two-dimensional recognition tools, completely bypassing the need for labor-intensive, expensive three-dimensional manual labeling or fragile three-dimensional bounding box detectors. For operational workflows, this significantly reduces data collection costs and mitigates the risk of compounding errors from complex multi-model pipelines. Organizations developing spatial computing, mapping, or simulation platforms can deploy this approach to rapidly reconstruct interactive digital twins from commodity cameras.

Technical leaders should consider piloting this lifting architecture within their existing spatial reconstruction workflows, particularly where photorealistic rendering and object manipulation are needed simultaneously. However, decision-makers should note that the framework currently assumes static scenes and requires training individual neural fields per environment, which takes approximately ten hours per scene on high-end hardware. Further work is required to extend the method to dynamic environments with moving objects and to reduce per-scene training overhead before deploying the pipeline into real-time operational systems.

Cover for Panoptic Lifting for 3D Scene Understanding with Neural Fields

Abstract

We propose Panoptic Lifting, a novel approach for learning panoptic 3D volumetric representations from images of in-the-wild scenes. Once trained, our model can render color images together with 3D-consistent panoptic segmentation from novel viewpoints. Unlike existing approaches which use 3D input directly or indirectly, our method requires only machine-generated 2D panoptic segmentation masks inferred from a pre-trained network. Our core contribution is a panoptic lifting scheme based on a neural field representation that generates a unified and multi-view consistent, 3D panoptic representation of the scene. To account for inconsistencies of 2D instance identifiers across views, we solve a linear assignment with a cost based on the model’s current predictions and the machine-generated segmentation masks, thus enabling us to lift 2D instances to 3D in a consistent way. We further propose and ablate contributions that make our method more robust to noisy, machine-generated labels, including test-time augmentations for confidence estimates, segment consistency loss, bounded segmentation fields, and gradient stopping. Experimental results validate our approach on the challenging Hypersim, Replica, and ScanNet datasets, improving by 8.4, 13.8, and 10.6% in scene-level PQ over state of the art.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Input Data
  • 3.2. Scene Representation and Rendering
  • 3.3. Loss Functions
  • 4. Experiments
  • 4.1. Results
  • 4.1.1 Ablations
  • 4.1.2 Scene Editing
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Panoptic Radiance Field Scene Representation and Rendering

    model/method

    A static 3D scene is modeled as an implicit volumetric function Φ(x,d)\Phi(x, d), mapping any 3D coordinate x∈R3x \in \mathbb{R}^3 and viewing direction d∈S2d \in \mathbb{S}^2 to volume density σ∈R≥0\sigma \in \mathbb{R}_{\ge 0}, an RGB color c∈R3c \in \mathbb{R}^3, a semantic class probability distribution κ\kappa over a set of categories K\mathcal{K}, and a surrogate instance identifier probability distribution π\pi over a fixed set of surrogate identifiers J\mathcal{J}.

    The radiance and density components (σ,c)(\sigma, c) are represented using a TensoRF tensor decomposition architecture, while the semantic field κ(x)\kappa(x) and surrogate instance field π(x)\pi(x) are parameterized by two small multi-layer perceptrons (MLPs).

    A continuous vector field f(x,d)f(x, d) (which can be color cc, semantic distribution κ\kappa, or surrogate instance distribution π\pi) is rendered along a ray r(t)=p0+tdrr(t) = p_0 + t d_r according to the volume rendering equation:

    R[f∣r,σ]:=∫0∞αt(r)σ(r(t))f(r(t),dr) dt\mathcal{R}[f \mid r, \sigma] := \int_{0}^{\infty} \alpha_t(r) \sigma(r(t)) f(r(t), d_r) \, dt

    where αt(r)\alpha_t(r) is the ray transmittance probability at depth tt:

    αt(r):=exp⁡(−∫0tσ(r(s)) ds)\alpha_t(r) := \exp\left(-\int_{0}^{t} \sigma(r(s)) \, ds\right)

    For ray rr, the rendered vector is denoted fr:=R[f∣r,σ]f_r := \mathcal{R}[f \mid r, \sigma]. At inference time, the rendered semantic class is selected as kr⋆:=arg⁡max⁡k∈Kκr(k)k_r^\star := \arg\max_{k \in \mathcal{K}} \kappa_r(k). If kr⋆k_r^\star is a "thing" class (kr⋆∈KT⊂Kk_r^\star \in \mathcal{K}_T \subset \mathcal{K}), the most probable surrogate identifier is selected as jr⋆:=arg⁡max⁡j∈Jπr(j)j_r^\star := \arg\max_{j \in \mathcal{J}} \pi_r(j), producing a multi-view consistent 3D instance identifier (kr⋆,jr⋆)∈H3D:=KT×J(k_r^\star, j_r^\star) \in \mathcal{H}_{3\text{D}} := \mathcal{K}_T \times \mathcal{J}.

  2. Knowl 2 — 2D-to-3D Instance Lifting via Linear Assignment

    model/method

    Because 2D instance segmentation predictions from off-the-shelf single-image models do not track consistent object identities across views, 2D instances are mapped to persistent 3D surrogate identifiers using linear assignment.

    Let RR denote a batch of rays sampled from an image II, h^r∈HI\hat{h}_r \in \mathcal{H}_I denote the 2D machine-generated instance ID for ray rr, and Rh⊆RR_h \subseteq R be the set of rays associated with 2D instance h∈HR⊆HIh \in H_R \subseteq \mathcal{H}_I. Each 2D instance hh is mapped to a 3D surrogate identifier j∈Jj \in \mathcal{J} via an injective mapping ΠR⋆\Pi_R^\star: HI→J\mathcal{H}_I \to \mathcal{J}, found by solving the following linear assignment problem:

    ΠR⋆:=arg⁡max⁡ΠR∑h∈HR∑r∈Rhπr(ΠR(h))∣Rh∣\Pi_R^\star := \arg\max_{\Pi_R} \sum_{h \in H_R} \sum_{r \in R_h} \frac{\pi_r(\Pi_R(h))}{|R_h|}

    where πr\pi_r is the volumetrically rendered surrogate identifier probability distribution along ray rr. Given the optimal mapping ΠR⋆\Pi_R^\star, the instance field is supervised with the log-loss:

    Lins(R):=−1∣R∣∑r∈Rwrlog⁡πr(ΠR⋆(h^r))L_{\text{ins}}(R) := -\frac{1}{|R|} \sum_{r \in R} w_r \log \pi_r\left(\Pi_R^\star(\hat{h}_r)\right)

    where wr∈[0,1]w_r \in [0, 1] is a confidence estimate for ray rr derived from test-time augmentations.

  3. Knowl 3 — Segment Consistency Loss

    equation

    To prevent multi-view inconsistent 2D semantic labels from creating a spatially fragmented blend of competing classes across an object's surface, the segment consistency loss encourages all rays corresponding to the same 2D predicted segment to predict the segment's dominant semantic class:

    Lcon(R):=−1∣R∣∑r∈Rwr∑i=1mlog⁡κr(Ki)L_{\text{con}}(R) := -\frac{1}{|R|} \sum_{r \in R} w_r \sum_{i=1}^m \log \kappa_r(K_i)

    where:

    • RR is a batch of rays sampled from a single image, partitioned into mm disjoint clusters {R1,…,Rm}\{R_1, \dots, R_m\} corresponding to individual 2D panoptic segments (distinct 2D instances or stuff regions).
    • wr∈[0,1]w_r \in [0, 1] is the prediction confidence weight for ray rr.
    • κr(k)\kappa_r(k) is the volumetrically rendered probability of class k∈Kk \in \mathcal{K} along ray rr.
    • KiK_i is the confidence-weighted majority semantic class predicted by the neural field across segment RiR_i:

    Ki:=arg⁡max⁡k∈K∑r∈Riwrκr(k)K_i := \arg\max_{k \in \mathcal{K}} \sum_{r \in R_i} w_r \kappa_r(k)

  4. Knowl 4 — Bounded Probability Field Rendering and Density Gradient Blocking

    model/method

    To prevent artifacts caused by noisy machine-generated labels during volumetric rendering, Panoptic Lifting incorporates two structural design constraints:

    1. Bounded Probability Field Rendering: Instead of rendering unbounded classification logits along rays and computing a softmax afterwards (which allows the network to predict strong semantic signals even in low-density space by scaling logits to large magnitudes), the network predicts a bounded probability distribution κ(x)∈Δ∣K∣−1\kappa(x) \in \Delta^{|\mathcal{K}|-1} and π(x)∈Δ∣J∣−1\pi(x) \in \Delta^{|\mathcal{J}|-1} at each 3D point xx. Volume integration directly renders these bounded probability vectors, enforcing geometric adherence.

    2. Semantics-to-Geometry Gradient Blocking: Gradients computed from the semantic loss LsemL_{\text{sem}}, instance loss LinsL_{\text{ins}}, and segment consistency loss LconL_{\text{con}} are stopped before backpropagating into the density field σ\sigma. This prevents the optimizer from degrading underlying scene geometry (creating cloudy or semi-transparent artifacts) to satisfy conflicting multi-view segmentation labels.

  5. Knowl 5 — Confidence-Weighted Pseudo-Labeling via Test-Time Augmentations

    model/method

    Machine-generated 2D panoptic segmentation models (such as Mask2Former) tend to produce overconfident probability peaks even on misclassified regions. To obtain calibrated supervision for the neural fields:

    1. Multiple stochastic test-time augmentations (TTA) — including horizontal flips, multi-scale resizing, brightness adjustments, and contrast variations — are applied to each input training image.
    2. The predicted 2D segments across augmented versions are fused by segment clustering.
    3. The fusion outputs a softened per-ray target semantic class probability distribution κ^r\hat{\kappa}_r over category set K\mathcal{K} and a per-pixel scalar confidence weight wr∈[0,1]w_r \in [0, 1], which weights all downstream semantic and instance loss terms.
  6. Knowl 6 — Panoptic Lifting Optimization Objective and Training Configuration

    model/method

    The complete network parameters are optimized jointly over a ray batch R=RS∪RIR = R_S \cup R_I, where RSR_S is a set of rays sampled uniformly at random across the entire scene and RIR_I is a set of rays sampled from a single image.

    The combined objective function is:

    Ltot(R):=λinsLins(RI)+λconLcon(RI)+λRGBLRGB(RS)+λsemLsem(RS)L_{\text{tot}}(R) := \lambda_{\text{ins}} L_{\text{ins}}(R_I) + \lambda_{\text{con}} L_{\text{con}}(R_I) + \lambda_{\text{RGB}} L_{\text{RGB}}(R_S) + \lambda_{\text{sem}} L_{\text{sem}}(R_S)

    where:

    • LRGB(RS):=1∣RS∣∑r∈RS∥cr−c^r∥2L_{\text{RGB}}(R_S) := \frac{1}{|R_S|} \sum_{r \in R_S} \|c_r - \hat{c}_r\|^2 is the photometric MSE loss against ground-truth color c^r\hat{c}_r.
    • Lsem(RS):=−1∣RS∣∑r∈RSwr∑k∈Kκ^r(k)log⁡κr(k)L_{\text{sem}}(R_S) := -\frac{1}{|R_S|} \sum_{r \in R_S} w_r \sum_{k \in \mathcal{K}} \hat{\kappa}_r(k) \log \kappa_r(k) is the cross-entropy semantic loss.
    • Lins(RI)L_{\text{ins}}(R_I) is the linear assignment instance loss.
    • Lcon(RI)L_{\text{con}}(R_I) is the segment consistency loss.
    • Hyperparameters: λcon=1.35\lambda_{\text{con}} = 1.35, λins=1.0\lambda_{\text{ins}} = 1.0, λsem=1.0\lambda_{\text{sem}} = 1.0, λRGB=1.0\lambda_{\text{RGB}} = 1.0.

    Optimization is performed with Adam for 450,000 iterations using a learning rate of 5×10−45 \times 10^{-4} for the semantic and instance MLPs, and 1×10−21 \times 10^{-2} for the TensoRF color and geometry components.

  7. Knowl 7 — Scene-level Panoptic Quality (PQscene) Metric

    definition

    Standard Panoptic Quality (PQ) assesses segmentation accuracy on individual 2D images independently, failing to measure whether unique instance identifiers remain consistent across multiple viewpoints.

    Scene-level Panoptic Quality (PQscene\text{PQ}^{\text{scene}}) extends the standard matching criterion to the entire multi-view scene:

    1. For a given scene, all predicted 2D segments across all test views sharing the same predicted 3D instance identifier (or stuff class) are merged into a global predicted segment subset Pscene⊂P\mathcal{P}_{\text{scene}} \subset \mathcal{P}.
    2. All ground-truth segments across all views sharing the same object instance (or stuff class) are merged into a global ground-truth subset Gscene⊂G\mathcal{G}_{\text{scene}} \subset \mathcal{G}.
    3. An intersection-over-union is computed on the unified subsets across views, recording a true-positive match if and only if:

    IoU(Pscene,Gscene)>0.5\text{IoU}(\mathcal{P}_{\text{scene}}, \mathcal{G}_{\text{scene}}) > 0.5

    This global formulation penalizes methods that switch object instance IDs between viewpoints, even if the per-frame 2D segmentations are internally accurate.

  8. Knowl 8 — 3D Scene Editing via Panoptic Radiance Fields

    model/method

    The panoptic radiance field representation supports 3D object manipulation without retraining:

    • Object Deletion: For a selected target 3D instance identifier (k⋆,j⋆)(k^\star, j^\star), points xx where the predicted instance ID matches the target have their volume density explicitly set to zero: σ(x)=0\sigma(x) = 0.
    • Object Duplication: For an object to be copied to a target rigid transform (rotation and translation), viewing rays passing through the destination volume query the radiance field at the corresponding transformed coordinates inside the original object's bounding region.
    • Affine Manipulation: Combined deletion and transformed re-querying allows translation, rotation, and rigid transformation of individual 3D objects in synthesized views.
  9. Knowl 9 — Novel View Panoptic and Semantic Benchmark Performance

    data/table

    Panoptic Lifting was evaluated on novel view synthesis, semantic segmentation (mIoU), and panoptic segmentation (PQscene\text{PQ}^{\text{scene}}) across the Hypersim, Replica, and ScanNet datasets against 2D Mask2Former and 3D NeRF baselines trained on identical machine-generated 2D labels:

    Method HyperSim Replica ScanNet
    mIoU↑\uparrow PQscene↑\text{PQ}^{\text{scene}}\uparrow PSNR↑\uparrow mIoU↑\uparrow PQscene↑\text{PQ}^{\text{scene}}\uparrow PSNR↑\uparrow mIoU↑\uparrow PQscene↑\text{PQ}^{\text{scene}}\uparrow PSNR↑\uparrow
    Mask2Former 53.9 – – 52.4 – – 46.7 – –
    SemanticNeRF 58.9 – 26.6 58.5 – 24.8 59.2 – 26.6
    DM-NeRF 57.6 51.6 28.1 56.0 44.1 26.9 49.5 41.7 27.5
    PNF 50.3 44.8 27.4 51.5 41.1 29.8 53.9 48.3 26.7
    PNF + GT Bounding Boxes 58.7 47.6 28.1 54.8 52.5 31.6 58.7 54.3 26.8
    Panoptic Lifting 67.8 60.1 30.1 67.2 57.9 29.6 65.2 58.9 28.5

    Panoptic Lifting outperforms all baselines across all three benchmarks, exceeding the strongest 3D baseline without ground truth bounding boxes by 8.5% to 17.2% in PQscene\text{PQ}^{\text{scene}} and 6.0% to 15.7% in mIoU.

  10. Knowl 10 — Ablation of Panoptic Lifting Robustness Components

    data/table

    An ablation study on the ScanNet dataset demonstrates the contribution of individual robustness components (Segment Consistency Loss, Test-Time Augmentations [TTA], Bounded Segmentation Fields, and Semantics-to-Geometry Gradient Blocking):

    Segment Consistency TTA Bounded Segm. Field Gradient Blocking mIoU↑\uparrow PQscene↑\text{PQ}^{\text{scene}}\uparrow PSNR↑\uparrow
    ✗ ✗ ✗ ✗ 57.3 47.9 27.1
    ✗ ✓ ✓ ✓ 60.9 54.3 28.3
    ✓ ✗ ✓ ✓ 63.1 55.2 28.4
    ✓ ✓ ✗ ✓ 62.9 52.5 28.4
    ✓ ✓ ✓ ✗ 61.6 53.7 27.2
    ✓ ✓ ✓ ✓ 65.2 58.9 28.5

    Key observations:

    • Disabling all components drops mIoU by 7.9% and PQscene\text{PQ}^{\text{scene}} by 11.0%.
    • Removing the segment consistency loss causes the largest individual mIoU reduction (-4.3%).
    • Removing bounded segmentation fields degrades PQscene\text{PQ}^{\text{scene}} by 6.4%.
    • Allowing gradients into density fields causes a 1.3 dB drop in novel view appearance PSNR due to geometric distortion.

Coverage note — None was omitted; all key contributions including neural field parameterization, linear assignment lifting, loss formulations, gradient blocking, bounded fields, TTA confidence weighting, scene-level panoptic quality evaluation, and quantitative experiments are fully covered.

References

  1. 1.Dejan Azinovic, Ricardo Martin-Brualla, Dan B Goldman, Matthias Nießner, and Justus Thies. Neural rgb-d surface reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 6290–6301, June 2022. 2
  2. 2.Jonathan T Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-Brualla, and Pratul P Srinivasan. Mip-nerf: A multiscale representation for anti-aliasing neural radiance fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5855–5864, 2021. 2
  3. 3.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In European conference on computer vision, pages 213–229. Springer, 2020. 3
  4. 4.Anpei Chen, Zexiang Xu, Andreas Geiger, Jingyi Yu, and Hao Su. Tensorf: Tensorial radiance fields. In European Conference on Computer Vision (ECCV), 2022. 2, 3, 4
  5. 5.Bowen Cheng, Maxwell D Collins, Yukun Zhu, Ting Liu, Thomas S Huang, Hartwig Adam, and Liang-Chieh Chen. Panoptic-deeplab: A simple, strong, and fast baseline for bottom-up panoptic segmentation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12475–12485, 2020. 3
  6. 6.Bowen Cheng, Ishan Misra, Alexander G Schwing, Alexander Kirillov, and Rohit Girdhar. Masked-attention mask transformer for universal image segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1290–1299, 2022. 1, 2, 3, 4, 5, 6
  7. 7.Bowen Cheng, Alex Schwing, and Alexander Kirillov. Per-pixel classification is not all you need for semantic segmentation. Advances in Neural Information Processing Systems, 34:17864–17875, 2021. 3
  8. 8.Manuel Dahnert, Ji Hou, Matthias Nießner, and Angela Dai. Panoptic 3d scene reconstruction from a single rgb image. Advances in Neural Information Processing Systems, 34:8282–8293, 2021. 3
  9. 9.Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5828–5839, 2017. 5, 6, 7
  10. 10.Sara Fridovich-Keil, Alex Yu, Matthew Tancik, Qinhong Chen, Benjamin Recht, and Angjoo Kanazawa. Plenoxels: Radiance fields without neural networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5501–5510, 2022. 2
  11. 11.Xiao Fu, Shangzhan Zhang, Tianrun Chen, Yichong Lu, Lanyun Zhu, Xiaowei Zhou, Andreas Geiger, and Yiyi Liao. Panoptic nerf: 3d-to-2d label transfer for panoptic urban scene segmentation. In International Conference on 3D Vision (3DV), 2022. 2, 5
  12. 12.Guy Gafni, Justus Thies, Michael Zollhofer, and Matthias Nießner. Dynamic neural radiance fields for monocular 4d facial avatar reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 8649–8658, June 2021. 2
  13. 13.Stefano Gasperini, Mohammad-Ali Nikouei Mahani, Alvaro Marcos-Ramiro, Nassir Navab, and Federico Tombari. Panoster: End-to-end panoptic segmentation of lidar point clouds. IEEE Robotics and Automation Letters, 6(2):3216–3223, 2021. 3
  14. 14.James T Kajiya and Brian P Von Herzen. Ray tracing volume densities. ACM SIGGRAPH computer graphics, 18(3):165–174, 1984. 3
  15. 15.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014. 5
  16. 16.Alexander Kirillov, Kaiming He, Ross Girshick, Carsten Rother, and Piotr Dollar. Panoptic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9404–9413, 2019. 3, 5
  17. 17.Sosuke Kobayashi, Eiichi Matsumoto, and Vincent Sitzmann. Decomposing nerf for editing via feature field distillation. In Advances in Neural Information Processing Systems, volume 35, 2022. 2
  18. 18.Abhijit Kundu, Kyle Genova, Xiaoqi Yin, Alireza Fathi, Caroline Pantofaru, Leonidas J Guibas, Andrea Tagliasacchi, Frank Dellaert, and Thomas Funkhouser. Panoptic neural fields: A semantic object-aware neural scene representation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12871–12881, 2022. 2, 4, 5, 6, 7, 8
  19. 19.Feng Li, Hao Zhang, Shilong Liu, Lei Zhang, Lionel M Ni, Heung-Yeung Shum, et al. Mask dino: Towards a unified transformer-based framework for object detection and segmentation. arXiv preprint arXiv:2206.02777, 2022. 1
  20. 20.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755. Springer, 2014. 3, 5
  21. 21.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10012–10022, 2021. 3
  22. 22.Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. Communications of the ACM, 65(1):99–106, 2021. 2
  23. 23.Andres Milioto, Jens Behley, Chris McCool, and Cyrill Stachniss. Lidar panoptic segmentation for autonomous driving. In 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 8505–8512. IEEE, 2020. 3
  24. 24.Norman Müller, Andrea Simonelli, Lorenzo Porzi, Samuel Rota Bulò, Matthias Nießner, and Peter Kontschieder. 2
  25. 25.Thomas Müller, Alex Evans, Christoph Schied, and Alexander Keller. Instant neural graphics primitives with a multiresolution hash encoding. arXiv preprint arXiv:2201.05989, 2022. 2
  26. 26.Gaku Narita, Takashi Seno, Tomoya Ishikawa, and Yohsuke Kaji. Panopticfusion: Online volumetric semantic mapping at the level of stuff and things. In 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 4205–4212. IEEE, 2019. 3
  27. 27.Michael Oechsle, Songyou Peng, and Andreas Geiger. Unisurf: Unifying neural implicit surfaces and radiance fields for multi-view reconstruction. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5589–5599, 2021. 2
  28. 28.Lorenzo Porzi, Samuel Rota Bulo, Aleksander Colovic, and Peter Kontschieder. Seamless scene segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8277–8286, 2019. 3
  29. 29.Mike Roberts, Jason Ramapuram, Anurag Ranjan, Atulit Kumar, Miguel Angel Bautista, Nathan Paczan, Russ Webb, and Joshua M Susskind. Hypersim: A photorealistic synthetic dataset for holistic indoor scene understanding. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10912–10922, 2021. 5, 6, 7
  30. 30.Antoni Rosinol, Arjun Gupta, Marcus Abate, Jingnan Shi, and Luca Carlone. 3d dynamic scene graphs: Actionable spatial perception with places, objects, and humans. arXiv preprint arXiv:2002.06289, 2020. 3
  31. 31.Danila Rukhovich, Anna Vorontsova, and Anton Konushin. Imvoxelnet: Image to voxels projection for monocular and multi-view general-purpose 3d object detection. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 2397–2406, 2022. 6
  32. 32.Johannes L Schonberger and Jan-Michael Frahm. Structure-from-motion revisited. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4104–4113, 2016. 1, 5
  33. 33.Kshitij Sirohi, Rohit Mohan, Daniel Buscher, Wolfram Burgard, and Abhinav Valada. Efficientlps: Efficient lidar panoptic segmentation. IEEE Transactions on Robotics, 2021. 3
  34. 34.Julian Straub, Thomas Whelan, Lingni Ma, Yufan Chen, Erik Wijmans, Simon Green, Jakob J Engel, Raul Mur-Artal, Carl Ren, Shobhit Verma, et al. The replica dataset: A digital replica of indoor spaces. arXiv preprint arXiv:1906.05797, 2019. 5, 6, 7
  35. 35.Cheng Sun, Min Sun, and Hwann-Tzong Chen. Direct voxel grid optimization: Super-fast convergence for radiance fields reconstruction. In CVPR, 2022. 2
  36. 36.Matthew Tancik*, Ethan Weber*, Evonne Ng*, Ruilong Li, Brent Yi, Terrance Wang, Alexander Kristoffersen, Jake Austin, Kamyar Salahi, Abhik Ahuja, David McAllister, and Angjoo Kanazawa. Nerfstudio: A framework for neural radiance field development, 2022. 4, 6
  37. 37.Vadim Tschernezki, Iro Laina, Diane Larlus, and Andrea Vedaldi. Neural feature fusion fields: 3d distillation of self-supervised 2d image representations. arXiv preprint arXiv:2209.03494, 2022. 2
  38. 38.Dor Verbin, Peter Hedman, Ben Mildenhall, Todd Zickler, Jonathan T Barron, and Pratul P Srinivasan. Ref-nerf: Structured view-dependent appearance for neural radiance fields. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 5481–5490. IEEE, 2022. 2
  39. 39.Suhani Vora, Noha Radwan, Klaus Greff, Henning Meyer, Kyle Genova, Mehdi SM Sajjadi, Etienne Pot, Andrea Tagliasacchi, and Daniel Duckworth. Nesf: Neural semantic fields for generalizable semantic segmentation of 3d scenes. arXiv preprint arXiv:2111.13260, 2021. 2
  40. 40.Bing Wang, Lu Chen, and Bo Yang. Dm-nerf: 3d scene geometry decomposition and manipulation from 2d images. arXiv preprint arXiv:2208.07227, 2022. 2, 4, 6
  41. 41.Peng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt, Taku Komura, and Wenping Wang. Neus: Learning neural implicit surfaces by volume rendering for multi-view reconstruction. arXiv preprint arXiv:2106.10689, 2021. 2
  42. 42.Qiangeng Xu, Zexiang Xu, Julien Philip, Sai Bi, Zhixin Shu, Kalyan Sunkavalli, and Ulrich Neumann. Point-nerf: Point-based neural radiance fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5438–5448, 2022. 2
  43. 43.Qihang Yu, Huiyu Wang, Siyuan Qiao, Maxwell Collins, Yukun Zhu, Hartwig Adam, Alan Yuille, and Liang-Chieh Chen. k-means mask transformer. In ECCV, 2022. 1
  44. 44.Wenwei Zhang, Jiangmiao Pang, Kai Chen, and Chen Change Loy. K-net: Towards unified image segmentation. Advances in Neural Information Processing Systems, 34:10326–10338, 2021. 3
  45. 45.Shuaifeng Zhi, Tristan Laidlow, Stefan Leutenegger, and Andrew J Davison. In-place scene labelling and understanding with implicit scene representation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 15838–15847, 2021. 2, 4, 6
  46. 46.Zixiang Zhou, Yang Zhang, and Hassan Foroosh. Panoptic-polarnet: Proposal-free lidar point cloud panoptic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13194–13203, 2021. 3

Citation

MLA
Siddiqui, Y., et al. “Panoptic Lifting for 3D Scene Understanding with Neural Fields”. arXiv, 2022, http://arxiv.org/abs/2212.09802v1.
APA
Siddiqui, Y., Porzi, L., Buló, S. R., Müller, N., Nießner, M., Dai, A., & Kontschieder, P. (2022). Panoptic Lifting for 3D Scene Understanding with Neural Fields. arXiv. http://arxiv.org/abs/2212.09802v1
Chicago
Siddiqui, Y., L. Porzi, S. R. Buló, et al. 2022. “Panoptic Lifting for 3D Scene Understanding with Neural Fields”. arXiv. http://arxiv.org/abs/2212.09802v1.
Harvard
Siddiqui, Y. et al. (2022) “Panoptic Lifting for 3D Scene Understanding with Neural Fields”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2212.09802v1.
Vancouver
1. Siddiqui Y, Porzi L, Buló SR, Müller N, Nießner M, Dai A, Kontschieder P (2022) Panoptic Lifting for 3D Scene Understanding with Neural Fields. arXiv

BibTeX

@article{siddiqui2022panoptic,
  title = {Panoptic Lifting for 3D Scene Understanding with Neural Fields},
  author = {Siddiqui, Yawar and Porzi, Lorenzo and Buló, Samuel Rota and Müller, Norman and Nießner, Matthias and Dai, Angela and Kontschieder, Peter},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2212.09802v1},
  eprint = {2212.09802}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE