DL3DV-10K: A Large-Scale Scene Dataset for Deep Learning-based 3D Vision

Lu LingYichen ShengZhi TuWentian ZhaoCheng XinKun WanLantao YuQianyu GuoZixun YuYawen Lu

article2024CVPR561 citations

Introduces a large-scale real-world dataset of over ten thousand 4K videos across diverse scene types to benchmark novel view synthesis methods and train generalizable neural radiance fields.

Listen

Deep learning techniques for 3D computer vision and novel view synthesis—the ability to generate new visual perspectives of a scene from captured images—are crucial for emerging applications in virtual reality, augmented reality, and spatial simulation. However, progress in this domain has been significantly constrained by existing scene-level datasets, which rely heavily on synthetic environments or small, domain-specific collections of real-world captures. These existing resources fail to capture complex real-world visual phenomena, such as complex lighting, transparency, reflections, and unbounded outdoor spaces, preventing rigorous benchmarking and limiting the ability of models to learn universal 3D visual representations.

The article addresses this bottleneck by introducing DL3DV-10K, a massive multi-view scene dataset, and evaluating state-of-the-art 3D reconstruction and rendering algorithms across diverse, real-world conditions. Specifically, the article aims to benchmark leading novel view synthesis techniques on a standardized set of challenging real-world scenes and demonstrate that large-scale pretraining on diverse real scenes enhances the generalizability of 3D representation models.

To construct this resource, the researchers collected 10,510 high-resolution (4K) videos containing over 51.2 million frames across 65 categories of everyday locations, ranging from restaurants and retail spaces to outdoor parks. The team employed standard consumer smartphones and drones, applying standardized capture paths and rigorous quality screening to protect personal privacy and minimize blur. Each scene was systematically annotated based on environmental settings (indoor versus outdoor), lighting conditions, surface reflectivity, material transparency, and texture frequency. From this collection, the authors established a balanced benchmark of 140 representative scenes, DL3DV-140, to rigorously test five leading view-synthesis methods under standardized experimental constraints, alongside pretraining experiments for generalizable neural rendering architectures.

The benchmark evaluation yielded four critical findings. First, Zip-NeRF and 3D Gaussian Splatting (3DGS) consistently outperformed older neural rendering baselines in image reconstruction quality across all standardized visual metrics. Second, 3DGS achieved rendering quality comparable to the top neural methods while requiring a fraction of the computational training time (2.1 hours versus 48 hours for Mip-NeRF 360), though Zip-NeRF achieved slightly higher absolute visual fidelity at the cost of higher memory demands. Third, across all evaluated methods, unbounded outdoor environments and scenes with high transparency proved to be the most challenging conditions, resulting in the lowest overall quality scores. Finally, pilot experiments showed that pretraining generalizable neural models on DL3DV-10K consistently improved their downstream performance on unseen target benchmarks, whereas pretraining on narrower indoor-only datasets failed to produce similar improvements.

These findings indicate that 3D vision systems can successfully transition toward foundational models capable of general scene understanding when trained on diverse real-world multi-view data. For decision-makers and technical leaders, the results highlight distinct trade-offs between rendering speed, compute cost, and visual accuracy. Teams prioritizing rapid training and real-time rendering can deploy 3DGS, while those prioritizing maximum detail on complex surfaces may prefer advanced grid-based neural radiance fields, provided they account for greater hardware memory requirements. Moreover, the failure of indoor-only datasets to generalize proves that training data diversity is non-negotiable for robust real-world performance.

Organizations developing 3D spatial platforms should leverage large-scale, diverse real-world datasets for pretraining rather than relying exclusively on synthetic or single-domain indoor captures. Future research and development should focus specifically on improving rendering accuracy in unbounded outdoor scenes and complex transparent environments, where current methods still struggle. Additionally, teams should explore dynamic novel view synthesis models that can natively handle transient elements, such as moving objects or pedestrians.

The findings are supported by comprehensive statistical evaluations across 140 diverse scenes and standard vision metrics. However, decision-makers should note that DL3DV-10K primarily targets static scenes, even though a minority of captures contain brief appearances of moving objects (between 3 and 10 seconds) inherent to real-world data collection. Overall, the evidence provides high confidence that large-scale, fine-grained real scene datasets are essential for building robust, generalizable 3D visual representations.

No sufficiently relevant recommendations were found.

Cover for DL3DV-10K: A Large-Scale Scene Dataset for Deep Learning-based 3D Vision

Abstract

We have witnessed significant progress in deep learning-based 3D vision, ranging from neural radiance field (NeRF) based 3D representation learning to applications in novel view synthesis (NVS). However, existing scene-level datasets for deep learning-based 3D vision, limited to either synthetic environments or a narrow selection of real-world scenes, are quite insufficient. This insufficiency not only hinders a comprehensive benchmark of existing methods but also caps what could be explored in deep learning-based 3D analysis. To address this critical gap, we present DL3DV-10K, a large-scale scene dataset, featuring 51.2 million frames from 10,510 videos captured from 65 types of point-of-interest (POI) locations, covering both bounded and unbounded scenes, with different levels of reflection, transparency, and lighting. We conducted a comprehensive benchmark of recent NVS methods on DL3DV-10K, which revealed valuable insights for future research in NVS. In addition, we have obtained encouraging results in a pilot study to learn generalizable NeRF from DL3DV-10K, which manifests the necessity of a large-scale scene-level dataset to forge a path toward a foundation model for learning 3D representation. Our DL3DV-10K dataset, benchmark results, and models will be publicly accessible.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Novel View Synthesis
  • 2.2. Multi-view Scene Dataset
  • 3. Data Acquisition and Processing
  • 3.1. Data Acquisition
  • 3.2. Data Processing
  • 3.3. Data Statistics
  • 3.4. Benchmark
  • 4. Experiment
  • 4.1. Evaluation on the NVS benchmark
  • 4.2. Generalizable NeRF
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — DL3DV-10K Dataset Specifications and Composition

    definition

    DL3DV-10K is a large-scale, multi-view real-world scene dataset designed for deep learning-based 3D vision and novel view synthesis (NVS). The dataset comprises 10,510 video sequences totaling 51.3 million frames captured at 4K resolution (3840×21603840 \times 2160) at 30 or 60 frames per second. The captures cover 16 primary and 65 secondary Point-of-Interest (POI) categories spanning everyday indoor and outdoor environments, including restaurants, tourist attractions, shopping centers, educational institutions, transportation hubs, and outdoor natural areas.

    Of the 10,510 scenes, 10,407 were captured using consumer mobile devices and 103 using drones. In terms of transient dynamic content, 8,064 scenes contain moving objects for less than 3 seconds, while 2,446 scenes contain moving objects for 3 to 10 seconds. Each scene includes camera pose annotations and fine-grained labels across five scene complexity dimensions:

    • Environment Setting: Indoor (bounded) vs. outdoor (unbounded).
    • Lighting Condition: Natural lighting (nlight), artificial lighting (alight), or mixed lighting (mlight).
    • Surface Reflectivity: Categorized into four classes (none, less, medium, more) based on the proportion of reflective pixels and their persistence across frames.
    • Material Transparency: Categorized into four classes (none, less, medium, more) based on transparent surface area and duration.
    • Texture Frequency: High vs. low frequency determined by high-frequency wavelet subband energy.
  2. Knowl 2 — Novel View Synthesis Performance Comparison on DL3DV-140

    data/table

    State-of-the-art novel view synthesis methods were evaluated on the 140 scenes of the DL3DV-140 benchmark at 960×560960 \times 560 resolution using a 7/87/8 train and 1/81/8 test split. The results report mean Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index Measure (SSIM), Learned Perceptual Image Patch Similarity (LPIPS), training time (in hours), and peak GPU memory usage (in gigabytes):

    Method PSNR ↑\uparrow SSIM ↑\uparrow LPIPS ↓\downarrow Train ↓\downarrow Mem ↓\downarrow
    Instant-NGP 25.01 0.834 0.228 1.2 hr 3.9 GB
    Nerfacto 24.61 0.848 0.211 2.6 hr 3.7 GB
    Mip-NeRF 360 30.98 0.911 0.132 48.0 hr 23.6 GB
    3DGS 29.82 0.919 0.120 2.1 hr 16.8 GB
    Zip-NeRF* 29.07 0.878 0.169 2.5 hr 23.8 GB
    Zip-NeRF 31.22 0.921 0.112 4.0 hr 38.2 GB

    Zip-NeRF uses its default batch size of 65,536 rays, whereas Zip-NeRF* uses a batch size of 4,096 identical to the other NeRF baselines. Zip-NeRF achieves the highest reconstruction fidelity (PSNR 31.22, SSIM 0.921, LPIPS 0.112), but requires the largest memory (38.2 GB). 3D Gaussian Splatting (3DGS) achieves competitive quality (SSIM 0.919, LPIPS 0.120) with fast training (2.1 hours) and moderate memory (16.8 GB). Mip-NeRF 360 achieves strong visual quality (PSNR 30.98) but requires 48 hours of training time. Instant-NGP and Nerfacto train quickly with low memory footprint but yield lower PSNR and SSIM.

  3. Knowl 3 — Effect of DL3DV-10K Pretraining on Generalizable IBRNet

    data/table

    The effect of pretraining the generalizable Image-Based Rendering Network (IBRNet) on subsets of the DL3DV-10K dataset was evaluated against training from scratch and pretraining on ScanNet++ (270 scenes). Models were evaluated on the Diffuse Synthetic 360∘360^\circ and Real Forward-Facing (LLFF) benchmarks:

    Diffuse Synthetic 360∘360^\circ Real Forward-Facing
    Method PSNR ↑\uparrow SSIM ↑\uparrow LPIPS ↓\downarrow PSNR ↑\uparrow SSIM ↑\uparrow LPIPS ↓\downarrow
    IBRNet (Scratch) 34.72 0.983 0.024 24.82 0.808 0.178
    IBRNet-S (ScanNet++ 270) 34.22 0.979 0.024 24.86 0.807 0.183
    IBRNet-270 (DL3DV 270) 35.18 0.984 0.024 25.00 0.812 0.180
    IBRNet-1K (DL3DV 1,000) 35.13 0.984 0.023 25.02 0.814 0.175
    IBRNet-2K (DL3DV 2,000) 35.34 0.984 0.024 25.08 0.815 0.176

    Pretraining on ScanNet++ (IBRNet-S) degrades PSNR on Diffuse Synthetic (34.22 vs. 34.72 scratch). In contrast, pretraining on DL3DV subsets (DL3DV-270, DL3DV-1K, DL3DV-2K) consistently improves PSNR, SSIM, and LPIPS over training from scratch across both test datasets, with performance scaling with the volume of DL3DV pretraining data (reaching 35.34 PSNR on Diffuse Synthetic and 25.08 PSNR on Real Forward-Facing with DL3DV-2K).

  4. Knowl 4 — Effect of DL3DV-10K Pretraining on Generalizable MVSNeRF

    data/table

    The effect of pretraining Multi-View Stereo NeRF (MVSNeRF) on DL3DV-10K subsets (DL3DV-270, DL3DV-1K, DL3DV-2K) was evaluated against training from scratch and pretraining on ScanNet++ (270 scenes). Evaluation was performed on the DTU validation set and the Real Forward-Facing (LLFF) dataset:

    DTU Validation Real Forward-Facing
    Method PSNR ↑\uparrow SSIM ↑\uparrow LPIPS ↓\downarrow PSNR ↑\uparrow SSIM ↑\uparrow LPIPS ↓\downarrow
    MVSNeRF (Scratch) 18.49 0.598 0.476 16.16 0.350 0.549
    MVSNeRF-S (ScanNet++ 270) 19.74 0.660 0.428 17.50 0.442 0.541
    MVSNeRF-270 (DL3DV 270) 20.00 0.671 0.397 17.64 0.449 0.492
    MVSNeRF-1K (DL3DV 1,000) 20.42 0.671 0.396 17.89 0.458 0.492
    MVSNeRF-2K (DL3DV 2,000) 20.85 0.671 0.394 17.90 0.459 0.505

    Pretraining on DL3DV-10K produces significant quantitative improvements across all metrics compared to training MVSNeRF from scratch (e.g., PSNR increases from 18.49 to 20.85 on DTU, and from 16.16 to 17.90 on Real Forward-Facing). Furthermore, pretraining on DL3DV-270 outperforms pretraining on an equal number of scenes from ScanNet++ (MVSNeRF-S) across both DTU (20.00 vs. 19.74 PSNR) and Real Forward-Facing (17.64 vs. 17.50 PSNR), with gains continuing to grow up to 2,000 pretraining scenes.

  5. Knowl 5 — DL3DV-140 Novel View Synthesis Benchmark Specification

    experimental setup

    DL3DV-140 is a standardized evaluation benchmark composed of 140 static scenes sampled from the DL3DV-10K dataset. The benchmark is designed to evaluate novel view synthesis (NVS) methods across balanced complexity attributes.

    The 140 scenes are balanced across four binary complexity axes:

    1. Environment: Indoor (bounded) vs. outdoor (unbounded).
    2. Texture Frequency: High-frequency (high-freq) vs. low-frequency (low-freq).
    3. Reflectivity: More reflection (more-ref, combining more and medium) vs. less reflection (less-ref, combining less and none).
    4. Transparency: More transparency (more-transp, combining more and medium) vs. less transparency (less-transp, combining less and none).

    Evaluation Protocol:

    • Image Resolution: All scenes are evaluated at a downsampled resolution of 960×560960 \times 560 (scale factor of 4 from 4K).
    • Dataset Split: Each scene contains 300 to 380 images. Views are split into 7/87/8 for training and 1/81/8 for testing.
    • Ray Sampling & Bounds: For NeRF-based methods, near bound is fixed to 0.050.05 and far bound to 10610^6. Ray batch size is set to 4,096 rays per batch for standardized comparison (with 65,536 also tested for Zip-NeRF).
    • Evaluation Metrics: Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index Measure (SSIM), and Learned Perceptual Image Patch Similarity (LPIPS).
  6. Knowl 6 — Scene Texture Frequency Estimation via Wavelet Subband Decomposition

    model/method

    To quantify texture complexity and high-frequency visual detail across scenes in DL3DV-10K, a wavelet-based frequency metric is computed over captured video frames:

    1. Frame Sampling: For each video, M=100M = 100 RGB frames are sampled uniformly across its duration.
    2. Grayscale Conversion: Each sampled RGB frame is converted to a single-channel grayscale image with intensities normalized to [0,1][0, 1].
    3. Wavelet Decomposition: A 2D bi-orthogonal wavelet transform is applied to each grayscale frame, generating four subbands: low-low (LL), low-high (LH), high-low (HL), and high-high (HH).
    4. High-Frequency Energy Computation: High-frequency energy is computed as the Frobenius norm of the ensemble of LH, HL, and HH subbands: Ehigh=∥Whigh∥F=∑k∈{LH,HL,HH}∑u,v∣Wk(u,v)∣2E_{\text{high}} = \|\mathbf{W}_{\text{high}}\|_F = \sqrt{\sum_{k \in \{\text{LH}, \text{HL}, \text{HH}\}} \sum_{u, v} |W_k(u, v)|^2}
    5. Scene Frequency Metric: The Frobenius norm is normalized by the total pixel count NpixN_{\text{pix}}, and the final frequency metric FsceneF_{\text{scene}} is the average quotient over the 100 sampled frames: Fscene=1M∑m=1M∥Whigh(m)∥FNpixF_{\text{scene}} = \frac{1}{M} \sum_{m=1}^{M} \frac{\|\mathbf{W}_{\text{high}}^{(m)}\|_F}{N_{\text{pix}}}

    Scenes are categorized into high-frequency and low-frequency groups based on their computed FsceneF_{\text{scene}} values.

  7. Knowl 7 — Impact of Scene Complexity Attributes on Novel View Synthesis Methods

    empirical result

    Statistical evaluation of novel view synthesis (NVS) methods across specific scene complexity dimensions in DL3DV-140 reveals the following systematic performance trends:

    • Indoor vs. Outdoor Environments: Outdoor (unbounded) scenes are the most challenging condition for all NVS methods, yielding the lowest PSNR and SSIM scores across all tested architectures due to large scale disparities between foregrounds and distant backgrounds.
    • Texture Frequency: Low-frequency scenes yield the highest PSNR and SSIM scores across all methods. In low-frequency settings, Mip-NeRF 360 demonstrates the highest performance and the lowest variance across scenes. High-frequency scenes pose severe aliasing challenges for both grid-based and ray-sampling methods.
    • Reflective Surfaces: High-reflectivity scenes (more-ref) exhibit reduced performance across all models compared to low-reflectivity scenes (less-ref). Zip-NeRF and Mip-NeRF 360 show smoother rendering of broad reflections, whereas 3DGS accurately resolves sharp, specular highlights but tends to oversimplify soft reflections. Instant-NGP and Nerfacto frequently generate floating artifacts around reflective regions.
    • Transparency: Highly transparent scenes (more-transp) result in lower PSNR and SSIM than less transparent scenes (less-transp). 3DGS effectively delineates the thin, subtle silhouettes of transparent objects, while volumetric NeRF methods struggle with accurate boundary placement.
  8. Knowl 8 — Visual Artifact Profiles of NeRF Variants and 3D Gaussian Splatting

    empirical result

    Evaluation on DL3DV-140 demonstrates distinct visual artifact modalities between coordinate-based Neural Radiance Field (NeRF) variants and 3D Gaussian Splatting (3DGS):

    • NeRF Variants (Instant-NGP, Nerfacto, Mip-NeRF 360, Zip-NeRF):

      • Produce high-frequency "grainy" microstructural noise and cloudy floaters when ray sampling is insufficient or depth is ambiguous.
      • Exhibit distance-scale sensitivity in unbounded outdoor scenes, resulting in blurry background structures and floating background artifacts (particularly in Instant-NGP).
      • Struggle with geometric aliasing on fine structures such as grass blades and foliage in high-frequency scenes, even with anti-aliasing cone tracing.
      • Render reflections and lighting variations as generalized diffuse hazes or floating artifacts rather than sharp view-dependent highlights.
    • 3D Gaussian Splatting (3DGS):

      • Generates elongated, anisotropic "splotchy" Gaussian splat artifacts in regions with sparse camera coverage or low geometric constraints.
      • Produces prominent artifacts in far-distance unbounded backgrounds, notably across open sky and distant building boundaries.
      • Mitigates aliasing on sharp, intricate foreground geometries (such as foliage) significantly better than NeRF variants.
      • Excels at rendering sharp specular reflections on polished metal and glass surfaces and capturing precise silhouettes of transparent objects, though it tends to oversimplify low-contrast, diffuse reflections (such as cloud reflections on glass).
  9. Knowl 9 — Video Capture Protocol and Quality Guidelines for DL3DV-10K

    model/method

    To collect diverse, high-density, multi-view real-world scenes while minimizing reconstruction artifacts, the DL3DV-10K data acquisition protocol specifies the following recording requirements:

    1. Scene Coverage: The capture trajectory follows a circle or semi-circle with a walking diameter of 30 to 45 seconds, covering at least five distinct objects/instances in a natural arrangement.
    2. Camera Lens/Focal Length: Recording is performed using standard 0.5×\times ultra-wide camera mode on consumer mobile phones to capture extensive background geometry.
    3. Multi-Angle Coverage: Trajectories encompass a horizontal angular sweep of at least 180∘180^\circ or 360∘360^\circ captured across multiple heights, including overhead and waist levels, providing dense viewpoints of objects within the scene.
    4. Resolution and Frame Rate: Videos are recorded at 4K resolution (3840×21603840 \times 2160) at 30 fps or 60 fps.
    5. Video Length: Minimum duration of 60 seconds for mobile phone captures and 45 seconds for drone captures.
    6. Dynamic Object Constraints: Transient moving objects are restricted to under 3 seconds duration (with an absolute ceiling of 10 seconds).
    7. Optical Quality: Frames must be stereoscopic, free of motion blur, and avoid extreme overexposure or underexposure.
    8. Privacy and Sensitive Data Screening: Audio and camera metadata are stripped, and all identifiable personal information (faces, names, license numbers) is detected and mosaiced.
  10. Knowl 10 — Transient Dynamic Elements in Handheld Scene Captures

    limitation

    While DL3DV-10K is designed primarily as a static scene dataset for 3D novel view synthesis and representation learning, the real-world nature of consumer mobile video recording in public and semi-public spaces results in occasional dynamic elements. Specifically, 2,446 of the 10,510 captured scenes contain transient moving objects (such as pedestrians or vehicles) persisting for 3 to 10 seconds (the remaining 8,064 scenes limit dynamic presence to under 3 seconds). While these transient elements introduce potential reconstruction artifacts for purely static novel view synthesis methods, they present opportunities and benchmarks for learning robust 3D representations and dynamic novel view synthesis algorithms.

Coverage note — No substantial contributed material was omitted; all key contributions including dataset statistics, acquisition guidelines, frequency estimation, benchmark results, complexity analyses, and generalizable pretraining findings are covered.

References

  1. 1.Adel Ahmadyan, Liangkai Zhang, Artsiom Ablavatski, Jianing Wei, and Matthias Grundmann. Objectron: A large scale dataset of object-centric videos in the wild with pose annotations. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 7822–7831, 2021.
  2. 2.Jonathan T Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-Brualla, and Pratul P Srinivasan. Mip-nerf: A multiscale representation for anti-aliasing neural radiance fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5855–5864, 2021.
  3. 3.Jonathan T Barron, Ben Mildenhall, Dor Verbin, Pratul P Srinivasan, and Peter Hedman. Mip-nerf 360: Unbounded anti-aliased neural radiance fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5470–5479, 2022.
  4. 4.Jonathan T Barron, Ben Mildenhall, Dor Verbin, Pratul P Srinivasan, and Peter Hedman. Zip-nerf: Anti-aliased grid-based neural radiance fields. arXiv preprint arXiv:2304.06706, 2023.
  5. 5.Gilad Baruch, Zhuoyuan Chen, Afshin Dehghan, Tal Dimry, Yuri Feigin, Peter Fu, Thomas Gebauer, Brandon Joffe, Daniel Kurz, Arik Schwartz, et al. Arkitscenes: A diverse real-world dataset for 3d indoor scene understanding using mobile rgb-d data. arXiv preprint arXiv:2111.08897, 2021.
  6. 6.Manush Bhatt, Rajesh Kalyanam, Gen Nishida, Liu He, Christopher May, Dev Niyogi, and Daniel Aliaga. Design and deployment of photo2building: A cloud-based procedural modeling tool as a service. In Practice and Experience in Advanced Research Computing, pages 132–138. 2020.
  7. 7.Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Halber, Matthias Niessner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang. Matterport3d: Learning from rgb-d data in indoor environments. arXiv preprint arXiv:1709.06158, 2017.
  8. 8.Angel X Chang, Thomas Funkhouser, Leonidas Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, et al. Shapenet: An information-rich 3d model repository. arXiv preprint arXiv:1512.03012, 2015.
  9. 9.Anpei Chen, Zexiang Xu, Fuqiang Zhao, Xiaoshuai Zhang, Fanbo Xiang, Jingyi Yu, and Hao Su. Mvsnerf: Fast generalizable radiance field reconstruction from multi-view stereo. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 14124–14133, 2021.
  10. 10.Julian Chibane, Aayush Bansal, Verica Lazova, and Gerard Pons-Moll. Stereo radiance fields (srf): Learning view synthesis for sparse views of novel scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7911–7920, 2021.
  11. 11.Albert Cohen, Ingrid Daubechies, and J-C Feauveau. Biorthogonal bases of compactly supported wavelets. Communications on pure and applied mathematics, 45(5):485–560, 1992.
  12. 12.Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5828–5839, 2017.
  13. 13.Matt Deitke, Ruoshi Liu, Matthew Wallingford, Huong Ngo, Oscar Michel, Aditya Kusupati, Alan Fan, Christian Laforte, Vikram Voleti, Samir Yitzhak Gadre, et al. Objaverse-xl: A universe of 10m+ 3d objects. arXiv preprint arXiv:2307.05663, 2023.
  14. 14.Yuchun Huang, Ping Ma, Zheng Ji, and Liu He. Part-based modeling of pole-like objects using divergence-incorporated 3-d clustering of mobile laser scanning point clouds. IEEE Transactions on Geoscience and Remote Sensing, 59(3): 2611–2626, 2020.
  15. 15.Rasmus Jensen, Anders Dahl, George Vogiatzis, Engin Tola, and Henrik Aanæs. Large scale multi-view stereopsis evaluation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 406–413, 2014.
  16. 16.Bernhard Kerbl, Georgios Kopanas, Thomas Leimkuhler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Transactions on Graphics (ToG), 42(4):1–14, 2023.
  17. 17.Arno Knapitsch, Jaesik Park, Qian-Yi Zhou, and Vladlen Koltun. Tanks and temples: Benchmarking large-scale scene reconstruction. ACM Transactions on Graphics (ToG), 36(4):1–13, 2017.
  18. 18.Marc Levoy and Pat Hanrahan. Light field rendering. In Seminal Graphics Papers: Pushing the Boundaries, Volume 2, pages 441–452. 2023.
  19. 19.Ben Mildenhall, Pratul P Srinivasan, Rodrigo Ortiz-Cayon, Nima Khademi Kalantari, Ravi Ramamoorthi, Ren Ng, and Abhishek Kar. Local light field fusion: Practical view synthesis with prescriptive sampling guidelines. ACM Transactions on Graphics (TOG), 38(4):1–14, 2019.
  20. 20.Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. Communications of the ACM, 65(1):99–106, 2021.
  21. 21.Thomas Muller, Alex Evans, Christoph Schied, and Alexander Keller. Instant neural graphics primitives with a multiresolution hash encoding. ACM Transactions on Graphics (ToG), 41(4):1–15, 2022.
  22. 22.Jeremy Reizenstein, Roman Shapovalov, Philipp Henzler, Luca Sbordone, Patrick Labatut, and David Novotny. Common objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10901–10911, 2021.
  23. 23.Konstantinos Rematas, Tobias Ritschel, Mario Fritz, Efstratios Gavves, and Tinne Tuytelaars. Deep reflectance maps. In Proceedings of the IEEE Conference on computer vision and pattern recognition, pages 4508–4516, 2016.
  24. 24.Konstantinos Rematas, Ricardo Martin-Brualla, and Vittorio Ferrari. Sharf: Shape-conditioned radiance fields from a single view. arXiv preprint arXiv:2102.08860, 2021.
  25. 25.Thomas Schops, Johannes L Schonberger, Silvano Galliani, Torsten Sattler, Konrad Schindler, Marc Pollefeys, and Andreas Geiger. A multi-view stereo benchmark with high-resolution images and multi-camera videos. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3260–3269, 2017.
  26. 26.Yichen Sheng, Jianming Zhang, and Bedrich Benes. Ssn: Soft shadow network for image compositing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4380–4390, 2021.
  27. 27.Yichen Sheng, Yifan Liu, Jianming Zhang, Wei Yin, A Cengiz Oztireli, He Zhang, Zhe Lin, Eli Shechtman, and Bedrich Benes. Controllable shadow generation using pixel height maps. In European Conference on Computer Vision, pages 240–256. Springer, 2022.
  28. 28.Yichen Sheng, Jianming Zhang, Julien Philip, Yannick Hold-Geoffroy, Xin Sun, He Zhang, Lu Ling, and Bedrich Benes. Pixht-lab: Pixel height based light effect generation for image compositing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16643–16653, 2023.
  29. 29.Vincent Sitzmann, Justus Thies, Felix Heide, Matthias Nießner, Gordon Wetzstein, and Michael Zollhofer. Deepvoxels: Learning persistent 3d feature embeddings. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2437–2446, 2019.
  30. 30.Yizhi Song, Zhifei Zhang, Zhe Lin, Scott Cohen, Brian Price, Jianming Zhang, Soo Ye Kim, and Daniel Aliaga. Object-stitch: Object compositing with diffusion model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18310–18319, 2023.
  31. 31.Yizhi Song, Zhifei Zhang, Zhe Lin, Scott Cohen, Brian Price, Jianming Zhang, Soo Ye Kim, He Zhang, Wei Xiong, and Daniel Aliaga. Imprint: Generative object compositing by learning identity-preserving representation. arXiv preprint arXiv:2403.10701, 2024.
  32. 32.Matthew Tancik, Ethan Weber, Evonne Ng, Ruilong Li, Brent Yi, Terrance Wang, Alexander Kristoffersen, Jake Austin, Kamyar Salahi, Abhik Ahuja, et al. Nerfstudio: A modular framework for neural radiance field development. In ACM SIGGRAPH 2023 Conference Proceedings, pages 1–12, 2023.
  33. 33.Alex Trevithick and Bo Yang. Grf: Learning a general radiance field for 3d representation and rendering. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 15182–15192, 2021.
  34. 34.Dan Wang, Xinrui Cui, Septimiu Salcudean, and Z Jane Wang. Generalizable neural radiance fields for novel view synthesis with transformer. arXiv preprint arXiv:2206.05375, 2022.
  35. 35.Qianqian Wang, Zhicheng Wang, Kyle Genova, Pratul P Srinivasan, Howard Zhou, Jonathan T Barron, Ricardo Martin-Brualla, Noah Snavely, and Thomas Funkhouser. Ibrnet: Learning multi-view image-based rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4690–4699, 2021.
  36. 36.Haozhe Xie, Hongxun Yao, Shengping Zhang, Shangchen Zhou, and Wenxiu Sun. Pix2vox++: Multi-scale context-aware 3d object reconstruction from single and multiple images. International Journal of Computer Vision, 128(12): 2919–2935, 2020.
  37. 37.Qiangeng Xu, Zexiang Xu, Julien Philip, Sai Bi, Zhixin Shu, Kalyan Sunkavalli, and Ulrich Neumann. Point-nerf: Point-based neural radiance fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5438–5448, 2022.
  38. 38.Hao Yang, Lanqing Hong, Aoxue Li, Tianyang Hu, Zhenguo Li, Gim Hee Lee, and Liwei Wang. Contranerf: Generalizable neural radiance fields for synthetic-to-real novel view synthesis via contrastive learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16508–16517, 2023.
  39. 39.Yao Yao, Zixin Luo, Shiwei Li, Jingyang Zhang, Yufan Ren, Lei Zhou, Tian Fang, and Long Quan. Blendedmvs: A large-scale dataset for generalized multi-view stereo networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 1790–1799, 2020.
  40. 40.Mao Ye, Peifeng Yin, Wang-Chien Lee, and Dik-Lun Lee. Exploiting geographical influence for collaborative point-of-interest recommendation. In Proceedings of the 34th international ACM SIGIR conference on Research and development in Information Retrieval, pages 325–334, 2011.
  41. 41.Chandan Yeshwanth, Yueh-Cheng Liu, Matthias Nießner, and Angela Dai. Scannet++: A high-fidelity dataset of 3d indoor scenes. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 12–22, 2023.
  42. 42.Alex Yu, Ruilong Li, Matthew Tancik, Hao Li, Ren Ng, and Angjoo Kanazawa. Plenoctrees for real-time rendering of neural radiance fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5752–5761, 2021.
  43. 43.Alex Yu, Vickie Ye, Matthew Tancik, and Angjoo Kanazawa. pixelnerf: Neural radiance fields from one or few images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4578–4587, 2021.
  44. 44.Xianggang Yu, Mutian Xu, Yidan Zhang, Haolin Liu, Chongjie Ye, Yushuang Wu, Zizheng Yan, Chenming Zhu, Zhangyang Xiong, Tianyou Liang, et al. Mvimgnet: A large-scale dataset of multi-view images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9150–9161, 2023.
  45. 45.Andy Zeng, Shuran Song, Matthias Nießner, Matthew Fisher, Jianxiong Xiao, and Thomas Funkhouser. 3dmatch: Learning local geometric descriptors from rgb-d reconstructions. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1802–1811, 2017.
  46. 46.Yinda Zhang, Shuran Song, Ersin Yumer, Manolis Savva, Joon-Young Lee, Hailin Jin, and Thomas Funkhouser. Physically-based rendering for indoor scene understanding using convolutional neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5287–5295, 2017.
  47. 47.Tinghui Zhou, Richard Tucker, John Flynn, Graham Fyffe, and Noah Snavely. Stereo magnification: Learning view synthesis using multiplane images. arXiv preprint arXiv:1805.09817, 2018.

Citation

MLA
Ling, L., et al. “DL3DV-10K: A Large-Scale Scene Dataset for Deep Learning-based 3D Vision”. arXiv, 2023, http://arxiv.org/abs/2312.16256v2.
APA
Ling, L., Sheng, Y., Tu, Z., Zhao, W., Xin, C., Wan, K., Yu, L., Guo, Q., Yu, Z., Lu, Y., Li, X., Sun, X., Ashok, R., Mukherjee, A., Kang, H., Kong, X., Hua, G., Zhang, T., Benes, B., & Bera, A. (2023). DL3DV-10K: A Large-Scale Scene Dataset for Deep Learning-based 3D Vision. arXiv. http://arxiv.org/abs/2312.16256v2
Chicago
Ling, L., Y. Sheng, Z. Tu, et al. 2023. “DL3DV-10K: A Large-Scale Scene Dataset for Deep Learning-based 3D Vision”. arXiv. http://arxiv.org/abs/2312.16256v2.
Harvard
Ling, L. et al. (2023) “DL3DV-10K: A Large-Scale Scene Dataset for Deep Learning-based 3D Vision”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2312.16256v2.
Vancouver
1. Ling L, Sheng Y, Tu Z, et al (2023) DL3DV-10K: A Large-Scale Scene Dataset for Deep Learning-based 3D Vision. arXiv

BibTeX

@article{ling2023dl3dv,
  title = {DL3DV-10K: A Large-Scale Scene Dataset for Deep Learning-based 3D Vision},
  author = {Ling, Lu and Sheng, Yichen and Tu, Zhi and Zhao, Wentian and Xin, Cheng and Wan, Kun and Yu, Lantao and Guo, Qianyu and Yu, Zixun and Lu, Yawen and Li, Xuanmao and Sun, Xingpeng and Ashok, Rohan and Mukherjee, Aniruddha and Kang, Hao and Kong, Xiangrui and Hua, Gang and Zhang, Tianyi and Benes, Bedrich and Bera, Aniket},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2312.16256v2},
  eprint = {2312.16256}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE