The Replica Dataset: A Digital Replica of Indoor Spaces
Julian StraubThomas WhelanLingni MaYufan ChenErik WijmansSimon GreenJakob J. EngelRaul Mur-ArtalCarl RenShobhit Verma
Presents a dataset of eighteen photorealistic 3D indoor reconstructions featuring dense meshes, high-dynamic-range textures, and rich semantic annotations to train embodied AI agents and vision models that transfer directly to real-world environments.
Training embodied artificial intelligence agents directly in physical environments is costly, slow, and operationally difficult. Simulated environments offer a scalable alternative through parallelization, but existing digital scene datasets suffer from visual artifacts, low lighting fidelity, incomplete geometric boundaries, and imprecise semantic labels. These flaws create a significant domain gap between simulation and the real world, limiting the ability of AI models to transfer their learned behaviors to physical reality.
The article demonstrates the creation and release of Replica, a high-fidelity dataset of 18 photo-realistic 3D indoor scene reconstructions. The primary objective is to provide visually, geometrically, and semantically accurate generative models of physical spaces to train and evaluate embodied AI systems and benchmark 3D perception algorithms.
To construct the dataset, the authors gathered time-aligned motion, visual, and infrared data using a custom handheld capture rig. They reconstructed dense geometric meshes, manually refined missing surfaces, and mapped mirror and glass planes. The workflow incorporated high dynamic range imaging with an 85,000:1 dynamic range (exceeding 16 f-stops) and a two-stage semantic segmentation process that mapped labels from 2D images back to 3D meshes for manual touch-up down to the mesh primitive level.
The article presents several key findings and specifications. First, Replica achieves superior rendering fidelity, incorporating high dynamic range textures and renderable glass and mirror reflectors that are absent from existing datasets like Matterport3D and ScanNet. Second, it delivers higher geometric detail (6,000 primitives per square meter compared to 700 in Matterport3D) and high color resolution (92,000 pixels per square meter). Third, it provides precise semantic labeling across 88 distinct object classes organized into a hierarchical segmentation forest. Fourth, the dataset spans 18 diverse room- and building-scale scenes, including 6 scans of a single apartment across different furniture arrangements to capture variations over time.
These findings indicate that high-fidelity simulations can significantly narrow the gap between virtual training and physical deployment. High dynamic range lighting and accurate surface reflections allow models to handle realistic lighting variations, while clean semantic boundaries improve the performance of geometric inference, 2D and 3D segmentation, and robotic navigation tasks.
The authors recommend that researchers use Replica within compatible platforms such as the AI Habitat simulator to train and evaluate embodied AI agents directly with deep learning frameworks. They also provide a minimal C++ software development kit for custom integration. For future work, expanding the scale and diversity of captured environments will be essential to provide broader training variety.
The primary limitation of this release is its relatively small scope: 18 scenes across 35 rooms compared to hundreds or thousands of scenes in lower-fidelity datasets. While this trade-off favors precision and visual realism over sheer volume, users should exercise caution regarding dataset variety when training models that require massive scene diversity.
- Paper: Matterport3D: Learning from RGB-D Data in Indoor Environments, Angel Chang et al. (2017). Matterport3D established the standard for building-scale, reconstructed RGB-D indoor environments with dense semantic annotations, providing the foundational benchmark and methodology that Replica directly refines with higher photorealism.
- Paper: ScanNet: Richly-Annotated 3D Reconstructions of Indoor Scenes, Angela Dai et al. (2017). ScanNet introduced large-scale semantic annotation of real-world indoor 3D mesh reconstructions, which Replica builds upon by adding high-dynamic-range textures and planar reflector modeling.
- Paper: SUN RGB-D: A RGB-D scene understanding benchmark suite, Shuran Song et al. (2015). SUN RGB-D laid the groundwork for RGB-D indoor scene understanding and 3D bounding annotation benchmarks that motivated the creation of higher-fidelity reconstructed indoor datasets like Replica.
- Paper: Indoor Segmentation and Support Inference from RGBD Images, Nathan Silberman et al. (2012). This seminal work introduced the NYU Depth v2 dataset and established early protocols for dense indoor RGB-D segmentation and geometric reasoning.
- Paper: ShapeNet: An Information-Rich 3D Model Repository, Angel X. Chang et al. (2015). ShapeNet created the standardized taxonomy and large-scale repository for 3D object models and semantics that underpins modern 3D semantic datasets.
- Paper: Habitat: A Platform for Embodied AI Research, Manolis Savva et al. (2019). Habitat establishes the simulation platform explicitly designed to ingest datasets like Replica for training and evaluating embodied AI navigation agents at scale.
- Paper: NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis, Ben Mildenhall et al. (2020). NeRF introduces neural radiance fields for continuous novel view synthesis, frequently utilizing high-quality multi-view indoor captures like Replica for benchmarking novel-view geometry and rendering.
- Paper: NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view Reconstruction, Peng Wang et al. (2021). NeuS improves multi-view neural surface reconstruction via volume rendering of signed distance functions, advancing the reconstruction fidelity of intricate indoor surfaces beyond classical textured meshes.
- Paper: 3D Gaussian Splatting for Real-Time Radiance Field Rendering, Bernhard Kerbl et al. (2023). 3D Gaussian Splatting provides real-time radiance field rendering of complex indoor and outdoor scenes, advancing the visual realism and interactive simulation capabilities originally targeted by Replica.
- Paper: Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance Fields, Jonathan T. Barron et al. (2022). Mip-NeRF 360 extends neural radiance fields to 360-degree unbounded indoor and outdoor scenes with anti-aliased view synthesis.
- Paper: Objaverse: A Universe of Annotated 3D Objects, Matt Deitke et al. (2022). Objaverse scales 3D asset diversity to hundreds of thousands of annotated objects, complementing scene datasets like Replica for open-vocabulary embodied simulation.
- Paper: DUSt3R: Geometric 3D Vision Made Easy, Shuzhe Wang et al. (2023). DUSt3R simplifies multi-view 3D reconstruction into end-to-end regression of dense pointmaps from uncalibrated images, advancing dense geometry recovery in indoor environments.
