The Replica Dataset: A Digital Replica of Indoor Spaces

Julian StraubThomas WhelanLingni MaYufan ChenErik WijmansSimon GreenJakob J. EngelRaul Mur-ArtalCarl RenShobhit Verma

article2019arXiv1,411 citations

Presents a dataset of eighteen photorealistic 3D indoor reconstructions featuring dense meshes, high-dynamic-range textures, and rich semantic annotations to train embodied AI agents and vision models that transfer directly to real-world environments.

Listen

Training embodied artificial intelligence agents directly in physical environments is costly, slow, and operationally difficult. Simulated environments offer a scalable alternative through parallelization, but existing digital scene datasets suffer from visual artifacts, low lighting fidelity, incomplete geometric boundaries, and imprecise semantic labels. These flaws create a significant domain gap between simulation and the real world, limiting the ability of AI models to transfer their learned behaviors to physical reality.

The article demonstrates the creation and release of Replica, a high-fidelity dataset of 18 photo-realistic 3D indoor scene reconstructions. The primary objective is to provide visually, geometrically, and semantically accurate generative models of physical spaces to train and evaluate embodied AI systems and benchmark 3D perception algorithms.

To construct the dataset, the authors gathered time-aligned motion, visual, and infrared data using a custom handheld capture rig. They reconstructed dense geometric meshes, manually refined missing surfaces, and mapped mirror and glass planes. The workflow incorporated high dynamic range imaging with an 85,000:1 dynamic range (exceeding 16 f-stops) and a two-stage semantic segmentation process that mapped labels from 2D images back to 3D meshes for manual touch-up down to the mesh primitive level.

The article presents several key findings and specifications. First, Replica achieves superior rendering fidelity, incorporating high dynamic range textures and renderable glass and mirror reflectors that are absent from existing datasets like Matterport3D and ScanNet. Second, it delivers higher geometric detail (6,000 primitives per square meter compared to 700 in Matterport3D) and high color resolution (92,000 pixels per square meter). Third, it provides precise semantic labeling across 88 distinct object classes organized into a hierarchical segmentation forest. Fourth, the dataset spans 18 diverse room- and building-scale scenes, including 6 scans of a single apartment across different furniture arrangements to capture variations over time.

These findings indicate that high-fidelity simulations can significantly narrow the gap between virtual training and physical deployment. High dynamic range lighting and accurate surface reflections allow models to handle realistic lighting variations, while clean semantic boundaries improve the performance of geometric inference, 2D and 3D segmentation, and robotic navigation tasks.

The authors recommend that researchers use Replica within compatible platforms such as the AI Habitat simulator to train and evaluate embodied AI agents directly with deep learning frameworks. They also provide a minimal C++ software development kit for custom integration. For future work, expanding the scale and diversity of captured environments will be essential to provide broader training variety.

The primary limitation of this release is its relatively small scope: 18 scenes across 35 rooms compared to hundreds or thousands of scenes in lower-fidelity datasets. While this trade-off favors precision and visual realism over sheer volume, users should exercise caution regarding dataset variety when training models that require massive scene diversity.

Cover for The Replica Dataset: A Digital Replica of Indoor Spaces

Abstract

We introduce Replica, a dataset of 18 highly photo-realistic 3D indoor scene reconstructions at room and building scale. Each scene consists of a dense mesh, high-resolution high-dynamic-range (HDR) textures, per-primitive semantic class and instance information, and planar mirror and glass reflectors. The goal of Replica is to enable machine learning (ML) research that relies on visually, geometrically, and semantically realistic generative models of the world - for instance, egocentric computer vision, semantic segmentation in 2D and 3D, geometric inference, and the development of embodied agents (virtual robots) performing navigation, instruction following, and question answering. Due to the high level of realism of the renderings from Replica, there is hope that ML systems trained on Replica may transfer directly to real world image and video data. Together with the data, we are releasing a minimal C++ SDK as a starting point for working with the Replica dataset. In addition, Replica is `Habitat-compatible', i.e. can be natively used with AI Habitat for training and testing embodied agents.

Table of Contents

  • I Introduction
  • II Related Work
  • II-A Synthetic Scenes
  • II-B Real Scenes
  • III Dataset Creation
  • III-A Mesh and Reflector Fixing
  • III-B Semantic Annotation
  • IV Dataset Description
  • IV-A Data Organization
  • V Conclusion
  • References

Knowls

  1. Knowl 1 — Replica 3D Indoor Scene Dataset Specifications and Scene Composition

    definition

    The Replica dataset is a collection of 18 photo-realistic 3D reconstructions of indoor environments comprising 35 rooms across a variety of architectural spaces:

    • Apartment and Residential Spaces: 6 different furniture configurations of an apartment (capturing the same physical space rearranged over time), 2 multi-room apartments, 3 single rooms within apartments, and 1 two-floor house.
    • Commercial and Hospitality Spaces: 5 office rooms and 1 hotel room.

    Each reconstructed scene provides high-density quad mesh geometry, high-dynamic-range (HDR) surface textures, renderable planar reflector annotations (glass and mirrors), and per-primitive semantic class and instance segmentations covering 88 semantic categories. Categories span permanent structural elements (such as floor, wall, and ceiling), standard furniture objects (such as chair, table, and sofa), and small tabletop or household entities (such as cup, coaster, book, and wall plug).

  2. Knowl 2 — Comparison of Reconstruction-Based 3D Indoor Scene Datasets

    data/table
    Metric Replica Matterport 3D (MP3D) ScanNet Stanford 2D-3D-S Gibson
    # scenes 18 90 1513 6 572
    # rooms 35 2056 707 270 ?
    Color res. [pixel/m2\text{pixel}/\text{m}^2] 92k 97k 20k ≈\approx MP3D ≈\approx MP3D
    Geometry res. [primitives/m2\text{primitives}/\text{m}^2] 6k 0.7k 20k ≈\approx MP3D ≈\approx MP3D
    HDR textures ✓ ×\times ×\times ×\times ×\times
    Reflectors ✓ ×\times ×\times ×\times ×\times
    Semantic classes 88 40 ≈1000\approx 1000 13 –
    Semantic annotation 3D Paint 3D Felsenszwalb 3D Felsenszwalb 3D –

    Color resolution and geometry resolution represent the median number of texture pixels and mesh primitives per square meter, respectively, computed across the semantically annotated meshes in each dataset.

    While Matterport3D achieves comparable color resolution (97k pixels/m297\text{k }\text{pixels}/\text{m}^2), its geometric density is substantially lower (0.7k primitives/m20.7\text{k }\text{primitives}/\text{m}^2) and its semantic labels rely on automated 3D Felsenszwalb pre-segmentation, which introduces boundary inaccuracies. ScanNet offers higher geometric primitive density (20k primitives/m220\text{k }\text{primitives}/\text{m}^2) but lower color resolution (20k pixels/m220\text{k }\text{pixels}/\text{m}^2), lacks complete room coverage, and contains surface holes. Replica provides a balanced resolution (92k pixels/m292\text{k }\text{pixels}/\text{m}^2 color and 6k primitives/m26\text{k }\text{primitives}/\text{m}^2 geometry), introduces 16-bit HDR textures and explicitly modeled reflectors, and provides primitive-level semantic segmentations refined via direct 3D painting.

  3. Knowl 3 — Multi-Sensor RGB-D Scanning Rig and Geometric Reconstruction Pipeline

    model/method

    Replica scene reconstructions are captured using a custom handheld multi-sensor rig equipped with:

    1. An active infrared (IR) pattern projector paired with an IR camera to produce structured-light depth frames.
    2. Wide-angle greyscale tracking cameras paired with an Inertial Measurement Unit (IMU).
    3. An RGB texture camera.

    Pose Estimation and Geometry Generation:

    • 6-DoF Tracking: Visual-inertial Simultaneous Localization and Mapping (SLAM) estimates 6-degree-of-freedom camera trajectories from synchronized wide-angle greyscale streams and IMU data.
    • Volumetric Fusion: Raw depth frames derived from the projected IR pattern are integrated into a Truncated Signed Distance Function (TSDF) volume using the computed 6-DoF poses.
    • Mesh Extraction & Simplification: An initial dense isosurface mesh is extracted from the TSDF volume via Marching Cubes, regularized and simplified into a quad mesh using Instant Meshes, and parameterized for texturing via a PTex-compatible per-face mapping system.
  4. Knowl 4 — High Dynamic Range Texture Generation via Exposure Cycling

    model/method

    To capture photometrically accurate indoor scenes spanning bright illumination sources and dark shadows, the RGB capture camera actively cycles through multiple exposure times throughout scanning.

    Using the estimated 6-DoF SLAM poses, measured camera radiance across varying exposures is fused per surface texel into 16-bit floating-point RGB representations. This radiance fusion yields an overall scene dynamic range of approximately 85,000:185{,}000:1 (exceeding 16 f-stops of dynamic range), avoiding color clamping and saturation present in standard 8-bit RGB vertex colorings or texture maps.

  5. Knowl 5 — Mesh Hole Repair and Planar Reflector Parameterization

    model/method

    Reconstructed meshes undergo two specialized geometry post-processing procedures:

    Planar Reflector Extraction and Modeling: Surfaces that violate standard diffuse structured-light depth sensing (such as glass windows and mirrors) are manually annotated as planar polygons on the mesh. Each reflector is parameterized by:

    • The coordinate transformation from world space to the reflector plane.
    • A 2D planar polygon defining the reflector boundary.
    • The surface normal vector.
    • A scalar reflectance coefficient r∈[0,1]r \in [0, 1], where r=1.0r = 1.0 designates a perfect planar mirror and r<1.0r < 1.0 denotes a partially transparent glass surface.

    Topological Hole Repair: Geometric holes resulting from sensor dropouts or occlusions are identified by searching for boundary edges that form closed topological cycles. An operator selects detected holes, which are automatically triangulated using CGAL polygon mesh processing implementing Liepa's hole-filling method, followed by subdivision refinement and fairing/smoothing against the surrounding mesh curvature.

  6. Knowl 6 — Two-Stage Semantic Annotation Pipeline via Multi-View Label Fusion and 3D Painting

    model/method

    To avoid boundary bleeding and pre-segmentation artifacts common in automated 3D clustering, semantic annotations in Replica are generated through a two-stage process:

    1. Multi-View 2D Annotation and Projection: A set of camera viewpoints is rendered from the mesh to guarantee that every surface primitive is observed at least once. Human annotators apply 2D instance-level masks across these rendered images in parallel. The 2D masks are projected back onto the 3D surface mesh and fused per primitive using a majority voting scheme. Small unobserved gaps and projection noise are resolved using superpixel-like neighborhood smoothing.
    2. Interactive 3D Refinement and Privacy Anonymization: Human annotators review the fused 3D segmentations using an interactive 3D mesh painting tool, manually correcting segment boundaries and mislabeled regions directly at the primitive level. During this stage, regions containing sensitive or personally identifiable information are annotated for privacy anonymization (blurring/pixelation).
  7. Knowl 7 — Segmentation Forest Data Structure for Hierarchical 3D Annotations

    definition

    Replica encodes scene segmentations using a multi-tree graph structure termed a segmentation forest:

    • Leaf Nodes (pip_i): Represent individual geometric primitives (quads/triangles) of the 3D mesh.
    • Intermediate Segmentation Nodes (seg): Group adjacent primitives into localized, continuous geometric surface fragments.
    • Root Nodes: Group intermediate segmentation nodes into discrete semantic object instances (e.g., individual chair, book, or table entities).

    Properties:

    • Instance Segmentation: Each individual tree within the forest corresponds to a unique object instance.
    • Class Segmentation: Derived by mapping all instance root nodes sharing the same semantic class label to a uniform identifier or color.
    • Hierarchical Queries: Traversing and rendering at intermediate levels of the forest enables extraction of fine-grained sub-part segmentations or coarse object groupings without altering the underlying mesh primitives.
    • Storage Format: Graph topology (parent pointers, child lists, instance IDs, and class labels) is stored in JSON format (semantic.json, preseg.json), while primitive membership lists per node are indexed in binary array files (semantic.bin, preseg.bin).
  8. Knowl 8 — Replica File Organization and AI Habitat Simulation Interface

    experimental setup

    Each dataset scene contains the following core files:

    • mesh.ply: Dense quad mesh geometry with fallback per-vertex RGB colors.
    • textures/*: PTex format high-dynamic-range texture archives.
    • glass.sur: Reflector parameters, plane definitions, boundary polygons, surface normals, and reflectance coefficients.
    • semantic.json / semantic.bin: Instance and class segmentation forest metadata and primitive mappings.
    • preseg.json / preseg.bin: Planar versus non-planar surface segmentations.
    • habitat/: Formatted assets for direct integration with the AI Habitat simulator, including:
      • mesh_semantic.ply: Mesh geometry with embedded per-primitive semantic instance IDs.
      • mesh_semantic.navmesh: Navigable surface occupancy mesh for agent motion planning and collision checking.
      • semantic.json: Lookup mapping each instance ID in mesh_semantic.ply to its semantic class name.

    AI Habitat renders RGB, depth, semantic class segmentation, and semantic instance segmentation tensors from Replica models directly into PyTorch tensors at rendering speeds up to 10,00010{,}000 frames per second.

Coverage note — No substantial contributed material was omitted. The knowls cover all dataset specifications, pipeline stages (capture, TSDF fusion, HDR texturing, reflector extraction, and hole filling), annotation mechanisms, data structures, and simulator integration.

References

  1. 1.Peter Anderson, Angel X. Chang, Devendra Singh Chaplot, Alexey Dosovitskiy, Saurabh Gupta, Vladlen Koltun, Jana Kosecka, Jitendra Malik, Roozbeh Mottaghi, Manolis Savva, and Amir Roshan Zamir. On evaluation of embodied navigation agents. arXiv:1807.06757, 2018.
  2. 2.Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko Sunderhauf, Ian Reid, Stephen Gould, and Anton van den Hen- ¨ gel. Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments. In CVPR, 2018.
  3. 3.Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. VQA: Visual Question Answering. In ICCV, 2015.
  4. 4.Iro Armeni, Ozan Sener, Amir R. Zamir, Helen Jiang, Ioannis Brilakis, Martin Fischer, and Silvio Savarese. 3D semantic parsing of large-scale indoor spaces. In CVPR, 2016.
  5. 5.Brent Burley and Dylan Lacewell. Ptex: Per-face texture mapping for production rendering. In Computer Graphics Forum, volume 27, pages 1155–1164. Wiley Online Library, 2008.
  6. 6.Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Halber, Matthias Niessner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang. Matterport3D: Learning from RGB-D data in indoor environments. In 3DV, 2017. https://niessner.github.io/Matterport/.
  7. 7.Kenneth J. W. Craik. The Nature of Explanation. Cambridge University Press, 1943.
  8. 8.Angela Dai, Angel X. Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3D reconstructions of indoor scenes. In CVPR, 2017. http://www.scan-net.org/.
  9. 9.Abhishek Das, Samyak Datta, Georgia Gkioxari, Stefan Lee, Devi Parikh, and Dhruv Batra. Embodied Question Answering. In CVPR, 2018.
  10. 10.Jakob Engel, Vladlen Koltun, and Daniel Cremers. Direct sparse odometry. TPAMI, 40(3):611–625, 2017.
  11. 11.Pedro F Felzenszwalb and Daniel P Huttenlocher. Efficient graph-based image segmentation. IJCV, 59(2):167–181, 2004.
  12. 12.Matthew Fisher, Daniel Ritchie, Manolis Savva, Thomas Funkhouser, and Pat Hanrahan. Example-based synthesis of 3D object arrange-ments. In ACM SIGGRAPH Asia, 2012.
  13. 13.Alberto Garcia-Garcia, Pablo Martinez-Gonzalez, Sergiu Oprea, John Alejandro Castro-Vargas, Sergio Orts-Escolano, Jose Garcia-Rodriguez, and Alvaro Jover-Alvarez. The robotrix: An extremely photorealistic and very-large-scale indoor dataset of sequences with robot trajectories and interactions. In IROS, pages 6790–6797. IEEE, 2018.
  14. 14.A Handa, V Patraucean, V Badrinarayanan, S Stent, and R Cipolla. Scenenet: understanding real world indoor scenes with synthetic data. arxiv preprint (2015). arXiv preprint arXiv:1511.07041, 2015.
  15. 15.Wenzel Jakob, Marco Tarini, Daniele Panozzo, and Olga Sorkine-Hornung. Instant field-aligned meshes. ACM Transactions on Graph-ics, 34(6), November 2015.
  16. 16.Alex Krizhevsky, Ilya Sutskever, and Geoff Hinton. ImageNet classi-fication with deep convolutional neural networks. In NIPS, 2012.
  17. 17.Wenbin Li, Sajad Saeedi, John McCormac, Ronald Clark, Dimos Tzoumanikas, Qing Ye, Yuzhong Huang, Rui Tang, and Stefan Leutenegger. Interiornet: Mega-scale multi-sensor photo-realistic in-door scenes dataset. In BMVC, 2018.
  18. 18.Peter Liepa. Filling holes in meshes. In ACM SIGGRAPH Symposium on Geometry Processing, pages 200–205, 2003.
  19. 19.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Per-ona, Deva Ramanan, Piotr Dollr, and C. Lawrence Zitnick. Microsoft COCO: Common objects in context. In ECCV, 2014.
  20. 20.William E. Lorensen and Harvey E. Cline. Marching cubes: A high resolution 3D surface construction algorithm. In Proceedings of the 14th Annual Conference on Computer Graphics and Interactive Techniques, SIGGRAPH ’87, pages 163–169, New York, NY, USA, 1987. ACM.
  21. 21.Sebastien Loriot, Jane Tournois, and Ilker O. Yaz. Polygon mesh ´ processing. In CGAL User and Reference Manual. CGAL Editorial Board, 4.14 edition, 2019.
  22. 22.Raul Mur-Artal, Jose Maria Martinez Montiel, and Juan D Tardos. ORB-SLAM: a versatile and accurate monocular SLAM system. TRO, 31(5):1147–1163, 2015.
  23. 23.Richard A Newcombe, Shahram Izadi, Otmar Hilliges, David Molyneaux, David Kim, Andrew J. Davison, Pushmeet Kohi, Jamie Shotton, Steve Hodges, and Andrew Fitzgibbon. Kinectfusion: Real-time dense surface mapping and tracking. In 2011 IEEE International Symposium on Mixed and Augmented Reality, pages 127–136. IEEE, 2011.
  24. 24.Manolis Savva*, Abhishek Kadian*, Oleksandr Maksymets*, Yili Zhao, Erik Wijmans, Bhavana Jain, Julian Straub, Jia Liu, Vladlen Koltun, Jitendra Malik, Devi Parikh, and Dhruv Batra. Habitat: A platform for embodied ai research. arXiv preprint arXiv:1904.01201, 2019.
  25. 25.Shuran Song, Fisher Yu, Andy Zeng, Angel X Chang, Manolis Savva, and Thomas Funkhouser. Semantic scene completion from a single depth image. In CVPR, 2017.
  26. 26.R. S. Sutton and A. G. Barto. An adaptive network that constructs and uses an internal model of its world. Cognition and Brain Theory, 1981.
  27. 27.Thomas Whelan, Michael Goesele, Steven J. Lovegrove, Julian Straub, Simon Green, Richard Szeliski, Steven Butterfield, Shobhit Verma, and Richard Newcombe. Reconstructing scenes with mirror and glass surfaces. ACM Transactions on Graphics (TOG), 37(4):102, 2018.
  28. 28.Fei Xia, Amir R. Zamir, Zhiyang He, Alexander Sax, Jitendra Malik, and Silvio Savarese. Gibson env: Real-world perception for embodied agents. In CVPR, 2018. http://gibsonenv.stanford.edu/database/.

Citation

MLA
Straub, J., et al. “The Replica Dataset: A Digital Replica of Indoor Spaces”. arXiv, 2019, http://arxiv.org/abs/1906.05797v1.
APA
Straub, J., Whelan, T., Ma, L., Chen, Y., Wijmans, E., Green, S., Engel, J. J., Mur-Artal, R., Ren, C., Verma, S., Clarkson, A., Yan, M., Budge, B., Yan, Y., Pan, X., Yon, J., Zou, Y., Leon, K., Carter, N., … Newcombe, R. (2019). The Replica Dataset: A Digital Replica of Indoor Spaces. arXiv. http://arxiv.org/abs/1906.05797v1
Chicago
Straub, J., T. Whelan, L. Ma, et al. 2019. “The Replica Dataset: A Digital Replica of Indoor Spaces”. arXiv. http://arxiv.org/abs/1906.05797v1.
Harvard
Straub, J. et al. (2019) “The Replica Dataset: A Digital Replica of Indoor Spaces”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1906.05797v1.
Vancouver
1. Straub J, Whelan T, Ma L, et al (2019) The Replica Dataset: A Digital Replica of Indoor Spaces. arXiv

BibTeX

@article{straub2019the,
  title = {The Replica Dataset: A Digital Replica of Indoor Spaces},
  author = {Straub, Julian and Whelan, Thomas and Ma, Lingni and Chen, Yufan and Wijmans, Erik and Green, Simon and Engel, Jakob J. and Mur-Artal, Raul and Ren, Carl and Verma, Shobhit and Clarkson, Anton and Yan, Mingfei and Budge, Brian and Yan, Yajie and Pan, Xiaqing and Yon, June and Zou, Yuyang and Leon, Kimberly and Carter, Nigel and Briales, Jesus and Gillingham, Tyler and Mueggler, Elias and Pesqueira, Luis and Savva, Manolis and Batra, Dhruv and Strasdat, Hauke M. and Nardi, Renzo De and Goesele, Michael and Lovegrove, Steven and Newcombe, Richard},
  year = {2019},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1906.05797v1},
  eprint = {1906.05797}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/