Breaking Bad: A Dataset for Geometric Fracture and Reassembly
Silvia SellánYun-Chun ChenZiyi WuAnimesh GargAlec Jacobson
Introduces a large-scale benchmark of over one million physically simulated fractured shapes to advance geometric reassembly beyond traditional semantic part composition.
Reassembling fractured objects into their original shapes is a critical capability across multiple domains, including cultural artifact preservation, digital heritage archiving, robotics, computer vision, and geometry processing. While machine learning offers promising avenues to automate this process, progress has been constrained by a lack of suitable training data. Existing shape assembly benchmarks focus predominantly on semantic part decomposition, reflecting how manufactured objects are constructed rather than how materials naturally fracture under physical impacts. In natural fractures, fragments lack distinct semantic identities, exhibit irregular and non-convex geometries, and vary widely in count and volume.
To address this gap, the article introduces Breaking Bad, a large-scale dataset designed to benchmark and advance geometric shape assembly. The primary objective is to model the physical destruction process across a diverse set of three-dimensional shapes and evaluate how modern deep learning architectures perform when reassembling physically realistic fragments. Using an efficient physics-based pre-fracture simulation framework, the authors generated over one million fractured objects derived from more than 10,000 base models. The dataset spans three main categories—everyday items, archaeological artifacts, and general 3D-printing geometries—providing roughly 100 distinct fracture patterns per base shape. To facilitate practical distribution, a lossless compression scheme reduces the raw storage requirement from over 1 terabyte to approximately 7.3 gigabytes.
Benchmarking state-of-the-art deep learning methods—specifically Global, Long Short-Term Memory, and Dynamic Graph Learning architectures—revealed several critical findings. First, existing models perform poorly on geometric fracture reassembly compared to semantic part assembly. Even the top-performing graph-based model achieved a part assembly accuracy of only 31.0% on everyday objects, largely because standard architectures rely on global shape priors rather than reasoning over fine-grained local fracture boundaries. Second, reassembly difficulty scales steeply with fragmentation: as the number of pieces increased from 20 up to 100, part accuracy dropped sharply below 8%. Third, pre-training models on everyday objects improved subsequent fine-tuning performance on archaeological artifacts, raising accuracy from 12.8% to 19.4%. Finally, generalization to entirely unseen object categories remains severely limited, with accuracy falling to single-digit percentages.
These findings indicate that fractured shape reassembly remains a largely unsolved challenge for current computer vision and machine learning frameworks. For decision-makers and researchers invested in automated restoration, heritage archiving, or robotic manipulation, the results demonstrate that deploying off-the-shelf assembly models carries high performance risk. Future investment should focus on developing purpose-built neural architectures that prioritize local geometric surface matching over global category priors, alongside sequential decision-making frameworks tailored for multi-part robotic assembly.
Readers should interpret these results within the scope of the underlying simulation assumptions. The dataset relies on brittle fracture models under isotropic, single-material conditions, meaning ductile deformations (such as bent metals or plastics) and progressive stress propagation are not captured. Despite these boundary conditions, the dataset provides a robust, high-confidence benchmark that establishes a standardized foundation for advancing automated geometric reassembly.
- Paper: A scalable active framework for region annotation in 3D shape collections, Li Yi et al. (2016). It provides the foundational framework and benchmarks for semantic 3D part decomposition, establishing the conventional assembly baseline that Breaking Bad seeks to challenge and replace with physical geometric fracture modeling.
- Paper: Learning Representations and Generative Models for 3D Point Clouds, Panos Achlioptas et al. (2017). It introduces essential representation learning paradigms and generative modeling techniques for unordered 3D point clouds upon which downstream geometric shape assembly architectures rely.
- Paper: PCT: Point cloud transformer, Meng-Hao Guo et al. (2020). It develops the Point Cloud Transformer architecture that serves as a cornerstone deep learning mechanism for capturing both global shape context and local geometric surface relationships.
- Paper: DeepSDF: Learning Continuous Signed Distance Functions for Shape Representation, Jeong Joon Park et al. (2019). It introduces continuous signed distance functions for representing 3D shapes implicitly, providing core concepts for fine-grained geometric representation and surface reasoning.
- Paper: 3D ShapeNets: A deep representation for volumetric shapes, Zhirong Wu et al. (2014). It establishes foundational deep geometric shape representations across large-scale 3D CAD datasets, contextualizing the evolution toward physical fracture benchmarks.
- Paper: Surface Reconstruction from Point Clouds by Learning Predictive Context Priors, Baorui Ma et al. (2022). It advances local geometric feature modeling by learning flexible spatial context priors on point clouds, directly addressing the local boundary reasoning limitations highlighted in Breaking Bad.
- Paper: Learning Local Displacements for Point Cloud Completion, Yida Wang et al. (2022). It extends geometric surface reconstruction by predicting fine local displacements to complete partial 3D point clouds, complementing fractured shape assembly tasks.
- Paper: Objaverse: A Universe of Annotated 3D Objects, Matt Deitke et al. (2022). It dramatically scales up the availability of diverse, artist-annotated 3D object repositories, providing a massive foundation for extending geometric fracture simulations across broader object domains.
