FunFact: Building Probabilistic Functional 3D Scene Graphs via Factor-Graph Reasoning
Zhengyu FuRené ZurbrüggKaixian QuMarc PollefeysMarco HutterHermann BlumZuria Bauer
Proposes a factor-graph reasoning framework that combines geometric constraints with language model priors to construct probabilistic functional 3D scene graphs with accurately calibrated relation predictions from RGB-D images.
Autonomous robots, assistive augmented reality, and virtual training environments increasingly require an understanding of how physical spaces function rather than just how they are shaped. Identifying which switch operates a specific light fixture or which knob controls a stove burner is difficult because static visual observations rarely expose underlying causal links. Prior systems typically evaluate pairs of objects in isolation, which leaves them vulnerable to visual ambiguities and produces poorly calibrated, often overconfident predictions that can mislead downstream autonomous decision-making.
To overcome these limitations, the article introduces FunFact, a framework designed to construct open-vocabulary functional 3D scene graphs from posed RGB-D image sequences. The pipeline combines a hierarchical 3D scene reconstruction process with a probabilistic factor-graph reasoning module to jointly infer interactive relationships and generate well-calibrated confidence scores across entire environments.
The approach operates in two main phases. In the first phase, a vision-language model generates object and part proposals that are verified using grounded 2D detectors, segmented, lifted into 3D, and fused across viewpoints to create an explicit object- and part-centric point cloud representation. In the second phase, a large language model proposes semantically plausible relationships alongside structural constraints, such as one-to-one pairings and proximity requirements. FunFact formulates these candidate relationships as binary variables in a factor graph, where proximity factors and cardinality constraints resolve ambiguities through probabilistic belief propagation. To thoroughly evaluate performance, the authors also introduce FunThor, a synthetic benchmark built on AI2-THOR containing 12 dense scenes across four environment types with ground-truth functional annotations across 720 images, alongside evaluations on real-world datasets including SceneFun3D and FunGraph3D.
Experiments show substantial improvements in discovering functional entities and relationships across complex scenes. On the FunThor benchmark, FunFact achieves an overall triplet discovery recall of 54.1% compared to 15.1% for the baseline OpenFunGraph, representing an absolute improvement of approximately 39 percentage points. The framework also improves overall triplet precision from 23.4% to 31.9% and boosts the functional F1 score from 16.0% to 38.7%. Crucially, the factor-graph formulation dramatically improves confidence calibration: on ambiguous classes such as light switches and stove knobs, the expected calibration error falls from 0.51 to 0.07. Furthermore, on the FunGraph3D dataset, overall triplet recall improves from 29.8% to 48.7%, while ablation studies confirm that both factor-graph optimization and hierarchical object-part proposals are essential for accurately resolving fine-grained interactive elements.
These findings demonstrate that modeling functional relationships jointly rather than independently provides significant operational value. Calibrated confidence estimates enable autonomous systems to reliably distinguish between certain and uncertain interactions, which is critical for planning safe real-world manipulation tasks, preventing accidental activations, and guiding targeted exploration. Moreover, the results highlight that rigid vocabulary matching benchmarks can penalize valid open-vocabulary predictions, underscoring the need for richer functional evaluation standards.
For real-world deployment, practitioners should adopt hierarchical object-part grounding combined with joint probabilistic inference when building semantic and functional maps. In the near term, developers should leverage calibrated uncertainty scores to trigger interactive verification actions—such as having a robot test a low-confidence switch—to update beliefs online. To prepare for time-critical operational deployments, engineering efforts should prioritize reducing query latencies through parallelized vision-language calls or local model caching.
Certain limitations should be noted. The framework exhibits occasional over-segmentation or under-segmentation on complex repetitive furniture, and its current reliance on large vision-language models incurs an average scene reconstruction time of several seconds per frame, making it unsuited for real-time applications without system optimization. Additionally, evaluation in real-world benchmarks remains somewhat constrained by sparse or generic ground-truth annotations, though high consistency across synthetic and real datasets provides strong confidence in the underlying methodology.
- Paper: Scene Graph Generation by Iterative Message Passing, Danfei Xu et al. (2017). This foundational paper establishes message-passing graph architectures for scene graph generation, providing the essential contextual graph reasoning principles that FunFact extends to 3D factor graphs.
- Paper: Learning to Generate Language-Supervised and Open-Vocabulary Scene Graph Using Pre-Trained Visual-Semantic Space, Yong Zhang et al. (2023). It introduces open-vocabulary scene graph generation leveraging pre-trained vision-language models, which informs FunFact's open-vocabulary functional relation proposals.
- Paper: GARField: Group Anything with Radiance Fields, Chung Min Kim et al. (2024). It details hierarchical and part-centric 3D scene decomposition from multi-view images, laying the structural groundwork for constructing FunFact's object- and part-level 3D maps.
- Paper: Indoor Segmentation and Support Inference from RGBD Images, Nathan Silberman et al. (2012). This seminal work demonstrates RGB-D scene segmentation and joint inference of physical support relationships under geometric constraints, setting the classic precedent for functional relation inference.
- Paper: Open3DIS: Open-Vocabulary 3D Instance Segmentation with 2D Mask Guidance, Phuc D. A. Nguyen et al. (2024). It establishes techniques for lifting 2D open-vocabulary mask guidance into 3D instance proposals, directly relevant to FunFact's front-end entity mapping.
- Paper: FunRec: Reconstructing Functional 3D Scenes from Egocentric Interaction Videos, Alexandros Delitzas et al. (2026). It builds upon functional and interactive 3D scene modeling by reconstructing dynamic, articulated digital twins directly from egocentric video interactions.
- Paper: SG2Loc: Sequential Visual Localization on 3D Scene Graphs, Nicole Damblon et al. (2026). It applies structured 3D scene graph representations as lightweight topological maps for visual sequential camera localization.
- Paper: OVI-MAP:Open-Vocabulary Instance-Semantic Mapping, Zilong Deng et al. (2026). It extends open-vocabulary instance mapping into an online, real-time SLAM setting by decoupling geometric mapping from viewpoint-selected vision-language querying.
- Paper: HiSpatial: Taming Hierarchical 3D Spatial Understanding in Vision-Language Models, Huizhi Liang et al. (2026). It generalizes multi-level spatial and relational reasoning in 3D into a hierarchical vision-language model architecture.
