FunFact: Building Probabilistic Functional 3D Scene Graphs via Factor-Graph Reasoning

Zhengyu FuRené ZurbrüggKaixian QuMarc PollefeysMarco HutterHermann BlumZuria Bauer

article2026arXiv8 citations

Proposes a factor-graph reasoning framework that combines geometric constraints with language model priors to construct probabilistic functional 3D scene graphs with accurately calibrated relation predictions from RGB-D images.

Listen

Autonomous robots, assistive augmented reality, and virtual training environments increasingly require an understanding of how physical spaces function rather than just how they are shaped. Identifying which switch operates a specific light fixture or which knob controls a stove burner is difficult because static visual observations rarely expose underlying causal links. Prior systems typically evaluate pairs of objects in isolation, which leaves them vulnerable to visual ambiguities and produces poorly calibrated, often overconfident predictions that can mislead downstream autonomous decision-making.

To overcome these limitations, the article introduces FunFact, a framework designed to construct open-vocabulary functional 3D scene graphs from posed RGB-D image sequences. The pipeline combines a hierarchical 3D scene reconstruction process with a probabilistic factor-graph reasoning module to jointly infer interactive relationships and generate well-calibrated confidence scores across entire environments.

The approach operates in two main phases. In the first phase, a vision-language model generates object and part proposals that are verified using grounded 2D detectors, segmented, lifted into 3D, and fused across viewpoints to create an explicit object- and part-centric point cloud representation. In the second phase, a large language model proposes semantically plausible relationships alongside structural constraints, such as one-to-one pairings and proximity requirements. FunFact formulates these candidate relationships as binary variables in a factor graph, where proximity factors and cardinality constraints resolve ambiguities through probabilistic belief propagation. To thoroughly evaluate performance, the authors also introduce FunThor, a synthetic benchmark built on AI2-THOR containing 12 dense scenes across four environment types with ground-truth functional annotations across 720 images, alongside evaluations on real-world datasets including SceneFun3D and FunGraph3D.

Experiments show substantial improvements in discovering functional entities and relationships across complex scenes. On the FunThor benchmark, FunFact achieves an overall triplet discovery recall of 54.1% compared to 15.1% for the baseline OpenFunGraph, representing an absolute improvement of approximately 39 percentage points. The framework also improves overall triplet precision from 23.4% to 31.9% and boosts the functional F1 score from 16.0% to 38.7%. Crucially, the factor-graph formulation dramatically improves confidence calibration: on ambiguous classes such as light switches and stove knobs, the expected calibration error falls from 0.51 to 0.07. Furthermore, on the FunGraph3D dataset, overall triplet recall improves from 29.8% to 48.7%, while ablation studies confirm that both factor-graph optimization and hierarchical object-part proposals are essential for accurately resolving fine-grained interactive elements.

These findings demonstrate that modeling functional relationships jointly rather than independently provides significant operational value. Calibrated confidence estimates enable autonomous systems to reliably distinguish between certain and uncertain interactions, which is critical for planning safe real-world manipulation tasks, preventing accidental activations, and guiding targeted exploration. Moreover, the results highlight that rigid vocabulary matching benchmarks can penalize valid open-vocabulary predictions, underscoring the need for richer functional evaluation standards.

For real-world deployment, practitioners should adopt hierarchical object-part grounding combined with joint probabilistic inference when building semantic and functional maps. In the near term, developers should leverage calibrated uncertainty scores to trigger interactive verification actions—such as having a robot test a low-confidence switch—to update beliefs online. To prepare for time-critical operational deployments, engineering efforts should prioritize reducing query latencies through parallelized vision-language calls or local model caching.

Certain limitations should be noted. The framework exhibits occasional over-segmentation or under-segmentation on complex repetitive furniture, and its current reliance on large vision-language models incurs an average scene reconstruction time of several seconds per frame, making it unsuited for real-time applications without system optimization. Additionally, evaluation in real-world benchmarks remains somewhat constrained by sparse or generic ground-truth annotations, though high consistency across synthetic and real datasets provides strong confidence in the underlying methodology.

arXiv: 2604.03696leggedrobotics/funfact-scenegraph
Cover for FunFact: Building Probabilistic Functional 3D Scene Graphs via Factor-Graph Reasoning

Abstract

Recent work in 3D scene understanding is moving beyond purely spatial analysis toward functional scene understanding. However, existing methods often consider functional relationships between object pairs in isolation, failing to capture the scene-wide interdependence that humans use to resolve ambiguity. We introduce FunFact, a framework for constructing probabilistic open-vocabulary functional 3D scene graphs from posed RGB-D images. FunFact first builds an object- and part-centric 3D map and uses foundation models to propose semantically plausible functional relations. These candidates are converted into factor graph variables and constrained by both LLM-derived common-sense priors and geometric priors. This formulation enables joint probabilistic inference over all functional edges and their marginals, yielding substantially better calibrated confidence scores. To benchmark this setting, we introduce FunThor, a synthetic dataset based on AI2-THOR with part-level geometry and rule-based functional annotations. Experiments on SceneFun3D, FunGraph3D, and FunThor show that FunFact improves node and relation discovery recall and significantly reduces calibration error for ambiguous relations, highlighting the benefits of holistic probabilistic modeling for functional scene understanding. See our project page at this https URL

Citation

MLA
Fu, Z., et al. “FunFact: Building Probabilistic Functional 3D Scene Graphs via Factor-Graph Reasoning”. arXiv, 2026, https://doi.org/10.48550/arxiv.2604.03696.
APA
Fu, Z., Zurbrügg, R., Qu, K., Pollefeys, M., Hutter, M., Blum, H., & Bauer, Z. (2026). FunFact: Building Probabilistic Functional 3D Scene Graphs via Factor-Graph Reasoning. arXiv. https://doi.org/10.48550/arxiv.2604.03696
Chicago
Fu, Z., R. Zurbrügg, K. Qu, et al. 2026. “FunFact: Building Probabilistic Functional 3D Scene Graphs via Factor-Graph Reasoning”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2604.03696.
Harvard
Fu, Z. et al. (2026) “FunFact: Building Probabilistic Functional 3D Scene Graphs via Factor-Graph Reasoning”. arXiv. Available at: https://doi.org/10.48550/arxiv.2604.03696.
Vancouver
1. Fu Z, Zurbrügg R, Qu K, Pollefeys M, Hutter M, Blum H, Bauer Z (2026) FunFact: Building Probabilistic Functional 3D Scene Graphs via Factor-Graph Reasoning. https://doi.org/10.48550/arxiv.2604.03696

BibTeX

@misc{https://doi.org/10.48550/arxiv.2604.03696,
  doi = {10.48550/ARXIV.2604.03696},
  url = {https://arxiv.org/abs/2604.03696},
  author = {Fu, Zhengyu and Zurbrügg, René and Qu, Kaixian and Pollefeys, Marc and Hutter, Marco and Blum, Hermann and Bauer, Zuria},
  keywords = {Computer Vision and Pattern Recognition (cs.CV), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {FunFact: Building Probabilistic Functional 3D Scene Graphs via Factor-Graph Reasoning},
  publisher = {arXiv},
  year = {2026},
  copyright = {Creative Commons Attribution 4.0 International}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/