pix2gestalt: Amodal Segmentation by Synthesizing Wholes
Ege OzgurogluRuoshi LiuDídac SurísDian ChenAchal DavePavel TokmakovCarl Vondrick
Presents a zero-shot framework that transfers pre-trained diffusion representations to reconstruct and segment entire occluded objects, outperforming fully supervised baselines and boosting downstream object recognition and 3D reconstruction.
Real-world computer vision applications in robotics, autonomous navigation, and augmented reality frequently fail when objects are partially hidden behind obstructions. While humans effortlessly perceive complete shapes and physical extents despite heavy occlusions, machine vision models have historically struggled with this task, known as amodal perception. Prior automated systems have been confined to narrow, closed-world categories or synthetic environments, severely limiting their practical deployment in complex real-world settings.
The article evaluates a new framework, named pix2gestalt, designed to perform zero-shot amodal completion and segmentation by synthesizing the whole, unoccluded appearance and shape of objects from partial visual inputs.
To achieve this, the researchers adapted pre-trained large-scale diffusion models, which implicitly encode rich representations of natural objects. The model was fine-tuned on an automatically curated dataset of 837,000 paired images. This dataset was constructed by identifying foreground objects in natural images using depth estimation and overlaying realistic synthetic occlusions to create ground-truth whole-part pairs. The system conditions on an input image and a prompt indicating the visible region, using iterative generation to synthesize complete images. This output serves as a direct input for downstream visual tasks, including segmentation, image classification, and three-dimensional reconstruction across standard benchmarks.
The evaluation yielded several key findings. First, pix2gestalt achieved state-of-the-art amodal segmentation accuracy, reaching an 82.87% mean intersection-over-union score on Amodal COCO without prior training on that dataset, outperforming fully supervised models. Second, incorporating the framework into object recognition pipelines substantially improved classification accuracy under occlusion; on the challenging Separated COCO benchmark, top-one accuracy rose from 26.04% using standard open-vocabulary classification to 31.15% with pix2gestalt, with even larger gains on standard occlusions where accuracy rose from 23.33% to 43.39%. Third, when serving as a drop-in module for three-dimensional reconstruction tools, it more than doubled volumetric accuracy and significantly reduced geometric error compared to standard baselines. Finally, the framework demonstrated robust zero-shot generalization across out-of-distribution inputs, including photographs, visual illusions, and artistic works.
These results demonstrate that generative image completion can act as a versatile foundation module to eliminate occlusion-related performance bottlenecks across diverse computer vision pipelines. By synthesizing complete objects rather than relying on narrow category-specific masks, organizations can improve system robustness in unstructured environments without incurring the high cost of collecting task-specific human annotations.
Organizations developing perception stacks should evaluate pix2gestalt as a front-end pre-processing module for existing recognition and three-dimensional reconstruction workflows. When handling ambiguous occlusions, engineering teams should deploy multi-sample generation with majority voting to select the most reliable completion. Further research should focus on integrating physical reasoning constraints into the generative process to prevent plausible-looking but physically impossible completions.
Confidence in these findings is supported by consistent quantitative improvements across multiple recognized benchmarks and tasks. However, decision-makers should note that the model relies on probabilistic sampling and can occasionally fail in scenarios requiring complex common-sense or physical reasoning, such as predicting correct motion directions or contact mechanics.
- Paper: Globally and locally consistent image completion, SATOSHI IIZUKA et al. (2017). It provides foundational principles for context-guided image completion and visual inpainting that directly motivate generative amodal completion.
- Paper: Image-to-Image Translation with Conditional Adversarial Networks, Phillip Isola et al. (2017). It establishes the foundational conditional image-to-image synthesis formulation that modern conditional generative frameworks build upon to synthesize full appearances from partial inputs.
- Paper: Panoptic Segmentation, Alexander Kirillov et al. (2018). It defines the modern benchmark formulations and metrics for holistic image segmentation against which amodal completion approaches are evaluated.
- Paper: Generative Semantic Segmentation, Jiaqi Chen et al. (2023). It introduces casting visual segmentation as an image-conditioned generative task rather than purely discriminative classification, a key perspective leveraged in generative amodal completion.
- Paper: Zero-1-to-3: Zero-shot One Image to 3D Object, Ruoshi Liu et al. (2023). It demonstrates how pre-trained 2D diffusion models encode rich geometric and appearance priors that can be fine-tuned to condition on partial visual observations for downstream reconstruction.
- Paper: Magic123: One Image to High-Quality 3D Object Generation Using Both 2D and 3D Diffusion Priors, Guocheng Qian et al. (2024). It leverages 2D and 3D diffusion priors to perform complete 3D object reconstruction from single images, advancing the downstream 3D generation capabilities demonstrated by amodal whole-synthesis front-ends.
- Paper: Free3D: Consistent Novel View Synthesis Without 3D Representation, Chuanxia Zheng et al. (2024). It extends single-image object synthesis to full 360-degree novel view consistency without 3D representations, building further on zero-shot 2D generative priors.
- Paper: SyncDreamer: Generating Multiview-consistent Images from a Single-view Image, Yuan Liu et al. (2024). It addresses multi-view visual consistency from single-view inputs using synchronized diffusion, advancing single-view visual completion into synchronized multi-angle synthesis.
- Paper: WildDet3D: Scaling Promptable 3D Detection in the Wild, Weikai Huang et al. (2026). It applies open-category promptable detection to uncontrolled in-the-wild environments, extending occlusion-robust spatial object understanding.
