pix2gestalt: Amodal Segmentation by Synthesizing Wholes

Ege OzgurogluRuoshi LiuDídac SurísDian ChenAchal DavePavel TokmakovCarl Vondrick

article2024CVPR73 citations

Presents a zero-shot framework that transfers pre-trained diffusion representations to reconstruct and segment entire occluded objects, outperforming fully supervised baselines and boosting downstream object recognition and 3D reconstruction.

Listen

Real-world computer vision applications in robotics, autonomous navigation, and augmented reality frequently fail when objects are partially hidden behind obstructions. While humans effortlessly perceive complete shapes and physical extents despite heavy occlusions, machine vision models have historically struggled with this task, known as amodal perception. Prior automated systems have been confined to narrow, closed-world categories or synthetic environments, severely limiting their practical deployment in complex real-world settings.

The article evaluates a new framework, named pix2gestalt, designed to perform zero-shot amodal completion and segmentation by synthesizing the whole, unoccluded appearance and shape of objects from partial visual inputs.

To achieve this, the researchers adapted pre-trained large-scale diffusion models, which implicitly encode rich representations of natural objects. The model was fine-tuned on an automatically curated dataset of 837,000 paired images. This dataset was constructed by identifying foreground objects in natural images using depth estimation and overlaying realistic synthetic occlusions to create ground-truth whole-part pairs. The system conditions on an input image and a prompt indicating the visible region, using iterative generation to synthesize complete images. This output serves as a direct input for downstream visual tasks, including segmentation, image classification, and three-dimensional reconstruction across standard benchmarks.

The evaluation yielded several key findings. First, pix2gestalt achieved state-of-the-art amodal segmentation accuracy, reaching an 82.87% mean intersection-over-union score on Amodal COCO without prior training on that dataset, outperforming fully supervised models. Second, incorporating the framework into object recognition pipelines substantially improved classification accuracy under occlusion; on the challenging Separated COCO benchmark, top-one accuracy rose from 26.04% using standard open-vocabulary classification to 31.15% with pix2gestalt, with even larger gains on standard occlusions where accuracy rose from 23.33% to 43.39%. Third, when serving as a drop-in module for three-dimensional reconstruction tools, it more than doubled volumetric accuracy and significantly reduced geometric error compared to standard baselines. Finally, the framework demonstrated robust zero-shot generalization across out-of-distribution inputs, including photographs, visual illusions, and artistic works.

These results demonstrate that generative image completion can act as a versatile foundation module to eliminate occlusion-related performance bottlenecks across diverse computer vision pipelines. By synthesizing complete objects rather than relying on narrow category-specific masks, organizations can improve system robustness in unstructured environments without incurring the high cost of collecting task-specific human annotations.

Organizations developing perception stacks should evaluate pix2gestalt as a front-end pre-processing module for existing recognition and three-dimensional reconstruction workflows. When handling ambiguous occlusions, engineering teams should deploy multi-sample generation with majority voting to select the most reliable completion. Further research should focus on integrating physical reasoning constraints into the generative process to prevent plausible-looking but physically impossible completions.

Confidence in these findings is supported by consistent quantitative improvements across multiple recognized benchmarks and tasks. However, decision-makers should note that the model relies on probabilistic sampling and can occasionally fail in scenarios requiring complex common-sense or physical reasoning, such as predicting correct motion directions or contact mechanics.

arXiv: 2401.14398
  • Paper: Globally and locally consistent image completion, SATOSHI IIZUKA et al. (2017). It provides foundational principles for context-guided image completion and visual inpainting that directly motivate generative amodal completion.
  • Paper: Image-to-Image Translation with Conditional Adversarial Networks, Phillip Isola et al. (2017). It establishes the foundational conditional image-to-image synthesis formulation that modern conditional generative frameworks build upon to synthesize full appearances from partial inputs.
  • Paper: Panoptic Segmentation, Alexander Kirillov et al. (2018). It defines the modern benchmark formulations and metrics for holistic image segmentation against which amodal completion approaches are evaluated.
  • Paper: Generative Semantic Segmentation, Jiaqi Chen et al. (2023). It introduces casting visual segmentation as an image-conditioned generative task rather than purely discriminative classification, a key perspective leveraged in generative amodal completion.
  • Paper: Zero-1-to-3: Zero-shot One Image to 3D Object, Ruoshi Liu et al. (2023). It demonstrates how pre-trained 2D diffusion models encode rich geometric and appearance priors that can be fine-tuned to condition on partial visual observations for downstream reconstruction.
Cover for pix2gestalt: Amodal Segmentation by Synthesizing Wholes

Abstract

We introduce pix2gestalt, a framework for zero-shot amodal segmentation, which learns to estimate the shape and appearance of whole objects that are only partially visible behind occlusions. By capitalizing on large-scale diffusion models and transferring their representations to this task, we learn a conditional diffusion model for reconstructing whole objects in challenging zero-shot cases, including examples that break natural and physical priors, such as art. As training data, we use a synthetically curated dataset containing occluded objects paired with their whole counterparts. Experiments show that our approach outperforms supervised baselines on established benchmarks. Our model can furthermore be used to significantly improve the performance of existing object recognition and 3D reconstruction methods in the presence of occlusions.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Amodal Completion and Segmentation
  • 2.2. Analysis by Synthesis
  • 2.3. Diffusion Models
  • 3. Amodal Completion via Generation
  • 3.1. Whole-Part Pairs
  • 3.2. Conditional Diffusion
  • 3.3. Amodal Base Representations
  • 4. Experiments
  • 4.1. Amodal Segmentation
  • 4.2. Occluded Object Recognition
  • 4.3. Amodal 3D Reconstruction
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — pix2gestalt Conditional Latent Diffusion Architecture

    model/method

    pix2gestalt formulates amodal completion as predicting a complete RGB image x^p\hat{x}_p of an occluded object given an input image xx and a prompt pp (such as a point prompt or a binary mask specifying the visible, modal region of the object). Built upon a pre-trained latent diffusion model (Stable Diffusion), the denoising U-Net ϵθ\epsilon_\theta incorporates conditioning via two separate streams:

    1. Semantic Conditioning Stream: The prompt-specified target in the input image xx is encoded using a CLIP vision encoder to produce semantic embedding vector C(x)\mathcal{C}(x), which is injected into the diffusion U-Net through cross-attention layers.
    2. Spatial/Visual Detail Stream: A Variational Autoencoder (VAE) encoder E\mathcal{E} encodes the input RGB image into latent feature map E(x)\mathcal{E}(x) and encodes the modal prompt into E(p)\mathcal{E}(p). These latent feature maps are concatenated along the channel dimension with the noised latent representation ztz_t of the target object.

    Classifier-Free Guidance (CFG) is incorporated during training by randomly replacing conditional inputs with null vectors, allowing the inference process to scale the influence of the conditioning signals during iterative denoising.

  2. Knowl 2 — pix2gestalt Diffusion Training Objective

    equation

    pix2gestalt fine-tunes a conditional latent diffusion model ϵθ\epsilon_\theta to reconstruct the unoccluded whole object x^p\hat{x}_p from an occluded image xx and modal prompt pp via the mean squared error loss:

    min⁡θ  Ez0∼E(x^p),t∼[0,1000),ϵ∼N(0,I)[∥ϵ−ϵθ(zt,E(x),t,E(p),C(x))∥22]\min _{\theta }\; \mathbb {E}_{z_0 \sim \mathcal {E}(\hat{x}_p), t \sim [0, 1000), \epsilon \sim \mathcal {N}(0, \mathbf{I})} \left [ \left\|\epsilon - \epsilon _{\theta }(z_t, \mathcal {E}(x), t, \mathcal {E}(p), \mathcal {C}(x))\right\|_2^2 \right ]

    where:

    • x^p\hat{x}_p is the ground-truth RGB image of the whole, unoccluded object.
    • z0=E(x^p)z_0 = \mathcal{E}(\hat{x}_p) is the VAE latent representation of the whole target object produced by a pre-trained encoder E\mathcal{E}.
    • t∈[0,1000)t \in [0, 1000) is the diffusion time step.
    • ϵ∼N(0,I)\epsilon \sim \mathcal{N}(0, \mathbf{I}) is standard Gaussian noise added to z0z_0 to form the noised latent zt=αˉtz0+1−αˉtϵz_t = \sqrt{\bar{\alpha}_t} z_0 + \sqrt{1 - \bar{\alpha}_t} \epsilon.
    • E(x)\mathcal{E}(x) is the VAE latent encoding of the occluded input image xx.
    • E(p)\mathcal{E}(p) is the VAE latent encoding of the prompt mask pp.
    • C(x)\mathcal{C}(x) is the CLIP embedding of the partially visible object in xx, provided via cross-attention.
    • ϵθ\epsilon_\theta is the parameterized denoising network that predicts the added noise ϵ\epsilon.
  3. Knowl 3 — Depth-Filtered Synthetic Occlusion Dataset Generation

    algorithm

    To train pix2gestalt without human-annotated whole-object counterparts, an automated pipeline constructs paired training samples (x,x^p,p)(x, \hat{x}_p, p) from natural images:

    Input: Unlabeled natural image dataset SA-1B, pre-trained Segment Anything (SAM), pre-trained MiDaS relative depth estimator
    Output: Paired dataset D = {(x, x_hat_p, p)} of 837K synthetic occlusion instances
    for each image I in SA-1B do
        Extract candidate object masks {M_i} using SAM on I
        Estimate relative monocular depth map D = MiDaS(I)
        for each candidate mask M_i do
            Compare estimated depth values inside M_i against neighboring background regions around the boundary of M_i
            if candidate M_i is closer to the camera than its adjacent neighbors then
                Designate object in M_i as a whole, unoccluded target x_hat_p
                Sample an occluder patch O from the dataset
                Superimpose O over x_hat_p to generate occluded image x
                Set prompt p as the remaining visible modal mask of the target object
                Add (x, x_hat_p, p) to D
            end if
        end for
    end for
    return D

    The depth-comparison heuristic ensures that synthetic occluders are only superimposed onto objects that are fully visible in the original image, preventing the model from learning to complete partially truncated objects.

  4. Knowl 4 — Amodal Base Representations for Downstream Vision Tasks

    model/method

    Because pix2gestalt outputs a complete RGB image x^p\hat{x}_p of an occluded object, it provides a unified amodal representation usable by standard off-the-shelf vision models:

    • Amodal Segmentation: An occluded image is completed into x^p\hat{x}_p using pix2gestalt, and the output is thresholded to produce an amodal segmentation mask. To handle under-constrained occlusion ambiguity, multiple independent completions are generated from the diffusion model and combined using majority voting over the resulting masks.
    • Occluded Object Recognition: Given an occluded object indicated by a bounding box or mask pp, pix2gestalt generates the whole object image x^p\hat{x}_p, which is then classified directly with an open-vocabulary vision-language model (CLIP).
    • Amodal 3D Reconstruction and Novel-View Synthesis: The synthesized whole image x^p\hat{x}_p and its derived amodal mask are fed into single-view 3D diffusion models (such as SyncDreamer) and neural implicit surface optimizers (NeuS/NeRF) to reconstruct textured 3D meshes without occlusion artifacts.
  5. Knowl 5 — Zero-Shot Amodal Segmentation Performance on COCO-A and BSDS-A

    data/table

    Zero-shot amodal segmentation accuracy evaluated on Amodal COCO (COCO-A, 13,000 objects in 2,500 images) and Amodal Berkeley Segmentation (BSDS-A, 650 objects in 200 images), measured by mean Intersection-over-Union (mIoU, %):

    Method Zero-shot COCO-A mIoU (%) ↑\uparrow BSDS-A mIoU (%) ↑\uparrow
    PCNet No 81.35 –
    PCNet-Sup No 82.53 –
    Segment Anything (SAM) Yes 67.21 65.25
    SD-XL Inpainting Yes 76.52 74.19
    pix2gestalt (Ours) Yes 82.87 80.76
    pix2gestalt: Best of 3 (Ours) Yes 87.10 85.68

    Without training on COCO-A, zero-shot pix2gestalt (82.87% mIoU) outperforms PCNet-Sup (82.53% mIoU), which was trained with ground-truth supervised amodal masks on COCO-A. Sampling 3 completions per instance and selecting the top-performing sample ("Best of 3") increases mIoU to 87.10% on COCO-A and 85.68% on BSDS-A, illustrating that multi-sample diffusion effectively captures multimodal uncertainty behind occlusions.

  6. Knowl 6 — Zero-Shot Occluded Object Recognition on Occluded and Separated COCO

    data/table

    Evaluation of open-vocabulary classification accuracy using CLIP across all 80 categories of the Occluded COCO and Separated COCO benchmarks (where the occluder splits the visible object region into multiple disjoint fragments):

    Method Top 1 Acc. (%) ↑\uparrow Top 3 Acc. (%) ↑\uparrow
    Occluded Separated Occluded Separated
    CLIP 23.33 26.04 43.84 43.19
    CLIP + Red Circle Prompt 23.46 25.64 43.86 43.24
    Visible Object Only + CLIP 34.00 21.10 49.26 34.70
    pix2gestalt + CLIP (Ours) 43.39 31.15 58.97 45.77

    Directly masking out the background to show only visible object fragments ("Visible Object Only + CLIP") improves Top-1 accuracy on simple occlusions (34.00% vs. 23.33%) but degrades performance on Separated COCO (21.10% vs. 26.04%). Reconstructing the whole object using pix2gestalt prior to classification achieves 43.39% Top-1 accuracy on Occluded COCO and 31.15% on Separated COCO.

  7. Knowl 7 — Single-View Amodal Novel-View Synthesis and 3D Reconstruction on GSO

    data/table

    Quantitative results for single-view 3D reconstruction and novel-view synthesis on 30 objects (60 rendered synthetic occlusion scenes) from Google Scanned Objects (GSO). Metrics evaluated include Chamfer Distance (CD, ↓\downarrow), Volumetric IoU (↑\uparrow), LPIPS (↓\downarrow), PSNR (↑\uparrow), and SSIM (↑\uparrow):

    Method CD ↓\downarrow Vol. IoU ↑\uparrow LPIPS ↓\downarrow PSNR ↑\uparrow SSIM ↑\uparrow
    SyncDreamer 0.1090 0.1920 0.3560 10.350 0.6530
    SAM Mask + SyncDreamer 0.1160 0.0938 0.3210 10.830 0.6950
    pix2gestalt (SAM Mask) + SyncDreamer 0.0760 0.3210 0.2880 12.150 0.6920
    GT Mask + SyncDreamer 0.1084 0.1027 0.2905 12.561 0.7322
    pix2gestalt (GT Mask) + SyncDreamer 0.0681 0.3639 0.2631 14.657 0.7328

    Using pix2gestalt as a drop-in front-end before SyncDreamer with an automated SAM prompt improves Volumetric IoU from 0.0938 to 0.3210 and reduces Chamfer Distance from 0.1160 to 0.0760. For novel-view synthesis, pix2gestalt improves PSNR from 10.83 dB to 12.15 dB and reduces perceptual error (LPIPS) from 0.3210 to 0.2880.

  8. Knowl 8 — Physical and Common-Sense Reasoning Limitations in pix2gestalt

    limitation

    While pix2gestalt accurately generates plausible visual texture and geometry for partially hidden objects, it exhibits systematic failure modes in settings requiring complex physical or commonsense scene reasoning:

    • Physical Support Constraints: When an occluder physically supports or holds the target object (such as a human hand holding a box of donuts), removing the occluder during completion can produce floating objects rather than inferring appropriate alternative contact or support points.
    • Semantic Orientation Ambiguity: For objects whose orientation is not completely constrained by visible cues (such as an occluded car), the model may generate shapes pointing in directions that contradict the logical context of the scene.

Coverage note — No substantial contributed material was omitted.

References

  1. 1.Rameen Abdal, Yipeng Qin, and Peter Wonka. Image2StyleGAN: How to embed images into the stylegan latent space? In ICCV, 2019. 3
  2. 2.Tomer Amit, Tal Shaharbany, Eliya Nachmani, and Lior Wolf. Segdiff: Image segmentation with diffusion probabilistic models. arXiv preprint arXiv:2112.00390, 2021. 3
  3. 3.Dmitry Baranchuk, Ivan Rubachev, Andrey Voynov, Valentin Khrulkov, and Artem Babenko. Label-efficient semantic segmentation with diffusion models. arXiv preprint arXiv:2112.03126, 2021. 3
  4. 4.Reiner Birkl, Diana Wofk, and Matthias Müller. Midas v3.1 – a model zoo for robust monocular relative depth estimation. arXiv preprint arXiv:2307.14460, 2023. 4
  5. 5.Volker Blanz and Thomas Vetter. A morphable model for the synthesis of 3d faces. In Seminal Graphics Papers: Pushing the Boundaries, Volume 2, pages 157–164. 2023. 3
  6. 6.Tim Brooks, Aleksander Holynski, and Alexei A. Efros. Instructpix2pix: Learning to follow image editing instructions. In CVPR, 2023. 3, 4
  7. 7.Matt Deitke, Ruoshi Liu, Matthew Wallingford, Huong Ngo, Oscar Michel, Aditya Kusupati, Alan Fan, Christian Laforte, Vikram Voleti, Samir Yitzhak Gadre, et al. Objaverse-xl: A universe of 10m+ 3d objects. arXiv preprint arXiv:2307.05663, 2023. 3, 7
  8. 8.Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. NeurIPS, 2021. 3
  9. 9.Laura Downs, Anthony Francis, Nate Koenig, Brandon Kinman, Ryan Hickman, Krista Reymann, Thomas B. McHugh, and Vincent Vanhoucke. Google scanned objects: A high-quality dataset of 3D scanned household items. In ICRA, 2022. 7, 8
  10. 10.Kiana Ehsani, Roozbeh Mottaghi, and Ali Farhadi. Segan: Segmenting and generating the invisible. In CVPR, 2018. 1, 3
  11. 11.Rinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik, Amit H Bermano, Gal Chechik, and Daniel Cohen-Or. An image is worth one word: Personalizing text-to-image generation using textual inversion. arXiv preprint arXiv:2208.01618, 2022. 3
  12. 12.Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets. NeurIPS, 2014. 3
  13. 13.Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598, 2022. 3, 4
  14. 14.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. NeurIPS, 33, 2020. 1, 3
  15. 15.Cheng-Yen Hsieh, Tarasha Khurana, Achal Dave, and Deva Ramanan. Tracking any object amodally, 2023. 1, 3
  16. 16.Y.-T. Hu, H.-S. Chen, K. Hui, J.-B. Huang, and A. G. Schwing. SAIL-VOS: Semantic Amodal Instance Level Video Object Segmentation – A Synthetic Dataset and Baselines. In Proc. CVPR, 2019. 4
  17. 17.Abhishek Kar, Shubham Tulsiani, Joao Carreira, and Jitendra Malik. Amodal completion and size constancy in natural scenes. In ICCV, 2015. 1, 3
  18. 18.Lei Ke, Yu-Wing Tai, and Chi-Keung Tang. Deep occlusion-aware instance segmentation with overlapping bilayers. In CVPR, 2021. 1, 3
  19. 19.Diederik P Kingma and Max Welling. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, 2013. 3
  20. 20.Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollár, and Ross Girshick. Segment anything. In ICCV, 2023. 4, 5, 7
  21. 21.Huan Ling, David Acuna, Karsten Kreis, Seung Wook Kim, and Sanja Fidler. Variational amodal object completion. NeurIPS, 2020. 1, 3
  22. 22.Ruoshi Liu and Carl Vondrick. Humans as light bulbs: 3d human reconstruction from thermal reflection. In CVPR, 2023. 3
  23. 23.Ruoshi Liu, Sachit Menon, Chengzhi Mao, Dennis Park, Simon Stent, and Carl Vondrick. Shadows shed light on 3d objects. arXiv preprint arXiv:2206.08990, 2022. 3
  24. 24.Ruoshi Liu, Chengzhi Mao, Purva Tendulkar, Hao Wang, and Carl Vondrick. Landscape learning for neural network inversion. In ICCV, 2023. 3
  25. 25.Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov, Sergey Zakharov, and Carl Vondrick. Zero-1-to-3: Zero-shot one image to 3d object. In ICCV, 2023. 3, 4, 7
  26. 26.Yuan Liu, Cheng Lin, Zijiao Zeng, Xiaoxiao Long, Lingjie Liu, Taku Komura, and Wenping Wang. Syncdreamer: Learning to generate multiview-consistent images from a single-view image. arXiv preprint arXiv:2309.03453, 2023. 7, 8
  27. 27.Wufei Ma, Angtian Wang, Alan Yuille, and Adam Kortylewski. Robust category-level 6D pose estimation with coarse-to-fine rendering of neural features. In ECCV, 2022. 3
  28. 28.D. Martin, C. Fowlkes, D. Tal, and J. Malik. A database of human segmented natural images and its application to evaluating segmentation algorithms and measuring ecological statistics. In ICCV, 2001. 5, 7
  29. 29.Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. In ECCV, 2020. 7
  30. 30.Jean Piaget. The construction of reality in the child. Routledge, 2013. 1
  31. 31.Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. Sdxl: Improving latent diffusion models for high-resolution image synthesis, 2023. 5, 7
  32. 32.Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Mildenhall. Dreamfusion: Text-to-3d using 2d diffusion. arXiv preprint arXiv:2209.14988, 2022. 4
  33. 33.Lu Qi, Li Jiang, Shu Liu, Xiaoyong Shen, and Jiaya Jia. Amodal instance segmentation with KINS dataset. In CVPR, 2019. 1, 3, 4
  34. 34.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision, 2021. 7
  35. 35.N Dinesh Reddy, Robert Tamburo, and Srinivasa G Narasimhan. Walt: Watch and learn 2d amodal representation from time-lapse imagery. In CVPR, 2022. 1
  36. 36.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In CVPR, 2022. 3, 4
  37. 37.Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. In CVPR, 2023. 3
  38. 38.Kyle Sargent, Zizhang Li, Tanmay Shah, Charles Herrmann, Hong-Xing Yu, Yunzhi Zhang, Eric Ryan Chan, Dmitry Lagun, Li Fei-Fei, Deqing Sun, and Jiajun Wu. ZeroNVS: Zero-shot 360-degree view synthesis from a single real image. arXiv preprint arXiv:2310.17994, 2023. 7
  39. 39.Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al. Laion-5B: An open large-scale dataset for training next generation image-text models. NeurIPS, 2022. 3
  40. 40.Aleksandar Shtedritski, Christian Rupprecht, and Andrea Vedaldi. What does clip know about a red circle? visual prompt engineering for vlms. In ICCV, 2023. 7
  41. 41.Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502, 2020. 3
  42. 42.Zhuowen Tu, Xiangrong Chen, Alan L Yuille, and Song-Chun Zhu. Image parsing: Unifying segmentation, detection, and recognition. International Journal of computer vision, 63:113–140, 2005. 3
  43. 43.Peng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt, Taku Komura, and Wenping Wang. Neus: Learning neural implicit surfaces by volume rendering for multi-view reconstruction. arXiv preprint arXiv:2106.10689, 2021. 7
  44. 44.Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. Image quality assessment: from error visibility to structural similarity. IEEE Transactions on Image Processing, 13(4):600–612, 2004. 7, 8
  45. 45.Rundi Wu, Ruoshi Liu, Carl Vondrick, and Changxi Zheng. Sin3dm: Learning a diffusion model from a single 3d textured shape. arXiv preprint arXiv:2305.15399, 2023. 3
  46. 46.Jiarui Xu, Sifei Liu, Arash Vahdat, Wonmin Byeon, Xiaolong Wang, and Shalini De Mello. Open-Vocabulary Panoptic Segmentation with Text-to-Image Diffusion Models. arXiv preprint arXiv:2303.04803, 2023. 3
  47. 47.Katherine Xu, Lingzhi Zhang, and Jianbo Shi. Amodal completion via progressive mixed context diffusion, 2023. 3
  48. 48.Alan Yuille and Daniel Kersten. Vision as bayesian inference: analysis by synthesis? Trends in cognitive sciences, 10(7):301–308, 2006. 3
  49. 49.Guanqi Zhan, Weidi Xie, and Andrew Zisserman. A tri-layer plugin to improve occluded detection. BMVC, 2022. 6, 7
  50. 50.Guanqi Zhan, Chuanxia Zheng, Weidi Xie, and Andrew Zisserman. Amodal ground truth and completion in the wild. In arXiv, 2023. 3
  51. 51.Xiaohang Zhan, Xingang Pan, Bo Dai, Ziwei Liu, Dahua Lin, and Chen Change Loy. Self-supervised scene de-occlusion. In CVPR, 2020. 1, 3, 5, 7
  52. 52.Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In CVPR, 2018. 7, 8
  53. 53.Yi Zhang, Pengliang Ji, Angtian Wang, Jieru Mei, Adam Kortylewski, and Alan Yuille. 3D-Aware neural body fitting for occlusion robust 3d human pose estimation. In ICCV, 2023. 3
  54. 54.Jiapeng Zhu, Yujun Shen, Deli Zhao, and Bolei Zhou. In-domain GAN inversion for real image editing. In ECCV, 2020. 3
  55. 55.Yan Zhu, Yuandong Tian, Dimitris Metaxas, and Piotr Dollár. Semantic amodal segmentation. In CVPR, 2017. 1, 3, 4, 5, 7

Citation

MLA
Ozguroglu, E., et al. “Pix2gestalt: Amodal Segmentation by Synthesizing Wholes”. arXiv, 2024, http://arxiv.org/abs/2401.14398v1.
APA
Ozguroglu, E., Liu, R., Surís, D., Chen, D., Dave, A., Tokmakov, P., & Vondrick, C. (2024). pix2gestalt: Amodal Segmentation by Synthesizing Wholes. arXiv. http://arxiv.org/abs/2401.14398v1
Chicago
Ozguroglu, E., R. Liu, D. Surís, et al. 2024. “Pix2gestalt: Amodal Segmentation by Synthesizing Wholes”. arXiv. http://arxiv.org/abs/2401.14398v1.
Harvard
Ozguroglu, E. et al. (2024) “pix2gestalt: Amodal Segmentation by Synthesizing Wholes”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2401.14398v1.
Vancouver
1. Ozguroglu E, Liu R, Surís D, Chen D, Dave A, Tokmakov P, Vondrick C (2024) pix2gestalt: Amodal Segmentation by Synthesizing Wholes. arXiv

BibTeX

@article{ozguroglu2024pix2gestalt,
  title = {pix2gestalt: Amodal Segmentation by Synthesizing Wholes},
  author = {Ozguroglu, Ege and Liu, Ruoshi and Surís, Dídac and Chen, Dian and Dave, Achal and Tokmakov, Pavel and Vondrick, Carl},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2401.14398v1},
  eprint = {2401.14398}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE