ACSeg: Adaptive Conceptualization for Unsupervised Semantic Segmentation

Kehan LiZhennan WangZesen ChengRunyi YuYian ZhaoGuoli SongChang LiuLi YuanJie Chen

article2023CVPR52 citations

Proposes an unsupervised semantic segmentation framework that dynamically maps learnable prototypes into image-specific semantic concepts using attention mechanisms and a modularity loss, overcoming over- and under-clustering issues without requiring manual annotations.

Listen

Semantic segmentation—the automated process of identifying and categorizing visual objects at the pixel level—is critical for modern computer vision applications such as autonomous navigation and medical imaging. However, conventional methods demand vast quantities of manually labeled data, which are slow and expensive to produce. While recent self-supervised vision models capture rich semantic information without human annotations, standard segmentation methods frequently suffer from over-clustering or under-clustering because they fail to adapt to varying scene complexity across different images.

The article introduces and evaluates Adaptive Conceptualization for Unsupervised Semantic Segmentation (ACSeg), a novel framework designed to segment images accurately without human labels. Its primary objective is to adaptively discover underlying visual concepts within an image's pixel representation space and classify them in an entirely unsupervised or text-guided manner.

The evaluated approach uses a self-supervised Vision Transformer to extract pixel-level features and introduces an Adaptive Concept Generator. This generator dynamically maps learnable initial prototypes into image-specific concepts via attention mechanisms. The system optimizes these concepts end-to-end without ground-truth labels using a newly formulated modularity loss, which estimates whether pixel pairs belong together based on graph affinity. Discovered concept regions are subsequently categorized using region clustering or zero-shot text matching via vision-language models. The authors validated this framework across standard benchmark datasets, including PASCAL VOC 2012 and COCO-Stuff.

The key findings demonstrate significant performance and efficiency gains over existing approaches. First, ACSeg achieved state-of-the-art unsupervised semantic segmentation accuracy on PASCAL VOC 2012 with a mean intersection over union of 47.1%, outperforming prior techniques without requiring network retraining or post-processing. Second, it delivered top results on the challenging COCO-Stuff benchmark with 16.4% mean intersection over union. Third, when integrated with vision-language models for text-supervised segmentation, ACSeg outperformed existing zero-shot baselines, reaching 53.9% on PASCAL VOC and 28.1% on COCO-Stuff. Finally, the framework processed 149.2 images per second during clustering evaluation, operating roughly 10 to 60 times faster than conventional clustering baselines while requiring only tens of minutes to train.

These results demonstrate that high-performance visual segmentation can be achieved efficiently by leveraging existing self-supervised models rather than training large segmentation networks from scratch. Eliminating manual annotation requirements and lengthy retraining cycles lowers development costs, shortens deployment timelines, and reduces computational overhead for complex computer vision systems.

Organizations developing computer vision systems should consider adopting adaptive conceptualization pipelines to minimize data annotation expenses and accelerate deployment. Where text labels are available, pairing the framework with vision-language models offers a robust option for zero-shot categorization. Stakeholders planning deployment should conduct domain-specific pilot testing, particularly to calibrate the initial prototype count, as setting it too low limits segmentation detail while setting it too high can introduce unnecessary visual fragmentation.

Cover for ACSeg: Adaptive Conceptualization for Unsupervised Semantic Segmentation

Abstract

Recently, self-supervised large-scale visual pre-training models have shown great promise in representing pixel-level semantic relationships, significantly promoting the development of unsupervised dense prediction tasks, e.g., unsupervised semantic segmentation (USS). The extracted relationship among pixel-level representations typically contains rich class-aware information that semantically identical pixel embeddings in the representation space gather together to form sophisticated concepts. However, leveraging the learned models to ascertain semantically consistent pixel groups or regions in the image is non-trivial since over/ under-clustering overwhelms the conceptualization procedure under various semantic distributions of different images. In this work, we investigate the pixel-level semantic aggregation in self-supervised ViT pre-trained models as image Segmentation and propose the Adaptive Conceptualization approach for USS, termed ACSeg. Concretely, we explicitly encode concepts into learnable prototypes and design the Adaptive Concept Generator (ACG), which adaptively maps these prototypes to informative concepts for each image. Meanwhile, considering the scene complexity of different images, we propose the modularity loss to optimize ACG independent of the concept number based on estimating the intensity of pixel pairs belonging to the same concept. Finally, we turn the USS task into classifying the discovered concepts in an unsupervised manner. Extensive experiments with state-of-the-art results demonstrate the effectiveness of the proposed ACSeg.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 3. The Proposed ACSeg
  • 3.1. Overall Approach
  • 3.2. Adaptive Concept Generator
  • 3.3. Pixel Assignment
  • 3.4. Modularity Loss
  • 3.5. Concept Classifier
  • 4. Experiment
  • 4.1. Implementation Details
  • 4.2. Qualitative Results
  • 4.3. Quantitative Results
  • 4.4. Ablation Study
  • 5. Conclusion and Discussion
  • References

Knowls

  1. Knowl 1 — Adaptive Concept Generator Architecture

    model/method

    The Adaptive Concept Generator (ACG) dynamically updates a set of learnable initial prototypes C0∈Rk×dC^0 \in \mathbb{R}^{k \times d} into image-specific concept embeddings by interacting with the spatial token representations X∈Rn×dX \in \mathbb{R}^{n \times d} extracted from a frozen self-supervised Vision Transformer (ViT), where nn is the number of patch tokens, dd is the feature dimension, and kk is the number of prototypes.

    The ACG consists of NN iterative update blocks (with N=6N = 6). Each block executes a cross-attention step, a self-attention step, and a feed-forward network (FFN), with multi-head attention, layer normalization, and residual connections:

    1. Cross-Attention Update (mapping image features to prototypes): Cˉl=Softmax(Cl−1Wq(XWk)Td)(XWv)\bar{C}^l = \mathrm{Softmax}\left(\frac{C^{l-1} W_q (X W_k)^T}{\sqrt{d}}\right)(X W_v) Cl=Cl−1+CˉlWoC^l = C^{l-1} + \bar{C}^l W_o where Wq,Wk,Wv,Wo∈Rd×dW_q, W_k, W_v, W_o \in \mathbb{R}^{d \times d} are learnable projection matrices and Cl−1C^{l-1} is the prototype matrix after update step l−1l-1.

    2. Self-Attention Update (modeling mutual relations between concept prototypes): Cˉl=Softmax(Cl−1Wq′(Cl−1Wk′)Td)(Cl−1Wv′)\bar{C}^l = \mathrm{Softmax}\left(\frac{C^{l-1} W_q' (C^{l-1} W_k')^T}{\sqrt{d}}\right)(C^{l-1} W_v') Cl=Cl−1+CˉlWo′C^l = C^{l-1} + \bar{C}^l W_o' where Wq′,Wk′,Wv′,Wo′∈Rd×dW_q', W_k', W_v', W_o' \in \mathbb{R}^{d \times d} are linear projections.

    After NN update steps, the resulting vectors CN=[c1,c2,…,ck]T∈Rk×dC^N = [c_1, c_2, \dots, c_k]^T \in \mathbb{R}^{k \times d} serve as the final concept representations for the input image.

  2. Knowl 2 — Modularity Loss for Unsupervised Prototype Optimization

    model/method

    The Adaptive Concept Generator is optimized in an unsupervised manner using a graph modularity loss that evaluates whether pairs of pixels belonging to the same cluster exhibit higher affinity than expected by random chance.

    For an image with pixel embeddings xi∈Rdx_i \in \mathbb{R}^d (i∈{1,…,n}i \in \{1, \dots, n\}), an undirected affinity graph is constructed with edge weights defined by non-negative cosine similarities: Ai,j=max⁡(0,cos⁡⟨xi,xj⟩)A_{i,j} = \max(0, \cos\langle x_i, x_j \rangle) The degree of node ii is ki=∑j=1nAi,jk_i = \sum_{j=1}^n A_{i,j}, and the total graph volume is 2m=∑i=1n∑j=1nAi,j2m = \sum_{i=1}^n \sum_{j=1}^n A_{i,j}. Under a random graph null model with matching node degrees, the expected edge weight between ii and jj is kikj2m\frac{k_i k_j}{2m}. The modularity intensity wijw_{ij} is: wij=Ai,j−kikj2mw_{ij} = A_{i,j} - \frac{k_i k_j}{2m}

    To couple this with the updated prototypes c1,…,ckc_1, \dots, c_k, a soft assignment score is computed as Sˉi,c=max⁡(0,cos⁡⟨xi,c⟩)\bar{S}_{i,c} = \max(0, \cos\langle x_i, c \rangle). The degree of co-membership δ(i,j)\delta(i, j) of pixels ii and jj assigned to the same prototype is: δ(i,j)=max⁡c∈{c1,…,ck}Sˉi,c⋅Sˉj,c\delta(i, j) = \max_{c \in \{c_1, \dots, c_k\}} \bar{S}_{i,c} \cdot \bar{S}_{j,c}

    The modularity loss L\mathcal{L} across all pixel pairs in the image is formulated as: L=−12m∑i=1n∑j=1nwijδ(i,j)=−12m∑i=1n∑j=1n(Ai,j−kikj2m)δ(i,j)\mathcal{L} = -\frac{1}{2m} \sum_{i=1}^n \sum_{j=1}^n w_{ij} \delta(i,j) = -\frac{1}{2m} \sum_{i=1}^n \sum_{j=1}^n \left(A_{i,j} - \frac{k_i k_j}{2m}\right) \delta(i,j)

    This loss is hyperparameter-free and its minimum does not depend on a predefined cluster count, allowing the network to adapt the active number of concepts to each image's semantic complexity.

  3. Knowl 3 — Pixel Assignment and Dynamic Concept Pruning

    model/method

    In the ACSeg framework, image regions are formed by mapping pixels to the nearest concept in the embedding space.

    1. Soft Assignment: For pixel embedding xi∈Rdx_i \in \mathbb{R}^d and concept vector cj∈Rdc_j \in \mathbb{R}^d (j∈{1,…,k}j \in \{1, \dots, k\}), the assignment matrix S∈Rn×kS \in \mathbb{R}^{n \times k} is defined as: Si,j=cos⁡⟨xi,cj⟩S_{i,j} = \cos\langle x_i, c_j \rangle During training, SS remains continuous and differentiable for computing the modularity loss.

    2. Inference Assignment: During inference, the soft assignment matrix SS is bilinearly upsampled to the original image resolution. The discrete concept assignment aia_i for pixel ii is determined via argmax: ai=arg⁡max⁡j∈{1,…,k}cos⁡⟨xi,cj⟩a_i = \arg\max_{j \in \{1, \dots, k\}} \cos\langle x_i, c_j \rangle

    3. Dynamic Pruning: The argmax operation partitions an image into m≤km \le k regions. Any prototype cjc_j to which no pixels are assigned (i.e., {i∣ai=j}=∅\{i \mid a_i = j\} = \emptyset) is pruned. This mechanism enables ACSeg to output variable concept counts across images of varying scene complexity while starting from a fixed initial prototype count kk.

  4. Knowl 4 — Attention-Based Background Region Separation

    model/method

    To identify background regions without human labels, ACSeg computes a foreground saliency score for each discovered concept region using attention maps from the self-supervised Vision Transformer backbone (DINO ViT).

    For every pixel, its attention value is taken as the minimum across all self-attention heads in the final layer of the self-supervised ViT, under the assumption that a true foreground pixel will be salient in at least one attention head. A foreground score is computed for each concept region by summing the pixel-level attention values across all pixels assigned to that region.

    The resulting regional foreground scores within an image are clustered into two groups using 1D clustering. The cluster possessing the lower mean foreground score is designated as the background category, and its constituent regions are assigned to the background class without requiring external models or annotations.

  5. Knowl 5 — Foreground Concept Classification via Feature Clustering and Zero-Shot Text Alignment

    model/method

    Once background regions are removed, ACSeg classifies foreground concept regions into semantic classes using one of two non-parametric approaches:

    1. Region-Level Visual Representation Clustering / Retrieval: Each foreground concept region is cropped by its bounding box and resized to 224×224224 \times 224 pixels. Non-region pixels are masked out. A discriminative region-level representation is extracted using the frozen self-supervised ViT. For fully unsupervised semantic segmentation, kk-means clustering is applied to all region embeddings across the dataset, and cluster identities are matched to ground-truth semantic classes using the Hungarian algorithm for evaluation. For retrieval evaluation, a kk-NN classifier assigns class labels to query regions based on nearest neighbors in the training set.

    2. Text-Prompt Zero-Shot Classification: Pixel-level visual features are extracted from the image using the visual encoder of CLIP modified via MaskCLIP. A single visual embedding for each concept is formed by average-pooling the MaskCLIP pixel features within the concept region mask. Semantic classification is performed by computing cosine similarity between this pooled visual embedding and CLIP text embeddings of predefined category prompt texts, assigning the concept to the highest-scoring category.

  6. Knowl 6 — Unsupervised Semantic Segmentation Benchmark Results on PASCAL VOC 2012 and COCO-Stuff-27

    data/table

    ACSeg was evaluated against existing unsupervised semantic segmentation methods on the PASCAL VOC 2012 and COCO-Stuff-27 benchmarks using mean Intersection over Union (mIoU). Foreground concept regions were clustered using kk-means (averaged over 10 runs with mean ±\pm standard deviation reported) and matched to ground truth via the Hungarian algorithm.

    Method (PASCAL VOC) mIoU (%) Method (COCO-Stuff-27) mIoU (%)
    IIC 9.8 MoCo v2 4.4
    MaskContrast 35.0 IIC 6.7
    DSM (with re-training) 37.2 ±\pm 3.8 ImageNet 8.9
    Leopart 41.7 DINO 9.6
    TransFGU 37.2 Modified DC 9.8
    MaskDistill 42.0 PiCIE 13.8
    MaskDistill (with re-training) 45.8 PiCIE+H 14.4
    ACSeg (Ours) 47.1 ±\pm 2.4 ACSeg (Ours) 16.4 ±\pm 0.9

    ACSeg outperforms prior approaches on both datasets without requiring pixel-level training or end-to-end segmentation network re-training.

  7. Knowl 7 — Performance of Zero-Shot Semantic Segmentation with Text Supervision

    data/table

    When combined with pre-trained vision-language models (CLIP) for zero-shot text-based classification of discovered concept regions, ACSeg was compared against vision-language segmentation frameworks on PASCAL VOC and COCO-Stuff-27.

    Method PASCAL VOC mIoU (%) COCO-Stuff mIoU (%)
    MaskCLIP – 19.6
    GroupViT 51.2 20.3
    ReCo – 26.3
    ACSeg (Ours) 53.9 28.1

    By leveraging the precise region boundaries discovered by the ACG, ACSeg achieves higher segmentation performance than methods relying directly on dense CLIP patch features or text-supervised grouping architectures.

  8. Knowl 8 — Segmentation Quality and Speed: ACG vs. Traditional Clustering Algorithms

    data/table

    The effectiveness and efficiency of the Adaptive Concept Generator were compared on PASCAL VOC against traditional clustering algorithms applied to the pixel representations extracted from a self-supervised ViT-Small backbone.

    Clustering Method mIoU (%) Speed (images/sec)
    K-Means 28.6 2.4
    Spectral Clustering 28.3 3.4
    Affinity Propagation 11.0 6.8
    Agglomerative Clustering 13.9 15.8
    ACG (Ours) 47.1 149.2

    ACG achieves an 18.5 mIoU improvement over the best traditional clustering baseline (K-Means) while executing over 60 times faster due to GPU-accelerated parallel neural attention operations.

  9. Knowl 9 — Sensitivity to Initial Prototype Number

    data/table

    The impact of the initial prototype count kk on unsupervised semantic segmentation was evaluated on PASCAL VOC using both kk-means clustering mIoU and kk-NN retrieval mIoU (K=1K=1).

    Number of Prototypes (kk) Clustering mIoU (%) Retrieval mIoU (%)
    2 35.6 43.7
    5 47.1 57.8
    7 46.9 56.9
    10 42.1 54.4
    15 34.5 51.1

    Setting k=2k=2 causes the ACG to degenerate into basic foreground-background segmentation, harming multi-class segmentation. Increasing kk beyond 5–7 causes mild over-clustering of fine object parts, slightly reducing mIoU. The method demonstrates stability within the range of 5 to 7 prototypes.

  10. Knowl 10 — Nearest-Neighbor Concept Retrieval Performance

    data/table

    The quality of region-level representations and concept localization was assessed via kk-NN retrieval (K=1K=1 and K=5K=5) on PASCAL VOC and COCO-Stuff-27. For each validation concept, the predicted category is matched to the category of the nearest training concept, whose label is defined by the ground-truth segment with maximum spatial overlap.

    Dataset Method K=1 mIoU (%) K=5 mIoU (%)
    VOC MaskContrast 43.3 –
    VOC DSM 32.1 31.9
    VOC K-means 45.1 49.1
    VOC Spectral Clustering 43.0 47.3
    VOC ACSeg (Ours) 57.8 61.0
    COCO-Stuff-27 K-means 29.9 33.1
    COCO-Stuff-27 Spectral Clustering 28.5 31.3
    COCO-Stuff-27 ACSeg (Ours) 30.4 34.0

    ACSeg outperforms prior self-supervised region extraction methods and standard clustering algorithms on both datasets in retrieval mIoU.

Coverage note — None was omitted; all contributed models, loss formulations, assignment mechanisms, classification pipelines, and experimental results have been covered.

References

  1. 1.Jiwoon Ahn and Suha Kwak. Learning pixel-level semantic affinity with image-level supervision for weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4981–4990, 2018. 1
  2. 2.Adam Bielski and Paolo Favaro. Move: Unsupervised movable object segmentation and detection. arXiv preprint arXiv:2210.07920, 2022. 2
  3. 3.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in Neural Information Processing Systems, 33:1877–1901, 2020. 2
  4. 4.Mathilde Caron, Hugo Touvron, Ishan Misra, Herv´e J´egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9650–9660, 2021. 2, 6, 7
  5. 5.Xinlei Chen, Haoqi Fan, Ross Girshick, and Kaiming He. Improved baselines with momentum contrastive learning. arXiv preprint arXiv:2003.04297, 2020. 7
  6. 6.Jang Hyun Cho, Utkarsh Mall, Kavita Bala, and Bharath Hariharan. Picie: Unsupervised semantic segmentation using invariance and equivariance in clustering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16794–16804, 2021. 2, 7
  7. 7.Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. The cityscapes dataset for semantic urban scene understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3213–3223, 2016. 1
  8. 8.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018. 2
  9. 9.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations, 2020. 2, 6
  10. 10.Delbert Dueck. Affinity propagation: clustering data by passing messages. University of Toronto Toronto, ON, Canada, 2009. 8
  11. 11.Mark Everingham and John Winn. The pascal visual object classes challenge 2012 (voc2012) development kit. Pattern Anal. Stat. Model. Comput. Learn., Tech. Rep, 2007:1–45, 2012. 1, 2, 6
  12. 12.Deng-Ping Fan, Ge-Peng Ji, Tao Zhou, Geng Chen, Huazhu Fu, Jianbing Shen, and Ling Shao. Pranet: Parallel reverse attention network for polyp segmentation. In International Conference on Medical Image Computing and Computer-assisted Intervention, pages 263–273. Springer, 2020. 1
  13. 13.Santo Fortunato. Community detection in graphs. Physics Reports, 486(3-5):75–174, 2010. 5
  14. 14.Andreas Geiger, Philip Lenz, Christoph Stiller, and Raquel Urtasun. Vision meets robotics: The kitti dataset. The International Journal of Robotics Research, 32(11):1231–1237, 2013. 1
  15. 15.Mark Hamilton, Zhoutong Zhang, Bharath Hariharan, Noah Snavely, and William T Freeman. Unsupervised semantic segmentation by distilling feature correspondences. In International Conference on Learning Representations, 2021. 2, 3
  16. 16.John A Hartigan. Clustering algorithms. John Wiley & Sons, Inc., 1975. 8
  17. 17.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 770–778, 2016. 7
  18. 18.Haiyang Huang, Zhi Chen, and Cynthia Rudin. Segdiscover: Visual concept discovery via unsupervised semantic segmentation. arXiv preprint arXiv:2204.10926, 2022. 2, 3
  19. 19.Jyh-Jing Hwang, Stella X Yu, Jianbo Shi, Maxwell D Collins, Tien-Ju Yang, Xiao Zhang, and Liang-Chieh Chen. Segsort: Segmentation by discriminative sorting of segments. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7334–7344, 2019. 1, 2
  20. 20.Xu Ji, Joao F Henriques, and Andrea Vedaldi. Invariant information clustering for unsupervised image classification and segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9865–9874, 2019. 1, 2, 6, 7
  21. 21.Peng Jin, JinFa Huang, Fenglin Liu, Xian Wu, Shen Ge, Guoli Song, David A. Clifton, and Jie Chen. Expectation-maximization contrastive learning for compact video-and-language representations. In Thirty-Sixth Conference on Neural Information Processing Systems, 2022. 2
  22. 22.Tsung-Wei Ke, Jyh-Jing Hwang, Yunhui Guo, Xudong Wang, and Stella X Yu. Unsupervised hierarchical semantic segmentation with multiview cosegmentation and clustering transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2571–2581, 2022. 2
  23. 23.Harold W Kuhn. The hungarian method for the assignment problem. Naval Research Logistics Quarterly, 2(1-2):83–97, 1955. 7
  24. 24.Neeraj Kumar, Ruchika Verma, Sanuj Sharma, Surabhi Bhargava, Abhishek Vahadane, and Amit Sethi. A dataset and a technique for generalized nuclear segmentation for computational pathology. IEEE Transactions on Medical Imaging, 36(7):1550–1560, 2017. 1
  25. 25.Hao Li, Jinfa Huang, Peng Jin, Guoli Song, Qi Wu, and Jie Chen. Toward 3d spatial reasoning for human-like text-based visual question answering. arXiv preprint arXiv:2209.10326, 2022. 2
  26. 26.Hao Li, Xu Li, Belhal Karimi, Jie Chen, and Mingming Sun. Joint learning of object graph and relation graph for visual question answering. arXiv preprint arXiv:2205.04188, 2022. 2
  27. 27.Di Lin, Jifeng Dai, Jiaya Jia, Kaiming He, and Jian Sun. Scribblesup: Scribble-supervised convolutional networks for semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3159–3167, 2016. 1
  28. 28.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll´ar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In European Conference on Computer Vision, pages 740–755. Springer, 2014. 1, 2, 6
  29. 29.Jonathan Long, Evan Shelhamer, and Trevor Darrell. Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3431–3440, 2015. 1
  30. 30.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In International Conference on Learning Representations, 2018. 6
  31. 31.Luke Melas-Kyriazi, Christian Rupprecht, Iro Laina, and Andrea Vedaldi. Finding an unsupervised image segmenter in each of your deep generative models. In International Conference on Learning Representations, 2021. 3
  32. 32.Luke Melas-Kyriazi, Christian Rupprecht, Iro Laina, and Andrea Vedaldi. Deep spectral methods: A surprisingly strong baseline for unsupervised semantic segmentation and localization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8364–8375, 2022. 2, 3, 7, 8
  33. 33.Fionn Murtagh and Pedro Contreras. Algorithms for hierarchical clustering: an overview. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 2(1):86–97, 2012. 8
  34. 34.Mark EJ Newman and Michelle Girvan. Finding and evaluating community structure in networks. Physical Review E, 69(2):026113, 2004. 2, 3, 4, 5
  35. 35.Youngmin Oh, Beomjun Kim, and Bumsub Ham. Background-aware pooling and noise-aware loss for weakly-supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6913–6922, 2021. 1
  36. 36.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pages 8748–8763. PMLR, 2021. 2, 3, 6, 8
  37. 37.Wei Shen, Zelin Peng, Xuehui Wang, Huayu Wang, Jiazhong Cen, Dongsheng Jiang, Lingxi Xie, Xiaokang Yang, and Qi Tian. A survey on label-efficient deep segmentation: Bridging the gap between weak supervision and dense prediction. arXiv preprint arXiv:2207.01223, 2022. 1
  38. 38.Gyungin Shin, Weidi Xie, and Samuel Albanie. Namedmask: Distilling segmenters from complementary foundation models. arXiv:2209.11228, 2022. 3
  39. 39.Gyungin Shin, Weidi Xie, and Samuel Albanie. Reco: Retrieve and co-segment for zero-shot transfer. arXiv preprint arXiv:2206.07045, 2022. 3, 8
  40. 40.Oriane Simeoni, Gilles Puy, Huy V Vo, Simon Roburin, Spyros Gidaris, Andrei Bursuc, Patrick P´erez, Renaud Marlet, and Jean Ponce. Localizing objects with self-supervised transformers and no labels. arXiv preprint arXiv:2109.14279, 2021. 2
  41. 41.Korsuk Sirinukunwattana, Josien PW Pluim, Hao Chen, Xiaojuan Qi, Pheng-Ann Heng, Yun Bo Guo, Li Yang Wang, Bogdan J Matuszewski, Elia Bruni, Urko Sanchez, et al. Gland segmentation in colon histology images: The glas challenge contest. Medical Image Analysis, 35:489–502, 2017. 1
  42. 42.Laurens Van der Maaten and Geoffrey Hinton. Visualizing data using t-sne. Journal of Machine Learning Research, 9(11), 2008. 6
  43. 43.Wouter Van Gansbeke, Simon Vandenhende, Stamatios Georgoulis, and Luc Van Gool. Unsupervised semantic segmentation by contrasting object mask proposals. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10052–10062, 2021. 2, 7, 8
  44. 44.Wouter Van Gansbeke, Simon Vandenhende, and Luc Van Gool. Discovering object masks with transformers for unsupervised semantic segmentation. arXiv preprint arXiv:2206.06363, 2022. 2, 3, 7
  45. 45.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in Neural Information Processing Systems, 30, 2017. 2, 3, 4
  46. 46.Ulrike Von Luxburg. A tutorial on spectral clustering. Statistics and Computing, 17(4):395–416, 2007. 8
  47. 47.Xinlong Wang, Zhiding Yu, Shalini De Mello, Jan Kautz, Anima Anandkumar, Chunhua Shen, and Jose M Alvarez. Freesolo: Learning to segment objects without annotations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14176–14186, 2022. 3
  48. 48.Yangtao Wang, Xi Shen, Shell Xu Hu, Yuan Yuan, James L Crowley, and Dominique Vaufreydaz. Self-supervised transformers for unsupervised object discovery using normalized cut. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14543–14553, 2022. 2
  49. 49.Yude Wang, Jie Zhang, Meina Kan, Shiguang Shan, and Xilin Chen. Self-supervised equivariant attention mechanism for weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12275–12284, 2020. 1
  50. 50.Yuan Wang, Wei Zhuo, Yucong Li, Zhi Wang, Qi Ju, and Wenwu Zhu. Fully self-supervised learning for semantic segmentation. arXiv preprint arXiv:2202.11981, 2022. 2
  51. 51.Xin Wen, Bingchen Zhao, Anlin Zheng, Xiangyu Zhang, and Xiaojuan Qi. Self-supervised visual representation learning with semantic grouping. arXiv preprint arXiv:2205.15288, 2022. 2
  52. 52.Jiarui Xu, Shalini De Mello, Sifei Liu, Wonmin Byeon, Thomas Breuel, Jan Kautz, and Xiaolong Wang. Groupvit: Semantic segmentation emerges from text supervision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18134–18144, 2022. 3, 8
  53. 53.Zhaoyuan Yin, Pichao Wang, Fan Wang, Xianzhe Xu, Hanling Zhang, Hao Li, and Rong Jin. Transfgu: a top-down approach to fine-grained unsupervised semantic segmentation. In European Conference on Computer Vision, pages 73–89. Springer, 2022. 2, 3, 7
  54. 54.Andrii Zadaianchuk, Matthaeus Kleindessner, Yi Zhu, Francesco Locatello, and Thomas Brox. Unsupervised semantic segmentation with self-supervised object-centric representations. arXiv preprint arXiv:2207.05027, 2022. 2
  55. 55.Feihu Zhang, Philip Torr, Ren´e Ranftl, and Stephan Richter. Looking beyond single images for contrastive semantic segmentation learning. Advances in Neural Information Processing Systems, 34:3285–3297, 2021. 2
  56. 56.Xiao Zhang and Michael Maire. Self-supervised visual representation learning from hierarchical grouping. Advances in Neural Information Processing Systems, 33:16579–16590, 2020. 2
  57. 57.Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Scene parsing through ade20k dataset. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 633–641, 2017. 1
  58. 58.Chong Zhou, Chen Change Loy, and Bo Dai. Extract free dense labels from clip. In European Conference on Computer Vision (ECCV), 2022. 3, 6, 8
  59. 59.Adrian Ziegler and Yuki M Asano. Self-supervised learning of object parts for semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14502–14511, 2022. 2, 7

Citation

MLA
Li, K., et al. “ACSeg: Adaptive Conceptualization for Unsupervised Semantic Segmentation”. arXiv, 2022, http://arxiv.org/abs/2210.05944v3.
APA
Li, K., Wang, Z., Cheng, Z., Yu, R., Zhao, Y., Song, G., Liu, C., Yuan, L., & Chen, J. (2022). ACSeg: Adaptive Conceptualization for Unsupervised Semantic Segmentation. arXiv. http://arxiv.org/abs/2210.05944v3
Chicago
Li, K., Z. Wang, Z. Cheng, et al. 2022. “ACSeg: Adaptive Conceptualization for Unsupervised Semantic Segmentation”. arXiv. http://arxiv.org/abs/2210.05944v3.
Harvard
Li, K. et al. (2022) “ACSeg: Adaptive Conceptualization for Unsupervised Semantic Segmentation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2210.05944v3.
Vancouver
1. Li K, Wang Z, Cheng Z, Yu R, Zhao Y, Song G, Liu C, Yuan L, Chen J (2022) ACSeg: Adaptive Conceptualization for Unsupervised Semantic Segmentation. arXiv

BibTeX

@article{li2022acseg,
  title = {ACSeg: Adaptive Conceptualization for Unsupervised Semantic Segmentation},
  author = {Li, Kehan and Wang, Zhennan and Cheng, Zesen and Yu, Runyi and Zhao, Yian and Song, Guoli and Liu, Chang and Yuan, Li and Chen, Jie},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2210.05944v3},
  eprint = {2210.05944}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE