Unsupervised Semantic Segmentation by Distilling Feature Correspondences

Mark HamiltonZhoutong ZhangBharath HariharanNoah SnavelyWilliam T. Freeman

article2022ICLR317 citations

Develops STEGO, a framework that distills self-supervised feature correspondences into compact semantic clusters, dramatically improving unsupervised image segmentation on complex benchmarks like CocoStuff and Cityscapes.

Listen

Semantic segmentation—the automated task of classifying every individual pixel in an image into meaningful categories—is vital for applications ranging from autonomous driving to medical diagnostics. However, training these systems traditionally requires exhaustive manual labeling, which can take over one hundred times more effort per image than simple classification. In specialized fields, precise ground-truth labels may even be ill-defined or impossible to obtain. Consequently, developing automated methods that can discover and segment objects entirely without human annotations addresses a critical scalability bottleneck in visual intelligence.

The article demonstrates that modern self-supervised vision models inherently capture dense, semantically consistent relationships across images, and introduces a framework named STEGO (Self-supervised Transformer with Energy-based Graph Optimization) to distill these relationships into discrete, high-quality pixel-level segmentations without human supervision.

To achieve this, the authors separate the learning of visual features from cluster formation. Instead of training a model from scratch, the approach keeps a pre-trained self-supervised vision transformer backbone frozen and trains a lightweight segmentation head using an energy-based contrastive loss. This training process contrasts feature correspondences across three relationship types: within the same image, between an image and its nearest visual neighbors, and across random image pairs. The architecture incorporates targeted design modifications, including spatial centering and zero-clamping, to stabilize optimization and better detect small objects. It finishes with a standard clustering step and spatial refinement. The methodology was evaluated on benchmark segmentation datasets, including CocoStuff and Cityscapes, and grounded theoretically as maximum likelihood estimation on graph Potts models.

The experimental findings show substantial improvements over existing unsupervised approaches. Most notably, STEGO improves segmentation accuracy by 14 mean Intersection over Union (mIoU) points on the challenging CocoStuff benchmark (reaching 28.2 mIoU compared to the prior state of the art's 14.4 mIoU). On the Cityscapes urban driving dataset, it achieves an 8.7 mIoU gain (reaching 21.0 mIoU versus the prior best of 12.3 mIoU). In linear probe evaluations, which test underlying feature utility, the framework achieved 41.0 mIoU on CocoStuff compared to the previous best of 14.8 mIoU. Furthermore, the approach proves computationally efficient, training in under two hours on a single graphical processing unit because the primary visual backbone remains frozen.

These findings demonstrate that high-resolution visual segmentation can be effectively automated without expensive human annotation pipelines, substantially lowering development costs and training timelines for computer vision deployments. By decoupling feature extraction from segmentation, organizations can directly leverage pre-trained foundational vision models for domain-specific visual parsing without redesigning end-to-end architectures.

Stakeholders and engineering teams should consider adopting this correspondence distillation approach when building vision systems for unlabeled or annotation-scarce environments. However, before deploying the framework in mission-critical applications, practitioners should conduct focused pilot studies. Technical teams must specifically address current limitations: unsupervised models can struggle with ambiguous or arbitrary class boundaries (such as distinguishing between walls and ceilings or generic food subtypes), and manual tuning is currently required to balance the attractive and repulsive optimization pressures without ground-truth validation data. Confidence in the reported performance gains is high across standard benchmarks, though readers should account for the remaining performance gap when comparing these fully unsupervised outputs to supervised alternatives.

arXiv: 2203.08414mhamilton723/STEGO

No sufficiently relevant recommendations were found.

Cover for Unsupervised Semantic Segmentation by Distilling Feature Correspondences

Abstract

Unsupervised semantic segmentation aims to discover and localize semantically meaningful categories within image corpora without any form of annotation. To solve this task, algorithms must produce features for every pixel that are both semantically meaningful and compact enough to form distinct clusters. Unlike previous works which achieve this with a single end-to-end framework, we propose to separate feature learning from cluster compactification. Empirically, we show that current unsupervised feature learning frameworks already generate dense features whose correlations are semantically consistent. This observation motivates us to design STEGO (S\textbf{S}elf-supervised T\textbf{T}ransformer with E\textbf{E}nergy-based G\textbf{G}raph O\textbf{O}ptimization), a novel framework that distills unsupervised features into high-quality discrete semantic labels. At the core of STEGO is a novel contrastive loss function that encourages features to form compact clusters while preserving their relationships across the corpora. STEGO yields a significant improvement over the prior state of the art, on both the CocoStuff (+14 mIoU\textbf{+14 mIoU}) and Cityscapes (+9 mIoU\textbf{+9 mIoU}) semantic segmentation challenges.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methods
  • 3.1 Feature Correspondences Predict Class Co-Occurrence
  • 3.2 Distilling Feature Correspondences
  • 3.3 STEGO Architecture
  • 3.4 Relation to Potts Models and Energy-based Graph Optimization
  • 4 Experiments
  • 4.1 Evaluation Details
  • 4.2 Results
  • 4.3 Ablation Study
  • 5 Conclusion
  • References
  • A Appendix
  • A.1 Video and Code
  • A.2 Additional Results on the Potsdam-3 Dataset
  • A.3 Additional Ablation Study
  • A.4 Additional Qualitative Results
  • A.5 Failure Cases
  • A.6 Feature Correspondences Predict STEGO’s Errors
  • A.7 Higher Resolution Confusion Matrices
  • A.8 Relationship with Graph Energy Minimization
  • A.9 Continuous, Unsupervised, and Mini-batch CRF
  • A.10 Implementation Details
  • A.11 A Heuristic for Setting Hyper-parameters
  • A.12 A note on 5-Crop Nearest Neighbors

Knowls

  1. Knowl 1 — STEGO Framework for Unsupervised Semantic Segmentation

    model/method

    STEGO (Self-supervised Transformer with Energy-based Graph Optimization) is an unsupervised semantic segmentation framework that separates representation learning from cluster compactification. Rather than training a backbone end-to-end with heuristic clustering objectives, STEGO utilizes frozen dense visual representations from a pretrained self-supervised visual backbone (such as DINO ViT) and trains a lightweight segmentation head to distill spatial feature correspondences into discrete, compact semantic clusters.

    The framework consists of four primary stages:

    1. Feature Extraction: A frozen self-supervised backbone network N\mathcal{N} extracts dense spatial feature representations f∈RC×H×Wf \in \mathbb{R}^{C \times H \times W} for an input image xx.
    2. Nonlinear Projection Head: A segmentation head S:RC×H×W→RK×H×WS: \mathbb{R}^{C \times H \times W} \to \mathbb{R}^{K \times H \times W} (where code dimension K<CK < C) maps backbone features ff to a continuous representation s=S(f)s = S(f). The segmentation head architecture consists of the sum of a single linear layer and a two-layer ReLU multilayer perceptron (MLP), outputting vectors of dimension K=70K = 70.
    3. Correspondence Distillation: The head SS is optimized via a contrastive correspondence distillation loss across three pairs: self-image pairs (x,x)(x, x), KK-nearest neighbor image pairs (x,xknn)(x, x^{knn}), and random negative image pairs (x,xrand)(x, x^{rand}).
    4. Clustering and Spatial Refinement: At inference, continuous segmentation codes ss are clustered across the dataset using minibatch KK-means with cosine distance to assign discrete semantic labels. Discrete labels are subsequently refined at full pixel resolution using a fully connected Conditional Random Field (CRF).
  2. Knowl 2 — Correspondence Distillation Loss with Clamping and Spatial Centering

    equation

    Let f∈RC×H×Wf \in \mathbb{R}^{C \times H \times W} and g∈RC×I×Jg \in \mathbb{R}^{C \times I \times J} be feature tensors extracted by a frozen backbone from images xx and yy, and let s=S(f)∈RK×H×Ws = S(f) \in \mathbb{R}^{K \times H \times W} and t=S(g)∈RK×I×Jt = S(g) \in \mathbb{R}^{K \times I \times J} denote their corresponding segmentation embeddings produced by segmentation head SS.

    The spatial feature correspondence tensor F∈RH×W×I×JF \in \mathbb{R}^{H \times W \times I \times J} measures pairwise cosine similarity between normalized backbone features:

    Fhwij:=∑c=1Cfchw∥fhw∥2gcij∥gij∥2F_{hwij} := \sum_{c=1}^C \frac{f_{chw}}{\|f_{hw}\|_2} \frac{g_{cij}}{\|g_{ij}\|_2}

    To balance learning signals for small objects with highly localized correspondences, Spatial Centering (SC) subtracts the average correlation across the target image spatial dimensions:

    FhwijSC:=Fhwij−1IJ∑i′=1I∑j′=1JFhwi′j′F^{SC}_{hwij} := F_{hwij} - \frac{1}{IJ}\sum_{i'=1}^I \sum_{j'=1}^J F_{hwi'j'}

    The segmentation correspondence tensor Shwij∈RH×W×I×JS_{hwij} \in \mathbb{R}^{H \times W \times I \times J} measures cosine similarity between distilled features:

    Shwij:=∑k=1Kskhw∥shw∥2tkij∥tij∥2S_{hwij} := \sum_{k=1}^K \frac{s_{khw}}{\|s_{hw}\|_2} \frac{t_{kij}}{\|t_{ij}\|_2}

    To avoid co-linearity and optimization instability caused by penalizing uncorrelated features with anti-alignment (pushing towards −1-1), segmentation similarities are clamped at 00 (0-Clamp), driving weakly correlated features toward orthogonality instead. The final correspondence distillation loss between image pair (x,y)(x, y) given negative pressure bias b∈Rb \in \mathbb{R} is defined as:

    Lcorr(x,y,b):=−∑h=1H∑w=1W∑i=1I∑j=1J(FhwijSC−b)max⁡(Shwij,0)\mathcal{L}_{corr}(x, y, b) := -\sum_{h=1}^H \sum_{w=1}^W \sum_{i=1}^I \sum_{j=1}^J (F^{SC}_{hwij} - b) \max(S_{hwij}, 0)

  3. Knowl 3 — STEGO Multi-Target Optimization Objective and Minibatch Sampling

    model/method

    STEGO optimizes its segmentation head SS by combining three instantiations of the correspondence distillation loss Lcorr\mathcal{L}_{corr}:

    L=λselfLcorr(x,x,bself)+λknnLcorr(x,xknn,bknn)+λrandLcorr(x,xrand,brand)\mathcal{L} = \lambda_{self} \mathcal{L}_{corr}(x, x, b_{self}) + \lambda_{knn} \mathcal{L}_{corr}(x, x^{knn}, b_{knn}) + \lambda_{rand} \mathcal{L}_{corr}(x, x^{rand}, b_{rand})

    where λself,λknn,λrand∈R≥0\lambda_{self}, \lambda_{knn}, \lambda_{rand} \in \mathbb{R}_{\ge 0} balance the loss components, and bself,bknn,brand∈Rb_{self}, b_{knn}, b_{rand} \in \mathbb{R} control the ratio of attractive (positive) to repulsive (negative) forces.

    Sampling and balancing mechanics operate as follows:

    • Nearest Neighbors (xknnx^{knn}): Prior to training, global image embeddings are constructed via global average pooling GAP(f)\text{GAP}(f) over backbone features. For each training crop, top-77 nearest neighbors are precomputed using cosine similarity. In each training step, xknnx^{knn} is sampled uniformly at random from the image's top-77 KNNs.
    • Random Images (xrandx^{rand}): Sampled by shuffling the minibatch such that no image matches with itself, providing repulsive negative pressure.
    • Loss Balancing: The loss weights are set such that λself≈λrand≈2λknn\lambda_{self} \approx \lambda_{rand} \approx 2\lambda_{knn}. Negative pressures are tuned per dataset so that mean KNN feature similarity settles near ≈0.3\approx 0.3 and mean random similarity near ≈0.0\approx 0.0.
    • Resolution Independence: Losses are calculated by randomly sampling 121 spatial coordinate pairs from source and target feature grids using bilinear grid sampling.
  4. Knowl 4 — Equivalence of Correspondence Distillation to Potts Model Maximum Likelihood Estimation

    theoretical result

    The correspondence distillation objective is mathematically equivalent to Maximum Likelihood Estimation (MLE) on a continuous Potts / Ising model defined on an undirected, fully connected graph G=(V,w)G = (V, w) of all dataset pixels.

    Let VV be the set of all spatial pixel locations across an entire dataset of images XX, where each vertex v=(n,h,w)v = (n, h, w) denotes spatial position (h,w)(h, w) in image nn. Let ϕ:V→RK\phi: V \to \mathbb{R}^K be parameterized by ϕ(v)=S(N(x))\phi(v) = S(\mathcal{N}(x)), mapping pixels to normalized code vectors sv/∥sv∥2s_v / \|s_v\|_2. Defining edge weights w(vi,vj)=F(vi,vj)−bw(v_i, v_j) = F(v_i, v_j) - b (where FF is backbone feature cosine similarity) and the compatibility function as cosine distance μ(ϕ(vi),ϕ(vj))=−⟨ϕ(vi),ϕ(vj)⟩\mu(\phi(v_i), \phi(v_j)) = - \langle \phi(v_i), \phi(v_j) \rangle, the Potts graph energy is:

    E(ϕ)=∑vi,vj∈Vw(vi,vj)μ(ϕ(vi),ϕ(vj))=−∑vi,vj∈V(F(vi,vj)−b)⟨ϕ(vi),ϕ(vj)⟩E(\phi) = \sum_{v_i, v_j \in V} w(v_i, v_j) \mu(\phi(v_i), \phi(v_j)) = - \sum_{v_i, v_j \in V} (F(v_i, v_j) - b) \langle \phi(v_i), \phi(v_j) \rangle

    The Boltzmann distribution over the function space Φ\Phi is:

    p(ϕ∣w,μ)=exp⁡(−E(ϕ))∫Φexp⁡(−E(ϕ′))dϕ′p(\phi \mid w, \mu) = \frac{\exp(-E(\phi))}{\int_\Phi \exp(-E(\phi')) d\phi'}

    Maximizing the likelihood arg⁡max⁡ϕ∈Φp(ϕ∣w,μ)\arg\max_{\phi \in \Phi} p(\phi \mid w, \mu) simplifies to minimizing the energy functional arg⁡min⁡SE(S∘N)\arg\min_S E(S \circ \mathcal{N}), which directly recovers the sum over all image pairs of the basic correspondence loss:

    arg⁡max⁡ϕ∈Φp(ϕ∣w,μ)=arg⁡min⁡S∑x,y∈XLsimple−corr(x,y,b)\arg\max_{\phi \in \Phi} p(\phi \mid w, \mu) = \arg\min_S \sum_{x, y \in X} \mathcal{L}_{simple-corr}(x, y, b)

  5. Knowl 5 — Unsupervised Feature Correspondences as Predictors of Semantic Label Co-occurrence

    empirical result

    Pretrained self-supervised vision transformer features (specifically DINO ViT) without any fine-tuning or supervision exhibit dense spatial feature correspondences that strongly predict ground truth semantic class co-occurrences.

    When treating pairwise cosine feature similarity FhwijF_{hwij} as a logit to classify whether two pixels share the same ground-truth class label (Lhwij=1L_{hwij} = 1 if lhw=kijl_{hw} = k_{ij}, else 00):

    • DINO self-correspondences achieve 90%90\% precision at 50%50\% recall on the 27-class CocoStuff benchmark, obtaining an Average Precision (AP) of 79%79\%.
    • DINO significantly outperforms self-supervised convolutional representations from MoCoV2 (AP=65%\text{AP} = 65\%) and spatial Gaussian CRF kernels (AP=60%\text{AP} = 60\%).
    • After training the STEGO segmentation head with correspondence distillation, the predicted label co-occurrence Average Precision improves to 86%86\%, demonstrating that distillation amplifies the supervisory correlation signal across the dataset.
  6. Knowl 6 — CocoStuff 27-Class Benchmark Performance

    data/table

    STEGO was evaluated on the 27 mid-level classes of the CocoStuff dataset using standard unsupervised clustering (evaluated via Hungarian matching against ground truth) and linear probe evaluation (a linear classifier trained on fixed segmentation features). Validation images were resized to 320 along the minor axis and center cropped to 320×320320 \times 320.

    Model Unsupervised Linear Probe
    Accuracy mIoU Accuracy mIoU
    ResNet50 24.6 8.9 41.3 10.2
    MoCoV2 25.2 10.4 44.4 13.2
    DINO 30.5 9.6 66.8 29.4
    Deep Cluster 19.9 - - -
    SIFT 20.2 - - -
    Doersch et al. 23.1 - - -
    Isola et al. 24.3 - - -
    AC 30.8 - - -
    InMARS 31.0 - - -
    IIC 21.8 6.7 44.5 8.4
    MDC 32.2 9.8 48.6 13.3
    PiCIE 48.1 13.8 54.2 13.9
    PiCIE + H 50.0 14.4 54.8 14.8
    STEGO (Ours) 56.9 28.2 76.1 41.0

    STEGO establishes a new state-of-the-art on CocoStuff-27, improving unsupervised mIoU by +13.8+13.8 points (+95.8%+95.8\% relative improvement over PiCIE+H) and linear probe mIoU by +26.2+26.2 points.

  7. Knowl 7 — Cityscapes and Potsdam-3 Unsupervised Segmentation Performance

    data/table

    STEGO was evaluated on the 27 classes of Cityscapes and the 3 classes of the Potsdam aerial segmentation dataset under fully unsupervised evaluation protocols.

    Cityscapes (27 Classes) Unsupervised Accuracy (%) Unsupervised mIoU
    IIC 47.9 6.4
    MDC 40.7 7.1
    PiCIE 65.5 12.3
    STEGO (Ours) 73.2 21.0
    Potsdam-3 Aerial Dataset Unsupervised Accuracy (%)
    Random CNN 38.2
    SIFT 38.2
    Deep Cluster 41.7
    K-Means 45.7
    Doersch et al. 49.6
    Isola et al. 63.9
    IIC 65.1
    STEGO (Ours) 77.0

    On Cityscapes-27, STEGO surpasses prior state-of-the-art (PiCIE) by +8.7+8.7 mIoU and +7.7%+7.7\% accuracy without fine-tuning backbone weights on driving data. On Potsdam-3, STEGO surpasses IIC by +11.9%+11.9\% accuracy.

  8. Knowl 8 — Ablation Analysis of STEGO Architectural Components and Losses

    data/table

    An ablation study on CocoStuff (27 classes) demonstrates the contribution of individual backbone architectures, loss adjustments, and preprocessing/post-processing modules:

    • 0-Clamp: Clamping negative segmentation correlation forces at 0 prevents co-linearity instabilities.
    • 5-Crop: Pre-cropping images into 5 crops (0.5h×0.5w0.5h \times 0.5w) prior to KNN mining quadruples local image detail resolution.
    • SC: Spatial centering balances attractive forces for small semantic objects.
    • CRF: Post-processing label maps with fully connected CRF aligns boundaries to high-resolution image edges.
    Backbone 0-Clamp 5-Crop SC CRF Unsupervised Linear Probe
    Acc. mIoU Acc. mIoU
    MoCoV2 ✓ 48.4 20.8 70.7 26.5
    ViT-S 34.2 7.3 54.9 15.6
    ViT-S ✓ 44.3 21.3 70.9 36.8
    ViT-S ✓ ✓ 47.6 23.4 72.2 36.8
    ViT-S ✓ ✓ ✓ 47.7 24.0 72.9 38.4
    ViT-S ✓ ✓ ✓ ✓ 48.3 24.5 74.4 38.3
    ViT-B ✓ ✓ ✓ 54.8 26.8 74.3 39.5
    ViT-B ✓ ✓ ✓ ✓ 56.9 28.2 76.1 41.0

    Ablation of individual loss terms shows that removing the random negative loss (Rand-Loss) causes the largest performance drop (Unsupervised mIoU falls from 24.524.5 to 12.812.8 on ViT-S), while removing KNN-Loss reduces mIoU to 20.220.2 and removing Self-Loss reduces mIoU to 22.222.2.

  9. Knowl 9 — Continuous and Unsupervised Formulation of Conditional Random Fields

    model/method

    The correspondence energy minimization formulation enables extending fully connected Gaussian Conditional Random Fields (CRFs) into continuous, unsupervised, minibatch formulations without unary potentials.

    In standard fully connected CRFs, the pairwise edge potential between pixel locations pi,pjp_i, p_j with RGB color vectors Ii,IjI_i, I_j is defined by:

    wcrf(vi,vj)=aexp⁡(−∥pi−pj∥22θα2−∥Ii−Ij∥22θβ2)+bexp⁡(−∥pi−pj∥22θγ2)w_{crf}(v_i, v_j) = a \exp\left( -\frac{\|p_i - p_j\|^2}{2\theta_\alpha^2} - \frac{\|I_i - I_j\|^2}{2\theta_\beta^2} \right) + b \exp\left( -\frac{\|p_i - p_j\|^2}{2\theta_\gamma^2} \right)

    When unary classification potentials are unavailable, strictly positive pairwise affinities cause the maximum likelihood solution to collapse to a constant single-cluster state. By subtracting a negative pressure constant bcrfb_{crf} from the edge weights (wcrf′=wcrf−bcrfw'_{crf} = w_{crf} - b_{crf}) analogous to negative sampling, the energy forces unrelated pixels apart.

    Lifting the state space from discrete probability simplices P(l)\mathcal{P}(l) to a continuous embedding space RK\mathbb{R}^K with cosine compatibility allows optimizing continuous superpixel representations with minibatch gradient descent, avoiding discrete local minima.

  10. Knowl 10 — Limitations in Unsupervised Semantic Segmentation and Ontology Ambiguities

    limitation

    STEGO and related unsupervised segmentation frameworks exhibit specific systematic limitations:

    1. Arbitrary Semantic Ontologies: Unsupervised clustering cannot resolve human-defined arbitrary taxonomical distinctions without external priors. For instance, STEGO systematically conflates 'food (things)' with 'food (stuff)', struggles to separate 'walls' from 'ceilings', and yields inconsistent segmentations for abstract categories like 'indoor', 'accessory', 'textile', and 'rawmaterial'.
    2. Error Inheritance from Backbone Correspondences: Confusion patterns in the trained model closely match error patterns already present in the raw DINO feature correspondences, showing that errors originate largely from representation structure rather than optimization failure.
    3. Unsupervised Hyperparameter Calibration: Hyperparameters (such as negative biases bself,bknn,brandb_{self}, b_{knn}, b_{rand}) must currently be tuned manually by monitoring whether the distribution of learned feature similarities exhibits a healthy bimodal distribution (peaking at 00 for orthogonality and 11 for alignment), as automated cross-validation without ground-truth labels remains an open challenge.

Coverage note — None was omitted; all main architectural designs, loss formulations, theoretical links to Potts models/CRFs, quantitative benchmark tables, ablation analyses, and core limitations are fully covered.

References

  1. 1.Martın Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dan Mane, Rajat Monga, Sherry Moore, Derek Murray, Chris Olah,  Mike Schuster, Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul Tucker, Vincent Vanhoucke, Vijay Vasudevan, Fernanda Viegas, Oriol Vinyals, Pete Warden, Martin Watten- berg, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng. TensorFlow: Large-scale machine learning on heterogeneous systems, 2015. URL http://tensorflow.org/. Software available from tensorflow.org.
  2. 2.Jiwoon Ahn, Sunghyun Cho, and Suha Kwak. Weakly supervised learning of instance segmentation with inter-pixel relations. CoRR, abs/1904.05044, 2019. URL http://arxiv.org/abs/1904.05044.
  3. 3.George A Baker Jr and John M Kincaid. Continuous-spin ising model and ̂̀: ́ 4: d field theory. Physical Review Letters, 42(22):1431, 1979.
  4. 4.Hakan Bilen, Rodrigo Benenson, and Seong Joon Oh. Eccv 2020 tutorial on weakly-supervised learning in computer vision. URL https://github.com/hbilen/wsl-eccv20. github.io.
  5. 5.Holger Caesar, Jasper Uijlings, and Vittorio Ferrari. Coco-stuff: Thing and stuff classes in context. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 1209–1218, 2018.
  6. 6.Mathilde Caron, Piotr Bojanowski, Armand Joulin, and Matthijs Douze. Deep clustering for unsupervised learning of visual features. In Proceedings of the European Conference on Computer Vision (ECCV), pp. 132–149, 2018.
  7. 7.Mathilde Caron, Hugo Touvron, Ishan Misra, Herve J  egou, Julien Mairal, Piotr Bojanowski, and  Armand Joulin. Emerging properties in self-supervised vision transformers. arXiv preprint arXiv:2104.14294, 2021.
  8. 8.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. Semantic image segmentation with deep convolutional nets and fully connected crfs. arXiv preprint arXiv:1412.7062, 2014.
  9. 9.Liang-Chieh Chen, George Papandreou, Florian Schroff, and Hartwig Adam. Rethinking atrous convolution for semantic image segmentation. arXiv preprint arXiv:1706.05587, 2017.
  10. 10.Ting Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi, and Geoffrey Hinton. Big selfsupervised models are strong semi-supervised learners. arXiv preprint arXiv:2006.10029, 2020a.
  11. 11.Xinlei Chen, Haoqi Fan, Ross Girshick, and Kaiming He. Improved baselines with momentum contrastive learning. arXiv preprint arXiv:2003.04297, 2020b.
  12. 12.Xinlei Chen, Haoqi Fan, Ross Girshick, and Kaiming He. Improved baselines with momentum contrastive learning. arXiv preprint arXiv:2003.04297, 2020c.
  13. 13.Jang Hyun Cho, U. Mall, K. Bala, and Bharath Hariharan. Picie: Unsupervised semantic segmentation using invariance and equivariance in clustering. ArXiv, abs/2103.17070, 2021.
  14. 14.Edo Collins, Radhakrishna Achanta, and Sabine Susstrunk. Deep feature factorization for concept discovery. In Proceedings of the European Conference on Computer Vision (ECCV), pp. 336–352, 2018.
  15. 15.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pp. 248–255. Ieee, 2009.
  16. 16.Carl Doersch, Abhinav Gupta, and Alexei A Efros. Unsupervised visual representation learning by context prediction. In Proceedings of the IEEE international conference on computer vision, pp. 1422–1430, 2015.
  17. 17.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  18. 18.William Falcon et al. Pytorch lightning. GitHub. Note: https://github.com/PyTorchLightning/pytorch-lightning, 3, 2019.
  19. 19.Spyros Gidaris, Praveer Singh, and Nikos Komodakis. Unsupervised representation learning by predicting image rotations. arXiv preprint arXiv:1803.07728, 2018.
  20. 20.Xavier Glorot, Antoine Bordes, and Yoshua Bengio. Deep sparse rectifier neural networks. In Proceedings of the fourteenth international conference on artificial intelligence and statistics, pp. 315–323. JMLR Workshop and Conference Proceedings, 2011.
  21. 21.Mark Hamilton, Scott Lundberg, Lei Zhang, Stephanie Fu, and William T Freeman. Model-agnostic explainability for visual search. arXiv preprint arXiv:2103.00370, 2021.
  22. 22.Yufei Han and Maurizio Filippone. Mini-batch spectral clustering. In 2017 International Joint Conference on Neural Networks (IJCNN), pp. 3888–3895. IEEE, 2017.
  23. 23.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, 2016.
  24. 24.Geoffrey E Hinton. Training products of experts by minimizing contrastive divergence. Neural computation, 14(8):1771–1800, 2002.
  25. 25.R Devon Hjelm, Alex Fedorov, Samuel Lavoie-Marchildon, Karan Grewal, Phil Bachman, Adam Trischler, and Yoshua Bengio. Learning deep representations by mutual information estimation and maximization. arXiv preprint arXiv:1808.06670, 2018.
  26. 26.Jyh-Jing Hwang, Stella X. Yu, Jianbo Shi, Maxwell D. Collins, Tien-Ju Yang, Xiao Zhang, and Liang-Chieh Chen. Segsort: Segmentation by discriminative sorting of segments. CoRR, abs/1910.06962, 2019. URL http://arxiv.org/abs/1910.06962.
  27. 27.Phillip Isola, Daniel Zoran, Dilip Krishnan, and Edward H Adelson. Learning visual groups from co-occurrences in space and time. arXiv preprint arXiv:1511.06811, 2015.
  28. 28.Max Jaderberg, Karen Simonyan, Andrew Zisserman, and Koray Kavukcuoglu. Spatial transformer networks. arXiv preprint arXiv:1506.02025, 2015.
  29. 29.Xu Ji, Joao F Henriques, and Andrea Vedaldi. Invariant information clustering for unsupervised ȷ image classification and segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 9865–9874, 2019.
  30. 30.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  31. 31.Philipp Krahenb  uhl and Vladlen Koltun. Efficient inference in fully connected crfs with gaussian  edge potentials. Advances in neural information processing systems, 24:109–117, 2011.
  32. 32.John Lafferty, Andrew McCallum, and Fernando CN Pereira. Conditional random fields: Probabilistic models for segmenting and labeling sequence data. 2001.
  33. 33.Omer Levy and Yoav Goldberg. Neural word embedding as implicit matrix factorization. Advances in neural information processing systems, 27:2177–2185, 2014.
  34. 34.Yunfan Li, Peng Hu, Zitao Liu, Dezhong Peng, Joey Tianyi Zhou, and Xi Peng. Contrastive clustering. arXiv preprint arXiv:2009.09687, 2020.
  35. 35.Yun Liu, Yu-Huan Wu, Peisong Wen, Yujun Shi, Yu Qiu, and Ming-Ming Cheng. Leveraging instance-, image- and dataset-level information for weakly supervised instance segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020. doi: 10.1109/TPAMI. 2020.3023152.
  36. 36.Jonathan Long, Evan Shelhamer, and Trevor Darrell. Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 3431–3440, 2015.
  37. 37.David G Lowe. Object recognition from local scale-invariant features. In Proceedings of the seventh IEEE international conference on computer vision, volume 2, pp. 1150–1157. Ieee, 1999.
  38. 38.James MacQueen et al. Some methods for classification and analysis of multivariate observations. In Proceedings of the fifth Berkeley symposium on mathematical statistics and probability, volume 1, pp. 281–297. Oakland, CA, USA, 1967.
  39. 39.Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg Corrado, and Jeffrey Dean. Distributed representations of words and phrases and their compositionality. arXiv preprint arXiv:1310.4546, 2013.
  40. 40.S Ehsan Mirsadeghi, Ali Royat, and Hamid Rezatofighi. Unsupervised image segmentation by mutual information maximization and adversarial regularization. IEEE Robotics and Automation Letters, 6(4):6931–6938, 2021.
  41. 41.Annamalai Narayanan, Mahinthan Chandramohan, Rajasekar Venkatesan, Lihui Chen, Yang Liu, and Shantanu Jaiswal. graph2vec: Learning distributed representations of graphs. arXiv preprint arXiv:1707.05005, 2017.
  42. 42.Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2018.
  43. 43.Yassine Ouali, Celine Hudelot, and Myriam Tami. Autoregressive unsupervised image segmenta-  tion. In European Conference on Computer Vision, pp. 142–158. Springer, 2020.
  44. 44.Shun-Yi Pan, Cheng-You Lu, Shih-Po Lee, and Wen-Hsiao Peng. Weakly-supervised image semantic segmentation using graph convolutional networks. In 2021 IEEE International Conference on Multimedia and Expo (ICME), pp. 1–6. IEEE, 2021.
  45. 45.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. Pytorch: An imperative style, high-performance deep learning library. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alche-Buc, E. Fox, and  R. Garnett (eds.), Advances in Neural Information Processing Systems 32, pp. 8024–8035. Curran Associates, Inc., 2019.
  46. 46.Deepak Pathak, Philipp Krahenb  uhl, Jeff Donahue, Trevor Darrell, and Alexei Efros. Context en-  coders: Feature learning by inpainting. In CVPR, 2016.
  47. 47.Fabian Pedregosa, Gael Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier  Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, et al. Scikit-learn: Machine learning in python. the Journal of machine Learning research, 12:2825–2830, 2011.
  48. 48.Pedro O Pinheiro, Amjad Almahairi, Ryan Y Benmalek, Florian Golemo, and Aaron Courville. Unsupervised learning of dense visual representations. arXiv preprint arXiv:2011.05499, 2020.
  49. 49.Renfrey Bernard Potts. Some generalized order-disorder transformations. In Mathematical proceedings of the cambridge philosophical society, volume 48, pp. 106–109. Cambridge University Press, 1952.
  50. 50.Rene Ranftl, Alexey Bochkovskiy, and Vladlen Koltun. Vision transformers for dense prediction.  ArXiv preprint, 2021.
  51. 51.Zhongzheng Ren, Zhiding Yu, Xiaodong Yang, Ming-Yu Liu, Alexander G Schwing, and Jan Kautz. Ufo2: A unified framework towards omni-supervised object detection. In European Conference on Computer Vision, pp. 288–313. Springer, 2020.
  52. 52.Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. Dropout: a simple way to prevent neural networks from overfitting. The journal of machine learning research, 15(1):1929–1958, 2014.
  53. 53.Zachary Teed and Jia Deng. RAFT: recurrent all-pairs field transforms for optical flow. CoRR, abs/2003.12039, 2020. URL https://arxiv.org/abs/2003.12039.
  54. 54.Marvin TT Teichmann and Roberto Cipolla. Convolutional crfs for semantic segmentation. arXiv preprint arXiv:1805.04777, 2018.
  55. 55.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herve J  egou. Training data-efficient image transformers & distillation through attention. In  International Conference on Machine Learning, pp. 10347–10357. PMLR, 2021.
  56. 56.Wouter Van Gansbeke, Simon Vandenhende, Stamatios Georgoulis, Marc Proesmans, and Luc Van Gool. Scan: Learning to classify images without labels. In European Conference on Computer Vision, pp. 268–285. Springer, 2020.
  57. 57.Wouter Van Gansbeke, Simon Vandenhende, Stamatios Georgoulis, and Luc Van Gool. Unsupervised semantic segmentation by contrasting object mask proposals. arxiv preprint arxiv:2102.06191, 2021.
  58. 58.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in neural information processing systems, pp. 5998–6008, 2017.
  59. 59.Pascal Vincent, Hugo Larochelle, Yoshua Bengio, and Pierre-Antoine Manzagol. Extracting and composing robust features with denoising autoencoders. In Proceedings of the 25th international conference on Machine learning, pp. 1096–1103, 2008.
  60. 60.Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He. Non-local neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 7794–7803, 2018.
  61. 61.Zhirong Wu, Yuanjun Xiong, Stella X Yu, and Dahua Lin. Unsupervised feature learning via nonparametric instance discrimination. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 3733–3742, 2018.
  62. 62.Donghui Yan, Ling Huang, and Michael I Jordan. Fast approximate spectral clustering. In Proceedings of the 15th ACM SIGKDD international conference on Knowledge discovery and data mining, pp. 907–916, 2009.
  63. 63.Hongshan Yu, Zhengeng Yang, Lei Tan, Yaonan Wang, Wei Sun, Mingui Sun, and Yandong Tang. Methods and datasets on semantic segmentation: A review. Neurocomputing, 304:82–103, 2018.
  64. 64.Han Zhang, Ian Goodfellow, Dimitris Metaxas, and Augustus Odena. Self-attention generative adversarial networks. In International conference on machine learning, pp. 7354–7363. PMLR, 2019.
  65. 65.Richard Zhang, Phillip Isola, and Alexei A Efros. Split-brain autoencoders: Unsupervised learning by cross-channel prediction. In CVPR, 2017.
  66. 66.Xuewen Zhang, Selene E Chew, Zhenlin Xu, and Nathan D Cahill. Slic superpixels for efficient graph-based dimensionality reduction of hyperspectral imagery. In Algorithms and Technologies for Multispectral, Hyperspectral, and Ultraspectral Imagery XXI, volume 9472, pp. 947209. International Society for Optics and Photonics, 2015.
  67. 67.Bolei Zhou, Aditya Khosla, Agata Lapedriza, Aude Oliva, and Antonio Torralba. Learning deep features for discriminative localization. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2921–2929, 2016.
  68. 68.Aleksandar Zlateski, Ronnachai Jaroensri, Prafull Sharma, and Fredo Durand. On the importance  of label quality for semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 1479–1487, 2018.

Citation

MLA
Hamilton, M., et al. “Unsupervised Semantic Segmentation by Distilling Feature Correspondences”. arXiv, 2022, http://arxiv.org/abs/2203.08414v1.
APA
Hamilton, M., Zhang, Z., Hariharan, B., Snavely, N., & Freeman, W. T. (2022). Unsupervised Semantic Segmentation by Distilling Feature Correspondences. arXiv. http://arxiv.org/abs/2203.08414v1
Chicago
Hamilton, M., Z. Zhang, B. Hariharan, N. Snavely, and W. T. Freeman. 2022. “Unsupervised Semantic Segmentation by Distilling Feature Correspondences”. arXiv. http://arxiv.org/abs/2203.08414v1.
Harvard
Hamilton, M. et al. (2022) “Unsupervised Semantic Segmentation by Distilling Feature Correspondences”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.08414v1.
Vancouver
1. Hamilton M, Zhang Z, Hariharan B, Snavely N, Freeman WT (2022) Unsupervised Semantic Segmentation by Distilling Feature Correspondences. arXiv

BibTeX

@article{hamilton2022unsupervised,
  title = {Unsupervised Semantic Segmentation by Distilling Feature Correspondences},
  author = {Hamilton, Mark and Zhang, Zhoutong and Hariharan, Bharath and Snavely, Noah and Freeman, William T.},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.08414v1},
  eprint = {2203.08414}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/