Decoupled Contrastive Multi-View Clustering with High-Order Random Walks

Yiding LuYijie LinMouxing YangDezhong PengPeng HuXi Peng

article2024AAAI144 citations

Proposes a decoupled multi-view clustering framework that uses high-order random walks to rectify false positive and false negative pairs globally while preserving view-specific information through cross-view reconstruction.

Listen

Modern data systems increasingly rely on multi-view clustering to automatically group complex, unlabeled information collected across multiple perspectives or modalities, such as multi-camera security feeds, multi-lingual texts, and medical imaging. Unsupervised contrastive learning has emerged as a leading technique to group these multi-view datasets by pulling representations of matching instances closer while pushing mismatched pairs apart. However, existing contrastive techniques routinely misidentify similar data points as negative pairs (false negatives) or, when attempting local corrections via immediate neighborhood checks, mistakenly label dissimilar points as positive pairs (false positives). Furthermore, conventional architectures often discard view-specific unique details while trying to enforce consistency, severely degrading clustering accuracy when data is incomplete or corrupted with missing views.

The article evaluates and demonstrates a new framework called DIVIDE (Decoupled Contrastive Multi-View Clustering with High-Order Random Walks). The primary objective of the article is to show how DIVIDE simultaneously eliminates false pairing errors across global graph connections and preserves view-specific information, thereby delivering superior clustering performance and robust recovery for both complete and incomplete multi-view data.

To evaluate this framework, the authors conducted extensive empirical experiments across four benchmark datasets spanning 2,100 to 18,758 instances: Scene-15, Caltech-101, Reuters, and LandUse-21. The approach relies on two core innovations: a multi-step random walk mechanism that traces probabilistic graph transitions to identify true global relationships, and a decoupled architecture combining Siamese encoders with cross-view decoders to handle intra-view and inter-view contrastive objectives separately. The researchers benchmarked DIVIDE against nine state-of-the-art multi-view clustering methods across fully observed datasets and partially observed environments where up to 90% of views were missing.

The article establishes several key findings. First, DIVIDE consistently outperformed all nine competing methods across every tested benchmark; for example, under a 50% missing-view setting on the Caltech-101 benchmark, DIVIDE reached a clustering accuracy of 63.4%, surpassing the next best competing baseline (56.2%) by over 7 percentage points. Second, the framework proved exceptionally resilient to extreme data loss, maintaining performance advantages on the Scene-15 dataset even as the view missing rate rose from 0% to 90%. Third, the multi-step random walk strategy generated significantly lower target error (achieving Kullback-Leibler divergence scores of 8.2 and 5.1 on Scene-15 and Caltech-101) compared to standard nearest-neighbor approaches (which scored 22.3 and 17.2), confirming its ability to filter false positives and recover false negatives without requiring fragile parameter tuning across 3 to 7 walk steps. Finally, ablation studies showed that every architectural component—including intra-view contrastive loss, the cross-view decoder, momentum updates, and target rectification—provided measurable accuracy and consistency gains.

These findings indicate that multi-view clustering systems can achieve high accuracy without sacrificing unique, modality-specific details or failing in the presence of missing inputs. For technical leadership and operational stakeholders, this means unsupervised machine learning pipelines can be deployed in complex, real-world data environments with reduced data curation costs, lower computational sensitivity to hyperparameter tuning, and stronger resilience against missing sensor feeds or data streams. The decoupled decoder design also eliminates the need to develop separate, complex imputation models to artificially complete missing data prior to clustering.

Based on these results, technical teams implementing unsupervised representation learning or multi-view analytics should adopt high-order graph propagation techniques over rigid local neighborhood thresholds to manage sample pair assignments. Organizations dealing with missing data streams should also leverage cross-view reconstruction decoders to preserve modality-specific signals. For future work, development teams should pilot this architecture within their domain-specific pipelines and explore adapting DIVIDE's decoupled random-walk approach to broader unsupervised contrastive representation learning tasks.

The conclusions of the article carry high confidence due to consistent, validated empirical results across multiple data modalities and comparative ablation tests. However, readers should note that the current evaluations primarily reflect two-view scenarios and balanced benchmark settings. Further validation is warranted before deploying the architecture in operational environments featuring highly imbalanced clusters or three or more simultaneous data views.

  • Paper: GCFAgg: Global and Cross-View Feature Aggregation for Multi-View Clustering, Weiqing Yan et al. (2023). It analyzes the negative-pair conflict and neighborhood alignment challenges in multi-view contrastive clustering that DIVIDE directly aims to resolve.
  • Paper: Contrastive Multiview Coding, Yonglong Tian et al. (2019). It provides the foundational framework for contrastive multi-view representation learning that underpins modern multi-view clustering objectives.
  • Paper: Co-regularized Multi-view Spectral Clustering, Abhishek Kumar et al. (2011). It establishes classic multi-view spectral graph clustering principles and inter-view agreement formulations adapted by deep multi-view clustering methods.
  • Paper: A tutorial on spectral clustering, Ulrike von Luxburg (2007). It details the foundational theory connecting graph Laplacians, random walks, and spectral clustering partitions utilized in graph-based clustering.
  • Paper: Higher-order organization of complex networks, Austin R. Benson et al. (2016). It introduces higher-order graph structures and motifs for network partitioning, providing core motivation for DIVIDE's higher-order random walk formulation.
  • Paper: Unsupervised Deep Embedding for Clustering Analysis, Junyuan Xie et al. (2015). It introduces deep embedded clustering, laying the baseline framework for jointly optimizing neural embeddings and cluster assignments.

No sufficiently relevant recommendations were found.

Cover for Decoupled Contrastive Multi-View Clustering with High-Order Random Walks

Abstract

In recent, some robust contrastive multi-view clustering (MvC) methods have been proposed, which construct data pairs from neighborhoods to alleviate the false negative issue, i.e., some intra-cluster samples are wrongly treated as negative pairs. Although promising performance has been achieved by these methods, the false negative issue is still far from addressed and the false positive issue emerges because all in- and out-of-neighborhood samples are simply treated as positive and negative, respectively. To address the issues, we propose a novel robust method, dubbed decoupled contrastive multi-view clustering with high-order random walks (DIVIDE). In brief, DIVIDE leverages random walks to progressively identify data pairs in a global instead of local manner. As a result, DIVIDE could identify in-neighborhood negatives and out-of-neighborhood positives. Moreover, DIVIDE embraces a novel MvC architecture to perform inter- and intra-view contrastive learning in different embedding spaces, thus boosting clustering performance and embracing the robustness against missing views. To verify the efficacy of DIVIDE, we carry out extensive experiments on four benchmark datasets comparing with nine state-of-the-art MvC methods in both complete and incomplete MvC settings. The code is released on https://github.com/XLearning-SCU/2024-AAAI-DIVIDE.

Table of Contents

  • Introduction
  • Related Work
  • Deep Multi-view Clustering
  • Contrastive Multi-view Clustering
  • Method
  • Decoupled Contrastive Learning Framework
  • False Negative Identification via Random Walks
  • Experiments
  • Experimental Setup
  • Comparisons with State of the Arts (Q1)
  • Performance under Different Missing Rates (Q2)
  • Analysis on Multi-step Random Walks (Q3)
  • Ablation Studies (Q4)
  • Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — DIVIDE’s decoupled high-order contrastive clustering design

    model/method

    DIVIDE is a robust deep multi-view clustering method with two coupled contributions: a decoupled contrastive architecture and a high-order random-walk rectifier. The architecture learns cross-view consistency in view-specific embedding spaces rather than forcing all views into one common space, while the rectifier replaces binary identity targets with soft targets derived from multi-step walks on an affinity graph. High-order affinities can identify both out-of-neighborhood samples that should be positive and in-neighborhood samples that should remain negative, thereby addressing false negatives and false positives caused by first-order neighborhood rules. The workflow illustration on page 3 shows these two branches: view-specific Siamese encoders and cross-view decoders on the left, and batch affinity construction, transition-matrix powering, and target rectification on the right.

  2. Knowl 2 — View-specific Siamese encoders and cross-view decoders

    model/method

    For each view v∈{1,…,V}v\in\{1,\ldots,V\}, DIVIDE uses an online encoder fq(v)f_q^{(v)} and a target encoder fk(v)f_k^{(v)} with the same architecture. For an instance xi(v)x_i^{(v)}, the two embeddings are

    zq,i(v)=fq(v)(xi(v)),zk,i(v)=fk(v)(xi(v)).z_{q,i}^{(v)}=f_q^{(v)}(x_i^{(v)}),\qquad z_{k,i}^{(v)}=f_k^{(v)}(x_i^{(v)}).

    The target encoder is updated as an exponential-moving-average version of the online encoder, and gradients are stopped through the target branch. For every ordered pair of distinct views v≠uv\ne u, a cross-view decoder g(v→u)g^{(v\to u)} maps the online embedding of view vv into the embedding space of view uu:

    pi(v→u)=g(v→u) ⁣(zq,i(v)).p_i^{(v\to u)}=g^{(v\to u)}\!\left(z_{q,i}^{(v)}\right).

    This preserves view-specific information in zq(v)z_q^{(v)} while allowing cross-view consistency to be learned through the decoded representation p(v→u)p^{(v\to u)}.

  3. Knowl 3 — Decoupled intra-view and inter-view contrastive objective

    equation

    For a mini-batch of nn instances, DIVIDE minimizes the sum of intra-view and inter-view contrastive losses:

    L=Lintra+Linter=∑v=1VH ⁣(T(v),ρ ⁣(zq(v),zk(v)))+∑v≠uH ⁣(T(u),ρ ⁣(p(v→u),zk(u))).\mathcal{L}=\mathcal{L}_{\mathrm{intra}}+\mathcal{L}_{\mathrm{inter}} =\sum_{v=1}^{V}H\!\left(T^{(v)},\rho\!\left(z_q^{(v)},z_k^{(v)}\right)\right) +\sum_{v\ne u}H\!\left(T^{(u)},\rho\!\left(p^{(v\to u)},z_k^{(u)}\right)\right).

    Here T(v)∈Rn×nT^{(v)}\in\mathbb{R}^{n\times n} is the soft pairwise target for view vv, HH is row-wise cross entropy, and ρ\rho converts pairwise similarities into row-normalized probabilities. For embeddings ai,bj∈Rda_i,b_j\in\mathbb{R}^{d} and scalar similarity function s(⋅,⋅)s(\cdot,\cdot),

    [ρ(a,b)]ij=exp⁡ ⁣(s(ai,bj)/τ)∑ℓ=1nexp⁡ ⁣(s(ai,bℓ)/τ),[\rho(a,b)]_{ij}=\frac{\exp\!\left(s(a_i,b_j)/\tau\right)}{\sum_{\ell=1}^{n}\exp\!\left(s(a_i,b_\ell)/\tau\right)},

    where τ\tau is the contrastive temperature, fixed to 0.50.5 in the experiments. The first loss learns discrimination within each view-specific space, whereas the second aligns a decoded representation from view vv with the target representation of view uu.

  4. Knowl 4 — High-order random-walk affinity and false-negative rectification

    model/method

    For every view vv, DIVIDE constructs a fully connected affinity graph over the nn target embeddings in a mini-batch. The edge weight between samples ii and jj is a heat-kernel similarity:

    Aij(v)=exp⁡ ⁣(−∥zk,i(v)−zk,j(v)∥22σ),A_{ij}^{(v)}=\exp\!\left(-\frac{\left\|z_{k,i}^{(v)}-z_{k,j}^{(v)}\right\|_2^2}{\sigma}\right),

    where σ=0.1\sigma=0.1 is the kernel-width hyperparameter. With di(v)=∑j=1nAij(v)d_i^{(v)}=\sum_{j=1}^{n}A_{ij}^{(v)}, the row-normalized transition probability is Mij(v)=Aij(v)/di(v)M_{ij}^{(v)}=A_{ij}^{(v)}/d_i^{(v)}. A walk distribution evolves as p(t)=p(0)(M(v))tp^{(t)}=p^{(0)}(M^{(v)})^t, so (M(v))ijt(M^{(v)})^t_{ij} is the probability that sample jj is reached from anchor ii after tt steps. DIVIDE uses t=5t=5 and treats these high-order probabilities as soft evidence that a nominally negative pair is actually a false negative. This global, multi-step relation is not restricted to the first-order nearest-neighbor set.

  5. Knowl 5 — Self- and swap-rectified contrastive targets

    equation

    For each view vv, DIVIDE combines the ordinary same-instance identity target with the tt-step random-walk target:

    T(v)=αIn+(1−α)(M(v))t,T^{(v)}=\alpha I_n+(1-\alpha)(M^{(v)})^t,

    where InI_n is the n×nn\times n identity matrix and α=0.5\alpha=0.5. The identity component retains the known correspondence between two embeddings of the same instance, while the random-walk component assigns nonzero target mass to likely same-cluster pairs. In intra-view learning, the affinity graph and target are constructed from the same view, which is the self strategy. In inter-view learning from vv to uu, the target associated with the destination view uu is used with p(v→u)p^{(v\to u)} and zk(u)z_k^{(u)}, which is the swap strategy. Thus the self target emphasizes within-view information and the swap target supplies target-space affinities for cross-view interaction.

  6. Knowl 6 — Cross-view recovery of missing representations

    model/method

    DIVIDE handles incomplete multi-view data using the cross-view decoders rather than requiring a separate missing-view clustering mechanism. If view vv is observed for instance ii and view uu is missing, the missing representation is recovered as

    z^i(u)=g(v→u) ⁣(zk,i(v)).\widehat{z}_{i}^{(u)}=g^{(v\to u)}\!\left(z_{k,i}^{(v)}\right).

    The recovered representation is combined with the observed-view representations before clustering. Because each decoder is trained to map between view-specific embedding spaces, the model can reconstruct a missing view while retaining the complementary information encoded by the observed view. The incomplete-data pathway is also illustrated in the page-3 workflow.

  7. Knowl 7 — Datasets, clustering protocol, and implementation configuration

    experimental setup

    DIVIDE was evaluated on four two-view benchmarks: Scene-15 with 4,485 images from 15 scene classes using PHOG and GIST; Caltech-101 with 8,677 images from 101 object classes using DECAF and VGG19 features; Reuters with 18,758 multilingual samples using 10-dimensional autoencoder representations of English and French text; and LandUse-21 with 2,100 satellite images from 21 classes using PHOG and LBP. The final embeddings from all views were concatenated and clustered with kk-means, where kk is the number of ground-truth classes. Performance was measured by ACC, NMI, and ARI, with larger values being better.

    The model was trained for 200 epochs with Adam, no weight decay, learning rate 2×10−32\times10^{-3}, and batch size 1024. The target was the identity matrix for the first 100 warm-up epochs and the rectified target thereafter. The fixed hyperparameters were contrastive temperature τ=0.5\tau=0.5, kernel width σ=0.1\sigma=0.1, walk length t=5t=5, and identity-target weight α=0.5\alpha=0.5. View-specific encoders were four-layer fully connected networks with batch normalization and ReLU; cross-view decoders were two-layer fully connected networks. In the incomplete setting, a missing rate η\eta was simulated by randomly selecting m=ηnm=\eta n samples and deleting one view from each.

  8. Knowl 8 — Performance on complete and incomplete multi-view clustering

    data/table

    The comparative results reported on page 5 show that DIVIDE is generally the strongest method among nine baselines in both complete data and data where 50% of samples have one view removed. Values are reported in percentage points as ACC/NMI/ARI; the comparison value is the best competing baseline for each metric, with its method named when the winning baseline differs across metrics.

    • Incomplete Scene-15: DIVIDE achieved 46.8/45.7/29.146.8/45.7/29.1, versus the best baseline values 41.6/42.9/25.341.6/42.9/25.3 from ProImp.
    • Incomplete Caltech-101: DIVIDE achieved 63.4/82.5/52.463.4/82.5/52.4, versus 56.2/78.0/45.356.2/78.0/45.3 from DAIMC for ACC and NMI and DCP for ARI.
    • Incomplete Reuters: DIVIDE achieved 54.7/37.3/28.654.7/37.3/28.6, versus 51.9/35.5/28.551.9/35.5/28.5 from ProImp.
    • Incomplete LandUse-21: DIVIDE achieved 30.0/35.8/16.030.0/35.8/16.0, versus 23.1/28.6/10.623.1/28.6/10.6 from SURE.
    • Complete Scene-15: DIVIDE achieved 49.1/48.7/31.649.1/48.7/31.6, versus 43.6/45.1/26.843.6/45.1/26.8 from ProImp for ACC and ARI and DCP for NMI.
    • Complete Caltech-101: DIVIDE achieved 62.2/83.0/50.562.2/83.0/50.5, versus 57.5/78.7/51.957.5/78.7/51.9 from DAIMC for ACC and NMI and DCP for ARI; therefore DIVIDE did not achieve the best ARI in this one condition.
    • Complete Reuters: DIVIDE achieved 59.3/39.5/29.059.3/39.5/29.0, versus 56.5/39.4/32.856.5/39.4/32.8 from ProImp; DIVIDE led ACC and NMI but not ARI.
    • Complete LandUse-21: DIVIDE achieved 32.3/39.7/18.132.3/39.7/18.1, versus 26.2/32.7/13.526.2/32.7/13.5 from DCP.

    When the Scene-15 missing rate was varied from 0% to 90% in 10% increments, the page-6 performance plots showed DIVIDE above all compared baselines at every tested missing rate. The curves used five random experiments, with colored variability regions representing standard deviations.

  9. Knowl 9 — Random walks improve false-negative correction over local neighborhoods

    empirical result

    DIVIDE’s false-negative rectification was compared with kk-nearest-neighbor identification, ϵ\epsilon-neighborhood identification, and no identification. The reported values are clustering ACC and the KL divergence DKL(TGT∥T)D_{\mathrm{KL}}(T^{\mathrm{GT}}\|T) between the rectified target TT and an oracle target. For an in-batch anchor ii, the oracle target is TijGT=1/∣Pi∣T^{\mathrm{GT}}_{ij}=1/|P_i| when jj belongs to the same ground-truth class as ii, and 00 otherwise, where PiP_i is the set of in-batch samples sharing ii’s class. Lower KL divergence indicates more accurate false-negative correction and fewer incorrectly introduced false positives.

    • On Scene-15, random walks obtained ACC 49.149.1 and KL 8.28.2; kk-nearest neighbors obtained 48.548.5 and 22.322.3; ϵ\epsilon-neighborhood obtained 48.448.4 and 19.119.1; and no identification obtained 48.048.0 and 23.023.0.
    • On Caltech-101, random walks obtained ACC 62.262.2 and KL 5.15.1; kk-nearest neighbors obtained 57.057.0 and 17.217.2; ϵ\epsilon-neighborhood obtained 60.760.7 and 8.58.5; and no identification obtained 54.354.3 and 22.122.1.

    The page-6 walking-step plots further showed that higher-order walks outperform lower-order walks, remain effective for approximately 3–7 steps, and are relatively insensitive to the exact step count, whereas nearest-neighbor performance drops sharply when its neighborhood size is poorly selected. The unrectified case corresponds to t=0t=0 and T=InT=I_n.

  10. Knowl 10 — Self-and-swap target construction is the strongest target strategy

    empirical result

    On Scene-15, DIVIDE was evaluated with four ways of constructing the rectified target in the decoupled contrastive objective. The three reported clustering metrics are ACC/NMI/ARI.

    • Self and swap targets: 49.1/48.7/31.649.1/48.7/31.6.
    • Self targets only: 48.7/48.5/31.248.7/48.5/31.2.
    • Swap targets only: 48.7/47.9/30.848.7/47.9/30.8.
    • A target from an affinity graph built by concatenating all views: 48.1/48.1/30.548.1/48.1/30.5.

    The combined strategy performs best because the self target preserves view-specific and intra-view information, while the swap target improves interaction between the source and destination view spaces.

  11. Knowl 11 — Momentum targets, stop-gradient, decoder, and rectification all contribute

    empirical result

    Component ablations on Scene-15 show that the decoupled architecture and its optimization details are jointly important. With the cross-view decoder included, using a shared online encoder as the target without stop-gradient produced ACC/NMI/ARI 43.1/44.1/25.943.1/44.1/25.9; adding stop-gradient produced 47.6/45.7/28.247.6/45.7/28.2; and replacing the shared target with a momentum target while retaining stop-gradient produced the full 49.1/48.7/31.649.1/48.7/31.6. Without the cross-view decoder, the corresponding values were 44.0/44.3/26.744.0/44.3/26.7, 46.7/46.3/29.746.7/46.3/29.7, and 48.4/47.3/30.448.4/47.3/30.4.

    Here, shared means that the target encoder directly reuses online-encoder parameters, momentum means exponential-moving-average updating, and stop-gradient blocks target-branch gradients. The broader component ablation reported that adding the intra-view loss to the inter-view baseline increased ACC from 39.639.6 to 43.743.7 and NMI from 39.739.7 to 43.643.6; the decoder and rectified objective provided further gains, with the complete model reaching 49.1/48.7/31.649.1/48.7/31.6. The authors attribute approximately 3.7 ACC points and 2.9 NMI points to the intra-view loss, approximately 0.6 ACC points and 1.6 NMI points to the cross-view decoder, and approximately 1 ACC point to target rectification in the corresponding ablations.

Coverage note — No substantial contributed material was omitted; proof-free intermediate details, background, related work, acknowledgements, and references were excluded.

References

  1. 1.Amini, M. R.; Usunier, N.; and Goutte, C. 2009. Learning from multiple partially observed views-an application to multilingual text categorization. Advances in neural information processing systems, 22.
  2. 2.Chen, T.; Kornblith, S.; Norouzi, M.; and Hinton, G. 2020. A Simple Framework for Contrastive Learning of Visual Representations. In Proc. Int. Conf. Mach. Learn., 1597–1607.
  3. 3.Chuang, C.-Y.; Robinson, J.; Lin, Y.-C.; Torralba, A.; and Jegelka, S. 2020. Debiased contrastive learning. Advances in neural information processing systems, 33: 8765–8775.
  4. 4.Fang, S. 2023. Incomplete Multi-view Clustering via Diffusion Completion. arXiv preprint arXiv:2305.11489.
  5. 5.Han, Z.; Zhang, C.; Fu, H.; and Zhou, J. T. 2021. Trusted Multi-View Classification. In Proc. Int. Conf. Learn. Representations.
  6. 6.He, K.; Fan, H.; Wu, Y.; Xie, S.; and Girshick, R. B. 2020. Momentum Contrast for Unsupervised Visual Representation Learning. In Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 9726–9735.
  7. 7.Hu, M.; and Chen, S. 2018. Doubly Aligned Incomplete Multi-view Clustering. In Proc. Int. Joint Conf. Artif. Intell., 2262–2268.
  8. 8.Huang, Z.; Zhou, J. T.; Peng, X.; Zhang, C.; Zhu, H.; and Lv, J. 2019. Multi-view Spectral Clustering Network. In Proc. Int. Joint Conf. Artif. Intell., 2563–2569.
  9. 9.Huynh, T.; Kornblith, S.; Walter, M. R.; Maire, M.; and Khademi, M. 2022. Boosting contrastive self-supervised learning with false negative cancellation. In Proceedings of the IEEE/CVF winter conference on applications of computer vision, 2785–2795.
  10. 10.Jiang, Y.; Xu, Q.; Yang, Z.; Cao, X.; and Huang, Q. 2019. DM2C: Deep Mixed-Modal Clustering. In Proc. Int. Conf. Neural Inf. Process. Syst., 5880–5890.
  11. 11.Krizhevsky, A.; Sutskever, I.; and Hinton, G. E. 2012. Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems, 25.
  12. 12.Li, F.; and Perona, P. 2005. A bayesian hierarchical model for learning natural scene categories. In Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 524–531.
  13. 13.Li, H.; Li, Y.; Yang, M.; Hu, P.; Peng, D.; and Peng, X. 2023. Incomplete Multi-view Clustering via Prototype-based Imputation. arXiv preprint arXiv:2301.11045.
  14. 14.Li, S.-Y.; Jiang, Y.; and Zhou, Z.-H. 2014. Partial multi-view clustering. In Proceedings of the AAAI conference on artificial intelligence, volume 28.
  15. 15.Li, Y.; Nie, F.; Huang, H.; and Huang, J. 2015. Large-Scale Multi-View Spectral Clustering via Bipartite Graph. In Proc. AAAI Conf. Artif. Intell., 2750–2756.
  16. 16.Lin, F.; Bai, B.; Bai, K.; Ren, Y.; Zhao, P.; and Xu, Z. 2022a. Contrastive multi-view hyperbolic hierarchical clustering. arXiv preprint arXiv:2205.02618.
  17. 17.Lin, Y.; Gou, Y.; Liu, X.; Bai, J.; Lv, J.; and Peng, X. 2022b. Dual contrastive prediction for incomplete multi-view representation learning. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(4): 4447–4461.
  18. 18.Lin, Y.; Gou, Y.; Liu, Z.; Li, B.; Lv, J.; and Peng, X. 2021. COMPLETER: Incomplete Multi-view Clustering via Contrastive Prediction. In Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 11174–11183.
  19. 19.Lin, Y.; Yang, M.; Yu, J.; Hu, P.; Zhang, C.; and Peng, X. 2023. Graph Matching with Bi-level Noisy Correspondence. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV).
  20. 20.Liu, J.; Lin, Y.; Jiang, L.; Liu, J.; Wen, Z.; and Peng, X. 2022. Improve Interpretability of Neural Networks via Sparse Contrastive Coding. In Findings of the Association for Computational Linguistics: EMNLP 2022.
  21. 21.Liu, X.; Li, M.; Tang, C.; Xia, J.; Xiong, J.; Liu, L.; Kloft, M.; and Zhu, E. 2020. Efficient and Effective Regularized Incomplete Multi-view Clustering. IEEE Trans. Pattern Anal. Mach. Intell., 43(8): 2634–2646.
  22. 22.Lovasz, L. 1993. Random walks on graphs. Combinatorics, Paul erdos is eighty, 2(1-46): 4.
  23. 23.Simonyan, K.; and Zisserman, A. 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556.
  24. 24.Sun, W.; Zhang, J.; Wang, J.; Liu, Z.; Zhong, Y.; Feng, T.; Guo, Y.; Zhang, Y.; and Barnes, N. 2023. Learning Audio-Visual Source Localization via False Negative Aware Contrastive Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 6420–6429.
  25. 25.Tang, H.; and Liu, Y. 2022. Deep safe incomplete multi-view clustering: Theorem and algorithm. In International Conference on Machine Learning, 21090–21110. PMLR.
  26. 26.Trosten, D. J.; Lokse, S.; Jenssen, R.; and Kampffmeyer, M. 2021. Reconsidering representation alignment for multi-view clustering. In Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 1255–1265.
  27. 27.Trosten, D. J.; Løkse, S.; Jenssen, R.; and Kampffmeyer, M. C. 2023. On the Effects of Self-supervision and Contrastive Alignment in Deep Multi-view Clustering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 23976–23985.
  28. 28.Tsai, Y.-H. H.; Wu, Y.; Salakhutdinov, R.; and Morency, L.-P. 2021. Self-supervised learning from a multi-view perspective. In ICLR.
  29. 29.Wang, H.; Guo, X.; Deng, Z.-H.; and Lu, Y. 2022. Rethinking minimal sufficient representation in contrastive learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 16041–16050.
  30. 30.Wang, W.; Arora, R.; Livescu, K.; and Bilmes, J. A. 2015. On Deep Multi-View Representation Learning. In Proc. Int. Conf. Mach. Learn., 1083–1092.
  31. 31.Xu, C.; Guan, Z.; Zhao, W.; Wu, H.; Niu, Y.; and Ling, B. 2019. Adversarial incomplete multi-view clustering. In IJCAI, volume 7, 3933–3939.
  32. 32.Yadav, S. K.; Tiwari, K.; Pandey, H. M.; and Akbar, S. A. 2021. A review of multimodal human activity recognition with special emphasis on classification, applications, challenges and future directions. Knowledge-Based Systems, 223: 106970.
  33. 33.Yang, M.; Huang, Z.; Hu, P.; Li, T.; Lv, J.; and Peng, X. 2022a. Learning with Twin Noisy Labels for Visible-Infrared Person Re-Identification. In Proc. IEEE Conf. Comput. Vis. Pattern Recognit.
  34. 34.Yang, M.; Li, Y.; Hu, P.; Bai, J.; Lv, J. C.; and Peng, X. 2022b. Robust Multi-view Clustering with Incomplete Information. IEEE Trans. Pattern Anal. Mach. Intell.
  35. 35.Yang, Y.; and Newsam, S. 2010. Bag-of-visual-words and spatial extensions for land-use classification. In Proc. ACM SIGSPATIAL Int. Conf. Adv. Inf., 270–279.
  36. 36.Zhang, C.; Cui, Y.; Han, Z.; Zhou, J. T.; Fu, H.; and Hu, Q. 2020. Deep Partial Multi-View Learning. IEEE Trans. Pattern Anal. Mach. Intell.
  37. 37.Zhang, C.; Han, Z.; Cui, Y.; Fu, H.; Zhou, J. T.; and Hu, Q. 2019a. CPM-Nets: Cross Partial Multi-View Networks. In Proc. Int. Conf. Neural Inf. Process. Syst., 557–567.
  38. 38.Zhang, C.; Liu, Y.; and Fu, H. 2019. AE2-Nets: Autoencoder in Autoencoder Networks. In Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2577–2585.
  39. 39.Zhang, Z.; Liu, L.; Shen, F.; Shen, H.; and Shao, L. 2019b. Binary Multi-View Clustering. IEEE Trans. Pattern Anal. Mach. Intell., 41(7): 1774–1782.
  40. 40.Zheng, L.; Xiong, J.; Zhu, Y.; and He, J. 2022. Contrastive learning with complex heterogeneity. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2594–2604.
  41. 41.Zhong, H.; Wu, J.; Chen, C.; Huang, J.; Deng, M.; Nie, L.; Lin, Z.; and Hua, X.-S. 2021. Graph Contrastive Clustering. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 9224–9233.
  42. 42.Zolfaghari, M.; Zhu, Y.; Gehler, P.; and Brox, T. 2021. Crossclr: Cross-modal contrastive learning for multi-modal video representations. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 1450–1459.

Citation

MLA
Lu, Y., et al. “Decoupled Contrastive Multi-View Clustering with High-Order Random Walks”. arXiv, 2023, http://arxiv.org/abs/2308.11164v2.
APA
Lu, Y., Lin, Y., Yang, M., Peng, D., Hu, P., & Peng, X. (2023). Decoupled Contrastive Multi-View Clustering with High-Order Random Walks. arXiv. http://arxiv.org/abs/2308.11164v2
Chicago
Lu, Y., Y. Lin, M. Yang, D. Peng, P. Hu, and X. Peng. 2023. “Decoupled Contrastive Multi-View Clustering with High-Order Random Walks”. arXiv. http://arxiv.org/abs/2308.11164v2.
Harvard
Lu, Y. et al. (2023) “Decoupled Contrastive Multi-View Clustering with High-Order Random Walks”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2308.11164v2.
Vancouver
1. Lu Y, Lin Y, Yang M, Peng D, Hu P, Peng X (2023) Decoupled Contrastive Multi-View Clustering with High-Order Random Walks. arXiv

BibTeX

@article{lu2023decoupled,
  title = {Decoupled Contrastive Multi-View Clustering with High-Order Random Walks},
  author = {Lu, Yiding and Lin, Yijie and Yang, Mouxing and Peng, Dezhong and Hu, Peng and Peng, Xi},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2308.11164v2},
  eprint = {2308.11164}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF