Decoupled Contrastive Multi-View Clustering with High-Order Random Walks
Yiding LuYijie LinMouxing YangDezhong PengPeng HuXi Peng
Proposes a decoupled multi-view clustering framework that uses high-order random walks to rectify false positive and false negative pairs globally while preserving view-specific information through cross-view reconstruction.
Modern data systems increasingly rely on multi-view clustering to automatically group complex, unlabeled information collected across multiple perspectives or modalities, such as multi-camera security feeds, multi-lingual texts, and medical imaging. Unsupervised contrastive learning has emerged as a leading technique to group these multi-view datasets by pulling representations of matching instances closer while pushing mismatched pairs apart. However, existing contrastive techniques routinely misidentify similar data points as negative pairs (false negatives) or, when attempting local corrections via immediate neighborhood checks, mistakenly label dissimilar points as positive pairs (false positives). Furthermore, conventional architectures often discard view-specific unique details while trying to enforce consistency, severely degrading clustering accuracy when data is incomplete or corrupted with missing views.
The article evaluates and demonstrates a new framework called DIVIDE (Decoupled Contrastive Multi-View Clustering with High-Order Random Walks). The primary objective of the article is to show how DIVIDE simultaneously eliminates false pairing errors across global graph connections and preserves view-specific information, thereby delivering superior clustering performance and robust recovery for both complete and incomplete multi-view data.
To evaluate this framework, the authors conducted extensive empirical experiments across four benchmark datasets spanning 2,100 to 18,758 instances: Scene-15, Caltech-101, Reuters, and LandUse-21. The approach relies on two core innovations: a multi-step random walk mechanism that traces probabilistic graph transitions to identify true global relationships, and a decoupled architecture combining Siamese encoders with cross-view decoders to handle intra-view and inter-view contrastive objectives separately. The researchers benchmarked DIVIDE against nine state-of-the-art multi-view clustering methods across fully observed datasets and partially observed environments where up to 90% of views were missing.
The article establishes several key findings. First, DIVIDE consistently outperformed all nine competing methods across every tested benchmark; for example, under a 50% missing-view setting on the Caltech-101 benchmark, DIVIDE reached a clustering accuracy of 63.4%, surpassing the next best competing baseline (56.2%) by over 7 percentage points. Second, the framework proved exceptionally resilient to extreme data loss, maintaining performance advantages on the Scene-15 dataset even as the view missing rate rose from 0% to 90%. Third, the multi-step random walk strategy generated significantly lower target error (achieving Kullback-Leibler divergence scores of 8.2 and 5.1 on Scene-15 and Caltech-101) compared to standard nearest-neighbor approaches (which scored 22.3 and 17.2), confirming its ability to filter false positives and recover false negatives without requiring fragile parameter tuning across 3 to 7 walk steps. Finally, ablation studies showed that every architectural component—including intra-view contrastive loss, the cross-view decoder, momentum updates, and target rectification—provided measurable accuracy and consistency gains.
These findings indicate that multi-view clustering systems can achieve high accuracy without sacrificing unique, modality-specific details or failing in the presence of missing inputs. For technical leadership and operational stakeholders, this means unsupervised machine learning pipelines can be deployed in complex, real-world data environments with reduced data curation costs, lower computational sensitivity to hyperparameter tuning, and stronger resilience against missing sensor feeds or data streams. The decoupled decoder design also eliminates the need to develop separate, complex imputation models to artificially complete missing data prior to clustering.
Based on these results, technical teams implementing unsupervised representation learning or multi-view analytics should adopt high-order graph propagation techniques over rigid local neighborhood thresholds to manage sample pair assignments. Organizations dealing with missing data streams should also leverage cross-view reconstruction decoders to preserve modality-specific signals. For future work, development teams should pilot this architecture within their domain-specific pipelines and explore adapting DIVIDE's decoupled random-walk approach to broader unsupervised contrastive representation learning tasks.
The conclusions of the article carry high confidence due to consistent, validated empirical results across multiple data modalities and comparative ablation tests. However, readers should note that the current evaluations primarily reflect two-view scenarios and balanced benchmark settings. Further validation is warranted before deploying the architecture in operational environments featuring highly imbalanced clusters or three or more simultaneous data views.
- Paper: GCFAgg: Global and Cross-View Feature Aggregation for Multi-View Clustering, Weiqing Yan et al. (2023). It analyzes the negative-pair conflict and neighborhood alignment challenges in multi-view contrastive clustering that DIVIDE directly aims to resolve.
- Paper: Contrastive Multiview Coding, Yonglong Tian et al. (2019). It provides the foundational framework for contrastive multi-view representation learning that underpins modern multi-view clustering objectives.
- Paper: Co-regularized Multi-view Spectral Clustering, Abhishek Kumar et al. (2011). It establishes classic multi-view spectral graph clustering principles and inter-view agreement formulations adapted by deep multi-view clustering methods.
- Paper: A tutorial on spectral clustering, Ulrike von Luxburg (2007). It details the foundational theory connecting graph Laplacians, random walks, and spectral clustering partitions utilized in graph-based clustering.
- Paper: Higher-order organization of complex networks, Austin R. Benson et al. (2016). It introduces higher-order graph structures and motifs for network partitioning, providing core motivation for DIVIDE's higher-order random walk formulation.
- Paper: Unsupervised Deep Embedding for Clustering Analysis, Junyuan Xie et al. (2015). It introduces deep embedded clustering, laying the baseline framework for jointly optimizing neural embeddings and cluster assignments.
No sufficiently relevant recommendations were found.
