keyword
canonical correlation analysis
Canonical correlation analysis is a multivariate statistical method used to identify and measure the linear relationships between two multidimensional sets of variables. Given two distinct sets of features observed on the same entities, the technique seeks optimal linear combinations from each set, termed canonical variates, that maximize the correlation between the two projections. Successive pairs of canonical variates are derived to capture the highest remaining correlation while remaining uncorrelated with previously identified pairs. Widely applied in multivariate statistics, data science, and machine learning, canonical correlation analysis provides a foundational framework for dimensionality reduction, cross-modal analysis, and multi-view representation learning by projecting disparate data sources into a shared latent space.
2 items

Co-regularized Multi-view Spectral Clustering
Abhishek Kumar, Piyush Rai, Hal Daumé
Why you should read this
Proposes a multi-view spectral clustering framework that co-regularizes graph Laplacians across different data representations to find consistent cluster assignments across diverse views.
In many clustering problems, we have access to multiple views of the data each of which could be individually used for clustering. Exploiting information from multiple views, one can hope to find a clustering that is more accurate than the ones obtained using the individual views. Often these different views admit same underlying clustering of the data, so we can approach this problem by looking for clusterings that are consistent across the views, i.e., corresponding data points in each view should have same cluster membership. We propose a spectral clustering framework that achieves this goal by co-regularizing the clustering hypotheses, and propose two co-regularization schemes to accomplish this. Experimental comparisons with a number of baselines on two synthetic and three real-world datasets establish the efficacy of our proposed approaches.
Added
2026-09-25

Connecting Modalities: Semi-supervised Segmentation and Annotation of Images Using Unaligned Text Corpora
Richard Socher, Li Fei-Fei
Why you should read this
Demonstrates how to train image segmentation and annotation systems using only a handful of labeled images and freely available news articles by discovering that visual regions and text words follow similar contextual patterns, enabling efficient learning without massive labeled datasets.
We propose a semi-supervised model which segments and annotates images using very few labeled images and a large unaligned text corpus to relate image regions to text labels. Given photos of a sports event, all that is necessary to provide a pixel-level labeling of objects and background is a set of newspaper articles about this sport and one to five labeled images. Our model is motivated by the observation that words in text corpora share certain context and feature similarities with visual objects. We describe images using visual words, a new region-based representation. The proposed model is based on kernelized canonical correlation analysis which finds a mapping between visual and textual words by projecting them into a latent meaning space. Kernels are derived from context and adjective features inside the respective visual and textual domains. We apply our method to a challenging dataset and rely on articles of the New York Times for textual features. Our model outperforms the state-of-the-art in annotation. In segmentation it compares favorably with other methods that use significantly more labeled training data.
Added
2026-02-21
