Built independently by an author, for readers. Read the story and support ChapterPal

keyword

contrastive self-supervised representation learning

Contrastive self-supervised representation learning is a machine learning paradigm that trains models to extract meaningful feature representations from unlabeled data by mapping similar samples close to one another in an embedding space while pushing dissimilar samples apart. Rather than relying on human annotations, the training process automatically creates supervisory signals by defining pairs of related inputs, such as different augmented views of the same instance, synchronized multimodal streams, or temporally adjacent observations, and contrasting them against unrelated negative examples. Through this contrastive objective, the model discovers intrinsic structures, invariances, and distinguishing characteristics within the data, producing rich, generalizable representations that enhance performance across diverse downstream tasks.

1 item

FOCAL: Contrastive Learning for Multimodal Time-Series Sensing Signals in Factorized Orthogonal Latent Space

FOCAL: Contrastive Learning for Multimodal Time-Series Sensing Signals in Factorized Orthogonal Latent Space

Shengzhong Liu, Tomoyoshi Kimura, Dongxin Liu, Ruijie Wang, Jinyang Li, Suhas N. Diggavi, Mani B. Srivastava, Tarek F. Abdelzaher

OrganizationsMetaShanghai Jiao Tong UniversityUniversity of California, Los AngelesUniversity of Illinois Urbana-Champaign

Why you should read this

Proposes a self-supervised contrastive learning framework that factorizes multimodal time-series signals into orthogonal shared and private latent spaces alongside statistical temporal constraints to achieve state-of-the-art representation quality across diverse sensing datasets.

This paper proposes a novel contrastive learning framework, called FOCAL, for extracting comprehensive features from multimodal time-series sensing signals through self-supervised training. Existing multimodal contrastive frameworks mostly rely on the shared information between sensory modalities, but do not explicitly consider the exclusive modality information that could be critical to understanding the underlying sensing physics. Besides, contrastive frameworks for time series have not handled the temporal information locality appropriately. FOCAL solves these challenges by making the following contributions: First, given multimodal time series, it encodes each modality into a factorized latent space consisting of shared features and private features that are orthogonal to each other. The shared space emphasizes feature patterns consistent across sensory modalities through a modal-matching objective. In contrast, the private space extracts modality-exclusive information through a transformation-invariant objective. Second, we propose a temporal structural constraint for modality features, such that the average distance between temporally neighboring samples is no larger than that of temporally distant samples. Extensive evaluations are performed on four multimodal sensing datasets with two backbone encoders and two classifiers to demonstrate the superiority of FOCAL. It consistently outperforms the state-of-the-art baselines in downstream tasks with a clear margin, under different ratios of available labels. The code and self-collected dataset are available at https://github.com/tomoyoshki/focal.

Added

2026-09-26