FOCAL: Contrastive Learning for Multimodal Time-Series Sensing Signals in Factorized Orthogonal Latent Space
Shengzhong LiuTomoyoshi KimuraDongxin LiuRuijie WangJinyang LiSuhas N. DiggaviMani B. SrivastavaTarek F. Abdelzaher
Proposes a self-supervised contrastive learning framework that factorizes multimodal time-series signals into orthogonal shared and private latent spaces alongside statistical temporal constraints to achieve state-of-the-art representation quality across diverse sensing datasets.
Modern sensing and Internet of Things systems rely on multiple sensor streams, such as acoustic, seismic, and motion readings, to perceive physical environments. Training effective artificial intelligence models across these sensors typically requires vast amounts of human-labeled data, which is expensive and time-consuming to obtain. While self-supervised contrastive learning enables models to learn representations from unlabeled data, existing methods face two critical flaws: they focus almost exclusively on shared information across sensors while discarding unique, modality-exclusive signals, and they enforce rigid temporal constraints that fail to accommodate periodic or cyclical physical phenomena.
The article develops and evaluates FOCAL, a self-supervised contrastive learning framework designed to extract comprehensive features from multimodal time-series signals. The main objective is to demonstrate that factorizing the representation space into distinct shared and private features, paired with a relaxed temporal constraint, substantially improves classification accuracy and data efficiency in sensing applications.
The authors evaluated the framework across four multimodal datasets: vehicle classification using acoustic and seismic data, ground vehicle identification across varied terrains, and two wearable human activity recognition benchmarks. Sensor inputs were converted into time-frequency representations and processed using two distinct backbone neural network architectures (DeepSense and Swin-Transformer). FOCAL projects sensor representations into a factorized space containing shared features across sensors and private features unique to each sensor, enforcing geometric orthogonality between them. In addition, the method introduces a loose sequence-level distance constraint ensuring that temporally adjacent samples are on average closer than distant samples, accommodating long-term periodic patterns without hard pair-matching.
The evaluation yielded several key findings. First, FOCAL consistently outperformed eleven state-of-the-art baseline frameworks across all four datasets, beating standard multiview benchmarks such as Contrastive Multiview Coding by 4.48% to 18.01% in classification accuracy. Second, the framework exhibited exceptional label efficiency: when downstream models were provided with only 1% of available labeled training data, FOCAL maintained high performance, achieving a 10.56% relative improvement over the strongest competing baseline. Third, ablation analyses confirmed that isolating private modality features, maintaining orthogonality, and applying the loose temporal constraint each provided distinct performance gains, with the removal of private features causing an accuracy drop of 5.20% to 5.90%. Finally, the temporal constraint demonstrated general plug-and-play applicability, lifting the accuracy of existing baseline frameworks by up to 18.99%.
These findings indicate that multimodal sensor systems can achieve high operational performance with minimal manual labeling, significantly reducing deployment costs, annotation timelines, and engineering risks for complex Internet of Things environments. Retaining sensor-exclusive physical signals rather than discarding them allows models to better separate classes and adapt to downstream recognition tasks, including vehicle speed and distance estimation across new domains.
Organizations developing multisensor intelligence systems should adopt factorized representation techniques and sequence-level temporal constraints rather than relying solely on shared cross-modal matching. Prior to enterprise deployment, teams should conduct domain-specific evaluations to address operational boundaries. The source highlights key areas for next-step development, notably incorporating signal synchronization mechanisms for modalities with varying propagation speeds (such as sound and light), reducing the computational complexity of evaluating all modality pairs, and building domain-invariant features to mitigate performance degradation under environmental shifts such as changing terrain and weather.
- Paper: Contrastive Multiview Coding, Yonglong Tian et al. (2019). It establishes the foundational multiview contrastive learning paradigm that maximizes shared cross-modal information, serving as the direct baseline and conceptual starting point that FOCAL critiques and extends with factorized private features.
- Paper: MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment Analysis, Devamanyu Hazarika et al. (2020). It introduces the core architectural concept of factorizing multimodal inputs into explicit shared and private subspaces constrained by orthogonality to preserve modality-specific signals.
- Paper: TS2Vec: Towards Universal Representation of Time Series, Zhihan Yue et al. (2022). It provides key insights into unsupervised contrastive representation learning over time-series sequences across multiscale temporal contexts.
- Paper: Representation Learning with Contrastive Predictive Coding, Aäron van den Oord et al. (2018). It outlines the foundational principles of contrastive predictive coding and temporal sequence modeling in latent spaces.
- Paper: A Simple Framework for Contrastive Learning of Visual Representations, Ting Chen et al. (2020). It details the standard framework and projection head mechanics of self-supervised contrastive learning that underlie modern multimodal representation pipelines.
- Paper: Multimodal Machine Learning: A Survey and Taxonomy, Tadas Baltrušaitis et al. (2017). It supplies the fundamental taxonomy and foundational challenges of multimodal alignment, fusion, and joint representation learning.
- Paper: Multi-View Causal Representation Learning with Partial Observability, Dingling Yao et al. (2024). It formalizes a causal identifiability theory for isolating shared versus view-specific latent factors under partial observability across multiple non-linear sensor modalities.
- Paper: ReconBoost: Boosting Can Achieve Modality Reconcilement, Cong Hua et al. (2024). It extends the challenge of balancing shared cross-modal interactions and private unimodal features by introducing a boosting framework to prevent dominant modality suppression.
- Paper: ImageBind One Embedding Space to Bind Them All, Rohit Girdhar et al. (2023). It scales joint multimodal contrastive alignment across six diverse sensory domains including IMU motion and audio time-series.
