FOCAL: Contrastive Learning for Multimodal Time-Series Sensing Signals in Factorized Orthogonal Latent Space

Shengzhong LiuTomoyoshi KimuraDongxin LiuRuijie WangJinyang LiSuhas N. DiggaviMani B. SrivastavaTarek F. Abdelzaher

article2023NeurIPS54 citations

Proposes a self-supervised contrastive learning framework that factorizes multimodal time-series signals into orthogonal shared and private latent spaces alongside statistical temporal constraints to achieve state-of-the-art representation quality across diverse sensing datasets.

Listen

Modern sensing and Internet of Things systems rely on multiple sensor streams, such as acoustic, seismic, and motion readings, to perceive physical environments. Training effective artificial intelligence models across these sensors typically requires vast amounts of human-labeled data, which is expensive and time-consuming to obtain. While self-supervised contrastive learning enables models to learn representations from unlabeled data, existing methods face two critical flaws: they focus almost exclusively on shared information across sensors while discarding unique, modality-exclusive signals, and they enforce rigid temporal constraints that fail to accommodate periodic or cyclical physical phenomena.

The article develops and evaluates FOCAL, a self-supervised contrastive learning framework designed to extract comprehensive features from multimodal time-series signals. The main objective is to demonstrate that factorizing the representation space into distinct shared and private features, paired with a relaxed temporal constraint, substantially improves classification accuracy and data efficiency in sensing applications.

The authors evaluated the framework across four multimodal datasets: vehicle classification using acoustic and seismic data, ground vehicle identification across varied terrains, and two wearable human activity recognition benchmarks. Sensor inputs were converted into time-frequency representations and processed using two distinct backbone neural network architectures (DeepSense and Swin-Transformer). FOCAL projects sensor representations into a factorized space containing shared features across sensors and private features unique to each sensor, enforcing geometric orthogonality between them. In addition, the method introduces a loose sequence-level distance constraint ensuring that temporally adjacent samples are on average closer than distant samples, accommodating long-term periodic patterns without hard pair-matching.

The evaluation yielded several key findings. First, FOCAL consistently outperformed eleven state-of-the-art baseline frameworks across all four datasets, beating standard multiview benchmarks such as Contrastive Multiview Coding by 4.48% to 18.01% in classification accuracy. Second, the framework exhibited exceptional label efficiency: when downstream models were provided with only 1% of available labeled training data, FOCAL maintained high performance, achieving a 10.56% relative improvement over the strongest competing baseline. Third, ablation analyses confirmed that isolating private modality features, maintaining orthogonality, and applying the loose temporal constraint each provided distinct performance gains, with the removal of private features causing an accuracy drop of 5.20% to 5.90%. Finally, the temporal constraint demonstrated general plug-and-play applicability, lifting the accuracy of existing baseline frameworks by up to 18.99%.

These findings indicate that multimodal sensor systems can achieve high operational performance with minimal manual labeling, significantly reducing deployment costs, annotation timelines, and engineering risks for complex Internet of Things environments. Retaining sensor-exclusive physical signals rather than discarding them allows models to better separate classes and adapt to downstream recognition tasks, including vehicle speed and distance estimation across new domains.

Organizations developing multisensor intelligence systems should adopt factorized representation techniques and sequence-level temporal constraints rather than relying solely on shared cross-modal matching. Prior to enterprise deployment, teams should conduct domain-specific evaluations to address operational boundaries. The source highlights key areas for next-step development, notably incorporating signal synchronization mechanisms for modalities with varying propagation speeds (such as sound and light), reducing the computational complexity of evaluating all modality pairs, and building domain-invariant features to mitigate performance degradation under environmental shifts such as changing terrain and weather.

Cover for FOCAL: Contrastive Learning for Multimodal Time-Series Sensing Signals in Factorized Orthogonal Latent Space

Abstract

This paper proposes a novel contrastive learning framework, called FOCAL, for extracting comprehensive features from multimodal time-series sensing signals through self-supervised training. Existing multimodal contrastive frameworks mostly rely on the shared information between sensory modalities, but do not explicitly consider the exclusive modality information that could be critical to understanding the underlying sensing physics. Besides, contrastive frameworks for time series have not handled the temporal information locality appropriately. FOCAL solves these challenges by making the following contributions: First, given multimodal time series, it encodes each modality into a factorized latent space consisting of shared features and private features that are orthogonal to each other. The shared space emphasizes feature patterns consistent across sensory modalities through a modal-matching objective. In contrast, the private space extracts modality-exclusive information through a transformation-invariant objective. Second, we propose a temporal structural constraint for modality features, such that the average distance between temporally neighboring samples is no larger than that of temporally distant samples. Extensive evaluations are performed on four multimodal sensing datasets with two backbone encoders and two classifiers to demonstrate the superiority of FOCAL. It consistently outperforms the state-of-the-art baselines in downstream tasks with a clear margin, under different ratios of available labels. The code and self-collected dataset are available at https://github.com/tomoyoshki/focal.

Table of Contents

  • 1 Introduction
  • 2 Problem Formulation
  • 3 FOCAL Framework
  • 3.1 Overview
  • 3.2 Multimodal Contrastive Learning in Factorized Orthogonal Space
  • 3.3 Temporal Structural Constraint
  • 3.4 Overall Training Objective
  • 4 Evaluation
  • 4.1 Experimental Setup
  • 4.2 Finetune Results
  • 4.3 Additional Downstream Tasks
  • 4.4 Ablation Studies
  • 4.5 General Applicability of the Temporal Constraint
  • 4.6 Sensitivity Test on Loss Weights
  • 5 Related Works
  • 6 Conclusion
  • Acknowledgments and Disclosure of Funding
  • References
  • Appendix
  • A Datasets
  • B Data Preprocessing
  • B.1 Data Augmentations
  • B.1.1 Time-Domain Augmentations
  • B.1.2 Frequency-Domain Augmentations
  • C Baselines
  • D Backbone Models
  • E Training Configurations
  • F Additional Evaluation Results
  • F.1 Finetuning: Complete Linear Classification Results
  • F.2 Finetuning: Complete KNN Classification Results
  • F.3 Complete Clustering Results
  • F.4 Complete Additional Downstream Task Results
  • F.5 Ablation Study Results
  • G Limitations and Potential Extensions
  • References

Knowls

  1. Knowl 1 — FOCAL Framework for Factorized Orthogonal Latent Multimodal Representation Learning

    model/method

    FOCAL is a self-supervised contrastive learning framework designed for multimodal time-series sensing signals. Given PP heterogeneous sensory modalities M={M1,M2,…,MP}\mathcal{M} = \{M_1, M_2, \dots, M_P\}, the multivariate time-series input from each modality is preprocessed using Short-Time Fourier Transform (STFT) into a time-frequency spectrogram xij∈RC×I×S\mathbf{x}_{ij} \in \mathbb{R}^{C \times I \times S}, where CC is the number of input channels, II is the number of time intervals in the sample window, and SS is the Fourier spectrum length.

    Each modality is processed by a dedicated backbone encoder EjE_j, yielding a latent embedding vector hij=Ej(xij)∈RK\mathbf{h}_{ij} = E_j(\mathbf{x}_{ij}) \in \mathbb{R}^K. To extract both cross-modal correlations and modality-exclusive physics, each modality representation hij\mathbf{h}_{ij} is factorized via multilayer perceptron (MLP) projection heads into two distinct components:

    1. Shared Feature Space (hijshared\mathbf{h}_{ij}^{\text{shared}}): Captures semantic information shared and consistent across collaborating modalities through a cross-modal matching contrastive objective.
    2. Private Feature Space (hijprivate\mathbf{h}_{ij}^{\text{private}}): Captures modality-exclusive, sensor-specific discriminative patterns through transformation invariance under stochastic data augmentations.

    To ensure that the private space does not redundantly encode shared information and that private features across different modalities remain statistically independent, FOCAL enforces geometric orthogonality constraints between the shared and private features of the same modality, as well as between the private features of different modalities.

  2. Knowl 2 — Contrastive Loss Functions and Pretraining Objectives in FOCAL

    equation

    Let B\mathcal{B} denote a mini-batch of multimodal samples, and τ>0\tau > 0 denote the contrastive temperature parameter. For a sample ii and modality Mj∈MM_j \in \mathcal{M}, the FOCAL pretraining objective optimizes four complementary loss terms:

    1. Shared Space InfoNCE Loss (Lshared\mathcal{L}_{\text{shared}}): Maximizes agreement between shared features hijshared\mathbf{h}_{ij}^{\text{shared}} and hij′shared\mathbf{h}_{ij'}^{\text{shared}} from different modalities j≠j′j \neq j' of the same sample ii, while contrasting against all other samples i′∈Bi' \in \mathcal{B}: Lshared=−∑i∑Mj,Mj′∈M,j≠j′log⁡exp⁡(⟨hijshared,hij′shared⟩/τ)∑i′∈Bexp⁡(⟨hijshared,hi′j′shared⟩/τ)\mathcal{L}_{\text{shared}} = -\sum_{i} \sum_{M_j, M_{j'} \in \mathcal{M}, j \neq j'} \log \frac{\exp(\langle \mathbf{h}_{ij}^{\text{shared}}, \mathbf{h}_{ij'}^{\text{shared}} \rangle / \tau)}{\sum_{i' \in \mathcal{B}} \exp(\langle \mathbf{h}_{ij}^{\text{shared}}, \mathbf{h}_{i'j'}^{\text{shared}} \rangle / \tau)}

    2. Private Space NT-Xent Loss (Lprivate\mathcal{L}_{\text{private}}): Enforces transformation consistency by matching the private embedding hijprivate\mathbf{h}_{ij}^{\text{private}} of sample ii under one augmentation with its representation h~ijprivate\tilde{\mathbf{h}}_{ij}^{\text{private}} under an alternate random augmentation, treating all other 2∣B∣−22|\mathcal{B}| - 2 representations as negatives: Lprivate=−∑i∑Mj∈Mlog⁡exp⁡(⟨hijprivate,h~ijprivate⟩/τ)∑i′∈B,i′≠iexp⁡(⟨hijprivate,hi′jprivate⟩/τ)+∑i′∈Bexp⁡(⟨hijprivate,h~i′jprivate⟩/τ)\mathcal{L}_{\text{private}} = -\sum_{i} \sum_{M_j \in \mathcal{M}} \log \frac{\exp(\langle \mathbf{h}_{ij}^{\text{private}}, \tilde{\mathbf{h}}_{ij}^{\text{private}} \rangle / \tau)}{\sum_{i' \in \mathcal{B}, i' \neq i} \exp(\langle \mathbf{h}_{ij}^{\text{private}}, \mathbf{h}_{i'j}^{\text{private}} \rangle / \tau) + \sum_{i' \in \mathcal{B}} \exp(\langle \mathbf{h}_{ij}^{\text{private}}, \tilde{\mathbf{h}}_{i'j}^{\text{private}} \rangle / \tau)}

    3. Orthogonality Loss (Lorthogonal\mathcal{L}_{\text{orthogonal}}): Minimizes cosine similarity between intra-modality shared-private pairs and inter-modality private-private pairs: Lorthogonal=∑i∑Mj∈M⟨hijshared,hijprivate⟩+∑i∑Mj,Mj′∈M,j≠j′⟨hijprivate,hij′private⟩\mathcal{L}_{\text{orthogonal}} = \sum_{i} \sum_{M_j \in \mathcal{M}} \langle \mathbf{h}_{ij}^{\text{shared}}, \mathbf{h}_{ij}^{\text{private}} \rangle + \sum_{i} \sum_{M_j, M_{j'} \in \mathcal{M}, j \neq j'} \langle \mathbf{h}_{ij}^{\text{private}}, \mathbf{h}_{ij'}^{\text{private}} \rangle

    4. Overall Training Objective: L=Lshared+λp⋅Lprivate+λo⋅Lorthogonal+λt⋅Ltemporal\mathcal{L} = \mathcal{L}_{\text{shared}} + \lambda_p \cdot \mathcal{L}_{\text{private}} + \lambda_o \cdot \mathcal{L}_{\text{orthogonal}} + \lambda_t \cdot \mathcal{L}_{\text{temporal}} where λp,λo,λt\lambda_p, \lambda_o, \lambda_t are weighting hyperparameters controlling the contribution of private contrastive learning, orthogonality regularization, and temporal structural constraints.

  3. Knowl 3 — Temporal Structural Constraint for Time-Series Locality Regularization

    model/method

    Instead of treating temporally proximate samples as strict positive pairs with binary similarity 1 (which fails under long-term periodicity or continuous physical transitions), FOCAL enforces a coarse-grained sequence-level distance ranking constraint on the holistic modality embeddings hij\mathbf{h}_{ij}.

    A training mini-batch B\mathcal{B} is constructed by sampling BB distinct short sequences, each consisting of LL temporally contiguous window samples (yielding a total batch size of B×LB \times L). For sample-level Euclidean distances Dpq=∥hp−hq∥2D_{pq} = \|\mathbf{h}_p - \mathbf{h}_q\|_2, the average distance between sequence ss and sequence s′s' is defined as: Dˉss′=1L2∑i∈s,i′∈s′Dii′\bar{D}_{ss'} = \frac{1}{L^2} \sum_{i \in s, i' \in s'} D_{ii'} where diagonal elements with distance 0 are omitted when s=s′s = s'.

    The temporal locality loss regularizes the latent space such that the average intra-sequence distance Dˉss\bar{D}_{ss} within a sequence ss is strictly smaller than the inter-sequence distance Dˉss′\bar{D}_{ss'} between sequence ss and any distinct sequence s′s' by a predefined margin: Ltemporal=∑s∑s′≠smax⁡(Dˉss−Dˉss′+margin,0)\mathcal{L}_{\text{temporal}} = \sum_{s} \sum_{s' \neq s} \max(\bar{D}_{ss} - \bar{D}_{ss'} + \text{margin}, 0)

    This loose ranking prevents model collapse on periodic signals while maintaining global temporal coherence and eliminating the O((BL)3)O((BL)^3) complexity of traversing sample triplets.

  4. Knowl 4 — Correlated Accuracy Metric for Ordinal Physical Downstream Tasks

    equation

    For downstream tasks where class categories follow an ordinal physical continuum (such as speed estimation in mph or distance binning), standard classification accuracy penalizes all misclassifications uniformly. The correlated accuracy metric (extcorr_acc ext{corr\_acc}) assigns graded penalties proportional to the numerical distance between the predicted label and the ground-truth label.

    Given N′N' evaluation samples with ground-truth class indices yis∈{0,1,…,C−1}y_i^s \in \{0, 1, \dots, C-1\} and predicted class indices yi∈{0,1,…,C−1}y_i \in \{0, 1, \dots, C-1\} among CC ordered classes, the correlated accuracy is defined as: corr_acc=1N′∑i=1N′(1−∣yi−yis∣max⁡(yis,C−1−yis))\text{corr\_acc} = \frac{1}{N'} \sum_{i=1}^{N'} \left( 1 - \frac{|y_i - y_i^s|}{\max(y_i^s, C - 1 - y_i^s)} \right)

    where max⁡(yis,C−1−yis)\max(y_i^s, C - 1 - y_i^s) represents the maximum possible integer distance from class yisy_i^s to any boundary class. The value lies in [0,1][0, 1], where 1 indicates perfect predictions and 0 indicates maximally distant errors across all test samples.

  5. Knowl 5 — Linear Probing Performance on Multimodal Time-Series Benchmarks

    data/table

    The linear probing classification results (accuracy and macro F1 score) across four sensing datasets—MOD (acoustic + seismic, 7 classes), ACIDS (acoustic + seismic, 9 classes), RealWorld-HAR (accelerometer, gyroscope, magnetometer, light, 8 classes), and PAMAP2 (accelerometer, gyroscope, magnetometer, 18 classes)—demonstrate the effectiveness of FOCAL. Two backbone encoders were tested: DeepSense and Swin-Transformer (SW-T).

    Framework MOD ACIDS RealWorld-HAR PAMAP2
    Acc F1 Acc F1 Acc F1 Acc F1
    DeepSense Backbone
    Supervised 0.9404 0.9399 0.9566 0.8407 0.9348 0.9388 0.8849 0.8761
    SimCLR 0.8855 0.8855 0.7438 0.6101 0.7138 0.6841 0.6802 0.6583
    MoCo 0.8808 0.8812 0.7717 0.6205 0.7859 0.7708 0.7559 0.7387
    CMC 0.9196 0.9186 0.8443 0.7244 0.7975 0.8116 0.7906 0.7706
    MAE 0.5981 0.5993 0.6644 0.5618 0.7565 0.7515 0.7114 0.6158
    Cosmo 0.8989 0.8998 0.8511 0.6929 0.8956 0.8888 0.8356 0.8135
    Cocoa 0.8774 0.8764 0.6644 0.5359 0.8465 0.8488 0.7603 0.7187
    MTSS 0.4153 0.3582 0.4352 0.2441 0.2989 0.1405 0.3541 0.1795
    TS2Vec 0.7669 0.7648 0.5224 0.3587 0.6595 0.5984 0.5729 0.4715
    GMC 0.9257 0.9267 0.9096 0.7929 0.8869 0.8948 0.8119 0.7860
    TNC 0.9518 0.9528 0.8237 0.6936 0.8892 0.8971 0.8387 0.8143
    TS-TCC 0.8707 0.8735 0.7667 0.6164 0.8073 0.8010 0.7776 0.7250
    FOCAL 0.9732 0.9729 0.9516 0.8580 0.9382 0.9290 0.8588 0.8463
    SW-T Backbone
    Supervised 0.8948 0.8931 0.9137 0.7770 0.9313 0.9278 0.8612 0.8384
    SimCLR 0.9250 0.9247 0.9128 0.8144 0.7046 0.7220 0.7705 0.7424
    MoCo 0.9390 0.9384 0.9174 0.8100 0.7813 0.8024 0.7717 0.7313
    CMC 0.9129 0.9105 0.8128 0.6857 0.8840 0.8955 0.8080 0.7901
    MAE 0.7803 0.7772 0.8516 0.7023 0.8829 0.8813 0.7910 0.7606
    Cosmo 0.3429 0.3378 0.7110 0.6086 0.8604 0.8169 0.7741 0.7366
    Cocoa 0.7040 0.7038 0.7096 0.5794 0.8892 0.8861 0.7689 0.7317
    MTSS 0.4206 0.4163 0.3429 0.2250 0.5136 0.4370 0.2847 0.1714
    TS2Vec 0.7254 0.7174 0.7183 0.5748 0.6151 0.5955 0.6195 0.5426
    GMC 0.8640 0.8611 0.9402 0.7766 0.9319 0.9379 0.8312 0.8083
    TNC 0.8533 0.8539 0.8352 0.7372 0.8817 0.8784 0.8013 0.7506
    TS-TCC 0.8734 0.8735 0.9041 0.7547 0.8731 0.8454 0.7997 0.7260
    FOCAL 0.9805 0.9800 0.9489 0.8262 0.9451 0.9503 0.8580 0.8401

    FOCAL consistently outperforms 11 self-supervised baselines. On the MOD dataset, FOCAL exceeds fully supervised performance by learning generalizable cross-modal representations from massive unlabeled data.

  6. Knowl 6 — Label Efficiency of FOCAL Under Low-Data Finetuning Regimes

    empirical result

    When evaluating linear probing accuracy under restricted label availability during finetuning (100%, 10%, and 1% of available labeled data), FOCAL shows substantial performance advantages over both self-supervised baselines and supervised training from scratch.

    Across datasets with the Swin-Transformer backbone:

    • MOD Dataset: At a 1% label ratio, FOCAL achieves an accuracy of 0.8840±0.02990.8840 \pm 0.0299 and an F1 score of 0.8776±0.03890.8776 \pm 0.0389, compared to 0.2028±0.01110.2028 \pm 0.0111 for supervised training and 0.7996±0.03310.7996 \pm 0.0331 for the best baseline (TNC).
    • ACIDS Dataset: At a 1% label ratio, FOCAL achieves 0.8669±0.02870.8669 \pm 0.0287 accuracy compared to 0.2666±0.03190.2666 \pm 0.0319 for supervised training and 0.7990±0.02990.7990 \pm 0.0299 for MoCo.
    • RealWorld-HAR Dataset: At a 1% label ratio, FOCAL achieves 0.8301±0.04280.8301 \pm 0.0428 accuracy compared to 0.4541±0.06940.4541 \pm 0.0694 for supervised training and 0.8061±0.02150.8061 \pm 0.0215 for TNC.
    • PAMAP2 Dataset: At a 1% label ratio, FOCAL attains 0.7371±0.03320.7371 \pm 0.0332 accuracy compared to 0.4048±0.03370.4048 \pm 0.0337 for supervised learning.

    On average across all evaluated tasks and datasets, FOCAL surpasses the supervised baseline by 1.37% with 100% labels, 15.04% with 10% labels, and 68.39% with 1% labels.

  7. Knowl 7 — Ablation Analysis of Latent Factorization and Regularization Components in FOCAL

    data/table

    Ablation experiments evaluate the contribution of individual components of FOCAL with the Swin-Transformer (SW-T) and DeepSense backbones across four benchmark datasets under linear probing:

    • FOCAL-noPrivate: Removes the private latent space and private contrastive loss.
    • FOCAL-noOrth: Keeps the private space but removes cosine orthogonality constraints.
    • FOCAL-wDistInd: Replaces geometric orthogonality with statistical distributional independence (minimizing mutual information / KL-divergence via a discriminator).
    • FOCAL-noTemp: Omits the temporal structural sequence ranking constraint.
    • FOCAL-wTempCon: Replaces the coarse temporal ranking constraint with a strict instance-level binary temporal contrastive objective.
    Variant (SW-T) MOD ACIDS RealWorld-HAR PAMAP2
    Acc F1 Acc F1 Acc F1 Acc F1
    FOCAL-noPrivate 0.9296 0.9284 0.7981 0.7100 0.8869 0.8768 0.7938 0.7787
    FOCAL-noOrth 0.9705 0.9692 0.9311 0.8261 0.9186 0.9257 0.8371 0.8233
    FOCAL-wDistInd 0.5773 0.5502 0.4926 0.4157 0.9099 0.9084 0.6518 0.5503
    FOCAL-noTemp 0.9671 0.9659 0.9456 0.8014 0.9361 0.9425 0.8367 0.8255
    FOCAL-wTempCon 0.9363 0.9359 0.9287 0.7587 0.8793 0.8842 0.8391 0.8242
    FOCAL (Full) 0.9805 0.9800 0.9489 0.8262 0.9451 0.9503 0.8580 0.8401

    Removing the private space degrades accuracy by 5.20% to 5.90%. Replacing geometric orthogonality with statistical independence causes optimization instability and severe performance drop. Replacing the sequence-level temporal constraint with strict temporal contrastive loss diminishes accuracy across three of the four datasets.

  8. Knowl 8 — Effectiveness of Temporal Structural Constraint as a Plugin for Other SSL Baselines

    data/table

    The temporal structural constraint (Ltemporal\mathcal{L}_{\text{temporal}}) can function as a modular plugin to enhance existing time-series and multimodal contrastive learning baselines.

    Evaluating linear classification accuracy and macro F1 on the ACIDS and PAMAP2 datasets with the DeepSense backbone before (Vanilla) and after adding the temporal ranking constraint (wTemp):

    Dataset Setup SimCLR MoCo CMC Cocoa GMC
    Acc F1 Acc F1 Acc F1 Acc F1 Acc F1
    ACIDS wTemp 0.7461 0.6938 0.7836 0.6618 0.8690 0.7090 0.8543 0.7665 0.9347 0.8109
    Vanilla 0.7438 0.6101 0.7717 0.6205 0.8443 0.7244 0.6644 0.5359 0.9096 0.7929
    PAMAP2 wTemp 0.7129 0.6884 0.7800 0.7602 0.7804 0.7583 0.8442 0.8146 0.8253 0.8114
    Vanilla 0.6802 0.6583 0.7559 0.7387 0.7906 0.7706 0.7603 0.7187 0.8119 0.7860

    Adding Ltemporal\mathcal{L}_{\text{temporal}} produces performance gains across baseline models, improving accuracy by up to 18.99% on ACIDS (Cocoa) and up to 8.39% on PAMAP2 (Cocoa).

  9. Knowl 9 — Non-Parametric Evaluation via KNN Classification and Latent Space Clustering

    empirical result

    To assess representation quality without training classification layers:

    1. KNN Classification (K=5): Pretrained representations are concatenated across modalities and evaluated directly with a K=5K=5 nearest neighbors classifier. With the Swin-Transformer encoder, FOCAL achieves test accuracies of 0.96650.9665 on MOD, 0.88260.8826 on ACIDS, 0.85860.8586 on RealWorld-HAR, and 0.85490.8549 on PAMAP2, outperforming all contrastive baselines by an average margin of 4.85%4.85\%.
    2. K-Means Clustering Quality: Modality embeddings are clustered into KK clusters matching the ground-truth class count, evaluated by Adjusted Rand Index (ARI) and Normalized Mutual Information (NMI). With Swin-Transformer, FOCAL attains an ARI of 0.4660±0.27370.4660 \pm 0.2737 and NMI of 0.5693±0.25790.5693 \pm 0.2579 on MOD, and ARI of 0.6050±0.10270.6050 \pm 0.1027 and NMI of 0.7389±0.07740.7389 \pm 0.0774 on ACIDS, outperforming CMC, Cosmo, Cocoa, and GMC by an average ARI margin of 8.33%8.33\% and NMI margin of 4.0%4.0\%.
  10. Knowl 10 — Limitations of the FOCAL Framework

    limitation

    The FOCAL framework has several identified limitations:

    1. Assumption of Modality Synchronization: It assumes signals simultaneously arrive across all sensory modalities. When signal propagation velocities differ substantially (e.g., optical vs. acoustic/seismic waves over distance), shared embeddings cannot be directly matched without explicit signal synchronization mechanisms.
    2. Quadratic Scaling in Modality Count: Pairwise calculation of shared space InfoNCE losses and private space orthogonality constraints yields an O(P2)O(P^2) computational complexity relative to the number of modalities PP.
    3. Dependence on Handcrafted Data Augmentations: Learning private representations relies on specific time-domain and frequency-domain augmentations, which require domain knowledge to avoid destroying label-critical semantics.
    4. Sensitivity to Environmental Domain Shift: Pretrained representations experience performance drops when tested on fine-grained regression/classification tasks across unseen environments and terrain (e.g., vehicle speed classification under domain shift drops to 0.6960 accuracy).

Coverage note — None was omitted; all key contributions including latent factorized space formulation, loss equations, temporal ranking regularization, experimental benchmarks across DeepSense and Swin-Transformer, ablation studies, and limitations have been captured.

References

  1. 1.Alessia Amelio and Clara Pizzuti. Is normalized mutual information a fair measure for comparing community detection methods? In IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining (ASONAM), 2015.
  2. 2.Razvan Brinzea, Bulat Khaertdinov, and Stylianos Asteriadis. Contrastive learning with crossmodal knowledge mining for multimodal human activity recognition. In International Joint Conference on Neural Networks (IJCNN), 2022.
  3. 3.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. In International Conference on Machine Learning (ICML), 2020.
  4. 4.Xinlei Chen and Kaiming He. Exploring simple siamese representation learning. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  5. 5.Xinlei Chen, Saining Xie, and Kaiming He. An empirical study of training self-supervised vision transformers. In IEEE/CVF International Conference on Computer Vision (CVPR), 2021.
  6. 6.Shohreh Deldari, Hao Xue, Aaqib Saeed, Daniel V. Smith, and Flora D. Salim. Cocoa: Cross modality contrastive learning for sensor data. ACM Interact. Mob. Wearable Ubiquitous Technol. (IMWUT), 6(3), 2022.
  7. 7.Emadeldeen Eldele, Mohamed Ragab, Zhenghua Chen, Min Wu, Chee Keong Kwoh, Xiaoli Li, and Cuntai Guan. Time-series representation learning via temporal and contextual contrasting. In Thirtieth International Joint Conference on Artificial Intelligence (IJCAI), 2021.
  8. 8.Emadeldeen Eldele, Mohamed Ragab, Zhenghua Chen, Min Wu, Chee Keong Kwoh, Xiaoli Li, and Cuntai Guan. Self-supervised contrastive representation learning for semi-supervised time-series classification. arXiv preprint arXiv:2208.06616, 2022.
  9. 9.Shengyu Feng, Baoyu Jing, Yada Zhu, and Hanghang Tong. Adversarial graph contrastive learning with information regularization. In International World Wide Web Conference (WWW), 2022.
  10. 10.Jean-Yves Franceschi, Aymeric Dieuleveut, and Martin Jaggi. Unsupervised scalable representation learning for multivariate time series. In Advances in Neural Information Processing Systems (NeurIPS), 2019.
  11. 11.Xinyang Geng, Hao Liu, Lisa Lee, Dale Schuurmans, Sergey Levine, and Pieter Abbeel. Multimodal masked autoencoders learn transferable representations. In First Workshop on Pre-training: Perspectives, Pitfalls, and Paths Forward at ICML, 2022.
  12. 12.Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Guo, Mohammad Gheshlaghi Azar, et al. Bootstrap your own latent-a new approach to self-supervised learning. In Advances in Neural Information Processing Systems (NeurIPS), 2020.
  13. 13.Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. Masked autoencoders are scalable vision learners. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  14. 14.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual representation learning. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020.
  15. 15.David T Hoffmann, Nadine Behrmann, Juergen Gall, Thomas Brox, and Mehdi Noroozi. Ranking info noise contrastive estimation: Boosting contrastive learning via ranked positives. In AAAI Conference on Artificial Intelligence (AAAI), 2022.
  16. 16.Umangi Jain, Alex Wilson, and Varun Gulshan. Multimodal contrastive learning for remote sensing tasks. In Self-Supervised Learning - Theory and Practice, NeurIPS Workshop, 2022.
  17. 17.Bulat Khaertdinov, Esam Ghaleb, and Stylianos Asteriadis. Contrastive self-supervised learning for sensor-based human activity recognition. In IEEE International Joint Conference on Biometrics (IJCB). IEEE, 2021.
  18. 18.Dhruv Khattar, Jaipal Singh Goud, Manish Gupta, and Vasudeva Varma. Mvae: Multimodal variational autoencoder for fake news detection. In International World Wide Web Conference (WWW), pages 2915–2921, 2019.
  19. 19.Gukyeong Kwon, Zhaowei Cai, Avinash Ravichandran, Erhan Bas, Rahul Bhotika, and Stefano Soatto. Masked vision and language modeling for multi-modal representation learning. In International Conference on Learning Representations (ICLR), 2023.
  20. 20.Kimin Lee, Kibok Lee, Honglak Lee, and Jinwoo Shin. A simple unified framework for detecting out-of-distribution samples and adversarial attacks. Advances in Neural Information Processing Systems (NeurIPS), 2018.
  21. 21.Oliver Limoyo, Trevor Ablett, and Jonathan Kelly. Learning sequential latent variable models from multimodal time series data. In Ivan Petrovic, Emanuele Menegatti, and Ivan Marković, editors, Intelligent Autonomous Systems 17, 2023.
  22. 22.Dongxin Liu, Tianshi Wang, Shengzhong Liu, Ruijie Wang, Shuochao Yao, and Tarek Abdelzaher. Contrastive self-supervised representation learning for sensing signals from the time-frequency perspective. In IEEE International Conference on Computer Communications and Networks (ICCCN), 2021.
  23. 23.Shengzhong Liu, Shuochao Yao, Jinyang Li, Dongxin Liu, Tianshi Wang, Huajie Shao, and Tarek Abdelzaher. Giobalfusion: A global attentional deep learning framework for multisensor information fusion. ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies (IMWUT), 4(1):1–27, 2020.
  24. 24.Xiao Liu, Fanjin Zhang, Zhenyu Hou, Li Mian, Zhaoyu Wang, Jing Zhang, and Jie Tang. Self-supervised learning: Generative or contrastive. IEEE Transactions on Knowledge and Data Engineering (TKDE), 35(1):857–876, 2021.
  25. 25.Yunze Liu, Qingnan Fan, Shanghang Zhang, Hao Dong, Thomas Funkhouser, and Li Yi. Contrastive multimodal fusion with tupleinfonce. In IEEE/CVF International Conference on Computer Vision (CVPR), 2021.
  26. 26.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In IEEE/CVF International Conference on Computer Vision (CVPR), 2021.
  27. 27.Ilya Loshchilov and Frank Hutter. SGDR: Stochastic gradient descent with warm restarts. In International Conference on Learning Representations (ICLR), 2017.
  28. 28.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In International Conference on Learning Representations (ICLR), 2019.
  29. 29.Shuang Ma, Sai Vemprala, Wenshan Wang, Jayesh K. Gupta, Yale Song, Daniel McDufft, and Ashish Kapoor. Compass: Contrastive multimodal pretraining for autonomous systems. In IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2022.
  30. 30.Hiroshi Morioka and Aapo Hyvarinen. Connectivity-contrastive learning: Combining causal discovery and representation learning for multimodal data. In International Conference on Artificial Intelligence and Statistics (AISTATS), 2023.
  31. 31.Basil Mustafa, Carlos Riquelme Ruiz, Joan Puigcerver, Rodolphe Jenatton, and Neil Houlsby. Multimodal contrastive learning with limoe: the language-image mixture of experts. In Advances in Neural Information Processing Systems (NeurIPS), 2022.
  32. 32.Ryumei Nakada, Halil Ibrahim Gulluk, Zhun Deng, Wenlong Ji, James Zou, and Linjun Zhang. Understanding multimodal contrastive learning and incorporating unpaired data. In International Conference on Artificial Intelligence and Statistics (AISTATS), 2023.
  33. 33.Natalia Neverova, Christian Wolf, Graham Taylor, and Florian Nebout. Moddrop: adaptive multimodal gesture recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 38(8):1692–1706, 2015.
  34. 34.Manuel T Nonnenmacher, Lukas Oldenburg, Ingo Steinwart, and David Reeb. Utilizing expert features for contrastive learning of time-series representations. In International Conference on Machine Learning (ICML), 2022.
  35. 35.Xiaomin Ouyang, Xian Shuai, Jiayu Zhou, Ivy Wang Shi, Zhiyuan Xie, Guoliang Xing, and Jianwei Huang. Cosmo: Contrastive fusion learning with small data for multimodal human activity recognition. In International Conference on Mobile Computing And Networking (MobiCom), 2022.
  36. 36.Yilmazcan Ozyurt, Stefan Feuerriegel, and Ce Zhang. Contrastive learning for unsupervised domain adaptation of time series. In International Conference on Learning Representations (ICLR), 2023.
  37. 37.Petra Poklukar, Miguel Vasco, Hang Yin, Francisco S Melo, Ana Paiva, and Danica Kragic. Geometric multimodal contrastive representation learning. In International Conference on Machine Learning (ICML), 2022.
  38. 38.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning (ICML), 2021.
  39. 39.Valentin Radu, Nicholas D Lane, Sourav Bhattacharya, Cecilia Mascolo, Mahesh K Marina, and Fahim Kawsar. Towards multimodal deep learning for activity recognition on mobile devices. In ACM International Joint Conference on Pervasive and Ubiquitous Computing (UbiComp), 2016.
  40. 40.Aniruddh Raghu, Payal Chandak, Ridwan Alam, John Guttag, and Collin Stultz. Contrastive pre-training for multimodal medical time series. In NeurIPS Workshop on Learning from Time Series for Health, 2022.
  41. 41.Attila Reiss and Didier Stricker. Introducing a new benchmarked dataset for activity monitoring. In International Symposium on Wearable Computers (ISWC), 2012.
  42. 42.Aaqib Saeed, Tanir Ozcelebi, and Johan Lukkien. Multi-task self-supervised learning for human activity detection. ACM Interact. Mob. Wearable Ubiquitous Technol. (IMWUT), 3(2), 2019.
  43. 43.Mathieu Salzmann, Carl Henrik Ek, Raquel Urtasun, and Trevor Darrell. Factorized orthogonal latent spaces. In International Conference on Artificial Intelligence and Statistics (AISTATS), 2010.
  44. 44.Ruiyuan Song, Dongheng Zhang, Zhi Wu, Cong Yu, Chunyang Xie, Shuai Yang, Yang Hu, and Yan Chen. Rf-url: unsupervised representation learning for rf sensing. In International Conference on Mobile Computing And Networking (MobiCom), 2022.
  45. 45.Douglas Steinley. Properties of the hubert-arable adjusted rand index. Psychological methods, 9(3):386, 2004.
  46. 46.Timo Sztyler and Heiner Stuckenschmidt. On-body localization of wearable devices: An investigation of position-aware activity recognition. In IEEE International Conference on Pervasive Computing and Communications (PerCom), 2016.
  47. 47.Aiham Taleb, Matthias Kirchler, Remo Monti, and Christoph Lippert. Contig: Self-supervised multimodal contrastive learning for medical imaging with genetics. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  48. 48.Yonglong Tian, Dilip Krishnan, and Phillip Isola. Contrastive multiview coding. In European Conference on Computer Vision (ECCV), 2020.
  49. 49.Sana Tonekaboni, Danny Eytan, and Anna Goldenberg. Unsupervised representation learning for time series with temporal neighborhood coding. In International Conference on Learning Representations (ICLR), 2021.
  50. 50.Xinming Tu, Zhi-Jie Cao, Sara Mostafavi, Ge Gao, et al. Cross-linked unified embedding for cross-modality representation learning. Advances in Neural Information Processing Systems (NeurIPS), 2022.
  51. 51.Laurens Van der Maaten and Geoffrey Hinton. Visualizing data using t-sne. Journal of Machine Learning Research (JMLR), 9(11), 2008.
  52. 52.Feng Wang and Huaping Liu. Understanding the behaviour of contrastive loss. In IEEE/CVF conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  53. 53.Tongzhou Wang and Phillip Isola. Understanding contrastive representation learning through alignment and uniformity on the hypersphere. In International Conference on Machine Learning (ICML), 2020.
  54. 54.Gerald Woo, Chenghao Liu, Doyen Sahoo, Akshat Kumar, and Steven Hoi. CoST: Contrastive learning of disentangled seasonal-trend representations for time series forecasting. In International Conference on Learning Representations (ICLR), 2022.
  55. 55.Zhirong Wu, Yuanjun Xiong, Stella X Yu, and Dahua Lin. Unsupervised feature learning via non-parametric instance discrimination. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018.
  56. 56.Han Xu, Zheng Yang, Zimu Zhou, Longfei Shangguan, Ke Yi, and Yunhao Liu. Indoor localization via multi-modal sensing on smartphones. In ACM International Joint Conference on Pervasive and Ubiquitous Computing (UbiComp), 2016.
  57. 57.Yihao Xue, Kyle Whitecross, and Baharan Mirzasoleiman. Investigating why contrastive learning benefits robustness against label noise. In International Conference on Machine Learning (ICML), pages 24851–24871, 2022.
  58. 58.Xinyu Yang, Zhenguo Zhang, and Rongyi Cui. Timeclr: A self-supervised contrastive learning framework for univariate time series representation. Knowledge-Based Systems, 245:108606, 2022.
  59. 59.Shuochao Yao, Shaohan Hu, Yiran Zhao, Aston Zhang, and Tarek Abdelzaher. Deepsense: A unified deep learning framework for time-series mobile sensing data processing. In International Conference on World Wide Web (WWW), 2017.
  60. 60.Shuochao Yao, Yiran Zhao, Huajie Shao, Dongxin Liu, Shengzhong Liu, Yifan Hao, Ailing Piao, Shaohan Hu, Su Lu, and Tarek F Abdelzaher. Sadeepsense: Self-attention deep learning framework for heterogeneous on-device sensors in internet of things applications. In IEEE International Conference on Computer Communications (INFOCOM), 2019.
  61. 61.Shuochao Yao, Yiran Zhao, Huajie Shao, Chao Zhang, Aston Zhang, Shaohan Hu, Dongxin Liu, Shengzhong Liu, Lu Su, and Tarek Abdelzaher. Sensegan: Enabling deep learning for internet of things with a semi-supervised framework. ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies (IMWUT), 2(3):1–21, 2018.
  62. 62.Xin Yuan, Zhe Lin, Jason Kuen, Jianming Zhang, Yilin Wang, Michael Maire, Ajinkya Kale, and Baldo Faieta. Multimodal contrastive training for visual representation learning. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  63. 63.Zhihan Yue, Yujing Wang, Juanyong Duan, Tianmeng Yang, Congrui Huang, Yunhai Tong, and Bixiong Xu. Ts2vec: Towards universal representation of time series. In AAAI Conference on Artificial Intelligence (AAAI), 2022.
  64. 64.Miaoran Zhang, Marius Mosbach, David Adelani, Michael Hedderich, and Dietrich Klakow. MCSE: Multimodal contrastive learning of sentence embeddings. In North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), 2022.
  65. 65.Xiang Zhang, Ziyuan Zhao, Theodoros Tsiligkaridis, and Marinka Zitnik. Self-supervised contrastive pre-training for time series via time-frequency consistency. In Advances in Neural Information Processing Systems (NeurIPS), 2022.
  66. 66.Nanxuan Zhao, Zhirong Wu, Rynson W. H. Lau, and Stephen Lin. What makes instance discrimination good for transfer learning? In International Conference on Learning Representations (ICLR), 2021.

Citation

MLA
Liu, S., et al. “FOCAL: Contrastive Learning for Multimodal Time-Series Sensing Signals in Factorized Orthogonal Latent Space”. Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 47309–38, https://proceedings.neurips.cc/paper_files/paper/2023/file/93e98ddf39a9beb0a97fbbe56a986c80-Paper-Conference.pdf.
APA
Liu, S., Kimura, T., Liu, D., Wang, R., Li, J., Diggavi, S., Srivastava, M., & Abdelzaher, T. (2023). FOCAL: Contrastive Learning for Multimodal Time-Series Sensing Signals in Factorized Orthogonal Latent Space. Advances in Neural Information Processing Systems, 36, 47309–47338. https://proceedings.neurips.cc/paper_files/paper/2023/file/93e98ddf39a9beb0a97fbbe56a986c80-Paper-Conference.pdf
Chicago
Liu, S., T. Kimura, D. Liu, et al. 2023. “FOCAL: Contrastive Learning for Multimodal Time-Series Sensing Signals in Factorized Orthogonal Latent Space”. Advances in Neural Information Processing Systems 36: 47309–38. https://proceedings.neurips.cc/paper_files/paper/2023/file/93e98ddf39a9beb0a97fbbe56a986c80-Paper-Conference.pdf.
Harvard
Liu, S. et al. (2023) “FOCAL: Contrastive Learning for Multimodal Time-Series Sensing Signals in Factorized Orthogonal Latent Space”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 47309–47338. Available at: https://proceedings.neurips.cc/paper_files/paper/2023/file/93e98ddf39a9beb0a97fbbe56a986c80-Paper-Conference.pdf.
Vancouver
1. Liu S, Kimura T, Liu D, Wang R, Li J, Diggavi S, Srivastava M, Abdelzaher T (2023) FOCAL: Contrastive Learning for Multimodal Time-Series Sensing Signals in Factorized Orthogonal Latent Space. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 47309–47338

BibTeX

@inproceedings{liu2023focal,
  title = {FOCAL: Contrastive Learning for Multimodal Time-Series Sensing Signals in Factorized Orthogonal Latent Space},
  author = {Liu, Shengzhong and Kimura, Tomoyoshi and Liu, Dongxin and Wang, Ruijie and Li, Jinyang and Diggavi, Suhas and Srivastava, Mani and Abdelzaher, Tarek},
  year = {2023},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {36},
  pages = {47309-47338},
  url = {https://proceedings.neurips.cc/paper_files/paper/2023/file/93e98ddf39a9beb0a97fbbe56a986c80-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors