ReconBoost: Boosting Can Achieve Modality Reconcilement

Cong HuaQianqian XuShilong BaoZhiyong YangQingming Huang

article2024ICML50 citations

Proposes a gradient-boosting-inspired alternating learning framework called ReconBoost that mitigates modality competition by dynamically updating individual modalities sequentially with regularization to reconcile uni-modal exploitation and cross-modal fusion.

Listen

Real-world machine learning systems increasingly rely on multi-modal data, combining sources such as audio, video, and text. However, current joint training paradigms suffer from a problem called modality competition, where a faster-converging dominant modality overpowers the learning process. This suppresses the optimization of weaker modalities, leading to suboptimal feature representation and degraded overall performance.

The article introduces and evaluates ReconBoost, a novel multi-modal alternating learning framework designed to achieve reconcilement between extracting single-modality features and exploring cross-modal interactions. The core objective is to prevent dominant modalities from inhibiting weaker ones by updating one modality at a time while dynamically regularizing learning to leverage complementary information.

To address this challenge, the authors design a framework that sequentially updates individual modality learners using a dynamic objective with a divergence-based reconcilement regularization term. Theoretically, this approach functions like gradient boosting by steering each updated modality to correct errors made by historical models. Unlike traditional boosting ensembles that accumulate large sets of learners, ReconBoost retains only the most recent model per modality to prevent overfitting in deep neural networks. It also incorporates a memory consolidation scheme to prevent catastrophic forgetting and a global rectification scheme to prevent models from getting trapped in poor local optima. The framework was evaluated across six public benchmark datasets covering tasks like audio-visual event localization, speech emotion recognition, object classification, and sentiment analysis.

The experimental findings show substantial improvements across all tested tasks. First, ReconBoost consistently outperformed existing joint-learning and modulation baselines on all six benchmarks, reaching 79.82% accuracy on the CREMA-D dataset compared to 59.50% for standard concatenation and 70.97% for the best prior competitor. Second, it markedly revitalized weak modality representations, boosting visual encoder accuracy on CREMA-D to 73.01% compared to 26.81% under standard joint training. Third, ReconBoost significantly reduced the degree of modality competition across benchmarks, achieving the largest relative performance gains on datasets that exhibited the most severe initial modality imbalance. Fourth, the framework maintained superior accuracy in cross-modal retrieval tasks and exhibited strong robustness against substantial Gaussian noise injected into audio and visual inputs.

These results indicate that synchronous joint training is fundamentally limited by gradient interference across modalities. Shifting to an alternating boosting-inspired optimization paradigm resolves this bottleneck without requiring complex architectural redesigns. For organizations deploying multi-modal artificial intelligence systems, this approach can enhance system accuracy, improve reliability in noisy environments, and mitigate unfair biases that arise when models systematically ignore weaker or under-represented data modalities.

Decision-makers should adopt alternating optimization strategies like ReconBoost in multi-modal pipelines where modality imbalance undermines performance. The framework is compatible with various decision-level fusion schemes, allowing practitioners to pair it with uncertainty-aware or learnable weighting mechanisms. Before wide-scale deployment, teams should conduct standard parameter tuning on the reconcilement trade-off factor and memory consolidation terms to suit domain-specific data characteristics.

While the empirical results are robust across multiple domains, the study's evaluations rely on standard academic benchmarks with predefined modalities. Practitioners should exercise caution when deploying the framework in real-time streaming contexts or environments with highly fluctuating modality availability until further pilot validations are completed.

Cover for ReconBoost: Boosting Can Achieve Modality Reconcilement

Abstract

This paper explores a novel multi-modal alternating learning paradigm pursuing a reconciliation between the exploitation of uni-modal features and the exploration of cross-modal interactions. This is motivated by the fact that current paradigms of multi-modal learning tend to explore multi-modal features simultaneously. The resulting gradient prohibits further exploitation of the features in the weak modality, leading to modality competition, where the dominant modality overpowers the learning process. To address this issue, we study the modality-alternating learning paradigm to achieve reconcilement. Specifically, we propose a new method called ReconBoost to update a fixed modality each time. Herein, the learning objective is dynamically adjusted with a reconcilement regularization against competition with the historical models. By choosing a KL-based reconcilement, we show that the proposed method resembles Friedman’s Gradient-Boosting (GB) algorithm, where the updated learner can correct errors made by others and help enhance the overall performance. The major difference with the classic GB is that we only preserve the newest model for each modality to avoid overfitting caused by ensembling strong learners. Furthermore, we propose a memory consolidation scheme and a global rectification scheme to make this strategy more effective. Experiments over six multi-modal benchmarks speak to the efficacy of the method. We release the code at https://github.com/huacong/ReconBoost.

Table of Contents

  • 1. Introduction
  • 2. Preliminary
  • 2.1. The Task of Multi-modal Learning
  • 2.2. The Challenge of Multi-modal Learning
  • 3. Methodology
  • 3.1. Modality-alternating Update with Dynamic Reconcilment
  • 3.2. Connection to the Boosting Strategy: Theoretical Guanratee
  • 3.3. Enhancement Schemes
  • 3.4. Final Goal
  • 4. Experiments
  • 4.1. Experimental Setup
  • 4.2. Overall Performance
  • 4.3. Quantitative Analysis
  • 4.4. Ablation Study
  • 4.5. Convergence
  • 5. Conclusion
  • Acknowledgements
  • Impact Statement
  • References
  • A. Prior Arts
  • A.1. Multi-modal Learning
  • A.2. Balanced Multi-modal Learning
  • A.3. Boosting
  • B. Proof of Theorem 3.1
  • C. Additional Experiment Setting
  • C.1. Dataset Description
  • C.2. Competitors
  • C.3. Implementation Details
  • C.3.1. NETWORK ARCHITECTURE
  • C.3.2. TRAINING DETAILS
  • D. Additional Experiment Analysis
  • D.1. Performance on Retrieval Task
  • D.2. Robustness Performance
  • D.3. Modality Competition Analysis
  • D.4. Analysis of Mutual Information in Modalities
  • D.5. Modality Selection Strategy
  • D.6. Applicable to Other Fusion Schemes
  • D.7. Impact of Different Classifiers
  • D.8. Sensitivity Analysis of λ
  • D.9. Latent Embedding Visualization

Knowls

  1. Knowl 1 — ReconBoost Modality-Alternating Learning with Dynamic Reconcilement Regularization

    model/method

    In supervised multi-modal learning with MM modalities, let a training sample be xi={mik}k=1Mx_i = \{m_i^k\}_{k=1}^M with one-hot class label yi∈{0,1}Yy_i \in \{0, 1\}^Y across YY categories. For each modality k∈{1,…,M}k \in \{1, \dots, M\}, a modality learner ϕk(ϑk;mik)=Wk⋅Fk(θk;mik)∈RY\phi_k(\vartheta_k; m_i^k) = W_k \cdot F_k(\theta_k; m_i^k) \in \mathbb{R}^Y maps raw feature inputs mikm_i^k to class logits through a neural feature extractor Fk(θk)F_k(\theta_k) and a linear classification head WkW_k, where ϑk={θk,Wk}\vartheta_k = \{\theta_k, W_k\}. Under standard joint synchronous training, a shared gradient ∂ℓ(ΦM(xi),yi)∂ΦM(xi)=σi−yi\frac{\partial \ell(\Phi_M(x_i), y_i)}{\partial \Phi_M(x_i)} = \sigma_i - y_i updates all modalities simultaneously (where σi∈RY\sigma_i \in \mathbb{R}^Y is the predicted probability vector), causing faster-converging dominant modalities to overpower weak modalities and suppress their optimization.

    To prevent modality competition, ReconBoost replaces synchronous joint optimization with an alternating learning paradigm. In each round ss, only one modality learner ϕk\phi_k is updated while parameters of all other modalities remain fixed. To promote cross-modal exploration and avoid isolating uni-modal learners, the objective is dynamically augmented with a reconcilement regularization term Ds\mathcal{D}_s that measures diversity between the updated modality and historical models:

    L~s(ϕk(mk),y)=1N∑i=1N[ℓ(ϕk(ϑk;mik),yi)−λDs(ΦM/k(xi),ϕk(ϑk;mik))]\tilde{\mathcal{L}}_s(\phi_k(m^k), y) = \frac{1}{N} \sum_{i=1}^N \left[ \ell(\phi_k(\vartheta_k; m_i^k), y_i) - \lambda \mathcal{D}_s(\Phi_{M/k}(x_i), \phi_k(\vartheta_k; m_i^k)) \right]

    where ℓ\ell is the cross-entropy loss, ΦM/k(xi)=∑j=1,j≠kMϕj(ϑj;mij)\Phi_{M/k}(x_i) = \sum_{j=1, j \neq k}^M \phi_j(\vartheta_j; m_i^j) is the aggregated prediction from all other modalities, and λ>0\lambda > 0 is a trade-off hyperparameter balancing task accuracy alignment against cross-modal diversity.

  2. Knowl 2 — Theoretical Equivalence of Reconcilement Regularization to Functional Gradient Boosting

    theoretical result

    Let ΦM/k(xi)=∑j≠kϕj(ϑj;mij)\Phi_{M/k}(x_i) = \sum_{j \neq k} \phi_j(\vartheta_j; m_i^j) denote the partial ensemble prediction excluding modality kk. When the reconcilement regularization term Ds\mathcal{D}_s satisfies the gradient condition:

    λ⋅∇ϕkDs(ΦM/k(xi),ϕk(ϑk;mik))=∇ϕkℓ(ϕk(ϑk;mik),yi)−∇ϕkℓ(ϕk(ϑk;mik),−∇ΦM/kℓ(ΦM/k(xi),yi))\lambda \cdot \nabla_{\phi_k} \mathcal{D}_s(\Phi_{M/k}(x_i), \phi_k(\vartheta_k; m_i^k)) = \nabla_{\phi_k} \ell(\phi_k(\vartheta_k; m_i^k), y_i) - \nabla_{\phi_k} \ell\left(\phi_k(\vartheta_k; m_i^k), -\nabla_{\Phi_{M/k}} \ell(\Phi_{M/k}(x_i), y_i)\right)

    the gradient updates of the dynamic loss objective L~s(ϕk(mk),y)\tilde{\mathcal{L}}_s(\phi_k(m^k), y) are mathematically equivalent to fitting the negative functional gradient of the remaining modalities:

    ∇ϑkL~s(ϕk(mk),y)  ⟺  ∇ϑkL(ϕk(mk),−∇ΦM/kℓ(ΦM/k(x),y))\nabla_{\vartheta_k} \tilde{\mathcal{L}}_s(\phi_k(m^k), y) \iff \nabla_{\vartheta_k} \mathcal{L}\left(\phi_k(m^k), -\nabla_{\Phi_{M/k}} \ell(\Phi_{M/k}(x), y)\right)

    Specifically, when ℓ\ell is the cross-entropy loss and Ds\mathcal{D}_s is chosen as the Kullback-Leibler (KL) divergence DKL(ΦM/k(xi)∥ϕk(ϑk;mik))\mathcal{D}_{KL}(\Phi_{M/k}(x_i) \parallel \phi_k(\vartheta_k; m_i^k)), the pseudo-label is:

    y~i=−∂ℓ(ΦM/k(xi),yi)∂ΦM/k(xi)=yi−σi\tilde{y}_i = -\frac{\partial \ell(\Phi_{M/k}(x_i), y_i)}{\partial \Phi_{M/k}(x_i)} = y_i - \sigma_i

    where σi\sigma_i is the softmax probability output of ΦM/k(xi)\Phi_{M/k}(x_i). Consequently, optimizing each modality sequentially with KL-based reconcilement regularization performs Friedman's gradient boosting in the functional space, allowing each learner to explicitly fit and correct the residual errors made by the other modalities.

  3. Knowl 3 — Memory Consolidation Regularization and Global Rectification Scheme

    model/method

    Because ReconBoost uses over-parameterized deep neural networks as modality learners and preserves only the newest model per modality rather than storing an infinite ensemble of weak learners, standard greedy boosting can lead to catastrophic forgetting of historical patterns and entrapment in bad local minima. ReconBoost introduces two enhancement schemes:

    1. Memory Consolidation Regularization (MCR): To ensure that an updated learner does not overfit to residuals at the expense of samples where previous learners performed well, MCR penalizes drastic divergence between the gradient directions of the current learner ϕk\phi_k and the previous learner ϕk−1\phi_{k-1}:

    Lmcr(−∇ϕk−1ℓ(ϕk−1(mk−1),y),−∇ϕkℓ(ϕk(mk),y))=1N∑i=1N∥∇ϕkℓ(ϕk(mik),yi)−∇ϕk−1ℓ(ϕk−1(mik−1),yi)∥2\mathcal{L}_{mcr}\left(-\nabla_{\phi_{k-1}} \ell(\phi_{k-1}(m^{k-1}), y), -\nabla_{\phi_k} \ell(\phi_k(m^k), y)\right) = \frac{1}{N} \sum_{i=1}^N \left\| \nabla_{\phi_k} \ell(\phi_k(m_i^k), y_i) - \nabla_{\phi_{k-1}} \ell(\phi_{k-1}(m_i^{k-1}), y_i) \right\|^2

    1. Global Rectification Scheme (GRS): Following each alternating-boosting stage, all MM modality learners are updated simultaneously on the overall multimodal ensemble loss L(ΦM(x),y)\mathcal{L}(\Phi_M(x), y) with learning rate η\eta:

    ϑmt=ϑmt−1−η∇ϑmt−1L(ΦMt−1(x),y),∀m∈[1,M]\vartheta_m^t = \vartheta_m^{t-1} - \eta \nabla_{\vartheta_m^{t-1}} \mathcal{L}(\Phi_M^{t-1}(x), y), \quad \forall m \in [1, M]

    Over a full cycle of MM stages, the unified optimization objective is:

    Lall=∑k=1ML(ϕk(mk),y)−λ∑k=1MDKL(ΦM/k(x)∥ϕk(mk))+α∑k=1MLmcr+∑k=1ML(ΦM(x),y)\mathcal{L}_{all} = \sum_{k=1}^M \mathcal{L}(\phi_k(m^k), y) - \lambda \sum_{k=1}^M \mathcal{D}_{KL}(\Phi_{M/k}(x) \parallel \phi_k(m^k)) + \alpha \sum_{k=1}^M \mathcal{L}_{mcr} + \sum_{k=1}^M \mathcal{L}(\Phi_M(x), y)

  4. Knowl 4 — ReconBoost Training Algorithm

    algorithm

    The ReconBoost optimization algorithm alternates between boosting a specific modality learner to fit historical residuals with regularization and executing global rectification across all learners.

    Input: Training dataset DtrainD_{train}, stage iterations TT, alternating-boosting learning rate γ\gamma, global rectification learning rate η\eta, regularization weights λ\lambda and α\alpha
    Output: Trained multimodal predictor ΦM\Phi_M
    repeat
        for k=1k = 1 to MM do
            # Alternating-boosting Strategy
            for t=0t = 0 to T−1T - 1 do
                Sample batch (xi,yi)∼Dtrain(x_i, y_i) \sim D_{train}
                Compute modality-specific loss ℓA(ϕkt)=ℓ(ϕkt(mik),yi)−λDKL,s+αℓmcr\ell_A(\phi_k^t) = \ell(\phi_k^t(m_i^k), y_i) - \lambda \mathcal{D}_{KL, s} + \alpha \ell_{mcr}
                Update ϑkt+1=ϑkt−γ∇ϑktℓA(ϕkt)\vartheta_k^{t+1} = \vartheta_k^t - \gamma \nabla_{\vartheta_k^t} \ell_A(\phi_k^t)
            Add updated model ϕkT\phi_k^T into ensemble ΦM,s\Phi_{M, s}
            
            # Global Rectification Scheme
            for t=0t = 0 to T−1T - 1 do
                Sample batch (xi,yi)∼Dtrain(x_i, y_i) \sim D_{train}
                Compute ensemble loss ℓG(ΦM,st)=ℓ(∑m=1Mϕmt(mim),yi)\ell_G(\Phi_{M, s}^t) = \ell(\sum_{m=1}^M \phi_m^t(m_i^m), y_i)
                for m=1m = 1 to MM do
                    Update ϑmt+1=ϑmt−η∇ϑmtℓG(ΦM,st)\vartheta_m^{t+1} = \vartheta_m^t - \eta \nabla_{\vartheta_m^t} \ell_G(\Phi_{M, s}^t)
    until convergence
    return ΦM\Phi_M

    In typical implementations, SGD optimizer is used with base learning rates γ=η=0.01\gamma = \eta = 0.01, stage lengths T1=T2=4T_1 = T_2 = 4 for datasets such as AVE, CREMA-D, and ModelNet40, and T1=T2=1T_1 = T_2 = 1 for MOSEI, MOSI, and CH-SIMS.

  5. Knowl 5 — Modality Imbalance Ratio and Degree of Modality Competition Metrics

    definition

    To quantitatively evaluate the severity of modality competition in multimodal systems relative to unimodal performance, two metrics are formulated based on accuracy evaluations Acc(⋅)\text{Acc}(\cdot) on linear probes trained on frozen encoders:

    1. Modality Imbalance Ratio (MIR): For strong modality XiX_i and weak modality XjX_j, the unimodal ratio MIRuni\text{MIR}^{uni} and multimodal ratio MIRmul\text{MIR}^{mul} are defined as:

    MIRuni(Xi,Xj)=Acc(fiuni∘φiuni(Xi))Acc(fjuni∘φjuni(Xj)),MIRmul(Xi,Xj)=Acc(fimul∘φimul(Xi))Acc(fjmul∘φjmul(Xj))\text{MIR}^{uni}(X_i, X_j) = \frac{\text{Acc}(f_i^{uni} \circ \varphi_i^{uni}(X_i))}{\text{Acc}(f_j^{uni} \circ \varphi_j^{uni}(X_j))}, \qquad \text{MIR}^{mul}(X_i, X_j) = \frac{\text{Acc}(f_i^{mul} \circ \varphi_i^{mul}(X_i))}{\text{Acc}(f_j^{mul} \circ \varphi_j^{mul}(X_j))}

    where φuni,funi\varphi^{uni}, f^{uni} denote encoders and classifiers trained via unimodal training, and φmul,fmul\varphi^{mul}, f^{mul} denote encoders trained via multimodal learning with separate post-hoc linear classifiers.

    1. Degree of Modality Competition (DMC): Compares the imbalance under multimodal joint training relative to unimodal baseline imbalance:

    DMC(Xi,Xj)=MIRmul(Xi,Xj)MIRuni(Xi,Xj)\text{DMC}(X_i, X_j) = \frac{\text{MIR}^{mul}(X_i, X_j)}{\text{MIR}^{uni}(X_i, X_j)}

    For trimodal datasets, DMC is computed via the geometric mean across all distinct pairs:

    DMC(Xi,Xj,Xk)=(∏m,n∈{i,j,k},m≠nDMC(Xm,Xn))1/3\text{DMC}(X_i, X_j, X_k) = \left( \prod_{m, n \in \{i, j, k\}, m \neq n} \text{DMC}(X_m, X_n) \right)^{1/3}

    A DMC>1.0\text{DMC} > 1.0 indicates that joint multimodal training has exacerbated the performance imbalance between modalities (active modality competition), whereas DMC<1.0\text{DMC} < 1.0 indicates successful modality reconcilement.

  6. Knowl 6 — Multimodal Classification Performance Across Benchmark Datasets

    data/table

    ReconBoost was evaluated across six multimodal classification datasets: AVE (28 event classes, audio-visual), CREMA-D (6 emotion classes, audio-visual), ModelNet40 (40 CAD object classes, front/rear visual views), CMU-MOSEI (3 sentiment classes, audio-visual-text), CMU-MOSI (3 sentiment classes, audio-visual-text), and CH-SIMS (3 sentiment classes, audio-visual-text). Backbones include ResNet-18 for image/spectrogram inputs and LSTM for text.

    Method AVE CREMA-D MN40 MOSEI MOSI CH-SIMS
    AudioNet 59.37 56.67 - 52.29 54.81 58.20
    VisualNet 30.46 50.14 80.51 50.35 57.87 63.02
    TextNet - - - 66.41 75.94 70.45
    Concat Fusion 62.68 59.50 83.18 66.71 76.23 71.55
    G-Blending 62.75 63.81 84.56 66.93 76.45 71.55
    OGM-GE 62.93 65.59 85.61 66.67 76.01 71.10
    PMR 64.20 66.10 86.20 66.41 76.12 70.90
    UME 66.92 68.41 85.37 63.88 76.97 71.77
    UMT 67.71 70.97 90.07 67.04 75.80 71.55
    Ours (ReconBoost) 71.35 79.82 91.78 68.61 77.96 73.88

    ReconBoost achieves state-of-the-art accuracy across all six benchmarks, outperforming naive concatenation by +8.67% on AVE, +20.32% on CREMA-D, and +8.60% on ModelNet40, while outperforming the strongest prior competitor (UMT) by +3.64% on AVE and +8.85% on CREMA-D.

  7. Knowl 7 — Evaluation of Modality-Specific Feature Encoders via Linear Probing

    data/table

    To evaluate whether multimodal models preserve and exploit unimodal representation quality, encoders trained under different multimodal frameworks were frozen, and separate linear classifiers were trained on their output representations for CREMA-D and AVE.

    CREMA-D AVE
    Method Visual Acc (%) Audio Acc (%) Visual Acc (%) Audio Acc (%)
    Uni-train 50.14 56.67 30.46 59.37
    Concat Fusion 26.81 54.86 23.96 55.47
    OGM-GE 29.17 55.42 25.52 56.51
    PMR 29.21 55.60 26.30 57.20
    UMT 45.69 58.47 31.25 60.70
    Ours (ReconBoost) 73.01 60.23 39.06 61.20

    Under standard concatenation joint fusion, the weak visual modality encoder severely degrades compared to standalone unimodal training (dropping from 50.14% to 26.81% on CREMA-D, and 30.46% to 23.96% on AVE). In contrast, ReconBoost substantially improves the weak modality representation (reaching 73.01% on CREMA-D visual and 39.06% on AVE visual), while also boosting the strong audio modality beyond unimodal levels, confirming effective modality reconcilement.

  8. Knowl 8 — Compatibility of ReconBoost with Decision-Level Fusion Strategies

    data/table

    ReconBoost is formulated around decision-level aggregation ΦM(xi)=∑k=1Mwk⋅ϕk(ϑk;mik)\Phi_M(x_i) = \sum_{k=1}^M w_k \cdot \phi_k(\vartheta_k; m_i^k), where wkw_k denotes the modality importance weight during inference. It can be integrated with various fusion mechanisms, such as Naive Averaging (wk=1w_k=1), Learnable Weighting (LW), Trusted Multi-view Classification (TMC), and Quality-aware Multimodal Fusion (QMF). In comparison, prior joint-learning methods were paired with the Multimodal Transfer Module (MMTM) feature-level fusion.

    AVE CREMA-D
    Method Overall Acc Audio Acc Visual Acc Overall Acc Audio Acc Visual Acc
    OGM-GE + MMTM 66.14 58.23 28.09 69.83 58.76 53.35
    PMR + MMTM 67.72 58.47 28.73 70.14 58.94 54.23
    UMT + MMTM 70.16 60.40 35.83 74.35 60.86 62.83
    Ours + NA 71.35 61.20 39.06 79.82 60.23 73.01
    Ours + LW 72.40 61.31 39.13 80.11 60.09 73.30
    Ours + TMC 72.96 61.51 40.20 80.68 60.37 73.86
    Ours + QMF 73.20 61.96 40.85 81.11 60.94 73.87

    ReconBoost consistently achieves superior classification accuracies under all decision aggregation schemes compared to feature-level MMTM competitors, with dynamic uncertainty-weighted decision fusion (QMF) achieving the top overall accuracies of 73.20% on AVE and 81.11% on CREMA-D.

  9. Knowl 9 — Cross-Modal and In-Domain Retrieval Performance

    data/table

    To evaluate multimodal feature representations on downstream information retrieval, pre-trained modality-specific encoders were tested on AVE and CREMA-D datasets using cosine similarity scoring and evaluated with Mean Average Precision (mAP in %).

    AVE CREMA-D
    Method Overall MAP Audio MAP Visual MAP Overall MAP Audio MAP Visual MAP
    Concat Fusion 35.25 37.23 18.82 36.43 34.71 20.08
    OGM-GE 36.92 35.43 20.04 38.50 36.59 24.42
    PMR 36.75 35.71 20.32 39.34 36.97 25.10
    UME 34.91 33.41 21.93 40.02 37.12 30.45
    UMT 36.72 34.64 21.76 42.58 35.65 32.41
    CMCL 40.21 38.15 23.41 53.31 38.31 41.25
    HSR 41.49 39.21 24.01 55.22 39.20 44.67
    Ours (ReconBoost) 43.85 42.71 25.22 60.52 40.58 54.26

    ReconBoost achieves the highest retrieval mAP across all settings, outperforming both general multimodal competition baselines (UMT by +7.13% overall mAP on AVE and +17.94% on CREMA-D) and dedicated cross-modal retrieval methods (CMCL and HSR), driven by large gains in weak modality feature representations.

  10. Knowl 10 — Robustness of ReconBoost Under Modality Gaussian Noise Corruption

    data/table

    To evaluate resilience to sensor degradation and noisy inputs, 50% of the testing data in CREMA-D was corrupted with zero-mean Gaussian noise ϵ∼N(0,σ2)\epsilon \sim \mathcal{N}(0, \sigma^2) across noise variances σ2∈{0.0,0.1,0.3,0.5,1.0}\sigma^2 \in \{0.0, 0.1, 0.3, 0.5, 1.0\} under three corruption setups:

    Corruption Setup σ2=0.0\sigma^2 = 0.0 σ2=0.1\sigma^2 = 0.1 σ2=0.3\sigma^2 = 0.3 σ2=0.5\sigma^2 = 0.5 σ2=1.0\sigma^2 = 1.0
    50% Visual Corrupted
    Concat Fusion 59.50 58.70 58.13 57.70 57.10
    OGM-GE 65.59 64.20 62.50 61.70 60.17
    UMT 70.97 68.76 64.92 63.01 62.23
    Ours 79.82 74.75 68.26 65.73 63.95
    50% Audio Corrupted
    Concat Fusion 59.50 57.01 55.56 54.74 52.27
    OGM-GE 65.59 63.28 62.09 59.56 56.49
    UMT 70.97 66.29 64.71 63.40 60.37
    Ours 79.82 73.65 68.19 65.05 63.24
    50% Both Corrupted
    Concat Fusion 59.50 55.31 52.34 49.42 47.23
    OGM-GE 65.59 60.14 57.36 54.26 50.38
    UMT 70.97 65.02 63.50 60.63 55.74
    Ours 79.82 71.83 67.17 63.09 57.60

    ReconBoost maintains superior accuracy over all baseline models across all noise levels σ2\sigma^2, demonstrating strong robustness even when both audio and visual streams are concurrently corrupted by severe noise (σ2=1.0\sigma^2 = 1.0).

Coverage note — Ablations on heuristic modality update order selection (S1 and S2) and qualitative t-SNE visual plots were omitted in favor of self-contained quantitative and methodological knowls.

References

  1. 1.Badirli, S., Liu, X., Xing, Z., Bhowmik, A., and Keerthi, S. S. Gradient boosting neural networks: Grownet. ArXiv, abs/2002.07971, 2020.
  2. 2.Baltrusaitis, T., Zadeh, A., Lim, Y. C., and Morency, L.-P. Openface 2.0: Facial behavior analysis toolkit. In IEEE International Conference on Automatic Face and Gesture Recognition, pp. 59–66, 2018.
  3. 3.Baltrušaitis, T., Ahuja, C., and Morency, L.-P. Multimodal machine learning: A survey and taxonomy. IEEE TPAMI, 41(2):423–443, 2019.
  4. 4.Cai, X., Nie, F., Huang, H., and Kamangar, F. Heterogeneous image feature integration via multi-modal spectral clustering. In CVPR, pp. 1977–1984, 2011.
  5. 5.Cao, H., Cooper, D. G., Keutmann, M. K., Gur, R. C., Nenkova, A., and Verma, R. Crema-d: Crowd-sourced emotional multimodal actors dataset. IEEE Trans. Affective Comput., 5(4):377–390, 2014.
  6. 6.Chen, R. J., Lu, M. Y., Williamson, D. F., Chen, T. Y., Lipkova, J., Noor, Z., Shaban, M., Shady, M., Williams, M., Joo, B., et al. Pan-cancer integrative histology-genomic analysis via multimodal deep learning. Cancer Cell, 40 (8):865–878, 2022.
  7. 7.Chen, T. and Guestrin, C. Xgboost: A scalable tree boosting system. In SIGKDD, pp. 785–794, 2016.
  8. 8.Chen, X., Lin, K.-Y., Wang, J., Wu, W., Qian, C., Li, H., and Zeng, G. Bi-directional cross-modality feature propagation with separation-and-aggregation gate for rgb-d semantic segmentation. In ECCV, pp. 561–577, 2020.
  9. 9.Degottex, G., Kane, J., Drugman, T., Raitio, T., and Scherer, S. Covarep — a collaborative voice analysis repository for speech technologies. In ICASSP, pp. 960–964, 2014.
  10. 10.Deng, X. and Dragotti, P. L. Deep convolutional neural network for multi-modal image restoration and fusion. IEEE TPAMI, 43(10):3333–3348, 2021.
  11. 11.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  12. 12.Du, C., Teng, J., Li, T., Liu, Y., Yuan, T., Wang, Y., Yuan, Y., and Zhao, H. On uni-modal feature learning in supervised multi-modal learning. In ICML, pp. 25, 2023.
  13. 13.Fan, Y., Xu, W., Wang, H., Wang, J., and Guo, S. Pmr: Prototypical modal rebalance for multimodal learning. In CVPR, pp. 20029–20038, 2023.
  14. 14.Feng, Y., Gao, Y., Zhao, X., Guo, Y., Bagewadi, N., Bui, N.-T., Dao, H., Gangisetty, S., Guan, R., Han, X., et al. Shrec’22 track: Open-set 3d object retrieval. Computers & Graphics, 107:231–240, 2022.
  15. 15.Freund, Y. Boosting a weak learning algorithm by majority. Inf. Comput., 121(2):256–285, 1995.
  16. 16.Freund, Y. and Schapire, R. E. A decision-theoretic generalization of on-line learning and an application to boosting. J. Comput. Syst. Sci., 55(1):119–139, 1997.
  17. 17.Freund, Y., Schapire, R. E., et al. Experiments with a new boosting algorithm. In ICML, pp. 148–156, 1996.
  18. 18.Friedman, J. H. Greedy function approximation: A gradient boosting machine. The Annals of Statistics, 29(5):1189 – 1232, 2001.
  19. 19.Gallego, G., Delbruck, T., Orchard, G., Bartolozzi, C., Taba, B., Censi, A., Leutenegger, S., Davison, A. J., Conradt, J., Daniilidis, K., and Scaramuzza, D. Event-based vision: A survey. IEEE TPAMI, 44(1):154–180, 2022.
  20. 20.Gao, Y., Li, S., Li, Y., Guo, Y., and Dai, Q. Superfast: 200× video frame interpolation via event camera. IEEE TPAMI, 45(6):7764–7780, 2023.
  21. 21.Guan, D., Cao, Y., Liang, J., Cao, Y., and Yang, M. Y. Fusion of multispectral data through illumination-aware deep neural networks for pedestrian detection. ArXiv, abs/1802.09972, 2018.
  22. 22.Han, Z., Zhang, C., Fu, H., and Zhou, J. T. Trusted multi-view classification. In ICLR, 2021.
  23. 23.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In CVPR, 2016.
  24. 24.Hessel, J. and Lee, L. Does my multimodal model learn cross-modal interactions? it’s harder to tell than you might think! In EMNLP, pp. 861–877, 2020.
  25. 25.Hu, W., Miyato, T., Tokui, S., Matsumoto, E., and Sugiyama, M. Learning discrete representations via information maximizing self-augmented training. In ICML, pp. 1558–1567, 2017.
  26. 26.Huang, F., Ash, J., Langford, J., and Schapire, R. Learning deep resnet blocks sequentially using boosting theory. ArXiv, abs/1706.04964, 2018.
  27. 27.Huang, Y., Lin, J., Zhou, C., Yang, H., and Huang, L. Modality competition: What makes joint training of multi-modal network fail in deep learning? (Provably). In ICML, pp. 9226–9259, 2022.
  28. 28.Ivanov, S. and Prokhorenkova, L. Boost then convolve: Gradient boosting meets graph neural networks. In ICLR, 2021.
  29. 29.Jiang, Y., Xu, Q., Yang, Z., Cao, X., and Huang, Q. Dm2c: Deep mixed-modal clustering. In NeurIPS, pp. 5880–5890, 2019.
  30. 30.Jiang, Y., Hua, C., Feng, Y., and Gao, Y. Hierarchical set-to-set representation for 3-d cross-modal retrieval. IEEE TNNLS, pp. 1–13, 2023a.
  31. 31.Jiang, Y., Wang, Y., Li, S., Zhang, Y., Zhao, M., and Gao, Y. Event-based low-illumination image enhancement. IEEE TMM, pp. 1–12, 2023b.
  32. 32.Jing, L., Vahdani, E., Tan, J., and Tian, Y. Cross-modal center loss for 3d cross-modal retrieval. In CVPR, pp. 3142–3151, 2021.
  33. 33.Joze, H. R. V., Shaban, A., Iuzzolino, M. L., and Koishida, K. Mmtm: Multimodal transfer module for cnn fusion. In CVPR, pp. 13289–13299, 2020.
  34. 34.Kullback, S. and Leibler, R. A. On information and sufficiency. The annals of mathematical statistics, 22(1): 79–86, 1951.
  35. 35.Li, J., Qiang, W., Zheng, C., Su, B., Razzak, F., Wen, J.-R., and Xiong, H. Modeling multiple views via implicitly preserving global consistency and local complementarity. IEEE TKDE, 2022.
  36. 36.Li, Z., Tang, J., and Mei, T. Deep collaborative embedding for social image understanding. IEEE TPAMI, 41(9): 2070–2083, 2019.
  37. 37.Liang, P. P., Deng, Z., Ma, M. Q., Zou, J., Morency, L.-P., and Salakhutdinov, R. Factorized contrastive learning: Going beyond multi-view redundancy. In NeurIPS, 2023a.
  38. 38.Liang, P. P., Deng, Z., Ma, M. Q., Zou, J. Y., Morency, L.-P., and Salakhutdinov, R. Factorized contrastive learning: Going beyond multi-view redundancy. In NeurIPS, pp. 32971–32998, 2023b.
  39. 39.Liang, P. P., Lyu, Y., Chhablani, G., Jain, N., Deng, Z., Wang, X., Morency, L.-P., and Salakhutdinov, R. Multiviz: Towards visualizing and understanding multimodal models. In ICLR, 2023c.
  40. 40.McFee, B., Raffel, C., Liang, D., Ellis, D. P., McVicar, M., Battenberg, E., and Nieto, O. librosa: Audio and music signal analysis in python. In SciPy, pp. 18–24, 2015.
  41. 41.Nagrani, A., Yang, S., Arnab, A., Jansen, A., Schmid, C., and Sun, C. Attention bottlenecks for multimodal fusion. In NeurIPS, pp. 14200–14213, 2021.
  42. 42.Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A. Automatic differentiation in pytorch. 2017.
  43. 43.Peng, X., Wei, Y., Deng, A., Wang, D., and Hu, D. Balanced multimodal learning via on-the-fly gradient modulation. In CVPR, pp. 8238–8247, 2022.
  44. 44.Robbins, H. and Monro, S. A stochastic approximation method. The annals of mathematical statistics, pp. 400–407, 1951.
  45. 45.Ruan, D., Ji, S., Yan, C., Zhu, J., Zhao, X., Yang, Y., Gao, Y., Zou, C., and Dai, Q. Exploring complex and heterogeneous correlations on hypergraph for the prediction of drug-target interactions. Patterns, 2(12), 2021.
  46. 46.Seichter, D., Köhler, M., Lewandowski, B., Wengefeld, T., and Gross, H.-M. Efficient rgb-d semantic segmentation for indoor scene analysis. ArXiv, abs/2011.06961, 2021.
  47. 47.Shahroudy, A., Ng, T.-T., Gong, Y., and Wang, G. Deep multimodal feature analysis for action recognition in rgb+ d videos. IEEE TPAMI, 40(5):1045–1058, 2017.
  48. 48.Shalev-Shwartz, S. Selfieboost: A boosting algorithm for deep learning. ArXiv, abs/1411.3436, 2014.
  49. 49.Shao, Z., Li, F., Zhou, Y., Chen, H., Zhu, H., and Yao, R. Identity-invariant representation and transformer-style relation for micro-expression recognition. Applied Intelligence, 53(17):19860–19871, 2023.
  50. 50.Shao, Z., Zhou, Y., Li, F., Zhu, H., and Liu, B. Joint facial action unit recognition and self-supervised optical flow estimation. Pattern Recognition Letters, 181:70–76, 2024.
  51. 51.Sun, K., Zhu, Z., and Lin, Z. Adagcn: Adaboosting graph convolutional networks into deep models. ArXiv, abs/1908.05081, 2019.
  52. 52.Tang, J., Shu, X., Qi, G.-J., Li, Z., Wang, M., Yan, S., and Jain, R. Tri-clustered tensor completion for social-aware image tag refinement. IEEE TPAMI, 39(8):1662–1674, 2017.
  53. 53.Tian, Y., Shi, J., Li, B., Duan, Z., and Xu, C. Audio-visual event localization in unconstrained videos. In ECCV, pp. 247–263, 2018.
  54. 54.Van der Maaten, L. and Hinton, G. Visualizing data using t-sne. JMLR, 9(11), 2008.
  55. 55.Wan, Z., Mao, Y., Zhang, J., and Dai, Y. Rpeflow: Multimodal fusion of rgb-pointcloud-event for joint optical flow and scene flow estimation. In ICCV, pp. 10030–10040, 2023.
  56. 56.Wang, W., Tran, D., and Feiszli, M. What makes training multi-modal classification networks hard? In CVPR, pp. 12692–12702, 2020a.
  57. 57.Wang, Y., Huang, W., Sun, F., Xu, T., Rong, Y., and Huang, J. Deep multimodal fusion by channel exchanging. In NeurIPS, pp. 4835–4845, 2020b.
  58. 58.Wang, Y., Zhang, Y., Guo, Q., Zhao, M., and Jiang, Y. Rnve: A real nighttime vision enhancement benchmark and dual-stream fusion network. IEEE Signal Process Lett., 31: 131–135, 2024.
  59. 59.Wei, Y., Hu, D., Tian, Y., and Li, X. Learning in audio-visual context: A review, analysis, and new perspective, 2022.
  60. 60.Williams, J., Kleinegesse, S., Comanescu, R., and Radu, O. Recognizing emotions in video using multimodal dnn feature fusion. In Proceedings of Grand Challenge and Workshop on Human Multimodal Language, 2018. doi: 10.18653/v1/W18-3302.
  61. 61.Wu, N., Jastrzebski, S., Cho, K., and Geras, K. J. Characterizing and overcoming the greedy nature of learning in multi-modal deep neural networks. In ICML, pp. 24043–24055, 2022.
  62. 62.Wu, Z., Song, S., Khosla, A., Yu, F., Zhang, L., Tang, X., and Xiao, J. 3d shapenets: A deep representation for volumetric shapes. In CVPR, pp. 1912–1920, 2015.
  63. 63.Yang, P., Wang, X., Duan, X., Chen, H., Hou, R., Jin, C., and Zhu, W. Avqa: A dataset for audio-visual question answering on videos. In ACM MM, pp. 3480–3491, 2022.
  64. 64.Yu, W., Xu, H., Meng, F., Zhu, Y., Ma, Y., Wu, J., Zou, J., and Yang, K. Ch-sims: A chinese multimodal sentiment analysis dataset with fine-grained annotation of modality. In Annual Meeting of the Association for Computational Linguistics, 2020.
  65. 65.Zadeh, A., Zellers, R., Pincus, E., and Morency, L.-P. Mosi: Multimodal corpus of sentiment intensity and subjectivity analysis in online opinion videos. ArXiv, abs/1606.06259, 2016.
  66. 66.Zadeh, A., Liang, P. P., Poria, S., Cambria, E., and Morency, L.-P. Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph. In Annual Meeting of the Association for Computational Linguistics, 2018.
  67. 67.Zhang, D., Zhang, H., Tang, J., Hua, X.-S., and Sun, Q. Causal intervention for weakly-supervised semantic segmentation. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H. (eds.), NeurIPS, volume 33, pp. 655–666, 2020.
  68. 68.Zhang, Q., Wu, H., Zhang, C., Hu, Q., Fu, H., Zhou, J. T., and Peng, X. Provable dynamic fusion for low-quality multimodal data. In ICML, pp. 17, 2023.
  69. 69.Zhang, Q., Wei, Y., Han, Z., Fu, H., Peng, X., Deng, C., Hu, Q., Xu, C., Wen, J., Hu, D., and Zhang, C. Multimodal fusion on low-quality data: A comprehensive survey, 2024.

Citation

MLA
Hua, C., et al. “ReconBoost: Boosting Can Achieve Modality Reconcilement”. arXiv, 2024, http://arxiv.org/abs/2405.09321v1.
APA
Hua, C., Xu, Q., Bao, S., Yang, Z., & Huang, Q. (2024). ReconBoost: Boosting Can Achieve Modality Reconcilement. arXiv. http://arxiv.org/abs/2405.09321v1
Chicago
Hua, C., Q. Xu, S. Bao, Z. Yang, and Q. Huang. 2024. “ReconBoost: Boosting Can Achieve Modality Reconcilement”. arXiv. http://arxiv.org/abs/2405.09321v1.
Harvard
Hua, C. et al. (2024) “ReconBoost: Boosting Can Achieve Modality Reconcilement”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2405.09321v1.
Vancouver
1. Hua C, Xu Q, Bao S, Yang Z, Huang Q (2024) ReconBoost: Boosting Can Achieve Modality Reconcilement. arXiv

BibTeX

@article{hua2024reconboost,
  title = {ReconBoost: Boosting Can Achieve Modality Reconcilement},
  author = {Hua, Cong and Xu, Qianqian and Bao, Shilong and Yang, Zhiyong and Huang, Qingming},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2405.09321v1},
  eprint = {2405.09321}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/