Characterizing and Overcoming the Greedy Nature of Learning in Multi-modal Deep Neural Networks

Nan WuStanislaw JastrzebskiKyunghyun ChoKrzysztof J. Geras

article2022ICML123 citations

Explains why multi-modal neural networks often over-rely on a single modality and introduces a training algorithm that balances learning speeds across modalities to improve overall generalization.

Listen

Modern artificial intelligence applications frequently rely on multi-modal deep neural networks to integrate diverse data streams, such as combining text with images or pairing standard video with depth sensors. Despite the intuitive expectation that combining multiple data sources should consistently yield better predictions than using a single source alone, multi-modal networks frequently perform worse than anticipated or even lag behind single-source models. The article investigates the root cause of this failure and introduces a practical method to overcome it.

The main objective of the article is to demonstrate that conventional joint training causes multi-modal neural networks to behave greedily by over-relying on the single data modality that is quickest to learn, while under-utilizing other valuable sources. To solve this, the article evaluates a new metric to measure real-time learning disparities across modalities and introduces a dynamic re-balancing training algorithm designed to ensure balanced utilization and improve overall generalization.

To conduct this evaluation, the researchers tested standard neural network architectures across three benchmark datasets representing different practical domains: a synthetic digit task (Colored-and-gray-MNIST), 3D object classification from multiple 2D camera angles (ModelNet40), and dynamic video gesture recognition using paired video and depth channels (NVGesture). Across these benchmarks, the researchers analyzed model behaviors under varying learning rates, weight regularizations, and architectural configurations, comparing their balanced training method against conventional training baselines and existing bias-reduction strategies.

The investigation produced several key findings. First, conventional multi-modal training consistently creates severe utilization imbalances; for example, in gesture recognition, the baseline model achieved a conditional utilization score of 0.63 for depth data but only 0.01 for standard video, effectively ignoring the visual channel. Second, applying stronger parameter regularization exacerbates this problem, causing the network to discard secondary modalities even more aggressively. Third, the proposed balanced multi-modal training algorithm successfully prevented this imbalance and achieved superior generalization across all benchmarks. Most notably, on the gesture recognition benchmark trained from scratch, the balanced method raised classification accuracy from 79.81% to 80.22%, and on the synthetic biased benchmark, accuracy increased dramatically from 45.26% under standard training to 91.01% with the proposed guided method.

These findings indicate that poor multi-modal performance is fundamentally an optimization failure rather than an inherent deficiency of the data or neural network architectures. Without targeted intervention during training, organizations deploying multi-modal systems risk wasting substantial resources on collecting and processing secondary sensor feeds that the deployed models quietly ignore. Adopting a dynamically balanced training approach mitigates this risk and ensures models extract meaningful signals across all available inputs.

For technical teams developing multi-modal systems, the article recommends incorporating real-time learning speed tracking into the training loop and applying targeted balancing steps whenever relative learning rates diverge beyond a specified threshold. When scaling to three or more data streams, teams should monitor pairwise differences and periodically accelerate the slowest-progressing modality. Further empirical testing on large-scale industrial datasets and diverse sensor combinations is recommended before organization-wide deployment.

The conclusions of the article carry high confidence across the tested classification benchmarks and intermediate fusion architectures. However, decision-makers should note that the evaluation primarily focused on two-modality classification pipelines and relatively small-to-moderate dataset scales. Additional validation may be required for late-fusion architectures, regression problems, or ultra-large foundation models.

Cover for Characterizing and Overcoming the Greedy Nature of Learning in Multi-modal Deep Neural Networks

Abstract

We hypothesize that due to the greedy nature of learning in multi-modal deep neural networks, these models tend to rely on just one modality while under-fitting the other modalities. Such behavior is counter-intuitive and hurts the models’ generalization, as we observe empirically. To estimate the model’s dependence on each modality, we compute the gain on the accuracy when the model has access to it in addition to another modality. We refer to this gain as the conditional utilization rate. In the experiments, we consistently observe an imbalance in conditional utilization rates between modalities, across multiple tasks and architectures. Since conditional utilization rate cannot be computed efficiently during training, we introduce a proxy for it based on the pace at which the model learns from each modality, which we refer to as the conditional learning speed. We propose an algorithm to balance the conditional learning speeds between modalities during training and demonstrate that it indeed addresses the issue of greedy learning.1 The proposed algorithm improves the model’s generalization on three datasets: Colored MNIST, ModelNet40, and NVIDIA Dynamic Hand Gesture.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Problem Setup
  • 4. The Greedy Learner Hypothesis
  • 4.1. Conditional Utilization Rate
  • 4.2. Multi-modal Learning Process is Greedy
  • 5. Making Multi-modal Learning Less Greedy
  • 5.1. Conditional Learning Speed
  • 5.2. Balanced Multi-modal Learning
  • 6. Experiments and Results
  • 6.1. Datasets, Tasks and Baselines
  • 6.2. Validating the Greedy Learner Hypothesis
  • 6.3. Strong Regularization Encourages Greediness
  • 6.4. Balanced Multi-modal Learning
  • 7. Discussion
  • Acknowledgements
  • References
  • A. Data Preparation
  • B. Supplementary Figures

Knowls

  1. Knowl 1 — Greedy Learner Hypothesis in Multi-Modal Deep Neural Networks

    assumption

    The greedy learner hypothesis posits that when a multi-modal deep neural network (DNN) is trained end-to-end to minimize the sum of modality-specific classification losses, the optimization process acts greedily: the network predominantly learns to rely on the single input modality from which it can learn the fastest, while failing to continue learning from and utilizing the other predictive modalities.

    This hypothesis assumes:

    1. Both modalities m0m_0 and m1m_1 contain predictive mutual information about target YY, i.e., I(Y;Xm0)>0I(Y; X_{m_0}) > 0 and I(Y;Xm1)>0I(Y; X_{m_1}) > 0.
    2. Different modalities provide varying predictive capabilities and the network learns from them at distinct rates.
    3. The greedy optimization leads to severe under-utilization of slower-to-learn modalities, which harms cross-modal synergy and out-of-distribution or test generalization.
  2. Knowl 2 — Conditional Utilization Rate for Multi-Modal Neural Networks

    definition

    For a multi-modal deep neural network ff taking two input modalities m0m_0 and m1m_1 via uni-modal branches ϕ0\phi_0 and ϕ1\phi_1 interconnected by intermediate fusion modules, the overall prediction is y^=12(y^0+y^1)\hat{y} = \frac{1}{2}(\hat{y}_0 + \hat{y}_1), where y^0=f0(xm0,xm1)\hat{y}_0 = f_0(x_{m_0}, x_{m_1}) and y^1=f1(xm0,xm1)\hat{y}_1 = f_1(x_{m_0}, x_{m_1}).

    To isolate each branch from receiving cross-modal information, the spatial pooled feature activations h0,h1h_0, h_1 passed into fusion gating function g([h0,h1])g([h_0, h_1]) are replaced with their training-set averages hˉ0=1n∑i=1nh0(xi)\bar{h}_0 = \frac{1}{n} \sum_{i=1}^n h_0(x^i) and hˉ1=1n∑i=1nh1(xi)\bar{h}_1 = \frac{1}{n} \sum_{i=1}^n h_1(x^i):

    w0,∗=g([hˉ0,hˉ1]),∗,w1=g([hˉ0,hˉ1])w_0, * = g([\bar{h}_0, \bar{h}_1]), \quad *, w_1 = g([\bar{h}_0, \bar{h}_1])

    This yields single-modality sub-networks f0′(xm0)f'_0(x_{m_0}) and f1′(xm1)f'_1(x_{m_1}). Let A(⋅)A(\cdot) denote classification accuracy on test dataset Dtest\mathcal{D}^{\text{test}}. The conditional utilization rates are defined as:

    u(m0∣m1)=A(f1)−A(f1′)A(f1)u(m_0|m_1) = \frac{A(f_1) - A(f'_1)}{A(f_1)}

    u(m1∣m0)=A(f0)−A(f0′)A(f0)u(m_1|m_0) = \frac{A(f_0) - A(f'_0)}{A(f_0)}

    The conditional utilization imbalance is quantified by:

    dutil(f)=u(m1∣m0)−u(m0∣m1)∈[−1,1]d_{\text{util}}(f) = u(m_1|m_0) - u(m_0|m_1) \in [-1, 1]

    A value of ∣dutil(f)∣|d_{\text{util}}(f)| close to 11 indicates extreme imbalance, where the network derives predictive capacity almost entirely from one modality.

  3. Knowl 3 — Conditional Learning Speed Metric

    definition

    Conditional learning speed is a real-time training proxy for conditional utilization rate that monitors the relative rate of parameter updates across modalities. For parameter group θ\theta at gradient step ii, the effective parameter update is:

    μ(θ;i)=∥G∥22∥θ(i)∥22,where G=∂L∂θ∣θ(i−1)\mu(\theta; i) = \frac{\|G\|_2^2}{\|\theta(i)\|_2^2}, \quad \text{where } G = \left.\frac{\partial L}{\partial \theta}\right|_{\theta(i-1)}

    Model parameters are partitioned into:

    • θ0\theta_0: parameters of uni-modal branch ϕ0\phi_0
    • θ0′\theta'_0: fusion module parameters contributing to output y^0\hat{y}_0
    • θ1\theta_1: parameters of uni-modal branch ϕ1\phi_1
    • θ1′\theta'_1: fusion module parameters contributing to output y^1\hat{y}_1

    After tt training steps, the conditional learning speeds are defined as:

    s(m1∣m0;t)=log⁡∑i=1tμ(θ0′;i)∑i=1tμ(θ0;i)s(m_1|m_0; t) = \log \frac{\sum_{i=1}^t \mu(\theta'_0; i)}{\sum_{i=1}^t \mu(\theta_0; i)}

    s(m0∣m1;t)=log⁡∑i=1tμ(θ1′;i)∑i=1tμ(θ1;i)s(m_0|m_1; t) = \log \frac{\sum_{i=1}^t \mu(\theta'_1; i)}{\sum_{i=1}^t \mu(\theta_1; i)}

    The learning speed difference is:

    dspeed(f;t)=s(m1∣m0;t)−s(m0∣m1;t)d_{\text{speed}}(f; t) = s(m_1|m_0; t) - s(m_0|m_1; t)

  4. Knowl 4 — Balanced Multi-Modal Learning Algorithm

    algorithm

    Balanced Multi-Modal Learning balances conditional learning speeds across modalities during training by dynamically interleaving regular multi-modal optimization steps with re-balancing steps targeted at the underutilized modality.

    Input: Re-balancing window size QQ, imbalance tolerance parameter α\alpha, number of epochs NN, number of updating steps per epoch nn
    Initiate: Mθ0,Mθ0′,Mθ1,Mθ1′←0,0,0,0M_{\theta_0}, M_{\theta'_0}, M_{\theta_1}, M_{\theta'_1} \leftarrow 0, 0, 0, 0; q←Qq \leftarrow Q
    for i=1i = 1 to nn do
        Take a regular step
        Mθ0←Mθ0+μ(θ0;i)M_{\theta_0} \leftarrow M_{\theta_0} + \mu(\theta_0; i)
        Mθ0′←Mθ0′+μ(θ0′;i)M_{\theta'_0} \leftarrow M_{\theta'_0} + \mu(\theta'_0; i)
        Mθ1←Mθ1+μ(θ1;i)M_{\theta_1} \leftarrow M_{\theta_1} + \mu(\theta_1; i)
        Mθ1′←Mθ1′+μ(θ1′;i)M_{\theta'_1} \leftarrow M_{\theta'_1} + \mu(\theta'_1; i)
    end for
    for i=1i = 1 to n×Nn \times N do
        if q==Qq == Q then
            Take a regular step
            Mθ0←Mθ0+μ(θ0;i)M_{\theta_0} \leftarrow M_{\theta_0} + \mu(\theta_0; i)
            Mθ0′←Mθ0′+μ(θ0′;i)M_{\theta'_0} \leftarrow M_{\theta'_0} + \mu(\theta'_0; i)
            Mθ1←Mθ1+μ(θ1;i)M_{\theta_1} \leftarrow M_{\theta_1} + \mu(\theta_1; i)
            Mθ1′←Mθ1′+μ(θ1′;i)M_{\theta'_1} \leftarrow M_{\theta'_1} + \mu(\theta'_1; i)
            dspeed←log⁡(Mθ0′/Mθ0)−log⁡(Mθ1′/Mθ1)d_{\text{speed}} \leftarrow \log(M_{\theta'_0}/M_{\theta_0}) - \log(M_{\theta'_1}/M_{\theta_1})
            if ∣dspeed∣>α|d_{\text{speed}}| > \alpha then
                q←1q \leftarrow 1
            end if
        else
            q←q+1q \leftarrow q + 1
            if dspeed>0d_{\text{speed}} > 0 then
                Take a re-balancing step to accelerate learning from m0m_0
            else
                Take a re-balancing step to accelerate learning from m1m_1
            end if
        end if
    end for

    The model executes standard regular steps in the first epoch as a warm-up. In subsequent epochs, whenever ∣dspeed∣>α|d_{\text{speed}}| > \alpha, the algorithm executes QQ consecutive re-balancing steps to boost the slower modality before returning to regular training.

  5. Knowl 5 — Fusion Feature Re-scaling for Modality Acceleration in Re-balancing Steps

    model/method

    In intermediate multi-modal fusion networks using Multi-Modal Transfer Modules (MMTM), feature maps A0∈RN1×⋯×NL×CA_0 \in \mathbb{R}^{N_1 \times \dots \times N_L \times C} and A1∈RM1×⋯×MJ×C′A_1 \in \mathbb{R}^{M_1 \times \dots \times M_J \times C'} are pooled into h0,h1h_0, h_1, passed to gating network [w0,w1]=g([h0,h1])[w_0, w_1] = g([h_0, h_1]), and rescaled during regular steps as:

    A~0=2×σ(w0)⊙A0,A~1=2×σ(w1)⊙A1\tilde{A}_0 = 2 \times \sigma(w_0) \odot A_0, \quad \tilde{A}_1 = 2 \times \sigma(w_1) \odot A_1

    where ⊙\odot is channel-wise multiplication and σ(⋅)\sigma(\cdot) is the sigmoid function.

    During a re-balancing step designed to accelerate learning for modality m0m_0 (when dspeed>0d_{\text{speed}} > 0), the activation vector w0w_0 applied to A0A_0 is replaced by the historical mean over the preceding ntn_t regular training samples wˉ0=1nt∑i=1ntw0(xi)\bar{w}_0 = \frac{1}{n_t} \sum_{i=1}^{n_t} w_0(x^i), while A1A_1 uses sample-specific gating:

    A~0=2×σ(wˉ0)⊙A0,A~1=2×σ(w1)⊙A1\tilde{A}_0 = 2 \times \sigma(\bar{w}_0) \odot A_0, \quad \tilde{A}_1 = 2 \times \sigma(w_1) \odot A_1

    Conversely, to accelerate learning from modality m1m_1 (when dspeed<0d_{\text{speed}} < 0), the feature rescaling replaces w1w_1 with wˉ1=1nt∑i=1ntw1(xi)\bar{w}_1 = \frac{1}{n_t} \sum_{i=1}^{n_t} w_1(x^i):

    A~0=2×σ(w0)⊙A0,A~1=2×σ(wˉ1)⊙A1\tilde{A}_0 = 2 \times \sigma(w_0) \odot A_0, \quad \tilde{A}_1 = 2 \times \sigma(\bar{w}_1) \odot A_1

  6. Knowl 6 — Empirical Modality Utilization Asymmetry in Multi-Modal Learning

    empirical result

    When multi-modal DNNs are trained with standard SGD across diverse tasks (Colored-and-gray-MNIST, ModelNet40, NVGesture):

    1. For networks supplied with two identical duplicated input modalities (e.g., duplicate gray-scale MNIST images, duplicated front-views in ModelNet40, or duplicated RGB in NVGesture), the empirical distribution of conditional utilization difference d^util\hat{d}_{\text{util}} is symmetric around zero, with expected value E[d^util]≈0.0\mathbb{E}[\hat{d}_{\text{util}}] \approx 0.0.
    2. For networks supplied with two distinct input modalities, d^util\hat{d}_{\text{util}} is strongly asymmetric, yielding E[d^util]\mathbb{E}[\hat{d}_{\text{util}}] of approximately 0.30.3 on Colored-and-gray-MNIST, 0.10.1 on ModelNet40, and 0.40.4 on NVGesture.
    3. The distribution of conditional learning speed difference d^speed\hat{d}_{\text{speed}} closely mirrors the asymmetry and distribution profile of d^util\hat{d}_{\text{util}}, confirming that unequal modality learning rates drive unequal conditional utilization.
  7. Knowl 7 — Effect of L1 Regularization on Multi-Modal Greediness

    empirical result

    Applying L1L_1 weight regularization L′=L+λ∥θ∥1L' = L + \lambda \|\theta\|_1 increases the greediness of multi-modal neural network optimization.

    In experiments on ModelNet40 with λ∈[10−9,10−3]\lambda \in [10^{-9}, 10^{-3}]:

    • Increasing λ\lambda increases model parameter sparsity R(f)R(f) (defined as the fraction of network parameters with absolute value <10−7< 10^{-7}).
    • For λ≥10−5\lambda \ge 10^{-5}, both the conditional utilization imbalance ∣dutil(f)∣|d_{\text{util}}(f)| and conditional learning speed difference ∣dspeed(f)∣|d_{\text{speed}}(f)| increase significantly and positively correlate with R(f)R(f), showing that stronger capacity regularization exacerbates the network's tendency to rely unilaterally on a single dominant modality.
  8. Knowl 8 — Generalization Performance of Balanced Multi-Modal Learning

    data/table

    Across synthetic and real-world multi-modal datasets, guided Balanced Multi-Modal Learning outperforms unimodal baselines, standard vanilla intermediate-fusion multi-modal training, RUBi re-weighting, and random re-balancing in test classification accuracy.

    Method Colored-and-gray-MNIST ModelNet40 NVGesture-scratch NVGesture-pretrained
    uni-modal (best) 99.14 ±\pm 0.11 89.34 ±\pm 0.39 77.59 ±\pm 0.55 78.98 ±\pm 2.02
    multi-modal (vanilla) 45.26 ±\pm 0.46 90.09 ±\pm 0.58 79.81 ±\pm 1.14 83.20 ±\pm 0.21
    + RUBi (Cadene et al., 2019) 44.79 ±\pm 0.62 90.45 ±\pm 0.58 79.95 ±\pm 0.12 81.60 ±\pm 1.28
    + random (proposed) 74.07 ±\pm 2.75 91.36 ±\pm 0.10 79.88 ±\pm 0.90 82.64 ±\pm 0.84
    + guided (proposed) 91.01 ±\pm 1.20 91.37 ±\pm 0.28 80.22 ±\pm 0.73 83.82 ±\pm 1.45

    On Colored-and-gray-MNIST, where monochromatic color is spurious and gray-scale digits are predictive, vanilla multimodal training collapses to 45.26%45.26\%, whereas guided balanced training recovers 91.01%91.01\% (with its gray-scale branch alone reaching 99.16±0.14%99.16 \pm 0.14\%). Guided balanced learning achieves the highest test accuracy across all benchmarks.

Coverage note — None was omitted; all key definitions, algorithmic mechanisms, theoretical conjectures, and empirical findings are fully covered.

References

  1. 1.Agrawal, A., Batra, D., and Parikh, D. Analyzing the behavior of visual question answering models. EMNLP, 2016.
  2. 2.Agrawal, A., Batra, D., Parikh, D., and Kembhavi, A. Don’t just assume; look and answer: Overcoming priors for visual question answering. In CVPR, 2018.
  3. 3.Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., and Zhang, L. Bottom-up and top-down attention for image captioning and visual question answering. CVPR, 2018.
  4. 4.Arnab, A., Dehghani, M., Heigold, G., Sun, C., Lućić, M., and Schmid, C. Vivit: A video vision transformer. arXiv:2103.15691, 2021.
  5. 5.Atrey, P. K., Hossain, M. A., El Saddik, A., and Kankanhalli, M. S. Multimodal fusion for multimedia analysis: a survey. Multimedia Systems, 2010.
  6. 6.Baltrušaitis, T., Ahuja, C., and Morency, L.-P. Multimodal machine learning: A survey and taxonomy. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2018.
  7. 7.Blum, A. and Mitchell, T. Combining labeled and unlabeled data with co-training. COLT, 1998.
  8. 8.Brock, A., De, S., Smith, S. L., and Simonyan, K. High-performance large-scale image recognition without normalization. ICML, 2021.
  9. 9.Cadene, R., Dancette, C., Ben-Younes, H., Cord, M., and Parikh, D. Rubi: Reducing unimodal biases in visual question answering. NeurIPS, 2019.
  10. 10.Cao, J., Gan, Z., Cheng, Y., Yu, L., Chen, Y.-C., and Liu, J. Behind the scene: Revealing the secrets of pre-trained vision-and-language models. ECCV, 2020.
  11. 11.Carreira, J. and Zisserman, A. Quo vadis, action recognition? a new model and the kinetics dataset. CVPR, 2017.
  12. 12.Gat, I., Schwartz, I., Schwing, A., and Hazan, T. Removing bias in multi-modal classifiers: Regularization by maximizing functional entropies. NeurIPS, 2020.
  13. 13.Gat, I., Schwartz, I., and Schwing, A. Perceptual score: What data modalities does your model perceive? NeurIPS, 34, 2021.
  14. 14.Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D. Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering. CVPR, 2017.
  15. 15.Han, X., Wang, S., Su, C., Huang, Q., and Tian, Q. Greedy gradient ensemble for robust visual question answering. ICCV, 2021a.
  16. 16.Han, Z., Zhang, C., Fu, H., and Zhou, J. T. Trusted multi-view classification. ICLR, 2021b.
  17. 17.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. CVPR, 2016.
  18. 18.Hessel, J. and Lee, L. Does my multimodal model learn cross-modal interactions? it’s harder to tell than you might think! EMNLP, 2020.
  19. 19.Hoffer, E., Banner, R., Golan, I., and Soudry, D. Norm matters: efficient and accurate normalization schemes in deep networks. NeurIPS, 2018.
  20. 20.Hu, J., Shen, L., and Sun, G. Squeeze-and-excitation networks. In CVPR, 2018.
  21. 21.Ioffe, S. and Szegedy, C. Batch normalization: Accelerating deep network training by reducing internal covariate shift. ICML, 2015.
  22. 22.Jabri, A., Joulin, A., and Van Der Maaten, L. Revisiting visual question answering baselines. ECCV, 2016.
  23. 23.Joze, H. R. V., Shaban, A., Iuzzolino, M. L., and Koishida, K. Mmtm: multimodal transfer module for cnn fusion. CVPR, 2020.
  24. 24.Kim, B., Kim, H., Kim, K., Kim, S., and Kim, J. Learning not to learn: Training deep neural networks with biased data. CVPR, 2019.
  25. 25.Lao, M., Guo, Y., Liu, Y., Chen, W., Pu, N., and Lew, M. S. From superficial to deep: Language bias driven curriculum learning for visual question answering. In Proceedings of the 29th ACM International Conference on Multimedia, pp. 3370–3379, 2021.
  26. 26.Lazaridou, A., Baroni, M., et al. Combining language and vision with a multimodal skip-gram model. HLT-NAACL, 2015.
  27. 27.LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 1998.
  28. 28.Li, L., Gan, Z., and Liu, J. A closer look at the robustness of vision-and-language pre-trained models. arXiv:2012.08673, 2020.
  29. 29.Liu, K., Li, Y., Xu, N., and Natarajan, P. Learn to combine modalities in multimodal deep learning. arXiv:1805.11730, 2018.
  30. 30.Molchanov, P., Gupta, S., Kim, K., and Kautz, J. Hand gesture recognition with 3d convolutional neural networks. CVPR, 2015.
  31. 31.Nair, V. and Hinton, G. E. Rectified linear units improve restricted boltzmann machines. In ICML, 2010.
  32. 32.Neverova, N., Wolf, C., Taylor, G., and Nebout, F. Moddrop: Adaptive multi-modal gesture recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2015.
  33. 33.Ngiam, J., Khosla, A., Kim, M., Nam, J., Lee, H., and Ng, A. Y. Multimodal deep learning. ICML, 2011.
  34. 34.Perez, E., Strub, F., De Vries, H., Dumoulin, V., and Courville, A. Film: Visual reasoning with a general conditioning layer. AAAI, 2018.
  35. 35.Pérez-Rúa, J.-M., Vielzeuf, V., Pateux, S., Baccouche, M., and Jurie, F. Mfas: Multimodal fusion architecture search. CVPR, 2019.
  36. 36.Sridharan, K. and Kakade, S. M. An information theoretic framework for multi-view learning. COLT, 2008.
  37. 37.Su, H., Maji, S., Kalogerakis, E., and Learned-Miller, E. G. Multi-view convolutional neural networks for 3d shape recognition. ICCV, 2015.
  38. 38.Sun, Y., Mai, S., and Hu, H. Learning to balance the learning rates between various modalities via adaptive tracking factor. IEEE Signal Processing Letters, 2021.
  39. 39.Tran, D., Bourdev, L., Fergus, R., Torresani, L., and Paluri, M. Learning spatiotemporal features with 3d convolutional networks. ICCV, 2015.
  40. 40.Van Laarhoven, T. L2 regularization versus batch and weight normalization. arXiv:1706.05350, 2017.
  41. 41.Wang, W., Tran, D., and Feiszli, M. What makes training multi-modal classification networks hard? CVPR, 2020a.
  42. 42.Wang, Y., Huang, W., Sun, F., Xu, T., Rong, Y., and Huang, J. Deep multimodal fusion by channel exchanging. NeurIPS, 2020b.
  43. 43.Weng, Z., Wu, Z., Li, H., and Jiang, Y.-G. Hms: Hierarchical modality selectionfor efficient video recognition. arXiv:2104.09760, 2021.
  44. 44.Winterbottom, T., Xiao, S., McLean, A., and Moubayed, N. A. On modality bias in the tvqa dataset. BMVC, 2020.
  45. 45.Wu, N., Jastrzębski, S., Park, J., Moy, L., Cho, K., and Geras, K. J. Improving the ability of deep networks to use information from multiple views in breast cancer screening. MIDL, 2020.
  46. 46.Wu, Z., Song, S., Khosla, A., Yu, F., Zhang, L., Tang, X., and Xiao, J. 3d shapenets: A deep representation for volumetric shapes. CVPR, 2015.
  47. 47.Zhang, G., Wang, C., Xu, B., and Grosse, R. Three mechanisms of weight decay regularization. ICLR, 2019.

Citation

MLA
Wu, N., et al. “Characterizing and Overcoming the Greedy Nature of Learning in Multi-modal Deep Neural Networks”. International Conference on Machine Learning, vol. 162, 2022, pp. 24043–55, https://proceedings.mlr.press/v162/wu22d.html.
APA
Wu, N., Jastrzebski, S., Cho, K., & Geras, K. J. (2022). Characterizing and Overcoming the Greedy Nature of Learning in Multi-modal Deep Neural Networks. International Conference on Machine Learning, 162, 24043–24055. https://proceedings.mlr.press/v162/wu22d.html
Chicago
Wu, N., S. Jastrzebski, K. Cho, and K. J. Geras. 2022. “Characterizing and Overcoming the Greedy Nature of Learning in Multi-modal Deep Neural Networks”. International Conference on Machine Learning 162: 24043–55. https://proceedings.mlr.press/v162/wu22d.html.
Harvard
Wu, N. et al. (2022) “Characterizing and Overcoming the Greedy Nature of Learning in Multi-modal Deep Neural Networks”, International Conference on Machine Learning. PMLR, pp. 24043–24055. Available at: https://proceedings.mlr.press/v162/wu22d.html.
Vancouver
1. Wu N, Jastrzebski S, Cho K, Geras KJ (2022) Characterizing and Overcoming the Greedy Nature of Learning in Multi-modal Deep Neural Networks. In: International Conference on Machine Learning. PMLR, pp 24043–24055

BibTeX

@InProceedings{pmlr-v162-wu22d,
  title = 	 {Characterizing and Overcoming the Greedy Nature of Learning in Multi-modal Deep Neural Networks},
  author =       {Wu, Nan and Jastrzebski, Stanislaw and Cho, Kyunghyun and Geras, Krzysztof J},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {24043--24055},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/wu22d/wu22d.pdf},
  url = 	 {https://proceedings.mlr.press/v162/wu22d.html},
  abstract = 	 {We hypothesize that due to the greedy nature of learning in multi-modal deep neural networks, these models tend to rely on just one modality while under-fitting the other modalities. Such behavior is counter-intuitive and hurts the models’ generalization, as we observe empirically. To estimate the model’s dependence on each modality, we compute the gain on the accuracy when the model has access to it in addition to another modality. We refer to this gain as the conditional utilization rate. In the experiments, we consistently observe an imbalance in conditional utilization rates between modalities, across multiple tasks and architectures. Since conditional utilization rate cannot be computed efficiently during training, we introduce a proxy for it based on the pace at which the model learns from each modality, which we refer to as the conditional learning speed. We propose an algorithm to balance the conditional learning speeds between modalities during training and demonstrate that it indeed addresses the issue of greedy learning. The proposed algorithm improves the model’s generalization on three datasets: Colored MNIST, ModelNet40, and NVIDIA Dynamic Hand Gesture.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/