MAP: Multimodal Uncertainty-Aware Vision-Language Pre-training Model

Yatai JiJunjie WangYuan GongLin ZhangYanru ZhuHongfa WangJiaxing ZhangTetsuya SakaiYujiu Yang

article2023CVPR65 citations

Proposes a vision-language pre-training framework that models multimodal features as Gaussian distributions instead of deterministic points to capture inter- and intra-modal semantic uncertainty across downstream tasks like visual reasoning and image-text retrieval.

Listen

Artificial intelligence systems for multimodal understanding face significant challenges with ambiguity and noise across visual and textual data. In real-world scenarios, an image region often contains multiple objects, and words can hold multiple meanings or synonyms, creating uncertainty both within individual modalities and between them. Most conventional vision-language models treat concepts as single fixed points in representation space, which fails to capture complex conceptual hierarchies and restricts the diversity of model predictions.

The main objective of the article is to develop and evaluate a vision-language pre-training framework that explicitly models semantic uncertainty by representing multimodal features as probability distributions rather than fixed points.

To achieve this, the article introduces the Multimodal Uncertainty-Aware Vision-Language Pre-training (MAP) model, centered around a newly designed Probability Distribution Encoder. This module models text tokens and image patches as multivariate Gaussian distributions by incorporating both sequence-level and feature-level interactions. The authors establish three distribution-based pre-training tasks to handle cross-modal alignment on large-scale unlabeled datasets: contrastive learning, masked language modeling, and image-text matching. Pre-trained on standard multimodal benchmarks—including MSCOCO, Visual Genome, SBU, and Conceptual Captions—the framework was subsequently fine-tuned and tested across multiple downstream tasks, including image-text retrieval, visual question answering, visual reasoning, and visual entailment.

The evaluations yielded several key findings. First, MAP achieved state-of-the-art performance across downstream tasks, notably outperforming comparable base-size models on visual question answering (78.03 on VQA2.0 test-dev) and visual reasoning (83.30 on NLVR2 dev). Second, in image retrieval on MSCOCO, MAP surpassed competitive baselines including ALBEF, even exceeding variants trained on significantly larger datasets of over 10 million images. Third, ablation analyses confirmed that distribution representations consistently outperformed deterministic point representations across all tasks, with masked language modeling proving to be the most vital pre-training objective. Finally, qualitative and toy analyses revealed that the learned distributions capture semantic overlap effectively and allow models to generate multiple diverse, plausible answers through sampling rather than being restricted to a single rigid prediction.

These results demonstrate that capturing semantic uncertainty substantially improves cross-modal understanding, representation robustness, and predictive flexibility. In practice, adopting distribution-based modeling allows organizations to achieve superior task accuracy without relying strictly on massive parameter scaling or exorbitantly large pre-training datasets, offering efficiency gains in data utilization. Furthermore, the capacity to produce diverse, valid outputs is particularly valuable for applications requiring nuanced reasoning or open-ended interactions.

Based on these findings, teams developing vision-language systems should consider adopting probabilistic encoders to handle real-world ambiguity in multimodal data. The article recommends utilizing Softmax-based sequence interactions within distribution encoders to best capture relational dependencies. Before broad commercial deployment, organizations should conduct pilot studies and further research exploring additional distribution subspaces and testing performance on larger, more diverse datasets. While confidence in the reported experimental improvements is high due to comprehensive baseline comparisons and statistical validation, current evaluations remain bounded by the specific pre-training corpora and distribution assumptions studied in the article.

Cover for MAP: Multimodal Uncertainty-Aware Vision-Language Pre-training Model

Abstract

Multimodal semantic understanding often has to deal with uncertainty, which means the obtained messages tend to refer to multiple targets. Such uncertainty is problematic for our interpretation, including inter- and intra-modal uncertainty. Little effort has studied the modeling of this uncertainty, particularly in pre-training on unlabeled datasets and fine-tuning in task-specific downstream datasets. In this paper, we project the representations of all modalities as probabilistic distributions via a Probability Distribution Encoder (PDE) by utilizing sequence-level interactions. Compared to the existing deterministic methods, such uncertainty modeling can convey richer multimodal semantic information and more complex relationships. Furthermore, we integrate uncertainty modeling with popular pre-training frameworks and propose suitable pre-training tasks: Distribution-based Vision-Language Contrastive learning (D-VLC), Distribution-based Masked Language Modeling (D-MLM), and Distribution-based Image-Text Matching (D-ITM). The fine-tuned models are applied to challenging downstream tasks, including image-text retrieval, visual question answering, visual reasoning, and visual entailment, and achieve state-of-the-art results.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 2.1. Probability Distribution Representations
  • 2.2. Vision-Language Pre-training (VLP)
  • 3. Approaches
  • 3.1. Model Overview
  • 3.1.1 Probability Distribution Encoder (PDE)
  • 3.1.2 Feature Extraction
  • 3.1.3 Cross-modal Transformer
  • 3.2. Distribution-based Pre-Training Tasks
  • 3.2.1 Coarse-grained Pre-training
  • 3.2.2 Fine-grained Pre-training
  • 3.2.3 Training Objectives
  • 4. Experiments
  • 4.1. Experimental settings
  • 4.2. Results of VL Downstream Tasks
  • 4.2.1 Evaluation on Image-Text Retrieval
  • 4.2.2 Evaluation on VQA2.0, NVLR2, and SNLI-VE
  • 4.3. Ablation Studies
  • 4.3.1 How do the probability distribution representations affect VL downstream task?
  • 4.3.2 How does the structure of PDE behave?
  • 4.3.3 What is the performance of different pre-training objectives?
  • 4.3.4 Does the number of layers of cross-modal transformer matter?
  • 4.4. Uncertainty Modeling Analysis
  • 5. Conclusions
  • 6. Acknowledgements
  • References

Knowls

  1. Knowl 1 — MAP architecture for multimodal uncertainty-aware pre-training

    model/method

    MAP, the Multimodal uncertainty-Aware vision-language Pre-training model, converts both visual and textual representations into token-level multivariate Gaussian distributions and trains them with distribution-based cross-modal objectives. A CLIP-ViT image encoder produces an image sequence {v[CLS],v1,…,vN}\{v_{[CLS]},v_1,\ldots,v_N\}, where v[CLS]v_{[CLS]} summarizes the image, and a RoBERTa-Base language encoder produces a text sequence {w[CLS],w1,…,wM}\{w_{[CLS]},w_1,\ldots,w_M\}.

    MAP uses a dual-stream cross-modal Transformer with NLN_L layers. In each layer, vision and language features first undergo separate self-attention, followed by bidirectional cross-attention: language queries attend to vision keys and values, while vision queries attend to language keys and values. A Probability Distribution Encoder (PDE) is applied after the unimodal encoders to produce distributions for coarse-grained alignment, and again after cross-modal fusion to produce distributions for masked language modeling, image-text matching, and downstream tasks.

    The dual-stream design is used because image patch sequences are substantially longer than text sequences; calculating self-attention separately avoids combining the two sequence lengths in a single attention operation.

  2. Knowl 2 — Probability Distribution Encoder with feature- and sequence-level interactions

    model/method

    For every visual or textual token, the Probability Distribution Encoder (PDE) predicts a mean vector μ\mu and a variance vector σ2\sigma^2, defining a diagonal-covariance Gaussian distribution N(μ,diag⁡(σ2))\mathcal N(\mu,\operatorname{diag}(\sigma^2)). The mean gives the distribution center, while the variance expresses uncertainty independently along each representation dimension.

    Given input hidden states H∈RT×DH\in\mathbb R^{T\times D}, where TT is sequence length and DD is hidden size, PDE splits the representation into kk attention heads and separates each head into a mean path and a variance path. For the mean path of head i∈{1,…,k}i\in\{1,\ldots,k\}, the input is Hμ(i)∈RT×D/(2k)H_\mu^{(i)}\in\mathbb R^{T\times D/(2k)}. With dk=D/(2k)d_k=D/(2k), the head computation is

    [Qμ(i),Kμ(i),Vμ(i)]=Hμ(i)Wqkv,[Q_\mu^{(i)},K_\mu^{(i)},V_\mu^{(i)}]=H_\mu^{(i)}W_{qkv}, Head⁡μ(i)=Act⁡ ⁣(Qμ(i)(Kμ(i))⊤dk)Vμ(i),\operatorname{Head}_\mu^{(i)}=\operatorname{Act}\!\left(\frac{Q_\mu^{(i)}(K_\mu^{(i)})^\top}{\sqrt{d_k}}\right)V_\mu^{(i)}, MHμ=concat⁡i=1k[Head⁡μ(i)]WO.MH_\mu=\operatorname{concat}_{i=1}^{k}\left[\operatorname{Head}_\mu^{(i)}\right]W_O.

    Here QQ, KK, and VV are query, key, and value matrices; Wqkv∈Rdk×3dkW_{qkv}\in\mathbb R^{d_k\times3d_k} projects each head into its three attention components; WO∈Rkdk×DW_O\in\mathbb R^{kd_k\times D} projects concatenated heads back to the hidden dimension; and Act⁡\operatorname{Act} consists of an activation and normalization operation. The variance path uses the same structure. A feed-forward transformation models feature-level interactions, while the multi-head operation models sequence-level interactions. The original point feature is added to the learned mean path because the point representation is expected to correlate with the Gaussian center.

  3. Knowl 3 — Distribution-based vision-language contrastive learning

    equation

    D-VLC aligns whole-image and whole-text distributions before cross-modal fusion. Let N(μ1,Σ1)\mathcal N(\mu_1,\Sigma_1) and N(μ2,Σ2)\mathcal N(\mu_2,\Sigma_2) be two multivariate Gaussian distributions, where μj\mu_j is a mean vector and Σj\Sigma_j is a covariance matrix. MAP measures their distributional separation using the squared 2-Wasserstein distance:

    D2W=∥μ1−μ2∥22+Tr⁡ ⁣(Σ1+Σ2−2(Σ11/2Σ2Σ11/2)1/2).D_{2W}=\|\mu_1-\mu_2\|_2^2+\operatorname{Tr}\!\left(\Sigma_1+\Sigma_2-2(\Sigma_1^{1/2}\Sigma_2\Sigma_1^{1/2})^{1/2}\right).

    MAP uses diagonal covariance matrices with diagonal standard-deviation vector σj\sigma_j, so the distance becomes

    D2W=∥μ1−μ2∥22+∥σ1−σ2∥22.D_{2W}=\|\mu_1-\mu_2\|_2^2+\|\sigma_1-\sigma_2\|_2^2.

    For an image II and text TT, the similarity between their unimodal [CLS][CLS] distributions is

    s(I,T)=a D2W(v[CLS],w[CLS])+b,s(I,T)=a\,D_{2W}(v_{[CLS]},w_{[CLS]})+b,

    where a<0a<0 is a learned or optimized negative scale factor, bb is a shift, and v[CLS]v_{[CLS]} and w[CLS]w_{[CLS]} denote the image and text Gaussian representations. For a batch {(Ii,Ti)}i=1N\{(I_i,T_i)\}_{i=1}^{N} of matched pairs, the bidirectional InfoNCE losses are

    LNCEI→T(i)=−log⁡exp⁡(s(Ii,Ti)/τ)∑n=1Nexp⁡(s(Ii,Tn)/τ),\mathcal L_{\mathrm{NCE}}^{I\to T}(i)=-\log\frac{\exp(s(I_i,T_i)/\tau)}{\sum_{n=1}^{N}\exp(s(I_i,T_n)/\tau)}, LNCET→I(i)=−log⁡exp⁡(s(Ti,Ii)/τ)∑n=1Nexp⁡(s(Ti,In)/τ),\mathcal L_{\mathrm{NCE}}^{T\to I}(i)=-\log\frac{\exp(s(T_i,I_i)/\tau)}{\sum_{n=1}^{N}\exp(s(T_i,I_n)/\tau)},

    where τ>0\tau>0 is a learned temperature. Summing the two directions gives the D-VLC loss, which contrasts the NN matched pairs against the N(N−1)N(N-1) mismatched pairs.

  4. Knowl 4 — Distribution-based masked language modeling and image-text matching

    model/method

    After fine-grained cross-modal interaction, MAP uses stochastic samples from the predicted Gaussian distributions for two objectives. In D-MLM, each input word is replaced by [MASK] with probability 15%15\%. If μ\mu is the predicted mean for a masked token, z(i)z^{(i)} is its ii-th sampled representation, KK is the number of Gaussian samples, yy is the original masked-word class, and ϕ\phi is the masked-language classifier, the training loss is

    LD-MLM=1K+1(CE⁡(ϕ(μ),y)+∑i=1KCE⁡(ϕ(z(i)),y)).\mathcal L_{\mathrm{D\text{-}MLM}}=\frac{1}{K+1}\left(\operatorname{CE}(\phi(\mu),y)+\sum_{i=1}^{K}\operatorname{CE}(\phi(z^{(i)}),y)\right).

    At inference, the token prediction is the mean of the classifier outputs from the mean representation and all sampled representations:

    P=1K+1(ϕ(μ)+∑i=1Kϕ(z(i))).P=\frac{1}{K+1}\left(\phi(\mu)+\sum_{i=1}^{K}\phi(z^{(i)})\right).

    D-ITM predicts whether an image-text pair is matched. Let vμv_\mu and wμw_\mu be the mean vectors of the image and text [CLS][CLS] distributions, and let v(i)v^{(i)} and w(i)w^{(i)} be corresponding samples. For binary label yy and image-text classifier ϕ\phi, the loss is

    LD-ITM=1K+1(CE⁡(ϕ(concat⁡[vμ,wμ]),y)+∑i=1KCE⁡(ϕ(concat⁡[v(i),w(i)]),y)).\mathcal L_{\mathrm{D\text{-}ITM}}=\frac{1}{K+1}\left(\operatorname{CE}(\phi(\operatorname{concat}[v_\mu,w_\mu]),y)+\sum_{i=1}^{K}\operatorname{CE}(\phi(\operatorname{concat}[v^{(i)},w^{(i)}]),y)\right).

    Matched dataset pairs are positive examples; negative examples are made by randomly replacing the image or the text. MAP obtains differentiable samples through the reparameterization rule z=μ+σϵz=\mu+\sigma\epsilon, where ϵ∼N(0,I)\epsilon\sim\mathcal N(0,I), σ\sigma is the standard-deviation vector, and II is the identity covariance matrix.

  5. Knowl 5 — Entropy regularization prevents Gaussian variance collapse

    equation

    Training D-MLM and D-ITM only with sampled representations can cause variance collapse: all samples converge toward the same optimal point, reducing Gaussian distributions to nearly deterministic point representations. MAP therefore penalizes distributions whose entropy is below a threshold γ\gamma:

    Lreg=max⁡(0,γ−h(N(μ,σ2))).\mathcal L_{\mathrm{reg}}=\max\left(0,\gamma-h(\mathcal N(\mu,\sigma^2))\right).

    For a dd-dimensional Gaussian with diagonal covariance Σ=diag⁡(σ2)\Sigma=\operatorname{diag}(\sigma^2), the entropy is

    h(N(μ,Σ))=12log⁡ ⁣(det⁡(2πeΣ))=12∑i=1dlog⁡(2πe σi2)=d2(log⁡(2π)+1)+∑i=1dlog⁡σi,h(\mathcal N(\mu,\Sigma))=\frac{1}{2}\log\!\left(\det(2\pi e\Sigma)\right) =\frac{1}{2}\sum_{i=1}^{d}\log(2\pi e\,\sigma_i^2) =\frac{d}{2}\bigl(\log(2\pi)+1\bigr)+\sum_{i=1}^{d}\log\sigma_i,

    where σi\sigma_i is the standard deviation in dimension ii and ee is Euler's number. The complete pre-training objective is

    Lpre=LD-MLM+LD-ITM+LD-VLC+αLreg,\mathcal L_{\mathrm{pre}}=\mathcal L_{\mathrm{D\text{-}MLM}}+\mathcal L_{\mathrm{D\text{-}ITM}}+\mathcal L_{\mathrm{D\text{-}VLC}}+\alpha\mathcal L_{\mathrm{reg}},

    where α\alpha weights the entropy regularizer.

  6. Knowl 6 — Pre-training and fine-tuning configuration

    experimental setup

    MAP uses hidden size 768768, 1212 attention heads in the multi-head attention modules, and NL=6N_L=6 cross-modal Transformer layers unless otherwise specified. Images are processed as square crops; the reported settings use 384×384384\times384 image inputs, while the pre-training procedure specifically resizes and crops images to 288×288288\times288. The image patch size is 1616, and text sequences are limited to 5050 tokens. The PDE uses k=6k=6 heads and Softmax as its default sequence-interaction activation.

    Pre-training uses MSCOCO, Visual Genome, SBU, and Conceptual Captions (CC-3M), with D-MLM, D-ITM, and D-VLC as the training tasks. MAP is fine-tuned and evaluated on image-text retrieval using MSCOCO and Flickr30K, visual question answering using VQA2.0, visual reasoning using NLVR2, and visual entailment using SNLI-VE. The experiments compare models grouped by whether their pre-training uses fewer or more than 1010 million images. The authors report randomized Tukey HSD pp-values and effect sizes based on one-way ANOVA for statistical analysis of the experimental results.

  7. Knowl 7 — MAP improves image-text retrieval and four vision-language tasks

    empirical result

    MAP achieves the strongest or near-strongest results across the evaluated downstream tasks. On the MSCOCO 55K test set, MAP obtains image-to-text retrieval scores of IR@1/5/10 =60.9/86.2/93.1=60.9/86.2/93.1 and text-to-image scores of TR@1/5/10 =79.3/94.8/97.6=79.3/94.8/97.6. On the Flickr30K 11K test set, it obtains IR@1/5/10 =83.8/97.2/98.7=83.8/97.2/98.7 and TR@1/5/10 =94.9/99.5/99.8=94.9/99.5/99.8. MAP outperforms ALBEF trained on 1414 million images on every MSCOCO retrieval metric and is either best or approximately 0.10.1 point behind the best result on Flickr30K.

    On VQA2.0, MAP scores 78.0378.03 on test-dev, 83.3083.30 on dev, and 83.4883.48 on test. On NLVR2, it scores 81.4081.40 on dev and 81.3981.39 on test-p. On SNLI-VE, it scores 81.4081.40 on validation and 81.3981.39 on test. Relative to METER, MAP improves VQA2.0 test-dev by 0.350.35 points and SNLI-VE validation by 0.540.54 points; relative to VLMo-Base, it improves NLVR2 dev by 0.530.53 points. The paper also reports that MAP exceeds SimVLM-Base, despite SimVLM using 1.81.8 billion pre-training images, on all three task families represented in the comparison.

  8. Knowl 8 — Distribution representations improve downstream understanding

    empirical result

    Removing the PDE and replacing distribution-based objectives with point-representation MLM and ITM reduces downstream performance. With random initialization, the point-based model scores 72.0972.09, 75.9175.91, 76.2876.28, 50.8650.86, and 51.0751.07 on VQA2.0 test-dev, SNLI-VE validation, SNLI-VE test, NLVR2 dev, and NLVR2 test-p, respectively; the full MAP scores 73.3573.35, 76.6776.67, 76.8676.86, 51.1251.12, and 51.0751.07 on the same metrics.

    With MSCOCO pre-training, the model without PDE scores 74.5774.57, 79.4279.42, 79.8479.84, 77.7277.72, and 79.3179.31, whereas MAP scores 75.0175.01, 80.0580.05, 80.3180.31, 78.9678.96, and 79.6479.64. Thus, modeling token and sentence representations as distributions rather than points improves most of the reported visual-language metrics under both random initialization and pre-trained initialization.

  9. Knowl 9 — Sequence interaction and complementary objectives are important ablations

    empirical result

    On VQA2.0 test-dev, the PDE design with Softmax sequence interaction scores 73.3573.35, compared with 72.0172.01 for an MLP-only PDE without sequence-level interaction. Alternative PDE variants score 69.7069.70 with ReLU plus normalization, 70.5370.53 with ReLU2^2 plus normalization, and 73.3473.34 with Sigmoid plus normalization. The results support using sequence-level token interaction and Softmax in PDE.

    When pre-training on MSCOCO, the reported scores on VQA2.0 test-dev, SNLI-VE test-p, and NLVR2 test are 73.35/76.86/51.0773.35/76.86/51.07 with random initialization; 75.01/80.31/79.6475.01/80.31/79.64 with D-MLM and D-ITM; 75.06/80.12/77.9075.06/80.12/77.90 with D-MLM and D-VLC; 71.02/78.54/73.6471.02/78.54/73.64 with D-ITM and D-VLC; and 75.16/80.39/79.4775.16/80.39/79.47 with all three objectives. Configurations without D-MLM are consistently weakest in this comparison. D-VLC helps VQA2.0 more than D-ITM, whereas D-ITM helps SNLI-VE and NLVR2 more than D-VLC.

    Changing the number of cross-modal Transformer layers gives VQA2.0 test-dev scores of 72.7172.71, 73.3273.32, 73.3573.35, and 73.3173.31 for 22, 44, 66, and 88 layers under random initialization, and 73.7873.78, 74.7374.73, 75.1675.16, and 75.2675.26 after pre-training. Pre-training therefore permits a small gain from increasing the depth from 66 to 88 layers, while random initialization peaks at 66 layers.

  10. Knowl 10 — Gaussian representations encode semantic overlap and diverse predictions

    empirical result

    In a two-dimensional visualization, MAP represents each image or caption distribution as an ellipse covering 95%95\% confidence. Distributions with similar semantics cluster together, and related image-caption ellipses have similar shapes and overlapping regions. In one example, the ellipse for an image containing a subset of the content of another image is almost fully contained within the latter image's ellipse; the shared intersections among several images and their captions correspond to the concept of a young boy.

    Sampling from the learned distributions also produces multiple plausible answers instead of a single deterministic output. For the question “What is this food?”, MAP samples “Dessert,” “Cake,” and “Muffin,” while MAP without PDE outputs only “Cake.” For “What is in the man's ear?”, MAP produces “Headphones,” “Headset,” and “Earbuds,” while the point model outputs only “Headphones.” For “Where are the people flying kites?”, MAP produces “Field,” “Park,” and “Grass,” while the point model outputs only “Field.” These examples demonstrate the paper's intended use of distributional uncertainty for representing semantic ambiguity and generating diverse valid predictions.

Coverage note — Detailed appendix implementation settings, statistical-test calculations, ethical discussion, and additional visualization cases were omitted because they support the central method and results rather than adding separate load-bearing contributions.

References

  1. 1.Haifa Alwahaby, Mutlu Cukurova, Zacharoula Papamitsiou, and Michail Giannakos. The evidence of impact and ethical considerations of multimodal learning analytics: A systematic literature review. The Multimodal Learning Analytics Handbook, pages 289–325, 2022. 14
  2. 2.Ben Athiwaratkun and Andrew Gordon Wilson. Multimodal word distributions. In Proc. of ACL, 2017. 2
  3. 3.Shruti Bhargava and David A. Forsyth. Exposing and correcting the gender bias in image captioning datasets and models. CoRR, abs/1912.00578, 2019. 14
  4. 4.Jie Chang, Zhonghao Lan, Changmao Cheng, and Yichen Wei. Data uncertainty learning in face recognition. In Proc. of CVPR, 2020. 2
  5. 5.Soravit Changpinyo, Piyush Sharma, Nan Ding, and Radu Soricut. Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. In CVPR, 2021. 11
  6. 6.Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. Uniter: Universal image-text representation learning. In Proc. of ECCV, 2020. 2, 4, 6, 12, 14
  7. 7.Sanghyuk Chun, Seong Joon Oh, Rafael Sampaio de Rezende, Yannis Kalantidis, and Diane Larlus. Probabilistic embeddings for cross-modal retrieval. In Proc. of CVPR, 2021. 1, 2, 7, 13
  8. 8.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: pre-training of deep bidirectional transformers for language understanding. In Proc. of NAACL, 2019. 2, 11, 12
  9. 9.Zi-Yi Dou, Yichong Xu, Zhe Gan, Jianfeng Wang, Shuohang Wang, Lijuan Wang, Chenguang Zhu, Pengchuan Zhang, Lu Yuan, Nanyun Peng, Zicheng Liu, and Michael Zeng. An empirical study of training end-to-end vision-and-language transformers. In Proc. of CVPR, 2022. 5, 6, 11, 12, 14
  10. 10.Zhe Gan, Yen-Chun Chen, Linjie Li, Chen Zhu, Yu Cheng, and Jingjing Liu. Large-scale adversarial training for vision-and-language representation learning. In NeurIPS, 2020. 6, 12
  11. 11.FranËois Garderes, Maryam Ziaeefard, Baptiste Abeloos, and `Freddy Lecue. ConceptBert: Concept-aware representation for visual question answering. In Proc. of EMNLP Findings, 2020. 2
  12. 12.Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering. In Proc. of CVPR, 2017. 6, 8, 13
  13. 13.Wenzhong Guo, Jianwen Wang, and Shiping Wang. Deep multimodal representation learning: A survey. IEEE Access, 2019. 1
  14. 14.Eyad Hakami and Davinia Hernandez Leo. How are learning  analytics considering the societal values of fairness, accountability, transparency and human well-being?: A literature review. MartÂınez-Mones A,  Alvarez A, Caeiro-Rodr  Âıguez M, Dimitriadis Y, editors. *LASI-SPAIN 2020: Learning Analytics Summer Institute Spain 2020: Learning Analytics. Time for Adoption?; 2020 Jun 15-16; Valladolid, Spain. Aachen: CEUR; 2020. p. 121-41, 2020. 14
  15. 15.Yicong Hong, Qi Wu, Yuankai Qi, Cristian Rodriguez Opazo, and Stephen Gould. VLN BERT: A recurrent vision-and-language BERT for navigation. In Proc. of CVPR, 2021. 3
  16. 16.Zhicheng Huang, Zhaoyang Zeng, Yupan Huang, Bei Liu, Dongmei Fu, and Jianlong Fu. Seeing out of the box: End-to-end pre-training for vision-language representation learning. In Proc. of CVPR, 2021. 3
  17. 17.Drew A. Hudson and Christopher D. Manning. GQA: A new dataset for real-world visual reasoning and compositional question answering. In CVPR, 2019. 11
  18. 18.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V. Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In Proc. of ICML, 2021. 2, 3, 11
  19. 19.Chenchen Jing, Yuwei Wu, Xiaoxun Zhang, Yunde Jia, and Qi Wu. Overcoming language priors in vqa via decomposed linguistic representations. In Proc. of AAAI, 2020. 2
  20. 20.Aishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve, Ishan Misra, and Nicolas Carion. MDETR - modulated detection for end-to-end multi-modal understanding. In Proc. of ICCV, 2021. 3
  21. 21.Leonid V Kantorovich. Mathematical methods of organizing and planning production. Management science, 1960. 4
  22. 22.Leonid V Kantorovich. On the translocation of masses. Journal of mathematical sciences, 2006. 4
  23. 23.Wonjae Kim, Bokyung Son, and Ildoo Kim. Vilt: Vision-and-language transformer without convolution or region supervision. In Proc. of ICML, 2021. 2, 3, 6, 11, 12
  24. 24.Diederik P. Kingma and Max Welling. Auto-encoding variational bayes. In Proc. of ICLR, 2014. 5
  25. 25.Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalan- tidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Li Fei-Fei. Visual genome: Connecting language and vision using crowdsourced dense image annotations. Int. J. Comput. Vis., 2017. 5, 11
  26. 26.Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi. Align before fuse: Vision and language representation learning with momentum distillation. Proc. of NeurIPS, 2021. 3, 4, 6, 12
  27. 27.Wei Li, Can Gao, Guocheng Niu, Xinyan Xiao, Hao Liu, Jiachen Liu, Hua Wu, and Haifeng Wang. UNIMO: towards unified-modal understanding and generation via cross-modal contrastive learning. In Proc. of ACL, 2021. 6, 12
  28. 28.Xiang Li, Luke Vilnis, Dongxu Zhang, Michael Boratko, and Andrew McCallum. Smoothing the geometry of probabilistic box embeddings. In Proc. of ICLR, 2019. 2
  29. 29.Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, Yejin Choi, and Jianfeng Gao. Oscar: Object-semantics aligned pre-training for vision-language tasks. In Proc. of ECCV, 2020. 6, 12
  30. 30.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence  Zitnick. Microsoft coco: Common objects in context. In Proc. of ECCV, 2014. 1, 5, 6, 8, 11, 13, 14
  31. 31.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized BERT pretraining approach. CoRR, 2019. 4
  32. 32.Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In Proc. of NeurIPS, 2019. 4, 14
  33. 33.Anton Mallasto and Aasa Feragen. Learning from uncertain curves: The 2-wasserstein metric for gaussian processes. Proc. of NeurIPS, 2017. 4
  34. 34.Vicente Ordonez, Girish Kulkarni, and Tamara L. Berg. Im2text: Describing images using 1 million captioned photographs. In NIPS, 2011. 5, 11
  35. 35.Bryan A. Plummer, Liwei Wang, Chris M. Cervantes, Juan C. Caicedo, Julia Hockenmaier, and Svetlana Lazebnik. Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models. In ICCV, 2015. 6, 11, 13
  36. 36.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In Proc. of ICML, 2021. 2, 3, 4
  37. 37.Tetsuya Sakai. Laboratory experiments in information retrieval: Sample Sizes, Effect Sizes, and Statistical Power. Springer, 2018. 6, 13
  38. 38.Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In ACL (1), 2018. 5, 11
  39. 39.Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. VL-BERT: pre-training of generic visual-linguistic representations. In Proc. of ICLR, 2020. 4
  40. 40.Yukun Su, Guosheng Lin, Ruizhou Sun, Yun Hao, and Qingyao Wu. Modeling the uncertainty for self-supervised 3d skeleton action representation learning. In Proc. of ACM MM, 2021. 2
  41. 41.Alane Suhr, Stephanie Zhou, Ally Zhang, Iris Zhang, Huajun Bai, and Yoav Artzi. A corpus for reasoning about natural language grounded in photographs. In Proc. of ACL, 2019. 6, 13
  42. 42.Jennifer J. Sun, Jiaping Zhao, Liang-Chieh Chen, Florian Schroff, Hartwig Adam, and Ting Liu. View-invariant probabilistic embedding for human pose. In Proc. of ECCV, 2020. 2
  43. 43.Hao Tan and Mohit Bansal. LXMERT: learning cross-modality encoder representations from transformers. In Proc. of EMNLP, 2019. 4
  44. 44.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Proc. of NeurIPS, 2017. 2, 11
  45. 45.Luke Vilnis and Andrew McCallum. Word representations via gaussian embedding. In Proc. of ICLR, 2015. 2
  46. 46.Junjie Wang, Yatai Ji, Jiaqi Sun, Yujiu Yang, and Tetsuya Sakai. Mirtt: Learning multimodal interaction representations from trilinear transformers for visual question answering. In Proc. of EMNLP Findings, 2021. 2
  47. 47.Wenhui Wang, Hangbo Bao, Li Dong, and Furu Wei. Vlmo: Unified vision-language pre-training with mixture-of-modality-experts. CoRR, 2021. 6, 12
  48. 48.Zirui Wang, Jiahui Yu, Adams Wei Yu, Zihang Dai, Yulia Tsvetkov, and Yuan Cao. Simvlm: Simple visual language model pretraining with weak supervision. CoRR, 2021. 6, 12
  49. 49.Ning Xie, Farley Lai, Derek Doran, and Asim Kadav. Visual entailment task for visually-grounded language learning. CoRR, 2018. 6, 13
  50. 50.Haiyang Xu, Ming Yan, Chenliang Li, Bin Bi, Songfang Huang, Wenming Xiao, and Fei Huang. E2E-VLP: end-to-end vision-language pre-training enhanced by visual learning. In Proc. of ACL, 2021. 2
  51. 51.Gengcong Yang, Jingyi Zhang, Yong Zhang, Baoyuan Wu, and Yujiu Yang. Probabilistic modeling of semantic ambiguity for scene graph generation. In Proc. of CVPR, 2021. 1, 2, 7
  52. 52.Gengcong Yang, Jingyi Zhang, Yong Zhang, Baoyuan Wu, and Yujiu Yang. Probabilistic modeling of semantic ambiguity for scene graph generation. In Proc. of CVPR, 2021. 2
  53. 53.Jinyu Yang, Jiali Duan, Son Tran, Yi Xu, Sampath Chanda, Liqun Chen, Belinda Zeng, Trishul Chilimbi, and Junzhou Huang. Vision-language pre-training with triple contrastive learning. In CVPR, pages 15650–15659. IEEE, 2022. 6, 12
  54. 54.Tianyuan Yu, Da Li, Yongxin Yang, Timothy M. Hospedales, and Tao Xiang. Robust person re-identification by modelling feature uncertainty. In Proc. of ICCV, 2019. 2, 7
  55. 55.Zhou Yu, Jun Yu, Yuhao Cui, Dacheng Tao, and Qi Tian. Deep modular co-attention networks for visual question answering. In Proc. of CVPR, 2019. 14
  56. 56.Pengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang, Lei Zhang, Lijuan Wang, Yejin Choi, and Jianfeng Gao. Vinvl: Revisiting visual representations in vision-language models. In Proc. of CVPR, 2021. 4, 6, 12

Citation

MLA
Ji, Y., et al. “MAP: Multimodal Uncertainty-Aware Vision-Language Pre-training Model”. arXiv, 2022, http://arxiv.org/abs/2210.05335v3.
APA
Ji, Y., Wang, J., Gong, Y., Zhang, L., Zhu, Y., Wang, H., Zhang, J., Sakai, T., & Yang, Y. (2022). MAP: Multimodal Uncertainty-Aware Vision-Language Pre-training Model. arXiv. http://arxiv.org/abs/2210.05335v3
Chicago
Ji, Y., J. Wang, Y. Gong, et al. 2022. “MAP: Multimodal Uncertainty-Aware Vision-Language Pre-training Model”. arXiv. http://arxiv.org/abs/2210.05335v3.
Harvard
Ji, Y. et al. (2022) “MAP: Multimodal Uncertainty-Aware Vision-Language Pre-training Model”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2210.05335v3.
Vancouver
1. Ji Y, Wang J, Gong Y, Zhang L, Zhu Y, Wang H, Zhang J, Sakai T, Yang Y (2022) MAP: Multimodal Uncertainty-Aware Vision-Language Pre-training Model. arXiv

BibTeX

@article{ji2022map,
  title = {MAP: Multimodal Uncertainty-Aware Vision-Language Pre-training Model},
  author = {Ji, Yatai and Wang, Junjie and Gong, Yuan and Zhang, Lin and Zhu, Yanru and Wang, Hongfa and Zhang, Jiaxing and Sakai, Tetsuya and Yang, Yujiu},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2210.05335v3},
  eprint = {2210.05335}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE