Tackling Ambiguity with Images: Improved Multimodal Machine Translation and Contrastive Evaluation

Matthieu FuteralCordelia SchmidIvan LaptevBenoît SagotRachel Bawden

article2023ACL58 citations

Proposes an adapter-based multimodal machine translation framework with guided self-attention alongside CoMMuTE, a contrastive evaluation benchmark designed to verify whether models effectively use visual context to resolve lexical ambiguity.

Listen

Translating ambiguous text accurately remains a persistent challenge for automated translation systems. While accompanying images can provide the visual context necessary to resolve ambiguities, existing multimodal translation systems have struggled to outperform strong text-only baselines. Furthermore, the field has been hampered by standard benchmarks that rarely require image context to translate correctly, making it difficult to assess whether models are genuinely using visual information.

The article evaluates a new multimodal translation architecture and introduces a targeted evaluation benchmark to determine whether integrating visual context can resolve lexical ambiguity across multiple languages without degrading overall translation quality.

To achieve this, the authors developed VGAMT (Visually Guided and Adapted Machine Translation). The approach freezes a strong, pretrained text-only translation model (mBART fine-tuned on parallel corpora) and adapts it using lightweight neural adapter modules and a novel guided self-attention mechanism that pairs relevant text words directly to image regions. The architecture was jointly trained on multimodal machine translation and visually conditioned masked language modeling using approximately 2 million image-text pairs from the Conceptual Captions dataset alongside the standard Multi30k dataset. To properly evaluate ambiguity resolution, the authors created CoMMuTE, a contrastive evaluation dataset comprising 155 ambiguous English sentences paired with alternative images and native translations across French, German, and Czech.

The experimental findings demonstrate significant improvements in visual disambiguation. First, VGAMT outperformed existing multimodal systems and balanced text-only baselines on the CoMMuTE benchmark, achieving an accuracy of 67.1% in English-to-French, 59.0% in English-to-German, and 55.6% in English-to-Czech, where text-only baselines achieve only the random chance baseline of 50.0%. Second, on standard Multi30k benchmarks where visual context is rarely required, VGAMT performed competitively with strong text-only models, demonstrating that multimodal integration did not cause translation quality degradation. Third, ablation analyses revealed that joint training with masked language modeling and the use of global image features were critical; omitting masked pretraining or image features caused accuracy on ambiguous sentences to drop sharply toward baseline levels.

These results demonstrate that multimodal machine translation can effectively leverage visual context without requiring expensive full-model retraining from scratch. Utilizing parameter-efficient adapters and guided cross-modal attention allows organizations to preserve the broad linguistic capabilities of large text-only systems while improving translation accuracy in visually grounded environments. The findings also underscore that standard evaluation metrics and captioning datasets are inadequate for measuring multimodal performance, emphasizing the need for targeted contrastive benchmarks.

Based on these findings, teams developing machine translation for image-grounded content (such as e-commerce, media subtitling, and catalog translation) should consider adapter-based multimodal architectures rather than training multimodal models from scratch. In addition, evaluation pipelines should integrate contrastive test suites like CoMMuTE rather than relying exclusively on standard text-overlap metrics. Prior to production deployments, further engineering is recommended to extend the underlying object-detection dependencies beyond English source texts and to optimize the computational cost associated with large caption pretraining datasets.

arXiv: 2212.10140MatthieuFP/CoMMuTE

No sufficiently relevant recommendations were found.

Cover for Tackling Ambiguity with Images: Improved Multimodal Machine Translation and Contrastive Evaluation

Abstract

One of the major challenges of machine translation (MT) is ambiguity, which can in some cases be resolved by accompanying context such as images. However, recent work in multimodal MT (MMT) has shown that obtaining improvements from images is challenging, limited not only by the difficulty of building effective cross-modal representations, but also by the lack of specific evaluation and training data. We present a new MMT approach based on a strong text-only MT model, which uses neural adapters, a novel guided self-attention mechanism and which is jointly trained on both visually-conditioned masking and MMT. We also introduce CoMMuTE, a Contrastive Multilingual Multimodal Translation Evaluation set of ambiguous sentences and their possible translations, accompanied by disambiguating images corresponding to each translation. Our approach obtains competitive results compared to strong text-only models on standard English→French, English→German and English→Czech benchmarks and outperforms baselines and state-of-the-art MMT systems by a large margin on our contrastive test set. Our code¹ and CoMMuTE² are freely available.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Our approach: VGAMT
  • 3.1 Combining training objectives
  • 3.2 Guided self-attention
  • 4 Contrastive Multilingual Multimodal Translation Evaluation (CoMMuTE)
  • 5 Experiments
  • 5.1 Text-only data
  • 5.2 Multimodal data
  • 5.3 Implementation details
  • 5.4 Baselines
  • 5.5 Results and Analysis
  • 6 Ablation Study
  • 7 Conclusion
  • Limitations
  • Acknowledgements
  • References
  • A CoMMuTE statistics
  • B Visual features
  • C Guided self-attention analysis
  • D Additional examples
  • E METEOR scores
  • F Translating CoMMuTE

Knowls

  1. Knowl 1 — VGAMT adapts a strong text-only MT model to image-conditioned translation

    model/method

    Visually Guided and Adapted Machine Translation (VGAMT) starts from an mBART text-only translation model that has been fine-tuned on parallel text, then freezes its original weights and adds trainable bottleneck adapters and visual projection layers. The adapters are inserted after each attention block and feed-forward layer; they use a reduction factor of 8 and ReLU activation. VGAMT represents each input image with 64 local feature vectors from MDETR decoder queries and one global 512-dimensional CLIP image-encoder [CLS] feature. Projected image features are concatenated with source-text embeddings in the encoder. The encoder applies shared self-attention to image and text positions, while the decoder uses ordinary self- and cross-attention over text positions. The resulting design preserves the text model's learned translation knowledge while adding trainable image-conditioning pathways.

  2. Knowl 2 — Guided self-attention restricts cross-modal links to MDETR alignments

    equation

    VGAMT uses a binary mask to constrain attention over the concatenated source-text and image-feature positions. For query position ii and key position jj, its attention weight is

    aij=cijexp⁡(qi⊤kj/dk)∑lcilexp⁡(qi⊤kl/dk).a_{ij}=\frac{c_{ij}\exp(q_i^\top k_j/\sqrt{d_k})}{\sum_l c_{il}\exp(q_i^\top k_l/\sqrt{d_k})}.

    Here qiq_i and kjk_j are the learned query and key vectors at positions ii and jj, respectively; dkd_k is the key-vector dimension; and ll ranges over all key positions. The binary value cijc_{ij} is 1 for positions in the same modality, allowing text-to-text and image-to-image attention, and for cross-modal pairs linked by MDETR; it is 0 for other cross-modal pairs. Thus unaligned text-image pairs receive no attention, while aligned pairs and all same-modality pairs remain available.

  3. Knowl 3 — Joint VMLM and translation training encourages use of images

    experimental setup

    VGAMT is trained jointly on multimodal machine translation (MMT) and visually conditioned masked language modeling (VMLM). For MMT, it translates a source sentence conditioned on its image; for VMLM, it predicts randomly masked source-text tokens conditioned on the image. Training batches are sampled with equal probability from the parallel multimodal task and a monolingual multimodal VMLM task. The VMLM task uses Conceptual Captions images and English text; the authors report collecting and using about 2 million images from that dataset. For MMT, the model uses Multi30k, which has about 29,000 training examples. A quarter of the text input is masked for VMLM, encouraging the model to rely on visual information. The authors report that equal task sampling and 25% masking gave the best CoMMuTE results among the settings they tested; more masking reduced performance. Training used batch size 512, Adam with β1=0.9\beta_1=0.9 and β2=0.99\beta_2=0.99, label smoothing 0.1, and learning rate 10−410^{-4} for English-to-French or 10−510^{-5} for English-to-German and English-to-Czech. Results are averaged over three random seeds.

  4. Knowl 4 — CoMMuTE pairs lexical ambiguity with images and alternative translations

    data/table

    CoMMuTE is a contrastive evaluation set for testing whether a translation model uses an image to resolve lexical ambiguity. It contains 155 English sentences, each built around a lexically ambiguous word and paired with two images depicting different meanings. Each sentence has two possible translations in French, German, and Czech, with each translation corresponding to one meaning; native speakers of the target languages produced the translations. The images were collected under Creative Commons licenses or taken by the authors, and the image-text relation is not restricted to literal image captions. Twenty-nine source examples were adapted from prior data and the remaining examples were created for the dataset.

    The dataset statistics below count unique sentences, average sentence length, and unique tokens for the English source and each target language. The two alternatives per source explain why target-side unique sentence counts exceed 155.

    English French German Czech
    Unique sentences 155 308 300 308
    Average sentence length 6.54 6.90 6.48 5.07
    Unique tokens 462 679 638 718
  5. Knowl 5 — CoMMuTE scores image-conditioned ranking of translation alternatives

    definition

    For a source sentence, image, and candidate translation y=(y1,…,yN)y=(y_1,\ldots,y_N), CoMMuTE uses the model's token probabilities to compute sequence perplexity, PPLq(y)=(∏i=1Nq(yi))−1/N\mathrm{PPL}_q(y)=\left(\prod_{i=1}^{N}q(y_i)\right)^{-1/N}, where q(yi)q(y_i) is the model probability assigned to token yiy_i when scoring the candidate given the source, image, and preceding target tokens. Lower perplexity means the model considers the candidate more likely. A comparison is correct when the image-matched translation has lower perplexity than the alternative. Each of the two images is used to compare the two translations, producing two comparisons per source example; accuracy is the fraction of correct comparisons. Because the alternatives are balanced, a model that ignores images receives 50% accuracy.

  6. Knowl 6 — VGAMT substantially improves contrastive disambiguation accuracy

    empirical result

    On CoMMuTE, VGAMT outperforms both the strong text-only mBART adapter baseline and the best-scoring retrained multimodal baseline for each language pair. The following accuracies are percentages, averaged over three runs where reported; the text-only system scores exactly 50% because it cannot distinguish the balanced alternatives using the image.

    English→\toFrench English→\toGerman English→\toCzech
    Text-only mBART + adapters 50.0 50.0 50.0
    Best retrained MMT baseline 50.2 ±\pm 3.5 50.0 ±\pm 0.2 51.0 ±\pm 1.9
    VGAMT 67.1 ±\pm 0.7 59.0 ±\pm 0.5 55.6 ±\pm 0.8

    The retrained MMT comparison is Graph-MMT for English-to-French, VTLM + MMT for English-to-German, and Gated Fusion for English-to-Czech. VGAMT's advantage on this contrastive test indicates that it uses images to rank meaning-appropriate translations, rather than merely matching the strength of text-only translation.

  7. Knowl 7 — VGAMT remains competitive on standard translation benchmarks

    empirical result

    The table compares VGAMT with the text-only mBART + MT adapter baseline on standard test sets. Each entry reports BLEU and COMET, respectively; results are means over three runs with the reported standard errors. VGAMT is approximately on par with the text-only baseline for English-to-French, and slightly behind it on most English-to-German and English-to-Czech test sets. These benchmarks generally require less image-based disambiguation than CoMMuTE.

    Language pair Test set mBART + MT adapters (BLEU; COMET) VGAMT (BLEU; COMET)
    English→\toFrench Test2016 67.2 ±\pm 0.3; 0.971 ±\pm 0.005 67.2 ±\pm 0.1; 0.968 ±\pm 0.002
    English→\toFrench Test2017 61.5 ±\pm 0.3; 0.918 ±\pm 0.004 61.6 ±\pm 0.1; 0.921 ±\pm 0.002
    English→\toFrench MSCOCO 51.5 ±\pm 0.7; 0.832 ±\pm 0.006 51.1 ±\pm 0.6; 0.811 ±\pm 0.003
    English→\toGerman Test2016 43.6 ±\pm 0.2; 0.697 ±\pm 0.003 43.3 ±\pm 0.2; 0.694 ±\pm 0.003
    English→\toGerman Test2017 38.9 ±\pm 0.5; 0.664 ±\pm 0.002 38.3 ±\pm 0.2; 0.653 ±\pm 0.005
    English→\toGerman MSCOCO 36.2 ±\pm 0.2; 0.574 ±\pm 0.004 35.7 ±\pm 0.3; 0.544 ±\pm 0.006
    English→\toCzech Test2016 37.3 ±\pm 0.1; 0.940 ±\pm 0.005 37.6 ±\pm 0.2; 0.934 ±\pm 0.004
    English→\toCzech Test2018 35.2 ±\pm 0.4; 0.876 ±\pm 0.002 34.2 ±\pm 0.1; 0.833 ±\pm 0.003
  8. Knowl 8 — Ablations identify the contributions of joint training and visual features

    empirical result

    English-to-French ablations show that VGAMT's image-disambiguation performance depends on its training objectives and visual components. CoMMuTE accuracy for the full model is 67.1±0.767.1\pm0.7%. Removing VMLM reduces it to 52.0±1.252.0\pm1.2%; replacing joint MMT/VMLM training with VMLM pretraining followed by MMT fine-tuning reduces it to 63.3±0.563.3\pm0.5%. Removing guided self-attention gives 64.6±1.664.6\pm1.6%, and removing MDETR local features gives 63.0±1.263.0\pm1.2%. Removing CLIP global features causes the largest drop, to 50.3±0.050.3\pm0.0%. Fine-tuning the full base model without adapters reaches 60.5±3.860.5\pm3.8%. These comparisons support the value of joint training, alignment-constrained attention, and both local and global image representations; CLIP features are especially important for this contrastive task.

  9. Knowl 9 — Standard Multi30k-related test sets contain few image-dependent examples

    empirical result

    The authors manually assessed English-to-French examples in commonly used multimodal test sets for whether the source sentence was ambiguous, multiple meanings were compatible with its text context, and the image could distinguish meanings requiring different target translations. Under this criterion, only 21 examples in Test2016 (2.1%), 20 in Test2017 (2%), and 6 in MSCOCO (1.3%) were image-dependent. The authors argue that these low proportions help explain why standard benchmark scores often do not reveal whether a multimodal MT system actually uses images: most examples can be translated correctly from text alone.

  10. Knowl 10 — VGAMT is limited to English-source translation and needs substantial caption data

    limitation

    The method depends on MDETR to extract source-language text-image features and alignments. At the time of the study, the required modulated object detector was available only for English, so the authors' approach was applicable only to English-to-other-language translation. They also state that VGAMT requires a large amount of image-caption data to perform well, making training computationally expensive.

Coverage note — Qualitative translation examples, the supplementary attention-map interpretation, and generation-based scores on CoMMuTE are omitted because they are illustrative or secondary to the core model, benchmark, and contrastive evaluation contributions.

References

  1. 1.Abien Fred Agarap. 2018. Deep learning using rectified linear units (ReLU). arXiv preprint arXiv:1803.08375.
  2. 2.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, et al. 2022. Flamingo: a visual language model for few-shot learning. arXiv preprint arXiv:2204.14198.
  3. 3.Satanjeev Banerjee and Alon Lavie. 2005. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, pages 65–72, Ann Arbor, Michigan. Association for Computational Linguistics.
  4. 4.Loïc Barrault, Fethi Bougares, Lucia Specia, Chiraag Lala, Desmond Elliott, and Stella Frank. 2018. Findings of the third shared task on multimodal machine translation. In Proceedings of the Third Conference on Machine Translation: Shared Task Papers, pages 304–323, Belgium, Brussels. Association for Computational Linguistics.
  5. 5.Rachel Bawden, Rico Sennrich, Alexandra Birch, and Barry Haddow. 2018. Evaluating discourse phenomena in neural machine translation. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1304–1313, New Orleans, Louisiana. Association for Computational Linguistics.
  6. 6.Ozan Caglayan, Menekse Kuyu, Mustafa Sercan Amac, Pranava Madhyastha, Erkut Erdem, Aykut Erdem, and Lucia Specia. 2021. Cross-lingual visual pre-training for multimodal machine translation. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 1317–1324, Online. Association for Computational Linguistics.
  7. 7.Ozan Caglayan, Pranava Madhyastha, Lucia Specia, and Loïc Barrault. 2019. Probing the need for visual context in multimodal machine translation. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4159–4170, Minneapolis, Minnesota. Association for Computational Linguistics.
  8. 8.Iacer Calixto and Qun Liu. 2017. Incorporating global visual features into attention-based neural machine translation. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 992–1003, Copenhagen, Denmark. Association for Computational Linguistics.
  9. 9.Iacer Calixto, Qun Liu, and Nick Campbell. 2017. Doubly-attentive decoder for multi-modal neural machine translation. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1913–1924, Vancouver, Canada. Association for Computational Linguistics.
  10. 10.Xi Chen, Xiao Wang, Soravit Changpinyo, AJ Piergiovanni, Piotr Padlewski, Daniel Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, et al. 2022. Pali: A jointly-scaled multilingual language-image model. arXiv preprint arXiv:2209.06794.
  11. 11.Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2020. Uniter: Universal image-text representation learning. In European conference on computer vision, pages 104–120. Springer.
  12. 12.Alexis Conneau and Guillaume Lample. 2019. Cross-lingual language model pretraining. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc.
  13. 13.Matthias De Lange, Rahaf Aljundi, Marc Masana, Sarah Parisot, Xu Jia, Aleš Leonardis, Gregory Slabaugh, and Tinne Tuytelaars. 2021. A continual learning survey: Defying forgetting in classification tasks. IEEE transactions on pattern analysis and machine intelligence, 44(7):3366–3385.
  14. 14.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009. Imagenet: A large-scale hierarchical image database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition, pages 248–255.
  15. 15.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  16. 16.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An image is worth 16x16 words: Transformers for image recognition at scale. In Proceedings of the 9th International Conference on Learning Representations. OpenReview.net.
  17. 17.Constantin Eichenberg, Sidney Black, Samuel Weinbach, Letitia Parcalabescu, and Anette Frank. 2021. Magma–multimodal augmentation of generative models through adapter-based finetuning. arXiv preprint arXiv:2112.05253.
  18. 18.Desmond Elliott. 2018. Adversarial evaluation of multimodal machine translation. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2974–2978, Brussels, Belgium. Association for Computational Linguistics.
  19. 19.Desmond Elliott, Stella Frank, Loïc Barrault, Fethi Bougares, and Lucia Specia. 2017. Findings of the second shared task on multimodal machine translation and multilingual image description. In Proceedings of the Second Conference on Machine Translation, Volume 2: Shared Task Papers, pages 215–233, Copenhagen, Denmark. Association for Computational Linguistics.
  20. 20.Desmond Elliott, Stella Frank, Khalil Sima’an, and Lucia Specia. 2016. Multi30K: Multilingual English-German image descriptions. In Proceedings of the 5th Workshop on Vision and Language, pages 70–74, Berlin, Germany. Association for Computational Linguistics.
  21. 21.Desmond Elliott and Ákos Kádár. 2017. Imagination improves multimodal translation. In Proceedings of the Eighth International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 130–141, Taipei, Taiwan. Asian Federation of Natural Language Processing.
  22. 22.Qingkai Fang and Yang Feng. 2022. Neural machine translation with phrase-level universal visual representations. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5687–5698, Dublin, Ireland. Association for Computational Linguistics.
  23. 23.Christiane Fellbaum. 1998. WordNet 1.6: An Electronic Lexical Database. Bradford Books. MIT Press.
  24. 24.Stig-Arne Grönroos, Benoit Huet, Mikko Kurimo, Jorma Laaksonen, Bernard Merialdo, Phu Pham, Mats Sjöberg, Umut Sulubacak, Jörg Tiedemann, Raphael Troncy, and Raúl Vázquez. 2018. The MeMAD submission to the WMT18 multimodal translation task. In Proceedings of the Third Conference on Machine Translation: Shared Task Papers, pages 603–611, Belgium, Brussels. Association for Computational Linguistics.
  25. 25.Kaiming He, X. Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 770–778.
  26. 26.Jindˇrich Helcl, Jindˇrich Libovický, and Dušan Variš. 2018. CUNI system for the WMT18 multimodal translation task. In Proceedings of the Third Conference on Machine Translation: Shared Task Papers, pages 616–623, Belgium, Brussels. Association for Computational Linguistics.
  27. 27.Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-efficient transfer learning for NLP. In International Conference on Machine Learning, pages 2790–2799. PMLR.
  28. 28.Haoyang Huang, Lin Su, Di Qi, Nan Duan, Edward Cui, Taroon Bharti, Lei Zhang, Lijuan Wang, Jianfeng Gao, Bei Liu, Jianlong Fu, Dongdong Zhang, Xin Liu, and Ming Zhou. 2021a. M3P: Learning universal representations via multitask multilingual multimodal pre-training. In 2021 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 3976–3985.
  29. 29.Xin Huang, Jiajun Zhang, and Chengqing Zong. 2021b. Entity-level cross-modal learning improves multimodal machine translation. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 1067–1080, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  30. 30.Julia Ive, Pranava Madhyastha, and Lucia Specia. 2019. Distilling translations with visual awareness. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 6525–6538, Florence, Italy. Association for Computational Linguistics.
  31. 31.Aishwarya Kamath, Mannat Singh, Yann LeCun, Ishan Misra, Gabriel Synnaeve, and Nicolas Carion. 2021. MDETR - modulated detection for end-to-end multimodal understanding. 2021 IEEE/CVF International Conference on Computer Vision (ICCV), pages 1760–1770.
  32. 32.Diederik P Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980.
  33. 33.Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, and Kentaro Inui. 2020. Attention is not only a weight: Analyzing transformers with vector norms. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7057–7075, Online. Association for Computational Linguistics.
  34. 34.Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, Chris Dyer, Ondˇrej Bojar, Alexandra Constantin, and Evan Herbst. 2007. Moses: Open source toolkit for statistical machine translation. In Proceedings of the 45th Annual Meeting of the Association for Computational Linguistics Companion Volume Proceedings of the Demo and Poster Sessions, pages 177–180, Prague, Czech Republic. Association for Computational Linguistics.
  35. 35.Chiraag Lala and Lucia Specia. 2018. Multimodal lexical translation. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan. European Language Resources Association (ELRA).
  36. 36.Jiaoda Li, Duygu Ataman, and Rico Sennrich. 2021. Vision matters when it should: Sanity checking multimodal machine translation models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 8556–8562, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  37. 37.Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. 2019. VisualBERT: A simple and performant baseline for vision and language. arXiv preprint arXiv:1908.03557.
  38. 38.Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, et al. 2020. Oscar: Object-semantics aligned pre-training for vision-language tasks. In European Conference on Computer Vision, pages 121–137. Springer.
  39. 39.Yi Li, Rameswar Panda, Yoon Kim, Chun-Fu (Richard) Chen, Rogerio Feris, David Cox, and Nuno Vasconcelos. 2022. VALHALLA: Visual Hallucination for Machine Translation. In 2022 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 5216–5226.
  40. 40.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014. Microsoft COCO: Common objects in context. In European conference on computer vision, pages 740–755. Springer.
  41. 41.Pierre Lison, Jörg Tiedemann, and Milen Kouylekov. 2018. OpenSubtitles2018: Statistical rescoring of sentence alignments in large, noisy parallel corpora. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan. European Language Resources Association (ELRA).
  42. 42.Pengbo Liu, Hailong Cao, and Tiejun Zhao. 2021. Gumbel-attention for multi-modal machine translation. arXiv preprint arXiv:2103.08862.
  43. 43.Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020. Multilingual denoising pre-training for neural machine translation. Transactions of the Association for Computational Linguistics, 8:726–742.
  44. 44.Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc.
  45. 45.Mathias Müller, Annette Rios, Elena Voita, and Rico Sennrich. 2018. A large-scale test set for the evaluation of context-aware pronoun translation in neural machine translation. In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 61–72, Brussels, Belgium. Association for Computational Linguistics.
  46. 46.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311–318, Philadelphia, Pennsylvania, USA. Association for Computational Linguistics.
  47. 47.Jonas Pfeiffer, Gregor Geigle, Aishwarya Kamath, Jan-Martin Steitz, Stefan Roth, Ivan Vulic, and Iryna Gurevych. 2022. xGQA: Cross-lingual visual question answering. In Findings of the Association for Computational Linguistics: ACL 2022, pages 2497–2511, Dublin, Ireland. Association for Computational Linguistics.
  48. 48.Jonas Pfeiffer, Andreas Rücklé, Clifton Poth, Aishwarya Kamath, Ivan Vulic, Sebastian Ruder, Kyunghyun Cho, and Iryna Gurevych. 2020. AdapterHub: A framework for adapting transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 46–54, Online. Association for Computational Linguistics.
  49. 49.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pages 8748–8763. PMLR.
  50. 50.Alessandro Raganato, Yves Scherrer, and Jörg Tiedemann. 2019. The MuCoW test suite at WMT 2019: Automatically harvested multilingual contrastive word sense disambiguation test sets for machine translation. In Proceedings of the Fourth Conference on Machine Translation (Volume 2: Shared Task Papers, Day 1), pages 470–480, Florence, Italy. Association for Computational Linguistics.
  51. 51.Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020. COMET: A neural framework for MT evaluation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2685–2702, Online. Association for Computational Linguistics.
  52. 52.Nils Reimers and Iryna Gurevych. 2020. Making monolingual sentence embeddings multilingual using knowledge distillation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4512–4525, Online. Association for Computational Linguistics.
  53. 53.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015. Faster R-CNN: Towards real-time object detection with region proposal networks. In Advances in Neural Information Processing Systems, volume 28. Curran Associates, Inc.
  54. 54.Annette Rios Gonzales, Laura Mascarell, and Rico Sennrich. 2017. Improving word sense disambiguation in neural machine translation with sense embeddings. In Proceedings of the Second Conference on Machine Translation, pages 11–19, Copenhagen, Denmark. Association for Computational Linguistics.
  55. 55.Rico Sennrich. 2017. How Grammatical is Character-level Neural Machine Translation? Assessing MT Quality with Contrastive Translation Pairs. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers, pages 376–382, Valencia, Spain. Association for Computational Linguistics.
  56. 56.Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2556–2565, Melbourne, Australia. Association for Computational Linguistics.
  57. 57.Lucia Specia, Stella Frank, Khalil Sima’an, and Desmond Elliott. 2016. A shared task on multimodal machine translation and crosslingual image description. In Proceedings of the First Conference on Machine Translation: Volume 2, Shared Task Papers, pages 543–553, Berlin, Germany. Association for Computational Linguistics.
  58. 58.Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. 2020. Vl-bert: Pre-training of generic visual-linguistic representations. In International Conference on Learning Representations.
  59. 59.Yi-Lin Sung, Jaemin Cho, and Mohit Bansal. 2022. VL-ADAPTER: Parameter-efficient transfer learning for vision-and-language tasks. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 5217–5227.
  60. 60.Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. 2016. Rethinking the inception architecture for computer vision. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition (CVPR), pages 2818–2826.
  61. 61.Jörg Tiedemann. 2012. Parallel data, tools and interfaces in OPUS. In Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC’12), pages 2214–2218, Istanbul, Turkey. European Language Resources Association (ELRA).
  62. 62.Maria Tsimpoukelli, Jacob L Menick, Serkan Cabi, SM Eslami, Oriol Vinyals, and Felix Hill. 2021. Multimodal few-shot learning with frozen language models. Advances in Neural Information Processing Systems, 34:200–212.
  63. 63.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc.
  64. 64.Elena Voita, Rico Sennrich, and Ivan Titov. 2019. When a good translation is wrong in context: Context-aware machine translation improves on deixis, ellipsis, and lexical cohesion. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 1198–1212, Florence, Italy. Association for Computational Linguistics.
  65. 65.Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, and Edouard Grave. 2020. CCNet: Extracting high quality monolingual datasets from web crawl data. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 4003–4012, Marseille, France. European Language Resources Association.
  66. 66.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  67. 67.Krzysztof Wołk and Krzysztof Marasek. 2014. Building subject-aligned comparable corpora and mining it for truly parallel sentence pairs. Procedia Technology, 18:126–132. International workshop on Innovations in Information and Communication Science and Technology, IICST 2014, 3-5 September 2014, Warsaw, Poland.
  68. 68.Zhiyong Wu, Lingpeng Kong, Wei Bi, Xiang Li, and Ben Kao. 2021. Good for misconceived reasons: An empirical revisiting on the need for visual context in multimodal machine translation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6153–6166, Online. Association for Computational Linguistics.
  69. 69.Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, and Cordelia Schmid. 2022. Zero-shot video question answering via frozen bidirectional language models. In Advances in Neural Information Processing Systems.
  70. 70.Shaowei Yao and Xiaojun Wan. 2020. Multimodal transformer for multimodal machine translation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4346–4350, Online. Association for Computational Linguistics.
  71. 71.Junjie Ye, Junjun Guo, Yan Xiang, Kaiwen Tan, and Zhengtao Yu. 2022. Noise-robust cross-modal interactive learning with Text2Image mask for multimodal neural machine translation. In Proceedings of the 29th International Conference on Computational Linguistics, pages 5098–5108, Gyeongju, Republic of Korea. International Committee on Computational Linguistics.
  72. 72.Yongjing Yin, Fandong Meng, Jinsong Su, Chulun Zhou, Zhengyuan Yang, Jie Zhou, and Jiebo Luo. 2020. A novel graph-based multi-modal fusion encoder for neural machine translation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 3025–3035, Online. Association for Computational Linguistics.
  73. 73.Mingyang Zhou, Luowei Zhou, Shuohang Wang, Yu Cheng, Linjie Li, Zhou Yu, and Jingjing Liu. 2021. UC2: Universal cross-lingual cross-modal vision-and-language pre-training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4155–4165, Nashville, TN, USA.

Citation

MLA
Futeral, M., et al. “Tackling Ambiguity with Images: Improved Multimodal Machine Translation and Contrastive Evaluation”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 5394–413, https://doi.org/10.18653/v1/2023.acl-long.295.
APA
Futeral, M., Schmid, C., Laptev, I., Sagot, B., & Bawden, R. (2023). Tackling Ambiguity with Images: Improved Multimodal Machine Translation and Contrastive Evaluation. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 5394–5413. https://doi.org/10.18653/v1/2023.acl-long.295
Chicago
Futeral, M., C. Schmid, I. Laptev, B. Sagot, and R. Bawden. 2023. “Tackling Ambiguity with Images: Improved Multimodal Machine Translation and Contrastive Evaluation”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 5394–5413. https://doi.org/10.18653/v1/2023.acl-long.295.
Harvard
Futeral, M. et al. (2023) “Tackling Ambiguity with Images: Improved Multimodal Machine Translation and Contrastive Evaluation”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 5394–5413. Available at: https://doi.org/10.18653/v1/2023.acl-long.295.
Vancouver
1. Futeral M, Schmid C, Laptev I, Sagot B, Bawden R (2023) Tackling Ambiguity with Images: Improved Multimodal Machine Translation and Contrastive Evaluation. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 5394–5413

BibTeX

@inproceedings{futeral-etal-2023-tackling,
    title = "Tackling Ambiguity with Images: Improved Multimodal Machine Translation and Contrastive Evaluation",
    author = "Futeral, Matthieu  and
      Schmid, Cordelia  and
      Laptev, Ivan  and
      Sagot, Beno{\^i}t  and
      Bawden, Rachel",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.295/",
    doi = "10.18653/v1/2023.acl-long.295",
    pages = "5394--5413"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/