On Vision Features in Multimodal Machine Translation

Bei LiChuanhao LvZefan ZhouTao ZhouTong XiaoAnxiang MaJingbo Zhu

article2022ACL89 citations

Reveals that stronger Transformer-based vision models improve multimodal machine translation on targeted probing tasks, while introducing a patch-level selective attention mechanism that directly correlates visual regions with words.

Listen

Multimodal machine translation combines text and images to translate across languages, but earlier research suggested that visual data provided little real value when complete text was present. Previous systems primarily relied on older visual encoders such as standard convolutional networks, assuming they were sufficient. The article investigates whether modern, high-capacity vision models—specifically Vision Transformers—and enhanced visual features can make images genuinely useful for translation, evaluating their impact through targeted testing scenarios.

The authors conducted an extensive empirical study across standard English-to-German and English-to-French datasets (Multi30K). They designed a selective attention mechanism that directly links text words with local image segments. To accurately assess visual contribution, they implemented probing tasks where key descriptive words—such as colors, characters, and nouns—were masked from the input text, forcing the translation system to extract missing information directly from the image. They also evaluated performance when models were presented with mismatched images.

The analysis yielded several critical findings. On complete text benchmarks, upgrading visual encoders produced only marginal score increases, confirming that standard metrics fail to reflect whether visual information is actually used. However, under probing conditions with incomplete text, Vision Transformer models dramatically outperformed traditional convolutional baselines, achieving accuracy improvements of over 20 percentage points in color- and character-recovery tasks. The selective attention mechanism proved essential, as finer-grained image patches and higher image resolutions allowed the model to accurately pinpoint and recover visual details. Additionally, when mismatched images were provided during decoding, models utilizing strong Vision Transformer features suffered substantial performance drops, proving that the system was actively relying on visual content rather than using it merely as generic noise or a training regularizer.

These results demonstrate that the visual modality is truly complementary and effective when powered by sufficiently strong visual architectures and fine-grained attention. Standard automated benchmark scores alone can be misleading, risking poor architectural decisions if systems are deployed without targeted probing. For organizations deploying multimodal translation in environments with noisy, ambiguous, or incomplete text inputs, utilizing modern Vision Transformer representations offers substantial quality improvements. Stakeholders should ensure that multimodal evaluation pipelines incorporate masked probing tasks before selecting models.

Future development should focus on creating unified architectures that jointly encode text and vision, as well as addressing the scarcity of large-scale multimodal translation datasets. Decision-makers should note that the current findings are primarily established on the Multi30K benchmark, so testing on specialized, domain-specific data remains recommended before full-scale operational rollout.

  • Paper: Transformers in Vision: A Survey, Salman Khan et al. (2021). This survey lays out Vision Transformer architectures and their use in multimodal tasks, making the source’s choice of visual encoder and feature representation easier to follow.
  • Paper: Do Vision Transformers See Like Convolutional Neural Networks?, Maithra Raghu et al. (2021). Its comparison of Vision Transformer and convolutional representations provides useful grounding for the source’s tests of whether ViT features recover visual details better than CNN features.

No sufficiently relevant recommendations were found.

Cover for On Vision Features in Multimodal Machine Translation

Abstract

Previous work on multimodal machine translation (MMT) has focused on the way of incorporating vision features into translation but little attention is on the quality of vision models. In this work, we investigate the impact of vision models on MMT. Given the fact that Transformer is becoming popular in computer vision, we experiment with various strong models (such as Vision Transformer) and enhanced features (such as object-detection and image captioning). We develop a selective attention model to study the patch-level contribution of an image in MMT. On detailed probing tasks, we find that stronger vision models are helpful for learning translation from the visual modality. Our results also suggest the need of carefully examining MMT models, especially when current benchmarks are small-scale and biased. Our code could be found at https://github.com/libeineu/fairseq_mmt.

Table of Contents

  • 1 Introduction
  • 2 Preliminary
  • 2.1 Insufficient Text Generation
  • 2.2 Various Vision Features
  • 2.3 Selective Attention
  • 3 Experiments
  • 3.1 Datasets
  • 3.2 Experimental Setups
  • 3.3 Results
  • 4 Analysis
  • 4.1 How Vision Features Improve the MMT
  • 4.2 Impact of Learning Objectives
  • 4.3 Impact of Resolution and Patch Size
  • 4.4 Incongruent Decoding
  • 4.5 Case Study
  • 5 Related Work
  • 6 Conclusions
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Selective attention lets translation attend to image patches

    model/method

    The selective-attention multimodal Transformer uses the source-text encoder states as queries and the vision-model patch representations as both keys and values. Let XtextX_{text} be a source-token sequence and XimgX_{img} an image; let Htext=TransformerEncoder⁡(Xtext)H_{text}=\operatorname{TransformerEncoder}(X_{text}), and let Himg=WViT⁡(Ximg)H_{img}=W\operatorname{ViT}(X_{img}), where WW projects the image representation to the text-state dimension. With one attention head, the patch information selected for each source position is

    Hattnimg=Softmax⁡ ⁣(HtextHimgTdk)Himg,H^{img}_{attn}=\operatorname{Softmax}\!\left(\frac{H_{text}H_{img}^{\mathsf T}}{\sqrt{d_k}}\right)H_{img},

    where dkd_k is the shared query/key feature dimension. The model then fuses HtextH_{text} with HattnimgH^{img}_{attn} using the gated-fusion rule λ=Sigmoid⁡(UHtext+VHattnimg)\lambda=\operatorname{Sigmoid}(UH_{text}+VH^{img}_{attn}) and Hout=(1−λ)Htext+λHattnimgH_{out}=(1-\lambda)H_{text}+\lambda H^{img}_{attn}; UU and VV are trainable projections and the gate λ\lambda controls the retained visual information. The fused sequence is passed to the translation decoder. Unlike fusion with a single global image vector, this method can select different image patches for different source words.

  2. Knowl 2 — Three insufficient-text probing tasks test visual contributions

    definition

    The paper evaluates whether an image helps translation by deliberately masking source words, requiring the translation system to infer missing content from the visual input. In color probing, every source color word is replaced by [MASK_C]; the training data contain 8,919 sentences with color words, and nearly one third contain multiple colors. In character probing, the words “man,” “woman,” “people,” “men,” “girl,” and “boy” are replaced by [MASK_P]; more than 60% of the training sentences contain one of these character words. For both tasks, strict accuracy requires the appropriate target-language form, including grammatical gender where relevant, while relaxed accuracy accepts any translation expressing the same color or character category. In noun probing, frequent nouns labeled in Flickr30K Entities—covering categories such as animals, clothing, and vehicles—are masked as [MASK_N] or [MASK_NS] for singular and plural nouns. The noun task varies the number of masked nouns from one to four, testing increasingly incomplete source sentences.

  3. Knowl 3 — ViT with selective attention substantially improves color and character recovery

    empirical result

    On the Multi30K English–German and English–French color and character probes, selective attention with pretrained ViT beats the text-only Transformer on every reported test set and both strict and relaxed criteria. Relative to the text-only system, color-probe accuracy rises by 9.19–32.81 percentage points on English–German and 14.43–35.73 points on English–French. Character-probe accuracy rises by 12.11–14.84 points on English–German and 14.45–17.12 points on English–French. For example, English–German color accuracy on Test2016 increases from 25.93% to 51.20% under the strict criterion and from 34.42% to 64.71% under the relaxed criterion. These gains are not explained by simply adding any visual feature: gated fusion with ResNet yields marginal color improvements and sometimes falls below the text-only character-probe accuracy, whereas gated fusion with ViT also improves over the text-only system.

  4. Knowl 4 — Visual-feature advantages grow as noun context is removed

    empirical result

    In noun-based probing, ViT features outperform ResNet features across all tested masking levels, language pairs, and test sets; the performance gap generally widens as more nouns are masked. Selective attention with ViT is especially effective in these incomplete-text settings. Experiments on English–German Test2016 also show that increasing ViT or Swin Transformer capacity generally improves BLEU in progressive noun-masking scenarios. However, Swin is inferior to ViT in the tested matched-capacity comparisons. The authors suggest that the shorter visual sequence may explain this result: at 384 × 384 resolution with 16 × 16 patches, ViT provides 576 patches plus a classification token, while the tested Swin representation has a fixed length of 49. They hypothesize that ViT’s longer sequence provides finer local features that better support patch selection.

  5. Knowl 5 — Complete-text benchmarks show only modest separation between vision models

    empirical result

    On the standard Multi30K translation tests, where the source text is complete, adding visual features generally produces only marginal improvements over a small text-only Transformer, and replacing ResNet with ViT in gated fusion does not yield a significant BLEU gain. For example, on Test2016 English–German, the text-only system scores 41.02 BLEU and 68.22 METEOR, while selective attention with ViT-Large scores 41.84 BLEU and 68.64 METEOR. On Test2016 English–French, the corresponding scores are 61.80 BLEU and 81.02 METEOR for text-only versus 62.24 BLEU and 81.41 METEOR for selective attention with ViT-Large. Thus, standard complete-text scores distinguish these systems much less clearly than the masked-word probes do; the paper cautions that automatic scores on current MMT benchmarks may not reliably indicate how much a model uses visual context.

  6. Knowl 6 — Detection and captioning features help complete-text scores but not masked-text probes

    empirical result

    The authors compare pretrained object-detection representations from DETR and QueryInst and image-captioning representations from CATR with ViT features. On complete-text Multi30K tests, these enhanced representations can outperform the ViT-Tiny comparison system. For instance, on Test2016 English–German, CATR obtains 42.50 BLEU and 68.81 METEOR, compared with 40.74 BLEU and 67.20 METEOR for the ViT-Tiny system; on English–French, CATR obtains 62.79 BLEU and 81.75 METEOR, compared with 61.44 BLEU and 80.91 METEOR. Their advantage does not persist in the insufficient-text probing scenarios: the plotted results show no corresponding consistent superiority over ViT-Tiny. The authors offer sensitivity to the quality of extracted objects as a possible explanation, not as a demonstrated cause.

  7. Knowl 7 — Image resolution and patch granularity affect probing accuracy

    empirical result

    On English–German Test2016, the authors vary ViT resolution and patch size and compare with Swin. The strongest character-probe scores occur with ViT using 16 × 16 patches at 384 × 384 resolution: 74.32% strict and 79.46% relaxed. That configuration also reaches 36.59 BLEU for one-noun masking and 27.29 BLEU for four-noun masking. By comparison, ViT with 32 × 32 patches at 224 × 224 resolution scores 68.19% and 73.47% on strict and relaxed character probing, and 35.14 and 25.19 BLEU for one- and four-noun masking. The pattern supports the authors’ conclusion that fine-grained visual features are useful for selective attention, although not every metric improves monotonically with resolution: for example, the highest strict color-probe score in the comparison is 50.11% for ViT with 16 × 16 patches at 224 × 224, versus 49.67% at 384 × 384. Attention-map visualizations further show that the higher-resolution model can identify a masked color region that the lower-resolution model misses.

  8. Knowl 8 — Incongruent decoding exposes the use of pretrained visual context

    empirical result

    For noun masking on English–German Test2016, the authors compare congruent and incongruent decoding as an additional test of visual reliance. Selective attention with pretrained ViT scores 36.59, 32.08, 29.47, and 27.29 BLEU under congruent decoding for one through four masked nouns, respectively; its incongruent scores are 32.88, 25.58, 20.42, and 15.80. The resulting BLEU drops grow from 3.71 points with one noun masked to 11.49 points with four masked, indicating stronger dependence on visual context as textual evidence is removed. Gated fusion with pretrained ResNet changes much less: its congruent versus incongruent scores are 34.90 versus 34.88, 28.94 versus 28.08, 24.18 versus 22.56, and 21.74 versus 20.79. The ViT-plus-CATR ensemble has the highest congruent scores in all four masking conditions (36.97, 32.45, 30.30, and 28.14 BLEU) and the paper reports that it also has the largest congruent–incongruent gaps among the compared systems. A ViT vision system learned from scratch is comparatively insensitive to the decoding condition, which the authors attribute to the limited training data available for learning strong visual representations.

  9. Knowl 9 — Multi30K experimental configuration

    experimental setup

    The experiments use Multi30K English–German and English–French translation, with 29,000 training instances and 1,014 validation instances. Results are reported on Test2016, Test2017, and MSCOCO; the paper describes MSCOCO as more challenging because it contains out-of-domain examples with ambiguous verbs. Source and target texts use joint BPE with 10,000 merge operations, producing vocabularies of 9,716 entries for English–German and 9,548 for English–French. The text translation backbone is a Transformer-Tiny with four encoder and four decoder layers, hidden size 128, feed-forward size 256, four attention heads, dropout 0.3, and label smoothing 0.1. Training uses Adam with β1=0.9\beta_1=0.9, β2=0.98\beta_2=0.98, and ϵ=10−8\epsilon=10^{-8}; the learning rate warms up linearly for 2,000 steps from 10−710^{-7} to 5×10−35\times10^{-3} and then decays with the inverse square root of the step. Batches contain 4,096 tokens, and early stopping is used. Evaluation averages the last 10 checkpoints, uses beam size 5, and reports BLEU and METEOR on standard tests and accuracy on probing tasks.

  10. Knowl 10 — Small and biased data limit conclusions and visual-model learning

    limitation

    The paper cautions that current MMT benchmarks are small-scale and potentially biased, so performance on the original complete-text task may understate or misrepresent whether a system uses the image; the authors therefore argue for careful probing in addition to automatic translation metrics. Multi30K’s limited data also constrains vision models trained from scratch, which appear comparatively insensitive to visual context. Finally, the authors identify a remaining difficulty in complex scenes: their qualitative examples show that even the ViT-based system can choose an incorrect noun, and they describe data scarcity and the need for more powerful cross-modal fusion as unresolved issues.

Coverage note — The two qualitative translation examples are omitted because they are anecdotal illustrations of the quantified probing results; no other substantial contributed material is deliberately omitted.

References

  1. 1.Ozan Caglayan, Loïc Barrault, and Fethi Bougares. 2016. Multimodal attention for neural machine translation. CoRR, abs/1609.03976.
  2. 2.Ozan Caglayan, Menekse Kuyu, Mustafa Sercan Amac, Pranava Madhyastha, Erkut Erdem, Aykut Erdem, and Lucia Specia. 2021. Cross-lingual visual pre-training for multimodal machine translation. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 1317–1324, Online. Association for Computational Linguistics.
  3. 3.Ozan Caglayan, Pranava Madhyastha, Lucia Specia, and Loïc Barrault. 2019. Probing the need for visual context in multimodal machine translation. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4159–4170, Minneapolis, Minnesota. Association for Computational Linguistics.
  4. 4.Iacer Calixto and Qun Liu. 2017. Incorporating global visual features into attention-based neural machine translation. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 992–1003, Copenhagen, Denmark. Association for Computational Linguistics.
  5. 5.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. 2020. End-to-end object detection with transformers. In Computer Vision - ECCV 2020 - 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part I, volume 12346 of Lecture Notes in Computer Science, pages 213–229. Springer.
  6. 6.Jean-Benoit Delbrouck and Stéphane Dupont. 2017. An empirical study on the effectiveness of images in multimodal neural machine translation. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 910–919, Copenhagen, Denmark. Association for Computational Linguistics.
  7. 7.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An image is worth 16x16 words: Transformers for image recognition at scale. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net.
  8. 8.Desmond Elliott. 2018. Adversarial evaluation of multimodal machine translation. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2974–2978, Brussels, Belgium. Association for Computational Linguistics.
  9. 9.Desmond Elliott, Stella Frank, Loïc Barrault, Fethi Bougares, and Lucia Specia. 2017. Findings of the second shared task on multimodal machine translation and multilingual image description. In Proceedings of the Second Conference on Machine Translation, pages 215–233, Copenhagen, Denmark. Association for Computational Linguistics.
  10. 10.Desmond Elliott, Stella Frank, Khalil Sima’an, and Lucia Specia. 2016. Multi30k: Multilingual english-german image descriptions. In Proceedings of the 5th Workshop on Vision and Language, hosted by the 54th Annual Meeting of the Association for Computational Linguistics, VL@ACL 2016, August 12, Berlin, Germany. The Association for Computer Linguistics.
  11. 11.Desmond Elliott and Ákos Kádár. 2017. Imagination improves multimodal translation. In Proceedings of the Eighth International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 130–141, Taipei, Taiwan. Asian Federation of Natural Language Processing.
  12. 12.Yuxin Fang, Shusheng Yang, Xinggang Wang, Yu Li, Chen Fang, Ying Shan, Bin Feng, and Wenyu Liu. 2021. Instances as queries. CoRR, abs/2105.01928.
  13. 13.Stig-Arne Grönroos, Benoit Huet, Mikko Kurimo, Jorma Laaksonen, Bernard Merialdo, Phu Pham, Mats Sjöberg, Umut Sulubacak, Jörg Tiedemann, Raphael Troncy, and Raúl Vázquez. 2018. The MeMAD submission to the WMT18 multimodal translation task. In Proceedings of the Third Conference on Machine Translation: Shared Task Papers, pages 603–611, Belgium, Brussels. Association for Computational Linguistics.
  14. 14.Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings.
  15. 15.Chiraag Lala, Pranava Swaroop Madhyastha, Carolina Scarton, and Lucia Specia. 2018. Sheffield submissions for WMT18 multimodal translation shared task. In Proceedings of the Third Conference on Machine Translation: Shared Task Papers, pages 624–631, Belgium, Brussels. Association for Computational Linguistics.
  16. 16.Bei Li, Hui Liu, Ziyang Wang, Yufan Jiang, Tong Xiao, Jingbo Zhu, Tongran Liu, and Changliang Li. 2020. Does multi-encoder help? a case study on context-aware neural machine translation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 3512–3518, Online. Association for Computational Linguistics.
  17. 17.Jindˇrich Libovický and Jindˇrich Helcl. 2017. Attention strategies for multi-source sequence-to-sequence learning. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 196–202, Vancouver, Canada. Association for Computational Linguistics.
  18. 18.Huan Lin, Fandong Meng, Jinsong Su, Yongjing Yin, Zhengyuan Yang, Yubin Ge, Jie Zhou, and Jiebo Luo. 2020. Dynamic context-guided capsule network for multimodal machine translation. In MM ’20: The 28th ACM International Conference on Multimedia, Virtual Event / Seattle, WA, USA, October 12-16, 2020, pages 1320–1329. ACM.
  19. 19.Pengbo Liu, Hailong Cao, and Tiejun Zhao. 2021a. Gumbel-attention for multi-modal machine translation. CoRR, abs/2103.08862.
  20. 20.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021b. Swin transformer: Hierarchical vision transformer using shifted windows. CoRR, abs/2103.14030.
  21. 21.Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019. fairseq: A fast, extensible toolkit for sequence modeling. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (Demonstrations), pages 48–53, Minneapolis, Minnesota. Association for Computational Linguistics.
  22. 22.Bryan A. Plummer, Liwei Wang, Chris M. Cervantes, Juan C. Caicedo, Julia Hockenmaier, and Svetlana Lazebnik. 2015. Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models. In 2015 IEEE International Conference on Computer Vision, ICCV 2015, Santiago, Chile, December 7-13, 2015, pages 2641–2649. IEEE Computer Society.
  23. 23.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research, pages 8748–8763. PMLR.
  24. 24.Lucia Specia, Stella Frank, Khalil Sima’an, and Desmond Elliott. 2016. A shared task on multimodal machine translation and crosslingual image description. In Proceedings of the First Conference on Machine Translation: Volume 2, Shared Task Papers, pages 543–553, Berlin, Germany. Association for Computational Linguistics.
  25. 25.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pages 5998–6008.
  26. 26.Dexin Wang and Deyi Xiong. 2021. Efficient object-level visual context modeling for multimodal machine translation: Masking irrelevant objects helps grounding. In Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI 2021, Thirty-Third Conference on Innovative Applications of Artificial Intelligence, IAAI 2021, The Eleventh Symposium on Educational Advances in Artificial Intelligence, EAAI 2021, Virtual Event, February 2-9, 2021, pages 2720–2728. AAAI Press.
  27. 27.Zhiyong Wu, Lingpeng Kong, Wei Bi, Xiang Li, and Ben Kao. 2021. Good for misconceived reasons: An empirical revisiting on the need for visual context in multimodal machine translation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6153–6166, Online. Association for Computational Linguistics.
  28. 28.Shaowei Yao and Xiaojun Wan. 2020. Multimodal transformer for multimodal machine translation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4346–4350, Online. Association for Computational Linguistics.
  29. 29.Yongjing Yin, Fandong Meng, Jinsong Su, Chulun Zhou, Zhengyuan Yang, Jie Zhou, and Jiebo Luo. 2020. A novel graph-based multi-modal fusion encoder for neural machine translation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 3025–3035, Online. Association for Computational Linguistics.
  30. 30.Zhuosheng Zhang, Kehai Chen, Rui Wang, Masao Utiyama, Eiichiro Sumita, Zuchao Li, and Hai Zhao. 2020. Neural machine translation with universal visual representation. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  31. 31.Yuting Zhao, Mamoru Komachi, Tomoyuki Kajiwara, and Chenhui Chu. 2020. Double attention-based multimodal neural machine translation with semantic image regions. In Proceedings of the 22nd Annual Conference of the European Association for Machine Translation, pages 105–114, Lisboa, Portugal. European Association for Machine Translation.

Citation

MLA
Li, B., et al. “On Vision Features in Multimodal Machine Translation”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 6327–37, https://doi.org/10.18653/V1/2022.ACL-LONG.438.
APA
Li, B., Lv, C., Zhou, Z., Zhou, T., Xiao, T., Ma, A., & Zhu, J. (2022). On Vision Features in Multimodal Machine Translation. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6327–6337. https://doi.org/10.18653/V1/2022.ACL-LONG.438
Chicago
Li, B., C. Lv, Z. Zhou, et al. 2022. “On Vision Features in Multimodal Machine Translation”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6327–37. https://doi.org/10.18653/V1/2022.ACL-LONG.438.
Harvard
Li, B. et al. (2022) “On Vision Features in Multimodal Machine Translation”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 6327–6337. Available at: https://doi.org/10.18653/V1/2022.ACL-LONG.438.
Vancouver
1. Li B, Lv C, Zhou Z, Zhou T, Xiao T, Ma A, Zhu J (2022) On Vision Features in Multimodal Machine Translation. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 6327–6337

BibTeX

@inproceedings{Li_2022, title={On Vision Features in Multimodal Machine Translation}, url={http://dx.doi.org/10.18653/V1/2022.ACL-LONG.438}, DOI={10.18653/v1/2022.acl-long.438}, booktitle={Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)}, publisher={Association for Computational Linguistics}, author={Li, Bei and Lv, Chuanhao and Zhou, Zefan and Zhou, Tao and Xiao, Tong and Ma, Anxiang and Zhu, JingBo}, year={2022}, pages={6327–6337} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/