On Vision Features in Multimodal Machine Translation
Bei LiChuanhao LvZefan ZhouTao ZhouTong XiaoAnxiang MaJingbo Zhu
Reveals that stronger Transformer-based vision models improve multimodal machine translation on targeted probing tasks, while introducing a patch-level selective attention mechanism that directly correlates visual regions with words.
Multimodal machine translation combines text and images to translate across languages, but earlier research suggested that visual data provided little real value when complete text was present. Previous systems primarily relied on older visual encoders such as standard convolutional networks, assuming they were sufficient. The article investigates whether modern, high-capacity vision models—specifically Vision Transformers—and enhanced visual features can make images genuinely useful for translation, evaluating their impact through targeted testing scenarios.
The authors conducted an extensive empirical study across standard English-to-German and English-to-French datasets (Multi30K). They designed a selective attention mechanism that directly links text words with local image segments. To accurately assess visual contribution, they implemented probing tasks where key descriptive words—such as colors, characters, and nouns—were masked from the input text, forcing the translation system to extract missing information directly from the image. They also evaluated performance when models were presented with mismatched images.
The analysis yielded several critical findings. On complete text benchmarks, upgrading visual encoders produced only marginal score increases, confirming that standard metrics fail to reflect whether visual information is actually used. However, under probing conditions with incomplete text, Vision Transformer models dramatically outperformed traditional convolutional baselines, achieving accuracy improvements of over 20 percentage points in color- and character-recovery tasks. The selective attention mechanism proved essential, as finer-grained image patches and higher image resolutions allowed the model to accurately pinpoint and recover visual details. Additionally, when mismatched images were provided during decoding, models utilizing strong Vision Transformer features suffered substantial performance drops, proving that the system was actively relying on visual content rather than using it merely as generic noise or a training regularizer.
These results demonstrate that the visual modality is truly complementary and effective when powered by sufficiently strong visual architectures and fine-grained attention. Standard automated benchmark scores alone can be misleading, risking poor architectural decisions if systems are deployed without targeted probing. For organizations deploying multimodal translation in environments with noisy, ambiguous, or incomplete text inputs, utilizing modern Vision Transformer representations offers substantial quality improvements. Stakeholders should ensure that multimodal evaluation pipelines incorporate masked probing tasks before selecting models.
Future development should focus on creating unified architectures that jointly encode text and vision, as well as addressing the scarcity of large-scale multimodal translation datasets. Decision-makers should note that the current findings are primarily established on the Multi30K benchmark, so testing on specialized, domain-specific data remains recommended before full-scale operational rollout.
- Paper: Transformers in Vision: A Survey, Salman Khan et al. (2021). This survey lays out Vision Transformer architectures and their use in multimodal tasks, making the source’s choice of visual encoder and feature representation easier to follow.
- Paper: Do Vision Transformers See Like Convolutional Neural Networks?, Maithra Raghu et al. (2021). Its comparison of Vision Transformer and convolutional representations provides useful grounding for the source’s tests of whether ViT features recover visual details better than CNN features.
No sufficiently relevant recommendations were found.
