Tackling Ambiguity with Images: Improved Multimodal Machine Translation and Contrastive Evaluation
Matthieu FuteralCordelia SchmidIvan LaptevBenoît SagotRachel Bawden
Proposes an adapter-based multimodal machine translation framework with guided self-attention alongside CoMMuTE, a contrastive evaluation benchmark designed to verify whether models effectively use visual context to resolve lexical ambiguity.
Translating ambiguous text accurately remains a persistent challenge for automated translation systems. While accompanying images can provide the visual context necessary to resolve ambiguities, existing multimodal translation systems have struggled to outperform strong text-only baselines. Furthermore, the field has been hampered by standard benchmarks that rarely require image context to translate correctly, making it difficult to assess whether models are genuinely using visual information.
The article evaluates a new multimodal translation architecture and introduces a targeted evaluation benchmark to determine whether integrating visual context can resolve lexical ambiguity across multiple languages without degrading overall translation quality.
To achieve this, the authors developed VGAMT (Visually Guided and Adapted Machine Translation). The approach freezes a strong, pretrained text-only translation model (mBART fine-tuned on parallel corpora) and adapts it using lightweight neural adapter modules and a novel guided self-attention mechanism that pairs relevant text words directly to image regions. The architecture was jointly trained on multimodal machine translation and visually conditioned masked language modeling using approximately 2 million image-text pairs from the Conceptual Captions dataset alongside the standard Multi30k dataset. To properly evaluate ambiguity resolution, the authors created CoMMuTE, a contrastive evaluation dataset comprising 155 ambiguous English sentences paired with alternative images and native translations across French, German, and Czech.
The experimental findings demonstrate significant improvements in visual disambiguation. First, VGAMT outperformed existing multimodal systems and balanced text-only baselines on the CoMMuTE benchmark, achieving an accuracy of 67.1% in English-to-French, 59.0% in English-to-German, and 55.6% in English-to-Czech, where text-only baselines achieve only the random chance baseline of 50.0%. Second, on standard Multi30k benchmarks where visual context is rarely required, VGAMT performed competitively with strong text-only models, demonstrating that multimodal integration did not cause translation quality degradation. Third, ablation analyses revealed that joint training with masked language modeling and the use of global image features were critical; omitting masked pretraining or image features caused accuracy on ambiguous sentences to drop sharply toward baseline levels.
These results demonstrate that multimodal machine translation can effectively leverage visual context without requiring expensive full-model retraining from scratch. Utilizing parameter-efficient adapters and guided cross-modal attention allows organizations to preserve the broad linguistic capabilities of large text-only systems while improving translation accuracy in visually grounded environments. The findings also underscore that standard evaluation metrics and captioning datasets are inadequate for measuring multimodal performance, emphasizing the need for targeted contrastive benchmarks.
Based on these findings, teams developing machine translation for image-grounded content (such as e-commerce, media subtitling, and catalog translation) should consider adapter-based multimodal architectures rather than training multimodal models from scratch. In addition, evaluation pipelines should integrate contrastive test suites like CoMMuTE rather than relying exclusively on standard text-overlap metrics. Prior to production deployments, further engineering is recommended to extend the underlying object-detection dependencies beyond English source texts and to optimize the computational cost associated with large caption pretraining datasets.
- Paper: On Vision Features in Multimodal Machine Translation, Bei Li et al. (2022). Read this study first to see how visual features and word-to-image attention can contribute to translation, and why targeted tests are needed to reveal that contribution.
No sufficiently relevant recommendations were found.
