ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval
Mengjun ChengYipeng SunLongchao WangXiongwei ZhuKun YaoJie ChenGuoli SongJunyu HanJingtuo LiuErrui Ding
Proposes a dual-encoder transformer architecture that fuses visual appearance with detected scene text via a shared fusion token and dual contrastive objectives, enabling unified cross-modal retrieval across both text-rich and text-free images while achieving faster inference.
Cross-modal retrieval systems, such as finding relevant images using natural language text queries, are vital for search engines and e-commerce platforms. While visual appearance is the primary signal for understanding images, text embedded directly within an image—known as scene text—often provides essential context for fine-grained identification. Existing systems largely struggle in this area: traditional models overlook embedded text entirely, while specialized text-aware models degrade in accuracy when images lack text. Furthermore, prevailing architectures with deep cross-modal interactions suffer from slow inference speeds that are impractical for large-scale production deployments.
The main objective of the article is to introduce and evaluate ViSTA (Vision and Scene Text Aggregation), a unified neural network framework that integrates visual appearance and scene text into a single model. The article demonstrates how ViSTA successfully handles both scene text-aware retrieval and conventional text-free retrieval within a fast, scalable architecture.
To achieve this, the authors designed a dual-encoder transformer architecture that processes images, scene text extracted via optical character recognition, and query text separately. Rather than merging modalities exhaustively, the system exchanges information between visual patches and recognized text exclusively through a specialized "fusion token." The network is trained end-to-end using dual contrastive learning losses—one matching image features with queries and another matching fused multimodal features with queries. The framework was evaluated across multiple industry-standard benchmarks, including the COCO-Text Captioned dataset for text-aware retrieval and the Flickr30K and MSCOCO datasets for conventional image-text retrieval, comparing retrieval accuracy and latency against leading methods.
The findings show that ViSTA significantly outperforms existing approaches across multiple settings. On scene text-aware retrieval using the CTC-1K benchmark, ViSTA improved top-1 image-to-text retrieval recall by 8.4% over previous state-of-the-art models. In conventional retrieval tasks where scene text is absent, ViSTA surpassed leading baselines while operating at least three times faster during inference than heavy single-encoder models, maintaining latency as low as 17 to 40 milliseconds. Ablation studies confirmed that isolating modality exchange to a shared fusion token and applying dual contrastive losses prevents performance drops when scene text is missing or noisy.
These results demonstrate that organizations can deploy a single, unified retrieval model across diverse enterprise search workflows without trading retrieval speed for semantic accuracy. By avoiding complex region-based object detectors and slow cross-attention mechanisms, the architecture reduces computational infrastructure costs while delivering superior search relevance across text-heavy and standard image catalogs alike.
Organizations planning to adopt this framework should implement it for catalog search and image retrieval pipelines where embedded text provides high business value, such as product labels or street scenes. Before deployment into production, engineering teams should conduct dataset cleaning and distribution auditing to mitigate potential risks associated with web-scraped training data, such as mislabeled pairs and demographic bias. Because the benefits of the fusion mechanism scale with the presence of text, stakeholders should verify the proportion and reliability of optical character recognition in their target domain to ensure optimal return on investment.
- Paper: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, Alexey Dosovitskiy et al. (2021). Provides the foundational Vision Transformer (ViT) patch-based visual architecture that serves as the visual encoder backbone in ViSTA.
- Paper: EAST: An Efficient and Accurate Scene Text Detector, Xinyu Zhou et al. (2017). Introduces fast scene text detection methodology that underpins extracting embedded visual text before multimodal fusion in scene text retrieval systems.
- Paper: Towards VQA Models That Can Read, Amanpreet Singh et al. (2019). Formalizes the integration of optical character recognition into vision-and-language tasks, establishing foundational architectures for reading embedded scene text.
- Paper: Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks, Xiujun Li et al. (2020). Pioneers the use of detected textual semantic anchor tags paired with contrastive cross-modal objectives to align visual and linguistic representations.
- Paper: UNITER: UNiversal Image-TExt Representation Learning, Yen-Chun Chen et al. (2020). Establishes universal cross-modal Transformer pre-training and alignment paradigms, against which fast dual-encoder designs are evaluated.
- Paper: Stacked Cross Attention for Image-Text Matching, Kuang-Huei Lee et al. (2018). Introduces cross-attention mechanisms for fine-grained image-text matching, representing the heavy interaction approaches that ViSTA accelerates via token-based aggregation.
- Paper: Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval, Max Bain et al. (2021). Demonstrates efficient, fast dual-encoder Transformer architectures optimized for scalable cross-modal retrieval.
- Paper: Vision-Language Models for Vision Tasks: A Survey, Jingyi Zhang et al. (2023). Surveys the broader evolution of vision-language foundation architectures, contextualizing contrastive cross-modal retrieval models within modern multi-task frameworks.
- Paper: Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning, Hao Shao et al. (2024). Extends multimodal reasoning over embedded visual text and fine details through step-by-step visual chain-of-thought localization.
- Paper: VisionZip: Longer is Better but Not Necessary in Vision Language Models, Senqiao Yang et al. (2025). Builds on efficient token-level visual representations by pruning redundant visual tokens to accelerate multimodal inference.
- Paper: Text Is MASS: Modeling as Stochastic Embedding for Text-Video Retrieval, Jiamian Wang et al. (2024). Generalizes cross-modal dual-encoder retrieval by modeling textual queries as stochastic embeddings rather than single deterministic points.
- Paper: Harnessing the Power of MLLMs for Transferable Text-to-Image Person ReID, Wentao Tan et al. (2024). Applies cross-modal text-to-image retrieval principles to specific domain transfer tasks using synthetic caption generation and token-level noise filtering.
