ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval

Mengjun ChengYipeng SunLongchao WangXiongwei ZhuKun YaoJie ChenGuoli SongJunyu HanJingtuo LiuErrui Ding

article2022CVPR79 citations

Proposes a dual-encoder transformer architecture that fuses visual appearance with detected scene text via a shared fusion token and dual contrastive objectives, enabling unified cross-modal retrieval across both text-rich and text-free images while achieving faster inference.

Listen

Cross-modal retrieval systems, such as finding relevant images using natural language text queries, are vital for search engines and e-commerce platforms. While visual appearance is the primary signal for understanding images, text embedded directly within an image—known as scene text—often provides essential context for fine-grained identification. Existing systems largely struggle in this area: traditional models overlook embedded text entirely, while specialized text-aware models degrade in accuracy when images lack text. Furthermore, prevailing architectures with deep cross-modal interactions suffer from slow inference speeds that are impractical for large-scale production deployments.

The main objective of the article is to introduce and evaluate ViSTA (Vision and Scene Text Aggregation), a unified neural network framework that integrates visual appearance and scene text into a single model. The article demonstrates how ViSTA successfully handles both scene text-aware retrieval and conventional text-free retrieval within a fast, scalable architecture.

To achieve this, the authors designed a dual-encoder transformer architecture that processes images, scene text extracted via optical character recognition, and query text separately. Rather than merging modalities exhaustively, the system exchanges information between visual patches and recognized text exclusively through a specialized "fusion token." The network is trained end-to-end using dual contrastive learning losses—one matching image features with queries and another matching fused multimodal features with queries. The framework was evaluated across multiple industry-standard benchmarks, including the COCO-Text Captioned dataset for text-aware retrieval and the Flickr30K and MSCOCO datasets for conventional image-text retrieval, comparing retrieval accuracy and latency against leading methods.

The findings show that ViSTA significantly outperforms existing approaches across multiple settings. On scene text-aware retrieval using the CTC-1K benchmark, ViSTA improved top-1 image-to-text retrieval recall by 8.4% over previous state-of-the-art models. In conventional retrieval tasks where scene text is absent, ViSTA surpassed leading baselines while operating at least three times faster during inference than heavy single-encoder models, maintaining latency as low as 17 to 40 milliseconds. Ablation studies confirmed that isolating modality exchange to a shared fusion token and applying dual contrastive losses prevents performance drops when scene text is missing or noisy.

These results demonstrate that organizations can deploy a single, unified retrieval model across diverse enterprise search workflows without trading retrieval speed for semantic accuracy. By avoiding complex region-based object detectors and slow cross-attention mechanisms, the architecture reduces computational infrastructure costs while delivering superior search relevance across text-heavy and standard image catalogs alike.

Organizations planning to adopt this framework should implement it for catalog search and image retrieval pipelines where embedded text provides high business value, such as product labels or street scenes. Before deployment into production, engineering teams should conduct dataset cleaning and distribution auditing to mitigate potential risks associated with web-scraped training data, such as mislabeled pairs and demographic bias. Because the benefits of the fusion mechanism scale with the presence of text, stakeholders should verify the proportion and reliability of optical character recognition in their target domain to ensure optimal return on investment.

arXiv: 2203.16778
Cover for ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Approach
  • 3.1. Vision and Scene Text Encoders
  • 3.2. Vision and Scene Text Aggregation
  • 3.3. Cross-Modal Contrastive Learning
  • 4. Experiments
  • 4.1. Scene Text aware Cross-Modal Retrieval
  • 4.2. Scene Text Free Cross-Modal Retrieval
  • 4.3. Ablations
  • 5. Conclusions and Discussions
  • References

Knowls

  1. Knowl 1 — Unified dual-encoder for visual and scene-text retrieval

    model/method

    ViSTA is a dual-encoder framework that supports both conventional scene-text-free retrieval and retrieval enhanced by text detected inside images. For an image II, an OCR system provides a set of detected words and bounding boxes; a vision encoder produces an image representation, and a scene-text encoder processes the OCR tokens. A separate text transformer encodes a query qq into a text representation. ViSTA produces two image-side representations: an image-token representation for visual-appearance matching and a fusion-token representation that combines visual appearance with relevant scene-text semantics. Image and text representations can therefore be computed independently and compared efficiently with a dot product, while the scene-text information is incorporated into the image representation rather than requiring query-time cross-attention.

  2. Knowl 2 — Patch-based vision and OCR-aware transformer encoders

    model/method

    The ViSTA image encoder directly processes image patches instead of detector regions. An image is divided into NpN_p patches, linearly projected into tokens, augmented with positional embeddings, and prepended or appended with a learnable image token [IMG]. For layer ll, with input token sequence VlV_l and intermediate sequence YlY_l, the transformer update is

    Yl=MHSA(LN(Vl))+VlY_l = MHSA(LN(V_l)) + V_l Vl+1=MLP(LN(Yl))+YlV_{l+1} = MLP(LN(Y_l)) + Y_l

    Here MHSAMHSA is multi-head self-attention, MLPMLP is the feed-forward network, and LNLN is layer normalization. The final [IMG] token supplies the visual representation.

    For each OCR record ojo_j, the scene-text encoder combines the word embedding, a modality-type embedding, and a token-position embedding. The resulting sequence is processed by BERT and augmented with a linear projection of the normalized four-coordinate OCR bounding box ojbboxo_j^{bbox}:

    Sinit=Embedding(ojword)+Stype+Stoken_idS_{init} = Embedding(o_j^{word}) + S^{type} + S^{token\_id} S0=BERT(Sinit)+Flinear(ojbbox)S_0 = BERT(S_{init}) + F_{linear}(o_j^{bbox})

    The text transformer independently encodes the retrieval query and uses its [CLS] token as the query representation. All three encoders are trained end to end.

  3. Knowl 3 — Fusion-token bottleneck for mid-level modality aggregation

    model/method

    ViSTA aggregates image patches and OCR tokens through a shared fusion token rather than allowing all visual and scene-text tokens to interact directly. Let VlV_l be the visual-token sequence, SlS_l the scene-text-token sequence, and FlF_l the shared fusion token at aggregation layer ll. The visual and scene-text branches independently apply transformer blocks to their own tokens together with the fusion token:

    YlV=MHSA(LN([Vl;Fl]))+[Vl;Fl]Y_l^V = MHSA(LN([V_l;F_l])) + [V_l;F_l] [Vl+1;VlFUS]=MLP(LN(YlV))+YlV[V_{l+1};V_l^{FUS}] = MLP(LN(Y_l^V)) + Y_l^V YlS=MHSA(LN([Sl;Fl]))+[Sl;Fl]Y_l^S = MHSA(LN([S_l;F_l])) + [S_l;F_l] [Sl+1;SlFUS]=MLP(LN(YlS))+YlS[S_{l+1};S_l^{FUS}] = MLP(LN(Y_l^S)) + Y_l^S

    The next fusion token is the element-wise sum Fl+1=VlFUS+SlFUSF_{l+1}=V_l^{FUS}+S_l^{FUS}. The initial fusion token is formed as F0=Finit+Ftype+Ftoken_idF_0=F^{init}+F^{type}+F^{token\_id}. Thus, visual and scene-text tokens exchange information only through the fusion-token bottleneck, while the fusion token collects information from both modalities and becomes the scene-text-aware image representation. The experiments use four aggregation layers for the main ViSTA-S, ViSTA-B, and ViSTA-L configurations.

  4. Knowl 4 — Dual symmetric contrastive supervision

    equation

    ViSTA trains both the image-token branch and the fusion-token branch against query text in a common cross-modal embedding space. For a batch of NN matched image-text pairs, let viv_i, fif_i, and tit_i be the normalized image, fusion, and text embeddings for pair ii, respectively. The image-text and fusion-text losses are symmetric image-to-text/text-to-image contrastive objectives:

    Ltotal=alphaLitc+(1−alpha)LftcL_{total} = alpha L_{itc} + (1-alpha)L_{ftc} Lftc=(Lf2t+Lt2f)/2L_{ftc} = (L_{f2t}+L_{t2f})/2 Lf2t=−1/Nsumilog(exp(fidotti/sigma)/sumjexp(fidottj/sigma))L_{f2t} = -1/N sum_i log(exp(f_i dot t_i / sigma) / sum_j exp(f_i dot t_j / sigma)) Lt2f=−1/Nsumilog(exp(tidotfi/sigma)/sumjexp(tidotfj/sigma))L_{t2f} = -1/N sum_i log(exp(t_i dot f_i / sigma) / sum_j exp(t_i dot f_j / sigma)) Litc=(Li2t+Lt2i)/2L_{itc} = (L_{i2t}+L_{t2i})/2 Li2t=−1/Nsumilog(exp(vidotti/sigma)/sumjexp(vidottj/sigma))L_{i2t} = -1/N sum_i log(exp(v_i dot t_i / sigma) / sum_j exp(v_i dot t_j / sigma)) Lt2i=−1/Nsumilog(exp(tidotvi/sigma)/sumjexp(tidotvj/sigma))L_{t2i} = -1/N sum_i log(exp(t_i dot v_i / sigma) / sum_j exp(t_i dot v_j / sigma))

    Here sigmasigma is a trainable temperature initialized to 0.070.07, and alpha=0.9alpha=0.9 by default. The matched pair is the positive example and the other N−1N-1 batch elements provide negatives. The fusion-text term teaches the model to use scene-text semantics, whereas the image-text term preserves a strong visual representation when scene text is absent or noisy.

  5. Knowl 5 — Explicit handling of missing scene text

    model/method

    ViSTA selects its final image representation according to whether OCR detects scene text. If an image has no detected OCR records, the aggregation branch degenerates to the pure vision transformer and the [IMG] representation is used for retrieval. If OCR records are available, the image and scene-text encoders run through the fusion-token aggregation layers and the final [FUS] representation is used for scene-text-aware retrieval. During training, the image-text contrastive loss remains available to strengthen the visual branch, while the fusion-text contrastive loss is omitted for samples whose OCR result is empty. This design prevents direct dependence on a missing modality and allows one model to operate in both scene-text-free and scene-text-aware settings.

  6. Knowl 6 — Scene-text-aware retrieval performance

    empirical result

    On the COCO-Text Captioned benchmark, ViSTA-S was evaluated on both the 1K and 5K candidate settings using image-to-text and text-to-image retrieval. Values below are Recall@1/Recall@5/Recall@10 percentages; the strongest earlier comparator shown is STARNet.

    On CTC-1K, ViSTA-S achieved 52.5/77.9/87.2 for image-to-text retrieval, compared with STARNet's 44.1/74.8/82.7, and 36.7/66.2/77.8 for text-to-image retrieval, compared with STARNet's 31.5/60.8/72.4. On CTC-5K, ViSTA-S achieved 31.8/56.6/67.8 for image-to-text retrieval, compared with STARNet's 26.4/51.1/63.9, and 20.0/42.9/54.4 for text-to-image retrieval, compared with STARNet's 17.1/37.4/48.3. ViSTA-S therefore improved CTC-1K Recall@1 by 8.4 percentage points for image-to-text retrieval and 5.2 points for text-to-image retrieval over STARNet, while also outperforming the SCAN and VSRN baselines in every reported direction and candidate setting.

  7. Knowl 7 — Zero-shot scene-text-free retrieval with low inference cost

    empirical result

    For conventional scene-text-free retrieval, ViSTA was pretrained on the combined SBU, GCC, Visual Genome, and deduplicated MSCOCO training data and evaluated without task-specific fine-tuning on Flickr30K and MSCOCO. Retrieval is reported as image-to-text followed by text-to-image, with each triplet giving Recall@1/Recall@5/Recall@10 percentages. ViSTA-B required approximately 17 ms per inference and obtained 75.3/93.8/97.5 and 59.5/84.3/90.3 on Flickr30K, and 60.7/85.8/92.3 and 44.8/72.8/82.5 on MSCOCO. ViSTA-L required approximately 40 ms and obtained 79.2/95.4/98.1 and 67.0/88.7/93.1 on Flickr30K, and 63.9/87.1/93.0 and 47.4/75.0/84.0 on MSCOCO.

    Against the efficient ViLT-B baseline, which required approximately 15 ms and obtained 73.2/93.6/96.5 and 55.0/82.5/89.8 on Flickr30K plus 56.5/82.6/89.6 and 40.4/70.0/81.1 on MSCOCO, ViSTA-B and ViSTA-L improved retrieval accuracy while retaining near-real-time dual-tower inference. The result demonstrates that the fusion-token design does not degrade performance when scene text is unavailable.

  8. Knowl 8 — Fine-tuned retrieval performance on Flickr30K and MSCOCO

    empirical result

    After fine-tuning on the standard Flickr30K and MSCOCO retrieval splits, ViSTA-B required approximately 17 ms per inference and achieved image-to-text Recall@1/5/10 of 84.8/97.4/99.0 on Flickr30K and 63.9/87.8/93.6 on MSCOCO. Its text-to-image scores were 68.9/91.1/95.1 on Flickr30K and 47.8/75.8/84.5 on MSCOCO.

    The larger ViSTA-L required approximately 40 ms and achieved 89.5/98.4/99.6 for Flickr30K image-to-text, 75.8/94.2/96.9 for Flickr30K text-to-image, 68.9/90.1/95.4 for MSCOCO image-to-text, and 52.6/79.6/87.6 for MSCOCO text-to-image. ViLT-B required approximately 15 ms and scored 83.5/96.7/98.6, 64.4/88.7/93.8, 61.5/86.3/92.7, and 42.7/72.9/83.1 in the same order. Thus ViSTA-B improves on the similarly efficient ViLT-B in all four retrieval directions, while ViSTA-L gives the strongest reported results among these efficient models and remains substantially faster than the approximately 900 ms single-encoder systems evaluated by the paper.

  9. Knowl 9 — Ablation evidence for aggregation and dual losses

    data/table

    Ablations were conducted on CTC-1K with a fixed BERT-mini text tower; the values below are image-to-text and text-to-image Recall@1 percentages.

    Adding the proposed scene-text aggregation to ViSTA-S improved image-to-text/text-to-image Recall@1 from 47.0/34.6 without scene text to 52.5/36.7 with scene text. The corresponding scene-text-only baselines were much weaker: GCN alone scored 10.8/4.4 and BERT-mini alone scored 24.3/9.6. RoI+GCN scored 44.1/31.5, and ViT-S+GCN scored 47.2/33.2.

    Increasing the number of fusion layers from one to two to four changed the scores from 48.2/35.6 to 52.2/35.4 to 52.5/36.7, respectively. Among fusion mechanisms, global attention, cross-attention, the proposed fusion token, and late fusion scored 48.4/34.7, 50.5/31.1, 52.5/36.7, and 49.2/34.9, respectively. Finally, using only the fusion-text loss produced 46.6/30.3, whereas using both fusion-text and image-text losses produced 52.5/36.7. These results support the fusion-token bottleneck and show that the additional image-text supervision is important for retaining effective visual features.

  10. Knowl 10 — Dependence on relevant scene text and web-data quality

    limitation

    The benefit of scene-text aggregation is conditional rather than universal: it depends on how many images contain scene text that is relevant to the visual semantics of the retrieval task and on the correlation between the detected text and the image appearance. The paper also notes that large-scale web image-text training can contain distribution bias and mislabeled data, so additional data analysis, balancing, and cleaning are needed before deployment.

Coverage note — The qualitative retrieval examples and broader-impact discussion were not made separate knowls because they illustrate the method or its deployment context without adding a distinct core algorithm or quantitative result.

References

  1. 1.Xiang Bai, Mingkun Yang, Pengyuan Lyu, Yongchao Xu, and Jiebo Luo. Integrating scene text and visual appearance for fine-grained image classification. IEEE Access, 6:66322–66335, 2018. 3
  2. 2.Ali Furkan Biten, Ruben Tito, Andrés Mafla, Lluís Gomez i Bigorda, Marçal Rusinol, C. V. Jawahar, Ernest Valveny, and Dimosthenis Karatzas. Scene text visual question answering. In ICCV, pages 4290–4300. IEEE, 2019. 3
  3. 3.Hui Chen, Guiguang Ding, Xudong Liu, Zijia Lin, Ji Liu, and Jungong Han. IMRAM: iterative matching with recurrent attention memory for cross-modal image-text retrieval. In CVPR, pages 12652–12660. Computer Vision Foundation / IEEE, 2020. 1, 2, 7
  4. 4.Jiacheng Chen, Hexiang Hu, Hao Wu, Yuning Jiang, and Changhu Wang. Learning the best pooling strategy for visual semantic embedding. In CVPR, pages 15789–15798. Computer Vision Foundation / IEEE, 2021. 7
  5. 5.Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. UNITER: universal image-text representation learning. In ECCV (30), volume 12375 of Lecture Notes in Computer Science, pages 104–120. Springer, 2020. 1, 2, 6, 7
  6. 6.Sanghyuk Chun, Seong Joon Oh, Rafael Sampaio de Rezende, Yannis Kalantidis, and Diane Larlus. Probabilistic embeddings for cross-modal retrieval. In CVPR, pages 8415–8424. Computer Vision Foundation / IEEE, 2021. 7
  7. 7.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: pre-training of deep bidirectional transformers for language understanding. In NAACL-HLT (1), pages 4171–4186. Association for Computational Linguistics, 2019. 1, 4
  8. 8.Haiwen Diao, Ying Zhang, Lin Ma, and Huchuan Lu. Similarity reasoning and filtration for image-text matching. In AAAI, pages 1218–1226. AAAI Press, 2021. 2, 7
  9. 9.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR. OpenReview.net, 2021. 4
  10. 10.Fartash Faghri, David J. Fleet, Jamie Ryan Kiros, and Sanja Fidler. VSE++: improving visual-semantic embeddings with hard negatives. In BMVC, page 12. BMVA Press, 2018. 1, 2
  11. 11.Andrea Frome, Gregory S. Corrado, Jonathon Shlens, Samy Bengio, Jeffrey Dean, Marc’Aurelio Ranzato, and Tomas Mikolov. Devise: A deep visual-semantic embedding model. In NIPS, pages 2121–2129, 2013. 1, 2
  12. 12.Lluís Gomez, Andrés Mafla, Marçal Rusiñol, and Dimosthenis Karatzas. Single shot scene text retrieval. In ECCV (14), volume 11218 of Lecture Notes in Computer Science, pages 728–744. Springer, 2018. 3
  13. 13.Google. Cloud Vision API, 2020(accessed June 3, 2020). 4
  14. 14.Ronghang Hu, Amanpreet Singh, Trevor Darrell, and Marcus Rohrbach. Iterative answer prediction with pointer-augmented multimodal transformers for textvqa. In CVPR, pages 9989–9999. Computer Vision Foundation / IEEE, 2020. 4
  15. 15.Zhicheng Huang, Zhaoyang Zeng, Yupan Huang, Bei Liu, Dongmei Fu, and Jianlong Fu. Seeing out of the box: End-to-end pre-training for vision-language representation learning. In CVPR, pages 12976–12985. Computer Vision Foundation / IEEE, 2021. 1, 2, 3, 7
  16. 16.Zhicheng Huang, Zhaoyang Zeng, Bei Liu, Dongmei Fu, and Jianlong Fu. Pixel-bert: Aligning image pixels with text by deep multi-modal transformers. CoRR, abs/2004.00849, 2020. 1, 2, 3, 6, 7
  17. 17.Yuqi Huo, Manli Zhang, Guangzhen Liu, Haoyu Lu, Yizhao Gao, Guoxing Yang, Jingyuan Wen, Heng Zhang, Baogui Xu, Weihao Zheng, Zongzheng Xi, Yueqian Yang, Anwen Hu, Jinming Zhao, Ruichen Li, Yida Zhao, Liang Zhang, Yuqing Song, Xin Hong, Wanqing Cui, Dan Yang Hou, Yingyan Li, Junyi Li, Peiyu Liu, Zheng Gong, Chuhao Jin, Yuchong Sun, Shizhe Chen, Zhiwu Lu, Zhicheng Dou, Qin Jin, Yanyan Lan, Wayne Xin Zhao, Ruihua Song, and Ji-Rong Wen. Wenlan: Bridging vision and language by large-scale multi-modal pre-training. CoRR, abs/2103.06561, 2021. 2
  18. 18.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V. Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In ICML, volume 139 of Proceedings of Machine Learning Research, pages 4904–4916. PMLR, 2021. 2, 5
  19. 19.Andrej Karpathy and Li Fei-Fei. Deep visual-semantic alignments for generating image descriptions. IEEE Trans. Pattern Anal. Mach. Intell., 39(4):664–676, 2017. 6
  20. 20.Wonjae Kim, Bokyung Son, and Ildoo Kim. Vilt: Vision-and-language transformer without convolution or region supervision. In ICML, volume 139 of Proceedings of Machine Learning Research, pages 5583–5594. PMLR, 2021. 1, 2, 3, 6, 7
  21. 21.Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Li Fei-Fei. Visual genome: Connecting language and vision using crowdsourced dense image annotations. Int. J. Comput. Vis., 123(1):32–73, 2017. 2, 6
  22. 22.Kuang-Huei Lee, Xi Chen, Gang Hua, Houdong Hu, and Xiadong He. Stacked cross attention for image-text matching. In ECCV (4), volume 11208 of Lecture Notes in Computer Science, pages 212–228. Springer, 2018. 1, 2, 6, 7
  23. 23.Gen Li, Nan Duan, Yuejian Fang, Ming Gong, and Daxin Jiang. Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training. In AAAI, pages 11336–11344. AAAI Press, 2020. 1, 2, 6, 7
  24. 24.Junnan Li, Ramprasaath R. Selvaraju, Akhilesh Deepak Gotmare, Shafiq R. Joty, Caiming Xiong, and Steven C. H. Hoi. Align before fuse: Vision and language representation learning with momentum distillation. CoRR, abs/2107.07651, 2021. 2
  25. 25.Kunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li, and Yun Fu. Visual semantic reasoning for image-text matching. In ICCV, pages 4653–4661. IEEE, 2019. 1, 2, 6, 7
  26. 26.Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, Yejin Choi, and Jianfeng Gao. Oscar: Object-semantics aligned pre-training for vision-language tasks. In ECCV (30), volume 12375 of Lecture Notes in Computer Science, pages 121–137. Springer, 2020. 1, 2
  27. 27.Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. Microsoft COCO: common objects in context. In ECCV (5), volume 8693 of Lecture Notes in Computer Science, pages 740–755. Springer, 2014. 6
  28. 28.Chunxiao Liu, Zhendong Mao, Tianzhu Zhang, Hongtao Xie, Bin Wang, and Yongdong Zhang. Graph structured network for image-text matching. In CVPR, pages 10918–10927. Computer Vision Foundation / IEEE, 2020. 1, 2, 7
  29. 29.Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In NeurIPS, pages 13–23, 2019. 1, 2, 3, 6, 7
  30. 30.Jiasen Lu, Vedanuj Goswami, Marcus Rohrbach, Devi Parikh, and Stefan Lee. 12-in-1: Multi-task vision and language representation learning. In CVPR, pages 10434–10443. Computer Vision Foundation / IEEE, 2020. 7
  31. 31.Andres Mafla, Rafael Sampaio de Rezende, Lluís Gomez, Diane Larlus, and Dimosthenis Karatzas. Stacmr: Scene-text aware cross-modal retrieval. In WACV, pages 2219–2229. IEEE, 2021. 2, 3, 4, 6, 7
  32. 32.Antoine Miech, Jean-Baptiste Alayrac, Ivan Laptev, Josef Sivic, and Andrew Zisserman. Thinking fast and slow: Efficient text-to-visual retrieval with transformers. In CVPR, pages 9826–9836. Computer Vision Foundation / IEEE, 2021. 7
  33. 33.Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. Efficient estimation of word representations in vector space. In ICLR (Workshop Poster), 2013. 2
  34. 34.Arsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen, Cordelia Schmid, and Chen Sun. Attention bottlenecks for multimodal fusion. In NIPS, volume 34, 2021. 5
  35. 35.Vicente Ordonez, Girish Kulkarni, and Tamara L. Berg. Im2text: Describing images using 1 million captioned photographs. In NIPS, pages 1143–1151, 2011. 6
  36. 36.Di Qi, Lin Su, Jia Song, Edward Cui, Taroon Bharti, and Arun Sacheti. Imagebert: Cross-modal pre-training with large-scale weak-supervised image-text data. CoRR, abs/2001.07966, 2020. 1, 2, 6
  37. 37.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In ICML, volume 139 of Proceedings of Machine Learning Research, pages 8748–8763. PMLR, 2021. 2, 3
  38. 38.Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2556–2565, 2018. 6
  39. 39.Oleksii Sidorov, Ronghang Hu, Marcus Rohrbach, and Amanpreet Singh. Textcaps: A dataset for image captioning with reading comprehension. In ECCV (2), volume 12347 of Lecture Notes in Computer Science, pages 742–758. Springer, 2020. 3
  40. 40.Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach. Towards VQA models that can read. In CVPR, pages 8317–8326. Computer Vision Foundation / IEEE, 2019. 3
  41. 41.Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. VL-BERT: pre-training of generic visual-linguistic representations. In ICLR. OpenReview.net, 2020. 2, 7
  42. 42.Hao Tan and Mohit Bansal. LXMERT: learning cross-modality encoder representations from transformers. In EMNLP/IJCNLP (1), pages 5099–5110. Association for Computational Linguistics, 2019. 2
  43. 43.Hao Wang, Xiang Bai, Mingkun Yang, Shenggao Zhu, Jing Wang, and Wenyu Liu. Scene text retrieval via joint text detection and similarity learning. In CVPR, pages 4558–4567. Computer Vision Foundation / IEEE, 2021. 3
  44. 44.Hongwei Xue, Yupan Huang, Bei Liu, Houwen Peng, Jianlong Fu, Houqiang Li, and Jiebo Luo. Probing inter-modality: Visual parsing with self-attention for vision-language pre-training. CoRR, abs/2106.13488, 2021. 1, 2, 3, 7
  45. 45.Zhengyuan Yang, Yijuan Lu, Jianfeng Wang, Xi Yin, Dinei Florencio, Lijuan Wang, Cha Zhang, Lei Zhang, and Jiebo Luo. TAP: text-aware pre-training for text-vqa and text-caption. In CVPR, pages 8751–8761. Computer Vision Foundation / IEEE, 2021. 3
  46. 46.Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. Trans. Assoc. Comput. Linguistics, 2:67–78, 2014. 6
  47. 47.Fei Yu, Jiji Tang, Weichong Yin, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang. Ernie-vil: Knowledge enhanced vision-language representations through scene graphs. In AAAI, pages 3208–3216. AAAI Press, 2021. 1, 2, 7
  48. 48.Gangyan Zeng, Yuan Zhang, Yu Zhou, and Xiaomeng Yang. Beyond OCR + VQA: involving OCR into the flow for robust and accurate textvqa. In ACM Multimedia, pages 376–385. ACM, 2021. 3
  49. 49.Pengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang, Lei Zhang, Lijuan Wang, Yejin Choi, and Jianfeng Gao. Vinvl: Revisiting visual representations in vision-language models. In CVPR, pages 5579–5588. Computer Vision Foundation / IEEE, 2021. 1, 2
  50. 50.Qi Zhu, Chenyu Gao, Peng Wang, and Qi Wu. Simple is not easy: A simple strong baseline for textvqa and textcaps. In AAAI, pages 3608–3615. AAAI Press, 2021. 3

Citation

MLA
Cheng, M., et al. “ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval”. arXiv, 2022, http://arxiv.org/abs/2203.16778v1.
APA
Cheng, M., Sun, Y., Wang, L., Zhu, X., Yao, K., Chen, J., Song, G., Han, J., Liu, J., Ding, E., & Wang, J. (2022). ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval. arXiv. http://arxiv.org/abs/2203.16778v1
Chicago
Cheng, M., Y. Sun, L. Wang, et al. 2022. “ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval”. arXiv. http://arxiv.org/abs/2203.16778v1.
Harvard
Cheng, M. et al. (2022) “ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.16778v1.
Vancouver
1. Cheng M, Sun Y, Wang L, et al (2022) ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval. arXiv

BibTeX

@article{cheng2022vista,
  title = {ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval},
  author = {Cheng, Mengjun and Sun, Yipeng and Wang, Longchao and Zhu, Xiongwei and Yao, Kun and Chen, Jie and Song, Guoli and Han, Junyu and Liu, Jingtuo and Ding, Errui and Wang, Jingdong},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.16778v1},
  eprint = {2203.16778}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE