PreFLMR: Scaling Up Fine-Grained Late-Interaction Multi-modal Retrievers
Weizhe LinJingbiao MeiJinghong ChenBill Byrne
Introduces the M2KR benchmark and PreFLMR, a pre-trained late-interaction multimodal retriever that achieves state-of-the-art results across knowledge-based visual question answering tasks while providing the first systematic study of scaling behaviors in multimodal retrieval.
Modern multimodal artificial intelligence systems perform strongly on basic visual and language tasks but often fail when required to answer complex questions that demand external, specialized knowledge. In visual question answering scenarios requiring factual evidence, unaugmented models struggle significantly, frequently scoring below 20% accuracy on challenging tasks. To solve this problem, systems must reliably locate and retrieve relevant external documents from vast knowledge bases before generating answers. However, existing retrieval models are typically designed for single, isolated tasks, and technical understanding regarding how to scale these vision-language retrieval systems effectively has remained limited.
The article addresses this challenge by evaluating how scaling different model components and training data impacts multi-modal retrieval performance. The primary objective is to demonstrate that a single, general-purpose vision-language retriever can be trained to achieve superior retrieval accuracy across a diverse range of visual and textual tasks.
To conduct this evaluation, the authors compiled the Multi-task Multi-modal Knowledge Retrieval (M2KR) benchmark by unifying nine distinct datasets across three core retrieval formats: image-to-text, question-to-text, and combined image-and-question-to-text. Using this benchmark, the authors developed PreFLMR (Pre-trained Fine-grained Late-interaction Multi-modal Retriever). The model integrates vision encoders, text encoders, and a cross-attention mapping structure that enables query-aware visual understanding, scoring relevance through token-level interactions. PreFLMR was trained across four structured stages on a multimodal corpus exceeding ten million items and evaluated across multiple encoder sizes and task configurations.
The findings establish new performance benchmarks while identifying clear boundaries for model scaling. First, the best-performing PreFLMR configuration outperforms baseline models on seven out of nine benchmark datasets without requiring task-specific fine-tuning. Second, scaling the visual encoder from smaller configurations (ViT-B at 86 million parameters) up to larger architectures (ViT-G at 1.8 billion parameters) yields major retrieval gains across tasks, with recall improving by roughly 10 percentage points on complex datasets such as WIT, KVQA, and OVEN, though performance gains begin to plateau beyond large vision backbones. Third, scaling the text encoder yields diminishing returns: a standard 110-million-parameter text model (BERT-Base) delivers competitive accuracy, whereas larger 340-million-parameter text encoders lead to severe overfitting and training instability. Fourth, intermediate pre-training on high-quality, knowledge-dense data (E-VQA) markedly enhances general retrieval accuracy across other visual knowledge tasks. Finally, integrating PreFLMR into downstream question answering systems produces substantial performance improvements, boosting downstream accuracy by approximately 6% on OKVQA, 9% on Infoseek, and 34% on E-VQA over retrieval-free baselines.
These results demonstrate that organizations deploying visual artificial intelligence systems can achieve top-tier retrieval performance without overspending on oversized text backbones. Instead, investment should be directed toward larger vision encoders and high-quality pre-training corpora. The findings also highlight that task-specific prompt instructions are vital for multi-task stability, preventing severe cross-task performance degradation.
For practical implementation, organizations should adopt moderately sized text encoders paired with robust vision encoders, utilizing structured, multi-stage pre-training pipelines on verified data sources. In addition, because the retriever surfaces third-party knowledge directly to user-facing applications, operational pipelines must incorporate content moderation and database sanitization to avoid surfacing inappropriate or inaccurate source texts. Future development should focus on testing pre-trained vision models directly on knowledge-intensive domains and refining data balancing strategies.
The confidence in these findings is high across standard vision-language benchmarks, supported by systematic ablation studies and reproducible training runs. However, decision-makers should note certain limitations: performance on simpler benchmarks like OKVQA showed narrower gains due to noisy ground-truth training documents, and the underlying vision models were not pre-trained specifically on specialized domain data, meaning specialized enterprise deployments may require additional domain-specific calibration.
- Paper: OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge, Kenneth Marino et al. (2019). This paper establishes the OK-VQA benchmark requiring external unstructured knowledge retrieval for visual queries, forming the foundational problem setting and evaluation target scaled up by PreFLMR.
- Paper: BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models, Nandan Thakur et al. (2021). This work introduces the standard benchmark and principles for zero-shot information retrieval evaluation across dense and late-interaction architectures that PreFLMR generalizes to multimodal retrieval.
- Paper: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, Patrick Lewis et al. (2020). This seminal paper introduces retrieval-augmented generation architectures that combine parametric language generation with non-parametric external retrieval, which PreFLMR extends into multimodal domains.
- Paper: When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories, Alex Troy Mallen et al. (2022). This study demonstrates why scaling language model parameters alone fails to resolve long-tail factual knowledge gaps without external retrieval, motivating PreFLMR's multimodal retrieval framework.
- Paper: UNITER: UNiversal Image-TExt Representation Learning, Yen-Chun Chen et al. (2020). This paper establishes foundational cross-modal transformer alignment objectives and multi-task pre-training strategies that underpin vision-language representation learning in PreFLMR.
- Paper: DocVQA: A Dataset for VQA on Document Images, Minesh Mathew et al. (2020). This paper introduces document visual question answering benchmarks, providing key task formulations and datasets integrated into multi-task multimodal knowledge retrieval benchmarks.
- Paper: Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution, Peng Wang et al. (2024). This work extends multimodal foundation models to handle dynamic visual resolutions and fine-grained visual comprehension across dense documents and visual QA scenarios.
- Paper: RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs, Yue Yu et al. (2024). This paper advances retrieval-augmented generation pipelines by unifying context reranking with answer generation into a single instruction-tuned model.
- Paper: InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models, Jinguo Zhu et al. (2025). This study explores advanced unified pre-training and test-time scaling recipes for large open-source multimodal models handling complex multidisciplinary reasoning.
- Paper: Qwen3-VL Technical Report, Shuai Bai et al. (2025). This report demonstrates scaling up multimodal reasoning through cross-layer visual token injection and multi-stage pre-training for complex vision-language understanding.
- Paper: Enhancing Noise Robustness of Retrieval-Augmented Language Models with Adaptive Adversarial Training, Feiteng Fang et al. (2024). This research addresses downstream robustness against noisy or imperfect retrieved contexts in retrieval-augmented language models.
