PreFLMR: Scaling Up Fine-Grained Late-Interaction Multi-modal Retrievers

Weizhe LinJingbiao MeiJinghong ChenBill Byrne

article2024ACL52 citations

Introduces the M2KR benchmark and PreFLMR, a pre-trained late-interaction multimodal retriever that achieves state-of-the-art results across knowledge-based visual question answering tasks while providing the first systematic study of scaling behaviors in multimodal retrieval.

Listen

Modern multimodal artificial intelligence systems perform strongly on basic visual and language tasks but often fail when required to answer complex questions that demand external, specialized knowledge. In visual question answering scenarios requiring factual evidence, unaugmented models struggle significantly, frequently scoring below 20% accuracy on challenging tasks. To solve this problem, systems must reliably locate and retrieve relevant external documents from vast knowledge bases before generating answers. However, existing retrieval models are typically designed for single, isolated tasks, and technical understanding regarding how to scale these vision-language retrieval systems effectively has remained limited.

The article addresses this challenge by evaluating how scaling different model components and training data impacts multi-modal retrieval performance. The primary objective is to demonstrate that a single, general-purpose vision-language retriever can be trained to achieve superior retrieval accuracy across a diverse range of visual and textual tasks.

To conduct this evaluation, the authors compiled the Multi-task Multi-modal Knowledge Retrieval (M2KR) benchmark by unifying nine distinct datasets across three core retrieval formats: image-to-text, question-to-text, and combined image-and-question-to-text. Using this benchmark, the authors developed PreFLMR (Pre-trained Fine-grained Late-interaction Multi-modal Retriever). The model integrates vision encoders, text encoders, and a cross-attention mapping structure that enables query-aware visual understanding, scoring relevance through token-level interactions. PreFLMR was trained across four structured stages on a multimodal corpus exceeding ten million items and evaluated across multiple encoder sizes and task configurations.

The findings establish new performance benchmarks while identifying clear boundaries for model scaling. First, the best-performing PreFLMR configuration outperforms baseline models on seven out of nine benchmark datasets without requiring task-specific fine-tuning. Second, scaling the visual encoder from smaller configurations (ViT-B at 86 million parameters) up to larger architectures (ViT-G at 1.8 billion parameters) yields major retrieval gains across tasks, with recall improving by roughly 10 percentage points on complex datasets such as WIT, KVQA, and OVEN, though performance gains begin to plateau beyond large vision backbones. Third, scaling the text encoder yields diminishing returns: a standard 110-million-parameter text model (BERT-Base) delivers competitive accuracy, whereas larger 340-million-parameter text encoders lead to severe overfitting and training instability. Fourth, intermediate pre-training on high-quality, knowledge-dense data (E-VQA) markedly enhances general retrieval accuracy across other visual knowledge tasks. Finally, integrating PreFLMR into downstream question answering systems produces substantial performance improvements, boosting downstream accuracy by approximately 6% on OKVQA, 9% on Infoseek, and 34% on E-VQA over retrieval-free baselines.

These results demonstrate that organizations deploying visual artificial intelligence systems can achieve top-tier retrieval performance without overspending on oversized text backbones. Instead, investment should be directed toward larger vision encoders and high-quality pre-training corpora. The findings also highlight that task-specific prompt instructions are vital for multi-task stability, preventing severe cross-task performance degradation.

For practical implementation, organizations should adopt moderately sized text encoders paired with robust vision encoders, utilizing structured, multi-stage pre-training pipelines on verified data sources. In addition, because the retriever surfaces third-party knowledge directly to user-facing applications, operational pipelines must incorporate content moderation and database sanitization to avoid surfacing inappropriate or inaccurate source texts. Future development should focus on testing pre-trained vision models directly on knowledge-intensive domains and refining data balancing strategies.

The confidence in these findings is high across standard vision-language benchmarks, supported by systematic ablation studies and reproducible training runs. However, decision-makers should note certain limitations: performance on simpler benchmarks like OKVQA showed narrower gains due to noisy ground-truth training documents, and the underlying vision models were not pre-trained specifically on specialized domain data, meaning specialized enterprise deployments may require additional domain-specific calibration.

arXiv: 2402.08327
Cover for PreFLMR: Scaling Up Fine-Grained Late-Interaction Multi-modal Retrievers

Abstract

Large Multimodal Models (LMMs) excel in natural language and visual understanding but are challenged by exacting tasks such as Knowledge-based Visual Question Answering (KB-VQA) which involve the retrieval of relevant information from document collections to use in shaping answers to questions. We present an extensive training and evaluation framework, M2KR, for KB-VQA. M2KR contains a collection of vision and language tasks which we have incorporated into a single suite of benchmark tasks for training and evaluating general-purpose multi-modal retrievers. We use M2KR to develop PreFLMR, a pre-trained version of the recently developed Fine-grained Late-interaction Multi-modal Retriever (FLMR) approach to KB-VQA, and we report new state-of-the-art results across a range of tasks. We also present investigations into the scaling behaviors of PreFLMR intended to be useful in future developments in general-purpose multi-modal retrievers. The code, demo, dataset, and pre-trained checkpoints are available at https://preflmr.github.io/.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 The M2KR Benchmark Suite
  • 3.1 Tasks and Datasets
  • 3.2 Evaluation
  • 3.3 Baselines
  • 4 PreFLMR Architecture and Training
  • 4.1 Training Procedures
  • 4.2 Training Configurations
  • 5 Experiments and Results
  • 5.1 Model Variants
  • 5.2 PreFLMR Performance
  • 5.3 Performance of Each PreFLMR Stage
  • 5.4 Ablation Studies
  • 5.5 Retrieval Augmented Visual Question Answering with PreFLMR
  • 5.6 Analysis of Intermediate Pre-training
  • 5.7 Summary of Findings
  • 6 Conclusion
  • Limitations
  • Ethics Statement
  • Acknowledgments
  • References
  • A Datasets details
  • A.1 I2T Retrieval
  • A.1.1 WIT
  • A.1.2 IGLUE
  • A.1.3 KVQA
  • A.1.4 CC3M
  • A.2 Q2T Retrieval
  • A.2.1 MSMARCO
  • A.3 IQ2T Retrieval
  • A.3.1 LLaVA
  • A.3.2 OVEN
  • A.3.3 E-VQA, Infoseek and OKVQA
  • B Implementation Details
  • B.1 Breakdown of Data Used in Training
  • B.2 Detailed Hyperparameters
  • B.3 Large-v1 Training
  • B.4 Model Design in Detail
  • C Ablation Study on Pre-training Stages
  • D V-Entropy-based Analysis of Intermediate Pre-training
  • E Qualitative Analysis for OKVQA and E-VQA
  • F Artifacts and License
  • H AI Assistance

Knowls

  1. Knowl 1 — PreFLMR Retriever Architecture

    model/method

    PreFLMR (Pre-trained Fine-grained Late-interaction Multi-modal Retriever) is an information retrieval model for Knowledge-based Visual Question Answering (KB-VQA) and multi-modal knowledge retrieval. Given a multi-modal query qˉ=(q,I)\bar{q} = (q, I) comprising a text string qq (with NqN_q tokens) and an image II, along with a document text dd (with lDl_D tokens):

    1. Text Query Encoding: A language model FLF_L encodes text qq into token embeddings Qq=FL(q)∈RNq×dLQ_q = F_L(q) \in \mathbb{R}^{N_q \times d_L}, where dLd_L is the text hidden dimension.
    2. Vision Query Encoding: A Vision Transformer (ViT) FVF_V encodes image II. The [CLS][\text{CLS}] token embedding from the final layer is extracted as QI,[CLS]=FV(I)∈R1×dVQ_{I,[\text{CLS}]} = F_V(I) \in \mathbb{R}^{1 \times d_V}. In addition, patch embeddings from ViT's penultimate layer are extracted as QI,PATCH=FV,−2(I)∈RNV×dVQ_{I,\text{PATCH}} = F_{V,-2}(I) \in \mathbb{R}^{N_V \times d_V}, where NVN_V is the number of image patches and dVd_V is the visual hidden dimension.
    3. Mapping Structure FMF_M: The visual features are mapped into the text embedding space via two parallel paths:
      • A 2-layer Multi-Layer Perceptron (MLP) FMMLPF_M^{\text{MLP}} converts QI,[CLS]Q_{I,[\text{CLS}]} into NvtN_{vt} visual token embeddings: QIMLP=FMMLP(QI,[CLS])∈RNvt×dhQ_I^{\text{MLP}} = F_M^{\text{MLP}}(Q_{I,[\text{CLS}]}) \in \mathbb{R}^{N_{vt} \times d_h}.
      • A Transformer cross-attention module FMTRF_M^{\text{TR}} (with NTRN_{\text{TR}} layers) takes projected patch features Fv(QI,PATCH)∈RNV×dLF_v(Q_{I,\text{PATCH}}) \in \mathbb{R}^{N_V \times d_L} and attends to the query text embeddings QqQ_q, producing query-aware patch embeddings QITR=FMTR(Fv(QI,PATCH),Qq)∈RNV×dhQ_I^{\text{TR}} = F_M^{\text{TR}}(F_v(Q_{I,\text{PATCH}}), Q_q) \in \mathbb{R}^{N_V \times d_h}, where FvF_v is a 1-layer MLP mapping from dVd_V to dLd_L, and dhd_h is the joint projection dimension.
    4. Query Representation: All token-level query embeddings are concatenated along the sequence length dimension: Q=[Qq∣QIMLP∣QITR]∈RlQ×dhQ = [Q_q \mid Q_I^{\text{MLP}} \mid Q_I^{\text{TR}}] \in \mathbb{R}^{l_Q \times d_h} where lQ=Nq+Nvt+NVl_Q = N_q + N_{vt} + N_V.
    5. Document Representation: For a document text dd, text encoder FLF_L and an output linear projection FlF_l produce document token embeddings: D=Fl(FL(d))∈RlD×dhD = F_l(F_L(d)) \in \mathbb{R}^{l_D \times d_h} where lDl_D is the number of tokens in dd.
  2. Knowl 2 — Multi-Modal Late-Interaction Relevance Scoring and Training Loss

    equation

    In late-interaction multi-modal retrieval, the relevance score r(qˉ,d)r(\bar{q}, d) between a multi-modal query qˉ=(q,I)\bar{q} = (q, I) with representation matrix Q∈RlQ×dhQ \in \mathbb{R}^{l_Q \times d_h} and a text document dd with representation matrix D∈RlD×dhD \in \mathbb{R}^{l_D \times d_h} is computed as the sum of maximum inner products across all query tokens:

    r(qˉ,d)=∑i=1lQmax⁡j=1lDQiDj⊤r(\bar{q}, d) = \sum_{i=1}^{l_Q} \max_{j=1}^{l_D} Q_i D_j^\top

    where Qi∈R1×dhQ_i \in \mathbb{R}^{1 \times d_h} denotes the ii-th query token embedding (encompassing text, global visual, and query-aware patch visual tokens), Dj∈R1×dhD_j \in \mathbb{R}^{1 \times d_h} denotes the jj-th document token embedding, and lQl_Q and lDl_D denote the sequence lengths of the query and document, respectively.

    The retriever is trained using in-batch negative examples N(q)\mathcal{N}(q) and random negatives sampled from the document collection. Given a gold positive document d∗d^*, the contrastive loss over training dataset D\mathcal{D} is defined as:

    L=−∑(q,d∗)∈Dlog⁡exp⁡(r(q,d∗))exp⁡(r(q,d∗))+∑z∈N(q)exp⁡(r(q,z))\mathcal{L} = -\sum_{(q, d^*) \in \mathcal{D}} \log \frac{\exp(r(q, d^*))}{\exp(r(q, d^*)) + \sum_{z \in \mathcal{N}(q)} \exp(r(q, z))}

  3. Knowl 3 — M2KR Benchmark Suite for Multi-Modal Knowledge Retrieval

    experimental setup

    The Multi-task Multi-modal Knowledge Retrieval (M2KR) benchmark suite unifies nine vision-language datasets into a standardized retrieval format covering three task types:

    1. Image-to-Text (I2T) Retrieval: Evaluates retrieval of relevant text documents given an image query paired with an instruction prompt.
      • WIT: 2.8M train, 20,102 validation, 5,120 test examples; 4.1M training passages, 40K validation/test passage corpus.
      • IGLUE-en: 685 test examples evaluated against a 1K Wikipedia passage corpus.
      • KVQA: Reformulated to retrieve biographical and entity details from images alone; 65K train, 13,365 validation, 5,120 test examples; 16.3K training passages, 4,648 validation/test passages.
      • CC3M: 595K training examples treated as image-caption retrieval to facilitate scene understanding (training only).
    2. Question-to-Text (Q2T) Retrieval: Evaluates preservation of text-only retrieval capability by pairing blank images with text questions.
      • MSMARCO: 400K train, 6,980 validation, 5,120 test queries evaluated against an 8.8M passage collection (sampled to 200K/400K for validation/test splits).
    3. Image & Question to Text (IQ2T) Retrieval: Requires joint multi-modal reasoning.
      • OVEN: 339K train, 20,000 validation, 5,120 test examples; 10K training passages, 3,192 validation/test passages.
      • LLaVA: Multi-modal conversation turns repurposed into question-image to response retrieval; 351K train, 5,120 test examples; 351K training passages, 6,006 test passages.
      • OKVQA: 9K train, 5,046 validation, 5,046 test examples; 110K passage corpus.
      • Infoseek: 100K train, 4,708 test examples; 100K passage corpus.
      • E-VQA: 212K (167K filtered) train, 9,852 validation, 3,750 test examples; 50K passage corpus based on WikiWeb2M.

    Evaluation is reported using Recall@K (R@K, measuring if a ground-truth document is retrieved in top KK) or Pseudo-Recall@K (PR@K, measuring if a target answer string is contained in the top KK retrieved documents).

  4. Knowl 4 — Four-Stage Pre-training Strategy for PreFLMR

    algorithm

    The PreFLMR pre-training pipeline consists of four sequential stages:

    Input: Vision encoder FVF_V, text encoder FLF_L, mapping module FM=(FMMLP,FMTR)F_M = (F_M^{\text{MLP}}, F_M^{\text{TR}}), document projection FlF_l, datasets in M2KR
    Output: Trained multi-task multi-modal retriever parameters
    Stage 0 (Text Encoder Pre-training):
      Train FLF_L and FlF_l as a ColBERT retriever on MSMARCO for up to 300k steps using Adam (learning rate 10−510^{-5}).
      Select checkpoint at best Recall@50 on MSMARCO validation.
    Stage 1 (Mapping Structure Training):
      Freeze FLF_L and FVF_V.
      Train FMF_M on a multi-task mix (LLaVA, OVEN, WIT, CC3M, KVQA, MSMARCO) for up to 220k steps with learning rate 10−410^{-4}.
      Mask late-interaction token embeddings produced directly by FLF_L to force FMTRF_M^{\text{TR}} cross-attention to integrate text features.
      Select checkpoint at best Recall@10 on WIT validation.
    Stage 2 (Intermediate KB-VQA Pre-training):
      Unfreeze FLF_L and train FLF_L and FMF_M jointly on E-VQA for 12k steps.
      Learning rate: 10−410^{-4} for FMF_M, 10−510^{-5} for FLF_L.
    Stage 3 (Full-Scale Multi-Task Fine-Tuning):
      Train FMF_M and decoupled query/document text encoders on all M2KR datasets for up to 50k steps.
      Keep FVF_V frozen. Rebalance dataset batch sampling proportions.
      Apply early stopping if WIT or E-VQA validation performance degrades for 3 consecutive validation checks.
  5. Knowl 5 — Encoder Scaling Properties in Late-Interaction Multi-Modal Retrieval

    empirical result

    Experiments evaluating vision and text encoder scaling across the 9 tasks of the M2KR benchmark reveal the following behaviors:

    1. Vision Encoder Scaling: Scaling the Vision Transformer backbone from ViT-B (86M/88M) →\to ViT-L (303M) →\to ViT-H (631M) →\to ViT-G (1.84B) while keeping the text encoder fixed at BERT-Base-v2 yields continuous gains on knowledge-intensive retrieval: Infoseek PR@5 increases from 48.8 (ViT-B) →\to 57.9 (ViT-L) →\to 59.5 (ViT-H) →\to 59.6 (ViT-G); E-VQA PR@5 increases from 67.9 (ViT-B) →\to 70.8 (ViT-L) →\to 71.7 (ViT-H) →\to 73.1 (ViT-G); and average rank across all tasks improves from 8.2 →\to 3.2 →\to 3.1 →\to 1.6. Gains are largest between ViT-B and ViT-L (approx10% approx 10\% recall gains on WIT, KVQA, OVEN, and Infoseek) and begin to plateau between ViT-H and ViT-G.
    2. Text Encoder Scaling: Scaling the text encoder from BERT-Small-v1 (28.8M) →\to BERT-Medium-v1 (41.1M) →\to BERT-Base-v1 (110M) improves multi-task average rank from 8.3 →\to 8.2 →\to 5.6. However, scaling further to BERT-Large-v1 (340M) degrades multi-task average rank to 6.6 (e.g., WIT R@10 drops from 58.2 to 49.9, E-VQA PR@5 drops from 70.7 to 58.2) due to severe overfitting and training instability during multi-modal tuning, even though BERT-Large achieves higher unimodal text retrieval accuracy on MSMARCO.
  6. Knowl 6 — Retrieval-Augmented Generation Performance with PreFLMR for KB-VQA

    empirical result

    When integrated into a Retrieval-Augmented Generation pipeline (RA-VQAv2) using a BLIP2-T5XL answer generator, PreFLMR (ViT-G + Base-v2) substantially outperforms systems without retrieval and prior retrieval-augmented models across downstream KB-VQA benchmarks:

    Model OKVQA (VQA Score) Infoseek (Accuracy) E-VQA (BEM)
    Generator w/o Retrieval 55.44 21.78 19.80
    RA-VQAv2 w/ FLMR 60.75 - -
    RA-VQAv2 w/ PreFLMR 61.88 30.65 54.45

    Compared to generation without retrieval, PreFLMR improves downstream question answering performance by +6.44%+6.44\% on OKVQA, +8.87%+8.87\% on Infoseek, and +34.65%+34.65\% on E-VQA. The larger gains on E-VQA and Infoseek demonstrate that specialized external knowledge retrieval provides greater benefit on expert-level domain entities compared to common-sense-dominated datasets like OKVQA.

  7. Knowl 7 — Ablation of Pre-training Stages, Prompts, and Mapping Depth

    empirical result

    Component-level ablations of PreFLMR establish the necessity of individual architectural and pre-training designs:

    1. Pre-training Stages: On a ViT-B + Base-v1 model, removing Stage 0 (ColBERT text pre-training) degrades average retrieval performance the most severely (WIT R@10 drops from 41.7 to 25.5, MSMARCO R@5 from 79.5 to 56.5, E-VQA PR@5 from 67.9 to 51.8). Removing Stage 1 (mapping pre-training) causes drops on WIT (41.7 →\to 38.2) and Infoseek (48.8 →\to 44.6). Removing Stage 2 (intermediate KB-VQA training) causes a large drop on E-VQA (67.9 →\to 57.3).
    2. Task-Specific Instructions: Omitting prompting instructions during multi-task training collapses multi-task performance (WIT R@10 drops from 60.5 to 13.3, IGLUE R@1 from 69.2 to 10.5 on ViT-L + Base-v2).
    3. Mapping Structure Depth: Increasing the number of Transformer cross-attention layers NTRN_{\text{TR}} in FMF_M from 1 to 4 marginally improves LLaVA R@1 (+0.5+0.5) but degrades WIT R@10 (from 49.6 to 45.9) and Infoseek PR@5 (from 48.7 to 46.8) for ViT-L + Base-v2, showing that a 1-layer cross-attention design is optimal.
  8. Knowl 8 — V-Entropy Formulation for Intermediate Transfer Gain and Overfitting

    theoretical result

    The benefit of Stage 2 intermediate pre-training on dataset D1D_1 (E-VQA) for target multi-task dataset D2D_2 (M2KR) under computation constraints is formalized using V\mathcal{V}-Entropy:

    IV[Nf](D1→D2)=HV[Nf](D2)−HV[Nf,D1,Nt](D2)I_{\mathcal{V}[N_f]}(D_1 \to D_2) = H_{\mathcal{V}[N_f]}(D_2) - H_{\mathcal{V}[N_f, D_1, N_t]}(D_2)

    where HV[Nf](D2)H_{\mathcal{V}[N_f]}(D_2) denotes the minimum negative log-likelihood (NLL) loss on the validation set of D2D_2 after NfN_f training steps on D2D_2 alone, and HV[Nf,D1,Nt](D2)H_{\mathcal{V}[N_f, D_1, N_t]}(D_2) denotes the minimum NLL loss on D2D_2 achieved after NfN_f steps when initialized from a model pre-trained on D1D_1 for NtN_t steps. Here, V\mathcal{V} defines the reachable model family under the specified architecture.

    Empirical evaluation of validation loss across intermediate steps Ninter=NtN_{\text{inter}} = N_t reveals that IV[Nf](D1→D2)I_{\mathcal{V}[N_f]}(D_1 \to D_2) is strongly positive for knowledge-intensive datasets (OKVQA, KVQA, OVEN), but an optimal NinterN_{\text{inter}} exists (Ninter≈10,000N_{\text{inter}} \approx 10{,}000 for BERT-Base, Ninter≈15,000N_{\text{inter}} \approx 15{,}000 for BERT-Medium), beyond which the model overfits to D1D_1 and validation loss on other tasks increases.

  9. Knowl 9 — Stage 3 Data Proportions and LoRA Regularization for Large Text Backbones

    model/method

    To prevent training loss collapse and over-specialization on simpler or smaller datasets during Stage 3 full-scale fine-tuning, training sample counts are rescaled as follows:

    • WIT: downsampled from 2.8M to 140K
    • CC3M: downsampled from 595K to 29.8K
    • MSMARCO: downsampled from 400K to 40K
    • OVEN: downsampled from 339K to 33.9K
    • LLaVA: downsampled from 351K to 35.1K
    • KVQA: downsampled from 65K to 6.5K
    • OKVQA: oversampled by factor of 10 from 9K to 90K
    • Infoseek: downsampled from 100K to 50K
    • E-VQA: maintained at 167K

    For large text backbones (BERT-Large-v1), standard full parameter fine-tuning in Stages 2 and 3 causes optimization collapse even at reduced learning rates (10−610^{-6}, 3×10−63 \times 10^{-6}). Stable training is achieved by applying Low-Rank Adaptation (LoRA) to BERT-Large with rank r=16r = 16, scaling factor α=32\alpha = 32, and dropout rate 0.050.05.

  10. Knowl 10 — Limitations of PreFLMR and M2KR

    limitation

    The authors identify three primary limitations:

    1. Unadapted Vision Pre-training: The CLIP-ViT vision backbones used in PreFLMR were pre-trained on general web image-text pairs rather than domain-specific in-domain data for knowledge-intensive retrieval tasks, potentially limiting the model's ability to discriminate fine-grained, specialized entities.
    2. Training Objective Scope: Training is restricted to contrastive loss with hard negatives; advanced optimization techniques such as late-interaction score distillation (as used in ColBERTv2) were not explored for the multi-modal retriever.
    3. Heuristic Dataset Mixing: The multi-task dataset proportions across M2KR during Stage 3 fine-tuning were adjusted empirically rather than through formal optimization of mixing ratios.

Coverage note — None was omitted; all contributed benchmark specifications, model architectures, mathematical formulas, four-stage pre-training algorithms, scaling laws/empirical comparisons, downstream RAG evaluations, and ablations are covered.

References

  1. 1.Ibrahim M Alabdulmohsin, Behnam Neyshabur, and Xiaohua Zhai. 2022. Revisiting neural scaling laws in language and vision. In Advances in Neural Information Processing Systems, volume 35, pages 22300–22312. Curran Associates, Inc.
  2. 2.Rohan Anil, Andrew M. Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, Eric Chu, Jonathan H. Clark, Laurent El Shafey, Yanping Huang, Kathy Meier-Hellstern, Gaurav Mishra, Erica Moreira, Mark Omernick, Kevin Robinson, Sebastian Ruder, Yi Tay, Kefan Xiao, Yuanzhong Xu, Yujing Zhang, Gustavo Hernandez Abrego, Junwhan Ahn, Jacob Austin, Paul Barham, Jan Botha, James Bradbury, Siddhartha Brahma, Kevin Brooks, Michele Catasta, Yong Cheng, Colin Cherry, Christopher A. Choquette-Choo, Aakanksha Chowdhery, Clément Crepy, Shachi Dave, Mostafa Dehghani, Sunipa Dev, Jacob Devlin, Mark Díaz, Nan Du, Ethan Dyer, Vlad Feinberg, Fangxiaoyu Feng, Vlad Fienber, Markus Freitag, Xavier Garcia, Sebastian Gehrmann, Lucas Gonzalez, Guy Gur-Ari, Steven Hand, Hadi Hashemi, Le Hou, Joshua Howland, Andrea Hu, Jeffrey Hui, Jeremy Hurwitz, Michael Isard, Abe Ittycheriah, Matthew Jagielski, Wenhao Jia, Kathleen Kenealy, Maxim Krikun, Sneha Kudugunta, Chang Lan, Katherine Lee, Benjamin Lee, Eric Li, Music Li, Wei Li, YaGuang Li, Jian Li, Hyeontaek Lim, Hanzhao Lin, Zhongtao Liu, Frederick Liu, Marcello Maggioni, Aroma Mahendru, Joshua Maynez, Vedant Misra, Maysam Moussalem, Zachary Nado, John Nham, Eric Ni, Andrew Nystrom, Alicia Parrish, Marie Pellat, Martin Polacek, Alex Polozov, Reiner Pope, Siyuan Qiao, Emily Reif, Bryan Richter, Parker Riley, Alex Castro Ros, Aurko Roy, Brennan Saeta, Rajkumar Samuel, Renee Shelby, Ambrose Slone, Daniel Smilkov, David R. So, Daniel Sohn, Simon Tokumine, Dasha Valter, Vijay Vasudevan, Kiran Vodrahalli, Xuezhi Wang, Pidong Wang, Zirui Wang, Tao Wang, John Wieting, Yuhuai Wu, Kelvin Xu, Yunhan Xu, Linting Xue, Pengcheng Yin, Jiahui Yu, Qiao Zhang, Steven Zheng, Ce Zheng, Weikang Zhou, Denny Zhou, Slav Petrov, and Yonghui Wu. 2023. Palm 2 technical report.
  3. 3.Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, Mir Rosenberg, Xia Song, Alina Stoica, Saurabh Tiwary, and Tong Wang. 2018. Ms marco: A human generated machine reading comprehension dataset. (arXiv:1611.09268). ArXiv:1611.09268 [cs].
  4. 4.Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool. 2014. Food-101 – mining discriminative components with random forests. In European Conference on Computer Vision.
  5. 5.Emanuele Bugliarello, Fangyu Liu, Jonas Pfeiffer, Siva Reddy, Desmond Elliott, Edoardo Maria Ponti, and Ivan Vulić. 2022. Iglue: A benchmark for transfer learning across modalities, tasks, and languages. In Proceedings of the 39th International Conference on Machine Learning, page 2370–2392. PMLR.
  6. 6.Jannis Bulian, Christian Buck, Wojciech Gajewski, Benjamin Börschinger, and Tal Schuster. 2022. Tomayto, tomahto. beyond token-level answer equivalence for question answering evaluation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 291–305, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  7. 7.Andrea Burns, Krishna Srinivasan, Joshua Ainslie, Geoff Brown, Bryan A. Plummer, Kate Saenko, Jianmo Ni, and Mandy Guo. 2023. Wikiweb2m: A page-level multimodal wikipedia dataset.
  8. 8.Danqi Chen, Adam Fisch, Jason Weston, and Antoine Bordes. 2017. Reading Wikipedia to answer open-domain questions. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1870–1879, Vancouver, Canada. Association for Computational Linguistics.
  9. 9.Xi Chen, Josip Djolonga, Piotr Padlewski, Basil Mustafa, Soravit Changpinyo, Jialin Wu, Carlos Riquelme Ruiz, Sebastian Goodman, Xiao Wang, Yi Tay, et al. 2023a. Pali-x: On scaling up a multilingual vision and language model. arXiv preprint arXiv:2305.18565.
  10. 10.Xi Chen, Xiao Wang, Soravit Changpinyo, A. J. Piergiovanni, Piotr Padlewski, Daniel Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, Alexander Kolesnikov, Joan Puigcerver, Nan Ding, Keran Rong, Hassan Akbari, Gaurav Mishra, Linting Xue, Ashish Thapliyal, James Bradbury, Weicheng Kuo, Mojtaba Seyedhosseini, Chao Jia, Burcu Karagol Ayan, Carlos Riquelme, Andreas Steiner, Anelia Angelova, Xiaohua Zhai, Neil Houlsby, and Radu Soricut. 2022. Pali: A jointly-scaled multilingual language-image model. (arXiv:2209.06794). ArXiv:2209.06794 [cs].
  11. 11.Yang Chen, Hexiang Hu, Yi Luan, Haitian Sun, Soravit Changpinyo, Alan Ritter, and Ming-Wei Chang. 2023b. Can pre-trained vision and language models answer visual information-seeking questions? In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 14948–14968, Singapore. Association for Computational Linguistics.
  12. 12.Yang Chen, Hexiang Hu, Yi Luan, Haitian Sun, Soravit Changpinyo, Alan Ritter, and Ming-Wei Chang. 2023c. Can pre-trained vision and language models answer visual information-seeking questions? (arXiv:2302.11713). ArXiv:2302.11713 [cs].
  13. 13.Zhuo Chen, Yufeng Huang, Jiaoyan Chen, Yuxia Geng, Yin Fang, Jeff Z. Pan, Ningyu Zhang, and Wen Zhang. 2023d. Lako: Knowledge-driven visual question answering via late knowledge-to-text injection. In Proceedings of the 11th International Joint Conference on Knowledge Graphs, IJCKG ’22, page 20–29, New York, NY, USA. Association for Computing Machinery.
  14. 14.Mehdi Cherti, Romain Beaumont, Ross Wightman, Mitchell Wortsman, Gabriel Ilharco, Cade Gordon, Christoph Schuhmann, Ludwig Schmidt, and Jenia Jitsev. 2023. Reproducible scaling laws for contrastive language-image learning. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 2818–2829.
  15. 15.Yin Cui, Yang Song, Chen Sun, Andrew Howard, and Serge Belongie. 2018. Large scale fine-grained categorization and domain-specific transfer learning. (arXiv:1806.06193). ArXiv:1806.06193 [cs].
  16. 16.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics. ACL.
  17. 17.Danny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, Wenlong Huang, Yevgen Chebotar, Pierre Sermanet, Daniel Duckworth, Sergey Levine, Vincent Vanhoucke, Karol Hausman, Marc Toussaint, Klaus Greff, Andy Zeng, Igor Mordatch, and Pete Florence. 2023. Palm-e: An embodied multimodal language model. In Proceedings of the 40th International Conference on Machine Learning, ICML’23. JMLR.org.
  18. 18.Julian Eisenschlos, Syrine Krichene, and Thomas Müller. 2020. Understanding tables with intermediate pre-training. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 281–296, Online. Association for Computational Linguistics.
  19. 19.Feng Gao, Qing Ping, Govind Thattai, Aishwarya Reganti, Ying Nian Wu, and Prem Natarajan. 2022. Transform-retrieve-generate: Natural language-centric outside-knowledge visual question answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5067–5077.
  20. 20.François Garderes, Maryam Ziaeefard, Baptiste Abeloos, and Freddy Lecue. 2020. Conceptbert: Concept-aware representation for visual question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: Findings, pages 489–498.
  21. 21.Google. Google lens: Image recognition and retrieval api. https://lens.google.com.
  22. 22.Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2017. Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering. In Conference on Computer Vision and Pattern Recognition (CVPR).
  23. 23.Liangke Gui, Borui Wang, Qiuyuan Huang, Alex Hauptmann, Yonatan Bisk, and Jianfeng Gao. 2021. Kat: A knowledge augmented transformer for vision-and-language. arXiv preprint arXiv:2112.08614.
  24. 24.Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. 2020. Realm: Retrieval-augmented language model pre-training. In Proceedings of the 37th International Conference on Machine Learning, ICML’20. JMLR.org.
  25. 25.Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-rank adaptation of large language models. In International Conference on Learning Representations.
  26. 26.Hexiang Hu, Yi Luan, Yang Chen, Urvashi Khandelwal, Mandar Joshi, Kenton Lee, Kristina Toutanova, and Ming-Wei Chang. 2023a. Open-domain visual entity recognition: Towards recognizing millions of wikipedia entities. (arXiv:2302.11154). ArXiv:2302.11154 [cs].
  27. 27.Ziniu Hu, Ahmet Iscen, Chen Sun, Kai-Wei Chang, Yizhou Sun, David Ross, Cordelia Schmid, and Alireza Fathi. 2024. Avis: Autonomous visual information seeking with large language model agent. Advances in Neural Information Processing Systems, 36.
  28. 28.Ziniu Hu, Ahmet Iscen, Chen Sun, Zirui Wang, Kai-Wei Chang, Yizhou Sun, Cordelia Schmid, David A. Ross, and Alireza Fathi. 2023b. Reveal: Retrieval-augmented visual-language pre-training with multi-source multimodal knowledge memory. page 23369–23379.
  29. 29.Gautier Izacard and Edouard Grave. 2021. Leveraging passage retrieval with generative models for open domain question answering. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 874–880, Online. Association for Computational Linguistics.
  30. 30.Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2019. Billion-scale similarity search with GPUs. IEEE Transactions on Big Data, 7(3):535–547.
  31. 31.Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361.
  32. 32.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6769–6781, Online. Association for Computational Linguistics.
  33. 33.Omar Khattab and Matei Zaharia. 2020. Colbert: Efficient and effective passage search via contextualized late interaction over bert. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’20, page 39–48, New York, NY, USA. Association for Computing Machinery.
  34. 34.Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings.
  35. 35.Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei. 2013. 3d object representations for fine-grained categorization. In 4th International IEEE Workshop on 3D Representation and Recognition (3dRR-13), Sydney, Australia.
  36. 36.Kenton Lee, Ming-Wei Chang, and Kristina Toutanova. 2019. Latent retrieval for weakly supervised open domain question answering. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 6086–6096, Florence, Italy. Association for Computational Linguistics.
  37. 37.Paul Lerner, Olivier Ferret, and Camille Guinaudeau. 2023. Multimodal inverse cloze task for knowledge-based visual question answering. In European Conference on Information Retrieval, pages 569–587. Springer.
  38. 38.Paul Lerner, Olivier Ferret, and Camille Guinaudeau. 2024. Cross-modal retrieval for knowledge-based visual question answering. In European Conference on Information Retrieval, pages 421–438. Springer.
  39. 39.Paul Lerner, Olivier Ferret, Camille Guinaudeau, Hervé Le Borgne, Romaric Besançon, José G Moreno, and Jesús Lovón Melgarejo. 2022. Viquae, a dataset for knowledge-based visual question answering about named entities. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 3108–3120.
  40. 40.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing Systems, 33:9459–9474.
  41. 41.Guohao Li, Xin Wang, and Wenwu Zhu. 2020. Boosting visual question answering with context-aware knowledge aggregation. In Proceedings of the 28th ACM International Conference on Multimedia, pages 1227–1235.
  42. 42.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. arXiv preprint arXiv:2301.12597.
  43. 43.Leroy Lin, Yujia Xie, Dongdong Chen, Yichong Xu, Chenguang Zhu, and Lu Yuan. 2022. REVIVE: Regional visual representation matters in knowledge-based visual question answering. In Advances in Neural Information Processing Systems.
  44. 44.Weizhe Lin, Rexhina Blloshmi, Bill Byrne, Adria de Gispert, and Gonzalo Iglesias. 2023a. LI-RAGE: Late interaction retrieval augmented generation with explicit signals for open-domain table question answering. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 1557–1566, Toronto, Canada. Association for Computational Linguistics.
  45. 45.Weizhe Lin and Bill Byrne. 2022. Retrieval augmented visual question answering with outside knowledge. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 11238–11254, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  46. 46.Weizhe Lin, Jinghong Chen, Jingbiao Mei, Alexandru Coca, and Bill Byrne. 2023b. Fine-grained late-interaction multi-modal retrieval for retrieval augmented visual question answering. In Thirty-seventh Conference on Neural Information Processing Systems.
  47. 47.Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2023a. Improved baselines with visual instruction tuning. arXiv preprint arXiv:2310.03744.
  48. 48.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023b. Visual instruction tuning. (arXiv:2304.08485). ArXiv:2304.08485 [cs].
  49. 49.Man Luo, Yankai Zeng, Pratyay Banerjee, and Chitta Baral. 2021. Weakly-supervised visual-retriever-reader for knowledge-based question answering. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6417–6431, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  50. 50.Kenneth Marino, Xinlei Chen, Devi Parikh, Abhinav Gupta, and Marcus Rohrbach. 2021. Krisp: Integrating implicit and symbolic knowledge for open-domain knowledge-based vqa. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14111–14121.
  51. 51.Thomas Mensink, Jasper Uijlings, Lluis Castrejon, Arushi Goel, Felipe Cadar, Howard Zhou, Fei Sha, André Araujo, and Vittorio Ferrari. 2023a. Encyclopedic vqa: Visual questions about detailed properties of fine-grained categories. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 3113–3124.
  52. 52.Thomas Mensink, Jasper Uijlings, Lluis Castrejon, Arushi Goel, Felipe Cadar, Howard Zhou, Fei Sha, André Araujo, and Vittorio Ferrari. 2023b. Encyclopedic vqa: Visual questions about detailed properties of fine-grained categories. (arXiv:2306.09224). ArXiv:2306.09224 [cs].
  53. 53.Medhini Narasimhan, Svetlana Lazebnik, and Alexander Schwing. 2018. Out of the box: Reasoning with graph convolution nets for factual visual question answering. Advances in neural information processing systems, 31.
  54. 54.Jianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai, Gustavo Hernandez Abrego, Ji Ma, Vincent Zhao, Yi Luan, Keith Hall, Ming-Wei Chang, and Yinfei Yang. 2022. Large dual encoders are generalizable retrievers. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 9844–9855, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  55. 55.OpenAI. 2023. Gpt-4 technical report.
  56. 56.Chen Qu, Hamed Zamani, Liu Yang, W Bruce Croft, and Erik Learned-Miller. 2021. Passage retrieval for outside-knowledge visual question answering. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 1753–1757.
  57. 57.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning transferable visual models from natural language supervision. arXiv.
  58. 58.Jiahua Rao, Zifei Shan, Longpo Liu, Yao Zhou, and Yuedong Yang. 2023. Retrieval-based knowledge augmented vision language pre-training. In Proceedings of the 31st ACM International Conference on Multimedia, MM ’23, page 5399–5409, New York, NY, USA. Association for Computing Machinery.
  59. 59.Keshav Santhanam, Omar Khattab, Christopher Potts, and Matei Zaharia. 2022a. Plaid: An efficient engine for late interaction retrieval. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management, CIKM ’22, page 1747–1756, New York, NY, USA. Association for Computing Machinery.
  60. 60.Keshav Santhanam, Omar Khattab, Jon Saad-Falcon, Christopher Potts, and Matei Zaharia. 2022b. ColBERTv2: Effective and efficient retrieval via lightweight late interaction. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3715–3734, Seattle, United States. Association for Computational Linguistics.
  61. 61.Dustin Schwenk, Apoorv Khandelwal, Christopher Clark, Kenneth Marino, and Roozbeh Mottaghi. 2022. A-okvqa: A benchmark for visual question answering using world knowledge. arXiv preprint arXiv:2206.01718.
  62. 62.Sanket Shah, Anand Mishra, Naganand Yadati, and Partha Pratim Talukdar. 2019. Kvqa: Knowledge-aware visual question answering. Proceedings of the AAAI Conference on Artificial Intelligence, 33(01):8876–8884.
  63. 63.Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), page 2556–2565, Melbourne, Australia. Association for Computational Linguistics.
  64. 64.Krishna Srinivasan, Karthik Raman, Jiecao Chen, Michael Bendersky, and Marc Najork. 2021. Wit: Wikipedia-based image text dataset for multimodal multilingual machine learning. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’21, page 2443–2449, New York, NY, USA. Association for Computing Machinery.
  65. 65.Cong Wei, Yang Chen, Haonan Chen, Hexiang Hu, Ge Zhang, Jie Fu, Alan Ritter, and Wenhu Chen. 2023. Uniir: Training and benchmarking universal multimodal information retrievers. arXiv preprint arXiv:2311.17136.
  66. 66.Tobias Weyand, Andre Araujo, Bingyi Cao, and Jack Sim. 2020. Google landmarks dataset v2 – a large-scale benchmark for instance-level recognition and retrieval. (arXiv:2004.01804). ArXiv:2004.01804 [cs].
  67. 67.Jialin Wu, Jiasen Lu, Ashish Sabharwal, and Roozbeh Mottaghi. 2022. Multi-modal answer validation for knowledge-based vqa. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pages 2712–2721.
  68. 68.Jialin Wu and Raymond Mooney. 2022. Entity-focused dense passage retrieval for outside-knowledge visual question answering. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 8061–8072, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  69. 69.Yilun Xu, Shengjia Zhao, Jiaming Song, Russell Stewart, and Stefano Ermon. 2020. A theory of usable information under computational constraints. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  70. 70.Lewei Yao, Runhui Huang, Lu Hou, Guansong Lu, Minzhe Niu, Hang Xu, Xiaodan Liang, Zhenguo Li, Xin Jiang, and Chunjing Xu. 2022. FILIP: Fine-grained interactive language-image pre-training. In International Conference on Learning Representations.
  71. 71.Da Yin, Feng Gao, Govind Thattai, Michael Johnston, and Kai-Wei Chang. 2023. Givl: Improving geographical inclusivity of vision-language models with pre-training methods. page 10951–10961.
  72. 72.Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2023. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592.

Citation

MLA
Lin, W., et al. “PreFLMR: Scaling Up Fine-Grained Late-Interaction Multi-modal Retrievers”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 5294–316, https://doi.org/10.18653/v1/2024.acl-long.289.
APA
Lin, W., Mei, J., Chen, J., & Byrne, B. (2024). PreFLMR: Scaling Up Fine-Grained Late-Interaction Multi-modal Retrievers. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 5294–5316. https://doi.org/10.18653/v1/2024.acl-long.289
Chicago
Lin, W., J. Mei, J. Chen, and B. Byrne. 2024. “PreFLMR: Scaling Up Fine-Grained Late-Interaction Multi-modal Retrievers”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 5294–5316. https://doi.org/10.18653/v1/2024.acl-long.289.
Harvard
Lin, W. et al. (2024) “PreFLMR: Scaling Up Fine-Grained Late-Interaction Multi-modal Retrievers”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 5294–5316. Available at: https://doi.org/10.18653/v1/2024.acl-long.289.
Vancouver
1. Lin W, Mei J, Chen J, Byrne B (2024) PreFLMR: Scaling Up Fine-Grained Late-Interaction Multi-modal Retrievers. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 5294–5316

BibTeX

@inproceedings{lin-etal-2024-preflmr,
    title = "{P}re{FLMR}: Scaling Up Fine-Grained Late-Interaction Multi-modal Retrievers",
    author = "Lin, Weizhe  and
      Mei, Jingbiao  and
      Chen, Jinghong  and
      Byrne, Bill",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.289/",
    doi = "10.18653/v1/2024.acl-long.289",
    pages = "5294--5316"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/