ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction

Keshav SanthanamOmar KhattabJon Saad-FalconChristopher PottsMatei Zaharia

article2022NAACL860 citations

Presents ColBERTv2, a neural retrieval model that combines denoised supervision with residual vector compression to cut indexing storage by 6–10× while outperforming existing dense retrievers across diverse in-domain and zero-shot benchmarks.

Listen

Modern neural search systems power critical language applications like open-domain question answering and internal enterprise search. Most state-of-the-art retrievers encode queries and documents into single vectors, but these models can struggle with complex semantic relationships. While late-interaction systems—which generate multi-vector representations at the token level—offer superior accuracy, they require an order-of-magnitude larger storage footprint to maintain billions of vectors for large-scale document collections.

The article introduces and evaluates ColBERTv2, a neural information retrieval system designed to eliminate this storage penalty while establishing state-of-the-art search quality. It aims to demonstrate that pairing late-interaction token scoring with an aggressive residual compression mechanism and denoised cross-encoder supervision delivers high accuracy and dramatic storage reductions across both familiar and novel domains.

The researchers evaluated the system through extensive empirical benchmarks. They trained the model on the standard MS MARCO collection using knowledge distillation from a lightweight cross-encoder teacher combined with hard-negative mining. To reduce storage, they implemented a residual compression method that clusters token vectors around centroids and quantizes the differences. Retrieval performance and out-of-domain transfer were evaluated across 28 diverse test suites, including standard benchmarks, Wikipedia question answering, and LoTTE, a newly introduced benchmark designed to evaluate natural, long-tail search queries across specialized online communities.

The primary finding is that ColBERTv2 establishes state-of-the-art retrieval quality while reducing index storage by a factor of 6 to 10 relative to the original late-interaction baseline. On the standard MS MARCO collection, the model compressed the index from 154 GB down to 16–25 GB, matching the storage footprint of single-vector models while achieving a leading 39.7% MRR@10. Furthermore, the model demonstrated superior generalization, achieving the highest quality score on 22 out of 28 out-of-domain test sets and outperforming competing models by up to 8% relative gain. The evaluation also showed that the residual compression preserved accuracy with minimal degradation while maintaining practical query latencies between 50 and 250 milliseconds.

These results demonstrate that organizations do not need to choose between search accuracy and infrastructure expense. By resolving the storage bottleneck of late-interaction architectures, high-precision neural retrieval becomes commercially viable for web-scale datasets and specialized business domains where labeled training data is unavailable. This reduces deployment costs and broadens the feasibility of grounding large language models in external knowledge bases without expensive retraining.

Organizations deploying search and retrieval infrastructure should evaluate late-interaction systems like ColBERTv2 as drop-in upgrades over standard single-vector architectures, especially in settings requiring high out-of-domain generalization. Before rolling out at extreme scale, engineering teams should pilot the pipeline to balance latency and compression trade-offs by selecting appropriate centroid probing parameters and bit encodings.

Confidence in these findings is supported by consistent performance across dozens of distinct benchmarks. However, decision-makers should note that evaluations were conducted exclusively on English-language corpora, and training complex multi-vector models with cross-encoder distillation requires substantial initial compute overhead compared to simpler lexical search methods.

Cover for ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction

Abstract

Neural information retrieval (IR) has greatly advanced search and other knowledge-intensive language tasks. While many neural IR methods encode queries and documents into single-vector representations, late interaction models produce multi-vector representations at the granularity of each token and decompose relevance modeling into scalable token-level computations. This decomposition has been shown to make late interaction more effective, but it inflates the space footprint of these models by an order of magnitude. In this work, we introduce ColBERTv2, a retriever that couples an aggressive residual compression mechanism with a denoised supervision strategy to simultaneously improve the quality and space footprint of late interaction. We evaluate ColBERTv2 across a wide range of benchmarks, establishing state-of-the-art quality within and outside the training domain while reducing the space footprint of late interaction models by 6–10×.

Table of Contents

  • 1 Introduction
  • 2 Background & Related Work
  • 2.1 Token-Decomposed Scoring in Neural IR
  • 2.2 Vector Compression for Neural IR
  • 2.3 Improving the Quality of Single-Vector Representations
  • 2.4 Out-of-Domain Evaluation in IR
  • 3 ColBERTv2
  • 3.1 Modeling
  • 3.2 Supervision
  • 3.3 Representation
  • 3.4 Indexing
  • 3.5 Retrieval
  • 4 LoTTE: Long-Tail, Cross-Domain Retrieval Evaluation
  • 5 Evaluation
  • 5.1 In-Domain Retrieval Quality
  • 5.2 Out-of-Domain Retrieval Quality
  • 5.3 Efficiency
  • 6 Conclusion
  • Acknowledgements
  • Broader Impact & Ethical Considerations
  • Research Limitations
  • References
  • A Analysis of ColBERT's Semantic Space
  • B Impact of Compression
  • C Retrieval Latency
  • D LoTTE
  • E Datasets in BEIR
  • F Implementation & Hyperparameters

Knowls

  1. Knowl 1 — Late Interaction Relevance Scoring via MaxSim

    equation

    ColBERTv2 computes the relevance score Sq,d∈RS_{q,d} \in \mathbb{R} between a query qq and a document or passage dd using a fine-grained late interaction mechanism over token-level contextual representations:

    Sq,d=∑i=1Nmax⁡j=1MQi⋅DjTS_{q,d} = \sum_{i=1}^{N} \max_{j=1}^{M} Q_i \cdot D_j^T

    where Q∈RN×ddimQ \in \mathbb{R}^{N \times d_{\text{dim}}} is the multi-vector representation of the query containing NN contextual token embeddings Qi∈RddimQ_i \in \mathbb{R}^{d_{\text{dim}}}, and D∈RM×ddimD \in \mathbb{R}^{M \times d_{\text{dim}}} is the multi-vector representation of the passage containing MM contextual token embeddings Dj∈RddimD_j \in \mathbb{R}^{d_{\text{dim}}}. The inner product Qi⋅DjTQ_i \cdot D_j^T evaluates cosine similarity between normalized vector representations. The "MaxSim" operator identifies the most contextually relevant token embedding in the passage for each query token embedding QiQ_i, and the total passage relevance score is obtained by summing these maximal similarity scores across all NN query tokens.

  2. Knowl 2 — Centroid-Based Residual Compression for Multi-Vector Embeddings

    model/method

    To reduce the space footprint of storing token embeddings (ddim=128d_{\text{dim}} = 128) without requiring re-training or altering model architecture, ColBERTv2 compresses each token vector v∈Rddimv \in \mathbb{R}^{d_{\text{dim}}} using a codebook of centroids C={C1,C2,…,C∣C∣}\mathcal{C} = \{C_1, C_2, \dots, C_{|\mathcal{C}|}\}.

    1. Centroid Assignment: The vector vv is assigned to its closest centroid Ct∈CC_t \in \mathcal{C} by Euclidean distance: t=arg⁡min⁡k∥v−Ck∥2t = \arg\min_{k} \|v - C_k\|_2.
    2. Residual Quantization: The residual error vector r=v−Ctr = v - C_t is computed, and each of its ddimd_{\text{dim}} dimensions is quantized into b∈{1,2}b \in \{1, 2\} bits, yielding an approximation r~\tilde{r}.
    3. Decompression: At query time, the original vector is reconstructed as v~=Ct+r~\tilde{v} = C_t + \tilde{r}.

    The total storage required per vector is ⌈log⁡2∣C∣⌉+b⋅ddim\lceil \log_2 |\mathcal{C}| \rceil + b \cdot d_{\text{dim}} bits. For ∣C∣≤232|\mathcal{C}| \le 2^{32} (4 bytes for the centroid index) and ddim=128d_{\text{dim}} = 128, each token vector requires 20 bytes when b=1b=1 bit or 36 bytes when b=2b=2 bits. This reduces per-vector storage by 7.1×7.1\times to 12.8×12.8\times relative to standard 16-bit floating-point representations (256 bytes per vector).

  3. Knowl 3 — Denoised Cross-Encoder Distillation with Negative Refresh

    model/method

    ColBERTv2 trains multi-vector retrieval encoders using a denoised distillation pipeline combined with dynamic negative sampling:

    1. Candidate Retrieval: An initial ColBERT model retrieves the top-kk (k=500k=500) candidate passages for each training query in the MS MARCO dataset.
    2. Teacher Scoring: A 22-million-parameter MiniLM cross-encoder reranker scores each query–passage pair.
    3. Tuple Construction: For each query, a ww-way tuple (w=64w=64) is formed consisting of the query, a single positive passage (either a human-labeled positive or the teacher's top-ranked passage), and 63 lower-ranked candidate passages.
    4. Distillation Loss: A Kullback–Leibler (KL) divergence loss transfers the soft score distribution from the cross-encoder teacher to the late-interaction student model, accommodating the differences in output score ranges between cross-encoders and late-interaction cosine sums.
    5. In-Batch Negatives: A cross-entropy loss is applied over in-batch negatives across GPU devices during training.
    6. Index and Negative Refresh: After an initial training cycle, the index is re-encoded and refreshed using the updated checkpoint to retrieve a renewed set of hard negatives for a second round of distillation.
  4. Knowl 4 — Three-Stage Offline Indexing Algorithm for ColBERTv2

    algorithm

    ColBERTv2 indexes a document collection in three offline stages to construct a compressed, searchable inverted index without keeping all raw uncompressed passage embeddings in memory.

    Input: Passage collection D\mathcal{D}, BERT encoder EE, dimension d=128d = 128, quantization bit-width b∈{1,2}b \in \{1, 2\}
    Output: Centroid codebook C\mathcal{C}, compressed passage representations {D~d}d∈D\{\tilde{D}_d\}_{d \in \mathcal{D}}, inverted lists I\mathcal{I}
    1. Sample subset Dsample⊂D\mathcal{D}_{\text{sample}} \subset \mathcal{D} proportional to ∣D∣\sqrt{|\mathcal{D}|}
    2. Compute token embeddings for Dsample\mathcal{D}_{\text{sample}} using EE
    3. Cluster sampled embeddings via kk-means to create codebook C\mathcal{C} with ∣C∣≈16Nembeddings|\mathcal{C}| \approx 16 \sqrt{N_{\text{embeddings}}} centroids
    4. Initialize inverted index mapping I[t]←∅\mathcal{I}[t] \leftarrow \emptyset for each centroid index t∈{1,…,∣C∣}t \in \{1, \dots, |\mathcal{C}|\}
    5. for each passage d∈Dd \in \mathcal{D} do
    6. Compute contextual embeddings D=E(d)∈RM×dD = E(d) \in \mathbb{R}^{M \times d}
    7. for each embedding vector v∈Dv \in D do
    8. Find centroid index t=arg⁡min⁡k∥v−Ck∥2t = \arg\min_{k} \|v - C_k\|_2
    9. Compute residual r=v−Ctr = v - C_t
    10. Quantize rr to bb bits per dimension, obtaining r~\tilde{r}
    11. Record compressed tuple (t,r~)(t, \tilde{r}) in passage entry D~d\tilde{D}_d
    12. Append vector identifier of vv to I[t]\mathcal{I}[t]
    13. Write codebook C\mathcal{C}, compressed representations {D~d}\{\tilde{D}_d\}, and inverted lists I\mathcal{I} to disk
  5. Knowl 5 — Two-Stage Candidate Generation and Scoring Algorithm

    algorithm

    At query time, ColBERTv2 performs retrieval by first deriving approximate lower bounds on token MaxSim scores to filter candidates, followed by exact late interaction scoring over candidate passages.

    Input: Query qq, encoder EE, codebook C\mathcal{C}, compressed index {D~d}\{\tilde{D}_d\}, inverted index I\mathcal{I}, parameters nproben_{\text{probe}}, ncandidaten_{\text{candidate}}
    Output: Top ranked passages for query qq
    1. Compute query embeddings Q=E(q)=[Q1,…,QN]∈RN×dQ = E(q) = [Q_1, \dots, Q_N] \in \mathbb{R}^{N \times d}
    2. Initialize candidate accumulator score array Scoreapprox[d]←0\text{Score}_{\text{approx}}[d] \leftarrow 0 for all d∈Dd \in \mathcal{D}
    3. for each query token vector Qi∈QQ_i \in Q do
    4. Find the nproben_{\text{probe}} nearest centroids in C\mathcal{C} to QiQ_i
    5. Retrieve passage token IDs indexed under these centroids via I\mathcal{I}
    6. for each retrieved token representation with centroid CtC_t and residual r~\tilde{r} from passage dd do
    7. Decompress token vector v~=Ct+r~\tilde{v} = C_t + \tilde{r}
    8. Compute cosine similarity s=Qi⋅v~Ts = Q_i \cdot \tilde{v}^T
    9. Update per-passage maximum: MaxSimToken[d]←max⁡(MaxSimToken[d],s)\text{MaxSimToken}[d] \leftarrow \max(\text{MaxSimToken}[d], s)
    10. for each touched passage dd do
    11. Scoreapprox[d]←Scoreapprox[d]+MaxSimToken[d]\text{Score}_{\text{approx}}[d] \leftarrow \text{Score}_{\text{approx}}[d] + \text{MaxSimToken}[d]
    12. Select the top ncandidaten_{\text{candidate}} passages according to Scoreapprox\text{Score}_{\text{approx}}
    13. for each passage dd in selected candidates do
    14. Load full compressed representation D~d\tilde{D}_d and decompress all passage token vectors
    15. Compute exact late-interaction score Sq,d=∑i=1Nmax⁡j=1MQi⋅D~d,jTS_{q,d} = \sum_{i=1}^N \max_{j=1}^M Q_i \cdot \tilde{D}_{d, j}^T
    16. Sort candidate passages by exact score Sq,dS_{q,d} and return top results
  6. Knowl 6 — LoTTE Benchmark for Long-Tail Zero-Shot Retrieval

    experimental setup

    The Long-Tail Topic-stratified Evaluation (LoTTE) benchmark assesses zero-shot out-of-domain neural retrieval on long-tail, non-Wikipedia topics across 5 domains from StackExchange: Writing, Recreation, Science, Technology, and Lifestyle. The passage corpora consist of StackExchange answer posts (with HTML removed and positive community scores).

    LoTTE divides queries into two distinct types per domain:

    1. Search Queries: Drawn from GooAQ, representing natural, concise Google search autocomplete queries whose answers were linked to StackExchange posts.
    2. Forum Queries: Organic StackExchange question titles reflecting complex, open-ended information-seeking intents.

    Passages and queries between development and test splits are strictly disjoint. LoTTE provides 12 test sets (500–2,000 queries and 100k–2M passages per domain) as well as a Pooled test setting aggregating all domains (2.8M passages, 3,869 search queries, 10,025 forum queries). Retrieval performance is evaluated using Success@5 (S@5S@5), which awards 1 point if an accepted or upvoted (score ≥1\ge 1) answer from the target thread is ranked within the top 5 retrieved results.

  7. Knowl 7 — In-Domain Passage Retrieval Performance on MS MARCO

    data/table

    When trained on MS MARCO Passage Ranking, ColBERTv2 outperforms existing sparse, dense single-vector, and late-interaction retrieval architectures on both the MS MARCO official development set (6,980 queries) and the 5,000-query Local Eval test set.

    Method Official Dev (7k) Local Eval (5k)
    MRR@10 R@50 R@1k MRR@10 R@50 R@1k
    Models without Distillation
    RepBERT 30.4 - 94.3 - - -
    DPR 31.1 - 95.2 - - -
    ANCE 33.0 - 95.9 - - -
    LTRe 34.1 - 96.2 - - -
    ColBERT (vanilla) 36.0 82.9 96.8 36.7 - -
    Models with Distillation / Pretraining
    TAS-B 34.7 - 97.8 - - -
    SPLADEv2 36.8 - 97.9 37.9 84.9 98.0
    PAIR 37.9 86.4 98.2 - - -
    coCondenser 38.2 - 98.4 - - -
    RocketQAv2 38.8 86.2 98.1 39.8 85.8 97.9
    ColBERTv2 39.7 86.8 98.4 40.8 86.3 98.3

    ColBERTv2 achieves 39.7% MRR@10 on the official dev set (a 3.7-point improvement over vanilla ColBERT and 0.9 points over RocketQAv2) and 40.8% MRR@10 on Local Eval using 2-bit residual compression.

  8. Knowl 8 — Zero-Shot Out-of-Domain Retrieval Performance across BEIR, LoTTE, and Open-QA

    data/table

    ColBERTv2 achieves the highest retrieval performance on 22 of 28 zero-shot out-of-domain test sets across the BEIR benchmark (nDCG@10), Wikipedia Open-QA tests (Success@5), and the LoTTE benchmark (Success@5).

    Benchmark / Dataset BM25 ANCE ColBERT RocketQAv2 SPLADEv2 ColBERTv2
    BEIR Search Tasks (nDCG@10)
    DBPedia 31.3 28.1 39.2 35.6 43.5 44.6
    FiQA 23.6 29.5 31.7 30.2 33.6 35.6
    NQ 30.6 44.6 52.4 50.5 52.1 56.2
    HotpotQA 60.3 45.6 59.3 53.3 68.4 66.7
    NFCorpus 32.2 23.7 30.5 29.3 33.4 33.8
    TREC-COVID 65.6 65.4 67.7 67.5 71.0 73.8
    Touché (v2) 36.7 - - 24.7 27.2 26.3
    OOD Wikipedia Open QA (Success@5)
    NQ-dev 44.6 - 65.7 - 65.6 68.9
    TQ-dev 67.6 - 72.6 - 74.7 76.7
    SQuAD-dev 50.6 - 60.0 - 60.4 65.0
    LoTTE Test Pooled (Success@5)
    Search Queries 48.3 66.4 67.3 69.8 68.9 71.6
    Forum Queries 47.2 55.7 58.2 57.7 60.1 63.4

    ColBERTv2 consistently leads on natural search queries across domains, outperforming SPLADEv2 on Open-QA by up to 4.6 points (SQuAD) and on LoTTE Pooled by 2.7 points (Search) and 3.3 points (Forum).

  9. Knowl 9 — Index Storage Footprint and Compression Efficiency

    empirical result

    On the MS MARCO passage collection (8.8M passages containing ~600M token embeddings with embedding dimension d=128d=128):

    1. Vanilla ColBERT: Requires 154 GiB of storage when storing 16-bit floating-point embeddings (256 bytes per token embedding).
    2. ColBERTv2 (2-bit Residuals): Requires 25 GiB total storage (20.5 GiB for compressed embeddings and centroids, plus 4.5 GiB for inverted lists), achieving a 6.2×6.2\times space reduction.
    3. ColBERTv2 (1-bit Residuals): Requires 16 GiB total storage (11.5 GiB for compressed embeddings and centroids, plus 4.5 GiB for inverted lists), achieving a 9.6×9.6\times space reduction.

    Applying 2-bit compression to vanilla ColBERT preserves ranking quality without loss: 36.2% MRR@10 and 82.3% Recall@50 (compared to 36.2% MRR@10 and 82.1% Recall@50 uncompressed). Applying 1-bit compression yields 35.5% MRR@10 and 81.6% Recall@50. Retrieval latency across collection scales ranges between 50 and 250 milliseconds per query.

  10. Knowl 10 — Semantic Clustering Properties of Late-Interaction Embeddings

    empirical result

    Clustering the ~600M token embeddings produced by ColBERT on the MS MARCO collection (covering ~27,000 unique vocabulary tokens) into k=218k = 2^{18} centroids demonstrates that late-interaction token vectors naturally occupy localized, sense-specific semantic regions:

    1. Cluster Purity: Approximately 90% of centroid clusters contain 16 or fewer distinct non-stopword tokens in ColBERT's embedding space, whereas fewer than 50% of clusters contain 16 or fewer distinct tokens when clustering randomly generated control vectors with identical token marginals.
    2. Sense Specialization: Vocabulary tokens distribute across a very small number of contextual sense clusters (e.g., photography-related tokens 'photo', 'photos', and 'pictures' consistently co-occur across specific overlapping clusters, as do meteorological tokens like 'tornado', 'storm', and 'hurricane').

    This structural clustering allows cluster centroids to summarize the semantic space, making low-bit residual quantization effective with minimal distortion.

Coverage note — None was omitted; all key contributions—late-interaction formulation, residual compression mechanism, denoised supervision and negative refresh, indexing and candidate retrieval algorithms, LoTTE benchmark, empirical results (in-domain, out-of-domain, Open-QA), compression efficiency analysis, and semantic space analysis—are represented.

References

  1. 1.Stack Exchange Data Dump.
  2. 2.Liefu Ai, Junqing Yu, Zebin Wu, Yunfeng He, and Tao Guan. 2017. Optimized Residual Vector Quantization for Efficient Approximate Nearest Neighbor Search. Multimedia Systems, 23(2):169–181.
  3. 3.Sören Auer, Christian Bizer, Georgi Kobilarov, Jens Lehmann, Richard Cyganiak, and Zachary Ives. 2007. DBpedia: A Nucleus for a Web of Open Data. In The semantic web, pages 722–735. Springer.
  4. 4.Christopher F Barnes, Syed A Rizvi, and Nasser M Nasrabadi. 1996. Advances in Residual Vector Quantization: A Review. IEEE transactions on image processing, 5(2):226–262.
  5. 5.Alexander Bondarenko, Maik Fröbe, Meriem Beloucif, Lukas Gienapp, Yamen Ajjour, Alexander Panchenko, Chris Biemann, Benno Stein, Henning Wachsmuth, Martin Potthast, et al. 2020. Overview of touché 2020: Argument Retrieval. In International Conference of the Cross-Language Evaluation Forum for European Languages, pages 384–395. Springer.
  6. 6.Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George van den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, et al. 2021. Improving language models by retrieving from trillions of tokens. arXiv preprint arXiv:2112.04426.
  7. 7.Vera Boteva, Demian Gholipour, Artem Sokolov, and Stefan Riezler. 2016. A Full-text Learning to Rank Dataset for Medical Information Retrieval. In European Conference on Information Retrieval, pages 716–722. Springer.
  8. 8.Chia-Yu Chen, Jungwook Choi, Daniel Brand, Ankur Agrawal, Wei Zhang, and Kailash Gopalakrishnan. 2018. Adacomp : Adaptive residual gradient compression for data-parallel distributed training. In Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, (AAAI-18), the 30th innovative Applications of Artificial Intelligence (IAAI-18), and the 8th AAAI Symposium on Educational Advances in Artificial Intelligence (EAAI-18), New Orleans, Louisiana, USA, February 2-7, 2018, pages 2827–2835. AAAI Press.
  9. 9.Arman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey, and Daniel Weld. 2020. SPECTER: Document-level representation learning using citation-informed transformers. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2270–2282, Online. Association for Computational Linguistics.
  10. 10.Nachshon Cohen, Amit Portnoy, Besnik Fetahu, and Amir Ingber. 2021. SDR: Efficient Neural Re-ranking using Succinct Document Representation. arXiv preprint arXiv:2110.02065.
  11. 11.Zhuyun Dai and Jamie Callan. 2020. Context-aware term weighting for first stage passage retrieval. In Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval, SIGIR 2020, Virtual Event, China, July 25-30, 2020, pages 1533–1536. ACM.
  12. 12.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  13. 13.Thomas Diggelmann, Jordan Boyd-Graber, Jannis Bulian, Massimiliano Ciaramita, and Markus Leippold. 2020. CLIMATE-FEVER: A Dataset for Verification of Real-World Climate Claims. arXiv preprint arXiv:2012.00614.
  14. 14.Thibault Formal, Carlos Lassance, Benjamin Piwowarski, and Stéphane Clinchant. 2021a. SPLADE v2: Sparse Lexical and Expansion Model for Information Retrieval. arXiv preprint arXiv:2109.10086.
  15. 15.Thibault Formal, Benjamin Piwowarski, and Stéphane Clinchant. 2021b. SPLADE: Sparse Lexical and Expansion Model for First Stage Ranking. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 2288–2292.
  16. 16.Luyu Gao and Jamie Callan. 2021. Unsupervised corpus aware language model pre-training for dense passage retrieval. arXiv preprint arXiv:2108.05540.
  17. 17.Luyu Gao, Zhuyun Dai, and Jamie Callan. 2020. Modularized transfomer-based ranking framework. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4180–4190, Online. Association for Computational Linguistics.
  18. 18.Luyu Gao, Zhuyun Dai, and Jamie Callan. 2021. COIL: Revisit exact lexical match in information retrieval with contextualized inverted list. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3030–3042, Online. Association for Computational Linguistics.
  19. 19.Robert Gray. 1984. Vector quantization. IEEE Assp Magazine, 1(2):4–29.
  20. 20.Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. 2020. Realm: Retrieval-augmented language model pre-training. arXiv preprint arXiv:2002.08909.
  21. 21.Sebastian Hofstätter, Sophia Althammer, Michael Schröder, Mete Sertkan, and Allan Hanbury. 2020. Improving Efficient Neural Ranking Models with Cross-Architecture Knowledge Distillation. arXiv preprint arXiv:2010.02666.
  22. 22.Sebastian Hofstätter, Sheng-Chieh Lin, Jheng-Hong Yang, Jimmy Lin, and Allan Hanbury. 2021. Efficiently Teaching an Effective Dense Retriever with Balanced Topic Aware Sampling. arXiv preprint arXiv:2104.06967.
  23. 23.Samuel Humeau, Kurt Shuster, Marie-Anne Lachaux, and Jason Weston. 2020. Poly-encoders: Architectures and pre-training strategies for fast and accurate multi-sentence scoring. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  24. 24.Gautier Izacard, Fabio Petroni, Lucas Hosseini, Nicola De Cao, Sebastian Riedel, and Edouard Grave. 2020. A memory efficient baseline for open domain question answering. arXiv preprint arXiv:2012.15156.
  25. 25.Herve Jegou, Matthijs Douze, and Cordelia Schmid. 2010. Product quantization for nearest neighbor search. IEEE transactions on pattern analysis and machine intelligence, 33(1):117–128.
  26. 26.Yichen Jiang, Shikha Bordia, Zheng Zhong, Charles Dognin, Maneesh Singh, and Mohit Bansal. 2020. HoVer: A dataset for many-hop fact extraction and claim verification. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3441–3460, Online. Association for Computational Linguistics.
  27. 27.Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2019. Billion-scale similarity search with gpus. IEEE Transactions on Big Data.
  28. 28.Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017. TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1601–1611, Vancouver, Canada. Association for Computational Linguistics.
  29. 29.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6769–6781, Online. Association for Computational Linguistics.
  30. 30.Daniel Khashabi, Amos Ng, Tushar Khot, Ashish Sabharwal, Hannaneh Hajishirzi, and Chris Callison-Burch. 2021. GooAQ: Open Question Answering with Diverse Answer Types. arXiv preprint arXiv:2104.08727.
  31. 31.Omar Khattab, Christopher Potts, and Matei Zaharia. 2021a. Baleen: Robust Multi-Hop Reasoning at Scale via Condensed Retrieval. In Thirty-Fifth Conference on Neural Information Processing Systems.
  32. 32.Omar Khattab, Christopher Potts, and Matei Zaharia. 2021b. Relevance-guided supervision for openqa with ColBERT. Transactions of the Association for Computational Linguistics, 9:929–944.
  33. 33.Omar Khattab and Matei Zaharia. 2020. Colbert: Efficient and effective passage search via contextualized late interaction over BERT. In Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval, SIGIR 2020, Virtual Event, China, July 25-30, 2020, pages 39–48. ACM.
  34. 34.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019. Natural questions: A benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:452–466.
  35. 35.Jinhyuk Lee, Mujeen Sung, Jaewoo Kang, and Danqi Chen. 2021a. Learning dense representations of phrases at scale. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6634–6647, Online. Association for Computational Linguistics.
  36. 36.Jinhyuk Lee, Alexander Wettig, and Danqi Chen. 2021b. Phrase retrieval learns passage retrieval, too. arXiv preprint arXiv:2109.08133.
  37. 37.Kenton Lee, Ming-Wei Chang, and Kristina Toutanova. 2019. Latent retrieval for weakly supervised open domain question answering. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 6086–6096, Florence, Italy. Association for Computational Linguistics.
  38. 38.Yue Li, Wenrui Ding, Chunlei Liu, Baochang Zhang, and Guodong Guo. 2021a. TRQ: Ternary Neural Networks With Residual Quantization. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 8538–8546.
  39. 39.Zefan Li, Bingbing Ni, Teng Li, Xiaokang Yang, Wenjun Zhang, and Wen Gao. 2021b. Residual Quantization for Low Bit-width Neural Networks. IEEE Transactions on Multimedia.
  40. 40.Jimmy Lin and Xueguang Ma. 2021. A Few Brief Notes on DeepImpact, COIL, and a Conceptual Framework for Information Retrieval Techniques. arXiv preprint arXiv:2106.14807.
  41. 41.Sheng-Chieh Lin, Jheng-Hong Yang, and Jimmy Lin. 2020. Distilling Dense Representations for Ranking using Tightly-Coupled Teachers. arXiv preprint arXiv:2010.11386.
  42. 42.Xiaorui Liu, Yao Li, Jiliang Tang, and Ming Yan. 2020. A double residual compression algorithm for efficient distributed learning. In The 23rd International Conference on Artificial Intelligence and Statistics, AISTATS 2020, 26-28 August 2020, Online [Palermo, Sicily, Italy], volume 108 of Proceedings of Machine Learning Research, pages 133–143. PMLR.
  43. 43.Yi Luan, Jacob Eisenstein, Kristina Toutanova, and Michael Collins. 2021. Sparse, Dense, and Attentional Representations for Text Retrieval. Transactions of the Association for Computational Linguistics, 9:329–345.
  44. 44.Sean MacAvaney, Franco Maria Nardini, Raffaele Perego, Nicola Tonellotto, Nazli Goharian, and Ophir Frieder. 2020. Efficient document re-ranking for transformers by precomputing term representations. In Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval, SIGIR 2020, Virtual Event, China, July 25-30, 2020, pages 49–58. ACM.
  45. 45.Craig Macdonald and Nicola Tonellotto. 2021. On approximate nearest neighbour selection for multistage dense retrieval. In Proceedings of the 30th ACM International Conference on Information & Knowledge Management, pages 3318–3322.
  46. 46.Macedo Maia, Siegfried Handschuh, André Freitas, Brian Davis, Ross McDermott, Manel Zarrouk, and Alexandra Balahur. 2018. WWW'18 Open Challenge: Financial Opinion Mining and Question Answering. In Companion Proceedings of the The Web Conference 2018, pages 1941–1942.
  47. 47.Antonio Mallia, Omar Khattab, Torsten Suel, and Nicola Tonellotto. 2021. Learning passage impacts for inverted indexes. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 1723–1727.
  48. 48.Aditya Krishna Menon, Sadeep Jayasumana, Seungyeon Kim, Ankit Singh Rawat, Sashank J. Reddi, and Sanjiv Kumar. 2022. In defense of dualencoders for neural ranking.
  49. 49.Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016. MS MARCO: A human-generated MAchine reading COmprehension dataset. arXiv preprint arXiv:1611.09268.
  50. 50.Rodrigo Nogueira and Kyunghyun Cho. 2019. Passage Re-ranking with BERT. arXiv preprint arXiv:1901.04085.
  51. 51.Barlas Oguz, Kushal Lakhotia, Anchit Gupta, Patrick ˘ Lewis, Vladimir Karpukhin, Aleksandra Piktus, Xilun Chen, Sebastian Riedel, Wen-tau Yih, Sonal Gupta, et al. 2021. Domain-matched Pre-training Tasks for Dense Retrieval. arXiv preprint arXiv:2107.13602.
  52. 52.Ashwin Paranjape, Omar Khattab, Christopher Potts, Matei Zaharia, and Christopher D Manning. 2022. Hindsight: Posterior-guided training of retrievers for improved open-ended generation. In International Conference on Learning Representations.
  53. 53.Yingqi Qu, Yuchen Ding, Jing Liu, Kai Liu, Ruiyang Ren, Wayne Xin Zhao, Daxiang Dong, Hua Wu, and Haifeng Wang. 2021. RocketQA: An optimized training approach to dense passage retrieval for open-domain question answering. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5835–5847, Online. Association for Computational Linguistics.
  54. 54.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. SQuAD: 100,000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2383–2392, Austin, Texas. Association for Computational Linguistics.
  55. 55.Ruiyang Ren, Shangwen Lv, Yingqi Qu, Jing Liu, Wayne Xin Zhao, QiaoQiao She, Hua Wu, Haifeng Wang, and Ji-Rong Wen. 2021a. PAIR: Leveraging passage-centric similarity relation for improving dense passage retrieval. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 2173–2183, Online. Association for Computational Linguistics.
  56. 56.Ruiyang Ren, Yingqi Qu, Jing Liu, Wayne Xin Zhao, Qiaoqiao She, Hua Wu, Haifeng Wang, and Ji-Rong Wen. 2021b. RocketQAv2: A Joint Training Method for Dense Passage Retrieval and Passage Re-ranking. arXiv preprint arXiv:2110.07367.
  57. 57.Stephen E Robertson, Steve Walker, Susan Jones, Micheline M Hancock-Beaulieu, Mike Gatford, et al. 1995. Okapi at TREC-3. NIST Special Publication.
  58. 58.Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. 2021. BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models. arXiv preprint arXiv:2104.08663.
  59. 59.James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2018. FEVER: a large-scale dataset for fact extraction and VERification. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 809–819, New Orleans, Louisiana. Association for Computational Linguistics.
  60. 60.Ellen Voorhees, Tasmeer Alam, Steven Bedrick, Dina Demner-Fushman, William R Hersh, Kyle Lo, Kirk Roberts, Ian Soboroff, and Lucy Lu Wang. 2021. TREC-COVID: Constructing a Pandemic Information Retrieval Test Collection. In ACM SIGIR Forum, volume 54, pages 1–12. ACM New York, NY, USA.
  61. 61.Henning Wachsmuth, Shahbaz Syed, and Benno Stein. 2018. Retrieval of the best counterargument without prior topic knowledge. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 241–251, Melbourne, Australia. Association for Computational Linguistics.
  62. 62.David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi. 2020. Fact or fiction: Verifying scientific claims. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7534–7550, Online. Association for Computational Linguistics.
  63. 63.Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou. 2020. MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers. arXiv preprint arXiv:2002.10957.
  64. 64.Benchang Wei, Tao Guan, and Junqing Yu. 2014. Projected Residual Vector Quantization for ANN Search. IEEE multimedia, 21(3):41–51.
  65. 65.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  66. 66.Ji Xin, Chenyan Xiong, Ashwin Srinivasan, Ankita Sharma, Damien Jose, and Paul N Bennett. 2021. Zero-Shot Dense Retrieval with Momentum Adversarial Domain Invariant Representations. arXiv preprint arXiv:2110.07581.
  67. 67.Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul N Bennett, Junaid Ahmed, and Arnold Overwijk. 2020. Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval. In International Conference on Learning Representations.
  68. 68.Ikuya Yamada, Akari Asai, and Hannaneh Hajishirzi. 2021a. Efficient passage retrieval with hashing for open-domain question answering. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pages 979–986, Online. Association for Computational Linguistics.
  69. 69.Ikuya Yamada, Akari Asai, and Hannaneh Hajishirzi. 2021b. Efficient passage retrieval with hashing for open-domain question answering. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pages 979–986, Online. Association for Computational Linguistics.
  70. 70.Peilin Yang, Hui Fang, and Jimmy Lin. 2018a. Anserini: Reproducible ranking baselines using lucene. Journal of Data and Information Quality (JDIQ), 10(4):1–20.
  71. 71.Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018b. HotpotQA: A dataset for diverse, explainable multi-hop question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2369–2380, Brussels, Belgium. Association for Computational Linguistics.
  72. 72.Jingtao Zhan, Jiaxin Mao, Yiqun Liu, Jiafeng Guo, Min Zhang, and Shaoping Ma. 2021a. Jointly Optimizing Query Encoder and Product Quantization to Improve Retrieval Performance. In Proceedings of the 30th ACM International Conference on Information & Knowledge Management, pages 2487–2496.
  73. 73.Jingtao Zhan, Jiaxin Mao, Yiqun Liu, Jiafeng Guo, Min Zhang, and Shaoping Ma. 2021b. Optimizing Dense Retrieval Model Training with Hard Negatives. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 1503–1512.
  74. 74.Jingtao Zhan, Jiaxin Mao, Yiqun Liu, Jiafeng Guo, Min Zhang, and Shaoping Ma. 2022. Learning discrete representations via constrained clustering for effective and efficient dense retrieval. In Proceedings of the Fifteenth ACM International Conference on Web Search and Data Mining, WSDM '22, page 1328–1336. Association for Computing Machinery.
  75. 75.Jingtao Zhan, Jiaxin Mao, Yiqun Liu, Min Zhang, and Shaoping Ma. 2020a. Learning to retrieve: How to train a dense retrieval model effectively and efficiently. arXiv preprint arXiv:2010.10469.
  76. 76.Jingtao Zhan, Jiaxin Mao, Yiqun Liu, Min Zhang, and Shaoping Ma. 2020b. Repbert: Contextualized text embeddings for first-stage retrieval. arXiv preprint arXiv:2006.15498.
  77. 77.Giulio Zhou and Jacob Devlin. 2021. Multi-vector attention models for deep re-ranking. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 5452–5456, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.

Citation

MLA
Santhanam, K., et al. “ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2022, pp. 3715–34, https://doi.org/10.18653/v1/2022.naacl-main.272.
APA
Santhanam, K., Khattab, O., Saad-Falcon, J., Potts, C., & Zaharia, M. (2022). ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 3715–3734. https://doi.org/10.18653/v1/2022.naacl-main.272
Chicago
Santhanam, K., O. Khattab, J. Saad-Falcon, C. Potts, and M. Zaharia. 2022. “ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 3715–34. https://doi.org/10.18653/v1/2022.naacl-main.272.
Harvard
Santhanam, K. et al. (2022) “ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction”, Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp. 3715–3734. Available at: https://doi.org/10.18653/v1/2022.naacl-main.272.
Vancouver
1. Santhanam K, Khattab O, Saad-Falcon J, Potts C, Zaharia M (2022) ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction. In: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp 3715–3734

BibTeX

@inproceedings{santhanam-etal-2022-colbertv2,
    title = "{C}ol{BERT}v2: Effective and Efficient Retrieval via Lightweight Late Interaction",
    author = "Santhanam, Keshav  and
      Khattab, Omar  and
      Saad-Falcon, Jon  and
      Potts, Christopher  and
      Zaharia, Matei",
    editor = "Carpuat, Marine  and
      de Marneffe, Marie-Catherine  and
      Meza Ruiz, Ivan Vladimir",
    booktitle = "Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
    month = jul,
    year = "2022",
    address = "Seattle, United States",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.naacl-main.272/",
    doi = "10.18653/v1/2022.naacl-main.272",
    pages = "3715--3734"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/