PromptReps: Prompting Large Language Models to Generate Dense and Sparse Representations for Zero-Shot Document Retrieval

Shengyao ZhuangXueguang MaBevan KoopmanJimmy LinGuido Zuccon

article2024EMNLP54 citations

Introduces a prompt-based approach that simultaneously extracts dense embeddings and sparse bag-of-words representations from large language models in a single forward pass, enabling zero-shot full-corpus document retrieval without expensive contrastive fine-tuning.

Listen

Modern search engines increasingly rely on large language models to identify and rank relevant documents across extensive databases. However, existing techniques face a critical operational trade-off: direct prompt-based re-ranking is computationally prohibitive across large collections because it requires running an inference pass for every document, while alternative search systems require resource-intensive contrastive training on massive text datasets. This training demands significant computing hardware, incurs high cloud costs, and consumes substantial energy and cooling water while risking poor performance when applied to unfamiliar data domains.

The article demonstrates that off-the-shelf generative language models can perform full-corpus document retrieval directly through prompt engineering, without requiring contrastive pre-training or human-labeled examples. It introduces PromptReps, a method that prompts a language model to represent text using a single word and extracts both dense embeddings (semantic vectors) and sparse representations (keyword weights) in a single model pass to construct a hybrid search index.

To evaluate this approach, the researchers conducted extensive empirical experiments using multiple language models, including Meta Llama 3 models ranging up to 70 billion parameters, Mistral 7B, and Phi-3-mini across established retrieval benchmarks such as MSMARCO, TREC Deep Learning, and the multi-domain BEIR benchmark comprising 13 distinct datasets. The study compared PromptReps against standard keyword baselines, heavily trained unsupervised embedding models, and fully supervised systems, while also testing variations in prompt phrasing, model scaling, downstream fine-tuning, and multi-word representation designs.

The findings show that combining dense and sparse outputs from PromptReps without any training outperforms standard keyword search (BM25) and previous unsupervised language-model-based encoders. On the BEIR benchmark, PromptReps paired with the 70-billion-parameter Llama 3 model achieved an average nDCG@10 score of 45.97, matching the performance of specialized embedding models trained on 1.3 billion text pairs. When combined with standard BM25 keyword search, performance rose to 50.66, rivaling fully supervised methods. Furthermore, using PromptReps as a starting initialization for downstream supervised fine-tuning required as few as 1,000 labeled examples to reach competitive performance, whereas conventional baselines without fine-tuning failed entirely.

These results demonstrate that large language models are inherently capable text encoders whose representation abilities can be unlocked through prompt engineering alone. For organizations deploying search systems, this significantly lowers computing costs, shortens development timelines, and mitigates the environmental impact associated with training specialized retrieval models. It also reduces risks tied to domain shifts in specialized or privacy-sensitive settings where labeled training data is scarce.

Organizations evaluating search architectures should consider prompt-based representation generation as a cost-effective alternative to contrastive pre-training, particularly when leveraging larger instruction-tuned models and combining dense and sparse signals. For teams with limited labeled data, using this prompting technique as an initialization before minimal fine-tuning offers a highly efficient path to high accuracy. Future work should explore automated prompt optimization, domain-specific instruction templates, and prompt compression techniques to reduce online query latency.

Decision-makers should note certain limitations: PromptReps introduces slight online latency during query processing due to prompt length and dual-index retrieval, though document indexing occurs offline and dense-sparse lookups can run in parallel. Additionally, because the approach uses foundation models in a black-box manner, retrieval outputs may inherit underlying model biases, and effectiveness depends heavily on prompt phrasing and the base model selected.

arXiv: 2404.18424
Cover for PromptReps: Prompting Large Language Models to Generate Dense and Sparse Representations for Zero-Shot Document Retrieval

Abstract

Utilizing large language models (LLMs) for zero-shot document ranking is done in one of two ways: (1) prompt-based re-ranking methods, which require no further training but are only feasible for re-ranking a handful of candidate documents due to computational costs; and (2) unsupervised contrastive trained dense retrieval methods, which can retrieve relevant documents from the entire corpus but require a large amount of paired text data for contrastive training. In this paper, we propose PromptReps, which combines the advantages of both categories: no need for training and the ability to retrieve from the whole corpus. Our method only requires prompts to guide an LLM to generate query and document representations for effective document retrieval. Specifically, we prompt the LLMs to represent a given text using a single word, and then use the last token’s hidden states and the corresponding logits associated with the prediction of the next token to construct a hybrid document retrieval system. The retrieval system harnesses both dense text embedding and sparse bag-of-words representations given by the LLM. Our experimental evaluation on the MSMARCO, TREC deep learning and BEIR zero-shot document retrieval datasets illustrates that this simple prompt-based LLM retrieval method can achieve a similar or higher retrieval effectiveness than state-of-the-art LLM embedding methods that are trained with large amounts of unsupervised data, especially when using a larger LLM.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Supervised neural retrievers
  • 2.2 Unsupervised neural retrievers
  • 2.3 Prompting LLMs for document ranking
  • 2.4 Prompting LLMs for sentence embedding
  • 3 PromptReps
  • 4 Experimental Setup
  • 5 Zero-shot Results
  • 5.1 Zero-shot retrieval effectiveness on BEIR
  • 5.2 Further hybrid with BM25
  • 5.3 Sensitivity to different prompts
  • 5.4 Impact of different LLMs
  • 6 Supervised Results
  • 7 Alternative Representation and Scoring
  • 8 Conclusion
  • 9 Limitations
  • 10 Ethical Considerations
  • Acknowledgments
  • References
  • A Python code example
  • B Impact of Hybrid weights
  • C Full results on TREC deep learning and MSMARCO
  • D Fine-tuning hyper-parameters
  • E Full results of different representation methods

Knowls

  1. Knowl 1 — PromptReps Representation Generation Framework

    model/method

    PromptReps is a zero-shot retrieval method that uses prompt engineering on decoder-only large language models (LLMs) to simultaneously generate dense vector embeddings and sparse bag-of-words representations in a single forward pass without requiring contrastive pre-training or fine-tuning.

    For a document passage, the input prompt is structured as:

    • System message: You are an AI assistant that can understand human language.
    • User message: Passage: "[text]". Use one word to represent the passage in a retrieval task. Make sure your word is in lowercase.
    • Assistant prompt prefix: The word is: "

    For a query, the placeholder Passage is replaced with Query and the word passage is replaced with query (for duplicate question retrieval tasks such as Quora, the query prompt is used for both queries and documents).

    From the forward pass on this formatted input:

    1. Dense representation: The last-layer hidden state vector h∈Rd\mathbf{h} \in \mathbb{R}^d corresponding to the final prompt token (the opening quote token ") is extracted and L2L_2-normalized to serve as the dense semantic embedding for nearest-neighbor search.
    2. Sparse representation: The unnormalized next-token prediction logits z∈R∣V∣\mathbf{z} \in \mathbb{R}^{|V|} output by the language model head over the vocabulary VV at that final token position are extracted and subsequently sparsified into a bag-of-words vector for inverted index retrieval.
  2. Knowl 2 — Logit Sparsification and Quantization in PromptReps

    algorithm

    To transform the dense vocabulary logits z∈R∣V∣\mathbf{z} \in \mathbb{R}^{|V|} produced by an LLM into an efficient sparse inverted-index representation, PromptReps executes a four-step filtering, activation, truncation, and quantization pipeline:

    Input: Input text TT, LLM next-token logits z∈R∣V∣\mathbf{z} \in \mathbb{R}^{|V|}, LLM tokenizer TokTok, vocabulary VV
    Output: Sparse weight dictionary W={(ti,wi)}W = \{(t_i, w_i)\}
    1. Lowercase text Tlower←lower(T)T_{lower} \leftarrow \text{lower}(T)
    2. Tokenize TlowerT_{lower} into words using NLTK and filter out standard English stopwords and punctuation, yielding word list UU
    3. Initialize token ID set I←∅I \leftarrow \emptyset
    4. for each word u∈Uu \in U do
    5. I←I∪Tok.encode(u,add_special_tokens=False)I \leftarrow I \cup Tok.\text{encode}(u, \text{add\_special\_tokens}=\text{False})
    6. end for
    7. Set all dimensions of z\mathbf{z} not in II to zero: ∀j∉I,zj←0\forall j \notin I, z_j \leftarrow 0
    8. Apply SPLADE transformation: z′←log⁡(1+ReLU(z))\mathbf{z}' \leftarrow \log(1 + \text{ReLU}(\mathbf{z}))
    9. if count of non-zero elements in z′>128\mathbf{z}' > 128 then
    10. Retain only the top 128 highest values in z′\mathbf{z}', setting all other values to 0
    11. end if
    12. Quantize weights to integers: ∀j∈I,wj←⌊100⋅zj′⌉\forall j \in I, w_j \leftarrow \lfloor 100 \cdot z'_j \rceil
    13. return W={(Tok.decode(j),wj)∣j∈I∧wj>0}W = \{(Tok.\text{decode}(j), w_j) \mid j \in I \land w_j > 0\}
  3. Knowl 3 — Hybrid Score Normalization and Score Fusion

    equation

    In PromptReps, document ranking scores from dense semantic retrieval and sparse lexical retrieval are combined at query time via min-max normalization and linear interpolation.

    Let Sdense(q,d)S_{\text{dense}}(q, d) be the dot-product similarity between the normalized dense query vector and normalized dense document vector in the Approximate Nearest Neighbor (ANN) index. Let Ssparse(q,d)S_{\text{sparse}}(q, d) be the lexical match score computed from the inverted index using the sparsified query and document weights.

    Each score distribution over candidate documents for a query qq is normalized to [0,1][0, 1] via min-max scaling: S~∗(q,d)=S∗(q,d)−min⁡d′∈DS∗(q,d′)max⁡d′∈DS∗(q,d′)−min⁡d′∈DS∗(q,d′)\tilde{S}_{*}(q, d) = \frac{S_{*}(q, d) - \min_{d' \in D} S_{*}(q, d')}{\max_{d' \in D} S_{*}(q, d') - \min_{d' \in D} S_{*}(q, d')} where ∗∈{dense,sparse}* \in \{\text{dense}, \text{sparse}\}.

    The combined hybrid score is calculated as: Shybrid(q,d)=αS~sparse(q,d)+(1−α)S~dense(q,d)S_{\text{hybrid}}(q, d) = \alpha \tilde{S}_{\text{sparse}}(q, d) + (1 - \alpha) \tilde{S}_{\text{dense}}(q, d) In zero-shot evaluation, equal weighting α=0.5\alpha = 0.5 is used by default.

    When further hybridized with BM25 scores SBM25(q,d)S_{\text{BM25}}(q, d), each of the three min-max normalized score components is linearly interpolated with equal weight: Shybrid+BM25(q,d)=13S~dense(q,d)+13S~sparse(q,d)+13S~BM25(q,d)S_{\text{hybrid+BM25}}(q, d) = \frac{1}{3} \tilde{S}_{\text{dense}}(q, d) + \frac{1}{3} \tilde{S}_{\text{sparse}}(q, d) + \frac{1}{3} \tilde{S}_{\text{BM25}}(q, d)

  4. Knowl 4 — Zero-Shot Retrieval Performance of PromptReps on BEIR

    data/table

    PromptReps was evaluated across 13 publicly available BEIR benchmark datasets in a zero-shot setting against unsupervised baselines (BM25, E5-PTlarge, LLM2Vec based on Llama-3-8B-Instruct) and supervised models (SPLADE++, DRAGON+). Metric reported is nDCG@10.

    Sup Contrastive Unsup Contrastive PromptReps (ours)
    LLM - BERT110M BERT110M BERT330M Llama3-8B-I Llama3-8B-I Llama3-70B-I
    Dataset BM25 SPLADE++ DRAGON+ E5-PTlarge LLM2Vec Dense Sparse Hybrid Dense Sparse Hybrid
    arguana 39.70 52.1 46.9 44.4 51.73 29.70 22.85 33.32 31.65 24.66 35.27
    climatefever 16.51 22.8 22.7 15.7 23.58 19.92 9.98 21.38 19.95 12.14 22.18
    dbpedia 31.80 44.2 41.7 37.1 26.78 31.53 28.84 37.71 31.12 28.30 37.59
    fever 65.13 79.6 78.1 68.6 53.42 56.28 52.35 71.11 42.06 51.75 63.97
    fiqa 23.61 35.1 35.6 43.3 28.56 27.11 20.33 32.40 30.80 22.16 34.66
    hotpotqa 63.30 68.6 66.2 52.2 52.37 19.64 44.75 47.05 24.32 42.12 48.51
    nfcorpus 32.18 34.5 33.9 33.7 26.28 29.56 28.18 32.98 33.84 29.74 36.08
    nq 30.55 54.4 53.7 41.7 37.65 34.43 29.55 43.14 38.25 30.37 46.97
    quora 78.86 81.4 87.5 86.1 84.64 81.77 70.35 84.24 81.18 67.69 83.70
    scidocs 14.90 15.9 15.9 21.9 10.39 18.51 11.57 17.59 20.59 13.25 19.10
    scifact 67.89 69.9 67.9 72.3 66.36 52.68 58.48 65.71 63.12 61.53 70.34
    trec-covid 59.47 71.1 75.9 62.1 63.34 59.52 54.59 69.25 67.64 63.00 76.85
    touche 44.22 24.4 26.3 19.8 12.82 14.85 18.47 21.65 15.56 18.65 22.35
    avg 43.70 50.3 50.2 46.06 41.38 36.58 34.64 44.43 38.47 35.80 45.97

    PromptReps Hybrid using Llama-3-8B-Instruct (44.43 avg nDCG@10) outperforms the unsupervised contrastive trained LLM2Vec model (41.38) and BM25 (43.70). Scaling the base model to Llama-3-70B-Instruct improves average Hybrid performance to 45.97, matching the performance of E5-PTlarge (46.06), which was trained on 1.3 billion paired text samples.

  5. Knowl 5 — Unsupervised LLM Hybridization with BM25 on BEIR

    data/table

    When dense and sparse representations generated by unsupervised LLM retrievers are combined with BM25 using min-max score normalization and equal weighting, retrieval effectiveness improves substantially across all datasets (nDCG@10 reported).

    LLM BERT-330M Llama3-8B-I Llama3-70B-I
    Dataset E5-PT+BM25 D+S+BM25 D+S+BM25
    arguana 47.34 38.13 39.53
    climatefever 21.92 22.95 23.34
    dbpedia 43.46 40.87 41.63
    fever 78.04 77.07 74.06
    fiqa 42.88 34.11 35.35
    hotpotqa 69.18 64.38 65.29
    nfcorpus 36.61 35.20 37.64
    nq 47.71 45.12 48.30
    quora 88.63 86.26 86.60
    scidocs 20.76 17.97 18.82
    scifact 76.37 70.92 73.58
    trec-covid 74.09 76.17 80.29
    touche 35.01 29.13 34.15
    avg 52.46 49.10 50.66

    D+S+BM25 represents the three-way hybrid of PromptReps dense, PromptReps sparse, and BM25. Scaling from Llama3-8B-I (49.10 avg) to Llama3-70B-I (50.66 avg) produces zero-shot results that exceed fully supervised contrastive models such as SPLADE++ (50.3) and DRAGON+ (50.2).

  6. Knowl 6 — Prompt Design Ablation and Instruction Sensitivity

    empirical result

    The wording of the prompt substantially impacts PromptReps retrieval effectiveness across TREC Deep Learning (DL2019, DL2020 nDCG@10) and MSMARCO dev passage retrieval (MRR@10) using Meta-Llama-3-8B-Instruct:

    1. Assistant Prefix Anchor: Omitting the trailing assistant string The word is: " (Prompt #4) causes dense retrieval performance to collapse completely to 0.00 across DL2019, DL2020, and MSMARCO dev, and hybrid retrieval drops to 13.47 / 11.81 / 5.06. Without this prefix, the model's next-token logits predict non-representative conversational tokens (such as The).
    2. Task Specification: Including the phrase in a retrieval task increases performance across all evaluation sets (e.g., comparing Prompt #1 with this phrase vs. Prompt #2 without it: DL2019 Dense increases from 43.32 to 49.26, MSMARCO Hybrid increases from 21.76 to 23.68).
    3. Casing Instruction: Adding Make sure your word is in lowercase. (Prompt #6 vs. Prompt #1) skews the logit distribution toward lowercase tokens, matching the downstream lowercase exact matching filter and increasing MSMARCO MRR@10 from 23.68 to 24.62 and DL2020 Hybrid nDCG@10 from 54.35 to 56.66.
    4. Superlative Adjective: Adding most important (Prompt #3 vs. #6) has no significant impact on retrieval quality (DL2019 Hybrid 55.64 vs. 55.58; MSMARCO MRR@10 23.86 vs. 24.62).
  7. Knowl 7 — Effect of Model Architecture, Instruction Tuning, and Scale

    empirical result

    Evaluating PromptReps (using Prompt #6) across five decoder-only base LLMs reveals key architectural and scaling behaviors:

    • Base Model Choice: On MSMARCO dev (MRR@10) and TREC DL2019 (nDCG@10), Mistral-7B-Instruct-v0.2 performs poorly on dense representations (5.61 MRR@10; 13.96 DL2019 Dense; 32.58 DL2019 Hybrid). In contrast, Phi-3-mini-4k-instruct (3.8B) achieves 22.06 MSMARCO MRR@10 and 51.10 DL2019 Hybrid despite having fewer parameters.
    • Instruction Tuning: Instruction-tuned models significantly outperform base pretrained models. Meta-Llama-3-8B-Instruct achieves 24.62 MRR@10 on MSMARCO dev and 55.58 DL2019 Hybrid, whereas base Meta-Llama-3-8B achieves 22.31 MRR@10 on MSMARCO dev and 51.13 DL2019 Hybrid.
    • Model Scale: Scaling to Meta-Llama-3-70B-Instruct achieves the highest performance: 25.66 MRR@10 on MSMARCO dev, 58.39 DL2019 Hybrid nDCG@10, and 59.17 DL2020 Hybrid nDCG@10 (reaching 63.18 DL2019 and 62.55 DL2020 when combined with BM25).
  8. Knowl 8 — PromptReps as an Initialization for Contrastive Supervised Fine-Tuning

    empirical result

    When fine-tuning LLMs using InfoNCE loss and hard negatives mined by a BM25 + dense hybrid system (following the RepLLaMA recipe with LoRA rank 8), PromptReps provides a superior starting representation compared to RepLLaMA prompt formatting:

    • Zero-Shot State: Without fine-tuning, RepLlama-3 scores 0.0 across MSMARCO dev MRR@10, DL2019 nDCG@10, and DL2020 nDCG@10, whereas zero-shot PromptReps-hybrid achieves 24.62, 55.58, and 56.66 respectively.
    • Low-Resource Regime (1k samples): When trained on only 1k MSMARCO passage pairs with early stopping, PromptReps-hybrid reaches 28.18 MRR@10 (MSMARCO dev), 65.23 nDCG@10 (DL2019), and 64.01 nDCG@10 (DL2020), outperforming RepLlama-3 (27.88 MRR@10, 63.91 DL2019, 63.10 DL2020).
    • Full Supervision (490k samples): When fine-tuned for 1 epoch on full MSMARCO data (490k), PromptReps-hybrid achieves 74.49 nDCG@10 on DL2019 and 73.87 nDCG@10 on DL2020, exceeding RepLlama-3 (73.19 on DL2019, 73.35 on DL2020).
  9. Knowl 9 — Comparison of Single-Token, Multi-Token, and Multi-Vector Representation Schemes

    empirical result

    Five representation strategies were tested with PromptReps on the BEIR benchmark (using Llama-3-8B-Instruct):

    1. First-Token Single-Representation (FTSR): Uses hidden states and logits of the first predicted token ("). Achieves 44.43 average hybrid nDCG@10.
    2. First-Word Single-Representation (FWSR): Generates the full single word until the closing quote ", applying mean pooling to dense states and max pooling to logits. Achieves 37.26 average hybrid nDCG@10 (the lowest among all variants, indicating sub-word representations dilute token signal).
    3. Multi-Token Single-Representation (MTSR): Prompted with Use three words, generating multiple tokens pooled with mean (dense) and max (sparse). Achieves 44.33 average hybrid nDCG@10 with higher query latency.
    4. Multi-Token Multi-Representation (MTMR): Indexes each generated token vector separately and ranks via ColBERT MaxSim late interaction. Achieves 39.97 average hybrid nDCG@10.
    5. Multi-Word Multi-Representation (MWMR): Groups tokens by word boundaries, pools word tokens, indexes each word vector, and ranks via ColBERT MaxSim. Achieves 39.70 average hybrid nDCG@10.

    Single-representation FTSR outperforms multi-representation ColBERT-style schemes in zero-shot retrieval without requiring additional generation steps.

  10. Knowl 10 — Query Latency and Index Overhead of PromptReps

    limitation

    PromptReps introduces specific operational trade-offs during online querying and indexing:

    1. Query Encoding Latency: PromptReps prepends prompt instruction tokens to the input query, increasing the sequence length processed during online LLM forward inference compared to short-prefix bi-encoders.
    2. Dual Index Storage: Optimal retrieval effectiveness requires hybrid retrieval, which requires storing and querying both an Approximate Nearest Neighbor (ANN) vector index for dense representations and an inverted index for sparse representations.
    3. Computation Overhead: While sparse extraction adds minimal cost (a GPU matrix multiplication of the last hidden state with token embeddings followed by truncation), running full forward passes through 8B to 70B parameter models offline for indexing and online for querying requires substantially higher memory and compute than standard bi-encoder architectures like BERT.

Coverage note — None was omitted; all key contributions, algorithms, empirical evaluations (BEIR, TREC DL, MSMARCO), prompt ablations, model comparisons, supervised fine-tuning results, representation variations, and stated limitations are covered.

References

  1. 1.Marah Abdin et al. 2024. Phi-3 technical report: A highly capable language model locally on your phone. Preprint, arXiv:2404.14219.
  2. 2.AI@Meta. 2024. Llama 3 model card.
  3. 3.Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, Mir Rosenberg, Xia Song, Alina Stoica, Saurabh Tiwary, and Tong Wang. 2018. MS MARCO: A human generated machine reading comprehension dataset. Preprint, arXiv:1611.09268.
  4. 4.Elias Bassani and Luca Romelli. 2022. ranx.fuse: A Python library for metasearch. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management, CIKM '22, pages 4808–4812, New York, NY, USA. Association for Computing Machinery.
  5. 5.Parishad BehnamGhader, Vaibhav Adlakha, Marius Mosbach, Dzmitry Bahdanau, Nicolas Chapados, and Siva Reddy. 2024. Llm2vec: Large language models are secretly powerful text encoders. Preprint, arXiv:2404.05961.
  6. 6.Steven Bird and Edward Loper. 2004. NLTK: The natural language toolkit. In Proceedings of the ACL Interactive Poster and Demonstration Sessions, pages 214–217, Barcelona, Spain. Association for Computational Linguistics.
  7. 7.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  8. 8.Xin Cheng, Xun Wang, Xingxing Zhang, Tao Ge, Si-Qing Chen, Furu Wei, Huishuai Zhang, and Dongyan Zhao. 2024. xrag: Extreme context compression for retrieval-augmented generation with one token. Preprint, arXiv:2405.13792.
  9. 9.Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, and Ellen M. Voorhees. 2020. Overview of the TREC 2019 deep learning track. Preprint, arXiv:2003.07820.
  10. 10.Matthijs Douze, Alexandr Guzhva, Chengqi Deng, Jeff Johnson, Gergely Szilvasy, Pierre-Emmanuel Mazaré, Maria Lomeli, Lucas Hosseini, and Hervé Jégou. 2024. The faiss library. Preprint, arXiv:2401.08281.
  11. 11.Yan Fang, Jingtao Zhan, Qingyao Ai, Jiaxin Mao, Weihang Su, Jia Chen, and Yiqun Liu. 2024. Scaling laws for dense retrieval. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR '24, page 1339–1349, New York, NY, USA. Association for Computing Machinery.
  12. 12.Chrisantha Fernando, Dylan Banarse, Henryk Michalewski, Simon Osindero, and Tim Rocktäschel. 2023. Promptbreeder: Self-referential self-improvement via prompt evolution. Preprint, arXiv:2309.16797.
  13. 13.Thibault Formal, Carlos Lassance, Benjamin Piwowarski, and Stéphane Clinchant. 2022. From distillation to hard negative sampling: Making sparse neural IR models more effective. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR '22, page 2353–2359, New York, NY, USA. Association for Computing Machinery.
  14. 14.Thibault Formal, Benjamin Piwowarski, and Stéphane Clinchant. 2021. SPLADE: Sparse lexical and expansion model for first stage ranking. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR '21, pages 2288–2292, New York, NY, USA. Association for Computing Machinery.
  15. 15.Isabel O. Gallegos, Ryan A. Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, and Nesreen K. Ahmed. 2024. Bias and fairness in large language models: A survey. Computational Linguistics, 50(3):1097–1179.
  16. 16.Luyu Gao, Xueguang Ma, Jimmy Lin, and Jamie Callan. 2023a. Precise zero-shot dense retrieval without relevance labels. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1762–1777, Toronto, Canada. Association for Computational Linguistics.
  17. 17.Luyu Gao, Xueguang Ma, Jimmy Lin, and Jamie Callan. 2023b. Tevatron: An efficient and flexible toolkit for neural retrieval. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR '23, pages 3120–3124, New York, NY, USA. Association for Computing Machinery.
  18. 18.Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021. SimCSE: Simple contrastive learning of sentence embeddings. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6894–6910, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  19. 19.Tao Ge, Hu Jing, Lei Wang, Xun Wang, Si-Qing Chen, and Furu Wei. 2024. In-context autoencoder for context compression in a large language model. In The Twelfth International Conference on Learning Representations.
  20. 20.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. Measuring massive multitask language understanding. In International Conference on Learning Representations.
  21. 21.Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-rank adaptation of large language models. In International Conference on Learning Representations.
  22. 22.Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bojanowski, Armand Joulin, and Edouard Grave. 2022. Unsupervised dense information retrieval with contrastive learning. Preprint, arXiv:2112.09118.
  23. 23.Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023a. Mistral 7b. Preprint, arXiv:2310.06825.
  24. 24.Ting Jiang, Shaohan Huang, Zhongzhi Luan, Deqing Wang, and Fuzhen Zhuang. 2023b. Scaling sentence embeddings with large language models. Preprint, arXiv:2307.16645.
  25. 25.Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. Scaling laws for neural language models. Preprint, arXiv:2001.08361.
  26. 26.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6769–6781, Online. Association for Computational Linguistics.
  27. 27.Omar Khattab and Matei Zaharia. 2020. ColBERT: Efficient and effective passage search via contextualized late interaction over bert. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR '20, pages 39–48, New York, NY, USA. Association for Computing Machinery.
  28. 28.Ekaterina Khramtsova, Shengyao Zhuang, Mahsa Baktashmotlagh, Xi Wang, and Guido Zuccon. 2023. Selecting which dense retriever to use for zero-shot search. In Proceedings of the Annual International ACM SIGIR Conference on Research and Development in Information Retrieval in the Asia Pacific Region, SIGIR-AP '23, page 223–233, New York, NY, USA. Association for Computing Machinery.
  29. 29.Ekaterina Khramtsova, Shengyao Zhuang, Mahsa Baktashmotlagh, and Guido Zuccon. 2024. Leveraging LLMs for unsupervised dense retriever ranking. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR '24, page 1307–1317, New York, NY, USA. Association for Computing Machinery.
  30. 30.Taku Kudo and John Richardson. 2018. SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 66–71, Brussels, Belgium. Association for Computational Linguistics.
  31. 31.Jinhyuk Lee, Zhuyun Dai, Xiaoqi Ren, Blair Chen, Daniel Cer, Jeremy R. Cole, Kai Hui, Michael Boratko, Rajvi Kapadia, Wen Ding, Yi Luan, Sai Meher Karthik Duddu, Gustavo Hernandez Abrego, Weiqiang Shi, Nithi Gupta, Aditya Kusupati, Prateek Jain, Siddhartha Reddy Jonnalagadda, Ming-Wei Chang, and Iftekhar Naim. 2024. Gecko: Versatile text embeddings distilled from large language models. Preprint, arXiv:2403.20327.
  32. 32.Yibin Lei, Di Wu, Tianyi Zhou, Tao Shen, Yu Cao, Chongyang Tao, and Andrew Yates. 2024. Meta-task prompting elicits embedding from large language models. Preprint, arXiv:2402.18458.
  33. 33.Chaofan Li, Zheng Liu, Shitao Xiao, and Yingxia Shao. 2023. Making large language models a better foundation for dense retrieval. Preprint, arXiv:2312.15503.
  34. 34.Jiayi Liao, Sihang Li, Zhengyi Yang, Jiancan Wu, Yancheng Yuan, Xiang Wang, and Xiangnan He. 2024. Llara: Large language-recommendation assistant. Preprint, arXiv:2312.02445.
  35. 35.Jimmy Lin and Xueguang Ma. 2021. A few brief notes on DeepImpact, COIL, and a conceptual framework for information retrieval techniques. Preprint, arXiv:2106.14807.
  36. 36.Jimmy Lin, Xueguang Ma, Sheng-Chieh Lin, Jheng-Hong Yang, Ronak Pradeep, and Rodrigo Nogueira. 2021. Pyserini: A Python toolkit for reproducible information retrieval research with sparse and dense representations. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR '21, pages 2356–2362, New York, NY, USA. Association for Computing Machinery.
  37. 37.Sheng-Chieh Lin, Akari Asai, Minghan Li, Barlas Oguz, Jimmy Lin, Yashar Mehdad, Wen-tau Yih, and Xilun Chen. 2023. How to train your Dragon: Diverse augmentation towards generalizable dense retrieval. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 6385–6400, Singapore. Association for Computational Linguistics.
  38. 38.Yiheng Liu, Tianle Han, Siyuan Ma, Jiayue Zhang, Yuanyuan Yang, Jiaming Tian, Hao He, Antong Li, Mengshen He, Zhengliang Liu, Zihao Wu, Lin Zhao, Dajiang Zhu, Xiang Li, Ning Qiang, Dingang Shen, Tianming Liu, and Bao Ge. 2023. Summary of ChatGPT-related research and perspective towards the future of large language models. Meta-Radiology, 1(2):100017.
  39. 39.Simon Lupart, Thibault Formal, and Stéphane Clinchant. 2022. MS-Shift: An analysis of MS MARCO distribution shifts on neural retrieval. In European Conference on Information Retrieval.
  40. 40.Xueguang Ma, Liang Wang, Nan Yang, Furu Wei, and Jimmy Lin. 2024. Fine-tuning LLaMA for multistage text retrieval. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR '24, page 2421–2425, New York, NY, USA. Association for Computing Machinery.
  41. 41.Xueguang Ma, Xinyu Zhang, Ronak Pradeep, and Jimmy Lin. 2023. Zero-shot listwise document reranking with a large language model. Preprint, arXiv:2305.02156.
  42. 42.Antonio Mallia, Omar Khattab, Torsten Suel, and Nicola Tonellotto. 2021. Learning passage impacts for inverted indexes. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR '21, pages 1723–1727, New York, NY, USA. Association for Computing Machinery.
  43. 43.OpenAI. 2024. GPT-4 technical report. Preprint, arXiv:2303.08774.
  44. 44.Ronak Pradeep, Sahel Sharifymoghaddam, and Jimmy Lin. 2023. RankVicuna: Zero-shot listwise document reranking with open-source large language models. Preprint, arXiv:2309.15088.
  45. 45.Zhen Qin, Rolf Jagerman, Kai Hui, Honglei Zhuang, Junru Wu, Le Yan, Jiaming Shen, Tianqi Liu, Jialu Liu, Donald Metzler, Xuanhui Wang, and Michael Bendersky. 2024. Large language models are effective text rankers with pairwise ranking prompting. Preprint, arXiv:2306.17563.
  46. 46.Ruiyang Ren, Yingqi Qu, Jing Liu, Xin Zhao, Qifei Wu, Yuchen Ding, Hua Wu, Haifeng Wang, and Ji-Rong Wen. 2023. A thorough examination on zero-shot dense retrieval. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 15783–15796, Singapore. Association for Computational Linguistics.
  47. 47.Devendra Sachan, Mike Lewis, Mandar Joshi, Armen Aghajanyan, Wen-tau Yih, Joelle Pineau, and Luke Zettlemoyer. 2022. Improving passage retrieval with zero-shot question generation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 3781–3797, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  48. 48.Harrisen Scells, Shengyao Zhuang, and Guido Zuccon. 2022. Reduce, reuse, recycle: Green information retrieval research. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR '22, page 2825–2837, New York, NY, USA. Association for Computing Machinery.
  49. 49.Weiwei Sun, Lingyong Yan, Xinyu Ma, Shuaiqiang Wang, Pengjie Ren, Zhumin Chen, Dawei Yin, and Zhaochun Ren. 2023. Is ChatGPT good at search? investigating large language models as re-ranking agents. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 14918–14937, Singapore. Association for Computational Linguistics.
  50. 50.Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. 2021. BEIR: A heterogeneous benchmark for zero-shot evaluation of information retrieval models. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2).
  51. 51.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023. Llama 2: Open foundation and fine-tuned chat models. Preprint, arXiv:2307.09288.
  52. 52.Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei. 2024a. Text embeddings by weakly-supervised contrastive pre-training. Preprint, arXiv:2212.03533.
  53. 53.Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei. 2024b. Improving text embeddings with large language models. Preprint, arXiv:2401.00368.
  54. 54.Shuai Wang, Shengyao Zhuang, and Guido Zuccon. 2021. BERT-based dense retrievers require interpolation with bm25 for effective passage retrieval. In Proceedings of the 2021 ACM SIGIR International Conference on Theory of Information Retrieval, ICTIR '21, pages 317–324, New York, NY, USA. Association for Computing Machinery.
  55. 55.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou. 2022. Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems, volume 35, pages 24824–24837. Curran Associates, Inc.
  56. 56.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  57. 57.Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul N. Bennett, Junaid Ahmed, and Arnold Overwijk. 2021. Approximate nearest neighbor negative contrastive learning for dense text retrieval. In International Conference on Learning Representations.
  58. 58.Chengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu, Quoc V Le, Denny Zhou, and Xinyun Chen. 2024. Large language models as optimizers. In The Twelfth International Conference on Learning Representations.
  59. 59.Bowen Zhang, Kehua Chang, and Chunping Li. 2024. Simple techniques for enhancing sentence embeddings in generative language models. Preprint, arXiv:2404.03921.
  60. 60.Shengyao Zhuang, Bing Liu, Bevan Koopman, and Guido Zuccon. 2023. Open-source large language models are strong zero-shot query likelihood models for document ranking. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 8807–8817, Singapore. Association for Computational Linguistics.
  61. 61.Shengyao Zhuang, Honglei Zhuang, Bevan Koopman, and Guido Zuccon. 2024. A setwise approach for effective and highly efficient zero-shot ranking with large language models. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR '24, page 38–47, New York, NY, USA. Association for Computing Machinery.
  62. 62.Shengyao Zhuang and Guido Zuccon. 2021a. Dealing with typos for BERT-based passage retrieval and ranking. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 2836–2842, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  63. 63.Shengyao Zhuang and Guido Zuccon. 2021b. Fast passage re-ranking with contextualized exact term matching and efficient passage expansion. Preprint, arXiv:2108.08513.
  64. 64.Shengyao Zhuang and Guido Zuccon. 2021c. TILDE: Term independent likelihood model for passage re-ranking. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR '21, pages 1483–1492, New York, NY, USA. Association for Computing Machinery.
  65. 65.Shengyao Zhuang and Guido Zuccon. 2022. Characterbert and self-teaching for improving the robustness of dense retrievers on queries with typos. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR '22, page 1444–1454, New York, NY, USA. Association for Computing Machinery.
  66. 66.Guido Zuccon, Harrisen Scells, and Shengyao Zhuang. 2023. Beyond CO2 emissions: The overlooked impact of water consumption of information retrieval models. In Proceedings of the 2023 ACM SIGIR International Conference on Theory of Information Retrieval, ICTIR '23, page 283–289, New York, NY, USA. Association for Computing Machinery.

Citation

MLA
Zhuang, S., et al. “PromptReps: Prompting Large Language Models to Generate Dense and Sparse Representations for Zero-Shot Document Retrieval”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 4375–91, https://doi.org/10.18653/v1/2024.emnlp-main.250.
APA
Zhuang, S., Ma, X., Koopman, B., Lin, J., & Zuccon, G. (2024). PromptReps: Prompting Large Language Models to Generate Dense and Sparse Representations for Zero-Shot Document Retrieval. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 4375–4391. https://doi.org/10.18653/v1/2024.emnlp-main.250
Chicago
Zhuang, S., X. Ma, B. Koopman, J. Lin, and G. Zuccon. 2024. “PromptReps: Prompting Large Language Models to Generate Dense and Sparse Representations for Zero-Shot Document Retrieval”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 4375–91. https://doi.org/10.18653/v1/2024.emnlp-main.250.
Harvard
Zhuang, S. et al. (2024) “PromptReps: Prompting Large Language Models to Generate Dense and Sparse Representations for Zero-Shot Document Retrieval”, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 4375–4391. Available at: https://doi.org/10.18653/v1/2024.emnlp-main.250.
Vancouver
1. Zhuang S, Ma X, Koopman B, Lin J, Zuccon G (2024) PromptReps: Prompting Large Language Models to Generate Dense and Sparse Representations for Zero-Shot Document Retrieval. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 4375–4391

BibTeX

@inproceedings{zhuang-etal-2024-promptreps,
    title = "{P}rompt{R}eps: Prompting Large Language Models to Generate Dense and Sparse Representations for Zero-Shot Document Retrieval",
    author = "Zhuang, Shengyao  and
      Ma, Xueguang  and
      Koopman, Bevan  and
      Lin, Jimmy  and
      Zuccon, Guido",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.250/",
    doi = "10.18653/v1/2024.emnlp-main.250",
    pages = "4375--4391"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/