PromptReps: Prompting Large Language Models to Generate Dense and Sparse Representations for Zero-Shot Document Retrieval
Shengyao ZhuangXueguang MaBevan KoopmanJimmy LinGuido Zuccon
Introduces a prompt-based approach that simultaneously extracts dense embeddings and sparse bag-of-words representations from large language models in a single forward pass, enabling zero-shot full-corpus document retrieval without expensive contrastive fine-tuning.
Modern search engines increasingly rely on large language models to identify and rank relevant documents across extensive databases. However, existing techniques face a critical operational trade-off: direct prompt-based re-ranking is computationally prohibitive across large collections because it requires running an inference pass for every document, while alternative search systems require resource-intensive contrastive training on massive text datasets. This training demands significant computing hardware, incurs high cloud costs, and consumes substantial energy and cooling water while risking poor performance when applied to unfamiliar data domains.
The article demonstrates that off-the-shelf generative language models can perform full-corpus document retrieval directly through prompt engineering, without requiring contrastive pre-training or human-labeled examples. It introduces PromptReps, a method that prompts a language model to represent text using a single word and extracts both dense embeddings (semantic vectors) and sparse representations (keyword weights) in a single model pass to construct a hybrid search index.
To evaluate this approach, the researchers conducted extensive empirical experiments using multiple language models, including Meta Llama 3 models ranging up to 70 billion parameters, Mistral 7B, and Phi-3-mini across established retrieval benchmarks such as MSMARCO, TREC Deep Learning, and the multi-domain BEIR benchmark comprising 13 distinct datasets. The study compared PromptReps against standard keyword baselines, heavily trained unsupervised embedding models, and fully supervised systems, while also testing variations in prompt phrasing, model scaling, downstream fine-tuning, and multi-word representation designs.
The findings show that combining dense and sparse outputs from PromptReps without any training outperforms standard keyword search (BM25) and previous unsupervised language-model-based encoders. On the BEIR benchmark, PromptReps paired with the 70-billion-parameter Llama 3 model achieved an average nDCG@10 score of 45.97, matching the performance of specialized embedding models trained on 1.3 billion text pairs. When combined with standard BM25 keyword search, performance rose to 50.66, rivaling fully supervised methods. Furthermore, using PromptReps as a starting initialization for downstream supervised fine-tuning required as few as 1,000 labeled examples to reach competitive performance, whereas conventional baselines without fine-tuning failed entirely.
These results demonstrate that large language models are inherently capable text encoders whose representation abilities can be unlocked through prompt engineering alone. For organizations deploying search systems, this significantly lowers computing costs, shortens development timelines, and mitigates the environmental impact associated with training specialized retrieval models. It also reduces risks tied to domain shifts in specialized or privacy-sensitive settings where labeled training data is scarce.
Organizations evaluating search architectures should consider prompt-based representation generation as a cost-effective alternative to contrastive pre-training, particularly when leveraging larger instruction-tuned models and combining dense and sparse signals. For teams with limited labeled data, using this prompting technique as an initialization before minimal fine-tuning offers a highly efficient path to high accuracy. Future work should explore automated prompt optimization, domain-specific instruction templates, and prompt compression techniques to reduce online query latency.
Decision-makers should note certain limitations: PromptReps introduces slight online latency during query processing due to prompt length and dual-index retrieval, though document indexing occurs offline and dense-sparse lookups can run in parallel. Additionally, because the approach uses foundation models in a black-box manner, retrieval outputs may inherit underlying model biases, and effectiveness depends heavily on prompt phrasing and the base model selected.
- Paper: Precise Zero-Shot Dense Retrieval without Relevance Labels, Luyu Gao et al. (2023). HyDE demonstrates how generative language models can perform zero-shot retrieval by producing hypothetical documents, providing foundational context for PromptReps's direct prompt-based representation approach.
- Paper: BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models, Nandan Thakur et al. (2021). BEIR establishes the standard multi-domain benchmark and zero-shot retrieval evaluation framework directly employed by PromptReps to validate dense and sparse retrieval effectiveness.
- Paper: Text Embeddings by Weakly-Supervised Contrastive Pre-training, Liang Wang et al. (2022). E5 details weakly-supervised contrastive pre-training for text embeddings, serving as a primary baseline and contrasting methodology that PromptReps seeks to replace via prompt engineering.
- Paper: ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction, Keshav Santhanam et al. (2022). ColBERTv2 explains multi-vector and lightweight index architectures in neural search, setting the background for PromptReps's hybrid single-pass dense and sparse indexing.
- Paper: Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking Agents, Weiwei Sun et al. (2023). This study analyzes the computational bottlenecks and accuracy of prompting LLMs as direct re-rankers, which PromptReps directly addresses by extracting reusable representations for full-corpus indexing.
- Paper: Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval, Lee Xiong et al. (2021). ANCE defines hard negative contrastive training for dense retrievers, highlighting the training-heavy paradigm that PromptReps bypasses with zero-shot prompting.
- Paper: Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm, Laria Reynolds et al. (2021). This paper examines how structured zero-shot prompting steers internal language model capabilities, underpinning PromptReps's technique of unlocking representations without fine-tuning.
- Paper: Making Text Embedders Few-Shot Learners, Chaofan Li et al. (2025). Extends prompt-based representation learning by showing how few-shot in-context learning examples can further adapt foundation model text embedders to novel tasks.
- Paper: M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation, Jianlv Chen et al. (2024). Extends multi-functional retrieval representations by training a unified model that integrates dense, sparse, and multi-vector search mechanisms across multilingual domains.
- Paper: Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models, Yanzhao Zhang et al. (2025). Generalizes foundation-model-driven representation learning by scaling text embedding and reranking systems across large-scale model families and downstream architectures.
- Paper: RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs, Yue Yu et al. (2024). Applies context-ranking capabilities within a unified retrieve-rerank-generate architecture to optimize downstream retrieval-augmented generation pipelines.
- Paper: Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale, Siddharth Gollapudi et al. (2026). Explores the alternative frontier of in-context retrieval where language models retrieve directly over large-scale contexts rather than utilizing separate prompt-extracted indices.
