Llama2Vec: Unsupervised Adaptation of Large Language Models for Dense Retrieval
Chaofan LiZheng LiuShitao XiaoYingxia ShaoDefu Lian
Develops an unsupervised adaptation method using embedding-based auto-encoding and auto-regression tasks to transform autoregressive language models into effective dense retrieval encoders, achieving state-of-the-art performance on MSMARCO and BEIR benchmarks.
Dense retrieval converts text queries and documents into mathematical representations, called embeddings, so search engines can match them based on underlying meaning rather than exact keywords. This capability is critical for search systems, question-answering applications, and generative artificial intelligence. While modern large language models offer exceptional language comprehension, they are natively trained for next-word text generation, which focuses heavily on local token relationships. This creates a fundamental gap when adapting them to dense retrieval, where an entire text passage must be compressed into a single global embedding.
The article introduces and evaluates Llama2Vec, a lightweight unsupervised training method designed to adapt large language models into highly accurate text encoders for dense retrieval. The primary objective is to demonstrate that an unsupervised intermediate adaptation phase enables large language models to capture whole-text semantics more effectively, setting new performance standards across standard information retrieval benchmarks.
To achieve this, the authors designed two complementary unsupervised training tasks: one where the model uses its text embedding to reconstruct the input text itself, and another where it predicts the content of the succeeding text passage. The authors applied this technique to a 7-billion-parameter open foundation model using an unlabeled Wikipedia text collection over 10,000 training steps, merging the prompt computations to cut processing overhead by nearly half. The adapted model was then fine-tuned on standard retrieval benchmarks and tested against various established search models.
The evaluations yielded several key findings. First, the adapted model established new state-of-the-art results for passage and document retrieval on the MS MARCO benchmark, achieving a Mean Reciprocal Rank at 10 of 43.1 on passage retrieval and 47.9 on document retrieval, notably outperforming unadapted base models and traditional smaller language models. Second, in zero-shot evaluations across diverse datasets in the BEIR benchmark, the method achieved an average score of 56.4, surpassing traditional keyword search (BM25) by approximately 31% relatively and consistently leading in 12 out of 14 domains. Third, analysis revealed that the unsupervised tasks significantly increased lexical alignment between queries and relevant answers prior to fine-tuning. Finally, when testing embedding compression strategies to reduce storage and compute overhead, embedding sparsification retained retrieval accuracy far more effectively than standard linear dimensionality reduction.
These findings demonstrate that directly fine-tuning large language models for dense retrieval leaves substantial performance gains untapped unless preceded by targeted global representation training. For organizations deploying search and retrieval systems, using properly adapted open models can deliver accuracy that surpasses proprietary commercial embedding services. While 7-billion-parameter models demand more compute and vector storage than legacy compact encoders, the approach significantly narrows the performance gap without requiring expensive labeled data generation or complex distillation pipelines.
Organizations developing high-performance search systems or retrieval-augmented generation pipelines should consider adopting unsupervised representation alignment prior to fine-tuning large language model backbones. When infrastructure cost or database memory is a constraint, technical teams should explore sparsification techniques over standard projection methods to downscale embedding dimensions. Further operational pilots are recommended before deploying these models in low-latency environments to assess hardware requirements and trade-offs.
The findings are subject to specific boundary conditions. The current evaluation focuses exclusively on a single 7-billion-parameter English-language architecture, leaving effectiveness on multilingual tasks or larger model scales unverified. Additionally, because the adapted system inherits the underlying foundation model's training data, stakeholders should exercise caution when deploying it in sensitive domains where biased or toxic source data could distort retrieval behavior.
- Paper: RetroMAE: Pre-Training Retrieval-oriented Language Models Via Masked Auto-Encoder, Shitao Xiao et al. (2022). RetroMAE shows how reconstruction from sentence embeddings can train retrieval-oriented representations, providing a direct precursor to Llama2Vec’s embedding-based auto-encoding task.
- Paper: Unsupervised Dense Information Retrieval with Contrastive Learning, Gautier Izacard et al. (2021). Contriever establishes unsupervised contrastive dense retrieval as a baseline, clarifying the retrieval setting that Llama2Vec targets with a different adaptation strategy.
- Paper: Llama 2: Open Foundation and Fine-Tuned Chat Models, Hugo Touvron et al. (2023). Llama2Vec adapts LLaMA-2-7B, so reading the Llama 2 report first explains the model family and its autoregressive foundation.
- Paper: Text Embeddings by Weakly-Supervised Contrastive Pre-training, Liang Wang et al. (2022). E5 demonstrates weakly supervised contrastive pre-training for general-purpose text embeddings, giving a useful retrieval-embedding baseline for Llama2Vec’s unsupervised adaptation.
- Paper: Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval, Lee Xiong et al. (2021). ANCE explains dense retrieval’s hard-negative training and benchmark context, helping situate the retrieval gains Llama2Vec reports.
- Paper: NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models, Chankyu Lee et al. (2025). NV-Embed carries forward the effort to turn decoder-only LLMs into strong general-purpose embedding models, testing further architectural and training changes after Llama2Vec’s adaptation approach.
- Paper: Making Text Embedders Few-Shot Learners, Chaofan Li et al. (2025). bge-en-icl extends LLM-based text embeddings toward few-shot task adaptation, building on the premise that large language models can serve as effective embedding backbones.
- Paper: Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models, Yanzhao Zhang et al. (2025). Qwen3 Embedding continues the use of foundation models for text embeddings with large-scale synthetic data and staged training, extending the adaptation agenda to newer models and broader tasks.
