Llama2Vec: Unsupervised Adaptation of Large Language Models for Dense Retrieval

Chaofan LiZheng LiuShitao XiaoYingxia ShaoDefu Lian

article2024ACL120 citations

Develops an unsupervised adaptation method using embedding-based auto-encoding and auto-regression tasks to transform autoregressive language models into effective dense retrieval encoders, achieving state-of-the-art performance on MSMARCO and BEIR benchmarks.

Listen

Dense retrieval converts text queries and documents into mathematical representations, called embeddings, so search engines can match them based on underlying meaning rather than exact keywords. This capability is critical for search systems, question-answering applications, and generative artificial intelligence. While modern large language models offer exceptional language comprehension, they are natively trained for next-word text generation, which focuses heavily on local token relationships. This creates a fundamental gap when adapting them to dense retrieval, where an entire text passage must be compressed into a single global embedding.

The article introduces and evaluates Llama2Vec, a lightweight unsupervised training method designed to adapt large language models into highly accurate text encoders for dense retrieval. The primary objective is to demonstrate that an unsupervised intermediate adaptation phase enables large language models to capture whole-text semantics more effectively, setting new performance standards across standard information retrieval benchmarks.

To achieve this, the authors designed two complementary unsupervised training tasks: one where the model uses its text embedding to reconstruct the input text itself, and another where it predicts the content of the succeeding text passage. The authors applied this technique to a 7-billion-parameter open foundation model using an unlabeled Wikipedia text collection over 10,000 training steps, merging the prompt computations to cut processing overhead by nearly half. The adapted model was then fine-tuned on standard retrieval benchmarks and tested against various established search models.

The evaluations yielded several key findings. First, the adapted model established new state-of-the-art results for passage and document retrieval on the MS MARCO benchmark, achieving a Mean Reciprocal Rank at 10 of 43.1 on passage retrieval and 47.9 on document retrieval, notably outperforming unadapted base models and traditional smaller language models. Second, in zero-shot evaluations across diverse datasets in the BEIR benchmark, the method achieved an average score of 56.4, surpassing traditional keyword search (BM25) by approximately 31% relatively and consistently leading in 12 out of 14 domains. Third, analysis revealed that the unsupervised tasks significantly increased lexical alignment between queries and relevant answers prior to fine-tuning. Finally, when testing embedding compression strategies to reduce storage and compute overhead, embedding sparsification retained retrieval accuracy far more effectively than standard linear dimensionality reduction.

These findings demonstrate that directly fine-tuning large language models for dense retrieval leaves substantial performance gains untapped unless preceded by targeted global representation training. For organizations deploying search and retrieval systems, using properly adapted open models can deliver accuracy that surpasses proprietary commercial embedding services. While 7-billion-parameter models demand more compute and vector storage than legacy compact encoders, the approach significantly narrows the performance gap without requiring expensive labeled data generation or complex distillation pipelines.

Organizations developing high-performance search systems or retrieval-augmented generation pipelines should consider adopting unsupervised representation alignment prior to fine-tuning large language model backbones. When infrastructure cost or database memory is a constraint, technical teams should explore sparsification techniques over standard projection methods to downscale embedding dimensions. Further operational pilots are recommended before deploying these models in low-latency environments to assess hardware requirements and trade-offs.

The findings are subject to specific boundary conditions. The current evaluation focuses exclusively on a single 7-billion-parameter English-language architecture, leaving effectiveness on multilingual tasks or larger model scales unverified. Additionally, because the adapted system inherits the underlying foundation model's training data, stakeholders should exercise caution when deploying it in sensitive domains where biased or toxic source data could distort retrieval behavior.

Cover for Llama2Vec: Unsupervised Adaptation of Large Language Models for Dense Retrieval

Abstract

Dense retrieval calls for discriminative embeddings to represent the semantic relationship between query and document. It may benefit from the using of large language models (LLMs), given LLMs’ strong capability on semantic understanding. However, the LLMs are learned by auto-regression, whose working mechanism is completely different from representing whole text as one discriminative embedding. Thus, it is imperative to study how to adapt LLMs properly so that they can be effectively initialized as the backbone encoder for dense retrieval.

In this paper, we propose a novel approach, called Llama2Vec, which performs unsupervised adaptation of LLM for its dense retrieval application. Llama2Vec consists of two pretext tasks: EBAE (Embedding-Based Auto-Encoding) and EBAR (Embedding-Based Auto-Regression), where the LLM is prompted to reconstruct the input sentence and predict the next sentence based on its text embeddings. Llama2Vec is simple, lightweight, but highly effective. It is used to adapt LLaMA-2-7B on the Wikipedia corpus. With a moderate steps of adaptation, it substantially improves the model’s fine-tuned performances on a variety of dense retrieval benchmarks. Notably, it results in the new state-of-the-art performances on popular benchmarks, such as passage and document retrieval on MSMARCO, and zero-shot retrieval on BEIR. The model and source code will be made publicly available to facilitate the future research. Our model is available at https://github.com/FlagOpen/FlagEmbedding.

Table of Contents

  • Abstract
  • 1 Introduction
  • 2 Related Works
  • 3 Llama2Vec
  • 3.1 Preliminary
  • 3.2 Unsupervised Adaptation
  • 4 Experiment
  • 4.1 Settings
  • 4.2 Supervised Performance
  • 4.3 Zero-shot Performance
  • 4.4 Technical Factors
  • 5 Conclusion
  • 6 Limitation
  • 7 Ethical Consideration
  • 8 Acknowledgements
  • References

Knowls

  1. Knowl 1 — Llama2Vec trains LLM embeddings to represent both the input and its continuation

    model/method

    Llama2Vec adapts a decoder-only language model for dense retrieval with two unsupervised pretext tasks. In EBAE (Embedding-Based Auto-Encoding), the model receives an input sentence followed by the prompt “The input sentence is:” and an end token; the hidden state at that final token is trained to predict the vocabulary items in the input sentence. This embedding is intended to encode the input’s global, inductive semantics. In EBAR (Embedding-Based Auto-Regression), the model receives the input followed by “The next sentence is:” and an end token; its final-token hidden state is trained to predict the vocabulary items in the following sentence, encouraging a deductive representation of related text. The two task-specific embeddings are intended to support retrieval both for correlation relations, such as question–answer matching, and paraphrase relations.

    The reported adaptation uses the base LLaMA-2-7B model and the unlabeled Wikipedia corpus curated for DPR. It runs for 10,000 steps with batch size 256, sequence length 1,024, and learning rate 10−510^{-5}.

  2. Knowl 2 — A shared prompt pass computes the two task embeddings efficiently

    model/method

    Instead of running the decoder-only LLM twice on the same input sentence, Llama2Vec concatenates the task prompts into one sequence: input, SELF, end token, NEXT, end token. It modifies causal attention so that the SELF and NEXT prompt branches cannot attend to one another, while each branch can use the shared input and its own prompt. The hidden states at the two end tokens then provide the EBAE and EBAR embeddings, respectively. Since the input tokens make up most of the sequence, sharing their processing saves almost 50% of the computation compared with separately evaluating both prompts.

  3. Knowl 3 — Vocabulary prediction trains each embedding to summarize its target context

    equation

    For either Llama2Vec pretext task, a single sentence embedding is trained to predict every token instance in its target context using a vocabulary softmax. Let e∈Rde\in\mathbb{R}^d be the task embedding, W∈Rd×∣V∣W\in\mathbb{R}^{d\times |V|} the trainable vocabulary projection head, VV the model vocabulary, and x1,…,xmx_1,\ldots,x_m the mm token instances in the target sentence. The target is the input sentence for EBAE and the following sentence for EBAR. The loss is

    L(e;x1:m)=−1m∑i=1mlog⁡exp⁡(e⊤Wxi)∑v∈Vexp⁡(e⊤Wv).\mathcal{L}(e;x_{1:m})=-\frac{1}{m}\sum_{i=1}^{m}\log\frac{\exp(e^\top W_{x_i})}{\sum_{v\in V}\exp(e^\top W_v)}.

    Thus, each target token is classified from the same sentence-level embedding rather than predicted autoregressively from a sequence of preceding target tokens. The paper’s stated rationale is that an embedding able to predict the target context’s vocabulary on its own must capture global information about that context.

  4. Knowl 4 — Retrieval fine-tuning pairs prompts with the semantic relation being retrieved

    model/method

    After unsupervised adaptation, Llama2Vec is fine-tuned with a contrastive retrieval objective. For a correlation task such as question answering, the query uses the NEXT prompt and the answer uses SELF. Given query qq, its positive answer a+a^+, and a candidate set A′A' containing the positive and negative answers, the loss is

    Lret=−∑qlog⁡exp⁡ ⁣(⟨eNEXT(q),eSELF(a+)⟩)∑a′∈A′exp⁡ ⁣(⟨eNEXT(q),eSELF(a′)⟩).\mathcal{L}_{\mathrm{ret}}=-\sum_q\log\frac{\exp\!\left(\langle e_{\mathrm{NEXT}}(q),e_{\mathrm{SELF}}(a^+)\rangle\right)}{\sum_{a'\in A'}\exp\!\left(\langle e_{\mathrm{NEXT}}(q),e_{\mathrm{SELF}}(a')\rangle\right)}.

    Here, eNEXT(q)e_{\mathrm{NEXT}}(q) and eSELF(a)e_{\mathrm{SELF}}(a) are the prompted embeddings and ⟨⋅,⋅⟩\langle\cdot,\cdot\rangle is their inner-product similarity. For paraphrase retrieval, the paper instead uses SELF for both long documents and NEXT for both short sentences. In the reported MS MARCO fine-tuning recipe, the model is trained with LoRA parameter-efficient adaptation and approximate-nearest-neighbor hard negatives.

  5. Knowl 5 — Llama2Vec improves MS MARCO passage retrieval, especially on the large development set

    data/table

    The adapted LLaMA-2-7B retriever was fine-tuned on MS MARCO passage queries using the hard-negative recipe described in the paper. The selected rows below compare it with the unadapted LLaMA-2-7B RepLLaMA baseline and strong previous passage retrievers. Metrics are MRR@10 and Recall@1000 on the development set, followed by NDCG@10 on DL’19 and DL’20.

    Method Dev MRR@10 Dev Recall@1000 DL’19 NDCG@10 DL’20 NDCG@10
    SimLM+distill 41.1 98.7 71.4 69.7
    RetroMAEv2+distill 42.6 98.9 75.1 –
    LLaMA2-RepLLaMA 41.2 99.4 74.3 72.1
    LLaMA2-Llama2Vec 43.1 99.5 73.4 72.9

    Llama2Vec obtains the best shown development-set MRR@10 and Recall@1000, exceeding the same-backbone RepLLaMA baseline by 1.9 and 0.1 points, respectively. Its DL’20 score also exceeds RepLLaMA’s, while its DL’19 score is lower than both RepLLaMA and RetroMAEv2+distill; the results therefore do not show a win on every evaluation column.

  6. Knowl 6 — Llama2Vec improves MS MARCO document retrieval over the unadapted LLaMA-2 baseline

    data/table

    The adapted LLaMA-2-7B retriever was fine-tuned for MS MARCO document retrieval. The table compares the strongest conventional baseline shown, COSTA, the same-backbone unadapted RepLLaMA model, and Llama2Vec. Metrics are development-set MRR@100 and Recall@100, followed by DL’19 and DL’20 NDCG@10.

    Method Dev MRR@100 Dev Recall@100 DL’19 NDCG@10 DL’20 NDCG@10
    COSTA 42.2 91.9 62.6 –
    LLaMA2-RepLLaMA 45.6 – 65.0 63.2
    LLaMA2-Llama2Vec 47.9 94.1 68.2 63.6

    Llama2Vec has the highest score among these comparisons in every reported column, with a 2.3-point development MRR@100 improvement over RepLLaMA and a 5.7-point improvement over COSTA. The paper attributes the feasibility of encoding full documents, rather than splitting them into smaller segments, partly to the LLM’s long context.

  7. Knowl 7 — MS MARCO-fine-tuned Llama2Vec transfers strongly to zero-shot retrieval benchmarks

    data/table

    The LLaMA-2-7B model adapted on Wikipedia and fine-tuned on MS MARCO was evaluated directly, without benchmark-specific fine-tuning, on BEIR and Llama Index. BEIR scores are NDCG@10; model sizes are included as reported. The comparison shows Llama2Vec’s average BEIR score is 56.4, above the next-highest listed average, RepLLaMA’s 53.9, and BM25’s 42.9.

    Dataset BM25 BERT GTR-XXL CPT-XL Ada-2 SGPT RepLLaMA Llama2Vec
    Size – 110M 4.8B 175B – 5.8B 7B 7B
    T-COVID 59.5 61.5 50.1 64.9 81.3 87.3 84.7 86.9
    NFCorpus 32.2 26.0 34.2 40.7 35.8 36.2 37.8 38.2
    NQ 30.6 46.7 56.8 – 48.2 52.4 62.4 64.6
    HotpotQA 63.3 48.8 59.9 68.8 65.4 59.3 68.5 70.1
    FiQA 23.6 25.2 46.7 51.2 41.1 37.2 45.8 48.5
    ArguAna 39.7 26.5 54.0 43.5 56.7 51.4 48.6 56.5
    Touche 44.2 25.9 25.6 29.1 28.0 25.4 30.5 34.2
    Quora 78.9 78.7 89.2 63.8 87.6 84.6 86.8 88.3
    DBPedia 31.8 31.4 40.8 43.2 40.2 39.9 43.7 45.9
    SCIDOCS 14.1 11.3 16.1 – 18.6 19.7 18.1 18.9
    FEVER 65.1 68.2 74.0 77.5 77.3 78.3 83.4 81.3
    C-FEVER 16.5 18.7 26.7 22.3 23.7 30.5 31.0 38.2
    SciFact 67.9 53.3 66.2 75.4 73.6 74.7 75.6 74.8
    CQA 32.5 28.2 39.9 – 41.7 38.1 37.4 43.2
    Average 42.9 39.3 48.6 – 51.4 52.1 53.9 56.4

    On the separate Llama Index evaluation, the reported MRR and Hit Rate were 70.56 and 92.31 for Llama2Vec, versus 68.43 and 91.35 for RepLLaMA, 69.67 and 88.94 for bge-m3, 65.69 and 89.42 for OpenAI-TE3-S, and 67.37 and 90.38 for OpenAI-TE3-L. Llama2Vec therefore has the highest score on both listed Llama Index metrics.

  8. Knowl 8 — Choosing prompts by relation type improves zero-shot BEIR retrieval

    data/table

    Llama2Vec can use different task prompts for different retrieval relations. N2S means NEXT for the query and SELF for the candidate; S2S means SELF for both; N2N means NEXT for both; “None” is the no-prompt baseline. Ada* selects N2S for correlation tasks, S2S for long-document paraphrase (ArguAna), and N2N for short-text paraphrase. The entries below are BEIR NDCG@10; dataset abbreviations are AR (ArguAna), CF (Climate-FEVER), DB (DBPedia), FV (FEVER), FQ (FiQA), HQ (HotpotQA), NF (NFCorpus), NQ (Natural Questions), QR (Quora), SD (SCIDOCS), SF (SciFact), TO (Touche), TC (T-COVID), and CQ (CQADupStack).

    Prompt Avg AR CF DB FV FQ HQ NF NQ QR SD SF TO TC CQ
    N2S 56.1 50.5 27.7 45.9 83.5 48.5 70.1 38.2 64.6 83.7 18.9 76.5 34.2 86.9 41.2
    S2S 30.2 56.5 20.6 10.4 33.1 16.7 27.4 18.3 8.8 88.0 7.1 67.8 3.1 34.2 26.0
    N2N 52.3 49.9 38.2 40.5 81.3 43.8 65.5 35.3 54.9 88.3 20.5 74.8 22.6 63.9 43.2
    None 47.6 53.8 26.7 40.0 72.3 38.9 63.8 32.3 46.1 88.3 17.7 74.3 15.0 49.6 41.1
    Ada* 57.4 56.5 38.2 45.9 81.3 48.5 70.1 38.2 64.6 88.3 18.9 74.8 34.2 86.9 43.2

    Adaptive prompt selection gives the highest average among the listed configurations (57.4), compared with 56.1 for the fixed correlation-oriented N2S prompt and 47.6 with no prompt. The table also shows why one prompt is not best for every relation type: N2S is strong across many correlation datasets, while S2S and N2N perform better on particular paraphrase datasets.

  9. Knowl 9 — Adaptation increases lexical overlap between vocabulary predictions for queries and answers

    data/table

    To examine what changes during adaptation, the authors projected MS MARCO query and answer embeddings through the LLM’s vocabulary decoding head, selected the top-NN predicted vocabulary items for each, and measured their lexical similarity with BM25. The scores below compare the initial LLaMA-2 backbone, the unsupervised-adapted model, and the model after retrieval fine-tuning.

    Top-NN Initial Adapt Fine-tune
    10 1.85 2.74 13.83
    100 18.01 47.68 84.08
    500 93.68 205.42 307.30
    1000 219.09 392.80 542.89

    For every tested NN, the lexical-similarity score rises after unsupervised adaptation and rises further after retrieval fine-tuning. The paper interprets this as evidence that the adapted embeddings predict vocabulary more consistently across related queries and answers, a change it considers beneficial for retrieval.

  10. Knowl 10 — Embedding compression reduces retrieval scores; sparsification preserves them better

    data/table

    The paper compared dense dimension reduction with sparsifying the embedding, evaluating reported retrieval scores at embedding dimensions 768, 1,024, 2,048, and 4,096. DimRed jointly learns the LLM and a projection head during fine-tuning. DimRed* fixes the fine-tuned LLM retriever and learns a projection head by distillation. Sparse retains the top entries of the embedding. The source table does not name the score metric.

    Dimension DimRed DimRed* Sparse
    768 41.0 41.2 41.5
    1024 41.0 41.3 41.9
    2048 41.1 41.4 42.3
    4096 43.1 43.1 43.1

    At 4,096 dimensions all three methods score 43.1. At smaller dimensions, both dense projection approaches score lower, whereas sparsification retains higher scores among the three options—for example, 42.3 at dimension 2,048 versus 41.1 and 41.4 for DimRed and DimRed*. The results identify a trade-off between vector size and retrieval performance, with sparsification the more effective tested option for preserving scores.

Coverage note — The paper’s stated limitations and ethical caveats were not made separate knowls: they concern the evaluation’s scope (a 7B English-centric model and unresolved serving efficiency) and inherited LLaMA-2 risks, rather than an additional method or experimental finding.

References

  1. 1.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  2. 2.Jianlv Chen, Shitao Xiao, Peitian Zhang, Kun Luo, Defu Lian, and Zheng Liu. 2024. Bge m3-embedding: Multi-lingual, multi-functionality, multi-granularity text embeddings through self-knowledge distillation.
  3. 3.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2023. Palm: Scaling language modeling with pathways. Journal of Machine Learning Research, 24(240):1–113.
  4. 4.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2022. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416.
  5. 5.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,, pages 4171–4186. Association for Computational Linguistics.
  6. 6.Tommaso Dolci, Fabio Azzalini, and Mara Tanelli. 2023. Improving gender-related fairness in sentence encoders: A semantics-based approach. Data Science and Engineering, 8(2):177–195.
  7. 7.Thibault Formal, Carlos Lassance, Benjamin Piwowarski, and Stephane Clinchant. 2021. Splade v2: Sparse lexical and expansion model for information retrieval. arXiv preprint arXiv:2109.10086.
  8. 8.Luyu Gao and Jamie Callan. 2021. Condenser: a pre-training architecture for dense retrieval. arXiv preprint arXiv:2104.08253.
  9. 9.Luyu Gao and Jamie Callan. 2022. Unsupervised corpus aware language model pre-training for dense passage retrieval. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2843–2853, Dublin, Ireland.
  10. 10.Luyu Gao, Zhuyun Dai, and Jamie Callan. 2021. Coil: Revisit exact lexical match in information retrieval with contextualized inverted list. arXiv preprint arXiv:2104.07186.
  11. 11.Sebastian Hofstatter, Sheng-Chieh Lin, Jheng-Hong Yang, Jimmy Lin, and Allan Hanbury. 2021. Efficiently teaching an effective dense retriever with balanced topic aware sampling. In SIGIR, pages 113–122.
  12. 12.Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685.
  13. 13.Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bojanowski, Armand Joulin, and Edouard Grave. 2021. Unsupervised dense information retrieval with contrastive learning. arXiv preprint arXiv:2112.09118.
  14. 14.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. arXiv preprint arXiv:2004.04906.
  15. 15.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. 2019. Natural questions: a benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:453–466.
  16. 16.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Kuttler, Mike Lewis, Wen-tau Yih, Tim Rocktaschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing Systems, 33:9459–9474.
  17. 17.Jimmy Lin, Xueguang Ma, Sheng-Chieh Lin, Jheng-Hong Yang, Ronak Pradeep, and Rodrigo Nogueira. 2021. Pyserini: A python toolkit for reproducible information retrieval research with sparse and dense representations. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 2356–2362.
  18. 18.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized BERT pretraining approach. CoRR, abs/1907.11692.
  19. 19.Zheng Liu and Yingxia Shao. 2022. Retromae: Pre-training retrieval-oriented transformers via masked auto-encoder. arXiv preprint arXiv:2205.12035.
  20. 20.Zheng Liu, Shitao Xiao, Yingxia Shao, and Zhao Cao. 2023. Retromae-2: Duplex masked auto-encoder for pre-training retrieval-oriented language models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2635–2648.
  21. 21.Zhenghao Liu, Han Zhang, Chenyan Xiong, Zhiyuan Liu, Yu Gu, and Xiaohua Li. 2022. Dimension reduction for efficient dense retrieval via conditional autoencoder. arXiv preprint arXiv:2205.03284.
  22. 22.Xinyu Ma, Jiafeng Guo, Ruqing Zhang, Yixing Fan, and Xueqi Cheng. 2022. Pre-train a discriminative text encoder for dense retrieval via contrastive span prediction. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 848–858.
  23. 23.Xinyu Ma, Jiafeng Guo, Ruqing Zhang, Yixing Fan, Xiang Ji, and Xueqi Cheng. 2021a. Prop: Pre-training with representative words prediction for ad-hoc retrieval. In Proceedings of the 14th ACM international conference on web search and data mining, pages 283–291.
  24. 24.Xinyu Ma, Jiafeng Guo, Ruqing Zhang, Yixing Fan, Yingyan Li, and Xueqi Cheng. 2021b. B-prop: bootstrapped pre-training with representative words prediction for ad-hoc retrieval. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 1513–1522.
  25. 25.Xueguang Ma, Liang Wang, Nan Yang, Furu Wei, and Jimmy Lin. 2023. Fine-tuning llama for multi-stage text retrieval. arXiv preprint arXiv:2310.08319.
  26. 26.Yuning Mao, Pengcheng He, Xiaodong Liu, Yelong Shen, Jianfeng Gao, Jiawei Han, and Weizhu Chen. 2020. Generation-augmented retrieval for open-domain question answering. arXiv preprint arXiv:2009.08553.
  27. 27.Niklas Muennighoff. 2022. Sgpt: Gpt sentence embeddings for semantic search. arXiv preprint arXiv:2202.08904.
  28. 28.Arvind Neelakantan, Tao Xu, Raul Puri, Alec Radford, Jesse Michael Han, Jerry Tworek, Qiming Yuan, Nikolas Tezak, Jong Wook Kim, Chris Hallacy, et al. 2022. Text and code embeddings by contrastive pre-training. arXiv preprint arXiv:2201.10005.
  29. 29.Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016. Ms marco: A human generated machine reading comprehension dataset. choice, 2640:660.
  30. 30.Jianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai, Gustavo Hernandez Abrego, Ji Ma, Vincent Y Zhao, Yi Luan, Keith B Hall, Ming-Wei Chang, et al. 2021. Large dual encoders are generalizable retrievers. arXiv preprint arXiv:2112.07899.
  31. 31.Rodrigo Nogueira, Jimmy Lin, and AI Epistemic. 2019. From doc2query to doctttttquery. Online preprint, 6:2.
  32. 32.Yingqi Qu, Yuchen Ding, Jing Liu, Kai Liu, Ruiyang Ren, Wayne Xin Zhao, Daxiang Dong, Hua Wu, and Haifeng Wang. 2020. Rocketqa: An optimized training approach to dense passage retrieval for open-domain question answering. arXiv preprint arXiv:2010.08191.
  33. 33.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21(1):5485–5551.
  34. 34.Ruiyang Ren, Yingqi Qu, Jing Liu, Wayne Xin Zhao, Qiaoqiao She, Hua Wu, Haifeng Wang, and Ji-Rong Wen. 2021. Rocketqav2: A joint training method for dense passage retrieval and passage re-ranking. arXiv preprint arXiv:2110.07367.
  35. 35.Nandan Thakur, Nils Reimers, Andreas Ruckle, Abhishek Srivastava, and Iryna Gurevych. 2021. Beir: A heterogenous benchmark for zero-shot evaluation of information retrieval models. arXiv preprint arXiv:2104.08663.
  36. 36.James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2018. Fever: a large-scale dataset for fact extraction and verification. arXiv preprint arXiv:1803.05355.
  37. 37.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models, 2023. URL https://arxiv.org/abs/2307.09288.
  38. 38.Jiachen Wang, Jiajie Xu, Wei Chen, and Lei Zhao. 2022a. When research topic trend prediction meets fact-based annotations. Data Science and Engineering, 7(4):316–327.
  39. 39.Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei. 2022b. Simlm: Pre-training with representation bottleneck for dense passage retrieval. arXiv preprint arXiv:2207.02578.
  40. 40.Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei. 2022c. Text embeddings by weakly-supervised contrastive pre-training. arXiv preprint arXiv:2212.03533.
  41. 41.Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. 2021. Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652.
  42. 42.Shitao Xiao, Zheng Liu, Peitian Zhang, and Niklas Muennighof. 2023. C-pack: Packaged resources to advance general chinese embedding. arXiv preprint arXiv:2309.07597.
  43. 43.Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul Bennett, Junaid Ahmed, and Arnold Overwijk. 2020. Approximate nearest neighbor negative contrastive learning for dense text retrieval. arXiv preprint arXiv:2007.00808.
  44. 44.Jingtao Zhan, Jiaxin Mao, Yiqun Liu, Jiafeng Guo, Min Zhang, and Shaoping Ma. 2021. Optimizing dense retrieval model training with hard negatives. In SIGIR, pages 1503–1512.
  45. 45.Xin Zhang, Zehan Li, Yanzhao Zhang, Dingkun Long, Pengjun Xie, Meishan Zhang, and Min Zhang. 2023. Language models are universal embedders. arXiv preprint arXiv:2310.08232.
  46. 46.Kun Zhou, Yeyun Gong, Xiao Liu, Wayne Xin Zhao, Yelong Shen, Anlei Dong, Jingwen Lu, Rangan Majumder, Ji-Rong Wen, Nan Duan, et al. 2022. Simans: Simple ambiguous negatives sampling for dense text retrieval. arXiv preprint arXiv:2210.11773.

Citation

MLA
Liu, Z., et al. “Llama2Vec: Unsupervised Adaptation of Large Language Models for Dense Retrieval”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 3490–500, https://doi.org/10.18653/v1/2024.acl-long.191.
APA
Liu, Z., Li, C., Xiao, S., Shao, Y., & Lian, D. (2024). Llama2Vec: Unsupervised Adaptation of Large Language Models for Dense Retrieval. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3490–3500. https://doi.org/10.18653/v1/2024.acl-long.191
Chicago
Liu, Z., C. Li, S. Xiao, Y. Shao, and D. Lian. 2024. “Llama2Vec: Unsupervised Adaptation of Large Language Models for Dense Retrieval”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3490–3500. https://doi.org/10.18653/v1/2024.acl-long.191.
Harvard
Liu, Z. et al. (2024) “Llama2Vec: Unsupervised Adaptation of Large Language Models for Dense Retrieval”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 3490–3500. Available at: https://doi.org/10.18653/v1/2024.acl-long.191.
Vancouver
1. Liu Z, Li C, Xiao S, Shao Y, Lian D (2024) Llama2Vec: Unsupervised Adaptation of Large Language Models for Dense Retrieval. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 3490–3500

BibTeX

@inproceedings{li-etal-2024-llama2vec,
    title = "{L}lama2{V}ec: Unsupervised Adaptation of Large Language Models for Dense Retrieval",
    author = "Liu, Zheng  and
      Li, Chaofan  and
      Xiao, Shitao  and
      Shao, Yingxia  and
      Lian, Defu",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.191/",
    doi = "10.18653/v1/2024.acl-long.191",
    pages = "3490--3500"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/