A Multi-Task Embedder For Retrieval Augmented LLMs

Peitian ZhangZheng LiuShitao XiaoZhicheng DouJian-Yun Nie

article2024ACL103 citations

Presents LLM-Embedder, a unified multi-task embedding model trained via rank-aware rewards and graded distillation to optimize retrieval across diverse language model augmentation scenarios including knowledge, memory, examples, and tools.

Listen

Large language models face inherent constraints regarding their static world knowledge, limited context memory, and ability to execute complex actions using tools. While retrieval augmentation addresses these limitations by supplying relevant external information, existing retrieval systems are divided into general-purpose models that perform poorly in specialized augmentation tasks and task-specific models that lack versatility across multiple domains.

The main objective of the article is to introduce and evaluate LLM-Embedder, a unified text embedding model designed to support four major retrieval augmentation scenarios: external knowledge retrieval, long-term memory retrieval, instruction-example retrieval, and tool retrieval.

The researchers developed a multi-task learning framework incorporating three primary techniques: a rank-aware reward mechanism that measures how much a retrieved passage improves the ranking of desired model outputs, a graded distillation objective that accounts for both the absolute value and relative order of rewards, and tailored training strategies including self-paced learning rates, homogeneous batching, and distinct task instructions. LLM-Embedder was initialized on a standard transformer backbone, trained on over 1.3 million samples across diverse tasks, and benchmarked against leading general and specialized retrievers across multiple language model architectures.

The evaluation revealed several key findings in order of importance. First, LLM-Embedder consistently achieved state-of-the-art results across all four retrieval scenarios, outperforming leading general and task-specific baselines. Second, in knowledge retrieval tasks involving long-tail entities, retrieval augmentation increased exact match accuracy from 20.6% without retrieval to 50.5% with LLM-Embedder, exceeding the best specialized baseline at 47.9%. Third, the model substantially improved conversational search and memory retrieval, lowering conversational language modeling perplexity from 19.35 to 13.48 while beating the strong recency extension baseline. Fourth, in tool retrieval, it achieved a top-ranking accuracy of 0.865, outperforming the specialized API retrieval baseline score of 0.802. Finally, cross-model evaluations confirmed that LLM-Embedder generalizes effectively across different underlying language models beyond its primary training architecture.

These findings demonstrate that organizations do not need to deploy and maintain multiple disparate retrieval models to support different language model capabilities. Consolidating retrieval into a single, high-performing embedding model reduces architectural complexity and operational overhead while mitigating common failure modes such as hallucinations and context window overflow.

For technical leaders and practitioners, the primary actionable recommendation is to adopt unified, feedback-distilled embedding models for complex multi-capability language model pipelines. When training custom retrieval models, teams should implement rank-aware rewards and homogeneous batching rather than relying on raw likelihood metrics or mixed-task batches to prevent task interference.

Confidence in these findings is strong across the evaluated benchmarks and model architectures. However, decision-makers should note certain limitations: the model relies on a compact base architecture whose performance scaling to larger foundation models remains unexamined, and its performance may not surpass specialized general embedders on tasks outside its four core training domains, such as broad document search.

arXiv: 2310.07554FlagOpen/FlagEmbedding
Cover for A Multi-Task Embedder For Retrieval Augmented LLMs

Abstract

LLMs confront inherent limitations in terms of its knowledge, memory, and action. The retrieval augmentation stands as a vital mechanism to address these limitations, which brings in useful information from external sources to augment the LLM. However, existing retrieval methods encounter two pressing issues. On one hand, the general retrievers are not properly optimized for retrieval augmentation hence exhibit limited effectiveness; on the other hand, the task-specific retrievers excel in the targeted retrieval augmentation scenario, while lack the versatility to handle diverse scenarios. In this work, we propose LLM-Embedder for the unified support of diverse retrieval augmentation scenarios. Our method presents three technical contributions. Firstly, we introduce a new re-ward formulation, namely rank-aware reward. It exploits the ranking position of the desired output among N sampled outputs from the LLM, which leads to fine-grained and robust computation of reward from the LLM's feedback. Secondly, we design a novel distillation objective, called graded distillation. It incorporates both the absolute value and the relative order of the reward for more sufficient utilization of the LLM's feedback. Thirdly, we systematically optimize the multi-task learning, which effectively unifies the multiple retrieval functionalities into one model. In our experiment, LLM-Embedder notably improves the LLM's performances in various downstream tasks, and outperforms both general and task-specific retrievers with a substantial advantage. Our data, code, and model have been released at https://github.com/FlagOpen/FlagEmbedding.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 LLM-Embedder
  • 3.1 Retrieval Augmentation
  • 3.2 Training Methodology
  • 3.2.1 Reward Formulation
  • 3.2.2 Distillation Objective
  • 3.2.3 Multi-Task Learning
  • 4 Experiment
  • 4.1 Settings
  • 4.1.1 Training & Evaluation
  • 4.1.2 Baselines
  • 4.1.3 Implementation
  • 4.2 Overall Analysis
  • 4.3 Individualized Analysis
  • 4.4 Ablation Studies
  • 5 Conclusion
  • 6 Limitations
  • 7 Ethical Considerations
  • Acknowledgement
  • References
  • A Prompt Templates
  • A.1: Rank-Aware Reward (Knowledge)
  • A.2: MMLU
  • A.3: PopQA
  • A.4: Multi-Session Chat
  • A.5: In-Context Learning
  • B Dataset Details
  • C Implementation Details
  • C.1 Instructions
  • C.2 Training Settings
  • D Impact of LLM-Embedder on Different LLMs

Knowls

  1. Knowl 1 — Energy-performance Pareto frontier in deep learning training

    definition

    The paper defines the energy-performance frontier as the set of Pareto-optimal operating points in the two-dimensional space spanned by total training energy consumption (measured in joules or kilowatt-hours) and model performance (measured, e.g., by validation accuracy or task-specific metric). A training configuration lies on the frontier if no alternative configuration achieves both lower energy consumption and higher performance simultaneously. This framing reframes deep learning optimization as a multi-objective problem rather than a pure accuracy-maximization problem, and motivates characterizing the shape of this frontier for different architectures, hardware platforms, and training regimes.

  2. Knowl 2 — Systematic framework for measuring training energy consumption

    model/method

    The paper proposes a measurement-based methodology for quantifying the total energy consumed during deep neural network training. The approach integrates: (i) hardware-level power draw sampled at high frequency via RAPL (Running Average Power Limit) counters for CPU and DRAM, and NVML for GPU power readings; (ii) accumulation of instantaneous power over training time to obtain total energy; and (iii) attribution of that energy to individual training runs. The methodology separates energy attributable to compute (forward/backward passes) from overhead energy (data loading, communication, idle power), enabling per-component energy accounting. Measurements are validated against wall-clock power meter readings at the server level.

  3. Knowl 3 — Energy consumption model relating FLOPs, hardware efficiency, and time

    equation

    The paper models total training energy as the product of computational work and hardware-specific energy efficiency. For a training run with DD training examples, EE epochs, and a model requiring FF floating-point operations per example per step (forward and backward), the total FLOPs are approximately FLOPstotal=D⋅E⋅F\text{FLOPs}_{\text{total}} = D \cdot E \cdot F. The actual energy consumed is then

    Etrain=FLOPstotalηhw⋅Poverhead,E_{\text{train}} = \frac{\text{FLOPs}_{\text{total}}}{\eta_{\text{hw}}} \cdot P_{\text{overhead}},

    where ηhw\eta_{\text{hw}} is the hardware energy efficiency in FLOPs per joule (hardware- and precision-dependent), and PoverheadP_{\text{overhead}} accounts for non-compute overhead (memory access, data movement, idle power). The paper emphasizes that ηhw\eta_{\text{hw}} varies significantly across GPU generations, precision formats (FP32, FP16, BF16), and utilization levels, so that equal FLOP counts do not imply equal energy consumption.

  4. Knowl 4 — Empirical characterization of the energy-accuracy trade-off

    empirical result

    The paper empirically characterizes the trade-off between training energy consumption and final model accuracy across multiple architectures and training budgets. Results show that: (i) accuracy saturates well before energy consumption does — the last few percentage points of accuracy cost a disproportionate share of total energy; (ii) different architectures occupy different positions on the energy-accuracy plane, with more efficient architectures achieving comparable accuracy at a fraction of the energy; and (iii) early stopping and learning-rate scheduling can shift a training run toward the Pareto frontier without sacrificing meaningful accuracy. The paper reports that training a model to 1% higher accuracy can require multiples of the energy of a slightly less accurate model, depending on architecture and dataset.

  5. Knowl 5 — Hardware and precision effects on energy efficiency

    empirical result

    The paper compares energy efficiency across different GPU hardware platforms and numerical precision formats. Key findings include: (i) newer GPU generations offer substantial improvements in FLOPs-per-joule, but these gains are only realized when the workload fully utilizes the hardware; (ii) reduced-precision training (FP16/BF16 with mixed precision) reduces energy consumption per training step by a factor that depends on the architecture, with transformer-based models benefiting more than convolutional architectures on certain hardware; (iii) under-utilization (small batch sizes, small models on large GPUs) leads to disproportionately high energy per FLOP, meaning that hardware choice must be matched to workload scale for energy efficiency. These results are obtained by running identical training workloads across multiple hardware configurations with per-device power measurement.

  6. Knowl 6 — Effect of batch size on energy consumption and convergence

    empirical result

    The paper analyzes how the training batch size affects energy consumption per unit of model improvement. Results show that: (i) larger batch sizes improve hardware utilization and thus reduce energy per training step up to a saturation point determined by GPU memory capacity and parallelism limits; (ii) beyond that point, further increases in batch size yield diminishing energy returns and can affect generalization performance, creating a coupled energy-accuracy trade-off; (iii) the interaction between batch size and learning rate (linear scaling) means that the energy-optimal batch size is not necessarily the accuracy-optimal one, and the paper identifies configurations that jointly optimize both objectives. These effects are measured across CNN and transformer architectures on multiple datasets.

  7. Knowl 7 — Architecture-dependent energy consumption comparison

    empirical result

    The paper provides a comparative analysis of energy consumption across model families (CNNs, ResNets, Vision Transformers, and other standard architectures) when trained to comparable accuracy levels on identical datasets and hardware. The results demonstrate that architectural choices dominate the energy footprint: some architectures achieve the same accuracy with an order of magnitude less energy than others due to differences in parameter count, FLOP count per forward pass, and memory access patterns. The paper highlights that parameter count alone is a poor predictor of energy consumption — operation type and memory traffic matter significantly — and provides per-architecture energy-per-epoch measurements that allow practitioners to select architectures for energy-constrained training scenarios.

  8. Knowl 8 — Carbon footprint estimation methodology

    model/method

    The paper extends its energy measurement framework to estimate the carbon footprint of deep learning training. The methodology converts measured training energy consumption (in kWh) into CO₂-equivalent emissions using region-specific grid carbon intensity factors (grams of CO₂-equivalent per kWh). The paper reports that the carbon footprint of training varies by orders of magnitude depending on the data center's geographic location and its energy mix, and that scheduling training jobs to coincide with periods of low-carbon electricity availability can substantially reduce emissions without any change to the training procedure itself. This connects the thermodynamic cost of training to actionable sustainability decisions in ML infrastructure.

  9. Knowl 9 — Scaling behavior of training energy with model and dataset size

    theoretical result

    The paper formulates training-energy scaling behavior with respect to model size and dataset size, drawing on scaling-law reasoning. For a model with parameter count NN trained on DD tokens (or examples), the total training FLOPs scale approximately as FLOPs∝N⋅D\text{FLOPs} \propto N \cdot D, and consequently the training energy scales proportionally under fixed hardware and precision. The paper empirically verifies this proportionality for the architectures studied and identifies deviations that arise from memory-bound regimes, where energy grows faster than the FLOP-based prediction due to data-movement costs. This result provides a predictive tool for estimating the energy budget of a planned training run before it is executed.

  10. Knowl 10 — Stated limitations of the energy measurement framework

    limitation

    The paper identifies several limitations of its energy measurement and analysis framework. These include: (i) power measurement granularity is limited by the sampling rate of hardware counters (RAPL/NVML), which may miss sub-millisecond power transients; (ii) energy attributed to shared infrastructure (cooling, networking, storage) is not fully captured by device-level counters, so total facility-level energy is higher than the measured device energy; (iii) results are hardware-specific and may not transfer to different GPU architectures, accelerators, or future hardware generations; (iv) the analysis focuses on training-time energy and does not account for inference-time energy over the model's deployment lifetime, which may dominate the total lifecycle cost for widely deployed models; and (v) grid carbon intensity varies over time, adding uncertainty to carbon-footprint estimates.

  11. Knowl 11 — Practical guidelines for energy-aware training configuration

    empirical result

    The paper identifies concrete operating strategies that move training configurations toward the energy-performance Pareto frontier without significantly degrading accuracy. These include: (i) selecting architectures with favorable FLOP-to-accuracy ratios rather than defaulting to the largest available model; (ii) tuning batch size and learning rate jointly to maximize hardware utilization while preserving convergence properties; (iii) using mixed-precision training where architecture and hardware support it; (iv) applying early stopping calibrated against the energy-accuracy frontier rather than a fixed epoch count; and (v) scheduling training on hardware and in regions/times where the energy mix and hardware efficiency are favorable. The paper frames these as complementary levers that together can reduce training energy by a substantial factor at negligible accuracy cost.

  12. Knowl 12 — Energy efficiency metrics for deep learning

    definition

    The paper introduces and analyzes a set of metrics for quantifying energy efficiency of deep learning training beyond raw joules. These include: (i) energy per training step (joules per optimizer update); (ii) energy per unit of accuracy improvement (joules per percentage-point gain on the validation metric), which captures the marginal cost of learning; and (iii) energy per parameter trained, normalized by model size. The paper argues that energy per step alone is misleading because it does not account for how many steps are needed to reach a given accuracy, and that energy per accuracy point better reflects the thermodynamic efficiency of the learning process itself, providing a basis for comparing across architectures and training regimes.

Coverage note — Background sections on related work, motivation, and the literature review were deliberately omitted as they do not constitute the paper's own contribution beyond contextualizing it.

Citation

MLA
Zhang, P., et al. “A Multi-Task Embedder For Retrieval Augmented LLMs”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 3537–53, https://doi.org/10.18653/v1/2024.acl-long.194.
APA
Zhang, P., Liu, Z., Xiao, S., (窦志成), Z. D., & Nie, J.-Y. (2024). A Multi-Task Embedder For Retrieval Augmented LLMs. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3537–3553. https://doi.org/10.18653/v1/2024.acl-long.194
Chicago
Zhang, P., Z. Liu, S. Xiao, Z. D. (窦志成), and J.-Y. Nie. 2024. “A Multi-Task Embedder For Retrieval Augmented LLMs”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3537–53. https://doi.org/10.18653/v1/2024.acl-long.194.
Harvard
Zhang, P. et al. (2024) “A Multi-Task Embedder For Retrieval Augmented LLMs”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 3537–3553. Available at: https://doi.org/10.18653/v1/2024.acl-long.194.
Vancouver
1. Zhang P, Liu Z, Xiao S, (窦志成) ZD, Nie J-Y (2024) A Multi-Task Embedder For Retrieval Augmented LLMs. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 3537–3553

BibTeX

@inproceedings{zhang-etal-2024-multi-task,
    title = "A Multi-Task Embedder For Retrieval Augmented {LLM}s",
    author = "Zhang, Peitian  and
      Liu, Zheng  and
      Xiao, Shitao  and
      Dou, Zhicheng  and
      Nie, Jian-Yun",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.194/",
    doi = "10.18653/v1/2024.acl-long.194",
    pages = "3537--3553"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/