An Efficient Memory-Augmented Transformer for Knowledge-Intensive NLP Tasks
Yuxiang WuYu ZhaoBaotian HuPasquale MinerviniPontus StenetorpSebastian Riedel
Proposes an efficient memory-augmented transformer that stores external question-answer knowledge in a fast key-value memory queried during a single forward pass, matching or exceeding the accuracy of retrieval-augmented models on open-domain NLP benchmarks while processing up to a thousand queries per second.
Modern natural language processing tasks—such as answering open-domain questions or engaging in dialogue—require access to vast amounts of factual knowledge. Current systems face a difficult trade-off: internal parametric models store facts directly within model weights and run very quickly but suffer from limited capacity and factual hallucinations, whereas retrieval-augmented models search external knowledge sources for high accuracy but suffer from severe computational overhead and slow inference speeds.
The article introduces and evaluates the Efficient Memory-Augmented Transformer (EMAT). The main objective is to demonstrate an architecture that combines the predictive accuracy of retrieval-augmented systems with the high inference throughput of compact parametric models.
To achieve this, the authors built an architecture based on a standard text-to-text transformer (T5-base) augmented with an external key-value memory derived from a corpus of 14 million to 65 million question-answer pairs (PAQ). The system computes query representations early in the transformer's encoder layers, performs an efficient maximum inner product search across the memory using host system memory (RAM and CPU), and integrates the retrieved dense key-value vectors in later layers during a single forward pass. The model was trained using multi-task pre-training objectives—including auto-encoding and generative answer integration—and evaluated across open-domain question answering benchmarks (NaturalQuestions, TriviaQA, WebQuestions), open-domain dialogue (Wizard-of-Wikipedia), and long-form question answering (ELI5).
The evaluation produced several key findings. First, augmenting a standard parametric baseline (T5-base) with EMAT yielded substantial accuracy improvements, boosting Exact Match scores by 18.5 percentage points on NaturalQuestions (from 25.8 to 44.3) and by 20.0 percentage points on TriviaQA (from 24.4 to 44.4). Second, EMAT maintained high inference throughput, processing 1,000 to 1,200 questions per second on NaturalQuestions, which is orders of magnitude faster than conventional retrieve-and-read models like FiD-base. Third, compared to competing memory architectures like QAMAT, EMAT ran 4.2 to 5 times faster while requiring significantly fewer hardware resources (operating on a single GPU instead of 32 specialized TPU chips). Finally, the approach successfully generalized to open-ended dialogue and long-form generation, outperforming retrieval-augmented baselines such as RAG and BART+DPR in both accuracy and speed.
These findings indicate that organizations do not need to choose between expensive, slow retrieval pipelines and inaccurate, parameter-heavy language models. By offloading memory retrieval to high-speed CPU search and overlapping it with model computation, production systems can serve knowledge-intensive requests with high factual reliability and low latency, substantially reducing the operational computing budget.
Decision-makers considering knowledge-intensive language systems should evaluate memory-augmented architectures as an efficient alternative to scaling model parameters or deploying multi-stage retrieval pipelines. Future work and pilot implementations should focus on automating the weakly-supervised retrieval training via end-to-end gradient methods, expanding memory sources beyond question-answer pairs to include diverse structured and unstructured knowledge bases, and applying memory caching techniques in memory-constrained environments.
Confidence in these findings is high for standard question answering and knowledge benchmarks. However, stakeholders should note two primary limitations: storing dense memories requires substantial system RAM (approximately 300 GB in the evaluated setup), and the fine-tuning process relies on task-specific heuristics for weak supervision, which may require adjustment when adapting the framework to specialized enterprise domains.
- Paper: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, Patrick Lewis et al. (2020). Read this foundational retrieval-augmented generation work first to understand the retrieve-and-generate approach EMAT seeks to make more efficient.
- Paper: Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering, Gautier Izacard et al. (2021). Its Fusion-in-Decoder retrieve-and-read pipeline provides the key comparison point for understanding EMAT’s single-pass memory integration and throughput gains.
- Paper: REALM: Retrieval-Augmented Language Model Pre-Training, Kelvin Guu et al. (2020). REALM establishes how a learned retriever and external document memory can support knowledge-intensive language modeling, clarifying the retrieval foundations EMAT adapts.
No sufficiently relevant recommendations were found.
