An Efficient Memory-Augmented Transformer for Knowledge-Intensive NLP Tasks

Yuxiang WuYu ZhaoBaotian HuPasquale MinerviniPontus StenetorpSebastian Riedel

article2022EMNLP69 citationsOutstanding Paper Award

Proposes an efficient memory-augmented transformer that stores external question-answer knowledge in a fast key-value memory queried during a single forward pass, matching or exceeding the accuracy of retrieval-augmented models on open-domain NLP benchmarks while processing up to a thousand queries per second.

Listen

Modern natural language processing tasks—such as answering open-domain questions or engaging in dialogue—require access to vast amounts of factual knowledge. Current systems face a difficult trade-off: internal parametric models store facts directly within model weights and run very quickly but suffer from limited capacity and factual hallucinations, whereas retrieval-augmented models search external knowledge sources for high accuracy but suffer from severe computational overhead and slow inference speeds.

The article introduces and evaluates the Efficient Memory-Augmented Transformer (EMAT). The main objective is to demonstrate an architecture that combines the predictive accuracy of retrieval-augmented systems with the high inference throughput of compact parametric models.

To achieve this, the authors built an architecture based on a standard text-to-text transformer (T5-base) augmented with an external key-value memory derived from a corpus of 14 million to 65 million question-answer pairs (PAQ). The system computes query representations early in the transformer's encoder layers, performs an efficient maximum inner product search across the memory using host system memory (RAM and CPU), and integrates the retrieved dense key-value vectors in later layers during a single forward pass. The model was trained using multi-task pre-training objectives—including auto-encoding and generative answer integration—and evaluated across open-domain question answering benchmarks (NaturalQuestions, TriviaQA, WebQuestions), open-domain dialogue (Wizard-of-Wikipedia), and long-form question answering (ELI5).

The evaluation produced several key findings. First, augmenting a standard parametric baseline (T5-base) with EMAT yielded substantial accuracy improvements, boosting Exact Match scores by 18.5 percentage points on NaturalQuestions (from 25.8 to 44.3) and by 20.0 percentage points on TriviaQA (from 24.4 to 44.4). Second, EMAT maintained high inference throughput, processing 1,000 to 1,200 questions per second on NaturalQuestions, which is orders of magnitude faster than conventional retrieve-and-read models like FiD-base. Third, compared to competing memory architectures like QAMAT, EMAT ran 4.2 to 5 times faster while requiring significantly fewer hardware resources (operating on a single GPU instead of 32 specialized TPU chips). Finally, the approach successfully generalized to open-ended dialogue and long-form generation, outperforming retrieval-augmented baselines such as RAG and BART+DPR in both accuracy and speed.

These findings indicate that organizations do not need to choose between expensive, slow retrieval pipelines and inaccurate, parameter-heavy language models. By offloading memory retrieval to high-speed CPU search and overlapping it with model computation, production systems can serve knowledge-intensive requests with high factual reliability and low latency, substantially reducing the operational computing budget.

Decision-makers considering knowledge-intensive language systems should evaluate memory-augmented architectures as an efficient alternative to scaling model parameters or deploying multi-stage retrieval pipelines. Future work and pilot implementations should focus on automating the weakly-supervised retrieval training via end-to-end gradient methods, expanding memory sources beyond question-answer pairs to include diverse structured and unstructured knowledge bases, and applying memory caching techniques in memory-constrained environments.

Confidence in these findings is high for standard question answering and knowledge benchmarks. However, stakeholders should note two primary limitations: storing dense memories requires substantial system RAM (approximately 300 GB in the evaluated setup), and the fine-tuning process relies on task-specific heuristics for weak supervision, which may require adjustment when adapting the framework to specialized enterprise domains.

arXiv: 2210.16773uclnlp/EMAT

No sufficiently relevant recommendations were found.

Cover for An Efficient Memory-Augmented Transformer for Knowledge-Intensive NLP Tasks

Abstract

Access to external knowledge is essential for many natural language processing tasks, such as question answering and dialogue. Existing methods often rely on a parametric model that stores knowledge in its parameters, or use a retrieval-augmented model that has access to an external knowledge source. Parametric and retrieval-augmented models have complementary strengths in terms of computational efficiency and predictive accuracy. To combine the strength of both approaches, we propose the Efficient Memory-Augmented Transformer (EMAT) – it encodes external knowledge into a key-value memory and exploits the fast maximum inner product search for memory querying. We also introduce pre-training tasks that allow EMAT to encode informative key-value representations, and to learn an implicit strategy to integrate multiple memory slots into the transformer. Experiments on various knowledge-intensive tasks such as question answering and dialogue datasets show that, simply augmenting parametric models (T5-base) using our method produces more accurate results (e.g., 25.8 → 44.3 EM on NQ) while retaining a high throughput (e.g., 1000 queries/s on NQ). Compared to retrieval-augmented models, EMAT runs substantially faster across the board and produces more accurate results on WoW and ELI5.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Efficient Memory-Augmented Transformer
  • 3.1 Key-Value Memory
  • 3.2 Memory Retrieval
  • 3.3 Key-Value Integration
  • 4 Training Pipeline of EMAT
  • 4.1 Pre-Training
  • 4.2 Fine-Tuning on Downstream Tasks
  • 4.3 Inference
  • 5 Experiments
  • 5.1 Experimental Setup
  • Pre-Training and Fine-Tuning Configurations
  • 5.2 Open-Domain Question Answering
  • 5.3 Generalisation to Open-Domain Dialogue and Long-Form QA
  • 6 Analysis
  • 6.1 Ablation Study
  • 6.2 Qualitative Analysis
  • Wizard-of-Wikipedia
  • Dialogue History
  • Response
  • Retrieved Key-Values
  • 7 Conclusions
  • Limitations
  • Acknowledgements
  • References
  • A Data Efficiency
  • B Hyperparameters
  • Natural Questions
  • Wizard-of-Wikipedia

Knowls

  1. Knowl 1 — EMAT stores question–answer knowledge as dense key–value memory

    model/method

    Efficient Memory-Augmented Transformer (EMAT) augments a T5-style encoder–decoder with an external memory of question–answer pairs: each question is represented by a key and its answer by a value. For a prefix of length PP and encoder hidden size hh, the key encoder takes a prefixed question’s hidden states at layer lkl_k, applies a convolutional layer, and uses the prefix positions as the key representation in RP×h\mathbb{R}^{P\times h}. The value representation is the prefix’s hidden states at encoder layer lvl_v, also in RP×h\mathbb{R}^{P\times h}. The same encoder parameters are used for encoding memory questions and input questions. At inference, the input produces a query, retrieved key–value representations are integrated into the encoder, and the decoder generates the output conditioned on the input and the augmented encoder states. The architecture diagram on page 3 depicts this flow from stored question–answer vectors through retrieval and encoder integration to generation.

  2. Knowl 2 — Mean-pooled inner-product search retrieves memory entries efficiently

    model/method

    EMAT forms a query from the input question using the same prefixed-question encoder used for memory keys. For each query or key matrix with PP prefix-position vectors, it averages those vectors to obtain a single retrieval vector. The similarity of query qq to key kk is their inner product:

    qˉ=1P∑j=1Pqj,kˉ=1P∑j=1Pkj,sim⁡(q,k)=⟨qˉ,kˉ⟩.\bar q=\frac{1}{P}\sum_{j=1}^{P}q_j,\qquad \bar k=\frac{1}{P}\sum_{j=1}^{P}k_j,\qquad \operatorname{sim}(q,k)=\langle\bar q,\bar k\rangle.

    Here qj,kj∈Rhq_j,k_j\in\mathbb{R}^{h} are prefix-position vectors, and PP and hh are the prefix length and hidden size. Maximum inner product search (MIPS) selects the highest-scoring memory entries. At inference EMAT uses a FAISS-generated HNSW index; the search can run on CPU, and, when the retrieval layer lkl_k precedes the key-integration layer lcl_c, it can overlap with the intervening encoder computation. This permits retrieval and generation in one encoder pass without adding GPU memory for the index.

  3. Knowl 3 — Retrieved keys and values enter the encoder at separate layers

    model/method

    Given the top retrieved question–answer pairs ordered by query–key similarity, EMAT concatenates their key representations and prepends them to the encoder hidden states at layer lcl_c. It adds relative positional encodings to distinguish the retrieved keys. At layer lvl_v, the corresponding value representations are concatenated and added at the positions occupied by their associated keys. The encoder then continues through its remaining layers, and the decoder generates from the resulting encoder states. This arrangement lets the model use both retrieved questions and answers as dense representations rather than inserting retrieved text passages.

  4. Knowl 4 — Three objectives pre-train memory representations and their use in generation

    model/method

    EMAT is initialized from T5-base; its prefix embeddings and key convolutional layer are initialized from scratch. Pre-training uses PAQ-L1, a 14-million-pair subset of PAQ, with two auto-encoding tasks and a generation task. Key auto-encoding (KAE) trains the decoder to reconstruct question tokens xix_i from key representation kk; value auto-encoding (VAE) trains it to reconstruct answer tokens yiy_i from value representation vv. For generation, each PAQ question xx is paired with 10 relevant question–answer entries retrieved by RePAQ, whose memory representations are Kx′K'_x and Vx′V'_x; the model learns to generate the answer yy from xx and those representations. The token-level objectives are:

    LKAE=−∑ilog⁡P(xi∣k,x<i),LVAE=−∑ilog⁡P(yi∣v,y<i),\mathcal{L}_{\mathrm{KAE}}=-\sum_i\log P(x_i\mid k,x_{<i}),\qquad \mathcal{L}_{\mathrm{VAE}}=-\sum_i\log P(y_i\mid v,y_{<i}), LGen=−∑ilog⁡P(yi∣x,Kx′,Vx′,y<i).\mathcal{L}_{\mathrm{Gen}}=-\sum_i\log P(y_i\mid x,K'_x,V'_x,y_{<i}).

    Here xx and yy are question and answer token sequences, and x<ix_{<i} and y<iy_{<i} denote preceding tokens. The pre-training objective combines both auto-encoding losses and the generation loss; the reported configuration uses weight 0.5 for the auto-encoding objective and 1.0 for generation. Pre-training runs for five epochs. For 10% of examples, the example itself is retained among the relevant pairs so the model is exposed to its own question–answer entry.

  5. Knowl 5 — Weak answer matching supervises retrieval during downstream fine-tuning

    algorithm

    For each downstream training example with input xx and target sequence yy, EMAT ranks retrieved memory pairs by inner-product similarity. Retrieved pairs whose answers lexically match the target are treated as positive retrieval examples. For short outputs such as ODQA answers, the retrieved answer is matched against the target answer; for long-form outputs, the target is lower-cased and stop words are removed, and a retrieved answer is positive if it occurs in that normalized target. EMAT samples a positive pair from the positive set with probability proportional to the exponential of its query–key similarity, and samples mm nonmatching pairs as negatives. A contrastive retrieval loss increases the positive pair’s similarity relative to those negatives. The generation loss is token-level negative log-likelihood conditioned on xx, the retrieved pairs, and preceding output tokens; the fine-tuning objective is the sum of retrieval and generation losses.

    To avoid re-encoding a large memory after every update, EMAT freezes the memory for an epoch, retrieves the top nn pairs for each training example against that cached memory, and re-encodes the knowledge source to refresh the memory at the end of the epoch. The downstream configuration uses a memory-cache size of 384 and sets retrieval and generation loss weights to 1.

  6. Knowl 6 — ODQA accuracy improves substantially over a same-size parametric baseline at high throughput

    empirical result

    EMAT was evaluated on NaturalQuestions (NQ), TriviaQA (TQA), and WebQuestions (WQ), using exact match (EM) and NQ queries per second (Q/s). FKSV uses encoder layers lk=3l_k=3, lc=3l_c=3, lv=7l_v=7; SKSV uses lk=3l_k=3, lc=10l_c=10, lv=11l_v=11. The following comparison reports the paper’s measured scores and throughput; QAMAT’s throughput was measured on 32 TPU-v3s with 1024 GB TPU memory, whereas EMAT throughput was measured on one 40 GB A100 GPU.

    Model NQ EM NQ Q/s TQA EM WQ EM
    T5-base 25.8 1600 24.4 26.6
    RePAQ-base 40.9 1400 39.7 29.4
    QAMAT 44.7 240* 48.0 39.4
    FiD-base 48.2 3.7 65.0 32.4
    EMAT-FKSV 44.3 1000 44.4 36.7
    EMAT-SKSV 43.3 1200 43.7 33.2

    Against T5-base, EMAT-FKSV gains 18.5, 20.0, and 10.1 EM points on NQ, TQA, and WQ, respectively, while achieving 1000 NQ Q/s. EMAT also exceeds RePAQ-base on all three listed EM scores. Retrieval-augmented systems can score higher, but the table shows much lower throughput for FiD-base; the paper reports EMAT as about 4.2 times faster than QAMAT for FKSV and five times faster for SKSV, with the hardware difference noted above.

  7. Knowl 7 — EMAT generalizes to dialogue and long-form question answering

    empirical result

    On Wizard-of-Wikipedia (WoW), EMAT was evaluated with F1, ROUGE-L (R-L), and generated utterances per second (U/s). On ELI5, it was evaluated with F1, R-L, and queries per second (Q/s). The datasets test knowledge-intensive generation beyond short-form ODQA; throughput units differ between tasks.

    Task Model F1 R-L Throughput Unit
    WoW T5-base 13.53 12.40 160 U/s
    WoW BART + DPR 15.19 13.23 0.7 U/s
    WoW RAG 13.11 11.57 3.4 U/s
    WoW EMAT-FKSV 15.78 14.73 141 U/s
    WoW EMAT-SKSV 15.35 14.68 150 U/s
    ELI5 T5-base 16.01 19.08 76 Q/s
    ELI5 BART-large 19.23 20.55 30 Q/s
    ELI5 RAG 14.51 14.05 0.4 Q/s
    ELI5 EMAT-FKSV 18.42 20.61 67 Q/s
    ELI5 EMAT-SKSV 19.03 20.91 71 Q/s

    Both EMAT configurations outperform T5-base on both reported metrics for both tasks while retaining similar throughput to the parametric baselines. EMAT also exceeds the listed retrieval-augmented systems on both metrics in WoW and ELI5. In particular, the paper reports EMAT-SKSV as 4.52 F1 points and 6.86 R-L points above RAG on ELI5, and more than 160 times faster.

  8. Knowl 8 — Removing either pre-training component sharply reduces ODQA performance

    empirical result

    The ablation compares EMAT-FKSV with variants trained without downstream fine-tuning, without the question and answer auto-encoding objectives, without the generation pre-training objective, or without any pre-training objectives. Scores are exact match on NQ, TQA, and WQ.

    Configuration NQ EM TQA EM WQ EM
    EMAT-FKSV 44.3 44.4 36.7
    Without fine-tuning 30.6 32.4 25.6
    Without auto-encoding tasks 28.5 34.6 12.9
    Without generation task 28.7 24.7 31.4
    Without all pre-training tasks 27.1 17.7 6.0

    Removing auto-encoding particularly harms WQ (36.7 to 12.9 EM), while removing generation pre-training particularly harms TQA (44.4 to 24.7 EM). Without pre-training, all three scores are below the full model, and the paper reports that the resulting model also underperforms the T5-base baseline.

  9. Knowl 9 — Qualitative cases show generation can select, ignore, and combine retrieved evidence

    empirical result

    The paper’s examples suggest that EMAT’s decoder does more than return the highest-ranked retrieved answer. In an NQ example asking who plays the judge in Drop Dead Diva, the correct answer, Lex Medlin, appears in a retrieved entry that is not the first result, and EMAT generates Lex Medlin. In a Wizard-of-Wikipedia example about when jazz originated, EMAT generates late 19th century and New Orleans, while T5-base gives the incorrect time late 18th century. The authors also report examples in which EMAT ignores irrelevant retrieved pairs and uses information from retrieved question/key representations, not only answer/value representations. These cases are qualitative evidence of selection and generation over dense memory entries, not a quantitative guarantee.

  10. Knowl 10 — Memory size and weak-supervision requirements limit deployment

    limitation

    The authors identify two deployment constraints. First, the dense key–value memory requires about 300 GB of CPU RAM in their setup. They suggest an LRU cache as an option when less RAM is available. Second, retriever fine-tuning depends on weak supervision derived from lexical matching between retrieved answers and task targets; the matching rule may need to change across downstream tasks. The paper leaves end-to-end training of the retriever using gradients from the decoder as future work.

Coverage note — The supplementary plots on page 12 show increasing retrieval/generation accuracy with more retrieved pairs and larger PAQ subsets, but are omitted as a secondary sensitivity analysis whose exact plotted values are not tabulated.

References

  1. 1.Jonathan Berant, Andrew Chou, Roy Frostig, and Percy Liang. 2013. Semantic parsing on Freebase from question-answer pairs. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, pages 1533–1544, Seattle, Washington, USA. Association for Computational Linguistics.
  2. 2.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  3. 3.Danqi Chen, Adam Fisch, Jason Weston, and Antoine Bordes. 2017. Reading Wikipedia to answer open-domain questions. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1870–1879, Vancouver, Canada. Association for Computational Linguistics.
  4. 4.Wenhu Chen, Pat Verga, Michiel de Jong, John Wieting, and William Cohen. 2022. Augmenting pre-trained language models with qa-memory for open-domain question answering. CoRR, abs/2204.04581.
  5. 5.Rajarshi Das, Patrick Lewis, Sewon Min, June Thai, and Manzil Zaheer, editors. 2022. Proceedings of the 1st Workshop on Semiparametric Methods in NLP: Decoupling Logic from Knowledge. Association for Computational Linguistics, Dublin, Ireland and Online.
  6. 6.Emily Dinan, Stephen Roller, Kurt Shuster, Angela Fan, Michael Auli, and Jason Weston. 2019. Wizard of wikipedia: Knowledge-powered conversational agents. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net.
  7. 7.Angela Fan, Claire Gardent, Chloé Braud, and Antoine Bordes. 2021. Augmenting transformers with KNN-based composite memory for dialog. Transactions of the Association for Computational Linguistics, 9:82–99.
  8. 8.Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, and Michael Auli. 2019. ELI5: Long form question answering. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3558–3567, Florence, Italy. Association for Computational Linguistics.
  9. 9.Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. 2021. Transformer feed-forward layers are key-value memories. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 5484–5495, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  10. 10.Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. 2020. REALM: retrieval-augmented language model pre-training. CoRR, abs/2002.08909.
  11. 11.Gautier Izacard and Edouard Grave. 2021. Leveraging passage retrieval with generative models for open domain question answering. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 874–880, Online. Association for Computational Linguistics.
  12. 12.Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2019. Billion-scale similarity search with GPUs. IEEE Transactions on Big Data, 7(3):535–547.
  13. 13.Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017. TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1601–1611, Vancouver, Canada. Association for Computational Linguistics.
  14. 14.Norman P. Jouppi, Cliff Young, Nishant Patil, David A. Patterson, Gaurav Agrawal, Raminder Bajwa, Sarah Bates, Suresh Bhatia, Nan Boden, Al Borchers, Rick Boyle, Pierre-luc Cantin, Clifford Chao, Chris Clark, Jeremy Coriell, Mike Daley, Matt Dau, Jeffrey Dean, Ben Gelb, Tara Vazir Ghaemmaghami, Rajendra Gottipati, William Gulland, Robert Hagmann, C. Richard Ho, Doug Hogberg, John Hu, Robert Hundt, Dan Hurt, Julian Ibarz, Aaron Jaffey, Alek Jaworski, Alexander Kaplan, Harshit Khaitan, Daniel Killebrew, Andy Koch, Naveen Kumar, Steve Lacy, James Laudon, James Law, Diemthu Le, Chris Leary, Zhuyuan Liu, Kyle Lucke, Alan Lundin, Gordon MacKean, Adriana Maggiore, Maire Mahony, Kieran Miller, Rahul Nagarajan, Ravi Narayanaswami, Ray Ni, Kathy Nix, Thomas Norrie, Mark Omernick, Narayana Penukonda, Andy Phelps, Jonathan Ross, Matt Ross, Amir Salek, Emad Samadiani, Chris Severn, Gregory Sizikov, Matthew Snelham, Jed Souter, Dan Steinberg, Andy Swing, Mercedes Tan, Gregory Thorsen, Bo Tian, Horia Toma, Erick Tuttle, Vijay Vasudevan, Richard Walter, Walter Wang, Eric Wilcox, and Doe Hyun Yoon. 2017. In-datacenter performance analysis of a tensor processing unit. In ISCA, pages 1–12. ACM.
  15. 15.Nal Kalchbrenner, Edward Grefenstette, and Phil Blunsom. 2014. A convolutional neural network for modelling sentences. In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 655–665, Baltimore, Maryland. Association for Computational Linguistics.
  16. 16.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6769–6781, Online. Association for Computational Linguistics.
  17. 17.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019. Natural questions: A benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:452–466.
  18. 18.Guillaume Lample, Alexandre Sablayrolles, Marc’Aurelio Ranzato, Ludovic Denoyer, and Hervé Jégou. 2019. Large memory layers with product keys. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pages 8546–8557.
  19. 19.Jinhyuk Lee, Mujeen Sung, Jaewoo Kang, and Danqi Chen. 2021. Learning dense representations of phrases at scale. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6634–6647, Online. Association for Computational Linguistics.
  20. 20.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020a. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online. Association for Computational Linguistics.
  21. 21.Patrick Lewis, Pontus Stenetorp, and Sebastian Riedel. 2021a. Question and answer test-train overlap in open-domain question answering datasets. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 1000–1008, Online. Association for Computational Linguistics.
  22. 22.Patrick Lewis, Yuxiang Wu, Linqing Liu, Pasquale Minervini, Heinrich Küttler, Aleksandra Piktus, Pontus Stenetorp, and Sebastian Riedel. 2021b. PAQ: 65 million probably-asked questions and what you can do with them. Transactions of the Association for Computational Linguistics, 9:1098–1115.
  23. 23.Patrick S. H. Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020b. Retrieval-augmented generation for knowledge-intensive NLP tasks. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  24. 24.Linqing Liu, Patrick S. H. Lewis, Sebastian Riedel, and Pontus Stenetorp. 2021. Challenges in generalization in open domain question answering. CoRR, abs/2109.01156.
  25. 25.Yury A. Malkov and Dmitry A. Yashunin. 2020. Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs. IEEE Trans. Pattern Anal. Mach. Intell., 42(4):824–836.
  26. 26.Sewon Min, Jordan L. Boyd-Graber, Chris Alberti, Danqi Chen, Eunsol Choi, Michael Collins, Kelvin Guu, Hannaneh Hajishirzi, Kenton Lee, Jennimaria Palomaki, Colin Raffel, Adam Roberts, Tom Kwiatkowski, Patrick S. H. Lewis, Yuxiang Wu, Heinrich Küttler, Linqing Liu, Pasquale Minervini, Pontus Stenetorp, Sebastian Riedel, Sohee Yang, Minjoon Seo, Gautier Izacard, Fabio Petroni, Lucas Hosseini, Nicola De Cao, Edouard Grave, Ikuya Yamada, Sonse Shimaoka, Masatoshi Suzuki, Shumpei Miyawaki, Shun Sato, Ryo Takahashi, Jun Suzuki, Martin Fajcik, Martin Docekal, Karel Ondrej, Pavel Smrz, Hao Cheng, Yelong Shen, Xiaodong Liu, Pengcheng He, Weizhu Chen, Jianfeng Gao, Barlas Oguz, Xilun Chen, Vladimir Karpukhin, Stan Peshterliev, Dmytro Okhonko, Michael Sejr Schlichtkrull, Sonal Gupta, Yashar Mehdad, and Wen-tau Yih. 2020. Neurips 2020 efficientqa competition: Systems, analyses and lessons learned. In NeurIPS (Competition and Demos), volume 133 of Proceedings of Machine Learning Research, pages 86–111. PMLR.
  27. 27.Sewon Min, Danqi Chen, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2019. A discrete hard EM approach for weakly supervised question answering. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2851–2864, Hong Kong, China. Association for Computational Linguistics.
  28. 28.Ashwin Paranjape, Omar Khattab, Christopher Potts, Matei Zaharia, and Christopher D. Manning. 2021. Hindsight: Posterior-guided training of retrievers for improved open-ended generation. ArXiv preprint, abs/2110.07752.
  29. 29.Fabio Petroni, Aleksandra Piktus, Angela Fan, Patrick Lewis, Majid Yazdani, Nicola De Cao, James Thorne, Yacine Jernite, Vladimir Karpukhin, Jean Maillard, Vassilis Plachouras, Tim Rocktäschel, and Sebastian Riedel. 2021. KILT: a benchmark for knowledge intensive language tasks. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2523–2544, Online. Association for Computational Linguistics.
  30. 30.Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019. Language models as knowledge bases? In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2463–2473, Hong Kong, China. Association for Computational Linguistics.
  31. 31.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21:140:1–140:67.
  32. 32.Adam Roberts, Colin Raffel, and Noam Shazeer. 2020. How much knowledge can you pack into the parameters of a language model? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5418–5426, Online. Association for Computational Linguistics.
  33. 33.Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. 2014. Sequence to sequence learning with neural networks. In Advances in Neural Information Processing Systems 27: Annual Conference on Neural Information Processing Systems 2014, December 8-13 2014, Montreal, Quebec, Canada, pages 3104–3112.
  34. 34.Zhiguo Wang, Patrick Ng, Xiaofei Ma, Ramesh Nallapati, and Bing Xiang. 2019. Multi-passage BERT: A globally normalized BERT model for open-domain question answering. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 5878–5882, Hong Kong, China. Association for Computational Linguistics.
  35. 35.Yuxiang Wu, Pasquale Minervini, Pontus Stenetorp, and Sebastian Riedel. 2021. Training adaptive computation for open-domain question answering with computational constraints. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pages 447–453, Online. Association for Computational Linguistics.
  36. 36.Yuxiang Wu, Sebastian Riedel, Pasquale Minervini, and Pontus Stenetorp. 2020. Don’t read too much into it: Adaptive computation for open-domain question answering. In EMNLP (1), pages 3029–3039. Association for Computational Linguistics.
  37. 37.Wei Yang, Yuqing Xie, Aileen Lin, Xingyu Li, Luchen Tan, Kun Xiong, Ming Li, and Jimmy Lin. 2019. End-to-end open-domain question answering with BERTserini. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (Demonstrations), pages 72–77, Minneapolis, Minnesota. Association for Computational Linguistics.
  38. 38.Yunzhi Yao, Shaohan Huang, Ningyu Zhang, Li Dong, Furu Wei, and Huajun Chen. 2022. Kformer: Knowledge injection in transformer feed-forward layers. CoRR, abs/2201.05742.

Citation

MLA
Wu, Y., et al. “An Efficient Memory-Augmented Transformer for Knowledge-Intensive NLP Tasks”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 5184–96, https://doi.org/10.18653/v1/2022.emnlp-main.346.
APA
Wu, Y., Zhao, Y., Hu, B., Minervini, P., Stenetorp, P., & Riedel, S. (2022). An Efficient Memory-Augmented Transformer for Knowledge-Intensive NLP Tasks. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 5184–5196. https://doi.org/10.18653/v1/2022.emnlp-main.346
Chicago
Wu, Y., Y. Zhao, B. Hu, P. Minervini, P. Stenetorp, and S. Riedel. 2022. “An Efficient Memory-Augmented Transformer for Knowledge-Intensive NLP Tasks”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 5184–96. https://doi.org/10.18653/v1/2022.emnlp-main.346.
Harvard
Wu, Y. et al. (2022) “An Efficient Memory-Augmented Transformer for Knowledge-Intensive NLP Tasks”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 5184–5196. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.346.
Vancouver
1. Wu Y, Zhao Y, Hu B, Minervini P, Stenetorp P, Riedel S (2022) An Efficient Memory-Augmented Transformer for Knowledge-Intensive NLP Tasks. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 5184–5196

BibTeX

@inproceedings{wu-etal-2022-efficient,
    title = "An Efficient Memory-Augmented Transformer for Knowledge-Intensive {NLP} Tasks",
    author = "Wu, Yuxiang  and
      Zhao, Yu  and
      Hu, Baotian  and
      Minervini, Pasquale  and
      Stenetorp, Pontus  and
      Riedel, Sebastian",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.346/",
    doi = "10.18653/v1/2022.emnlp-main.346",
    pages = "5184--5196"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/