Meta-Learning Online Adaptation of Language Models

Nathan HuEric MitchellChristopher D. ManningChelsea Finn

article2023EMNLP51 citations

Introduces Context-aware Meta-learned Loss Scaling (CaMeLS), a meta-learning method that trains a lightweight model to dynamically upweight informative tokens during online document streams, substantially boosting factual knowledge uptake in language models over standard fine-tuning.

Listen

Large language models store extensive world knowledge within their parameters, but this information quickly becomes outdated as real-world facts change. While streaming new documents directly into models through online fine-tuning could theoretically keep them current, standard fine-tuning strategies achieve negligible improvement in retaining new facts because the standard training loss treats all words equally, allowing noisy words to drown out the learning signal from critical factual updates.

The article develops and evaluates a meta-learning method called Context-aware Meta-learned Loss Scaling, or CaMeLS, designed to automatically identify and upweight the most informative words in a text stream to improve factual retention during online model updating.

To accomplish this, the authors trained a small secondary model to dynamically assign importance weights to words in incoming texts. This secondary model was optimized through a two-step process that rewarded weight assignments enabling a lightweight stand-in language model to correctly answer downstream questions about the text after just a single update step. The authors evaluated the approach across streams containing thousands of articles from three benchmark question-answering datasets (StreamingQA, SQuAD, and ArchivalQA) using base language models ranging from smaller architectures up to a 6-billion-parameter model.

The primary finding is that the proposed method substantially outperforms standard fine-tuning and common heuristic baselines, such as keyword frequency and text-span filtering, delivering substantial relative gains in question-answering accuracy across all tested models. Crucially, the learned weighting model transfers effectively across architectures: training once on a small 82-million-parameter model successfully guided the online adaptation of a model more than 70 times larger. The weighting model also generalized across entirely different datasets without retraining. Analysis of the learned weights revealed a bimodal, context-dependent distribution that prioritizes informative parts of speech, especially proper nouns and numbers, leading to faster initial learning and longer knowledge retention without requiring substantial additional computational overhead during deployment.

These findings indicate that language models can be kept up-to-date efficiently in continuous operational settings without resorting to costly, frequent full retraining or relying exclusively on external retrieval systems. By focusing parameter updates strictly on high-value tokens, organizations can maintain parametric model accuracy on edge devices and streaming pipelines while minimizing the risk of degrading unrelated existing knowledge. Additionally, preliminary tests show the technique is complementary to retrieval-augmented generation and in-context learning, boosting the performance of both.

Organizations looking to maintain dynamic, up-to-date language models should consider adopting context-aware loss weighting rather than uniform online fine-tuning. Before wide enterprise deployment, decision-makers should run pilot programs within their specific operational domains, as the method currently requires paired document-question training examples during initial setup to teach the weighting model which facts matter. Further validation is also recommended to evaluate performance over much longer data streams and on extreme-scale models exceeding 100 billion parameters.

No sufficiently relevant recommendations were found.

No sufficiently relevant recommendations were found.

Cover for Meta-Learning Online Adaptation of Language Models

Abstract

Large language models encode impressively broad world knowledge in their parameters. However, the knowledge in static language models falls out of date, limiting the model’s effective “shelf life.” While online fine-tuning can reduce this degradation, we find that naively fine-tuning on a stream of documents leads to a low level of information uptake. We hypothesize that online fine-tuning does not sufficiently attend to important information. That is, the gradient signal from important tokens representing factual information is drowned out by the gradient from inherently noisy tokens, suggesting that a dynamic, context-aware learning rate may be beneficial. We therefore propose learning which tokens to upweight. We meta-train a small, autoregressive model to reweight the language modeling loss for each token during online fine-tuning, with the objective of maximizing the out-of-date base question-answering model’s ability to answer questions about a document after a single weighted gradient step. We call this approach Context-aware Meta-learned Loss Scaling (CaMeLS). Across three different distributions of documents, our experiments find that CaMeLS provides substantially improved information uptake on streams of thousands of documents compared with standard fine-tuning and baseline heuristics for reweighting token losses.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Meta-Learning Improved Online Adaptation of Large Language Models
  • 3.1 Unsupervised Online Adaptation
  • 3.2 CaMeLS: Context-aware Meta-learned Loss Scaling
  • 3.3 Mitigating Train-Test Shift
  • 3.4 Compute Requirements of CaMeLS.
  • 4 Experiments
  • 4.1 Datasets
  • 4.2 Experimental protocol details
  • 4.3 CaMeLS improves knowledge retention
  • 4.4 Analysis of learned weights
  • 4.5 Ablations
  • 4.6 Cross Dataset Transfer
  • 4.7 Forgetting and plasticity
  • 5 Discussion
  • References
  • A Dataset Details
  • B Larger Proxy Models
  • C Combining CaMeLS with other online Adaptation Methods

Knowls

  1. Knowl 1 — Unsupervised online adaptation is evaluated by query performance

    definition

    The paper studies updating an outdated question-answering language model from a stream of recent documents when the questions that will later be asked about those documents are unavailable during adaptation. An adapted model is evaluated on queries and answers associated with the stream’s documents. The adaptation objective is therefore to retain information useful for typical downstream queries, rather than to reproduce every document verbatim. Training uses separate document–query–answer examples sampled from a process intended to resemble the test process, providing supervision for learning a general update strategy.

  2. Knowl 2 — CaMeLS learns context-sensitive weights for online language-model loss

    model/method

    Context-aware Meta-learned Loss Scaling (CaMeLS) trains a small autoregressive weighting model wϕw_\phi to produce an importance weight for each token in an adaptation document. The weights rescale the document’s token-level negative log-likelihood (NLL), changing which tokens contribute most to the language model’s fine-tuning gradient. During meta-training, the weighting model is optimized so that a proxy language model, after a weighted document update, answers a query about that document more accurately. At deployment, the learned weighting model is applied to each incoming document to reweight the NLL used to update the target model; the target model need not be the same size as the proxy.

  3. Knowl 3 — Bilevel training rewards useful updates while regularizing behavior changes

    equation

    For a meta-training example consisting of document xx, query qq, and answer yy, let θbase\theta_{\mathrm{base}} be the proxy model’s parameters before adaptation, fθf_\theta the language model with parameters θ\theta, and L(fθ,x,a)L(f_\theta,x,a) the document NLL weighted by token weights aa. The inner update is

    θ′=θbase−α∇θbaseL(fθbase,x,wϕ(x)).\theta' = \theta_{\mathrm{base}} - \alpha\nabla_{\theta_{\mathrm{base}}} L\bigl(f_{\theta_{\mathrm{base}}},x,w_\phi(x)\bigr).

    The outer objective evaluates the adapted proxy on the document’s query and answer while penalizing changes to its predictions on unrelated text:

    Louter=−log⁡pθ′(y∣q)+clocLloc(θbase,θ′,xloc),Lloc=∑iKL ⁣(pθbase(⋅∣xloci) ∥ pθ′(⋅∣xloci)).L_{\mathrm{outer}} = -\log p_{\theta'}(y\mid q) + c_{\mathrm{loc}}L_{\mathrm{loc}}(\theta_{\mathrm{base}},\theta',x_{\mathrm{loc}}), \qquad L_{\mathrm{loc}} = \sum_i \mathrm{KL}\!\left(p_{\theta_{\mathrm{base}}}(\cdot\mid x^i_{\mathrm{loc}})\,\middle\|\,p_{\theta'}(\cdot\mid x^i_{\mathrm{loc}})\right).

    Here, wϕ(x)w_\phi(x) is the vector of token weights predicted for xx; α\alpha is the inner-loop learning rate; xlocx_{\mathrm{loc}} is an unrelated locality text; and xlocix^i_{\mathrm{loc}} is its iith prefix. Each pθ(⋅∣⋅)p_\theta(\cdot\mid\cdot) is a next-token probability distribution, and KL\mathrm{KL} is Kullback–Leibler divergence. The weighting-model parameters ϕ\phi are updated to reduce this outer objective. Experiments use α=5×10−4\alpha=5\times10^{-4}, cloc=0.1c_{\mathrm{loc}}=0.1, OpenWebText for locality text, and Adam with outer learning rate 10−510^{-5}. Outer gradients are accumulated over 24 examples in four batches of six.

  4. Knowl 4 — Multi-document meta-training addresses adaptation over long streams

    model/method

    To make the learned update strategy reflect a sequence of online updates rather than an isolated document, CaMeLS meta-training performs an inner update on each of k=6k=6 document–query–answer examples in sequence. If θ0\theta_0 is the episode’s starting proxy-model state and xix_i is the iith document, the update is θi=θi−1−α∇θi−1L(fθi−1,xi,wϕ(xi))\theta_i=\theta_{i-1}-\alpha\nabla_{\theta_{i-1}}L(f_{\theta_{i-1}},x_i,w_\phi(x_i)) for i=1,…,6i=1,\ldots,6. The outer query loss is averaged across the six examples, with the locality penalty retained. To avoid training the weighting model only on one proxy-model state, the next episode normally starts from the previous episode’s final state; every four training episodes, the starting state is reset to the original proxy-model parameters.

  5. Knowl 5 — Evaluation covers three document streams and transfers from a small proxy

    experimental setup

    CaMeLS is evaluated on StreamingQA, SQuAD, and ArchivalQA. The online streams average approximately 510 tokens across 1,665 articles for StreamingQA, 150 tokens across 1,170 paragraphs for SQuAD, and 80 tokens across 3,001 paragraphs for ArchivalQA. SQuAD and ArchivalQA answers are spans in their associated documents; StreamingQA answers are not necessarily document spans. A separate QA-tuned DistilGPT-2 (82 million parameters) serves as the proxy during meta-training, and DistilGPT-2 with an MLP head of hidden size 128 serves as the weighting model. Evaluation uses QA-tuned GPT-2, GPT-Neo, and GPT-J models, including GPT-J 6B; the evaluated models share a tokenizer. Test-time learning rates are selected using validation streams. Reported stream-level comparisons average over four sampled test streams.

  6. Knowl 6 — CaMeLS improves information uptake over fine-tuning and heuristic weighting

    empirical result

    Across the three document distributions and the evaluated base-model sizes, CaMeLS generally yields larger relative F1 improvements on questions about the adaptation documents than uniform fine-tuning, salient-span weighting, or TF-IDF weighting. Uniform fine-tuning provides little improvement in many tested settings, even when its learning rate is tuned; adding QA fine-tuning after uniform adaptation does not consistently solve this. The learned weights transfer from the 82-million-parameter DistilGPT-2 proxy to GPT-J 6B, a model over 70 times larger. These results support the practical use of a weighting model trained once on a small proxy to guide updates to larger target models.

  7. Knowl 7 — Learned weights transfer across datasets, but transfer is not uniformly best

    data/table

    The entries below are relative F1 improvements when adapting GPT-2 XL. Each CaMeLS row identifies the dataset used to train the weighting model; columns identify the evaluation stream. The bottom rows are baseline adaptation methods. CaMeLS weighting models trained on a different dataset often outperform the baselines, while the values also show that transfer is not uniformly superior—for example, TF-IDF is higher than two CaMeLS results on SQuAD.

    Adaptation method / CaMeLS training data StreamingQA SQuAD ArchivalQA
    CaMeLS (StreamingQA) 0.298 0.089 0.233
    CaMeLS (SQuAD) 0.249 0.174 0.254
    CaMeLS (ArchivalQA) 0.153 0.040 0.268
    Uniform + QA-tune 0.096 0.043 0.060
    Salient Spans 0.072 0.028 0.210
    TF-IDF + 5% Cutoff 0.086 0.158 0.181
  8. Knowl 8 — Context matters beyond token part of speech

    empirical result

    On the StreamingQA validation data, CaMeLS’s importance weights are sparse and bimodal, and proper nouns and numbers are especially likely to receive high weights. Restricting the weighting model to choose between only the lowest and highest values in its learned weight range slightly reduces performance but retains substantial gains over baseline methods. In contrast, two ablations that remove contextual information—sampling a weight based only on a token’s part of speech or assigning the mean weight for that part of speech—perform worse than uniform weighting or fail to produce a significant F1 increase. Thus, part of speech correlates with the learned weights but does not by itself identify which tokens are useful in context.

  9. Knowl 9 — CaMeLS improves acquisition during adaptation while new knowledge still fades

    empirical result

    During GPT-2 XL adaptation on StreamingQA, evaluated every 200 document updates, CaMeLS increases performance on questions about the stream as adaptation proceeds. In the comparison at learning rate 2.5×10−52.5\times10^{-5}, uniform fine-tuning with post-adaptation QA tuning realizes its gains only after that additional QA-tuning step, whereas CaMeLS improves test-query performance during online adaptation. All evaluated adaptation methods gradually reduce performance on unrelated QA-validation questions; at 2.5×10−52.5\times10^{-5}, CaMeLS has the smallest such decline among the compared methods. When query improvement is measured against the number of document updates since the answer-bearing document was seen, CaMeLS has both a larger initial improvement and a larger later improvement than the compared methods, consistent with improved immediate acquisition and retention of newly learned information.

  10. Knowl 10 — CaMeLS requires paired query supervision and has limited evaluated scope

    limitation

    Meta-training CaMeLS requires training triples containing a document, a query about that document, and its answer; this supervision may be costly or unavailable in some domains. The experiments evaluate question answering on streams of thousands of documents and base models up to 6 billion parameters. They do not establish performance on substantially longer streams, models at 100 billion parameters or more, or downstream tasks other than question answering. The paper identifies learning from a purely unlabeled stream and testing other tasks and model types as open directions.

Coverage note — The preliminary experiments combining CaMeLS with five-shot in-context learning and document retrieval are omitted because they are secondary and lack comparisons against other adaptation baselines; acknowledgements and background are outside the contribution scope.

References

  1. 1.Sid Black, Leo Gao, Phil Wang, Connor Leahy, and Stella Biderman. 2021. GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow.
  2. 2.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared J Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  3. 3.Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, Harsha Nori, Hamid Palangi, Marco Tulio Ribeiro, and Yi Zhang. 2023. Sparks of artificial general intelligence: Early experiments with GPT-4. ArXiv preprint arXiv:2303.12712.
  4. 4.Arslan Chaudhry, Marc’Aurelio Ranzato, Marcus Rohrbach, and Mohamed Elhoseiny. 2019. Efficient lifelong learning with a-GEM. In International Conference on Learning Representations.
  5. 5.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  6. 6.Kevin Clark, Kelvin Guu, Ming-Wei Chang, Panupong Pasupat, Geoffrey Hinton, and Mohammad Norouzi. 2022. Meta-learning fast weight language models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 9751–9757, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  7. 7.Bhuwan Dhingra, Jeremy R. Cole, Julian Martin Eisenschlos, Daniel Gillick, Jacob Eisenstein, and William W. Cohen. 2022. Time-Aware Language Models as Temporal Knowledge Bases. Transactions of the Association for Computational Linguistics, 10:257–273.
  8. 8.Shibhansh Dohare, Richard S. Sutton, and A. Rupam Mahmood. 2022. Continual backprop: Stochastic gradient descent with persistent randomness.
  9. 9.Danny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, Wenlong Huang, Yevgen Chebotar, Pierre Sermanet, Daniel Duckworth, Sergey Levine, Vincent Vanhoucke, Karol Hausman, Marc Toussaint, Klaus Greff, Andy Zeng, Igor Mordatch, and Pete Florence. 2023. Palm-e: An embodied multimodal language model. In arXiv preprint arXiv:2303.03378.
  10. 10.Georgi Gerganov. 2023. llama.cpp. https://github.com/ggerganov/llama.cpp.
  11. 11.Aaron Gokaslan, Vanya Cohen, Ellie Pavlick, and Stefanie Tellex. 2019. Openwebtext corpus. http://Skylion007.github.io/OpenWebTextCorpus.
  12. 12.Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. 2020. REALM: retrieval-augmented language model pre-training. CoRR, abs/2002.08909.
  13. 13.William L. Hamilton, Jure Leskovec, and Dan Jurafsky. 2016. Diachronic word embeddings reveal statistical laws of semantic change. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1489–1501, Berlin, Germany. Association for Computational Linguistics.
  14. 14.Joel Jang, Seonghyeon Ye, Sohee Yang, Joongbo Shin, Janghoon Han, Gyeonghun KIM, Stanley Jungkyu Choi, and Minjoon Seo. 2022. Towards continual knowledge learning of language models. In International Conference on Learning Representations.
  15. 15.F. Jelinek, B. Merialdo, S. Roukos, and M. Strauss. 1991. A dynamic language model for speech recognition. In Speech and Natural Language: Proceedings of a Workshop Held at Pacific Grove, California, February 19-22, 1991.
  16. 16.James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell. 2016. Overcoming catastrophic forgetting in neural networks. Proceedings of the National Academy of Sciences, 114.
  17. 17.Roland Kuhn. 1988. Speech recognition and the frequency of recently used words: A modified Markov model for natural language. In Coling Budapest 1988 Volume 1: International Conference on Computational Linguistics.
  18. 18.Vivek Kulkarni, Rami Al-Rfou, Bryan Perozzi, and Steven Skiena. 2015. Statistically significant detection of linguistic change. In Proceedings of the 24th International Conference on World Wide Web, WWW ’15, page 625–635, Republic and Canton of Geneva, CHE. International World Wide Web Conferences Steering Committee.
  19. 19.Angeliki Lazaridou, Adhiguna Kuncoro, Elena Gribovskaya, Devang Agrawal, Adam Liska, Tayfun Terzi, Mai Gimenez, Cyprien de Masson d’Autume, Tomáš Kociský, Sebastian Ruder, Dani Yogatama, Kris Cao, Susannah Young, and Phil Blunsom. 2021. Mind the gap: Assessing temporal generalization in neural language models. In Advances in Neural Information Processing Systems.
  20. 20.Wei Li, Wenhao Wu, Moye Chen, Jiachen Liu, Xinyan Xiao, and Hua Wu. 2022. Faithfulness in natural language generation: A systematic survey of analysis, evaluation and optimization methods. ArXiv, abs/2203.05227.
  21. 21.Adam Liška, Tomáš Kociský, Elena Gribovskaya, Tayfun Terzi, Eren Sezener, Devang Agrawal, Cyprien de Masson d’Autume, Tim Scholtes, Manzil Zaheer, Susannah Young, Ellen Gilsenan-McMahon, Sophia Austin, Phil Blunsom, and Angeliki Lazaridou. 2022. Streamingqa: A benchmark for adaptation to new knowledge over time in question answering models.
  22. 22.Shayne Longpre, Kartik Perisetla, Anthony Chen, Nikhil Ramesh, Chris DuBois, and Sameer Singh. 2021. Entity-based knowledge conflicts in question answering. arXiv preprint arXiv:2109.05052.
  23. 23.David Lopez-Paz and Marc’Aurelio Ranzato. 2017. Gradient episodic memory for continual learning. Advances in neural information processing systems, 30.
  24. 24.Divyam Madaan, Jaehong Yoon, Yuanchun Li, Yunxin Liu, and Sung Ju Hwang. 2022. Representational continuity for unsupervised continual learning. In International Conference on Learning Representations.
  25. 25.Michael McCloskey and Neal J. Cohen. 1989. Catastrophic interference in connectionist networks: The sequential learning problem. In Gordon H. Bower, editor, Psychology of Learning and Motivation, volume 24 of Psychology of Learning and Motivation, pages 109–165. Academic Press.
  26. 26.Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022. Locating and editing factual associations in GPT. ArXiv:2202.05262.
  27. 27.Eric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn, and Christopher D. Manning. 2021. Fast model editing at scale. CoRR.
  28. 28.Eric Mitchell, Charles Lin, Antoine Bosselut, Christopher D Manning, and Chelsea Finn. 2022. Memory-based model editing at scale. In Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, pages 15817–15831. PMLR.
  29. 29.T. Mitchell, W. Cohen, E. Hruschka, P. Talukdar, B. Yang, J. Betteridge, A. Carlson, B. Dalvi, M. Gardner, B. Kisiel, J. Krishnamurthy, N. Lao, K. Mazaitis, T. Mohamed, N. Nakashole, E. Platanios, A. Ritter, M. Samadi, B. Settles, R. Wang, D. Wijaya, A. Gupta, X. Chen, A. Saparov, M. Greaves, and J. Welling. 2018. Never-ending learning. Commun. ACM, 61(5):103–115.
  30. 30.Miles Osborne, Ashwin Lall, and Benjamin Van Durme. 2014. Exponential reservoir sampling for streaming language models. In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 687–692, Baltimore, Maryland. Association for Computational Linguistics.
  31. 31.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2018. Language models are unsupervised multitask learners. Ms., OpenAI.
  32. 32.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. Squad: 100,000+ questions for machine comprehension of text.
  33. 33.Dushyant Rao, Francesco Visin, Andrei Rusu, Razvan Pascanu, Yee Whye Teh, and Raia Hadsell. 2019. Continual unsupervised representation learning. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc.
  34. 34.Marek Rei. 2015. Online representation learning in recurrent neural language models. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 238–243, Lisbon, Portugal. Association for Computational Linguistics.
  35. 35.Mengye Ren, Michael Louis Iuzzolino, Michael Curtis Mozer, and Richard Zemel. 2021. Wandering within a world: Online contextualized few-shot learning. In International Conference on Learning Representations.
  36. 36.G. Salton and M. J. McGill. 1986. Introduction to Modern Information Retrieval. McGraw-Hill, Inc., New York, NY, USA.
  37. 37.Sandhaus, Evan. 2008. The new york times annotated corpus.
  38. 38.Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019. Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter. In NeurIPS EMC2 Workshop.
  39. 39.Hanul Shin, Jung Kwon Lee, Jaehong Kim, and Jiwon Kim. 2017. Continual learning with deep generative replay. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc.
  40. 40.Chenglei Si, Zhe Gan, Zhengyuan Yang, Shuohang Wang, Jianfeng Wang, Jordan Lee Boyd-Graber, and Lijuan Wang. 2023. Prompting GPT-3 to be reliable. In The Eleventh International Conference on Learning Representations.
  41. 41.Anton Sinitsin, Vsevolod Plokhotnyuk, Dmitry Pyrkin, Sergei Popov, and Artem Babenko. 2020. Editable neural networks. In ICLR.
  42. 42.Sebastian Thrun and Tom M. Mitchell. 1995. Lifelong robot learning. Robotics and Autonomous Systems, 15(1):25–46. The Biology and Technology of Intelligent Autonomous Agents.
  43. 43.Ben Wang and Aran Komatsuzaki. 2021. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https://github.com/kingoflolz/mesh-transformer-jax.
  44. 44.Jiexin Wang, Adam Jatowt, and Masatoshi Yoshikawa. 2022. Archivalqa: A large-scale benchmark dataset for open domain question answering over historical news collections.
  45. 45.Dani Yogatama, Chong Wang, Bryan R. Routledge, Noah A. Smith, and Eric P. Xing. 2014. Dynamic language models for streaming text. Transactions of the Association for Computational Linguistics, 2:181–192.

Citation

MLA
Hu, N., et al. “Meta-Learning Online Adaptation of Language Models”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 4418–32, https://doi.org/10.18653/v1/2023.emnlp-main.268.
APA
Hu, N., Mitchell, E., Manning, C. D., & Finn, C. (2023). Meta-Learning Online Adaptation of Language Models. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 4418–4432. https://doi.org/10.18653/v1/2023.emnlp-main.268
Chicago
Hu, N., E. Mitchell, C. D. Manning, and C. Finn. 2023. “Meta-Learning Online Adaptation of Language Models”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 4418–32. https://doi.org/10.18653/v1/2023.emnlp-main.268.
Harvard
Hu, N. et al. (2023) “Meta-Learning Online Adaptation of Language Models”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 4418–4432. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.268.
Vancouver
1. Hu N, Mitchell E, Manning CD, Finn C (2023) Meta-Learning Online Adaptation of Language Models. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 4418–4432

BibTeX

@inproceedings{hu-etal-2023-meta,
    title = "Meta-Learning Online Adaptation of Language Models",
    author = "Hu, Nathan  and
      Mitchell, Eric  and
      Manning, Christopher  and
      Finn, Chelsea",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.268/",
    doi = "10.18653/v1/2023.emnlp-main.268",
    pages = "4418--4432"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/