Can We Edit Factual Knowledge by In-Context Learning?

Ce ZhengLei LiQingxiu DongYuxuan FanZhiyong WuJingjing XuBaobao Chang

article2023EMNLP377 citations

Proposes an in-context knowledge editing framework that updates facts in large language models using structured demonstration prompts, achieving competitive editing performance without costly parameter updates or catastrophic forgetting.

Listen

Large language models store immense amounts of factual knowledge in their parameters, but this information frequently becomes outdated, incorrect, or biased. Correcting these errors traditionally requires parameter-tuning or gradient-based methods, which identify and adjust internal model weights. As models scale up and are increasingly delivered as closed, black-box cloud services, these traditional editing approaches become computationally expensive, hard to scale, and risky due to parameter damage that alters unrelated behaviors.

The article evaluates whether factual knowledge can be effectively updated without changing model weights by using in-context learning, a method called In-Context Knowledge Editing (IKE). The primary objective is to demonstrate that structured demonstration prompts can balance generalization—applying the new fact to paraphrased questions—with specificity, which ensures unrelated facts remain unchanged.

The authors tested this approach across multiple autoregressive language models spanning 1.5 billion to 175 billion parameters, including GPT-J, GPT-NeoX, and OPT-175B, as well as instruction-tuned models like LLaMA and Vicuna. The core evaluation used the COUNTERFACT benchmark consisting of over 21,000 records designed to test difficult counterfactual edits and evaluate side effects. The method constructs prompts containing three distinct demonstration types: copy examples to introduce the target fact, update examples to generalize across different phrasings, and retain examples to preserve unrelated knowledge retrieved via sentence similarity.

The investigation produced four central findings. First, IKE achieved competitive editing efficacy without altering any model parameters, matching or closely trailing complex parameter-modifying methods like ROME and outperforming baseline hyper-networks by roughly 10% in editing success on GPT-J. Second, the method dramatically reduced side effects: it prevented catastrophic forgetting of historical facts, retaining an 88% memorization ratio on sequential time-aware edits compared to less than 0.1% for ROME. Third, contrastive evaluations showed that IKE substantially lowered over-editing on similar, unrelated relations. Finally, the approach scaled positively with model size, delivering its highest accuracy, generalization (98.8%), and specificity (85.1%) on the largest 175-billion-parameter model.

These results indicate that organizations deploying large language models can correct facts with lower computational expense, reduced deployment risk, and greater operational transparency. Because the underlying model remains untouched, edits are fully reversible and easily calibrated through natural language. Furthermore, this method bypasses the access restrictions of black-box model-as-a-service APIs where parameter modification is impossible.

For practical implementation, technical teams should consider pairing this in-context approach with external retrieval memory systems, allowing systems to dynamically fetch relevant edit demonstrations per query rather than overloading input windows. Further engineering work and pilot testing are recommended to refine demonstration retrieval and validate performance across diverse query formats before full production rollout.

The primary limitations include increased per-query inference costs caused by longer input context lengths and the current restriction of the evaluation to factual knowledge rather than broader commonsense reasoning. Nevertheless, the evidence provides high confidence that demonstration-based in-context editing offers a viable, scalable alternative to direct parameter manipulation for large language models.

arXiv: 2305.12740
Cover for Can We Edit Factual Knowledge by In-Context Learning?

Abstract

Previous studies have shown that large language models (LLMs) like GPTs store massive factual knowledge in their parameters. However, the stored knowledge could be false or out-dated. Traditional knowledge editing methods refine LLMs via fine-tuning on texts containing specific knowledge. However, with the increasing scales of LLMs, these gradient-based approaches bring large computation costs. The trend of model-as-a-service also makes it impossible to modify knowledge in black-box LLMs. Inspired by in-context learning (ICL), a new paradigm based on demonstration contexts without parameter updating, we explore whether ICL can edit factual knowledge. To answer this question, we give a comprehensive empirical study of ICL strategies. Experiments show that in-context knowledge editing (IKE), without any gradient and parameter updating, achieves a competitive success rate compared to gradient-based methods on GPT-J (6B) but with much fewer side effects, including less over-editing on similar but unrelated facts and less knowledge forgetting on previously stored knowledge. We also apply the method to larger LMs with tens or hundreds of parameters like OPT-175B, which shows the scalability of our method. The code is available at https://github.com/pkunlp-icler/IKE.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Task Formulation
  • 4 Method: IKE
  • 4.1 In-Context Learning
  • 4.2 In-Context Knowledge Editing
  • 4.2.1 Demonstration Formatting
  • 4.2.2 Demonstration Organization
  • 4.3 Discussion: Gradient-based methods and gradient-free methods
  • 5 Experiment
  • 5.1 Experimental Setting
  • 5.1.1 Baselines
  • 5.1.2 Evaluation Setup
  • 5.2 Main Results
  • 5.3 Analysis
  • 5.3.1 Ablation on Demonstration
  • 5.3.2 IKE Benefits from Model Scaling
  • 5.3.3 Resilience to Over-Editing
  • 5.3.4 Maintenance for Original Knowledge
  • 6 Discussions
  • 7 Conclusion
  • Limitations
  • Acknowledgements
  • References
  • A Implementation Details
  • A.1 IKE
  • A.2 Demonstration Designing
  • A.2.1 Demonstration Formatting
  • A.3 Other Baselines
  • B Details of COUNTERFACT Dataset
  • C Model Details
  • D Time-aware Knowledge Editing
  • E Detailed Discussions
  • E.1 Scale up to more factual edits
  • E.2 Generalization on facts and prompts

Knowls

  1. Knowl 1 — In-Context Knowledge Editing Formulation

    model/method

    In-Context Knowledge Editing (IKE) performs factual knowledge editing on large language models (LLMs) without altering model parameters θ\theta. Given a target knowledge update f=(x∗,y∗)f = (x^*, y^*), where x∗x^* represents the probing prompt and y∗y^* represents the new target prediction (replacing the original fact prediction yoy^o), IKE constructs a set of kk demonstration contexts C={c1,c2,…,ck}C = \{c_1, c_2, \dots, c_k\}.

    The editing objective satisfies two criteria:

    1. Generalization: For any input prompt xx within the edit scope of the target prompt (x∈Dx∗x \in \mathcal{D}_{x^*}), maximize the conditional probability of the new target answer: max⁡PM(y∗∣x,f,C)\max P_M(y^* \mid x, f, C)

    2. Specificity: For any input prompt xx outside the edit scope (x∉Dx∗x \notin \mathcal{D}_{x^*}), preserve the original distribution by minimizing the divergence between pre-edit and post-edit outputs: min⁡D(PM(y∣x,f,C),  PM(y∣x))\min \mathcal{D}\big(P_M(y \mid x, f, C), \; P_M(y \mid x)\big)

    Unlike gradient-based methods that optimize Δθ=−∇θlog⁡PM(y∗∣x∗)\Delta \theta = -\nabla_\theta \log P_M(y^* \mid x^*) to produce an updated model Mθ+ΔθM_{\theta + \Delta \theta}, IKE evaluates predictions via PM(y∣x,f,C)P_M(y \mid x, f, C), preserving base model weights entirely and enabling execution on black-box or service-hosted LLMs.

  2. Knowl 2 — Demonstration Formatting for In-Context Knowledge Editing

    model/method

    In In-Context Knowledge Editing (IKE), each demonstration example ci=(fi,xi,yi)c_i = (f_i, x_i, y_i) comprises an injected fact fi=(xi∗,yi∗)f_i = (x^*_i, y^*_i), a probing prompt xix_i, and an expected answer yiy_i. Demonstrations are generated under three distinct functional roles:

    • Copy demonstrations: Instruct the language model to adopt the new fact for verbatim queries, where xi=xi∗x_i = x^*_i and yi=yi∗y_i = y^*_i.
    • Update demonstrations: Instruct the language model to generalize the injected fact to semantically rephrased prompts within the edit scope, where xi∈Dxi∗x_i \in \mathcal{D}_{x^*_i} and yi=yi∗y_i = y^*_i.
    • Retain demonstrations: Instruct the language model to maintain its pre-existing factual knowledge for out-of-scope prompts, where xi∉Dxi∗x_i \notin \mathcal{D}_{x^*_i} and yi=yioy_i = y^o_i (the original unedited answer).

    Demonstrations and the final query are formatted via a natural language template T(f,x,y)T(f, x, y) defined as: New Fact: f. Prompt: x  y\text{New Fact: } f. \text{ Prompt: } x \; y

    The proportion of demonstration types in context is set to a ratio of 1 (copy) : 3 (update) : 4 (retain), with demonstration types interleaved across the prompt to achieve a uniform distribution.

  3. Knowl 3 — Demonstration Selection and Retrieval Algorithm in IKE

    algorithm

    To construct the in-context demonstration prefix C={c1,…,ck}C = \{c_1, \dots, c_k\} for a target factual edit f=(x∗,y∗)f = (x^*, y^*) with original prediction yoy^o, demonstrations are retrieved and ordered from a training corpus using dense sentence representations.

    Input: Target edit fact f=(x∗,y∗)f = (x^*, y^*), original output yoy^o, demonstration candidate pool Dtrain\mathcal{D}_{train}, pretrained sentence encoder E\mathcal{E} (e.g., all-MiniLM-L6-v2), demonstration count kk
    Output: Ordered in-context prompt context C={c1,c2,…,ck}C = \{c_1, c_2, \dots, c_k\}
    vf←E(concat(x∗,yo,y∗))v_f \leftarrow \mathcal{E}(\text{concat}(x^*, y^o, y^*))
    for each candidate ci=(fi,xi,yi)∈Dtrainc_i = (f_i, x_i, y_i) \in \mathcal{D}_{train} do
        vci←E(concat(xi∗,yio,yi∗))v_{c_i} \leftarrow \mathcal{E}(\text{concat}(x^*_i, y^o_i, y^*_i))
        si←vf⋅vci∥vf∥∥vci∥s_i \leftarrow \frac{v_f \cdot v_{c_i}}{\|v_f\| \|v_{c_i}\|}
    end for
    Ctop←Select k candidates with highest cosine similarities si from DtrainC_{top} \leftarrow \text{Select } k \text{ candidates with highest cosine similarities } s_i \text{ from } \mathcal{D}_{train}
    Sort CtopC_{top} in ascending order of similarity: s(1)<s(2)<⋯<s(k)s_{(1)} < s_{(2)} < \dots < s_{(k)}
    C←(c(1),c(2),…,c(k))C \leftarrow (c_{(1)}, c_{(2)}, \dots, c_{(k)})
    return CC

    For models with a maximum input sequence length of 2048 tokens, k=32k = 32; for models with a context length limit of 1024 tokens, k=16k = 16. Placing the examples in ascending similarity order ensures that the most semantically relevant demonstrations appear closest to the query prompt at the end of the context.

  4. Knowl 4 — Factual Knowledge Editing Metrics on CounterFact

    definition

    Knowledge editing performance on the COUNTERFACT benchmark is evaluated on edit records modifying a factual triplet (s∗,r∗,oc)→(s∗,r∗,o∗)(s^*, r^*, o^c) \to (s^*, r^*, o^*) across three dimensions:

    1. Efficacy (Target Prompts):

      • Efficacy Score: ES=E[I[P(o∗)>P(oc)]]\text{ES} = \mathbb{E}\big[\mathbb{I}[P(o^*) > P(o^c)]\big]
      • Efficacy Magnitude: EM=E[P(o∗)−P(oc)]\text{EM} = \mathbb{E}[P(o^*) - P(o^c)]
    2. Generalization (In-Scope Paraphrase Prompts PP∈Dx∗P^P \in \mathcal{D}_{x^*}):

      • Paraphrase Score: PS=E[I[P(o∗)>P(oc)]]\text{PS} = \mathbb{E}\big[\mathbb{I}[P(o^*) > P(o^c)]\big]
      • Paraphrase Magnitude: PM=E[P(o∗)−P(oc)]\text{PM} = \mathbb{E}[P(o^*) - P(o^c)]
    3. Specificity (Out-of-Scope Neighborhood Prompts PN∉Dx∗P^N \notin \mathcal{D}_{x^*} sharing (s′,r∗,oc)(s', r^*, o^c)):

      • Neighborhood Score: NS=E[I[P(oc)>P(o∗)]]\text{NS} = \mathbb{E}\big[\mathbb{I}[P(o^c) > P(o^*)]\big]
      • Neighborhood Magnitude: NM=E[P(oc)−P(o∗)]\text{NM} = \mathbb{E}[P(o^c) - P(o^*)]
    4. Overall Score (SS): The harmonic mean of Efficacy Score, Paraphrase Score, and Neighborhood Score: S=31ES+1PS+1NSS = \frac{3}{\frac{1}{\text{ES}} + \frac{1}{\text{PS}} + \frac{1}{\text{NS}}}

  5. Knowl 5 — Knowledge Editing Benchmark Results on GPT-J and OPT-175B

    data/table

    Performance of knowledge editing approaches evaluated on the 2,000-example test split of the COUNTERFACT benchmark. Methods compared include base unedited models, fine-tuning (FT with Adam and early stopping), MEND, ROME, zero-shot fact prefixing (PROMPT), and IKE with k=32k=32 demonstrations.

    Editing Method #Edited Params #Extra Params Score (SS)↑\uparrow ES↑\uparrow EM↑\uparrow PS↑\uparrow PM↑\uparrow NS↑\uparrow NM↑\uparrow
    GPT-J (6B) 0 0 22.0 16.2 -7.4 15.9 -7.5 83.2 7.4
    FT 64M 0 28.7 99.9 98.6 96.4 67.0 11.9 -48.6
    MEND 384M 896M 63.6 90.4 53.9 53.4 14.3 57.6 -3.3
    ROME 64M 256M 91.5 100.0 99.4 99.6 78.0 78.5 5.0
    PROMPT 0 0 63.3 99.7 80.9 91.0 32.9 37.9 -2.8
    IKE (32 examples) 0 20M 89.6 100.0 91.7 95.2 64.5 77.0 35.2
    OPT (175B) 0 0 18.7 12.6 -8.4 14.3 -8.1 86.9 8.4
    PROMPT 0 0 58.1 99.6 77.2 94.1 37.4 32.3 -7.8
    IKE (32 examples) 0 20M 94.1 100.0 92.5 98.8 83.6 85.1 45.5

    IKE modifies 0 model parameters and requires only 20M parameters for the sentence embedding retriever. On GPT-J (6B), IKE achieves an overall score of 89.6 (competitive with ROME's 91.5) and achieves an NM of 35.2 (compared to ROME's 5.0 and FT's -48.6). On OPT-175B, where weight-editing methods are computationally prohibitive, IKE scales directly and achieves a score of 94.1.

  6. Knowl 6 — Ablation of Demonstration Quantity, Organization, and Types in IKE

    data/table

    Ablation experiments conducted on GPT-J (6B) using the COUNTERFACT benchmark, evaluating variations in demonstration count kk, retrieval/ordering strategies, and demonstration formatting roles.

    Editing Configuration Score (SS)↑\uparrow ES↑\uparrow PS↑\uparrow NS↑\uparrow
    IKE (32 examples) 89.6 100.0 95.2 77.0
    - 4 examples 81.5 99.6 83.5 67.5
    - 8 examples 84.2 100.0 85.6 71.7
    - 16 examples 87.0 100.0 91.7 73.6
    - random selection 70.3 100.0 95.8 45.0
    - random ordering 88.9 100.0 95.4 75.1
    - w/o copy 88.6 100.0 96.9 73.9
    - w/o update 84.4 100.0 73.8 83.4
    - w/o retain 28.0 100.0 99.8 11.5

    Key observations:

    • Increasing kk from 4 to 32 monotonically improves the overall score SS from 81.5 to 89.6.
    • Unsupervised nearest-neighbor retrieval is critical: replacing cosine-retrieved demonstrations with randomly selected demonstrations drops Neighborhood Score (NS) from 77.0 to 45.0.
    • Retain demonstrations are essential for specificity: removing retain demonstrations causes NS to collapse to 11.5 (and NM drops from 35.2 to -47.6).
    • Update demonstrations govern generalization: removing update demonstrations drops Paraphrase Score (PS) from 95.2 to 73.8.
  7. Knowl 7 — Scaling of IKE Across Model Sizes and Instruction-Tuned Architectures

    data/table

    Generalization and specificity of IKE (k=32k=32, except k=16k=16 for GPT-2 XL due to maximum sequence length limits) evaluated across decoder-only language models of varying scale and tuning paradigms on COUNTERFACT:

    Model Architecture PS↑\uparrow PM↑\uparrow NS↑\uparrow NM↑\uparrow
    GPT-2 XL (1.5B) 85.1 42.8 72.0 21.0
    GPT-NEO (2.7B) 96.3 73.5 70.7 28.0
    GPT-J (6B) 95.2 64.5 77.0 35.2
    GPT-NEOX (20B) 97.5 78.3 79.8 41.3
    OPT (175B) 98.8 83.6 85.1 45.5
    LLaMA (7B) 95.0 34.9 74.1 19.9
    Vicuna (7B) 88.9 31.5 82.6 33.5
    LLaMA-2 Chat (7B) 95.7 49.8 83.7 35.1

    Performance metrics scale monotonically with model parameter capacity from 1.5B to 175B parameters (OPT-175B achieves the highest PS of 98.8 and NS of 85.1). Furthermore, IKE applies directly to instruction-aligned models (Vicuna 7B, LLaMA-2 Chat 7B), yielding higher specificity (NS 82.6–83.7) than base LLaMA (NS 74.1).

  8. Knowl 8 — Contrastive Knowledge Assessment and Over-Editing Resistance

    data/table

    To evaluate over-editing beyond the standard out-of-scope neighborhood set, Contrastive Knowledge Assessment (CKA) measures whether injecting an edit (s∗,r∗,oc)→(s∗,r∗,o∗)(s^*, r^*, o^c) \to (s^*, r^*, o^*) erroneously increases the probability of the new object o∗o^* when probed with unrelated relations r′∈Rr' \in \mathcal{R} on the same subject s∗s^*.

    The CKA score is defined as: CKA=P(o∗∣s∗,r∗)Er′∈RP(o∗∣s∗,r′)\text{CKA} = \frac{P(o^* \mid s^*, r^*)}{\mathbb{E}_{r' \in \mathcal{R}} P(o^* \mid s^*, r')}

    A record constitutes an over-editing failure if its CKA score is less than a threshold α\alpha, measured by the False Rate (P(CKA<α)P(\text{CKA} < \alpha)).

    Method CKA Score↑\uparrow False Rate (α=1.0\alpha=1.0)↓\downarrow False Rate (α=1.1\alpha=1.1)↓\downarrow
    FT 1.8 0.6% 19.5%
    ROME 1.7 0.4% 24.1%
    PROMPT 2.3 0.2% 1.0%
    IKE 2.1 0.1% 1.7%

    Parameter-updating methods suffer from substantial predicate leakage (ROME exhibits a 24.1% False Rate at α=1.1\alpha=1.1 and the lowest CKA score of 1.7), indicating over-generalization to unrelated properties of s∗s^*. In contrast, IKE achieves a CKA score of 2.1 and maintains a False Rate of 1.7% at α=1.1\alpha=1.1.

  9. Knowl 9 — Knowledge Forgetting and Sequential Temporal Editing Retention

    data/table

    Knowledge editing can induce catastrophic forgetting of original facts P(oc∣s∗,r∗)P(o^c \mid s^*, r^*). On COUNTERFACT with GPT-J (6B), probability drop is measured as ΔP(oc∣s∗,r)=Ppre(oc∣s∗,r)−Ppost(oc∣s∗,r)\Delta P(o^c \mid s^*, r) = P_{pre}(o^c \mid s^*, r) - P_{post}(o^c \mid s^*, r). An original fact is considered forgotten when ΔP(oc∣s∗,r∗)>0.5×Ppre(oc∣s∗,r∗)\Delta P(o^c \mid s^*, r^*) > 0.5 \times P_{pre}(o^c \mid s^*, r^*):

    Method Probability Drop↓\downarrow Forgetting Rate↓\downarrow
    FT 7.6 94.1%
    ROME 7.7 99.3%
    PROMPT 6.2 64.1%
    IKE 6.1 50.5%

    For sequential time-aware editing over temporal sequences (t1,s,r,ot1),…,(tn,s,r,otn)(t_1, s, r, o_{t_1}), \dots, (t_n, s, r, o_{t_n}) on TEMPLAMA (across relations member of sports team, position held, employer), the retention of the earliest injected fact ot1o_{t_1} after nn sequential edits is measured by the memorization ratio: Memorization Ratio=Pt=tn(ot1∣s,r,t1)Pt=t1(ot1∣s,r,t1)\text{Memorization Ratio} = \frac{P_{t=t_n}(o_{t_1} \mid s, r, t_1)}{P_{t=t_1}(o_{t_1} \mid s, r, t_1)}

    Method Memorization Ratio↑\uparrow
    ROME 0.08%
    IKE 88.0%

    ROME suffers near-complete catastrophic forgetting (0.08% retention) due to feed-forward network parameter overwrite collisions, whereas IKE retains 88.0% of the initial temporal knowledge state.

  10. Knowl 10 — Limitations of In-Context Knowledge Editing

    limitation

    The methodology and evaluation of In-Context Knowledge Editing (IKE) possess four primary limitations:

    1. Scope of Knowledge Types: Investigations are restricted exclusively to factual entity-relation-object triplets; editing other forms of knowledge, such as commonsense reasoning, is not addressed.
    2. Inference Context Overhead: Incorporating up to 32 demonstration examples into the prompt increases input sequence length, introducing additional memory and computational overhead during inference relative to standard zero-shot generation.
    3. Interpretability of ICL Mechanisms: While avoiding weight modifications eliminates parameter distortion, the internal operational mechanisms of in-context learning remain incompletely understood, leaving potential latent edge-case behaviors uncharacterized.
    4. Real-World Deployment Gap: In practical applications where facts and queries diverge in formatting, or when thousands of factual edits must be supported simultaneously, IKE requires integration with external retrieval databases and classifiers, which introduces retrieval error sensitivity.

Coverage note — None was omitted; the knowls comprehensively capture all task formulations, demonstration engineering strategies, algorithms, metric definitions, empirical benchmarks, ablation studies, model scaling analyses, over-editing assessments, memory retention evaluations, and limitations.

References

  1. 1.Sid Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, Michael Pieler, USVSN Sai Prashanth, Shivanshu Purohit, Laria Reynolds, Jonathan Tow, Ben Wang, and Samuel Weinbach. 2022. Gpt-neox-20b: An open-source autoregressive language model. CoRR, abs/2204.06745.
  2. 2.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  3. 3.Boxi Cao, Hongyu Lin, Xianpei Han, Le Sun, Lingyong Yan, Meng Liao, Tong Xue, and Jin Xu. 2021a. Knowledgeable or educated guess? revisiting language models as knowledge bases. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1860–1874, Online. Association for Computational Linguistics.
  4. 4.Nicola De Cao, Wilker Aziz, and Ivan Titov. 2021b. Editing factual knowledge in language models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, pages 6491–6506. Association for Computational Linguistics.
  5. 5.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harrison Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374.
  6. 6.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.
  7. 7.Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, and Furu Wei. 2022. Knowledge neurons in pretrained transformers. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022, pages 8493–8502. Association for Computational Linguistics.
  8. 8.Bhuwan Dhingra, Jeremy R. Cole, Julian Martin Eisenschlos, Daniel Gillick, Jacob Eisenstein, and William W. Cohen. 2022. Time-aware language models as temporal knowledge bases. Trans. Assoc. Comput. Linguistics, 10:257–273.
  9. 9.Qingxiu Dong, Damai Dai, Yifan Song, Jingjing Xu, Zhifang Sui, and Lei Li. 2022. Calibrating factual knowledge in pretrained language models. CoRR, abs/2210.03329.
  10. 10.Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Zhiyong Wu, Baobao Chang, Xu Sun, Jingjing Xu, Lei Li, and Zhifang Sui. 2023. A survey for in-context learning. CoRR, abs/2301.00234.
  11. 11.Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, Eduard Hovy, Hinrich Schütze, and Yoav Goldberg. 2021. Measuring and improving consistency in pretrained language models. Transactions of the Association for Computational Linguistics, 9:1012–1031.
  12. 12.FAIR, Anton Bakhtin, Noam Brown, Emily Dinan, Gabriele Farina, Colin Flaherty, Daniel Fried, Andrew Goff, Jonathan Gray, Hengyuan Hu, et al. 2022. Human-level play in the game of diplomacy by combining language models with strategic reasoning. Science (New York, NY), 378(6624):1067–1074.
  13. 13.Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. 2021. The pile: An 800gb dataset of diverse text for language modeling. CoRR, abs/2101.00027.
  14. 14.Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020. Realtoxicityprompts: Evaluating neural toxic degeneration in language models. In Findings of the Association for Computational Linguistics: EMNLP 2020, Online Event, 16-20 November 2020, volume EMNLP 2020 of Findings of ACL, pages 3356–3369. Association for Computational Linguistics.
  15. 15.Omer Levy, Minjoon Seo, Eunsol Choi, and Luke Zettlemoyer. 2017. Zero-shot relation extraction via reading comprehension. In Proceedings of the 21st Conference on Computational Natural Language Learning (CoNLL 2017), pages 333–342, Vancouver, Canada. Association for Computational Linguistics.
  16. 16.Jiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, Lawrence Carin, and Weizhu Chen. 2022. What makes good in-context examples for GPT-3? In Proceedings of Deep Learning Inside Out (DeeLIO 2022): The 3rd Workshop on Knowledge Extraction and Integration for Deep Learning Architectures, pages 100–114, Dublin, Ireland and Online. Association for Computational Linguistics.
  17. 17.Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. 2022. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8086–8098, Dublin, Ireland. Association for Computational Linguistics.
  18. 18.Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022a. Locating and editing factual associations in GPT. Advances in Neural Information Processing Systems, 35.
  19. 19.Kevin Meng, Arnab Sen Sharma, Alex Andonian, Yonatan Belinkov, and David Bau. 2022b. Mass-editing memory in a transformer. CoRR, abs/2210.07229.
  20. 20.Eric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn, and Christopher D. Manning. 2022a. Fast model editing at scale. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net.
  21. 21.Eric Mitchell, Charles Lin, Antoine Bosselut, Christopher D. Manning, and Chelsea Finn. 2022b. Memory-based model editing at scale. In International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA, volume 162 of Proceedings of Machine Learning Research, pages 15817–15831. PMLR.
  22. 22.Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al. 2021. Webgpt: Browser-assisted question-answering with human feedback. arXiv preprint arXiv:2112.09332.
  23. 23.OpenAI. 2023. Gpt-4 technical report.
  24. 24.TB OpenAI. 2022. Chatgpt: Optimizing language models for dialogue. OpenAI.
  25. 25.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744.
  26. 26.Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer. 2017. Automatic differentiation in pytorch.
  27. 27.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  28. 28.Nils Reimers and Iryna Gurevych. 2019. Sentence-bert: Sentence embeddings using siamese bert-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics.
  29. 29.Ohad Rubin, Jonathan Herzig, and Jonathan Berant. 2022. Learning to retrieve prompts for in-context learning. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2655–2671, Seattle, United States. Association for Computational Linguistics.
  30. 30.Emily Sheng, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng. 2019. The woman worked as a babysitter: On biases in language generation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, pages 3405–3410. Association for Computational Linguistics.
  31. 31.Chenglei Si, Zhe Gan, Zhengyuan Yang, Shuohang Wang, Jianfeng Wang, Jordan L. Boyd-Graber, and Lijuan Wang. 2022. Prompting GPT-3 to be reliable. CoRR, abs/2210.09150.
  32. 32.Tianxiang Sun, Yunfan Shao, Hong Qian, Xuanjing Huang, and Xipeng Qiu. 2022. Black-box tuning for language-model-as-a-service. arXiv preprint arXiv:2201.03514.
  33. 33.James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2018. FEVER: a large-scale dataset for fact extraction and verification. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2018, New Orleans, Louisiana, USA, June 1-6, 2018, Volume 1 (Long Papers), pages 809–819. Association for Computational Linguistics.
  34. 34.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023a. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
  35. 35.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023b. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  36. 36.Ben Wang and Aran Komatsuzaki. 2021. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https://github.com/kingoflolz/mesh-transformer-jax.
  37. 37.Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. 2021. Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652.
  38. 38.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, and Jamie Brew. 2019. Huggingface’s transformers: State-of-the-art natural language processing. CoRR, abs/1910.03771.
  39. 39.Yunzhi Yao, Peng Wang, Bo Tian, Siyuan Cheng, Zhoubo Li, Shumin Deng, Huajun Chen, and Ningyu Zhang. 2023. Editing large language models: Problems, methods, and opportunities. ArXiv, abs/2305.13172.
  40. 40.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona T. Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. 2022. OPT: open pre-trained transformer language models. CoRR, abs/2205.01068.
  41. 41.Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. Calibrate before use: Improving few-shot performance of language models. In Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research, pages 12697–12706. PMLR.

Citation

MLA
Zheng, C., et al. “Can We Edit Factual Knowledge by In-Context Learning?”. arXiv, 2023, http://arxiv.org/abs/2305.12740v1.
APA
Zheng, C., Li, L., Dong, Q., Fan, Y., Wu, Z., Xu, J., & Chang, B. (2023). Can We Edit Factual Knowledge by In-Context Learning?. arXiv. http://arxiv.org/abs/2305.12740v1
Chicago
Zheng, C., L. Li, Q. Dong, et al. 2023. “Can We Edit Factual Knowledge by In-Context Learning?”. arXiv. http://arxiv.org/abs/2305.12740v1.
Harvard
Zheng, C. et al. (2023) “Can We Edit Factual Knowledge by In-Context Learning?”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2305.12740v1.
Vancouver
1. Zheng C, Li L, Dong Q, Fan Y, Wu Z, Xu J, Chang B (2023) Can We Edit Factual Knowledge by In-Context Learning?. arXiv

BibTeX

@article{zheng2023can,
  title = {Can We Edit Factual Knowledge by In-Context Learning?},
  author = {Zheng, Ce and Li, Lei and Dong, Qingxiu and Fan, Yuxuan and Wu, Zhiyong and Xu, Jingjing and Chang, Baobao},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2305.12740v1},
  eprint = {2305.12740}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/