AnyEdit: Edit Any Knowledge Encoded in Language Models

Houcheng JiangJunfeng FangNingyu ZhangMingyang WanGuojun MaXiang WangXiangnan HeTat-Seng Chua

article2025ICML95 citations

Proposes AnyEdit, an autoregressive framework that breaks down complex, long-form knowledge into sequential chunks to iteratively edit key tokens, enabling existing model editing techniques to update diverse formats like code and mathematics without degrading accuracy over long outputs.

Listen

Large language models frequently generate outdated or inaccurate text, creating an urgent operational need for fast and reliable knowledge correction without the massive computational expense of retraining entire models. Existing editing techniques operate on simple factual statements by identifying and modifying a single internal token representation. However, this approach fails on long-form content and diverse formats—such as code, mathematical proofs, poetry, and multi-sentence explanations—creating an efficacy barrier where edits cannot reliably propagate across lengthy or complex outputs.

The article introduces and evaluates AnyEdit, a modular editing framework designed to update arbitrary knowledge within language models regardless of output length or format. The core objective is to demonstrate that an iterative, step-by-step editing process can overcome the limitations of traditional single-token editing methods.

To address this challenge, the authors developed a collaborative editing paradigm grounded in information theory. Instead of modifying an entire generation through a single point, AnyEdit splits long target outputs into smaller sequential segments. It iteratively pinpoints the final token of each segment and modifies its internal representation to ensure the model naturally generates the subsequent segment. The researchers evaluated this approach on prominent open-source language models (Llama3-8B-Instruct and Qwen2.5-7B-Instruct) across standard unstructured benchmarks and a newly created dataset containing texts up to 458 tokens across domains such as programming, mathematics, chemistry, and news.

The experimental findings show substantial performance advantages. AnyEdit consistently outperformed existing editing baselines, achieving an average accuracy improvement of 21.5% across tested benchmarks. On diverse knowledge types, the framework delivered major gains, improving text overlap metrics by 60.58% in programming code and 52.38% in news content compared to prior methods. Furthermore, the approach exhibited strong robustness, maintaining high accuracy when queries were rephrased and retaining stability as target lengths expanded well beyond the 100-token threshold where conventional techniques fail. When deployed as an enhancement to existing baseline algorithms, it boosted their editing quality with only a modest 24.7% increase in processing time per sample.

These results demonstrate that direct model editing can realistically handle complex, multi-sentence technical content rather than just isolated facts. For organizations operating language models, this capability reduces reliance on costly retraining cycles, mitigates hallucination risks in technical workflows, and supports rapid maintenance of dynamic enterprise knowledge bases.

Organizations seeking to implement model editing on long-form or non-triplet text should adopt sequential multi-segment editing rather than traditional single-token modifications. In operational settings, practitioners should select a balanced segment size—around 40 to 50 tokens—to balance editing accuracy with computational efficiency.

Confidence in these findings is high for text-based knowledge updates up to several hundred tokens. However, decision-makers should note key limitations: the current framework is not yet optimized for continuous lifelong updates where models undergo hundreds of sequential edits over extended periods, nor does it support multimodal inputs such as images or audio. Further research and piloting will be required to validate its long-term stability in continuous production environments.

arXiv: 2502.05628jianghoucheng/AnyEdit

No sufficiently relevant recommendations were found.

Cover for AnyEdit: Edit Any Knowledge Encoded in Language Models

Abstract

Large language models (LLMs) often produce incorrect or outdated information, necessitating efficient and precise knowledge updates. Current model editing methods, however, struggle with long-form knowledge in diverse formats, such as poetry, code snippets, and mathematical derivations. These limitations arise from their reliance on editing a single token’s hidden state, a limitation we term as “efficacy barrier”. To solve this, we propose AnyEdit, a new autoregressive editing paradigm. It decomposes long-form knowledge into sequential chunks and iteratively edits the key token in each chunk, ensuring consistent and accurate outputs. Theoretically, we ground AnyEdit in the Chain Rule of Mutual Information, showing its ability to update any knowledge within LLMs. Empirically, it outperforms strong baselines by 21.5% on benchmarks including UnKEBench, AKEW, and our new EditEverything dataset for long-form diverse-formatted knowledge. Additionally, AnyEdit serves as a plug-and-play framework, enabling current editing methods to update knowledge with arbitrary length and format, significantly advancing the scope and practicality of LLM knowledge editing. Our code is available at: https://github.com/jianghoucheng/AnyEdit.

Table of Contents

  • 1. Introduction
  • 2. Preliminary
  • 3. Limitations of Single-token Editing
  • 3.1. Editing Diverse-formatted Knowledge
  • 3.2. Editing Long-form Knowledge
  • 4. AnyEdit: Autoregressive Model Editing
  • 4.1. Theoretical Foundation
  • 4.2. Implementation Details
  • 5. Experiments
  • 5.1. Experimental Setup
  • 5.2. Long-Form Knowledge Editing (RQ1)
  • 5.3. Diverse-Formatted Knowledge Editing (RQ2)
  • 5.4. Boosting Current Editing Methods (RQ3)
  • 5.5. Impact of Chunk Size (RQ4)
  • 6. Related Work
  • 7. Conclusion & Limitations
  • Acknowledgments
  • References
  • A. Experimental Setup
  • A.1. Datasets
  • A.2. Evaluation Metrics
  • A.3. Baseline Methods
  • A.4. Implementation Details
  • B. Locate-Then-Edit Paradigm & Related Proof
  • B.1. Locate-Then-Edit Paradigm
  • B.2. Proof of Optimization-Conditional Mutual Information Equivalence
  • B.3. Proof of the Decomposition of Mutual Information
  • C. More Experimental Results
  • C.1. Case Study
  • C.1.1. CASE 1
  • C.2. Supplementary Experimental Results on RQ1 & RQ2
  • C.3. Supplementary Experimental Results on RQ4

Knowls

  1. Knowl 1 — AnyEdit edits long outputs sequentially by chunk

    algorithm

    AnyEdit takes a prompt XX, a desired long-form response YY, a decoder-only language model, and a locate-then-edit parameter-update method as inputs; it returns a model updated to generate YY in order. It divides YY into chunks, either using a fixed-width sliding window or semantic sentence boundaries. For each chunk YkY_k, it selects the chunk’s final token and uses causal tracing to identify influential model layers. It then conditions on XX and all earlier target chunks Y1,…,Yk−1Y_1,\ldots,Y_{k-1}, optimizes the selected token’s hidden-state residual by gradient descent to increase the likelihood of generating YkY_k, and updates model parameters to align the selected hidden state with its optimized state using the editor’s parameter-update procedure. It repeats until every chunk has been processed. This sequential conditioning is intended to avoid conflicting simultaneous edits and to preserve coherence across chunks.

    In the reported configurations, sliding-window chunks had no overlap: the chunk size was 40 tokens for Llama3-8B-Instruct and 50 for Qwen2.5-7B-Instruct. For both models, hidden-state optimization used 25 gradient steps at learning rate 0.5, weight decay 0.001, and a KL-regularization factor of 0.0625; the loss was applied at layer 31 for Llama3 and layer 27 for Qwen2.5. The selected editing layers were 4–8, with a clamp-norm factor of 4. The paper does not state a computational-complexity bound.

  2. Knowl 2 — Autoregressive chunk editing gives a single-state conditional mutual-information objective

    theoretical result

    Let XX be an input prompt and let the target response YY be divided, in generation order, into chunks Y1,…,YKY_1,\ldots,Y_K. Let hk′h'_k be the edited hidden state of the final token in chunk YkY_k, and let I(A;B∣C)I(A;B\mid C) denote conditional mutual information. AnyEdit’s theoretical argument uses the causal property of autoregressive generation that later-token hidden states do not affect earlier chunks, and the premise that conditioning on a generated chunk subsumes conditioning on the edited state associated with that chunk. Under these premises, the paper expresses the editing objective as

    I(X;Y∣h1′,…,hK′)=∑k=1KI(X,Y1,…,Yk−1;Yk∣hk′).I(X;Y\mid h'_1,\ldots,h'_K)=\sum_{k=1}^{K} I(X,Y_1,\ldots,Y_{k-1};Y_k\mid h'_k).

    The result motivates editing one chunk at a time: each term conditions on a single edited hidden state and on the prompt plus already specified target chunks, rather than jointly conditioning on multiple perturbed states. AnyEdit uses this decomposition as a rationale for reducing interference and applying the same editing strategy to outputs of different lengths and formats; the claim depends on the stated autoregressive and conditioning premises.

  3. Knowl 3 — Single-token editing has a proposed efficacy barrier

    theoretical result

    For NN requested edits to a language model ff, let (Xi,Yi)(X_i,Y_i) be the iith prompt and desired output, let hih_i be the hidden state of its perturbed token, and let Y\mathcal{Y} be the set of possible model outputs. The paper formalizes the efficacy barrier of single-token editing using the average

    η=1N∑i=1Nsign⁡ ⁣(Pf(hi+δi)(Yi∣Xi)−max⁡Y′∈Y∖{Yi}Pf(hi+δi)(Y′∣Xi)),\eta=\frac{1}{N}\sum_{i=1}^{N}\operatorname{sign}\!\left(P_{f(h_i+\delta_i)}(Y_i\mid X_i)-\max_{Y'\in\mathcal{Y}\setminus\{Y_i\}}P_{f(h_i+\delta_i)}(Y'\mid X_i)\right),

    where Pf(hi+δi)(⋅∣Xi)P_{f(h_i+\delta_i)}(\cdot\mid X_i) is the output distribution after perturbing the selected hidden state, and δi\delta_i is chosen to minimize the negative log-likelihood of YiY_i given XiX_i. The sign function returns positive, zero, or negative according to the sign of its argument, so the expression tracks whether the desired output is more probable than every alternative. The paper presents this as an upper-bound formulation for single-token editing and argues that its efficacy deteriorates for knowledge that is difficult to generate initially or that requires long, interdependent outputs.

  4. Knowl 4 — Single-token editing struggles with low-probability and long targets

    empirical result

    The paper tested MEMIT on Llama3-8B-Instruct to examine how target format and length relate to single-token editing efficacy. For the format comparison, it sampled 200 knowledge instances per category and found that diverse-formatted targets, including code and mathematical derivations, tended to have lower probability before editing and poorer editing efficacy than triplet-structured targets. For the length analysis, it truncated sampled targets to different token lengths and used causal tracing to measure the output-probability shift caused by random perturbations. As target length increased, the probability shift and editing efficacy generally declined. These experiments support the paper’s diagnosis that a single hidden-state edit has difficulty overcoming low initial target probability and influencing distant parts of a long output; they show associations under these experimental conditions, not a universal law.

  5. Knowl 5 — EditEverything covers long, diverse-format targets

    experimental setup

    EditEverything is a dataset introduced to evaluate editing of long-form knowledge in diverse formats. It combines mathematics samples selected from Orca-Math, Python programming problems from MBPP, and chemistry problem–solution pairs from Camel-Chemistry. Its news and poetry examples were synthetically generated with GPT-4o so that the information would not already be known to the tested models. The dataset covers mathematics, code, chemistry, news, and poetry, and target responses reach 458 tokens. Evaluation compares edited generations with target responses for original questions, paraphrased questions, and subquestions, using BERT Score for semantic similarity and ROUGE scores for lexical similarity.

  6. Knowl 6 — AnyEdit improves long-form editing, with results varying by baseline and metric

    empirical result

    On UnKEBench, the paper compared editing methods on Llama3-8B-Instruct and Qwen2.5-7B-Instruct. Scores below are reported on a 0–100 scale as BERT Score and ROUGE-L, respectively; within each pair, original-question results precede paraphrase-question results.

    • Llama3-8B-Instruct: AnyEdit scored 97.76±0.11 and 92.96±0.24 on original questions, and 96.60±0.19 and 95.60±0.35 on paraphrases. MEMIT scored 76.21±0.36 and 30.49±0.52, and 74.25±0.31 and 28.65±0.61, respectively. UnKE scored 98.34±0.15 and 93.33±0.26, and 93.38±0.21 and 78.42±0.32. AnyEdit*—the variant that updates all parameters in the selected layer by gradient descent—scored 99.86±0.08 and 99.68±0.21, and 94.70±0.12 and 85.75±0.23.
    • Qwen2.5-7B-Instruct: AnyEdit scored 98.05±0.16 and 94.89±0.29 on original questions, and 93.56±0.15 and 79.98±0.28 on paraphrases. MEMIT scored 78.03±0.30 and 38.04±0.47, and 76.50±0.31 and 28.65±0.50. UnKE scored 96.97±0.18 and 91.01±0.24, and 89.17±0.15 and 67.00±0.29. AnyEdit* scored 99.35±0.12 and 98.82±0.24, and 94.81±0.13 and 82.60±0.26.

    These results show large gains over MEMIT on this benchmark and strong paraphrase performance for AnyEdit. They also show that UnKE or AnyEdit* exceeds AnyEdit on some individual metrics, so the advantage is not uniform across all methods and measures. The long-form experiments used batch size 1 and decoding temperature 0.001.

  7. Knowl 7 — AnyEdit performs well across diverse formats and target lengths

    empirical result

    On EditEverything, AnyEdit was evaluated against editing baselines across mathematics, poetry, news, code, and chemistry using lexical and semantic similarity. The paper reports its largest ROUGE-L gains over baselines in code and news: 60.58% and 52.38%, respectively. In a separate analysis of long targets, MEMIT and AlphaEdit showed marked performance declines when target length exceeded 30 tokens, while AnyEdit remained more stable as the number of target tokens increased. The plotted length comparison extends to roughly 200 tokens; the dataset itself includes targets as long as 458 tokens. The reported evidence supports robustness over the tested lengths and formats, rather than establishing that performance is invariant to length.

  8. Knowl 8 — The autoregressive procedure boosts existing editors at modest added time

    empirical result

    The paper integrated AnyEdit’s autoregressive editing procedure with MEMIT, AlphaEdit, and UnKE, then evaluated the modified methods on UnKEBench. ROUGE-L scores below are original-question/paraphrase-question results, with unmodified results followed by the AnyEdit-integrated results:

    • Llama3-8B-Instruct: MEMIT 30.49/28.65, MEMIT+ 92.96/85.91; AlphaEdit 26.59/25.92, AlphaEdit+ 80.92/72.22; UnKE 93.33/78.42, UnKE+ 99.68/85.75.
    • Qwen2.5-7B-Instruct: MEMIT 38.04/34.07, MEMIT+ 94.89/79.98; AlphaEdit 42.77/38.26, AlphaEdit+ 98.14/83.40; UnKE 91.01/67.00, UnKE+ 98.82/82.60.

    The authors report that integration increased editing time by an average of 24.7% across the evaluated settings. For example, on UnKEBench, MEMIT took 16.37 seconds per sample and MEMIT+ took 21.14 seconds; AlphaEdit took 17.25 seconds and AlphaEdit+ took 22.03 seconds; UnKE took 22.91 seconds and UnKE+ took 28.32 seconds. These results demonstrate substantial performance improvements in the reported comparisons alongside added editing time.

  9. Knowl 9 — Chunk size trades editing quality against iteration time

    empirical result

    In a sliding-window analysis on long-form diverse-formatted targets, AnyEdit’s ROUGE-L performance declined as chunk size increased beyond a threshold. Smaller chunks made each individual edit more manageable, but required more editing iterations and therefore more time on long responses. Larger chunks reduced the number of iterations but made each edit harder to achieve effectively. The paper recommends a balanced chunk size of 40 tokens for most scenarios; its main experiments used 40 tokens for Llama3-8B-Instruct and 50 for Qwen2.5-7B-Instruct.

  10. Knowl 10 — AnyEdit is limited to textual editing and is not optimized for lifelong updates

    limitation

    The paper identifies two limitations of AnyEdit. First, the framework is not explicitly optimized for lifelong editing, where knowledge must be updated continuously over time; adapting it to repeated, dynamic updates remains unresolved. Second, the method handles textual knowledge only and does not support multimodal edits, such as coordinated updates across text, images, and audio.

Coverage note — The supplementary qualitative case studies and exhaustive metric tables for AKEW and subquestion evaluations are omitted because they provide additional examples and measurements rather than distinct contributions beyond the reported benchmark results.

References

  1. 1.Austin, J., Odena, A., Nye, M. I., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C. J., Terry, M., Le, Q. V., and Sutton, C. Program synthesis with large language models. CoRR, abs/2108.07732, 2021.
  2. 2.Bi, B., Liu, S., Mei, L., Wang, Y., Ji, P., and Cheng, X. Decoding by contrasting knowledge: Enhancing llms’ confidence on edited facts. CoRR, abs/2405.11613, 2024.
  3. 3.Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. Language models are few-shot learners. In NeurIPS, 2020.
  4. 4.Cao, N. D., Aziz, W., and Titov, I. Editing factual knowledge in language models. In EMNLP (1), pp. 6491–6506. Association for Computational Linguistics, 2021.
  5. 5.Chen, C., Huang, B., Li, Z., Chen, Z., Lai, S., Xu, X., Gu, J., Gu, J., Yao, H., Xiao, C., Yan, X., Wang, W. Y., Torr, P., Song, D., and Shu, K. Can editing llms inject harm? CoRR, abs/2407.20224, 2024.
  6. 6.Deng, J., Wei, Z., Pang, L., Ding, H., Shen, H., and Cheng, X. Everything is editable: Extend knowledge editing to unstructured data in large language models. In ICLR. OpenReview.net, 2025.
  7. 7.Dobrushin, R. L. General formulation of shannon’s main theorem in information theory. American mathematical society translations, 33:323–438, 1963.
  8. 8.Dong, Q., Dai, D., Song, Y., Xu, J., Sui, Z., and Li, L. Calibrating factual knowledge in pretrained language models. In EMNLP (Findings), pp. 5937–5947. Association for Computational Linguistics, 2022.
  9. 9.Fang, J., Jiang, H., Wang, K., Ma, Y., Shi, J., Wang, X., He, X., and Chua, T. Alphaedit: Null-space constrained knowledge editing for language models. In ICLR. OpenReview.net, 2025.
  10. 10.Geva, M., Schuster, R., Berant, J., and Levy, O. Transformer feed-forward layers are key-value memories. In EMNLP (1), pp. 5484–5495. Association for Computational Linguistics, 2021.
  11. 11.Gu, J., Xu, H., Ma, J., Lu, P., Ling, Z., Chang, K., and Peng, N. Model editing harms general abilities of large language models: Regularization to the rescue. In EMNLP, pp. 16801–16819. Association for Computational Linguistics, 2024.
  12. 12.Hartvigsen, T., Sankaranarayanan, S., Palangi, H., Kim, Y., and Ghassemi, M. Aging with GRACE: lifelong model editing with discrete key-value adaptors. In NeurIPS, 2023.
  13. 13.Huang, B., Chen, C., Xu, X., Payani, A., and Shu, K. Can knowledge editing really correct hallucinations? In ICLR. OpenReview.net, 2025.
  14. 14.Huang, X., Wang, Y., Zhao, J., and Liu, K. Commonsense knowledge editing based on free-text in llms. In EMNLP, pp. 14870–14880. Association for Computational Linguistics, 2024.
  15. 15.Huang, Z., Shen, Y., Zhang, X., Zhou, J., Rong, W., and Xiong, Z. Transformer-patcher: One mistake worth one neuron. In ICLR. OpenReview.net, 2023.
  16. 16.Jiang, H., Fang, J., Zhang, T., Zhang, A., Wang, R., Liang, T., and Wang, X. Neuron-level sequential editing for large language models. CoRR, abs/2410.04045, 2024.
  17. 17.Kullback, S. Information theory and statistics. Courier Corporation, 1997.
  18. 18.Lang, S. Introduction to linear algebra. Springer Science & Business Media, 2012.
  19. 19.Li, G., Hammoud, H. A. A. K., Itani, H., Khizbullin, D., and Ghanem, B. CAMEL: communicative agents for ”mind” exploration of large scale language model society. CoRR, abs/2303.17760, 2023.
  20. 20.Li, Z., Jiang, H., Chen, H., Bi, B., Zhou, Z., Sun, F., Fang, J., and Wang, X. Reinforced lifelong editing for language models. CoRR, abs/2502.05759, 2025.
  21. 21.Lin, C.-Y. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pp. 74–81, 2004.
  22. 22.Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., and Liang, P. Lost in the middle: How language models use long contexts. Trans. Assoc. Comput. Linguistics, 12:157–173, 2024a.
  23. 23.Liu, Y., He, X., Xiong, M., Fu, J., Deng, S., and Hooi, B. Flipattack: Jailbreak llms via flipping. arXiv preprint arXiv:2410.02832, 2024b.
  24. 24.Meng, K., Bau, D., Andonian, A., and Belinkov, Y. Locating and editing factual associations in GPT. In NeurIPS, 2022.
  25. 25.Meng, K., Sharma, A. S., Andonian, A. J., Belinkov, Y., and Bau, D. Mass-editing memory in a transformer. In ICLR. OpenReview.net, 2023.
  26. 26.Mitchell, E., Lin, C., Bosselut, A., Finn, C., and Manning, C. D. Fast model editing at scale. In ICLR. OpenReview.net, 2022a.
  27. 27.Mitchell, E., Lin, C., Bosselut, A., Manning, C. D., and Finn, C. Memory-based model editing at scale. In ICML, volume 162 of Proceedings of Machine Learning Research, pp. 15817–15831. PMLR, 2022b.
  28. 28.Mitra, A., Khanpour, H., Rosset, C., and Awadallah, A. Orca-math: Unlocking the potential of slms in grade school math. CoRR, abs/2402.14830, 2024.
  29. 29.Papineni, K., Roukos, S., Ward, T., and Zhu, W. Bleu: a method for automatic evaluation of machine translation. In ACL, pp. 311–318. ACL, 2002.
  30. 30.Qinggang Zhang, Hao Chen, J. D. S. C. F. H. X. H. Structure-guided large language models for text-to-SQL generation. In Forty-second International Conference on Machine Learning, 2025.
  31. 31.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  32. 32.Wang, B. and Komatsuzaki, A. Gpt-j-6b: A 6 billion parameter autoregressive language model, 2021.
  33. 33.Wang, P., Li, Z., Zhang, N., Xu, Z., Yao, Y., Jiang, Y., Xie, P., Huang, F., and Chen, H. Wise: Rethinking the knowledge memory for lifelong model editing of large language models. arXiv preprint arXiv:2405.14768, 2024.
  34. 34.Wu, X., Pan, L., Wang, W. Y., and Luu, A. T. AKEW: assessing knowledge editing in the wild. In EMNLP, pp. 15118–15133. Association for Computational Linguistics, 2024.
  35. 35.Xie, J., Zhang, K., Chen, J., Lou, R., and Su, Y. Adaptive chameleon or stubborn sloth: Unraveling the behavior of large language models in knowledge clashes. CoRR, abs/2305.13300, 2023.
  36. 36.Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D., Huang, F., Wei, H., et al. Qwen2. 5 technical report. arXiv preprint arXiv:2412.15115, 2024.
  37. 37.Zhang, N., Tian, B., Cheng, S., Liang, X., Hu, Y., Xue, K., Gou, Y., Chen, X., and Chen, H. Instructedit: Instruction-based knowledge editing for large language models. In IJCAI, pp. 6633–6641. ijcai.org, 2024.
  38. 38.Zhang, T., Fang, J., Jiang, H., Bi, B., Wang, X., and He, X. Explainable and efficient editing for large language models. In THE WEB CONFERENCE 2025.
  39. 39.Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., and Artzi, Y. Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675, 2019.
  40. 40.Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., Min, Y., Zhang, B., Zhang, J., Dong, Z., Du, Y., Yang, C., Chen, Y., Chen, Z., Jiang, J., Ren, R., Li, Y., Tang, X., Liu, Z., Liu, P., Nie, J.-Y., and Wen, J.-R. A survey of large language models, 2024. URL https://arxiv.org/abs/2303.18223.
  41. 41.Zheng, C., Li, L., Dong, Q., Fan, Y., Wu, Z., Xu, J., and Chang, B. Can we edit factual knowledge by in-context learning? In EMNLP, pp. 4862–4876. Association for Computational Linguistics, 2023.
  42. 42.Zhu, C., Rawat, A. S., Zaheer, M., Bhojanapalli, S., Li, D., Yu, F. X., and Kumar, S. Modifying memories in transformer models. CoRR, abs/2012.00363, 2020.

Citation

MLA
Jiang, H., et al. “AnyEdit: Edit Any Knowledge Encoded in Language Models”. arXiv, 2025, https://doi.org/10.48550/arxiv.2502.05628.
APA
Jiang, H., Fang, J., Zhang, N., Ma, G., Wan, M., Wang, X., He, X., & Chua, T.-. seng . (2025). AnyEdit: Edit Any Knowledge Encoded in Language Models. arXiv. https://doi.org/10.48550/arxiv.2502.05628
Chicago
Jiang, H., J. Fang, N. Zhang, et al. 2025. “AnyEdit: Edit Any Knowledge Encoded in Language Models”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2502.05628.
Harvard
Jiang, H. et al. (2025) “AnyEdit: Edit Any Knowledge Encoded in Language Models”. arXiv. Available at: https://doi.org/10.48550/arxiv.2502.05628.
Vancouver
1. Jiang H, Fang J, Zhang N, Ma G, Wan M, Wang X, He X, Chua T-seng (2025) AnyEdit: Edit Any Knowledge Encoded in Language Models. https://doi.org/10.48550/arxiv.2502.05628

BibTeX

@misc{https://doi.org/10.48550/arxiv.2502.05628,
  doi = {10.48550/ARXIV.2502.05628},
  url = {https://arxiv.org/abs/2502.05628},
  author = {Jiang, Houcheng and Fang, Junfeng and Zhang, Ningyu and Ma, Guojun and Wan, Mingyang and Wang, Xiang and He, Xiangnan and Chua, Tat-seng},
  keywords = {Computation and Language (cs.CL), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {AnyEdit: Edit Any Knowledge Encoded in Language Models},
  publisher = {arXiv},
  year = {2025},
  copyright = {Creative Commons Attribution 4.0 International}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/