PMET: Precise Model Editing in a Transformer

Xiaopeng LiShasha LiShezheng SongJing YangJun MaJie Yu

article2024AAAI250 citations

Proposes PMET, a model editing framework that jointly optimizes attention and feed-forward hidden states while selectively updating feed-forward weights, preventing the transfer of irrelevant attention signals for more accurate factual updates in large language models.

Listen

Large language models often output outdated or incorrect facts, but retraining or full fine-tuning to correct minor errors is computationally expensive and slow. Recent model editing methods modify internal weights directly to update facts cheaply without full retraining. However, existing techniques conflate the information flows of internal subcomponents by using combined layer representations to update feed-forward network modules. This causes imprecise weight adjustments, degrades editing reliability, and risks unintended side effects on unrelated knowledge.

The article develops and evaluates Precise Model Editing in a Transformer (PMET), an optimization method designed to accurately update factual knowledge in language models. PMET isolates subcomponent representations to optimize target facts while restricting weight modifications exclusively to feed-forward network parameters.

To understand internal roles, the researchers analyzed information flow across multi-head self-attention and feed-forward subcomponents using 1,209 factual queries on a 6-billion parameter language model. Based on findings that attention mechanisms act primarily as extractors while feed-forward networks store factual associations, the researchers designed PMET to simultaneously optimize hidden representations for both components while applying updates strictly to the feed-forward network across critical layers. They evaluated this method against leading baselines on benchmark datasets containing up to 10,000 edits across 6-billion and 20-billion parameter models.

The evaluation yielded several key findings. First, internal analysis confirmed that attention subcomponents encode general knowledge extraction patterns rather than primary factual storage, meaning attention weights do not require updates during factual edits. Second, PMET achieved state-of-the-art overall editing performance, improving reliability on the primary counterfactual benchmark by an average of 3.3% over the prior leading method and boosting generalization by 4.2 percentage points on the 6-billion model. Third, scaling tests to 10,000 edits on a 20-billion parameter model showed PMET maintaining superior editing accuracy (98.4%) and generalization (89.4%) compared to earlier optimization baselines. Fourth, ablation studies confirmed that optimizing both component representations while updating only feed-forward weights struck the best balance across editing success, fluent generation, and preservation of unrelated facts.

These findings demonstrate that targeted, mathematically precise weight adjustments can correct factual errors at scale without retraining models from scratch. Organizations deploying large language models can leverage component-specific editing to significantly reduce operational compute costs, accelerate update cycles, and fix factual inaccuracies. Because PMET avoids modifying attention weights, it lowers the risk of corrupting general extraction capabilities during large-scale factual updates.

Decision-makers and engineering teams seeking to maintain accurate production models should adopt component-specific optimization over legacy fine-tuning methods for batch factual updates. However, teams should balance trade-offs, as prioritizing maximum editing reliability can slightly reduce specificity on unrelated facts. Organizations should implement validation protocols to verify that critical baseline behaviors remain intact after mass edits.

A primary limitation is that post-edited models often struggle to perform complex multi-step reasoning over newly inserted facts, meaning the updated knowledge is recalled but not deeply internalized. Additionally, direct editing tools carry safety risks if misused to inject false information, necessitating governance frameworks before broad operational deployment. Overall, confidence in PMET's performance is high for factual updates across modern model architectures, provided organizations recognize the boundary condition regarding downstream reasoning.

Cover for PMET: Precise Model Editing in a Transformer

Abstract

Model editing techniques modify a minor proportion of knowledge in Large Language Models (LLMs) at a relatively low cost, which have demonstrated notable success. Existing methods assume Transformer Layer (TL) hidden states are values of key-value memories of the Feed-Forward Network (FFN). They usually optimize the TL hidden states to memorize target knowledge and use it to update the weights of the FFN in LLMs. However, the information flow of TL hidden states comes from three parts: Multi-Head Self-Attention (MHSA), FFN, and residual connections. Existing methods neglect the fact that the TL hidden states contains information not specifically required for FFN. Consequently, the performance of model editing decreases. To achieve more precise model editing, we analyze hidden states of MHSA and FFN, finding that MHSA encodes certain general knowledge extraction patterns. This implies that MHSA weights do not require updating when new knowledge is introduced. Based on above findings, we introduce PMET, which simultaneously optimizes Transformer Component (TC, namely MHSA and FFN) hidden states, while only using the optimized TC hidden states of FFN to precisely update FFN weights. Our experiments demonstrate that PMET exhibits state-of-the-art performance on both the COUNTERFACT and zsRE datasets. Our ablation experiments substantiate the effectiveness of our enhancements, further reinforcing the finding that the MHSA encodes certain general knowledge extraction patterns and indicating its storage of a small amount of factual knowledge. Our code is available at https://github.com/xpq-tech/PMET.

Table of Contents

  • Introduction
  • Related Work
  • Model Editing
  • Post-Hoc Explanation of Transformers
  • Methodology
  • Preliminaries
  • Investigating the Role of MHSA and FFN in LLMs' Knowledge Recall
  • PMET Method
  • Experiments
  • Baselines and Datasets
  • Editing Experiments
  • Ablation Study
  • Conclusion
  • Limitations
  • Ethical Statement
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Subject-centric formulation of model editing

    definition

    PMET defines model editing around a subject rather than around isolated subject–relation–object triples. Let FθF_\theta be an autoregressive language model, let SS be a subject, and let KS={(xiS,yiS)}i=1N\mathcal K_S=\{(x_i^S,y_i^S)\}_{i=1}^{N} denote NN existing knowledge items about SS, where xiSx_i^S is a clue sequence and yiSy_i^S is the corresponding knowledge-point sequence. The edit replaces N′N' of these items with target pairs KSt={(xiS,t,yiS,t)}i=1N′\mathcal K_S^t=\{(x_i^{S,t},y_i^{S,t})\}_{i=1}^{N'}, while preserving M≫N′M\gg N' unrelated items K¬St={(xj¬S,yj¬S)}j=1M\mathcal K_{\neg S}^t=\{(x_j^{\neg S},y_j^{\neg S})\}_{j=1}^{M}.

    The desired edited model Fθ∗F_{\theta^*} should satisfy

    Fθ∗(xiS,t)=yiS,tandFθ∗(xj¬S)=yj¬SF_{\theta^*}(x_i^{S,t})=y_i^{S,t}\quad\text{and}\quad F_{\theta^*}(x_j^{\neg S})=y_j^{\neg S}

    for every edited item ii and unrelated item jj. The subject-centric formulation is intended to make the edited model preserve enough subject-associated information to reason over the edited knowledge, rather than merely memorize the exact edited prompts.

  2. Knowl 2 — MHSA acts primarily as a knowledge extractor

    empirical result

    The paper analyzes the last-token hidden states before and after each Multi-Head Self-Attention (MHSA) and Feed-Forward Network (FFN) block in GPT-J (6B), using 1,209 factual statements. For each layer, it measures cosine similarity between the two hidden states and the Jaccard similarity between their top-50 vocabulary-token projections. If Tk(h)T_k(h) is the set of the kk highest-scoring vocabulary tokens obtained by projecting hidden state hh, the vocabulary-space similarity is

    Jk(Tk(h1),Tk(h2))=∣Tk(h1)∩Tk(h2)∣∣Tk(h1)∪Tk(h2)∣,k=50.J_k\big(T_k(h_1),T_k(h_2)\big)=\frac{|T_k(h_1)\cap T_k(h_2)|}{|T_k(h_1)\cup T_k(h_2)|},\qquad k=50.

    Both components change frequently during approximately the first 15 layers. After that point, FFN hidden states change more slowly and gradually stabilize in a particular direction, whereas MHSA hidden states continue to change frequently and do not settle on a stable direction during knowledge recall. The authors interpret this as evidence that FFN representations become consistent with the knowledge they retrieve, while MHSA continuously extracts different types of information from the context. Combined with their ablations, the result supports the view that MHSA encodes general knowledge-extraction patterns and stores only a small amount of factual knowledge; therefore, introducing new facts should not normally require changing MHSA weights.

  3. Knowl 3 — PMET separates component optimization from weight updating

    model/method

    PMET treats the MHSA and FFN hidden states in a Transformer layer as separate Transformer-Component (TC) representations. For a target clue, let aiLa_i^L and miLm_i^L be the last-token MHSA and FFN hidden states at the last critical layer LL, respectively. PMET introduces independent optimizable shifts δia\delta_i^a and δim\delta_i^m:

    aiL′=aiL+δia,miL′=miL+δim.a_i^{L\prime}=a_i^L+\delta_i^a,\qquad m_i^{L\prime}=m_i^L+\delta_i^m.

    The shifts are optimized jointly so that the two components can represent the target knowledge. However, PMET retains only the optimized FFN representation

    vim=miL+δimv_i^m=m_i^L+\delta_i^m

    as the target value for the weight update. The optimized MHSA representation is used only to expand the hidden-state optimization space; PMET does not update MHSA weights. Consequently, the subsequent weight modification is restricted to FFN output weights, avoiding the unrelated information carried by full Transformer-Layer hidden states and avoiding damage to MHSA's general extraction patterns.

  4. Knowl 4 — Target-value optimization balances reliability and specificity

    equation

    For each edited knowledge item, PMET obtains the optimized FFN target value vim=miL+δimv_i^m=m_i^L+\delta_i^m by minimizing a loss that combines preservation of the original output distribution with likelihood of the desired target sequence. Let FθF_\theta be the original model, Fθ†F_\theta^\dagger the model with δia\delta_i^a and δim\delta_i^m inserted at the critical layer, yiS,ty_i^{S,t} the desired target sequence, p′p' the fixed prompt template `{S} is a', p(xi)p(x_i) the prompt formed from clue xix_i, and pref⁡j\operatorname{pref}_j one of PP prefixes used to improve generalization. With ⊕\oplus denoting sequence concatenation, the optimization objective is

    L(vim)=μ DKL ⁣(PFθ†(y∣p′) ∥ PFθ(y∣p′))+ϕP∑j=1P[−log⁡PFθ† ⁣(yiS,t∣pref⁡j⊕p(xi))].\mathcal L(v_i^m)=\mu\,D_{\mathrm{KL}}\!\left(P_{F_\theta^\dagger}(y\mid p')\,\middle\|\,P_{F_\theta}(y\mid p')\right) +\frac{\phi}{P}\sum_{j=1}^{P}\left[-\log P_{F_\theta^\dagger}\!\left(y_i^{S,t}\mid \operatorname{pref}_j\oplus p(x_i)\right)\right].

    Here DKLD_{\mathrm{KL}} is the Kullback–Leibler divergence between output-sequence distributions, μ\mu weights preservation of the original behavior, and ϕ\phi weights fitting the target knowledge. Optimizing both component shifts while retaining only vimv_i^m for the FFN update is the mechanism by which PMET separates target representation construction from precise parameter modification.

  5. Knowl 5 — Closed-form FFN update with square-root residual spreading

    equation

    PMET updates an FFN weight by treating its input hidden states as keys and its output representations as values. Let W0∈Rdo×dkW_0\in\mathbb R^{d_o\times d_k} be the original FFN weight, W1=W0+ΔW_1=W_0+\Delta the edited weight, ki∈Rdkk_i\in\mathbb R^{d_k} the key for target item ii, and vi∈Rdov_i\in\mathbb R^{d_o} its optimized FFN target value. Stack the keys and values as K1=[k1∣⋯∣kn]K_1=[k_1|\cdots|k_n] and V1=[v1∣⋯∣vn]V_1=[v_1|\cdots|v_n]. The least-squares update has the closed form

    Δ=RK1T(C0+K1K1T)−1,R=V1−W0K1,\Delta=R K_1^{\mathsf T}\left(C_0+K_1K_1^{\mathsf T}\right)^{-1}, \qquad R=V_1-W_0K_1,

    where RR is the target residual and C0=λ E[kkT]C_0=\lambda\,\mathbb E[kk^{\mathsf T}] estimates the covariance of previously memorized keys from sampled model inputs; λ\lambda controls the trade-off between modifying target knowledge and preserving existing knowledge.

    For a set L\mathcal L of critical layers, PMET distributes the residual using a square-root schedule. If L∗=max⁡(L)L_*=\max(\mathcal L) and W0lW_0^l is the original FFN weight at critical layer ll, the residual assigned to that layer is

    Rl=V1−W0lK1L∗−l+1.R^l=\frac{V_1-W_0^lK_1}{\sqrt{L_*-l+1}}.

    The key for a weight is the average hidden state immediately before that weight over PP prefixed subject prompts:

    kil=1P∑j=1Pprev⁡ ⁣(Wl,pref⁡j⊕S),k_i^l=\frac{1}{P}\sum_{j=1}^{P}\operatorname{prev}\!\left(W^l,\operatorname{pref}_j\oplus S\right),

    where prev⁡(Wl,x)\operatorname{prev}(W^l,x) denotes the hidden state just before weight WlW^l processes input xx. For an FFN output weight WO,FFNlW_{O,\mathrm{FFN}}^l, this preceding state is σ(WIlγ(hl−1(x)))\sigma(W_I^l\gamma(h^{l-1}(x))), with WIlW_I^l the FFN input weight, hl−1(x)h^{l-1}(x) the previous-layer hidden state, γ\gamma layer normalization, and σ\sigma the FFN activation. The resulting updates are applied to FFN weights only.

  6. Knowl 6 — Evidence for PMET on 10,000 COUNTERFACT edits

    data/table

    The paper evaluates 10,000 counterfactual edits on GPT-J (6B) and GPT-NeoX (20B), comparing PMET with fine-tuning, MEND, ROME, and MEMIT. Efficacy and generalization use counterfactual targets, specificity uses factual information, and score is the harmonic mean of efficacy, generalization, and specificity. Fluency and consistency measure generation quality. Numbers in parentheses are 95% confidence intervals.

    PMET has the highest score on both models. On GPT-J, it improves generalization, fluency, and consistency over MEMIT while being slightly less specific. On GPT-NeoX, it improves reliability-related metrics and consistency over MEMIT but again has slightly lower specificity.

    Could not parse LaTeX table

    Across 17 logarithmically spaced COUNTERFACT edit counts, PMET outperformed the baselines on every reported metric except specificity, where MEMIT was slightly better. The paper reports a 3.3% average reliability improvement over the previous state of the art on COUNTERFACT.

  7. Knowl 7 — Evidence for PMET on 10,000 zsRE edits

    data/table

    On GPT-J (6B), PMET is evaluated on 10,000 counterfactual edits from the Zero-Shot Relation Extraction (zsRE) dataset. It achieves the best efficacy, generalization, and specificity among the compared editors. The unedited model has low efficacy and generalization but high specificity; PMET raises efficacy and generalization while preserving specificity at approximately the original model's level.

    Could not parse LaTeX table

    PMET's reported average improvement in reliability over the previous state of the art on zsRE is 0.4%.

  8. Knowl 8 — Ablations validate component optimization and residual spreading

    empirical result

    Ablations on GPT-J (6B) with COUNTERFACT edits test three PMET design choices: optimizing both MHSA and FFN hidden states, updating FFN rather than MHSA weights, and using square-root rather than even residual spreading. The variant w/o $\delta_i^a$' optimizes only the FFN shift; w/ WO,MHSAlW_{O,\mathrm{MHSA}}^l' updates both MHSA and FFN output weights; and `Even spread' distributes the residual uniformly across critical layers.

    Could not parse LaTeX table

    Jointly optimizing MHSA and FFN hidden states improves overall reliability over optimizing the FFN state alone, especially as the number of edits grows. Updating MHSA weights provides a small generalization gain but consistently reduces specificity, supporting the claim that MHSA contains reusable extraction patterns that should not be overwritten. Even spreading preserves more specificity and fluency but substantially reduces efficacy, generalization, and consistency; square-root spreading retains more target-update information at the cost of larger side effects.

  9. Knowl 9 — Limits of benchmark success and edited-model reasoning

    limitation

    The paper states that high efficacy and generalization scores on existing evaluations do not necessarily demonstrate true internalization of edited knowledge. PMET- and MEMIT-edited models may fail to reason compositionally with a new fact: after changing a statement such as The Prime Minister of the UK is Theresa' to Rishi Sunak', an edited model might still produce an incoherent related statement such as `The Prime Minister of England is British, not Indian'.

    The subject-centric editing definition is also not strictly followed by the available benchmarks; the zsRE and COUNTERFACT experiments are adapted to existing evaluation procedures rather than constructed specifically to test subject-level reasoning. The authors therefore identify more sophisticated editing methods and reasoning-oriented evaluations as necessary future work.

Coverage note — No substantial contributed material was omitted; background, related work, acknowledgements, ethics, and reference material were excluded as non-contributory.

References

  1. 1.Black, S.; Biderman, S.; Hallahan, E.; Anthony, Q.; Gao, L.; Golding, L.; He, H.; Leahy, C.; McDonell, K.; Phang, J.; Pieler, M.; Prashanth, U. S.; Purohit, S.; Reynolds, L.; Tow, J.; Wang, B.; and Weinbach, S. 2022. GPT-NeoX-20B: An Open-Source Autoregressive Language Model. In Proceedings of BigScience Episode #5 – Workshop on Challenges & Perspectives in Creating Large Language Models, 95–136. virtual+Dublin: Association for Computational Linguistics.
  2. 2.Cao, B.; Lin, H.; Han, X.; Sun, L.; Yan, L.; Liao, M.; Xue, T.; and Xu, J. 2021. Knowledgeable or Educated Guess? Revisiting Language Models as Knowledge Bases. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), 1860–1874. Online: Association for Computational Linguistics.
  3. 3.Cohen, R.; Biran, E.; Yoran, O.; Globerson, A.; and Geva, M. 2023. Evaluating the Ripple Effects of Knowledge Editing in Language Models. arXiv preprint arXiv:2307.12976.
  4. 4.De Cao, N.; Aziz, W.; and Titov, I. 2021. Editing Factual Knowledge in Language Models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 6491–6506. Online and Punta Cana, Dominican Republic: Association for Computational Linguistics.
  5. 5.Geva, M.; Bastings, J.; Filippova, K.; and Globerson, A. 2023. Dissecting Recall of Factual Associations in Autoregressive Language Models. arXiv:2304.14767.
  6. 6.Geva, M.; Caciularu, A.; Wang, K.; and Goldberg, Y. 2022. Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 30–45. Abu Dhabi, United Arab Emirates: Association for Computational Linguistics.
  7. 7.Geva, M.; Schuster, R.; Berant, J.; and Levy, O. 2021. Transformer Feed-Forward Layers Are Key-Value Memories. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 5484–5495. Online and Punta Cana, Dominican Republic: Association for Computational Linguistics.
  8. 8.Hao, Y.; Dong, L.; Wei, F.; and Xu, K. 2021. Self-Attention Attribution: Interpreting Information Interactions Inside Transformer. Proceedings of the AAAI Conference on Artificial Intelligence, 35(14): 12963–12971.
  9. 9.Hassid, M.; Peng, H.; Rotem, D.; Kasai, J.; Montero, I.; Smith, N. A.; and Schwartz, R. 2022. How much does attention actually attend? Questioning the Importance of Attention in Pretrained Transformers. arXiv:2211.03495.
  10. 10.Heinzerling, B.; and Inui, K. 2021. Language Models as Knowledge Bases: On Entity Representations, Storage Capacity, and Paraphrased Queries. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, 1772–1791. Online: Association for Computational Linguistics.
  11. 11.Hernandez, E.; Li, B. Z.; and Andreas, J. 2023. Inspecting and Editing Knowledge Representations in Language Models. arXiv:2304.00740.
  12. 12.Ji, Z.; Lee, N.; Frieske, R.; Yu, T.; Su, D.; Xu, Y.; Ishii, E.; Bang, Y. J.; Madotto, A.; and Fung, P. 2023. Survey of Hallucination in Natural Language Generation. ACM Computing Surveys, 55(12): 1–38.
  13. 13.Kobayashi, G.; Kuribayashi, T.; Yokoi, S.; and Inui, K. 2023. Feed-Forward Blocks Control Contextualization in Masked Language Models. arXiv:2302.00456.
  14. 14.Kovaleva, O.; Romanov, A.; Rogers, A.; and Rumshisky, A. 2019. Revealing the Dark Secrets of BERT. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 4365–4374. Hong Kong, China: Association for Computational Linguistics.
  15. 15.Levy, O.; Seo, M.; Choi, E.; and Zettlemoyer, L. 2017. Zero-Shot Relation Extraction via Reading Comprehension. In Proceedings of the 21st Conference on Computational Natural Language Learning (CoNLL 2017), 333–342. Vancouver, Canada: Association for Computational Linguistics.
  16. 16.Li, X.; Li, S.; Song, S.; Yang, J.; Ma, J.; and Yu, J. 2023. PMET: Precise Model Editing in a Transformer. arXiv:2308.08742.
  17. 17.Liang, K.; Liu, Y.; Zhou, S.; Tu, W.; Wen, Y.; Yang, X.; Dong, X.; and Liu, X. 2023a. Knowledge Graph Contrastive Learning Based on Relation-Symmetrical Structure. IEEE Transactions on Knowledge and Data Engineering, 1–12.
  18. 18.Liang, K.; Meng, L.; Liu, M.; Liu, Y.; Tu, W.; Wang, S.; Zhou, S.; and Liu, X. 2023b. Learn from relational correlations and periodic events for temporal knowledge graph reasoning. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval, 1559–1568.
  19. 19.Liang, R.; Li, T.; Li, L.; Wang, J.; and Zhang, Q. 2020. Knowledge Consistency between Neural Networks and Beyond. arXiv:1908.01581.
  20. 20.Meng, K.; Bau, D.; Andonian, A.; and Belinkov, Y. 2022a. Locating and Editing Factual Associations in GPT. In Koyejo, S.; Mohamed, S.; Agarwal, A.; Belgrave, D.; Cho, K.; and Oh, A., eds., Advances in Neural Information Processing Systems, volume 35, 17359–17372. Curran Associates, Inc.
  21. 21.Meng, K.; Sharma, A. S.; Andonian, A.; Belinkov, Y.; and Bau, D. 2022b. Mass-Editing Memory in a Transformer. arXiv:2210.07229.
  22. 22.Mitchell, E.; Lin, C.; Bosselut, A.; Finn, C.; and Manning, C. D. 2022a. Fast Model Editing at Scale. In International Conference on Learning Representations.
  23. 23.Mitchell, E.; Lin, C.; Bosselut, A.; Manning, C. D.; and Finn, C. 2022b. Memory-Based Model Editing at Scale. arXiv:2206.06520.
  24. 24.Murphy, A. H. 1996. The Finley affair: A signal event in the history of forecast verification. Weather and forecasting, 11(1): 3–20.
  25. 25.Petroni, F.; Rocktaschel, T.; Riedel, S.; Lewis, P.; Bakhtin, A.; Wu, Y.; and Miller, A. 2019. Language Models as Knowledge Bases? In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 2463–2473. Hong Kong, China: Association for Computational Linguistics.
  26. 26.Sinitsin, A.; Plokhotnyuk, V.; Pyrkin, D.; Popov, S.; and Babenko, A. 2020. Editable Neural Networks. arXiv:2004.00345.
  27. 27.Wang, B.; and Komatsuzaki, A. 2021. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https://github.com/kingoflolz/mesh-transformer-jax. Accessed: 2023-12-21.
  28. 28.Wang, K.; Variengien, A.; Conmy, A.; Shlegeris, B.; and Steinhardt, J. 2022. Interpretability in the wild: a circuit for indirect object identification in gpt-2 small. arXiv:2211.00593.
  29. 29.Yao, Y.; Wang, P.; Tian, B.; Cheng, S.; Li, Z.; Deng, S.; Chen, H.; and Zhang, N. 2023. Editing Large Language Models: Problems, Methods, and Opportunities. arXiv:2305.13172.
  30. 30.Zhao, W. X.; Zhou, K.; Li, J.; Tang, T.; Wang, X.; Hou, Y.; Min, Y.; Zhang, B.; Zhang, J.; Dong, Z.; Du, Y.; Yang, C.; Chen, Y.; Chen, Z.; Jiang, J.; Ren, R.; Li, Y.; Tang, X.; Liu, Z.; Liu, P.; Nie, J.-Y.; and Wen, J.-R. 2023. A Survey of Large Language Models. arXiv:2303.18223.
  31. 31.Zheng, C.; Li, L.; Dong, Q.; Fan, Y.; Wu, Z.; Xu, J.; and Chang, B. 2023. Can We Edit Factual Knowledge by In-Context Learning? arXiv:2305.12740.
  32. 32.Zhu, C.; Rawat, A. S.; Zaheer, M.; Bhojanapalli, S.; Li, D.; Yu, F.; and Kumar, S. 2020. Modifying Memories in Transformer Models. arXiv:2012.00363.

Citation

MLA
Li, X., et al. “PMET: Precise Model Editing in a Transformer”. arXiv, 2023, http://arxiv.org/abs/2308.08742v6.
APA
Li, X., Li, S., Song, S., Yang, J., Ma, J., & Yu, J. (2023). PMET: Precise Model Editing in a Transformer. arXiv. http://arxiv.org/abs/2308.08742v6
Chicago
Li, X., S. Li, S. Song, J. Yang, J. Ma, and J. Yu. 2023. “PMET: Precise Model Editing in a Transformer”. arXiv. http://arxiv.org/abs/2308.08742v6.
Harvard
Li, X. et al. (2023) “PMET: Precise Model Editing in a Transformer”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2308.08742v6.
Vancouver
1. Li X, Li S, Song S, Yang J, Ma J, Yu J (2023) PMET: Precise Model Editing in a Transformer. arXiv

BibTeX

@article{li2023pmet,
  title = {PMET: Precise Model Editing in a Transformer},
  author = {Li, Xiaopeng and Li, Shasha and Song, Shezheng and Yang, Jing and Ma, Jun and Yu, Jie},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2308.08742v6},
  eprint = {2308.08742}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF