From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs after Instruction Tuning

Xuansheng WuWenlin YaoJianshu ChenXiaoman PanXiaoyang WangNinghao LiuDong Yu

article2024NAACL59 citations

Explains the internal mechanisms of instruction tuning by analyzing how it redirects attention heads toward action verbs and rotates feed-forward representations toward task-oriented concepts to sustain prompt conditioning during generation.

Listen

Large language models have rapidly become foundational to modern artificial intelligence applications, where their ability to follow user instructions is critical. While fine-tuning pre-trained models on instruction-response pairs is known to establish this capability, the internal behavioral mechanisms driving this transformation have remained largely unexplained.

The article investigates how instruction tuning modifies the internal representations and operational behaviors of pre-trained language models. Specifically, it evaluates changes in prompt-to-response influence, self-attention mechanisms, and knowledge organization in feed-forward networks to understand how models transition from basic text continuation to helpful instruction following.

The researchers developed an interpretability toolbox to conduct a comparative analysis between foundation models (such as LLaMA and Mistral) and their instruction-tuned counterparts (such as Vicuna and Mistral-Instruct). They evaluated these models using human-written and benchmark datasets, including Self-Instruct, LIMA, and MT-Bench. The approach combined a gradient-based attribution method to track prompt token influence across responses, a co-occurrence analysis to interpret self-attention patterns, and principal component analysis paired with automated concept extraction to map internal knowledge shifts across linguistic and task dimensions.

The analysis produced three primary findings. First, instruction tuning enables models to reliably identify instruction words and maintain continuous conditioning on them throughout response generation. Quantitatively, the importance density score on instruction words was significantly higher in successfully followed instructions (e.g., 1.62 in followed vs. 1.28 in unfollowed tasks on the LIMA dataset). Second, self-attention heads in tuned models encode significantly more relationships tied to instruction verbs (e.g., "write", "create") than general verbs, especially in the initial eight layers where roughly 66% to 69% of modified heads favored instruction verbs. Third, feed-forward networks realign their encoded knowledge toward user-oriented tasks—such as coding, math, and writing—by slightly rotating their internal representation space without altering their foundational linguistic structure across phonology, morphology, syntax, and semantics.

These findings indicate that instruction tuning does not fundamentally reconstruct linguistic knowledge but instead acts as an internal routing mechanism that directs pre-existing capabilities toward user intentions. This distinction carries practical implications for model optimization: it validates targeted fine-tuning strategies that prioritize attention layers for rapid instruction alignment, while showing that full fine-tuning of feed-forward layers is necessary to adapt domain-specific functional concepts. Furthermore, the persistence of a "lost-in-the-middle" effect across models underscores systemic prompt-processing risks that system architects must mitigate.

Organizations developing or deploying aligned models should prioritize broad prompt diversity during training to maximize the coverage of instruction-following triggers. Engineering teams should also integrate internal attribution metrics, rather than relying solely on superficial output evaluation, to prevent models from generating seemingly compliant but unguided outputs. Future development should extend these interpretability tools to evaluate other alignment strategies, including reinforcement learning from human feedback, and address position-aware attribution to better detect repetitive response failures.

The conclusions are drawn with high confidence regarding the studied model families and open-weight architectures. However, decision-makers should note that the evaluation framework relies on white-box access to model weights and gradients, meaning these specific diagnostic methods cannot be directly applied to closed, commercial black-box models without further methodological adaptation.

  • Paper: Finetuned Language Models Are Zero-Shot Learners, Jason Wei et al. (2022). Establishes the foundational concept and methodology of instruction tuning to unlock zero-shot generalization in large language models, providing the baseline paradigm that the source mechanistically analyzes.
  • Paper: What Does BERT Look at? An Analysis of BERT’s Attention, Kevin Clark et al. (2019). Introduces foundational methods for dissecting attention heads and interpreting transformer internals, directly informing the source's approach to probing self-attention shifts.
  • Paper: Representation Engineering: A Top-Down Approach to AI Transparency, Andy Zou et al. (2023). Develops top-down techniques for interpreting and manipulating internal representations in language models, providing essential background for understanding how representations rotate during post-training.
  • Paper: Self-Instruct: Aligning Language Models with Self-Generated Instructions, Yizhong Wang et al. (2023). Demonstrates how synthetic instruction datasets align base language models into instruction followers, representing the core fine-tuning transition examined in the source.
  • Paper: LIMA: Less Is More for Alignment, Chunting Zhou et al. (2023). Proposes the hypothesis that alignment primarily surfaces existing pre-trained knowledge via superficial format learning rather than learning new capabilities from scratch.
  • Paper: BERT Rediscovers the Classical NLP Pipeline, Ian Tenney et al. (2019). Presents layer-wise probing techniques for transformer architectures to localize linguistic processing across feed-forward and self-attention components.
Cover for From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs after Instruction Tuning

Abstract

Large Language Models (LLMs) have achieved remarkable success, where instruction tuning is the critical step in aligning LLMs with user intentions. In this work, we investigate how the instruction tuning adjusts pre-trained models with a focus on intrinsic changes. Specifically, we first develop several local and global explanation methods, including a gradient-based method for input-output attribution, and techniques for interpreting patterns and concepts in self-attention and feed-forward layers. The impact of instruction tuning is then studied by comparing the explanations derived from the pre-trained and instruction-tuned models. This approach provides an internal perspective of the model shifts on a human-comprehensible level. Our findings reveal three significant impacts of instruction tuning: 1) It empowers LLMs to recognize the instruction parts of user prompts, and promotes the response generation constantly conditioned on the instructions. 2) It encourages the self-attention heads to capture more word-word relationships about instruction verbs. 3) It encourages the feed-forward networks to rotate their pre-trained knowledge toward user-oriented tasks. These insights contribute to a more comprehensive understanding of instruction tuning and lay the groundwork for future work that aims at explaining and optimizing LLMs for various applications. Our code and data are publicly available at https://github.com/JacksonWuxs/Interpret_Instruction_Tuning_LLMs.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Preliminary
  • 3.1 Notations
  • 3.2 General Experimental Settings
  • 4 Impact of User Prompts for Human Alignment
  • 4.1 Quantifying Prompt Influence on Generation Process
  • 4.2 Assessing Instruction Following Capability with Importance Density
  • 5 Shift within Instruction-tuned Models
  • 5.1 Analyzing Self-Attention Heads
  • 5.2 Analyzing Feed-forward Networks
  • 6 Discussion
  • 7 Conclusion
  • 8 Limitations
  • 9 Ethical Impact
  • References
  • A Proof of Linear Approximation to Importance Scores
  • B.2 Case Study on Outliers
  • B Analyzing Importance Density
  • B.1 Experiment Settings
  • B.3 Exploring Prompt Position with Importance Density
  • C Visualizing Salient Maps
  • C.1 Experiment Settings
  • C.2 Experiment Results
  • D Scaling up with Automated Tools
  • D.1 Experiment Settings
  • D.2 Experiment Results
  • E Interpreting Feed-Forward Networks
  • E.1 Details of the PCA Results
  • E.2 Concept Distribution Analysis with Mistral Family
  • E.3 Qualitative Analysis to Interpretability of Principal Components
  • F Interpreting Self-Attention Heads

Knowls

  1. Knowl 1 — Instruction tuning makes generation consistently depend on instruction words

    empirical result

    The paper compares prompt-to-response attribution patterns in pretrained LLaMA and instruction-tuned Vicuna models. In qualitative salient maps, instruction tokens such as task verbs and directives influence many response tokens across different output positions, whereas ordinary context tokens usually affect only localized spans. Context tokens often produce diagonal patterns corresponding to copying or repeating the input. In a tone-analysis prompt, Vicuna assigned stronger and more widespread influence to the instruction portion and successfully analyzed the email, while LLaMA relied more on the email context and failed to perform the requested analysis. These observations support the paper’s central claim that instruction tuning teaches models to identify instruction words and condition generation on them throughout the response.

  2. Knowl 2 — Normalized sparse gradient attribution for prompt-to-response influence

    model/method

    For a prompt X=(x_1,x_N) with NN tokens and a response Y=(y_1,y_M) with MM tokens, let Z_m=[X,y_1,y_{m-1}] be the context used to generate response token ymy_m. Let f(ym∣Zm)f(y_m\mid Z_m) be the language model’s conditional probability for ymy_m, and let Ei[xn]∈RDE_i[x_n]\in\mathbb{R}^{D} be the input embedding of prompt token xnx_n. The direct importance of xnx_n for ymy_m is defined as the probability change caused by removing xnx_n from the context:

    In,m=f(ym∣Zm)−f(ym∣Zm,∖n).I_{n,m}=f(y_m\mid Z_m)-f(y_m\mid Z_{m,\setminus n}).

    Here, Zm,∖nZ_{m,\setminus n} is the same context with the embedding of xnx_n removed or zeroed. To avoid repeatedly evaluating the model, the paper uses the first-order approximation

    In,m≈∂f(ym∣Zm)∂Ei[xn]Ei[xn]⊤.I_{n,m}\approx \frac{\partial f(y_m\mid Z_m)}{\partial E_i[x_n]}E_i[x_n]^{\top}.

    Because raw importance depends on the model’s confidence in the particular output token, the score is rescaled separately for each ymy_m and sparsified. Given an integer scale L>0L>0, threshold b∈[0,L]b\in[0,L], and prompt length NN, the normalized pairwise attribution is

    I~n,m=⌈LIn,mmax⁡1≤n′≤NIn′,m⌉,Sn,m={I~n,m,I~n,m>b,0,otherwise.\widetilde I_{n,m}=\left\lceil \frac{L I_{n,m}}{\max_{1\le n'\le N} I_{n',m}}\right\rceil, \qquad S_{n,m}=\begin{cases} \widetilde I_{n,m},&\widetilde I_{n,m}>b,\\ 0,&\text{otherwise.} \end{cases}

    The resulting matrix S∈RN×MS\in\mathbb{R}^{N\times M} is used to visualize which prompt tokens guide each generated token. The paper uses L=10L=10 and b=0b=0 for salient-map visualization.

  3. Knowl 3 — Importance density aggregates token influence over an entire response

    equation

    To measure whether a prompt token influences generation repeatedly rather than only at one output position, the paper aggregates its normalized pairwise attributions with an ℓ1/ℓp\ell_1/\ell_p density statistic. For prompt token xnx_n, let Sn,mS_{n,m} be its normalized sparse attribution to response token ymy_m, let Sn=[Sn,1,0˘07fSn,M]S_n=[S_{n,1},\u007fS_{n,M}], and let p>0p>0 be a tunable exponent. The importance density is

    an=∥Sn∥1∥Sn∥p.a_n=\frac{\lVert S_n\rVert_1}{\lVert S_n\rVert_p}.

    The statistic is larger when the token has concentrated, high attribution to particular response positions. In particular, among two prompt tokens with the same total attribution, the token with the larger maximum attribution receives the larger density score. For the quantitative experiments, the paper uses L=10L=10, sparsity threshold b=7b=7, and p=4p=4, normalizes scores across instances, and excludes responses shorter than five tokens because their density estimates are unstable.

  4. Knowl 4 — Instruction-word importance density tracks instruction following

    data/table

    The paper manually marks instruction sentences in prompts and labels generated responses as “followed” when they provide useful information relevant to the user’s intention, regardless of factual correctness. Responses that merely repeat or randomly produce content are labeled “unfollowed.” For each marked instruction token, it computes the importance density an=∥Sn∥1/∥Sn∥pa_n=\lVert S_n\rVert_1/\lVert S_n\rVert_p from normalized prompt-to-response attributions and averages the scores over instruction tokens.

    The results show that followed responses consistently assign greater density to instruction words than unfollowed responses. They also show that instruction-tuned Vicuna assigns greater instruction-word density than pretrained LLaMA on all three datasets.

    Could not parse LaTeX table
    Could not parse LaTeX table

    The measurements support the paper’s interpretation that successful instruction following involves repeatedly using instruction words to guide response generation, and that instruction tuning increases this behavior relative to the pretrained model.

  5. Knowl 5 — Co-occurrence-constrained word-pair explanations for self-attention heads

    model/method

    The paper interprets a self-attention head through word pairs rather than isolated projected weight vectors. For a head hh with query and key matrices Wqh,Wkh∈RD×D′W_q^h,W_k^h\in\mathbb{R}^{D\times D'}, input embedding Ei[w]∈RDE_i[w]\in\mathbb{R}^{D}, and neuron index d∈{1,0˘07fD′}d\in\{1,\u007fD'\}, define the query and key activations of word ww as

    qh,d(w)=Ei[w]⋅Wqh[:,d],kh,d(w)=Ei[w]⋅Wkh[:,d].q_{h,d}(w)=E_i[w]\cdot W_q^h[:,d],\qquad k_{h,d}(w)=E_i[w]\cdot W_k^h[:,d].

    The word-pair relation contributed by neuron dd is proportional to qh,d(wa)kh,d(wb)q_{h,d}(w_a)k_{h,d}(w_b). For each query neuron and key neuron, the method collects the top KK vocabulary words with the largest activation. It then retains a candidate pair (wq,wk)(w_q,w_k) only when the pretrained GloVe embeddings of the two words have cosine similarity above a threshold θ\theta. This local co-occurrence constraint removes many polysemantic or implausible pairs that arise when the entire vocabulary is searched without regard to words appearing together in the same input.

    The final explanation of a head is its frequent retained neuron-level word pairs. The experiments use K=100K=100. The threshold is word-dependent: for each word, cosine similarities to 1,000 frequent words are computed, and the threshold is the mean similarity plus 1.961.96 standard deviations; the larger threshold of the two words is used for a pair.

  6. Knowl 6 — Instruction tuning changes self-attention toward instruction verbs

    empirical result

    The paper compares the top-100 word-pair explanations of corresponding pretrained and instruction-tuned self-attention heads. If EptE_{pt} and EftE_{ft} are the two pair sets, their similarity is measured by

    M=∣Ept∩Eft∣∣Ept∪Eft∣.M=\frac{|E_{pt}\cap E_{ft}|}{|E_{pt}\cup E_{ft}|}.

    The difference 1−M1-M increases with layer depth, showing that instruction tuning substantially changes self-attention word relations, especially in deeper layers. The analysis also examines 45 instruction verbs, including “write,” “create,” and “classify,” against 3,000 frequent general verbs. It considers only heads whose number of associated pairs for a verb changes after tuning.

    Could not parse LaTeX table

    Instruction verbs become more prevalent in the lower LLaMA layers, with a statistically significant difference in layers 1–8, and the same tendency is present in the middle layers. General verbs remain close to a neutral 50% tendency. Mistral shows an increase for instruction verbs across all layer groups, with the strongest evidence in layers 1–8. The result indicates that self-attention specifically learns more relations associated with detailed user instructions rather than simply changing verb-related patterns in general.

  7. Knowl 7 — Principal-component decomposition gives feed-forward networks concept-level explanations

    model/method

    The paper treats each transformer feed-forward network as a key-value memory whose rows in the output matrix Wp∈RD′′×DW_p\in\mathbb{R}^{D''\times D} encode textual patterns. Because individual rows can be polysemantic, it centers the rows of WpW_p to form WpcW_p^c and analyzes orthogonal principal directions of their covariance matrix:

    C=(Wpc)⊤Wpc,Cvr=λrvr,C=(W_p^c)^{\top}W_p^c,\qquad C v_r=\lambda_r v_r,

    where vr∈RDv_r\in\mathbb{R}^{D} is a unit-length principal direction, λr≥0\lambda_r\ge0 is its eigenvalue, and the directions are ordered so that λ1≥λ2≥⋯≥λD\lambda_1\ge\lambda_2\ge\cdots\ge\lambda_D. For each of the leading directions, the method ranks words using the output embedding matrix Eo∈R∣V∣×DE_o\in\mathbb{R}^{|V|\times D}:

    score⁡r(w)=vr⊤Eo[w],\operatorname{score}_r(w)=v_r^{\top}E_o[w],

    where Eo[w]∈RDE_o[w]\in\mathbb{R}^{D} is the output embedding of vocabulary word ww. The top-ranked words are presented as a compact description of the concept represented by vrv_r.

    For the large-scale analysis, the paper builds a more interpretable vocabulary from ShareGPT rather than using highly fragmented model sub-tokens, examines the first 300 principal directions of every feed-forward layer, and uses ChatGPT-turbo-3.5-0613 to summarize each top-15 word list. The summarized concepts are classified into user-oriented tasks—daily, literary, or professional writing, coding, solving math problems, and translation—and into phonology, morphology, syntax, or semantic linguistic levels. Concepts may belong to multiple user-oriented tasks.

  8. Knowl 8 — Instruction tuning rotates feed-forward knowledge toward user tasks

    data/table

    The principal-component concepts show that instruction tuning changes the user-oriented applications represented by feed-forward networks while preserving their broad linguistic organization. Values below are percentages of concepts; task assignments can overlap, so scenario percentages need not sum to 100.

    Could not parse LaTeX table

    For the LLaMA family, Vicuna contains significantly more writing and coding concepts and slightly fewer translation concepts; the math difference is not statistically significant. For the Mistral family, instruction tuning significantly increases coding- and math-related concepts, while the other scenario changes are not statistically significant. Across both model families, none of the four linguistic-level distributions changes significantly. The result supports the interpretation that feed-forward networks adapt existing knowledge to user-oriented tasks by rotating or reallocating representational directions rather than by fundamentally changing the model’s linguistic hierarchy.

  9. Knowl 9 — Feed-forward concepts vary systematically across layers and are partly machine-interpretable

    empirical result

    The principal-component analysis yields several additional properties of feed-forward representations. Across Vicuna’s 32 layers, the first 300 principal directions account for 22.49% of accumulated variance, while approximately half of the directions account for about 80%, indicating that the encoded knowledge is distributed across many directions rather than concentrated in a few components.

    Automated concept descriptions are most interpretable in middle-to-late layers: for the first 30 ranked concepts in layers 24–28, the average proportion of word lists assigned a concise description reaches 91.67%. Interpretability becomes harder near the output layers 28–32 and generally declines for lower-ranked concepts. In the middle layers of Vicuna, approximately 60% of the first 300 components can be described by the machine annotator.

    The linguistic-level distribution also changes systematically with depth. Semantic concepts are most prevalent in lower and middle layers and then decrease toward the top, whereas morphology follows the opposite U-shaped trend and becomes more prominent in the final layers. The paper conjectures that the late-layer morphology pattern may help the generative model represent prefix and suffix structures efficiently, but this explanation is presented as a conjecture rather than a demonstrated mechanism. Frequency changes in concept descriptions additionally suggest that lower layers encode more basic behavior and word properties, middle layers encode more abstract programming and methodological knowledge, and higher layers encode patterns useful for efficient text completion.

  10. Knowl 10 — Comparative evaluation uses paired pretrained and instruction-tuned open models

    experimental setup

    The study evaluates two pretrained/instruction-tuned pairs with accessible weights and gradients: LLaMA versus Vicuna, and Mistral versus Mistral-Instruct. Responses are generated by greedy decoding with a maximum of 300 tokens per prompt. The prompts come from the test material of Self-Instruct, LIMA, and MT-Bench: Self-Instruct contributes 252 human-written prompt-response pairs, LIMA contributes 300 testing pairs in addition to its training set, and MT-Bench contributes 80 human-written pairs spanning eight categories.

    The analyses compare the same model families before and after instruction tuning. Prompt influence is studied with gradients and generated responses; self-attention is studied through extracted word-pair patterns; and feed-forward knowledge is studied through principal directions and concept descriptions. The quantitative instruction-following analysis manually marks instruction sentences and labels responses as followed or unfollowed, using helpfulness and relevance to the user’s intention rather than factual correctness as the criterion.

  11. Knowl 11 — The explanation toolbox is limited to white-box models and has known attribution failure cases

    limitation

    All proposed explanations require access to model weights and gradients, so they are designed for white-box analysis and may not transfer directly to black-box systems such as ChatGPT or Claude. The study also analyzes supervised instruction tuning but does not investigate reinforcement learning from human feedback, leaving the distinct contributions of instruction tuning and RLHF unresolved.

    The importance-density statistic has a positional limitation: it can assign high instruction density to a response that merely repeats the prompt. In one case where the entire input was treated as instruction, repeated output created diagonal attribution patterns that looked like successful instruction following even though the model was echoing the prompt. The paper therefore identifies position-aware density functions and early repetition detection as necessary directions for improving the method.

Coverage note — The paper’s extensive qualitative word-pair appendices, individual salient-map examples, automated-annotation prompt templates, and long lists of principal-component word examples were omitted because they instantiate the reported methods without adding separate load-bearing findings.

References

  1. 1.Anthropic. 2023. Model Card and Evaluations for Claude Models.
  2. 2.Sanjeev Arora, Yuanzhi Li, Yingyu Liang, Tengyu Ma, and Andrej Risteski. 2018. Linear algebraic structure of word senses, with applications to polysemy. Transactions of the Association for Computational Linguistics, 6:483–495.
  3. 3.Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. 2022. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862.
  4. 4.Oren Barkan, Edan Hauon, Avi Caciularu, Ori Katz, Itzik Malkiel, Omri Armstrong, and Noam Koenigstein. 2021. Grad-sam: Explaining transformers via gradient self-attention maps. In Proceedings of the 30th ACM International Conference on Information & Knowledge Management, pages 2882–2887.
  5. 5.Yonatan Belinkov, Lluís Màrquez, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James Glass. 2018. Evaluating layers of representation in neural machine translation on part-of-speech and semantic tagging tasks. arXiv preprint arXiv:1801.07772.
  6. 6.Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nicholas L Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Alex Tamkin, Karina Nguyen, Brayden McLean, Josiah E Burke, Tristan Hume, Shan Carter, Tom Henighan, and Chris Olah. 2023a. Decomposing Language Models With Dictionary Learning.
  7. 7.Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Adly Templeton, Amanda Askell, et al. 2023b. Towards monosemanticity: Decomposing language models with dictionary learning. transformer circuits thread, 2023.
  8. 8.Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey. 2023. Sparse autoencoders find highly interpretable features in language models. arXiv preprint arXiv:2309.08600.
  9. 9.Guy Dar, Mor Geva, Ankit Gupta, and Jonathan Berant. 2022. Analyzing transformers in embedding space. arXiv preprint arXiv:2209.02535.
  10. 10.Jinhao Duan, Hao Cheng, Shiqi Wang, Chenan Wang, Alex Zavalny, Renjing Xu, Bhavya Kailkhura, and Kaidi Xu. 2023. Shifting attention to relevance: Towards the uncertainty estimation of large language models. arXiv preprint arXiv:2307.01379.
  11. 11.Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, et al. 2022. Toy models of superposition. arXiv preprint arXiv:2209.10652.
  12. 12.Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, et al. 2021. A mathematical framework for transformer circuits. Transformer Circuits Thread, 1.
  13. 13.Shi Feng, Eric Wallace, Alvin Grissom II, Mohit Iyyer, Pedro Rodriguez, and Jordan Boyd-Graber. 2018. Pathologies of neural models make interpretations difficult. arXiv preprint arXiv:1804.07781.
  14. 14.Edward Fredkin. 1960. Trie memory. Communications of the ACM, 3(9):490–499.
  15. 15.Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. 2020. Transformer feed-forward layers are key-value memories. arXiv preprint arXiv:2012.14913.
  16. 16.Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. 2021. Transformer feed-forward layers are key-value memories. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 5484–5495.
  17. 17.Raffaele Giancarlo. 1995. A generalization of the suffix tree to square matrices, with applications. SIAM Journal on Computing, 24(3):520–562.
  18. 18.Jing Huang, Atticus Geiger, Karel D’Oosterlinck, Zhengxuan Wu, and Christopher Potts. 2023. Rigorously assessing natural language explanations of neurons. arXiv preprint arXiv:2309.10312.
  19. 19.Niall Hurley and Scott Rickard. 2009. Comparing measures of sparsity. IEEE Transactions on Information Theory, 55(10):4723–4741.
  20. 20.Ganesh Jawahar, Benoît Sagot, and Djamé Seddah. 2019. What does bert learn about the structure of language? In ACL 2019-57th Annual Meeting of the Association for Computational Linguistics.
  21. 21.Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023. Mistral 7b. arXiv preprint arXiv:2310.06825.
  22. 22.Juletx. 2023. Alpaca-lora. https://github.com/tloen/alpaca-lora.
  23. 23.Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, et al. 2018. Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav). In International conference on machine learning, pages 2668–2677. PMLR.
  24. 24.Robert Kirk, Ishita Mediratta, Christoforos Nalmpantis, Jelena Luketina, Eric Hambro, Edward Grefenstette, and Roberta Raileanu. 2023. Understanding the effects of rlhf on llm generalisation and diversity. arXiv preprint arXiv:2310.06452.
  25. 25.Enja Kokalj, Blaž Škrlj, Nada Lavrac, Senja Pollak, and Marko Robnik-Šikonja. 2021. Bert meets shapley: Extending shap explanations to transformer-based classifiers. In Proceedings of the EACL Hackashop on News Media Content Analysis and Automated Report Generation, pages 16–21.
  26. 26.Po-Nien Kung and Nanyun Peng. 2023. Do models really learn to follow instructions? an empirical study of instruction tuning. arXiv preprint arXiv:2305.11383.
  27. 27.Jiwei Li, Xinlei Chen, Eduard Hovy, and Dan Jurafsky. 2015. Visualizing and understanding neural models in nlp. arXiv preprint arXiv:1506.01066.
  28. 28.Zongxia Li, Paiheng Xu, Fuxiao Liu, and Hyemi Song. 2023. Towards understanding in-context learning with contrastive demonstrations and saliency maps. arXiv preprint arXiv:2307.05052.
  29. 29.Shihao Liang, Kunlun Zhu, Runchu Tian, Yujia Qin, Huadong Wang, Xin Cong, Zhiyuan Liu, Xiaojiang Liu, and Maosong Sun. 2023. Exploring format consistency for instruction tuning. arXiv preprint arXiv:2307.15504.
  30. 30.Nelson F Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2023. Lost in the middle: How language models use long contexts. arXiv preprint arXiv:2307.03172.
  31. 31.Beren Millidge and Sid Black. 2022. The singular value decompositions of transformer weight matrices are highly interpretable. https://www.alignmentforum.org/.
  32. 32.Jesse Mu and Jacob Andreas. 2020. Compositional explanations of neurons. Advances in Neural Information Processing Systems, 33:17153–17163.
  33. 33.Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, et al. 2022. In-context learning and induction heads. arXiv preprint arXiv:2209.11895.
  34. 34.R OpenAI. 2023. Gpt-4 technical report. arXiv, pages 2303–08774.
  35. 35.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744.
  36. 36.Sibabrata Paladhi and Sivaji Bandyopadhyay. 2008. Generation of referring expression using prefix tree structure. In Proceedings of the Third International Joint Conference on Natural Language Processing: Volume-II.
  37. 37.Baolin Peng, Chunyuan Li, Pengcheng He, Michel Galley, and Jianfeng Gao. 2023. Instruction tuning with gpt-4. arXiv preprint arXiv:2304.03277.
  38. 38.Jeffrey Pennington, Richard Socher, and Christopher D Manning. 2014. Glove: Global vectors for word representation. In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), pages 1532–1543.
  39. 39.Fabio Petroni, Tim Rocktäschel, Patrick Lewis, Anant Bakhtin, Yuxiang Wu, Alexander H Miller, and Sebastian Riedel. 2019. Language models as knowledge bases? arXiv preprint arXiv:1909.01066.
  40. 40.Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016. " why should i trust you?" explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining, pages 1135–1144.
  41. 41.RyokoAI. 2023. Sharegpt52k. Huggingface Datasets.
  42. 42.Adam Scherlis, Kshitij Sachan, Adam S Jermyn, Joe Benton, and Buck Shlegeris. 2022. Polysemanticity and capacity in neural networks. arXiv preprint arXiv:2210.01892.
  43. 43.Ramprasaath R Selvaraju, Abhishek Das, Ramakrishna Vedantam, Michael Cogswell, Devi Parikh, and Dhruv Batra. 2016. Grad-cam: Why did you say that? arXiv preprint arXiv:1611.07450.
  44. 44.Y Shan, X Chen, Y Shi, and J Liu. 2012. Fast language model look-ahead algorithm using extended n-gram model. Acta Automatica Sinica, 38(10):1618–1626.
  45. 45.Robyn Speer. 2022. rspeer/wordfreq: v3.0.
  46. 46.Bills Steven, Cammarata Nick, Mossing Dan, Tillman Henk, Gao Leo, Goh Gabriel, Sutskever Ilya, Leike Jan, Wu Jeff, and Saunders William. 2022. Language models can explain neurons in language models. https://openaipublic.blob.core.windows.net/neuron-explainer/paper/index.html.
  47. 47.Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano. 2020. Learning to summarize with human feedback. Advances in Neural Information Processing Systems, 33:3008–3021.
  48. 48.Sainbayar Sukhbaatar, Edouard Grave, Guillaume Lample, Herve Jegou, and Armand Joulin. 2019. Augmenting self-attention with persistent memory. arXiv preprint arXiv:1907.01470.
  49. 49.Xianghui Sun, Yunjie Ji, Baochang Ma, and Xiangang Li. 2023. A comparative study between full-parameter and lora-based fine-tuning on chinese instruction data for instruction following large language model. arXiv preprint arXiv:2304.08109.
  50. 50.Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017. Axiomatic attribution for deep networks. In International conference on machine learning, pages 3319–3328. PMLR.
  51. 51.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto. 2023. Alpaca: A strong, replicable instruction-following model. Stanford Center for Research on Foundation Models. https://crfm. stanford. edu/2023/03/13/alpaca. html, 3(6):7.
  52. 52.James J Thomas. 2005. Illuminating the path:[the research and development agenda for visual analytics]. IEEE Computer Society.
  53. 53.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  54. 54.Neeraj Varshney, Wenlin Yao, Hongming Zhang, Jianshu Chen, and Dong Yu. 2023. A stitch in time saves nine: Detecting and mitigating hallucinations of llms by validating low-confidence generation. arXiv preprint arXiv:2307.03987.
  55. 55.Jesse Vig. 2019. Bertviz: A tool for visualizing multihead self-attention in the bert model. In ICLR workshop: Debugging machine learning models, volume 23.
  56. 56.Elena Voita, Javier Ferrando, and Christoforos Nalmpantis. 2023. Neurons in large language models: Dead, n-gram, positional. arXiv preprint arXiv:2309.04827.
  57. 57.Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2022. Self-instruct: Aligning language model with self generated instructions. arXiv preprint arXiv:2212.10560.
  58. 58.Jerry Wei, Jason Wei, Yi Tay, Dustin Tran, Albert Webson, Yifeng Lu, Xinyun Chen, Hanxiao Liu, Da Huang, Denny Zhou, et al. 2023. Larger language models do in-context learning differently. arXiv preprint arXiv:2303.03846.
  59. 59.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al. 2019. Huggingface’s transformers: State-of-the-art natural language processing. arXiv preprint arXiv:1910.03771.
  60. 60.Tongshuang Wu, Marco Tulio Ribeiro, Jeffrey Heer, and Daniel S Weld. 2021. Polyjuice: Generating counterfactuals for explaining, evaluating, and improving models. arXiv preprint arXiv:2101.00288.
  61. 61.Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma. 2021. An explanation of in-context learning as implicit bayesian inference. arXiv preprint arXiv:2111.02080.
  62. 62.Miao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li, Jie Fu, Junxian He, and Bryan Hooi. 2023. Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms. arXiv preprint arXiv:2306.13063.
  63. 63.Matthew D Zeiler and Rob Fergus. 2014. Visualizing and understanding convolutional networks. In Computer Vision–ECCV 2014: Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part I 13, pages 818–833. Springer.
  64. 64.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2024. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in Neural Information Processing Systems, 36.
  65. 65.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric. P Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena.
  66. 66.Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, et al. 2023. Lima: Less is more for alignment. arXiv preprint arXiv:2305.11206.

Citation

MLA
Wu, X., et al. “From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs After Instruction Tuning”. arXiv, 2023, http://arxiv.org/abs/2310.00492v3.
APA
Wu, X., Yao, W., Chen, J., Pan, X., Wang, X., Liu, N., & Yu, D. (2023). From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs after Instruction Tuning. arXiv. http://arxiv.org/abs/2310.00492v3
Chicago
Wu, X., W. Yao, J. Chen, et al. 2023. “From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs After Instruction Tuning”. arXiv. http://arxiv.org/abs/2310.00492v3.
Harvard
Wu, X. et al. (2023) “From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs after Instruction Tuning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2310.00492v3.
Vancouver
1. Wu X, Yao W, Chen J, Pan X, Wang X, Liu N, Yu D (2023) From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs after Instruction Tuning. arXiv

BibTeX

@article{wu2023from,
  title = {From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs after Instruction Tuning},
  author = {Wu, Xuansheng and Yao, Wenlin and Chen, Jianshu and Pan, Xiaoman and Wang, Xiaoyang and Liu, Ninghao and Yu, Dong},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2310.00492v3},
  eprint = {2310.00492}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/