Built independently by an author, for readers. Read the story and support ChapterPal

keyword

knowledge injection

Knowledge injection is the process of integrating external, factual, or domain-specific information into an artificial intelligence model to expand its knowledge base and improve its reasoning capabilities. In natural language processing and machine learning, this is typically achieved through parametric techniques, such as continual pre-training, fine-tuning, and direct model editing, or through non-parametric approaches, such as retrieval-augmented prompting and knowledge graph integration. By supplementing an existing system with facts that were absent, outdated, or incomplete in the original training data, knowledge injection helps mitigate model hallucinations and enables language models to reliably utilize and reason over newly acquired information across diverse tasks.

7 items

Subagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks

Subagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks

Wasu Top Piriyakulkij, Rachel Lawrence, Alicia Curth, Sushrut Karmalkar, Niranjani Prasad

OrganizationsCornell UniversityMicrosoft

Why you should read this

Demonstrates that executing reusable skills through dedicated subagents with fresh context windows outperforms standard in-context skill loading on long-horizon tasks by preventing reasoning degradation.

How can language model agents effectively leverage libraries of reusable knowledge to solve long-horizon tasks? Recent work has increasingly focused on agent skills: reusable capabilities represented as skill packages, i.e., multi-file bundles containing instructions, scripts, and other resources that help agents perform specific tasks. Agent skills are typically executed by loading their skill instructions into an agent's context and relying on the agent to follow them. As task horizons grow, however, this approach becomes increasingly brittle, because reasoning quality degrades as more information accumulates in the context window. We investigate an alternative approach in which skill packages are instead invoked as subagents. Rather than loading skill instructions into the main context, subagent execution spawns fresh context windows dedicated to solving individual subtasks. We show that subagent execution outperforms agent-skill execution when skill packages expose clear input-output contracts and their instructions encode the procedural knowledge needed to fulfill those contracts. The tradeoff is additional communication overhead, as extra tokens are required to coordinate between the main agent and its subagents. Our results show that the benefit of reusable knowledge depends not only on its content, but also on how it is organized and invoked.

Added

2026-10-05

Learning or Self-aligning? Rethinking Instruction Fine-tuning

Learning or Self-aligning? Rethinking Instruction Fine-tuning

Mengjie Ren, Boxi Cao, Hongyu Lin, Cao Liu, Xianpei Han, Ke Zeng, Guanglu Wan, Xunliang Cai, Le Sun

OrganizationsChinese Information Processing LaboratoryInstitute of Software, Chinese Academy of SciencesMeituanUniversity of Chinese Academy of Sciences

Why you should read this

Demonstrates that instruction fine-tuning succeeds primarily by aligning outputs with a language model's existing internal parameter knowledge rather than teaching new facts, warning that introducing inconsistent world knowledge during tuning degrades model performance.

Instruction Fine-tuning (IFT) is a crucial phase in building large language models (LLMs). Previous works mainly focus on the IFT’s role in the transfer of behavioral norms and the learning of additional world knowledge. However, the understanding of the underlying mechanisms of IFT remains significantly limited. In this paper, we design a knowledge intervention framework to decouple the potential underlying factors of IFT, thereby enabling individual analysis of different factors. Surprisingly, our experiments reveal that attempting to learn additional world knowledge through IFT often struggles to yield positive impacts and can even lead to markedly negative effects. Further, we discover that maintaining internal knowledge consistency before and after IFT is a critical factor for achieving successful IFT. Our findings reveal the underlying mechanisms of IFT and provide robust support for some very recent and potential future works. We release our experimental dataset and codes to facilitate future work¹.

Added

2026-10-04

Stance Detection on Social Media with Background Knowledge

Stance Detection on Social Media with Background Knowledge

Ang Li, Bin Liang, Jingqian Zhao, Bowen Zhang, Min Yang, Ruifeng Xu

OrganizationsGuangdong Provincial Key Laboratory of Novel Security Intelligence TechnologiesHarbin Institute of TechnologyPeng Cheng LaboratoryShenzhen Institute of Advanced Technology, Chinese Academy of SciencesShenzhen Technology UniversityThe Chinese University of Hong Kong

Why you should read this

Proposes a knowledge-augmented stance detection framework that integrates Wikipedia-derived episodic context and language-model-interpreted social discourse to achieve state-of-the-art accuracy in both in-target and zero-shot settings.

Identifying users’ stances regarding specific targets/topics is a significant route to learning public opinion from social media platforms. Most existing studies of stance detection strive to learn stance information about specific targets from the context, in order to determine the user’s stance on the target. However, in real-world scenarios, we usually have a certain understanding of a target when we express our stance on it. In this paper, we investigate stance detection from a novel perspective, where the background knowledge of the targets is taken into account for better stance detection. To be specific, we categorize background knowledge into two categories: episodic knowledge and discourse knowledge, and propose a novel Knowledge-Augmented Stance Detection (KASD) framework. For episodic knowledge, we devise a heuristic retrieval algorithm based on the topic to retrieve the Wikipedia documents relevant to the sample. Further, we construct a prompt for ChatGPT to filter the Wikipedia documents to derive episodic knowledge. For discourse knowledge, we construct a prompt for ChatGPT to paraphrase the hashtags, references, etc., in the sample, thereby injecting discourse knowledge into the sample. Experimental results on four benchmark datasets demonstrate that our KASD achieves state-of-the-art performance in in-target and zero-shot stance detection.

Added

2026-10-03

Boosting Language Models Reasoning with Chain-of-Knowledge Prompting

Boosting Language Models Reasoning with Chain-of-Knowledge Prompting

Jianing Wang, Qiushi Sun, Xiang Li, Ming Gao

OrganizationsEast China Normal UniversityUniversity of Hong Kong

Why you should read this

Proposes Chain-of-Knowledge prompting to reduce hallucinations in large language models by generating structured knowledge triples alongside natural language explanations and using an F²-Verification mechanism to evaluate factuality and faithfulness before triggering iterative correction.

Recently, Chain-of-Thought (CoT) prompting has delivered success on complex reasoning tasks, which aims at designing a simple prompt like “Let’s think step by step” or multiple in-context exemplars with well-designed rationales to elicit Large Language Models (LLMs) to generate intermediate reasoning steps. However, the generated rationales often come with hallucinations, making unfactual and unfaithful reasoning chains. To mitigate this brittleness, we propose a novel Chain-of-Knowledge (CoK) prompting, where we aim at eliciting LLMs to generate explicit pieces of knowledge evidence in the form of structure triple. This is inspired by our human behaviors, i.e., we can draw a mind map or knowledge map as the reasoning evidence in the brain before answering a complex question. Benefiting from CoK, we additionally introduce a F²-Verification method to estimate the reliability of the reasoning chains in terms of factuality and faithfulness. For the unreliable response, the wrong evidence can be indicated to prompt the LLM to rethink. Extensive experiments demonstrate that our method can further improve the performance of commonsense, factual, symbolic, and arithmetic reasoning tasks¹.

Added

2026-10-01

Editing Large Language Models: Problems, Methods, and Opportunities

Editing Large Language Models: Problems, Methods, and Opportunities

Yunzhi Yao, Peng Wang, Bozhong Tian, Siyuan Cheng, Zhoubo Li, Shumin Deng, Huajun Chen, Ningyu Zhang

OrganizationsDonghai LaboratoryNational University of SingaporeNUS-NCS Joint LabZhejiang UniversityZhejiang University - Ant Group Joint Laboratory of Knowledge Graph

Why you should read this

Presents a unified taxonomy and empirical evaluation of large language model editing techniques across multiple architectures and settings, providing a standardized benchmark to assess reliability, generalization, locality, and computational efficiency.

Despite the ability to train capable LLMs, the methodology for maintaining their relevancy and rectifying errors remains elusive. To this end, the past few years have witnessed a surge in techniques for editing LLMs, the objective of which is to efficiently alter the behavior of LLMs within a specific domain without negatively impacting performance across other inputs. This paper embarks on a deep exploration of the problems, methods, and opportunities related to model editing for LLMs. In particular, we provide an exhaustive overview of the task definition and challenges associated with model editing, along with an in-depth empirical analysis of the most progressive methods currently at our disposal. We also build a new benchmark dataset to facilitate a more robust evaluation and pinpoint enduring issues intrinsic to existing techniques. Our objective is to provide valuable insights into the effectiveness and feasibility of each editing technique, thereby assisting the community in making informed decisions on the selection of the most appropriate method for a specific task or context¹.

Added

2026-09-29