Built independently by an author, for readers. Read the story and support ChapterPal

keyword

task learning

Task learning is the process by which a machine learning model acquires a new skill, rule, or input-output mapping at inference time directly from demonstration examples, without modifying its underlying parameters. In the context of in-context learning within large language models, task learning occurs when the model dynamically infers novel patterns or mappings presented in a prompt rather than merely retrieving previously memorized concepts. This mechanism contrasts with task retrieval, where the model simply identifies and activates a relevant skill that was already acquired during pretraining. As additional demonstration samples are provided, task learning allows the model to progressively refine its understanding of the new task, improving its ability to generalize and make accurate predictions on unseen queries.

1 item

Dual Operating Modes of In-Context Learning

Dual Operating Modes of In-Context Learning

Ziqian Lin, Kangwook Lee

OrganizationsUniversity of Wisconsin Madison

Why you should read this

Establishes a generalized probabilistic framework to mathematically explain how large language models transition between retrieving pretrained skills and learning new ones, providing theoretical foundations for puzzling empirical behaviors like the initial risk increase and the bounded efficacy of biased demonstrations.

In-context learning (ICL) exhibits dual operating modes: task learning, i.e. acquiring a new skill from in-context samples, and task retrieval, i.e., locating and activating a relevant pretrained skill. Recent theoretical work proposes various mathematical models to analyze ICL, but they cannot fully explain the duality. In this work, we analyze a generalized probabilistic model for pretraining data, obtaining a quantitative understanding of the two operating modes of ICL. Leveraging our analysis, we provide the first explanation of an unexplained phenomenon observed with real-world large language models (LLMs). Under some settings, the ICL risk initially increases and then decreases with more in-context examples. Our analysis offers a plausible explanation for this “early ascent” phenomenon: a limited number of in-context samples may lead to the retrieval of an incorrect skill, thereby increasing the risk, which will eventually diminish as task learning takes effect with more in-context samples. We also analyze ICL with biased labels, e.g., zero-shot ICL, where in-context examples are assigned random labels, and predict the bounded efficacy of such approaches. We corroborate our analysis and predictions with extensive experiments with Transformers and LLMs. The code is available at: https://github.com/UW-Madison-Lee-Lab/Dual_Operating_Modes_of_ICL.

Added

2026-10-03