Built independently by an author, for readers. Read the story and support ChapterPal

keyword

linear regression

Linear regression is a foundational statistical and machine learning method used to model the relationship between one or more independent predictor variables and a continuous dependent target variable by fitting a linear equation to observed data. In its simplest form, it estimates the expected outcome as a weighted sum of input features plus a constant bias term. The model parameters are commonly estimated using optimization techniques such as ordinary least squares, which minimizes the sum of squared differences between observed and predicted values, or iterative algorithms such as gradient descent. Linear regression is also frequently extended with regularization penalties, such as ridge or lasso regression, to prevent overfitting and improve generalization when dealing with correlated or high-dimensional data, making it a standard tool for predictive modeling, trend analysis, and quantifying the strength of relationships between variables.

8 items

What learning algorithm is in-context learning? Investigations with linear models

What learning algorithm is in-context learning? Investigations with linear models

Ekin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, Denny Zhou

OrganizationsGoogleMassachusetts Institute of TechnologyStanford University

Why you should read this

Demonstrates through constructive proofs and empirical analysis that transformers perform in-context learning by implicitly executing standard algorithms, such as gradient descent and ridge regression, directly within their internal activations.

Neural sequence models, especially transformers, exhibit a remarkable capacity for in-context learning. They can construct new predictors from sequences of labeled examples (x,f(x))(x, f(x)) presented in the input without further parameter updates. We investigate the hypothesis that transformer-based in-context learners implement standard learning algorithms implicitly, by encoding smaller models in their activations, and updating these implicit models as new examples appear in the context. Using linear regression as a prototypical problem, we offer three sources of evidence for this hypothesis. First, we prove by construction that transformers can implement learning algorithms for linear models based on gradient descent and closed-form ridge regression. Second, we show that trained in-context learners closely match the predictors computed by gradient descent, ridge regression, and exact least-squares regression, transitioning between different predictors as transformer depth and dataset noise vary, and converging to Bayesian estimators for large widths and depths. Third, we present preliminary evidence that in-context learners share algorithmic features with these predictors: learners' late layers non-linearly encode weight vectors and moment matrices. These results suggest that in-context learning is understandable in algorithmic terms, and that (at least in the linear case) learners may rediscover standard estimation algorithms. Code and reference implementations are released at this https URL.

Added

2026-10-05

Deep Regression Unlearning

Deep Regression Unlearning

Ayush Kumar Tarun, Vikram Singh Chundawat, Murari Mandal, Mohan S. Kankanhalli

OrganizationsKalinga Institute of Industrial TechnologyMavex LabsNational University of Singapore

Why you should read this

Presents the first machine unlearning frameworks designed specifically for deep regression and forecasting models, introducing a weight-optimization method called Blindspot alongside adapted evaluation metrics to effectively scrub targeted data while defending against privacy attacks.

With the introduction of data protection and privacy regulations, it has become crucial to remove the lineage of data on demand from a machine learning (ML) model. In the last few years, there have been notable developments in machine unlearning to remove the information of certain training data efficiently and effectively from ML models. In this work, we explore unlearning for the regression problem, particularly in deep learning models. Unlearning in classification and simple linear regression has been considerably investigated. However, unlearning in deep regression models largely remains an untouched problem till now. In this work, we introduce deep regression unlearning methods that generalize well and are robust to privacy attacks. We propose the Blindspot unlearning method which uses a novel weight optimization process. A randomly initialized model, partially exposed to the retain samples and a copy of the original model are used together to selectively imprint knowledge about the data that we wish to keep and scrub off the information of the data we wish to forget. We also propose a Gaussian fine tuning method for regression unlearning. The existing unlearning metrics for classification are not directly applicable to regression unlearning. Therefore, we adapt these metrics for the regression setting. We conduct regression unlearning experiments for computer vision, natural language processing and forecasting applications. Our methods show excellent performance for all these datasets across all the metrics. Source code: https://github.com/ayu987/deep-regression-unlearning

Added

2026-10-03

Quantifying the Persona Effect in LLM Simulations

Quantifying the Persona Effect in LLM Simulations

Tiancheng Hu, Nigel Collier

OrganizationsUniversity of Cambridge

Why you should read this

Quantifies the limits and efficacy of persona prompting across subjective NLP tasks, establishing that demographic variables explain under ten percent of annotation variance yet enable large language models to recover most predictable human variation when strong correlations exist.

Large language models (LLMs) have shown remarkable promise in simulating human language and behavior. This study investigates how integrating persona variables—demographic, social, and behavioral factors—impacts LLMs’ ability to simulate diverse perspectives. We find that persona variables account for <10% variance in annotating existing subjective NLP datasets. Nonetheless, incorporating persona variables via prompting in LLMs provides modest but statistically significant improvements. Persona prompting is most effective in samples where many annotators disagree, but their disagreements are relatively minor. Notably, we find a linear relationship in our setting: the stronger the correlation between persona variables and human annotations, the more accurate the LLM predictions are using persona prompting. In a zero-shot setting, a powerful 70b model with persona prompting captures 81% of the annotation variance achievable by linear regression trained on ground truth annotations. However, for most subjective NLP datasets, where persona variables have limited explanatory power, the benefits of persona prompting are limited.

Added

2026-09-30

On Model Selection Consistency of Lasso

On Model Selection Consistency of Lasso

Peng Zhao, Bin Yu

OrganizationsUniversity of California Berkeley

Why you should read this

Establishes the Irrepresentable Condition as an almost necessary and sufficient criterion for the Lasso to achieve consistent model selection in both classical and high-dimensional linear regression settings.

Sparsity or parsimony of statistical models is crucial for their proper interpretations, as in sciences and social sciences. Model selection is a commonly used method to find such models, but usually involves a computationally heavy combinatorial search. Lasso (Tibshirani, 1996) is now being used as a computationally feasible alternative to model selection. Therefore it is important to study Lasso for model selection purposes. In this paper, we prove that a single condition, which we call the Irrepresentable Condition, is almost necessary and sufficient for Lasso to select the true model both in the classical fixed p setting and in the large p setting as the sample size n gets large. Based on these results, sufficient conditions that are verifiable in practice are given to relate to previous works and help applications of Lasso for feature selection and sparse representation. This Irrepresentable Condition, which depends mainly on the covariance of the predictor variables, states that Lasso selects the true model consistently if and (almost) only if the predictors that are not in the true model are “irrepresentable” (in a sense to be clarified) by predictors that are in the true model. Furthermore, simulations are carried out to provide insights and understanding of this result.

Added

2026-09-14

The Hundred-Page Reinforcement Learning Book

Andriy Burkov

Added

2025-05-28

License

Read first, buy later