keyword
linear regression
Linear regression is a foundational statistical and machine learning method used to model the relationship between one or more independent predictor variables and a continuous dependent target variable by fitting a linear equation to observed data. In its simplest form, it estimates the expected outcome as a weighted sum of input features plus a constant bias term. The model parameters are commonly estimated using optimization techniques such as ordinary least squares, which minimizes the sum of squared differences between observed and predicted values, or iterative algorithms such as gradient descent. Linear regression is also frequently extended with regularization penalties, such as ridge or lasso regression, to prevent overfitting and improve generalization when dealing with correlated or high-dimensional data, making it a standard tool for predictive modeling, trend analysis, and quantifying the strength of relationships between variables.
8 items

Atari-5: Distilling the Arcade Learning Environment down to Five Games
Matthew Aitchison, Penny Sweetser, Marcus Hutter
Why you should read this
Presents a principled feature-selection method to extract a compact five-game subset of the Arcade Learning Environment that estimates full 57-game median performance within ten percent, dramatically reducing the computational barrier and carbon footprint of reinforcement learning research.
The Arcade Learning Environment (ALE) has become an essential benchmark for assessing the performance of reinforcement learning algorithms. However, the computational cost of generating results on the entire 57-game dataset limits ALE's use and makes the reproducibility of many results infeasible. We propose a novel solution to this problem in the form of a principled methodology for selecting small but representative subsets of environments within a benchmark suite. We applied our method to identify a subset of five ALE games, we call Atari-5, which generally produces 57-game median score estimates to within 10% of their true values. Extending the subset to 10-games recovers 80% of the variance for log-scores for all games within the 57-game set. We show this level of compression is possible due to a high degree of correlation between many of the games in ALE.
Added
2026-10-05

What learning algorithm is in-context learning? Investigations with linear models
Ekin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, Denny Zhou
Why you should read this
Demonstrates through constructive proofs and empirical analysis that transformers perform in-context learning by implicitly executing standard algorithms, such as gradient descent and ridge regression, directly within their internal activations.
Neural sequence models, especially transformers, exhibit a remarkable capacity for in-context learning. They can construct new predictors from sequences of labeled examples presented in the input without further parameter updates. We investigate the hypothesis that transformer-based in-context learners implement standard learning algorithms implicitly, by encoding smaller models in their activations, and updating these implicit models as new examples appear in the context. Using linear regression as a prototypical problem, we offer three sources of evidence for this hypothesis. First, we prove by construction that transformers can implement learning algorithms for linear models based on gradient descent and closed-form ridge regression. Second, we show that trained in-context learners closely match the predictors computed by gradient descent, ridge regression, and exact least-squares regression, transitioning between different predictors as transformer depth and dataset noise vary, and converging to Bayesian estimators for large widths and depths. Third, we present preliminary evidence that in-context learners share algorithmic features with these predictors: learners' late layers non-linearly encode weight vectors and moment matrices. These results suggest that in-context learning is understandable in algorithmic terms, and that (at least in the linear case) learners may rediscover standard estimation algorithms. Code and reference implementations are released at this https URL.
Added
2026-10-05

Deep Regression Unlearning
Ayush Kumar Tarun, Vikram Singh Chundawat, Murari Mandal, Mohan S. Kankanhalli
Why you should read this
Presents the first machine unlearning frameworks designed specifically for deep regression and forecasting models, introducing a weight-optimization method called Blindspot alongside adapted evaluation metrics to effectively scrub targeted data while defending against privacy attacks.
With the introduction of data protection and privacy regulations, it has become crucial to remove the lineage of data on demand from a machine learning (ML) model. In the last few years, there have been notable developments in machine unlearning to remove the information of certain training data efficiently and effectively from ML models. In this work, we explore unlearning for the regression problem, particularly in deep learning models. Unlearning in classification and simple linear regression has been considerably investigated. However, unlearning in deep regression models largely remains an untouched problem till now. In this work, we introduce deep regression unlearning methods that generalize well and are robust to privacy attacks. We propose the Blindspot unlearning method which uses a novel weight optimization process. A randomly initialized model, partially exposed to the retain samples and a copy of the original model are used together to selectively imprint knowledge about the data that we wish to keep and scrub off the information of the data we wish to forget. We also propose a Gaussian fine tuning method for regression unlearning. The existing unlearning metrics for classification are not directly applicable to regression unlearning. Therefore, we adapt these metrics for the regression setting. We conduct regression unlearning experiments for computer vision, natural language processing and forecasting applications. Our methods show excellent performance for all these datasets across all the metrics. Source code: https://github.com/ayu987/deep-regression-unlearning
Added
2026-10-03

Quantifying the Persona Effect in LLM Simulations
Tiancheng Hu, Nigel Collier
Why you should read this
Quantifies the limits and efficacy of persona prompting across subjective NLP tasks, establishing that demographic variables explain under ten percent of annotation variance yet enable large language models to recover most predictable human variation when strong correlations exist.
Large language models (LLMs) have shown remarkable promise in simulating human language and behavior. This study investigates how integrating persona variables—demographic, social, and behavioral factors—impacts LLMs’ ability to simulate diverse perspectives. We find that persona variables account for <10% variance in annotating existing subjective NLP datasets. Nonetheless, incorporating persona variables via prompting in LLMs provides modest but statistically significant improvements. Persona prompting is most effective in samples where many annotators disagree, but their disagreements are relatively minor. Notably, we find a linear relationship in our setting: the stronger the correlation between persona variables and human annotations, the more accurate the LLM predictions are using persona prompting. In a zero-shot setting, a powerful 70b model with persona prompting captures 81% of the annotation variance achievable by linear regression trained on ground truth annotations. However, for most subjective NLP datasets, where persona variables have limited explanatory power, the benefits of persona prompting are limited.
Added
2026-09-30

Stacked regressions
LEO BREIMAN
Why you should read this
Demonstrates how combining multiple regression models using cross-validation under non-negativity constraints reliably outperforms single-model selection across diverse algorithms like decision trees, subset selection, and ridge regression.
Stacking regressions is a method for forming linear combinations of different predictors to give improved prediction accuracy. The idea is to use cross-validation data and least squares under non-negativity constraints to determine the coefficients in the combination. Its effectiveness is demonstrated in stacking regression trees of different sizes and in a simulation stacking linear subset and ridge regressions. Reasons why this method works are explored. The idea of stacking originated with Wolpert (1992).
Added
2026-09-25

On Model Selection Consistency of Lasso
Peng Zhao, Bin Yu
Why you should read this
Establishes the Irrepresentable Condition as an almost necessary and sufficient criterion for the Lasso to achieve consistent model selection in both classical and high-dimensional linear regression settings.
Sparsity or parsimony of statistical models is crucial for their proper interpretations, as in sciences and social sciences. Model selection is a commonly used method to find such models, but usually involves a computationally heavy combinatorial search. Lasso (Tibshirani, 1996) is now being used as a computationally feasible alternative to model selection. Therefore it is important to study Lasso for model selection purposes. In this paper, we prove that a single condition, which we call the Irrepresentable Condition, is almost necessary and sufficient for Lasso to select the true model both in the classical fixed p setting and in the large p setting as the sample size n gets large. Based on these results, sufficient conditions that are verifiable in practice are given to relate to previous works and help applications of Lasso for feature selection and sparse representation. This Irrepresentable Condition, which depends mainly on the covariance of the predictor variables, states that Lasso selects the true model consistently if and (almost) only if the predictors that are not in the true model are “irrepresentable” (in a sense to be clarified) by predictors that are in the true model. Furthermore, simulations are carried out to provide insights and understanding of this result.
Added
2026-09-14

Dive into Deep Learning (with PyTorch)
Aston Zhang, Zachary Lipton, Mu Li, Alexander Smola
Why you should read this
Combines rigorous mathematical foundations with hands-on, executable code examples in a free, interactive format that takes you from deep learning fundamentals to state-of-the-art techniques, making it ideal whether you're a student, researcher, or practitioner looking to truly understand and implement neural networks.
Source
https://d2l.ai/Added
2025-09-22

The Hundred-Page Reinforcement Learning Book
Andriy Burkov
Added
2025-05-28
License
Read first, buy later
