keyword
Independent component analysis
Independent component analysis is a statistical and computational technique used to separate a multivariate signal into additive, mutually independent underlying sources. It operates under the premise that the observed data consists of linear mixtures of latent source variables that are statistically independent of one another and possess non-Gaussian distributions. Unlike methods such as principal component analysis that only decorrelate variables using second-order statistics, independent component analysis exploits higher-order statistical properties to recover both the unobserved source signals and the unknown mixing process. The technique is widely utilized across machine learning, digital signal processing, biomedical data analysis such as electroencephalography, and causal discovery to isolate distinct subcomponents and eliminate noise or artifacts from complex datasets.
17 items

Synergies between Disentanglement and Sparsity: Generalization and Identifiability in Multi-Task Learning
Sébastien Lachapelle, Tristan Deleu, Divyat Mahajan, Ioannis Mitliagkas, Yoshua Bengio, Simon Lacoste-Julien, Quentin Bertrand
Why you should read this
Proves that pairing disentangled representations with sparse predictors improves generalization, and leverages this finding to introduce an identifiable multi-task bi-level optimization method that achieves competitive few-shot classification performance using only a fraction of learned features per task.
Although disentangled representations are often said to be beneficial for downstream tasks, current empirical and theoretical understanding is limited. In this work, we provide evidence that disentangled representations coupled with sparse task-specific predictors improve generalization. In the context of multi-task learning, we prove a new identifiability result that provides conditions under which maximally sparse predictors yield disentangled representations. Motivated by this theoretical result, we propose a practical approach to learn disentangled representations based on a sparsity-promoting bi-level optimization problem. Finally, we explore a meta-learning version of this algorithm based on group Lasso multiclass SVM predictors, for which we derive a tractable dual formulation. It obtains competitive results on standard few-shot classification benchmarks, while each task is using only a fraction of the learned representations.
Added
2026-10-04

Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael Gutmann, Aapo Hyvönen
Why you should read this
Introduces a novel estimation principle for unnormalized statistical models that avoids complex integration by treating parameter estimation as a classification task between data and artificial noise.
We present a new estimation principle for parameterized statistical models. The idea is to perform nonlinear logistic regression to discriminate between the observed data and some artificially generated noise, using the model log-density function in the regression nonlinearity. We show that this leads to a consistent (convergent) estimator of the parameters, and analyze the asymptotic variance. In particular, the method is shown to directly work for unnormalized models, i.e. models where the density function does not integrate to one. The normalization constant can be estimated just like any other parameter. For a tractable ICA model, we compare the method with other estimation methods that can be used to learn unnormalized models, including score matching, contrastive divergence, and maximum-likelihood where the normalization constant is estimated with importance sampling. Simulations show that noise-contrastive estimation offers the best trade-off between computational and statistical efficiency. The method is then applied to the modeling of natural images: We show that the method can successfully estimate a large-scale two-layer model and a Markov random field.
Added
2026-10-04
License
Published with permission

Provably Learning Object-Centric Representations
Jack Brady, Roland S. Zimmermann, Yash Sharma, Bernhard Schölkopf, Julius von Kügelgen, Wieland Brendel
Why you should read this
Establishes the first theoretical identifiability guarantees for unsupervised object-centric representation learning by proving that invertible, compositional inference models can recover ground-truth object slots even when objects exhibit statistical dependencies.
Learning structured representations of the visual world in terms of objects promises to significantly improve the generalization abilities of current machine learning models. While recent efforts to this end have shown promising empirical progress, a theoretical account of when unsupervised object-centric representation learning is possible is still lacking. Consequently, understanding the reasons for the success of existing object-centric methods as well as designing new theoretically grounded methods remains challenging. In the present work, we analyze when object-centric representations can provably be learned without supervision. To this end, we first introduce two assumptions on the generative process for scenes comprised of several objects, which we call compositionality and irreducibility. Under this generative process, we prove that the ground-truth object representations can be identified by an invertible and compositional inference model, even in the presence of dependencies between objects. We empirically validate our results through experiments on synthetic data. Finally, we provide evidence that our theory holds predictive power for existing object-centric models by showing a close correspondence between models’ compositionality and invertibility and their empirical identifiability.
Added
2026-10-03

Desiderata for Representation Learning: A Causal Perspective
Yixin Wang, Michael I. Jordan
Why you should read this
Establishes a rigorous causal inference framework that formalizes key representation learning desiderata—efficiency, non-spuriousness, and disentanglement—into calculable metrics and practical algorithms directly applicable to observational data.
Added
2026-10-02

Interventional Causal Representation Learning
Kartik Ahuja, Divyat Mahajan, Yixin Wang, Yoshua Bengio
Why you should read this
Proves that interventional data enables provable identification of latent causal factors without parametric distribution or graph structure assumptions by exploiting the geometric support shifts induced by perfect and imperfect interventions.
Causal representation learning seeks to extract high-level latent factors from low-level sensory data. Most existing methods rely on observational data and structural assumptions (e.g., conditional independence) to identify the latent factors. However, interventional data is prevalent across applications. Can interventional data facilitate causal representation learning? We explore this question in this paper. The key observation is that interventional data often carries geometric signatures of the latent factors’ support (i.e. what values each latent can possibly take). For example, when the latent factors are causally connected, interventions can break the dependency between the intervened latents’ support and their ancestors’. Leveraging this fact, we prove that the latent causal factors can be identified up to permutation and scaling given data from perfect do interventions. Moreover, we can achieve block affine identification, namely the estimated latent factors are only entangled with a few other latents if we have access to data from imperfect interventions. These results highlight the unique power of interventional data in causal representation learning; they can enable provable identification of latent factors without any assumptions about their distributions or dependency structure.
Added
2026-10-01

Weakly supervised causal representation learning
Johann Brehmer, Pim de Haan, Phillip Lippe, Taco S. Cohen
Why you should read this
Proves that high-level causal variables and mechanisms can be identified from pixel-level data paired across unknown interventions, and introduces implicit latent causal models to learn these structures without optimizing discrete graphs.
Learning high-level causal representations together with a causal model from unstructured low-level data such as pixels is impossible from observational data alone. We prove under mild assumptions that this representation is however identifiable in a weakly supervised setting. This involves a dataset with paired samples before and after random, unknown interventions, but no further labels. We then introduce implicit latent causal models, variational autoencoders that represent causal variables and causal structure without having to optimize an explicit discrete graph structure. On simple image data, including a novel dataset of simulated robotic manipulation, we demonstrate that such models can reliably identify the causal structure and disentangle causal variables.
Added
2026-09-30

Brain Network Transformer
Xuan Kan, Wei Dai, Hejie Cui, Zilong Zhang, Ying Guo, Carl Yang
Why you should read this
Proposes the Brain Network Transformer, a specialized graph architecture that leverages ROI connection profiles for natural positional encoding and introduces an orthonormal clustering readout to identify functional brain modules, achieving superior predictive accuracy on standard fMRI benchmarks.
Human brains are commonly modeled as networks of Regions of Interest (ROIs) and their connections for the understanding of brain functions and mental disorders. Recently, Transformer-based models have been studied over different types of data, including graphs, shown to bring performance gains widely. In this work, we study Transformer-based models for brain network analysis. Driven by the unique properties of data, we model brain networks as graphs with nodes of fixed size and order, which allows us to (1) use connection profiles as node features to provide natural and low-cost positional information and (2) learn pair-wise connection strengths among ROIs with efficient attention weights across individuals that are predictive towards downstream analysis tasks. Moreover, we propose an ORTHONORMAL CLUSTERING READOUT operation based on self-supervised soft clustering and orthonormal projection. This design accounts for the underlying functional modules that determine similar behaviors among groups of ROIs, leading to distinguishable cluster-aware node embeddings and informative graph embeddings. Finally, we re-standardize the evaluation pipeline on the only one publicly available large-scale brain network dataset of ABIDE, to enable meaningful comparison of different models. Experiment results show clear improvements of our proposed BRAIN NETWORK TRANSFORMER on both the public ABIDE and our restricted ABCD datasets. The implementation is available at https://github.com/Wayfear/BrainNetworkTransformer.
Added
2026-09-26

Linear Causal Disentanglement via Interventions
Chandler Squires, Anna Seigal, Salil S. Bhate, Caroline Uhler
Why you should read this
Establishes necessary and sufficient interventional conditions for identifiable causal disentanglement under linear mixing, proving that exactly one single-node intervention per latent variable guarantees full model recovery without requiring sparsity or graph structure restrictions.
Causal disentanglement seeks a representation of data involving latent variables that are related via a causal model. A representation is identifiable if both the latent model and the transformation from latent to observed variables are unique. In this paper, we study observed variables that are a linear transformation of a linear latent causal model. Data from interventions are necessary for identifiability: if one latent variable is missing an intervention, we show that there exist distinct models that cannot be distinguished. Conversely, we show that a single intervention on each latent variable is sufficient for identifiability. Our proof uses a generalization of the RQ decomposition of a matrix that replaces the usual orthogonal and upper triangular conditions with analogues depending on a partial order on the rows of the matrix, with partial order determined by a latent causal model. We corroborate our theoretical results with a method for causal disentanglement. We show that the method accurately recovers a latent causal model on synthetic and semi-synthetic data and we illustrate a use case on a dataset of single-cell RNA sequencing measurements.
Added
2026-09-26

Introduction to Machine Learning
Laurent Younes
Why you should read this
Establishes a rigorous mathematical foundation for core machine learning methods by directly connecting optimization theory, reproducing kernel Hilbert spaces, and concentration inequalities to supervised, generative, and unsupervised algorithms.
This book introduces the mathematical foundations and techniques that lead to the development and analysis of many of the algorithms that are used in machine learning. It starts with an introductory chapter that describes notation used throughout the book and serve at a reminder of basic concepts in calculus, linear algebra and probability and also introduces some measure theoretic terminology, which can be used as a reading guide for the sections that use these tools. The introductory chapters also provide background material on matrix analysis and optimization. The latter chapter provides theoretical support to many algorithms that are used in the book, including stochastic gradient descent, proximal methods, etc. After discussing basic concepts for statistical prediction, the book includes an introduction to reproducing kernel theory and Hilbert space techniques, which are used in many places, before addressing the description of various algorithms for supervised statistical learning, including linear methods, support vector machines, decision trees, boosting, or neural networks. The subject then switches to generative methods, starting with a chapter that presents sampling methods and an introduction to the theory of Markov chains. The following chapter describe the theory of graphical models, an introduction to variational methods for models with latent variables, and to deep-learning based generative models. The next chapters focus on unsupervised learning methods, for clustering, factor analysis and manifold learning. The final chapter of the book is theory-oriented and discusses concentration inequalities and generalization bounds.
Added
2026-09-19


Challenging Common Assumptions in the Unsupervised Learning of Disentangled Representations
Francesco Locatello, Stefan Bauer, Mario Lučić, Sylvain Gelly, Bernhard Schölkopf, Olivier Bachem
Why you should read this
Proves that unsupervised disentanglement is fundamentally impossible without inductive biases and demonstrates across 12,000 trained models that disentangled representations cannot be reliably identified or expected to improve downstream learning without supervision.
The key idea behind the unsupervised learning of disentangled representations is that real-world data is generated by a few explanatory factors of variation which can be recovered by unsupervised learning algorithms. In this paper, we provide a sober look at recent progress in the field and challenge some common assumptions. We first theoretically show that the unsupervised learning of disentangled representations is fundamentally impossible without inductive biases on both the models and the data. Then, we train more than 12000 models covering most prominent methods and evaluation metrics in a reproducible large-scale experimental study on seven different data sets. We observe that while the different methods successfully enforce properties ``encouraged'' by the corresponding losses, well-disentangled models seemingly cannot be identified without supervision. Furthermore, increased disentanglement does not seem to lead to a decreased sample complexity of learning for downstream tasks. Our results suggest that future work on disentanglement learning should be explicit about the role of inductive biases and (implicit) supervision, investigate concrete benefits of enforcing disentanglement of the learned representations, and consider a reproducible experimental setup covering several data sets.
Added
2026-09-18

A Linear Non-Gaussian Acyclic Model for Causal Discovery
Shohei Shimizu, Patrik O. Hoyer, Aapo Hyvärinen, Antti Kerminen
Why you should read this
Introduces the LiNGAM framework, which uses independent component analysis and non-Gaussian error distributions to identify the complete, directed causal structure of continuous observational data without requiring prior variable ordering.
In recent years, several methods have been proposed for the discovery of causal structure from non-experimental data. Such methods make various assumptions on the data generating process to facilitate its identification from purely observational data. Continuing this line of research, we show how to discover the complete causal structure of continuous-valued data, under the assumptions that (a) the data generating process is linear, (b) there are no unobserved confounders, and (c) disturbance variables have non-Gaussian distributions of non-zero variances. The solution relies on the use of the statistical method known as independent component analysis, and does not require any pre-specified time-ordering of the variables. We provide a complete Matlab package for performing this LiNGAM analysis (short for Linear Non-Gaussian Acyclic Model), and demonstrate the effectiveness of the method using artificially generated data and real-world data.
Added
2026-09-16

Estimation of Non-Normalized Statistical Models by Score Matching
Aapo Hyvärinen
One often wants to estimate statistical models where the probability density function is known only up to a multiplicative normalization constant. Typically, one then has to resort to Markov Chain Monte Carlo methods, or approximations of the normalization constant. Here, we propose that such models can be estimated by minimizing the expected squared distance between the gradient of the log-density given by the model and the gradient of the log-density of the observed data. While the estimation of the gradient of log-density function is, in principle, a very difficult non-parametric problem, we prove a surprising result that gives a simple formula for this objective function. The density function of the observed data does not appear in this formula, which simplifies to a sample average of a sum of some derivatives of the log-density given by the model. The validity of the method is demonstrated on multivariate Gaussian and independent component analysis models, and by estimating an overcomplete filter set for natural image data.
Added
2026-09-16

Machine learning for neuroimaging with scikit-learn
Alexandre Abraham, Fabian Pedregosa, Michael Eickenberg, Philippe Gervais, Andreas Muller, Jean Kossaifi, Alexandre Gramfort, Bertrand Thirion, Gäel Varoquaux
Why you should read this
Demonstrates how to apply scikit-learn to functional neuroimaging datasets, providing practical implementations for high-dimensional brain decoding, encoding, and resting-state fMRI analysis.
Statistical machine learning methods are increasingly used for neuroimaging data analysis. Their main virtue is their ability to model high-dimensional datasets, e.g. multivariate analysis of activation images or resting-state time series. Supervised learning is typically used in decoding or encoding settings to relate brain images to behavioral or clinical observations, while unsupervised learning can uncover hidden structures in sets of images (e.g. resting state functional MRI) or find sub-populations in large cohorts. By considering different functional neuroimaging applications, we illustrate how scikit-learn, a Python machine learning library, can be used to perform some key analysis steps. Scikit-learn contains a very large set of statistical learning algorithms, both supervised and unsupervised, and its application to neuroimaging data provides a versatile tool to study the brain.
Added
2026-09-16

A New Learning Algorithm for Blind Signal Separation
S. Amari, A. Cichocki, H. Yang
Why you should read this
Derives an equivariant online blind signal separation algorithm by minimizing mutual information through a Gram-Charlier expansion and the natural gradient method, establishing effective non-monotonic activation functions for neural network implementations.
A new on-line learning algorithm which minimizes a statistical dependency among outputs is derived for blind separation of mixed signals. The dependency is measured by the average mutual information (MI) of the outputs. The source signals and the mixing matrix are unknown except for the number of the sources. The Gram-Charlier expansion instead of the Edgeworth expansion is used in evaluating the MI. The natural gradient approach is used to minimize the MI. A novel activation function is proposed for the on-line learning algorithm which has an equivariant property and is easily implemented on a neural network like model. The validity of the new learning algorithm are verified by computer simulations.
Added
2026-09-16

Independent Component Analysis of Electroencephalographic Data
S. Makeig, A. Bell, T. Jung, T. Sejnowski
Why you should read this
Demonstrates how independent component analysis can effectively separate mixed scalp electroencephalographic signals into distinct neurobiological sources and artifacts without requiring prior spatial information.
Because of the distance between the skull and brain and their different resistivities, electroencephalographic (EEG) data collected from any point on the human scalp includes activity generated within a large brain area. This spatial smearing of EEG data by volume conduction does not involve significant time delays, however, suggesting that the Independent Component Analysis (ICA) algorithm of Bell and Sejnowski [1] is suitable for performing blind source separation on EEG data. The ICA algorithm separates the problem of source identification from that of source localization. First results of applying the ICA algorithm to EEG and event-related potential (ERP) data collected during a sustained auditory detection task show: (1) ICA training is insensitive to different random seeds. (2) ICA may be used to segregate obvious artifactual EEG components (line and muscle noise, eye movements) from other sources. (3) ICA is capable of isolating overlapping EEG phenomena, including alpha and theta bursts and spatially-separable ERP components, to separate ICA channels. (4) Nonstationarities in EEG and behavioral state can be tracked using ICA via changes in the amount of residual correlation between ICA-filtered output channels.
Added
2026-09-14

Learning deep representations by mutual information estimation and maximization
R. Devon Hjelm, Alex Fedorov, Samuel Lavoie-Marchildon, Karan Grewal, Adam Trischler, Yoshua Bengio
Why you should read this
Introduces Deep InfoMax (DIM), an unsupervised representation learning method that maximizes mutual information between local input features and global encoder representations while using adversarial prior matching to rival supervised classification performance.
In this work, we perform unsupervised learning of representations by maximizing mutual information between an input and the output of a deep neural network encoder. Importantly, we show that structure matters: incorporating knowledge about locality of the input to the objective can greatly influence a representation's suitability for downstream tasks. We further control characteristics of the representation by matching to a prior distribution adversarially. Our method, which we call Deep InfoMax (DIM), outperforms a number of popular unsupervised learning methods and competes with fully-supervised learning on several classification tasks. DIM opens new avenues for unsupervised learning of representations and is an important step towards flexible formulations of representation-learning objectives for specific end-goals.
Added
2026-09-13

Pen and Paper Exercises in Machine Learning
Michael U. Gutmann
Why you should read this
This book would help anyone who wants to deeply understand the mathematical foundations of machine learning through hands-on practice, covering essential topics from linear algebra and optimization to graphical models and variational inference with structured exercises and solutions.
This is a collection of (mostly) pen-and-paper exercises in machine learning. The exercises are on the following topics: linear algebra, optimisation, directed graphical models, undirected graphical models, expressive power of graphical models, factor graphs and message passing, inference for hidden Markov models, model-based learning (including ICA and unnormalised models), sampling and Monte-Carlo integration, and variational inference.
Added
2025-10-08

