topic
applied mathematics
Applied mathematics is the branch of mathematics focused on developing and utilizing mathematical methods, models, and computational techniques to solve practical problems across science, engineering, business, and computer science. Unlike pure mathematics, which investigates abstract concepts independently of practical use, applied mathematics integrates theoretical principles with domain-specific knowledge to formulate algorithms, analyze numerical data, and optimize complex systems. Within computing, it supplies the quantitative foundation for areas such as numerical analysis, statistical modeling, discrete mathematics, and high-performance computation, driving the design, analysis, and execution of practical computational solutions.
24 items

Policy Gradient Methods Find the Nash Equilibrium in N-player General-sum Linear-quadratic Games
Ben M. Hambly, Renyuan Xu, Huining Yang
Why you should read this
Proves that natural policy gradient methods achieve global linear convergence to the Nash equilibrium in finite-horizon N-player general-sum linear-quadratic games by showing that sufficient system noise prevents optimization failures common in deterministic settings.
We consider a general-sum N-player linear-quadratic game with stochastic dynamics over a finite horizon and prove the global convergence of the natural policy gradient method to the Nash equilibrium. In order to prove convergence of the method we require a certain amount of noise in the system. We give a condition, essentially a lower bound on the covariance of the noise in terms of the model parameters, in order to guarantee convergence. We illustrate our results with numerical experiments to show that even in situations where the policy gradient method may not converge in the deterministic setting, the addition of noise leads to convergence.
Added
2026-10-06


Using the Nyström Method to Speed Up Kernel Machines
Christopher K. I. Williams, Matthias Seeger
Why you should read this
Cannot be determined due to unreadable document content.
The provided text consists of corrupted character encodings and symbols, preventing the extraction of a coherent abstract.
Added
2026-10-04
License
Published with permission

Introduction to Advanced Engineering Mathematics and Analysis
Brian D. Wood
Why you should read this
Builds a rigorous yet intuitive foundation in applied mathematics for modeling and solving engineering and scientific problems.
An introduction to applied mathematics written for students in engineering and science. Focus is on a rigorous presentation that also builds understanding by discussion, analogy, and examples. Discussion of concepts involved in modeling physical processes is a central theme in the text. Updated with new chapter on feedforward neural networks.<br /><br />A full version (1.3) of this textbook can be <a href="https://open.oregonstate.education/app/uploads/sites/246/2023/06/Introduction_to_Advanced_Engineering_Mathematics_and_AnalysisA.pdf">downloaded here.</a><br /><a class="link-button" href="https://forms.office.com/Pages/ResponsePage.aspx?id=4QVtzl48Yk2HqExKJxPBE_Ob1FlDquBLm_7cjZkcdXpUMU9XMElGQ1Q3TFFTTUI1TzREVVU2TkZWMi4u">Adoption Form</a>
Added
2026-09-25


Hidden physics models: Machine learning of nonlinear partial differential equations
Maziar Raissi, George Em Karniadakis
Why you should read this
Introduces hidden physics models, a Gaussian process-based machine learning framework that identifies and discovers governing nonlinear partial differential equations from sparse experimental data.
While there is currently a lot of enthusiasm about "big data", useful data is usually "small" and expensive to acquire. In this paper, we present a new paradigm of learning partial differential equations from {\em small} data. In particular, we introduce \emph{hidden physics models}, which are essentially data-efficient learning machines capable of leveraging the underlying laws of physics, expressed by time dependent and nonlinear partial differential equations, to extract patterns from high-dimensional data generated from experiments. The proposed methodology may be applied to the problem of learning, system identification, or data-driven discovery of partial differential equations. Our framework relies on Gaussian processes, a powerful tool for probabilistic inference over functions, that enables us to strike a balance between model complexity and data fitting. The effectiveness of the proposed approach is demonstrated through a variety of canonical problems, spanning a number of scientific domains, including the Navier-Stokes, Schrödinger, Kuramoto-Sivashinsky, and time dependent linear fractional equations. The methodology provides a promising new direction for harnessing the long-standing developments of classical methods in applied mathematics and mathematical physics to design learning machines with the ability to operate in complex domains without requiring large quantities of data.
Added
2026-09-25

Stochastic Gradient Descent over P2
Maria Oprea, Qin Li, Yunan Yang
Why you should read this
Establishes a rigorous diffusion approximation framework for stochastic gradient descent over the Wasserstein space of probability measures, proving that a Gaussian random field captures discrete optimization dynamics with second-order weak accuracy.
Stochastic gradient descent (SGD) admits diffusion approximations that replace the complicated randomness of stochastic gradients by Gaussian noise, providing a powerful tool for understanding its dynamics and long-time behavior. We investigate whether an analogous approximation principle holds for optimization over probability measures, where the objective is a functional defined on the Wasserstein space P2. The nonlinear geometry and infinite-dimensional nature of P2 prevent a direct extension of the classical Euclidean theory. Using Lions differentiability, we lift the problem to a linear Hilbert space, where higher-order differential calculus becomes available. We then construct a Gaussian random-field approximation whose velocity field matches the mean and covariance of the original stochastic gradient. By exploiting this moment matching through higher-order Taylor expansions, we show that the Gaussian approximation captures the SGD dynamics with second-order weak accuracy. Our result provides a rigorous foundation for replacing sample-driven randomness by analytically tractable Gaussian fluctuations in stochastic optimization over probability measures.
Added
2026-09-17


Implicit Neural Representations with Periodic Activation Functions
Vincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman, David B. Lindell, Gordon Wetzstein
Why you should read this
Introduces sinusoidal representation networks with periodic activation functions to model complex continuous signals and their derivatives, enabling high-fidelity representation of audio, video, and 3D shapes as well as direct solutions to partial differential equations.
Implicitly defined, continuous, differentiable signal representations parameterized by neural networks have emerged as a powerful paradigm, offering many possible benefits over conventional representations. However, current network architectures for such implicit neural representations are incapable of modeling signals with fine detail, and fail to represent a signal's spatial and temporal derivatives, despite the fact that these are essential to many physical signals defined implicitly as the solution to partial differential equations. We propose to leverage periodic activation functions for implicit neural representations and demonstrate that these networks, dubbed sinusoidal representation networks or Sirens, are ideally suited for representing complex natural signals and their derivatives. We analyze Siren activation statistics to propose a principled initialization scheme and demonstrate the representation of images, wavefields, video, sound, and their derivatives. Further, we show how Sirens can be leveraged to solve challenging boundary value problems, such as particular Eikonal equations (yielding signed distance functions), the Poisson equation, and the Helmholtz and wave equations. Lastly, we combine Sirens with hypernetworks to learn priors over the space of Siren functions.
Added
2026-09-16
License
Published with permission
Introduction to Engineering Mathematics and Analysis: Modeling Physical Systems Using the Language of Mathematics
Brian D Wood
Why you should read this
Presents a systematic framework for modeling physical systems by treating mathematics as a formal language, bridging rigorous analytical methods with practical engineering intuition and modern topics like feedforward neural networks.
An introduction to applied mathematics written for students in engineering and science. Focus is on a rigorous presentation that also builds understanding by discussion, analogy, and examples. Discussion of concepts involved in modeling physical processes is a central theme in the text. Updated with new chapter on feedforward neural networks.
Added
2026-09-14


Storing Infinite Dynamical Attractors in Nonreciprocal Associative Neural Networks
Miguel Aguilera, Daniele De Martino
Why you should read this
Establishes a dynamical mean-field theory for nonreciprocal associative networks, proving that coupling eigenvalue decoherence cancels destructive retarded noise to allow the storage of an extensive number of limit cycles and chaotic attractors.
We develop a dynamical mean-field theory for nonreciprocal associative networks that store an extensive number of dynamical attractors, from limit cycles to strange attractors. Using a path integral calculation under quenched disorder, we derive self-consistent dynamical mean-field equations for pattern overlaps, autocorrelations and response functions. Memory retrieval capacity is governed by the spectral structure of the coupling matrices encoding stored patterns. When their eigenvalues are coherently aligned, retarded self-interactions and quenched noise feed back destructively: at zero eigenphase (fixed point attractors) the classical equilibrium capacity bound is recovered, while for limit cycles retrieval collapses far below it. In contrast, for uniformly distributed eigenphases, retarded self-interactions and much of the quenched noise cancels, reducing the dynamics to an effective single-spin process and amplifying capacity substantially. We validate the theory against microscopic simulations for limit-cycle and chaotic attractors, identifying eigenvalue decoherence as the mechanism enabling enhanced storage of dynamical memories.
Added
2026-09-12


An elementary introduction to information geometry
Frank Nielsen
Why you should read this
Presents a self-contained foundation of information geometry, connecting differential-geometric concepts on probability distributions to concrete applications in statistical decision-making, hypothesis testing, and machine learning.
In this survey, we describe the fundamental differential-geometric structures of information manifolds, state the fundamental theorem of information geometry, and illustrate some use cases of these information manifolds in information sciences. The exposition is self-contained by concisely introducing the necessary concepts of differential geometry, but proofs are omitted for brevity.
Added
2026-09-08


Correlated initialization of deep residual networks
Felix Benning, Ivan Nourdin, Giovanni Peccati
Why you should read this
Proves that layer-correlated weight initializations enable deep residual networks in the infinite-depth limit to converge to Young differential equations driven by Hermite processes, establishing correlation decay as a tunable hyperparameter that bridges the gap between deterministic and Brownian scaling regimes.
We study the large-depth behavior of residual networks whose weights are correlated across layers at initialization. Our results confirm and extend a conjecture of Marion et al. [2025], according to which correlated initializations should interpolate continuously between the Brownian stochastic differential equation arising from independent initialization and the ordinary differential equation arising from perfectly correlated initialization. When the initialization is obtained from the application of a feature function to a stationary Gaussian sequence with regularly varying correlation, we prove that there exists a unique critical scaling such that the infinite-depth limit is the solution of a Young differential equation driven by a Hermite process. Hermite processes reduce to the fractional Brownian motion if the feature function generating the initialization has Hermite rank one, which is the case for the identity function, for example. We show that the critical scaling and asymptotic limit are uniquely determined by the decay of correlations together with the Hermite rank of the feature function. Consequently, the correlation structure and Hermite rank of the initialization represent meaningful hyperparameters in the asymptotic regime. By contrast, under finite-variance iid initialization, the asymptotic driver is universally Brownian up to normalization regardless of the choice of distribution. Our proofs rely on a collection of novel results establishing a robust stability theory for Young differential equations in Banach spaces.
Added
2026-09-05


Nora: Normalized Orthogonal Row Alignment for Scalable Matrix Optimizer
Jinghui Yuan, Jiaxuan Zou, Shuo Wang, Yong Liu, Feiping Nie
Why you should read this
Proposes Nora, a matrix optimizer for large language models that achieves second-order preconditioning efficiency and scale-invariant training stability at optimal linear computational complexity by projecting row-wise momentum onto the orthogonal complement of the weights.
Matrix-based optimizers have demonstrated immense potential in training Large Language Models (LLMs), however, designing an ideal optimizer remains a formidable challenge. A superior optimizer must satisfy three core desiderata: efficiency, achieving Muon-like preconditioning to accelerate optimization; stability, strictly adhering to the scale-invariance inherent in neural networks; and speed, minimizing computational overhead. While existing methods address these aspects to varying degrees, they often fail to unify them, either incurring prohibitive computational costs like Muon, or allowing radial jitters that compromise stability like RMNP. To bridge this gap, we propose Nora, an optimizer that rigorously satisfies all three requirements. Nora achieves training stability by explicitly stabilizing weight norms and angular velocities through row-wise momentum projection onto the orthogonal complement of the weights. Simultaneously, by leveraging the block-diagonal dominance of the Transformer Hessian, Nora effectively approximates structured preconditioning while maintaining an optimal computational complexity of . Furthermore, we prove that Nora is a scalable optimizer and establish its corresponding scaling theorems. With a streamlined implementation requiring only two lines of code, our preliminary experiments validate Nora as an efficient and highly promising optimizer for large-scale training.
Added
2026-09-03


Accelerating Scientific Research with Gemini in the Real-World
Samuel Schmidgall, Xiaokai Zhu, Marian Shaw, Lin Yang, Valentin Liévin, Jingyun Yang, Yuchen Zhuang, Tim Strother, Alex Bijamov, Min Woo Sun, Anil Palepu, Justin Chen, David Steiner, Jacqueline Shreibati, Wei-Hung Weng, Yilin Zhao, Xingjian Hu, Nicholas Zahn, Sadhya Garg, Julia Kirby, Yuxiang Gan, Jiaoli Li, Divy Thakkar, Shekoofeh Azizi, David Racz, Juraj Gottweis, Vivek Natarajan, Chenglin Wu, Tal Danino, Keran Rong, Haozhe Wang, Benoit Schillings, Yong Cheng, Quoc V. Le, Tao Tu
Why you should read this
Demonstrates how a Gemini-powered multi-agent system conducts closed-loop scientific research across materials synthesis, bacterial phenotype prediction, and medical AI design with real-world physical and clinical validation.
We present an extension and comprehensive real-world validation of Co-Scientist, a Gemini-based multi-agent system designed to accelerate end-to-end scientific research across hypothesis generation, experimentation, and manuscript generation. Moving beyond in silico hypothesis generation, this specialized configuration transitions Co-Scientist into an execution-grounded research partner advancing closed-loop scientific workflows across materials science, biology, and computer science. In materials science, Co-Scientist interfaced with a semi-automated chemical vapor deposition reactor to design a safe precursor route for MXenes; experimental execution produced a lamellar 2D material sharing key structural similarities with the Ti3C2Tx MXene lattice, although further experiments are needed to confirm the atomic structure. Leveraging Gemini 3 Deep Think for rapid, lab-in-the-loop execution, it also tailored growth recipes to laboratory constraints in minutes, enabling single-attempt growth of monolayer MoS2, MoSe2, and WS2 semiconductors. In biology, Co-Scientist predicted emergent swarming phenotypes of engineered E. coli across inducer (IPTG) gradients from sparse imaging data, quantitatively matching unpublished wet-lab morphological measurements. In computer science, Co-Scientist autonomously discovered an inference-time scaling architecture that outperformed six frontier models on HealthBench (Hard and Professional) while reducing potential clinical harm under blinded physician evaluation. Finally, a double-blind study of end-to-end generated papers with 30 domain experts across 450 reviews demonstrates that Co-Scientist's reliability modules reduce hallucination and plagiarism while improving research safety. Together, these results demonstrate progress toward closed-loop multi-agent scientific AI systems capable of accelerating real-world scientific discovery.
Added
2026-08-29
License
Published with permission

Convex Optimization: Algorithms and Complexity
Sébastien Bubeck
Why you should read this
Presents a unified theoretical treatment of convex optimization algorithms and their complexity bounds, linking classical black-box methods with modern structural and stochastic techniques tailored for machine learning.
This monograph presents the main complexity theorems in convex optimization and their corresponding algorithms. Starting from the fundamental theory of black-box optimization, the material progresses towards recent advances in structural optimization and stochastic optimization. Our presentation of black-box optimization, strongly influenced by Nesterov's seminal book and Nemirovski's lecture notes, includes the analysis of cutting plane methods, as well as (accelerated) gradient descent schemes. We also pay special attention to non-Euclidean settings (relevant algorithms include Frank-Wolfe, mirror descent, and dual averaging) and discuss their relevance in machine learning. We provide a gentle introduction to structural optimization with FISTA (to optimize a sum of a smooth and a simple non-smooth term), saddle-point mirror prox (Nemirovski's alternative to Nesterov's smoothing), and a concise description of interior point methods. In stochastic optimization we discuss stochastic gradient descent, mini-batches, random coordinate descent, and sublinear algorithms. We also briefly touch upon convex relaxation of combinatorial problems and the use of randomness to round solutions, as well as random walks based methods.
Added
2026-08-13
License
Published with permission

When is Routing Meaningful? Diversity and Robustness in Language Model Societies
Fantine Huot, Michael Kaisers, Mirella Lapata
Why you should read this
Introduces novel metrics for evaluating the behavioral diversity and robustness of language model routing policies, demonstrating that high task accuracy can conceal fundamental issues in multi-model system design and providing practical heuristics for creating meaningful societies.
Routing policies for multi-model systems are evaluated almost exclusively on task accuracy and inference cost. We argue that two properties, orthogonal to performance, determine whether routing is meaningful. First, the society of actors must be behaviourally differentiated: if all actors respond identically, routing is vacuous. Second, the routing policy must be stable: surface-form variants of a query should be assigned to the same actor. High task accuracy is compatible with violating both properties, since a router can operate over a redundant society or assign queries inconsistently, preventing specialisation regardless of performance. We adapt Hierarchic Social Entropy (HSE) to language-model societies and introduce a perturbation-based robustness metric to diagnose these failure modes. Applied to EmbedLLM and RouterBench, we find that HSE exhibits strong diminishing returns, suggesting that a curated subset of fewer than ten agents recovers most available diversity in a large pool -- a practical coreset heuristic for society design. We further find that KNN routers gain accuracy from specialist societies but collapse in robustness under perturbation, while prompted routing remains stable across all perturbation types -- illustrating that accuracy and meaningfulness can sharply diverge.
Added
2026-07-20


Stochastic Interpolants: A Unifying Framework for Flows and Diffusions
Michael S. Albergo, Nicholas M. Boffi, Eric Vanden-Eijnden
Why you should read this
Introduces a unifying framework of "stochastic interpolants" that bridges flow-based and diffusion-based generative models, allowing exact probability density function transformations in finite time with tunable noise levels.
A class of generative models that unifies flow-based and diffusion-based methods is introduced. These models extend the framework proposed in Albergo and Vanden-Eijnden (2023), enabling the use of a broad class of continuous-time stochastic processes called stochastic interpolants to bridge any two probability density functions exactly in finite time. These interpolants are built by combining data from the two prescribed densities with an additional latent variable that shapes the bridge in a flexible way. The time-dependent density function of the interpolant is shown to satisfy a transport equation as well as a family of forward and backward Fokker-Planck equations with tunable diffusion coefficient. Upon consideration of the time evolution of an individual sample, this viewpoint leads to both deterministic and stochastic generative models based on probability flow equations or stochastic differential equations with an adjustable level of noise. The drift coefficients entering these models are time-dependent velocity fields characterized as the unique minimizers of simple quadratic objective functions, one of which is a new objective for the score. We show that minimization of these quadratic objectives leads to control of the likelihood for generative models built upon stochastic dynamics, while likelihood control for deterministic dynamics is more stringent. We also construct estimators for the likelihood and the cross entropy of interpolant-based generative models, and we discuss connections with other methods such as score-based diffusion models, stochastic localization, probabilistic denoising, and rectifying flows. In addition, we demonstrate that stochastic interpolants recover the Schrödinger bridge between the two target densities when explicitly optimizing over the interpolant. Finally, algorithmic aspects are discussed and the approach is illustrated on numerical examples.
Added
2026-05-14


KAN: Kolmogorov-Arnold Networks
Ziming Liu, Yixuan Wang, Sachin Vaidya, Fabian Ruehle, James Halverson, Marin Soljačić, Thomas Y. Hou, Max Tegmark
Why you should read this
Introduces Kolmogorov-Arnold Networks (KANs), a novel deep learning architecture that replaces fixed activation functions with learnable, spline-parameterized functions on network edges, achieving superior accuracy and interpretability compared to MLPs, particularly for scientific discovery tasks.
Inspired by the Kolmogorov-Arnold representation theorem, we propose Kolmogorov-Arnold Networks (KANs) as promising alternatives to Multi-Layer Perceptrons (MLPs). While MLPs have fixed activation functions on nodes ("neurons"), KANs have learnable activation functions on edges ("weights"). KANs have no linear weights at all -- every weight parameter is replaced by a univariate function parametrized as a spline. We show that this seemingly simple change makes KANs outperform MLPs in terms of accuracy and interpretability. For accuracy, much smaller KANs can achieve comparable or better accuracy than much larger MLPs in data fitting and PDE solving. Theoretically and empirically, KANs possess faster neural scaling laws than MLPs. For interpretability, KANs can be intuitively visualized and can easily interact with human users. Through two examples in mathematics and physics, KANs are shown to be useful collaborators helping scientists (re)discover mathematical and physical laws. In summary, KANs are promising alternatives for MLPs, opening opportunities for further improving today's deep learning models which rely heavily on MLPs.
Added
2026-05-13


Score-Based Generative Modeling through Stochastic Differential Equations
Yang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar, Stefano Ermon, Ben Poole
Why you should read this
Develops a unified stochastic differential equation (SDE) framework that not only encapsulates existing score-based and diffusion models but also achieves record-breaking performance in unconditional image generation and solves various inverse problems, offering new sampling procedures and modeling capabilities.
Creating noise from data is easy; creating data from noise is generative modeling. We present a stochastic differential equation (SDE) that smoothly transforms a complex data distribution to a known prior distribution by slowly injecting noise, and a corresponding reverse-time SDE that transforms the prior distribution back into the data distribution by slowly removing the noise. Crucially, the reverse-time SDE depends only on the time-dependent gradient field (\aka, score) of the perturbed data distribution. By leveraging advances in score-based generative modeling, we can accurately estimate these scores with neural networks, and use numerical SDE solvers to generate samples. We show that this framework encapsulates previous approaches in score-based generative modeling and diffusion probabilistic modeling, allowing for new sampling procedures and new modeling capabilities. In particular, we introduce a predictor-corrector framework to correct errors in the evolution of the discretized reverse-time SDE. We also derive an equivalent neural ODE that samples from the same distribution as the SDE, but additionally enables exact likelihood computation, and improved sampling efficiency. In addition, we provide a new way to solve inverse problems with score-based models, as demonstrated with experiments on class-conditional generation, image inpainting, and colorization. Combined with multiple architectural improvements, we achieve record-breaking performance for unconditional image generation on CIFAR-10 with an Inception score of 9.89 and FID of 2.20, a competitive likelihood of 2.99 bits/dim, and demonstrate high fidelity generation of 1024 x 1024 images for the first time from a score-based generative model.
Added
2026-04-10
License
Published with permission

Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution
Aaron Lou, Chenlin Meng, Stefano Ermon
Why you should read this
Proposes "score entropy," a novel loss function that fundamentally extends diffusion models to discrete data, enabling them to significantly outperform existing language diffusion paradigms and even GPT-2, while offering superior text generation quality and controllable infilling capabilities.
Despite their groundbreaking performance for many generative modeling tasks, diffusion models have fallen short on discrete data domains such as natural language. Crucially, standard diffusion models rely on the well-established theory of score matching, but efforts to generalize this to discrete structures have not yielded the same empirical gains. In this work, we bridge this gap by proposing score entropy, a novel loss that naturally extends score matching to discrete spaces, integrates seamlessly to build discrete diffusion models, and significantly boosts performance. Experimentally, we test our Score Entropy Discrete Diffusion models (SEDD) on standard language modeling tasks. For comparable model sizes, SEDD beats existing language diffusion paradigms (reducing perplexity by -\%) and is competitive with autoregressive models, in particular outperforming GPT-2. Furthermore, compared to autoregressive mdoels, SEDD generates faithful text without requiring distribution annealing techniques like temperature scaling (around - better generative perplexity than un-annealed GPT-2), can trade compute and quality (similar quality with fewer network evaluations), and enables controllable infilling (matching nucleus sampling quality while enabling other strategies besides left to right prompting).
Added
2026-03-12


Composable Text Controls in Latent Space with ODEs
Guangyi Liu, Zeyu Feng, Yuan Gao, Zichao Yang, Xiaodan Liang, Junwei Bao, Xiaodong He, Shuguang Cui, Zhen Li, Zhiting Hu
Why you should read this
Proposes an efficient and flexible ODE-based sampling approach for composable text operations in a compact latent space, significantly improving the quality and efficiency of generating and editing text with diverse controls.
Real-world text applications often involve composing a wide range of text control operations, such as editing the text w.r.t. an attribute, manipulating keywords and structure, and generating new text of desired properties. Prior work typically learns/finetunes a language model (LM) to perform individual or specific subsets of operations. Recent research has studied combining operations in a plug-and-play manner, often with costly search or optimization in the complex sequence space. This paper proposes a new efficient approach for composable text operations in the compact latent space of text. The low-dimensionality and differentiability of the text latent vector allow us to develop an efficient sampler based on ordinary differential equations (ODEs) given arbitrary plug-in operators (e.g., attribute classifiers). By connecting pretrained LMs (e.g., GPT2) to the latent space through efficient adaption, we then decode the sampled vectors into desired text sequences. The flexible approach permits diverse control operators (sentiment, tense, formality, keywords, etc.) acquired using any relevant data from different domains. Experiments show that composing those operators within our approach manages to generate or edit high-quality text, substantially improving over previous methods in terms of generation quality and efficiency.
Added
2026-03-12


Variational Inference with Normalizing Flows
Danilo Jimenez Rezende, Shakir Mohamed
Why you should read this
Presents a novel approach using normalizing flows, this paper demonstrates how to construct arbitrarily complex and scalable approximate posterior distributions, significantly enhancing the accuracy and applicability of variational inference.
The choice of approximate posterior distribution is one of the core problems in variational inference. Most applications of variational inference employ simple families of posterior approximations in order to allow for efficient inference, focusing on mean-field or other simple structured approximations. This restriction has a significant impact on the quality of inferences made using variational methods. We introduce a new approach for specifying flexible, arbitrarily complex and scalable approximate posterior distributions. Our approximations are distributions constructed through a normalizing flow, whereby a simple initial density is transformed into a more complex one by applying a sequence of invertible transformations until a desired level of complexity is attained. We use this view of normalizing flows to develop categories of finite and infinitesimal flows and provide a unified view of approaches for constructing rich posterior approximations. We demonstrate that the theoretical advantages of having posteriors that better match the true posterior, combined with the scalability of amortized variational approaches, provides a clear improvement in performance and applicability of variational inference.
Added
2026-03-09

