Built independently by an author, for readers. Read the story and support ChapterPal

keyword

prior knowledge

In machine learning and statistical modeling, prior knowledge refers to information, domain expertise, or structural assumptions integrated into a computational model before it processes or learns from new empirical data. This knowledge can be represented in various forms, including informative prior probability distributions over parameters, specialized model architectures, explicit domain constraints, or pre-configured weight initializations. By incorporating preexisting information into the learning algorithm, models can effectively constrain the search space, accelerate convergence, enhance sample efficiency, and improve generalization performance when training on limited or noisy datasets.

2 items

Exact learning dynamics of deep linear networks with prior knowledge

Exact learning dynamics of deep linear networks with prior knowledge

Lukas Braun, Clémentine C. J. Dominé, James Fitzgerald, Andrew M. Saxe

OrganizationsCIFARHoward Hughes Medical InstituteUniversity College LondonUniversity of Oxford

Why you should read this

Derives exact analytical solutions for the learning dynamics of deep linear networks from structured initial weights, revealing how prior knowledge changes training speed and weight alignment across tasks such as continual and reversal learning.

Learning in deep neural networks is known to depend critically on the knowledge embedded in the initial network weights. However, few theoretical results have precisely linked prior knowledge to learning dynamics. Here we derive exact solutions to the dynamics of learning with rich prior knowledge in deep linear networks by generalising Fukumizu's matrix Riccati solution [1]. We obtain explicit expressions for the evolving network function, hidden representational similarity, and neural tangent kernel over training for a broad class of initialisations and tasks. The expressions reveal a class of task-independent initialisations that radically alter learning dynamics from slow non-linear dynamics to fast exponential trajectories while converging to a global optimum with identical representational similarity, dissociating learning trajectories from the structure of initial internal representations. We characterise how network weights dynamically align with task structure, rigorously justifying why previous solutions successfully described learning from small initial weights without incorporating their fine-scale structure. Finally, we discuss the implications of these findings for continual learning, reversal learning and learning of structured knowledge. Taken together, our results provide a mathematical toolkit for understanding the impact of prior knowledge on deep learning.

Added

2026-09-26

Learning Bayesian networks: The combination of knowledge and statistical data

Learning Bayesian networks: The combination of knowledge and statistical data

David Heckerman, Dan Geiger, David Maxwell Chickering

OrganizationsMicrosoftTechnion – Israel Institute of Technology

Why you should read this

Develops the foundational BDe scoring metric for learning Bayesian networks, establishing the principles of event equivalence and parameter modularity to integrate observational data with prior domain knowledge represented as a single prior network and equivalent sample size.

We describe a Bayesian approach for learning Bayesian networks from a combination of prior knowledge and statistical data. First and foremost, we develop a methodology for assessing informative priors needed for learning. Our approach is derived from a set of assumptions made previously as well as the assumption of likelihood equivalence, which says that data should not help to discriminate network structures that represent the same assertions of conditional independence. We show that likelihood equivalence when combined with previously made assumptions implies that the user’s priors for network parameters can be encoded in a single Bayesian network for the next case to be seen—a prior network—and a single measure of confidence for that network. Second, using these priors, we show how to compute the relative posterior probabilities of network structures given data. Third, we describe search methods for identifying network structures with high posterior probabilities. We describe polynomial algorithms for finding the highest-scoring network structures in the special case where every node has at most k = 1 parent. For the general case (k > 1), which is NP-hard, we review heuristic search algorithms including local search, iterative local search, and simulated annealing. Finally, we describe a methodology for evaluating Bayesian-network learning algorithms, and apply this approach to a comparison of various approaches.

Added

2026-09-11