keyword
Parameter estimation
Parameter estimation is the process of using observed data to determine values for unknown parameters in a statistical or other mathematical model. It may involve estimating one or more parameters that govern the model, using methods such as maximum likelihood, and the resulting values are used to describe or fit the model to the data.
6 items

Convergence Rates for Gaussian Mixtures of Experts
Nhat Ho, Chiao-Yu Yang, Michael I. Jordan
Why you should read this
Establishes the convergence rates of maximum likelihood estimation for over-specified Gaussian mixtures of experts by connecting the algebraic independence of expert functions to partial differential equations and generalized optimal transport distances.
We provide a theoretical treatment of over-specified Gaussian mixtures of experts with covariate-free gating networks. We establish the convergence rates of the maximum likelihood estimation (MLE) for these models. Our proof technique is based on a novel notion of algebraic independence of the expert functions. Drawing on optimal transport, we establish a connection between the algebraic independence of the expert functions and a certain class of partial differential equations (PDEs) with respect to the parameters. Exploiting this connection allows us to derive convergence rates for parameter estimation.
Added
2026-10-03

GibbsDDRM: A Partially Collapsed Gibbs Sampler for Solving Blind Inverse Problems with Denoising Diffusion Restoration
Naoki Murata, Koichi Saito, Chieh-Hsin Lai, Yuhta Takida, Toshimitsu Uesaka, Yuki Mitsufuji, Stefano Ermon
Why you should read this
Proposes a partially collapsed Gibbs sampling framework that enables pre-trained diffusion models to solve blind inverse problems like image deblurring and vocal dereverberation without requiring fine-tuning or specialized priors for the unknown measurement operator.
Pre-trained diffusion models have been successfully used as priors in a variety of linear inverse problems, where the goal is to reconstruct a signal from noisy linear measurements. However, existing approaches require knowledge of the linear operator. In this paper, we propose GibbsDDRM, an extension of Denoising Diffusion Restoration Models (DDRM) to a blind setting in which the linear measurement operator is unknown. GibbsDDRM constructs a joint distribution of the data, measurements, and linear operator by using a pre-trained diffusion model for the data prior, and it solves the problem by posterior sampling with an efficient variant of a Gibbs sampler. The proposed method is problem-agnostic, meaning that a pre-trained diffusion model can be applied to various inverse problems without fine-tuning. In experiments, it achieved high performance on both blind image deblurring and vocal dereverberation tasks, despite the use of simple generic priors for the underlying linear operators.
Added
2026-10-02

Enriching the Knowledge Sources Used in a Maximum Entropy Part-of-Speech Tagger
Kristina Toutanvoa, Christopher D. Manning
Why you should read this
Demonstrates how maximum entropy part-of-speech tagging accuracy on unseen words can be substantially improved by engineering rich features targeting capitalization patterns, verb tense forms, and particle-preposition distinctions.
This paper presents results for a maximum-entropy-based part of speech tagger, which achieves superior performance principally by enriching the information sources used for tagging. In particular, we get improved results by incorporating these features: (i) more extensive treatment of capitalization for unknown words; (ii) features for the disambiguation of the tense forms of verbs; (iii) features for disambiguating particles from prepositions and adverbs. The best resulting accuracy for the tagger on the Penn Treebank is 96.86% overall, and 86.91% on previously unseen words.
Added
2026-09-24

Self-Paced Learning for Latent Variable Models
M. P. Kumar, Ben Packer, D. Koller
Why you should read this
Introduces a self-paced learning framework that dynamically selects training samples from easiest to hardest to prevent latent variable models from getting trapped in poor local optima.
Latent variable models are a powerful tool for addressing several tasks in machine learning. However, the algorithms for learning the parameters of latent variable models are prone to getting stuck in a bad local optimum. To alleviate this problem, we build on the intuition that, rather than considering all samples simultaneously, the algorithm should be presented with the training data in a meaningful order that facilitates learning. The order of the samples is determined by how easy they are. The main challenge is that often we are not provided with a readily computable measure of the easiness of samples. We address this issue by proposing a novel, iterative self-paced learning algorithm where each iteration simultaneously selects easy samples and learns a new parameter vector. The number of samples selected is governed by a weight that is annealed until the entire training data has been considered. We empirically demonstrate that the self-paced learning algorithm outperforms the state of the art method for learning a latent structural SVM on four applications: object localization, noun phrase coreference, motif finding and handwritten digit recognition.
Added
2026-09-24

Discriminative Training Methods for Hidden Markov Models: Theory and Experiments with Perceptron Algorithms
Michael Collins
Why you should read this
Introduces a computationally efficient, discriminative perceptron training algorithm for sequence labeling that matches or outperforms maximum-entropy models and CRFs on natural language processing tasks while providing theoretical convergence guarantees.
We describe new algorithms for training tagging models, as an alternative to maximum-entropy models or conditional random fields (CRFs). The algorithms rely on Viterbi decoding of training examples, combined with simple additive updates. We describe the theory justifying the algorithms through a modification of the proof of convergence of the perceptron algorithm for classification problems. We give experimental results on part-of-speech tagging and base noun phrase chunking, in both cases showing improvements over results for a maximum-entropy tagger.
Added
2026-09-16

On Model Selection Consistency of Lasso
Peng Zhao, Bin Yu
Why you should read this
Establishes the Irrepresentable Condition as an almost necessary and sufficient criterion for the Lasso to achieve consistent model selection in both classical and high-dimensional linear regression settings.
Sparsity or parsimony of statistical models is crucial for their proper interpretations, as in sciences and social sciences. Model selection is a commonly used method to find such models, but usually involves a computationally heavy combinatorial search. Lasso (Tibshirani, 1996) is now being used as a computationally feasible alternative to model selection. Therefore it is important to study Lasso for model selection purposes. In this paper, we prove that a single condition, which we call the Irrepresentable Condition, is almost necessary and sufficient for Lasso to select the true model both in the classical fixed p setting and in the large p setting as the sample size n gets large. Based on these results, sufficient conditions that are verifiable in practice are given to relate to previous works and help applications of Lasso for feature selection and sparse representation. This Irrepresentable Condition, which depends mainly on the covariance of the predictor variables, states that Lasso selects the true model consistently if and (almost) only if the predictors that are not in the true model are “irrepresentable” (in a sense to be clarified) by predictors that are in the true model. Furthermore, simulations are carried out to provide insights and understanding of this result.
Added
2026-09-14
