An Empirical Study of Smoothing Techniques for Language Modeling
Stanley F. ChenJoshua Goodman
Evaluates prominent n-gram smoothing methods across diverse corpora and training set sizes while introducing novel interpolation techniques that consistently outperform classical baselines.
This paper presents a large-scale empirical comparison of smoothing methods used to build n-gram language models, which assign probabilities to word sequences and are central to applications such as speech recognition. Without effective smoothing, maximum-likelihood estimates from limited training data produce many zero probabilities that harm performance; earlier comparisons had examined only a few methods on single corpora and data sizes, leaving practitioners without clear guidance on which approach to choose.
The work set out to measure how the relative accuracy of established smoothing techniques varies with training-set size, corpus, and n-gram order (bigram versus trigram), while also testing two new methods. Performance was quantified by cross-entropy on held-out test text.
The authors implemented eight smoothing families—including additive smoothing, Katz smoothing, Church-Gale smoothing, and Jelinek-Mercer interpolation—on corpora ranging from one million to more than 100 million words. They trained models on data sets from 100 sentences to several million, optimized free parameters on separate development sets, and repeated runs on smaller data to assess statistical significance.
Katz and standard Jelinek-Mercer interpolation performed consistently well across conditions, with Katz showing a modest edge on bigrams and on large-data trigrams. Church-Gale smoothing was best on the largest bigram sets but lagged elsewhere. Additive smoothing performed poorly in all settings. The two new techniques—one that re-buckets interpolation weights by average count per observed word and one that adds a count proportional to the number of singletons—matched or exceeded the best prior methods on bigrams and were clearly superior on trigrams. Suboptimal parameter choices or use of deleted rather than held-out interpolation could increase cross-entropy by several tenths of a bit, corresponding to roughly 10–20 percent higher perplexity.
These differences matter because language-model quality directly affects error rates and search effort in downstream systems. The results indicate that no single existing method is optimal for every operating regime, so developers should match the smoother to expected data size and n-gram order rather than defaulting to the most common choice.
The clearest next step is to measure whether the observed cross-entropy gains translate into word-error-rate reductions in a full speech recognizer or other end application. Additional work could also test whether the new methods remain advantageous when extended to higher-order n-grams or to tasks such as tagging and parsing.
The study covers multiple corpora and data scales, yet all test sets are drawn from newswire or balanced written text; results may shift for other domains or for vocabularies an order of magnitude larger. Parameter optimization was limited on the biggest training sets for computational reasons, so the reported figures for those conditions carry somewhat lower confidence.
- Paper: Word Association Norms, Mutual Information, and Lexicography, Kenneth Ward Church et al. (1989). This paper establishes foundational corpus-based statistical estimation methods and mutual information metrics that underpin early n-gram statistical modeling and count estimation.
- Paper: A study of smoothing methods for language models applied to Ad Hoc information retrieval, ChengXiang Zhai et al. (2001). This work extends Chen and Goodman's systematic analysis of n-gram smoothing techniques to document-level language models used in information retrieval.
- Paper: TnT – A Statistical Part-of-Speech Tagger, Thorsten Brants (2000). This paper directly applies linear interpolation and n-gram smoothing principles to build an efficient, highly competitive statistical part-of-speech tagger.
- Paper: A Neural Probabilistic Language Model, Yoshua Bengio et al. (2003). This seminal work uses classical smoothed n-gram language models as the primary baseline to demonstrate how distributed neural representations overcome traditional n-gram sparsity.
