A study of smoothing methods for language models applied to Ad Hoc information retrieval

ChengXiang ZhaiJ. Lafferty

article2001SIGIR1,371 citationsTest of Time Award

Reveals how language model smoothing directly connects to traditional TF-IDF weighting and document length normalization, providing systematic empirical guidance on selecting and tuning smoothing techniques across different query types and retrieval collections.

Listen

Language modeling represents a promising statistical approach to search and information retrieval by estimating a probability distribution for each document and ranking results by how likely they are to produce a user's search query. However, because individual documents contain limited text, standard probability estimates assign zero probability to words not explicitly written in a document. To resolve this data sparseness problem, language models rely on "smoothing," which adjusts probabilities to account for unseen words. Although smoothing is fundamental to retrieval quality, previous systems applied different smoothing techniques heuristically without systematic evaluation. Understanding the behavior and parameter sensitivity of these methods is essential for deploying reliable, high-performing search systems.

The article evaluates the sensitivity of retrieval performance to smoothing parameters and compares three primary smoothing methods—Jelinek-Mercer, Dirichlet prior, and absolute discounting—across diverse document collections and query types.

The researchers conducted an extensive empirical evaluation across five standard text collections, including both smaller specialized datasets and large-scale web and news corpora. They tested two distinct query formats across 100 standardized search topics: short, concise keyword queries (2–3 words) and long, verbose sentence-based queries. The study deliberately preserved all terms without stop-word removal to examine baseline model behavior, testing each smoothing method across its entire parameter spectrum under both linear interpolation and traditional backoff frameworks.

The analysis revealed several key findings regarding retrieval performance. First, retrieval effectiveness is highly sensitive to the choice of smoothing parameters, showing that smoothing acts as the statistical equivalent of term weighting and document length normalization in traditional search models. Second, query type strongly dictates optimal parameter settings; long, verbose queries require significantly heavier smoothing than concise keyword queries. Third, Dirichlet prior smoothing delivered the strongest performance on short keyword queries across nearly all test collections, achieving an average precision of 0.256 compared to 0.227 for Jelinek-Mercer and 0.236 for absolute discounting. Fourth, Jelinek-Mercer smoothing was the most effective method for long, verbose queries, yielding an average precision of 0.280 compared to 0.279 for Dirichlet and 0.261 for absolute discounting. Finally, interpolation strategies consistently and significantly outperformed backoff strategies across all evaluated methods and datasets.

These findings indicate that smoothing serves a dual purpose in information retrieval: improving document model accuracy (estimation) and filtering out common, uninformative words in verbose requests (query modeling). Dirichlet prior smoothing excels at document estimation because it automatically adapts penalties based on document length. Conversely, Jelinek-Mercer smoothing applies a uniform collection background model across all documents, making it superior at discounting non-informative query terms. Contrary to speech recognition conventions where backoff methods are standard, search systems perform worse under backoff smoothing because it penalizes long documents too aggressively.

For practical implementation, search system designers should select Dirichlet prior smoothing for short, title-style user searches (targeting an initial prior parameter around 2,000) and Jelinek-Mercer smoothing for long, descriptive queries (using a parameter around 0.7). To optimize operational performance without manual tuning across query types, future engineering efforts should develop a two-stage smoothing architecture that combines Dirichlet smoothing for document estimation with Jelinek-Mercer smoothing for query term handling, alongside automated parameter training based on past relevance data.

The findings are supported with high confidence across multiple standardized benchmark datasets. However, decision-makers should note key limitations: the parameters were optimized globally across collections via exhaustive search rather than learned dynamically in real time, no stop-word filtering was applied, and web collections showed anomalous behavior under certain configurations that requires further study before deploying in specialized web environments.

  • Paper: Relevance-Based Language Models, Victor Lavrenko et al. (2001). This paper extends the foundational language modeling and smoothing framework of information retrieval to incorporate temporal recency and relevance models directly into the scoring distribution.
  • Paper: The Probabilistic Relevance Framework: BM25 and Beyond, Stephen Robertson et al. (2009). This comprehensive monograph synthesizes classical probabilistic ranking models and modern developments, contextualizing statistical document estimation and length normalization alongside language modeling approaches.
  • Paper: SPLADE: Sparse Lexical and Expansion Model for First Stage Ranking, Thibault Formal et al. (2021). This work advances lexical and term-weighting retrieval by employing neural representations to expand queries and documents while maintaining efficient inverted-index matching.
Cover for A study of smoothing methods for language models applied to Ad Hoc information retrieval

Abstract

Language modeling approaches to information retrieval are attractive and promising because they connect the problem of retrieval with that of language model estimation, which has been studied extensively in other application areas such as speech recognition. The basic idea of these approaches is to estimate a language model for each document, and then rank documents by the likelihood of the query according to the estimated language model. A core problem in language model estimation is smoothing, which adjusts the maximum likelihood estimator so as to correct the inaccuracy due to data sparseness. In this paper, we study the problem of language model smoothing and its influence on retrieval performance. We examine the sensitivity of retrieval performance to the smoothing parameters and compare several popular smoothing methods on different test collections.

Table of Contents

  • 1. INTRODUCTION
  • 2. THE LANGUAGE MODELING APPROACH
  • 3. SMOOTHING METHODS
  • 5. BEHAVIOR OF INDIVIDUAL METHODS
  • 6. COMPARISON OF METHODS
  • 7. INTERPOLATION VS. BACKOFF
  • 8. CONCLUSIONS AND FUTURE WORK
  • ACKNOWLEDGEMENTS
  • REFERENCES

Knowls

  1. Knowl 1 — General Query-Likelihood Retrieval Score Decomposition under Language Model Smoothing

    theoretical result

    Given a query q=(q1,q2,…,qn)q = (q_1, q_2, \dots, q_n) and a document dd, under a unigram language model with a uniform document prior p(d)p(d), document ranking by posterior probability p(d∣q)p(d \mid q) is equivalent to ranking by query log-likelihood:

    log⁡p(q∣d)=∑i=1nlog⁡p(qi∣d)\log p(q \mid d) = \sum_{i=1}^n \log p(q_i \mid d)

    When the document language model is smoothed using the collection language model p(w∣C)p(w \mid C) as a fallback reference distribution, the word probability is partitioned into a seen-word model ps(w∣d)p_s(w \mid d) for words with document count c(w;d)>0c(w; d) > 0 and an unseen-word model pu(w∣d)=αdp(w∣C)p_u(w \mid d) = \alpha_d p(w \mid C) for words with c(w;d)=0c(w; d) = 0, where:

    αd=1−∑w:c(w;d)>0ps(w∣d)1−∑w:c(w;d)>0p(w∣C)\alpha_d = \frac{1 - \sum_{w: c(w; d) > 0} p_s(w \mid d)}{1 - \sum_{w: c(w; d) > 0} p(w \mid C)}

    The query log-likelihood decomposes into:

    log⁡p(q∣d)=∑i:c(qi;d)>0log⁡ps(qi∣d)αdp(qi∣C)+nlog⁡αd+∑i=1nlog⁡p(qi∣C)\log p(q \mid d) = \sum_{i: c(q_i; d) > 0} \log \frac{p_s(q_i \mid d)}{\alpha_d p(q_i \mid C)} + n \log \alpha_d + \sum_{i=1}^n \log p(q_i \mid C)

    Because ∑i=1nlog⁡p(qi∣C)\sum_{i=1}^n \log p(q_i \mid C) depends only on the query and collection, it can be dropped during ranking. This demonstrates that smoothed unigram query-likelihood retrieval implements two traditional heuristic mechanisms:

    1. Term Weighting (TF-IDF analogue): Each matched term receives the weight log⁡ps(qi∣d)αdp(qi∣C)\log \frac{p_s(q_i \mid d)}{\alpha_d p(q_i \mid C)}, which increases with document term frequency c(qi;d)c(q_i; d) and decreases with collection term frequency p(qi∣C)p(q_i \mid C).
    2. Document Length Normalization: The component nlog⁡αdn \log \alpha_d scales with query length nn and penalizes documents that allocate less probability mass to unseen terms (as longer documents typically require less smoothing and have smaller αd\alpha_d, resulting in a larger length penalty).
  2. Knowl 2 — Dual Roles of Smoothing in Language Modeling Retrieval: Estimation vs. Query Modeling

    theoretical result

    Smoothing in query-likelihood language modeling retrieval serves two distinct functions:

    1. Estimation Role: Correcting the maximum likelihood estimator for document language models to handle data sparseness from limited document lengths. This role requires the smoothing amount to adapt dynamically to document length ∣d∣|d| (shorter documents need more smoothing, longer documents need less).
    2. Query Modeling Role: Accounting for common, non-informative words present in queries by treating them as generated by the background collection language model rather than the document-specific topic. This role requires uniform smoothing across all candidate documents to discount common query terms consistently.

    Concise keyword/title queries contain few non-informative words and are heavily dominated by the estimation role; methods that adapt naturally to document length (such as Dirichlet prior smoothing) perform best on them. Verbose, long queries contain many non-informative words and are heavily influenced by the query modeling role; methods with a document-independent smoothing coefficient (such as Jelinek-Mercer smoothing with a high λ\lambda) perform best on them.

  3. Knowl 3 — Bayesian Dirichlet Prior Smoothing for Information Retrieval

    model/method

    Bayesian Dirichlet prior smoothing models the document term multinomial distribution using a conjugate Dirichlet prior with pseudo-count parameters (μp(w1∣C),μp(w2∣C),…,μp(wV∣C))(\mu p(w_1 \mid C), \mu p(w_2 \mid C), \dots, \mu p(w_V \mid C)), where μ>0\mu > 0 is the prior sample size parameter, p(w∣C)p(w \mid C) is the collection language model, and VV is vocabulary size:

    pμ(w∣d)=c(w;d)+μp(w∣C)∣d∣+μp_\mu(w \mid d) = \frac{c(w; d) + \mu p(w \mid C)}{|d| + \mu}

    where c(w;d)c(w; d) is the count of word ww in document dd and ∣d∣=∑wc(w;d)|d| = \sum_w c(w; d) is the document length.

    In the general smoothing framework:

    • Seen word model: ps(w∣d)=c(w;d)+μp(w∣C)∣d∣+μp_s(w \mid d) = \frac{c(w; d) + \mu p(w \mid C)}{|d| + \mu}
    • Unseen word coefficient: αd=μ∣d∣+μ\alpha_d = \frac{\mu}{|d| + \mu}
    • Matched term weight: log⁡(1+c(qi;d)μp(qi∣C))=log⁡(1+∣d∣μpml(qi∣d)p(qi∣C))\log \left( 1 + \frac{c(q_i; d)}{\mu p(q_i \mid C)} \right) = \log \left( 1 + \frac{|d|}{\mu} \frac{p_{ml}(q_i \mid d)}{p(q_i \mid C)} \right)

    Because αd\alpha_d is inversely related to document length ∣d∣|d|, the nlog⁡αdn \log \alpha_d term provides document length normalization. Smaller μ\mu places more emphasis on relative term weights, while as μ→∞\mu \to \infty, αd→1\alpha_d \to 1 and term weights approach coordination-level matching (counting the number of matched query terms).

  4. Knowl 4 — Jelinek-Mercer Interpolation Smoothing for Information Retrieval

    model/method

    Jelinek-Mercer smoothing linearly interpolates the document maximum likelihood model pml(w∣d)=c(w;d)∣d∣p_{ml}(w \mid d) = \frac{c(w; d)}{|d|} with the collection language model p(w∣C)p(w \mid C) using a constant interpolation parameter λ∈[0,1]\lambda \in [0, 1]:

    pλ(w∣d)=(1−λ)pml(w∣d)+λp(w∣C)p_\lambda(w \mid d) = (1 - \lambda) p_{ml}(w \mid d) + \lambda p(w \mid C)

    where c(w;d)c(w; d) is the count of word ww in document dd and ∣d∣=∑wc(w;d)|d| = \sum_w c(w; d) is the document length.

    In the general smoothing framework:

    • Seen word model: ps(w∣d)=(1−λ)pml(w∣d)+λp(w∣C)p_s(w \mid d) = (1 - \lambda) p_{ml}(w \mid d) + \lambda p(w \mid C)
    • Unseen word coefficient: αd=λ\alpha_d = \lambda (constant for all documents)
    • Matched term weight: log⁡(1+1−λλpml(qi∣d)p(qi∣C))\log \left( 1 + \frac{1 - \lambda}{\lambda} \frac{p_{ml}(q_i \mid d)}{p(q_i \mid C)} \right)

    Because αd=λ\alpha_d = \lambda is identical across all documents, Jelinek-Mercer smoothing applies no document-dependent length normalization penalty term (nlog⁡λn \log \lambda is constant across documents). Smaller λ\lambda emphasizes relative term weighting, whereas as λ→1\lambda \to 1, term weights vanish and scoring reduces to coordination-level matching.

  5. Knowl 5 — Absolute Discounting Smoothing for Information Retrieval

    model/method

    Absolute discounting subtracts a fixed discount constant δ∈[0,1]\delta \in [0, 1] from the count of each seen word in a document and redistributes the accumulated probability mass proportionally to the collection language model p(w∣C)p(w \mid C):

    pδ(w∣d)=max⁡(c(w;d)−δ,0)∣d∣+δ∣d∣u∣d∣p(w∣C)p_\delta(w \mid d) = \frac{\max(c(w; d) - \delta, 0)}{|d|} + \frac{\delta |d|_u}{|d|} p(w \mid C)

    where c(w;d)c(w; d) is the term count, ∣d∣=∑wc(w;d)|d| = \sum_w c(w; d) is total document length, and ∣d∣u|d|_u is the number of unique vocabulary terms occurring in document dd.

    In the general smoothing framework:

    • Seen word model: ps(w∣d)=max⁡(c(w;d)−δ,0)∣d∣+δ∣d∣u∣d∣p(w∣C)p_s(w \mid d) = \frac{\max(c(w; d) - \delta, 0)}{|d|} + \frac{\delta |d|_u}{|d|} p(w \mid C)
    • Unseen word coefficient: αd=δ∣d∣u∣d∣\alpha_d = \frac{\delta |d|_u}{|d|}
    • Matched term weight: log⁡(1+c(qi;d)−δδ∣d∣up(qi∣C))\log \left( 1 + \frac{c(q_i; d) - \delta}{\delta |d|_u p(q_i \mid C)} \right)

    The unseen coefficient αd\alpha_d is larger for documents with a higher proportion of unique terms ∣d∣u/∣d∣|d|_u / |d|, penalizing documents with term distributions concentrated on a small set of words. Increasing δ\delta amplifies the weight differences for rare terms satisfying p(w∣C)<1/∣d∣up(w \mid C) < 1 / |d|_u, while flattening weight differences for frequent terms satisfying p(w∣C)>1/∣d∣up(w \mid C) > 1 / |d|_u.

  6. Knowl 6 — Comparative Retrieval Performance of Smoothing Methods on TREC Collections

    data/table

    Retrieval performance evaluated on five TREC document collections (FBIS, Financial Times [FT], Los Angeles Times [LA], TREC7&8 ad hoc, and TREC8 Web) using topics 351–400 (Trec7) and 401–450 (Trec8), across both concise Title queries and verbose Long queries (title + description + narrative). Parameters (λ,μ,δ)(\lambda, \mu, \delta) were individually optimized to maximize non-interpolated average precision.

    Collection Jelinek-Mercer Dirichlet Prior Absolute Discounting
    avgpr, pr@10d, pr@20d (λ\lambda) avgpr, pr@10d, pr@20d (μ\mu) avgpr, pr@10d, pr@20d (δ\delta)
    Title Queries
    fbis7T 0.172, 0.284, 0.220 (0.05) 0.197, 0.282, 0.238 (2000) 0.177, 0.284, 0.233 (0.8)
    ft7T 0.199, 0.263, 0.195 (0.5) 0.236, 0.283, 0.213 (4000) 0.215, 0.271, 0.196 (0.8)
    la7T 0.179, 0.238, 0.205 (0.4) 0.220, 0.294, 0.233 (2000) 0.194, 0.268, 0.216 (0.8)
    fbis8T 0.306, 0.344, 0.282 (0.01) 0.334, 0.367, 0.292 (500) 0.319, 0.363, 0.288 (0.5)
    ft8T 0.310, 0.359, 0.283 (0.3) 0.324, 0.367, 0.297 (800) 0.326, 0.367, 0.296 (0.7)
    la8T 0.231, 0.264, 0.211 (0.2) 0.258, 0.271, 0.216 (500) 0.238, 0.282, 0.224 (0.8)
    trec7T 0.167, 0.366, 0.315 (0.3) 0.186, 0.412, 0.342 (2000) 0.172, 0.382, 0.333 (0.7)
    trec8T 0.239, 0.438, 0.378 (0.2) 0.256, 0.448, 0.398 (800) 0.245, 0.466, 0.406 (0.6)
    web8T 0.243, 0.348, 0.293 (0.01) 0.294, 0.448, 0.374 (3000) 0.242, 0.370, 0.323 (0.7)
    Avg. 0.227, 0.323, 0.265 0.256, 0.352, 0.289 0.236, 0.339, 0.279
    Long Queries
    fbis7L 0.224, 0.339, 0.279 (0.7) 0.232, 0.313, 0.249 (5000) 0.185, 0.321, 0.259 (0.6)
    ft7L 0.279, 0.331, 0.244 (0.7) 0.281, 0.329, 0.248 (2000) 0.249, 0.317, 0.236 (0.8)
    la7L 0.264, 0.350, 0.286 (0.7) 0.265, 0.354, 0.285 (2000) 0.251, 0.340, 0.279 (0.7)
    fbis8L 0.341, 0.349, 0.283 (0.5) 0.347, 0.349, 0.290 (2000) 0.343, 0.356, 0.274 (0.7)
    ft8L 0.375, 0.427, 0.320 (0.8) 0.347, 0.380, 0.297 (2000) 0.351, 0.398, 0.309 (0.8)
    la8L 0.290, 0.296, 0.238 (0.7) 0.277, 0.282, 0.231 (500) 0.267, 0.287, 0.222 (0.6)
    trec7L 0.222, 0.476, 0.401 (0.8) 0.224, 0.456, 0.383 (3000) 0.204, 0.460, 0.396 (0.7)
    trec8L 0.265, 0.504, 0.434 (0.8) 0.260, 0.484, 0.400 (2000) 0.248, 0.518, 0.428 (0.8)
    web8L 0.259, 0.422, 0.348 (0.5) 0.275, 0.410, 0.343 (10000) 0.253, 0.414, 0.333 (0.6)
    Avg. 0.280, 0.388, 0.315 0.279, 0.373, 0.303 0.261, 0.379, 0.304

    For title queries, Dirichlet prior consistently outperforms the other methods across all collections (mean average precision 0.2560.256 vs. 0.2360.236 for absolute discounting and 0.2270.227 for Jelinek-Mercer). For long queries, Jelinek-Mercer and Dirichlet prior achieve the highest average precision (0.2800.280 and 0.2790.279, respectively), substantially outperforming absolute discounting (0.2610.261).

  7. Knowl 7 — Parameter Sensitivity and Optimal Setting Characteristics across Smoothing Methods

    empirical result

    Across TREC ad hoc collections, smoothing parameters exhibit distinct sensitivity and optimality patterns:

    • Sensitivity to Query Verbosity: Retrieval precision and recall are substantially more sensitive to parameter variations for long, verbose queries than for short title queries across all smoothing methods.
    • Jelinek-Mercer λ\lambda: The optimal interpolation parameter is heavily correlated with query type. Title queries require a very small λ≈0.1\lambda \approx 0.1 (maximizing term discrimination), whereas long verbose queries require a much larger λ≈0.7\lambda \approx 0.7 (heavily smoothing to discount background noise words).
    • Dirichlet Prior μ\mu: The optimal prior sample size μ\mu typically centers around 20002000 (with effective ranges from 500500 to 40004000) across both title and long queries on most collections. Larger μ\mu values provide stable performance for verbose queries without causing steep declines on title queries.
    • Absolute Discounting δ\delta: The optimal discount parameter δ\delta demonstrates high stability, consistently centering around δ≈0.7\delta \approx 0.7 across both query types and across all test collections.
  8. Knowl 8 — Comparison of Interpolation vs. Backoff Smoothing Strategies in Retrieval

    theoretical result

    Smoothing methods can be formulated under an interpolation strategy or a backoff strategy:

    • Interpolation formulation: Discounts seen word counts and allows extra probability mass to be shared by both seen and unseen words: ps(w∣d)=pdml(w∣d)+αdp(w∣C)p_s(w \mid d) = p_{dml}(w \mid d) + \alpha_d p(w \mid C) and pu(w∣d)=αdp(w∣C)p_u(w \mid d) = \alpha_d p(w \mid C). Matched term weights are bounded in (0,+∞)(0, +\infty).
    • Backoff formulation: Re-allocates discounted mass exclusively to unseen words: ps′(w∣d)=pdml(w∣d)p'_s(w \mid d) = p_{dml}(w \mid d) and pu′(w∣d)=αdp(w∣C)1−∑w:c(w;d)>0p(w∣C)p'_u(w \mid d) = \frac{\alpha_d p(w \mid C)}{1 - \sum_{w: c(w; d) > 0} p(w \mid C)}. The effective unseen normalization parameter αd′=αd1−∑w:c(w;d)>0p(w∣C)\alpha'_d = \frac{\alpha_d}{1 - \sum_{w: c(w; d) > 0} p(w \mid C)} penalizes long documents more severely, and matched term weights span (−∞,+∞)(-\infty, +\infty).

    Empirically, the backoff strategy performs worse than interpolation across Jelinek-Mercer, Dirichlet prior, and Absolute Discounting methods and is significantly more sensitive to smoothing parameter selection. Backoff performance approaches interpolation only in the limit as αd→0\alpha_d \to 0, where the mathematical distinction between the two strategies vanishes.

  9. Knowl 9 — Computational Scoring Complexity of Smoothed Language Models in Inverted Index Retrieval

    model/method

    When a smoothing parameter is fixed, scoring documents using a smoothed unigram query-likelihood language model can be executed with an inverted index in O(k∣q∣)O(k |q|) time, where ∣q∣|q| is the number of terms in query qq and kk is the average posting list length (number of documents containing a query term).

    Because the collection language model term ∑i=1∣q∣log⁡p(qi∣C)\sum_{i=1}^{|q|} \log p(q_i \mid C) is document-independent, it is omitted during ranking. The document-dependent factor αd\alpha_d is pre-computed and stored for each document at index time. Scoring a document requires summing matched term weights log⁡ps(qi∣d)αdp(qi∣C)\log \frac{p_s(q_i \mid d)}{\alpha_d p(q_i \mid C)} across posting lists containing query terms and adding the document length penalty nlog⁡αdn \log \alpha_d, matching the scoring efficiency of vector-space TF-IDF retrieval models.

Coverage note — No substantial contributed material was omitted.

References

  1. 1.A. Berger and J. Lafferty (1999). "Information retrieval as statistical translation," In Proceedings of the 1999 ACM SIGIR Conference on Research and Development in Information Retrieval, pp. 222–229.
  2. 2.S. F. Chen and J. Goodman (1998). "An empirical study of smoothing techniques for language modeling," Tech. Rep. TR-10-98, Harvard University.
  3. 3.N. Fuhr (1992). "Probabilistic models in information retrieval", The Computer Journal, Vol.35, No.3, pp. 243–255.
  4. 4.I. J. Good (1953). "The Population Frequencies of Species and the Estimation of Population Parameters," Biometrika, Volume 40, parts 3,4, pp. 237–264.
  5. 5.D. Hiemstra and W. Kraaij (1998). "Twenty-one at TREC-7: Ad-hoc and cross-language track," in Proc. of Seventh Text REtrieval Conference (TREC-7), Gaithersburg, MD.
  6. 6.F. Jelinek and R. Mercer (1980). "Interpolated estimation of Markov source parameters from sparse data". In Pattern Recognition in Practice, E. S. Gelsema and L. N. Kanal (editors), pages 381–402. North Holland, Amsterdam.
  7. 7.S. M. Katz (1987). "Estimation of probabilities from sparse data for the language model component of a speech recognizer," IEEE Transactions on Acoustics, Speech and Signal Processing, volume ASSP-35, pages 400–401, March 1987.
  8. 8.R. Kneser and H. Ney (1995). "Improved smoothing for m-gram language modeling," in Proceedings of the International Conference on Acoustics, Speech and Signal Processing, Detroit, MI.
  9. 9.MacKay, D. and Peto, L. (1995). "A hierarchical Dirichlet language model." Natural Language Engineering, 1(3), pp. 289–307.
  10. 10.D. H. Miller, T. Leek, and R. Schwartz (1999). "A hidden Markov model information retrieval system," In Proceedings of the 1999 ACM SIGIR Conference on Research and Development in Information Retrieval, pp. 214–221.
  11. 11.H. Ney, U. Essen, and R. Kneser (1994). "On structuring probabilistic dependencies in stochastic language modeling," Computer Speech and Language, 8:1-38.
  12. 12.J. Ponte (1998). A language modeling approach to information retrieval. Ph.D. thesis, University of Massachusetts at Amherst.
  13. 13.J. Ponte and W. B. Croft (1998). "A language modeling approach to information retrieval," Proceedings of the ACM SIGIR, pp. 275–281.
  14. 14.C. J. van Rijsbergen (1986). "A Non-classical Logic for Information Retrieval," The Computer Journal, 29(6).
  15. 15.S. E. Robertson, C. J. van-Rijsbergen, and M. F. Porter (1981). "Probabilistic models of indexing and searching", in Oddy R. N. et al. (Eds.)Information Retrieval Research, Butterworths, London, 1981, pp. 35–56.
  16. 16.S. E. Robertson, S. Walker, S. Jones, M. M. Hancock-Beaulieu, and M. Gatford (1995). "Okapi at TREC-3," The Third Text REtrieval Conference (TREC-3), in D. K. Harman (ed), NIST Special Publication.
  17. 17.G. Salton and C. Buckley (1988). "Term-weighting approaches in automatic text retrieval," Information Processing and Management, 24, pp. 513–523.
  18. 18.G. Salton and C. Buckley (1990), "Improving retrieval performance by relevance feedback", Journal of the American Society for Information Science, Vol. 44, No. 4, 288–297.
  19. 19.A. Singhal, C. Buckley, and M. Mitra (1996). "Pivoted document length normalization," in Proceedings of the 1996 ACM SIGIR Conference on Research and Development in Information Retrieval, pp. 21–29.
  20. 20.F. Song and B. Croft (1999). "A general language model for information retrieval," in Proceedings of the 1999 ACM SIGIR Conference on Research and Development in Information Retrieval, pp. 279–280.
  21. 21.K. Sparck Jones (1997). Readings in Information Retrieval, P. Willett, ed., Morgan Kaufmann Publishers.
  22. 22.S. K. M. Wong and Y. Y. Yao (1995), "On modeling information retrieval with probabilistic inference," ACM Transactions on Information Systems, 13(1), pp. 69–99.

Citation

MLA
Zhai, C., and J. Lafferty. “A Study of Smoothing Methods for Language Models Applied to Ad Hoc Information Retrieval”. Proceedings of the 24th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, 2001, pp. 334–42, https://doi.org/10.1145/383952.384019.
APA
Zhai, C., & Lafferty, J. (2001). A study of smoothing methods for language models applied to Ad Hoc information retrieval. Proceedings of the 24th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, 334–342. https://doi.org/10.1145/383952.384019
Chicago
Zhai, C., and J. Lafferty. 2001. “A Study of Smoothing Methods for Language Models Applied to Ad Hoc Information Retrieval”. Proceedings of the 24th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, 334–42. https://doi.org/10.1145/383952.384019.
Harvard
Zhai, C. and Lafferty, J. (2001) “A study of smoothing methods for language models applied to Ad Hoc information retrieval”, Proceedings of the 24th annual international ACM SIGIR conference on Research and development in information retrieval. ACM, pp. 334–342. Available at: https://doi.org/10.1145/383952.384019.
Vancouver
1. Zhai C, Lafferty J (2001) A study of smoothing methods for language models applied to Ad Hoc information retrieval. In: Proceedings of the 24th annual international ACM SIGIR conference on Research and development in information retrieval. ACM, pp 334–342

BibTeX

@inproceedings{Zhai_2001, series={SIGIR01}, title={A study of smoothing methods for language models applied to Ad Hoc information retrieval}, url={http://dx.doi.org/10.1145/383952.384019}, DOI={10.1145/383952.384019}, booktitle={Proceedings of the 24th annual international ACM SIGIR conference on Research and development in information retrieval}, publisher={ACM}, author={Zhai, Chengxiang and Lafferty, John}, year={2001}, month=Sept, pages={334–342}, collection={SIGIR01} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF