PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 1, 2004ACM Transactions on Information Systems1,216 citations

A study of smoothing methods for language models applied to information retrieval

View Full Paper
CZChengXiang ZhaiJLJohn Lafferty

Key Points

  • The central aim is to explore the effect of smoothing methods on language models in information retrieval and how they impact retrieval performance.
  • Examined several popular smoothing methods across different test collections.
  • Analyzed the sensitivity of retrieval performance to smoothing parameters based on query types.
  • Proposed a two-stage smoothing strategy with automated parameter estimation.
  • Retrieval performance was found to be sensitive to smoothing parameters, especially for verbose queries.
  • Verbose queries required more aggressive smoothing to optimize performance.
  • The two-stage smoothing method outperformed traditional single smoothing methods in retrieval accuracy.

Abstract

Language modeling approaches to information retrieval are attractive and promising because they connect the problem of retrieval with that of language model estimation, which has been studied extensively in other application areas such as speech recognition. The basic idea of these approaches is to estimate a language model for each document, and to then rank documents by the likelihood of the query according to the estimated language model. A central issue in language model estimation is smoothing , the problem of adjusting the maximum likelihood estimator to compensate for data sparseness. In this article, we study the problem of language model smoothing and its influence on retrieval performance. We examine the sensitivity of retrieval performance to the smoothing parameters and compare several popular smoothing methods on different test collections. Experimental results show that not only is the retrieval performance generally sensitive to the smoothing parameters, but also the sensitivity pattern is affected by the query type, with performance being more sensitive to smoothing for verbose queries than for keyword queries. Verbose queries also generally require more aggressive smoothing to achieve optimal performance. This suggests that smoothing plays two different role---to make the estimated document language model more accurate and to "explain" the noninformative words in the query. In order to decouple these two distinct roles of smoothing, we propose a two-stage smoothing strategy, which yields better sensitivity patterns and facilitates the setting of smoothing parameters automatically. We further propose methods for estimating the smoothing parameters automatically. Evaluation on five different databases and four types of queries indicates that the two-stage smoothing method with the proposed parameter estimation methods consistently gives retrieval performance that is close to---or better than---the best results achieved using a single smoothing method and exhaustive parameter search on the test data.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhai et al. (2004) studied this question.

synapsesocial.com/papers/69da2a940d540cafc5838bf0https://doi.org/10.1145/984321.984322
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1THE POPULATION FREQUENCIES OF SPECIES AND THE ESTIMATION OF POPULATION PARAMETERS1953 · 3,334 citations
  2. 2Model-based feedback in the language modeling approach to information retrieval2001 · 804 citations
  3. 3A Language Modeling Approach to Information Retrieval2017 · 2,544 citations
  4. 4Probabilistic models of indexing and searching1980 · 316 citations
  5. 5Interpolated estimation of Markov source parameters from sparse data1980 · 824 citations