Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
May 20, 2026Discover ComputingOpen Access

Common TF–IDF variants arise as key components in the test statistic of a penalized likelihood-ratio test for word burstiness

View Full Paper
Ask AI
Bookmark
Share

Authors

ZAZeyad AhmedPSPaul SheridanMMMichael McIsaac

Discussion

Loading...

Member takes

Overview

Randomized trial explores TF-IDF-like scores for document classification, suggesting new insights for term-weighting schemes.

Key Points

  • The aim is to demonstrate how TF-IDF-like scores emerge in the context of a penalized likelihood-ratio test for modeling word burstiness.
  • Utilized a penalized likelihood-ratio test capturing word burstiness through beta-binomial distributions with a gamma penalty.
  • Compared term-weighting schemes derived from the test statistic against standard TF-IDF on document classification tasks.
  • The test statistic-based term-weighting scheme performs comparably to TF-IDF for document classification tasks.
  • Insights provided on the statistical foundations of TF-IDF enhance understanding of term-weighting development.

Cite This Study

Ahmed et al. (2026) studied this question.

synapsesocial.com/papers/6a0d5100f03e14405aa9d467https://doi.org/10.1007/s10791-026-10090-4
View Full Paper
Ask AI
Bookmark
Share