Well-calibrated predictions of user preferences are essential for many applications. For example, displaying the predicted values alongside recommended items can support user decisions on whether to consume the items. Since recommender systems typically select the top-N items for users, it is crucial to ensure calibration for those top-N items, rather than for all items. In the aforementioned scenario, users only examine the top-N items and their predictions. In this work, we first demonstrate that previous calibration methods result in miscalibrated predictions for the top-N items, despite exhibiting excellent calibration performance when evaluated on all items. We address the issue of miscalibration in the top-N recommended items by first defining evaluation metrics tailored to this objective. Subsequently, we propose a generic method to optimize calibration models with a focus on the top-N items. This method groups the top-N items by their ranks and optimizes distinct calibration models for each group using rank-dependent training weights. We verify the effectiveness of the proposed method for both explicit and implicit feedback datasets, utilizing diverse classes of recommender models.
Masahiro Sato (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: