Key points are not available for this paper at this time.
User decision-making behavior in recommender systems is jointly driven by a large number of underlying factors. Learning and revealing the representations of these latent factors can provide more robustness and interpretability. However, mining the latent intentions of user decision-making behaviors in existing multimodal recommendation studies faces the following two key challenges: i) Modal noise pollution: in multimodal user intent modeling, inputs from individual modalities are inevitably corrupted by noise of varying severity. During message-passing, a large proportion of irrelevant or even contradictory signals are propagated and injected into item representations, which impedes the model's ability to achieve pure semantic alignment at the content level. ii) User Intent Confounding: real-world items naturally possess multiple attributes, with different attributes of the same item influencing distinct potential user intentions. However, in existing modelling designs for user intent, such intents are mapped onto user-item interaction labels of the same coarse granularity. This many-to-one mapping relationship between intentions and items leads to significant confounding and dilution of users’ fine-grained intentions. To address the above challenges, this work pays special attention to the implied user intent behind pure multimodal features. Specifically, we construct a dynamic adaptive multimodal intent disentanglement model(MIDM). This model adopts a non-ID paradigm and mines the distribution of user intents in multimodal scenarios directly from users’ decision-making behaviors. A comprehensive experimental study on the amazon dataset shows that the method is effective and provides a novel learning scheme for mining user intent in multimodal scenarios.
Wang et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: