Composed Video Retrieval (CVR) is a novel video retrieval paradigm. Unlike traditional single-modal video retrieval paradigms ( e.g. , text to video or video to video), CVR employs multi-modal queries (including both a reference video and a natural language modification) to retrieve the target video that best matches the modified reference video. Existing CVR methods primarily rely on generalized knowledge from vision-language pre-trained models or utilize caption expansions to enhance video comprehension. However, these approaches overlook the benefits offered by the shareability and variability of videos for multi-modal query understanding. To overcome this limitation, we introduce a novel CVR framework named shaREd and diFferential semantIcs eNhancement nEtwork ( REFINE ). REFINE is the first framework to exploit the shareability and variability of videos to improve multi-modal query comprehension. Specifically, REFINE leverages learnable tokens to achieve enhanced shared feature representation. Moreover, it introduces a carefully designed Differential Block to disentangle differential semantics between frames and employs modification associations to guide multi-modal query feature fusion. Additionally, REFINE has been extended to the Composed Image Retrieval task, making it effectively generalize across existing composed multi-modal retrieval scenarios and outperform existing methods. Extensive qualitative and quantitative evaluations on four benchmark datasets validate the superiority of the proposed REFINE framework.
Hu et al. (Sat,) studied this question.