Long-term sports assessment is a challenging task in video understanding, since it requires judging subtle movement variations over minutes and evaluating action–music coordination. However, in many sporting events the background music is only weakly related to the performed movements, and the cues that matter for synchrony are often temporal and structural, such as small phase or tempo deviations that occur around decisive moments, rather than semantic correspondences between audio content and action categories. Prior approaches typically rely on implicit cross-modal fusion over dense sequences to learn such weak associations, which can smooth out near-miss misalignment and become brittle under tempo or phase shifts. To address this issue, we propose BEATSCORE, a beat-guided audio–visual learning framework that explicitly models action–music alignment at the beat level and performs event-centric sparse grading for long videos. In our framework, we first convert audio and motion into beat-synchronous tokens, enabling direct comparison on a unified rhythmic timeline. We then introduce a beat-level contrastive objective with near-offset hard negatives to sharpen sensitivity to misalignment. To handle the sparsity of decisive moments, we further design an event proposal and grading module that scores a small set of key segments and aggregates them via learnable multiple-instance pooling into a final assessment score. We evaluate BEATSCORE on public long-term sports benchmarks to demonstrate improved accuracy with competitive efficiency.
Wang et al. (Tue,) studied this question.