PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 5, 2026Scientific Reports0 citationsOpen Access

MoSA-Det: motion state adaptive object detection for sports videos

LYLulu YangWSWei SunJRJinkui Ren

Key Points

  • The aim is to enhance object detection in sports videos by overcoming limitations caused by motion blur and frame aggregation.
  • Developed MoSA-Det framework utilizing motion states for feature extraction and temporal fusion.
  • Implemented Motion-Aware Adaptive Feature Module (MAAF) for feature degradation mitigation.
  • Constructed State-Guided Temporal Aggregation Module (SGTA) for efficient multi-frame feature aggregation.
  • MoSA-Det improved mean Average Precision (mAP) at thresholds of 0.5 and 0.75 by 1.7% and 2.6%, respectively.
  • Achieved significant enhancement in detection accuracy over strong baseline methods across two datasets.

Abstract

Object detection in sports videos serves as a fundamental task for applications such as intelligent broadcasting, tactical analysis, and athlete tracking. Existing methods face two critical challenges when processing sports scenarios: (1) Motion blur-induced feature degradation, where fast-moving objects generate severe motion blur that leads to indistinct boundaries, texture loss, and significantly reduced feature response intensity, severely weakening the discriminative capability of detectors; (2) Temporal aggregation failure, where existing temporal methods assume small inter-frame displacements that enable effective alignment, yet fast-moving objects often exhibit excessive inter-frame displacement that causes alignment failure, making multi-frame aggregation introduce noise and degrade detection accuracy. To address these challenges, we propose MoSA-Det, a framework whose core idea is to leverage motion states as regulatory signals for detection strategies, achieving joint adaptive optimization of feature extraction and temporal fusion. The framework comprises two core modules: the Motion-Aware Adaptive Feature Module (MAAF) and the State-Guided Temporal Aggregation Module (SGTA). Specifically, MAAF constructs a lightweight motion state estimator through inter-frame feature differencing and local correlation analysis to generate fine-grained motion state priors, and employs a state-conditioned dynamic convolution mechanism to perform weighted fusion of multi-scale receptive field features, introducing deformable convolution for adaptive spatial sampling in high-speed motion regions to effectively alleviate feature degradation caused by motion blur. SGTA embeds motion state priors into the temporal feature fusion process, achieving selective aggregation of inter-frame features through state-aware adaptive weights, ensuring that static regions fully exploit multi-frame information to enhance feature stability while enabling fast-moving regions to avoid noise interference from misaligned frames. Experiments on the SoccerNet-Tracking and SportsMOT datasets demonstrate that MoSA-Det improves mAP@0.5 by 1.7% and 1.6%, and mAP@0.75 by 2.6% and 1.8% over the strongest baselines, respectively, validating the effectiveness of the motion state adaptive strategy.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Yang et al. (2026) studied this question.

synapsesocial.com/papers/69d1fde4a79560c99a0a4489https://doi.org/10.1038/s41598-026-43231-2
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1C2T-Net: Channel-Aware Cross-Fused Transformer-Style Networks for Pedestrian Attribute Recognition2024 · 41 citations
  2. 2Use of deep learning in soccer videos analysis: survey2022 · 47 citations
  3. 3Deep soccer analytics: learning an action-value function for evaluating soccer players2020 · 105 citations
  4. 4FoT: an efficient transformer framework for real-time small object detection in football videos2025 · 4 citations
  5. 5Hybrid multi-attention transformer for robust video object detection2024 · 38 citations