PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 27, 2026Discover Artificial Intelligence0 citationsOpen Access

Real-time highlight clip recognition and storage for live sports events based on YOLOv11

DHDu HuanranYFYuming FengJWJiaxin Wu

Key Points

  • This research aims to develop a system for real-time recognition and storage of sports highlights.
  • Proposed an Inner-CIOU loss function to enhance small-object detection accuracy in YOLOv11.
  • Developed an audio-visual dual-modal mechanism using GMM-HMM for speech-to-text conversion.
  • Implemented a redundancy optimization and hierarchical storage architecture for efficient data management.
  • Classification accuracy for exciting events exceeds 85%.
  • Outperformed CNN-LSTM and Vision Transformer models in performance.
  • Achieved significantly faster training convergence and reduced overfitting.

Abstract

With the rapid development of the sports live‑streaming industry, the demand for real‑time extraction and instant replay of exciting moments has become increasingly urgent. However, the traditional manual editing method is time‑consuming, labor‑intensive, and lacks real‑time performance. Existing automated recognition schemes suffer from technical drawbacks such as single‑modal limitations, low accuracy in small‑object detection, and an imbalance between storage efficiency and access performance, making it difficult to meet the practical application requirements of sports live broadcasting. To address these issues, this paper proposes a multi‑modal collaborative optimized system for recognizing and storing exciting moments in sports live streams based on YOLOv11. Firstly, to tackle the insufficient accuracy of small‑object recognition and bounding‑box regression bias in YOLOv11, an Inner‑CIOU loss function incorporating a scale factor Sratio is proposed to enhance the model’s ability to capture balls and local player movements. Secondly, an audio‑visual dual‑modal collaborative decision mechanism is constructed. The GMM‑HMM model is used to realize real‑time speech‑to‑text conversion of match commentary, and a keyword library of exciting events is applied to coarsely locate potential highlight time windows. The optimized YOLOv11 is then employed to extract visual and temporal features, and a weighted fusion strategy with a visual confidence weight of 0.7 and an audio matching weight of 0.3 is adopted to achieve accurate decision‑making. Finally, a redundancy optimization and hierarchical storage architecture is designed. Redundant frames are filtered via the inter‑frame IOU threshold, and the balance between storage efficiency and access performance is achieved by storing hot data on high‑speed SSDs and cold data in distributed storage with H.265 compression. Experimental results demonstrate that the proposed system achieves a significantly faster training convergence speed and effectively alleviates overfitting. The classification accuracy for all types of exciting events exceeds 85%, and the overall performance outperforms comparative models including CNN‑LSTM and Vision Transformer. This study not only expands the application scenarios of real‑time object detection in the sports domain but also provides an automated and intelligent tool for processing exciting moments on sports live‑streaming platforms, offering important technical support for the implementation of downstream industries such as smart sports and match analysis.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Huanran et al. (2026) studied this question.

synapsesocial.com/papers/69eefd43fede9185760d4015https://doi.org/10.1007/s44163-026-01268-2
Ask AI
Helpful
Bookmark
Share
View Full Paper