PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 17, 2026ACM Transactions on Multimedia Computing Communications and Applications0 citations

Semantic Prototype Guided Sparse Temporal Interaction for Weakly Supervised Temporal Action Localization

View Full Paper
YWYì WángBGI Group (China)DKDehui KongBeijing University of TechnologyJLJinghua Li

Key Points

  • The research aims to enhance weakly supervised temporal action localization by improving action completeness and accuracy.
  • Utilized semantic prototypes to enrich video representations.
  • Employed a prototype contrastive loss for better feature discriminability.
  • Designed a sparse temporal interaction unit to model short-term context and long-range dependencies.
  • Applied a boundary-guided loss for precise action boundaries.
  • S2Net achieved more accurate action localization compared to previous methods.
  • Demonstrated improved completeness of action segment detection on THUMOS14 and ActivityNet1.3.

Abstract

Weakly supervised temporal action localization (WTAL) aims to detect action segments in untrimmed videos only using video-level labels. Existing methods typically follow the multi-instance learning (MIL) paradigm with a top-k strategy, often resulting in incomplete action localization. Moreover, the local and discontinuous nature of actions causes action segments to be isolated and lack sufficient temporal interaction. To address these issues, this paper introduces semantic prototypes to enrich video representations, enabling the model to aggregate category-level action cues across videos and recover semantically relevant but weakly activated segments, thereby improving action completeness. A prototype contrastive loss is further employed to improve feature discriminability. Moreover, a sparse temporal interaction unit is designed to jointly model short-term context and long-range dependencies. The boundary-guided loss utilizes the temporal interaction outputs to explicitly constrain semantic responses around action boundaries, promoting sharp and temporally consistent transitions. Based on these, this paper proposes a semantic prototype guided sparse temporal interaction network (S2Net), achieving a unified video modeling from full semantic understanding to fine-grained boundary perception. Extensive experiments on THUMOS14 and ActivityNet1.3 demonstrate that S2Net achieves more accurate and complete action localization.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Wáng et al. (2026) studied this question.

synapsesocial.com/papers/69e1cffa5cdc762e9d8590fchttps://doi.org/10.1145/3807956
Ask AI
Helpful
Bookmark
Share
View Full Paper