PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 27, 2026Journal of King Saud University - Computer and Information Sciences0 citationsOpen Access

Video feature guidance for temporal action segmentation

YLYue LiZFZixun FuRGRui Gao

Key Points

  • The aim is to enhance temporal action segmentation in long-form videos by reducing uncertainty in predictions.
  • Propose a video feature-guided temporal action segmentation method (VFG) using semantically rich features.
  • Construct feature-guided latent representation by combining encoded video features with ground-truth guidance.
  • Design a decoding architecture to enhance sensitivity to local details and model long-range contextual dependencies.
  • VFG shows improved accuracy in action boundary prediction compared to prior methods.
  • Achieves competitive computational efficiency across three benchmark datasets.

Abstract

High-level semantic understanding of long-form videos critically depends on accurate temporal action segmentation. Although generative diffusion models have recently introduced a new modeling paradigm for this task, most existing methods initialize the diffusion process with random gaussian noise, which induces substantial uncertainty in the generated outputs and makes it difficult to ensure stable and reproducible predictions. To address this limitation, we propose a video feature-guided temporal action segmentation method (VFG), which uses semantically rich video features as a deterministic prior for the generation process to guide the model toward learning task-relevant discriminative representations. Specifically, instead of relying on purely random gaussian initialization as in prior diffusion-based methods, we construct a feature-guided latent representation during training by combining encoded video features with a scheduled amount of ground-truth guidance. This provides stronger task-aligned supervision and improves the accuracy and consistency of action boundary prediction. In addition, we design a dedicated decoding architecture that enhances sensitivity to local details while also modeling long-range contextual dependencies. Extensive experiments on three benchmark datasets show that VFG is competitive in accuracy and computational efficiency.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Li et al. (2026) studied this question.

synapsesocial.com/papers/69eefdd1fede9185760d49e4https://doi.org/10.1007/s44443-026-00734-2
Ask AI
Helpful
Bookmark
Share
View Full Paper