PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 20, 2026IEEE Transactions on Image Processing0 citations

Vision-Language Collaborative Representation Learning for Action Quality Assessment

View Full Paper
KGKumie GedamuYJYanli JiSun Yat-sen UniversityWZWangmeng Zuo

Key Points

  • The aim is to enhance action quality assessment by integrating vision and language features in a unified representation.
  • Developed a Vision-Language Collaboration Representation Learning framework (VLC-Net).
  • Implemented bidirectional knowledge distillation for collaboration learning between modalities.
  • Designed alignment guidance to unify action features across vision and language.
  • Utilized multimodal contrastive learning for aligning subactions with textual descriptions.
  • Demonstrated superior performance compared to state-of-the-art methods.
  • Achieved improved accuracy in predicting AQA scores.
  • Successfully aligned action semantics across video and text modalities.

Abstract

Action Quality Assessment (AQA) has gained significant attention due to its potential real-world applications, which require a fine-grained understanding of action sequences. Recent works have attempted to utilize multimodal video features and address some existing challenges. However, these approaches primarily focus on leveraging textual information from language models only, leading to instability and suboptimal performance due to directional bias in a vision-language joint embedding space. To tackle these issues, we propose a Vision-Language Collaboration Representation Learning approach (VLC-Net) to understand fine-grained action sequences and create a unified feature representation along with their temporal dependencies for accurate AQA score prediction. Specifically, we design a bidirectional knowledge distillation operation to perform collaboration learning between vision-language pre-trained knowledge and visual action knowledge for fine-grained action feature learning. Furthermore, we design vision-language alignment guidance to explicitly align action features with the same action semantics across modalities, thereby unifying their joint representation. Leveraging these aligned features, we propose multimodal contrastive learning to relate modalities and align subactions with textual descriptions, ensuring accurate action representation. We conduct experiments on the FineDiving, MTL-AQA, FineFS, and Fis-V datasets, demonstrating the effectiveness of our approach, which outperforms state-of-the-art methods.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Gedamu et al. (2026) studied this question.

synapsesocial.com/papers/69e5c22d03c29399140288fchttps://doi.org/10.1109/tip.2026.3683256
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Vision-Language Action Knowledge Learning for Semantic-Aware Action Quality Assessment2024 · 15 citations
  2. 2Localization-assisted Uncertainty Score Disentanglement Network for Action Quality Assessment2023 · 25 citations
  3. 3Assessing the Quality of Actions2014 · 252 citations
  4. 4HiGCIN: Hierarchical Graph-Based Cross Inference Network for Group Activity Recognition2020 · 222 citations
  5. 5Assessing the aesthetic quality of photographs using generic image descriptors2011 · 413 citations