PulseExploreJournal ClubResearchersJournals
Instagram
HomeJournal ClubExplore
Synapse
⌘+K
Synapse
April 30, 2026ElectronicsOpen Access

Bridging Cross-Modal Semantic Gaps with Multi-Source Semantic Anchors in Knowledge-Based Visual Question Answering

View Full Paper
Ask AI
Bookmark
Share

Authors

HJHu JunMingJZJinxiong ZhangFZFeng Zhan

Discussion

Loading...

Member takes

Overview

Proposed framework enhances reasoning in KB-VQA by bridging semantic gaps, integrating multiple sources of information.

Key Points

  • The aim is to improve knowledge-based visual question answering by addressing cross-modal semantic gaps.
  • Developed a unified framework using multi-source semantic anchors for aligning vision and textual features.
  • Implemented a cross-residual gating mechanism to reduce modality noise during cross-modal fusion.
  • Adopted contrastive learning to enhance cross-modal alignment and used a retrieve-then-read pipeline for answer generation.
  • Framework demonstrates improved performance on OK-VQA, FVQA, and A-OKVQA datasets.
  • Outperforms state-of-the-art methods across multiple evaluation metrics.

Cite This Study

JunMing et al. (2026) studied this question.

synapsesocial.com/papers/69f2a42a8c0f03fd67763206https://doi.org/10.3390/electronics15091837
View Full Paper
Ask AI
Bookmark
Share