PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 18, 20245 citationsOpen Access

Prompting Large Language Models with Fine-Grained Visual Relations from Scene Graph for Visual Question Answering

View Full Paper
JLJiapeng LiuXinjiang Normal UniversityCFChengyang FangMultimedia UniversityLLLiang LiNingbo University

Key Points

Key points are not available for this paper at this time.

Abstract

Visual Question Answering (VQA) is a task that requires models to comprehend both questions and images. An increasing number of works are leveraging the strong reasoning capabilities of Large Language Models (LLMs) to address VQA. These methods typically utilize image captions as visual text description to aid LLMs in comprehending images. However, these captions often overlooking the relations of fine-grained objects, which will limit the reasoning capability of LLMs. In this paper, we present PFVR, a modular framework that Prompts LLMs with Fine-grained Visual Relationships for VQA. PFVR primarily consists of an answer-guided generation module (AGG) and a question-guided filtering module (QGF). The two modules can combine to extract the fine-grained visual relations from scene graph, which will finally serve as crucial context for LLMs to comprehend the image. Extensive experiments conducted on the popular VQA dataset, GQA, confirm PFVR achieves state-of-the-art results compared to other strong VQA competitors, demonstrating its exceptional effectiveness.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Liu et al. (2024) studied this question.

synapsesocial.com/papers/68e7397eb6db6435876b2a28https://doi.org/10.1109/icassp48485.2024.10448321
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Prompting Large Language Models with Answer Heuristics for Knowledge-Based Visual Question Answering2023 · 214 citations
  2. 2ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks2019 · 1,677 citations
  3. 3Training language models to follow instructions with human feedback2022 · 4,346 citations
  4. 4REVIVE: Regional Visual Representation Matters in Knowledge-Based Visual Question Answering2022 · 44 citations
  5. 5Declaration-based Prompt Tuning for Visual Question Answering2022 · 23 citations