PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 6, 20240 citationsOpen Access

Joint Visual and Text Prompting for Improved Object-Centric Perception with Multimodal Large Language Models

View Full Paper
SJSongtao JiangYZYan ZhangCZChenyi Zhou

Key Points

Key points are not available for this paper at this time.

Abstract

Multimodal Large Language Models (MLLMs) such as GPT-4V and Gemini Pro face challenges in achieving human-level perception in Visual Question Answering (VQA), particularly in object-oriented perception tasks which demand fine-grained understanding of object identities, locations or attributes, as indicated by empirical findings. This is mainly due to their limited capability to effectively integrate complex visual cues with textual information and potential object hallucinations. In this paper, we present a novel approach, Joint Visual and Text Prompting (VTPrompt), that employs fine-grained visual information to enhance the capability of MLLMs in VQA, especially for object-oriented perception. VTPrompt merges visual and text prompts to extract key concepts from textual questions and employs a detection model to highlight relevant objects as visual prompts in images. The processed images alongside text prompts are subsequently fed into MLLMs to produce more accurate answers. Our experiments with GPT-4V and Gemini Pro, on three benchmarks, i.e., MME , MMB and POPE, demonstrate significant improvements. Particularly, our method led to a score improvement of up to 183.5 for GPT-4V on MME and enhanced MMB performance by 8.17\% for GPT-4V and 15.69\% for Gemini Pro.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Jiang et al. (2024) studied this question.

synapsesocial.com/papers/68e7031db6db64358767ceeahttps://doi.org/10.48550/arxiv.2404.04514
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Enhancing Multimodal Large Language Models with Multi-instance Visual Prompt Generator for Visual Representation Enrichment2024
  2. 2Rethinking Visual Prompting for Multimodal Large Language Models with External Knowledge2024 · 1 citations
  3. 3Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want2024 · 1 citations
  4. 4Mini-Gemini: Mining the Potential of Multi-Modality Vision Language Models2025 · 32 citations
  5. 5Exploring the Capabilities of Large Multimodal Models on Dense Text2024