March 29, 2024Open Access

Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want

Key Points

Key points are not available for this paper at this time.

Abstract

The interaction between humans and artificial intelligence (AI) is a crucial factor that reflects the effectiveness of multimodal large language models (MLLMs). However, current MLLMs primarily focus on image-level comprehension and limit interaction to textual instructions, thereby constraining their flexibility in usage and depth of response. In this paper, we introduce the Draw-and-Understand project: a new model, a multi-domain dataset, and a challenging benchmark for visual prompting. Specifically, we propose SPHINX-V, a new end-to-end trained Multimodal Large Language Model (MLLM) that connects a vision encoder, a visual prompt encoder and an LLM for various visual prompts (points, bounding boxes, and free-form shape) and language understanding. To advance visual prompting research for MLLMs, we introduce MDVP-Data and MDVP-Bench. MDVP-Data features a multi-domain dataset containing 1.6M unique image-visual prompt-text instruction-following samples, including natural images, document images, OCR images, mobile screenshots, web screenshots, and multi-panel images. Furthermore, we present MDVP-Bench, a comprehensive and challenging benchmark to assess a model's capability in understanding visual prompting instructions. Our experiments demonstrate SPHINX-V's impressive multimodal interaction capabilities through visual prompting, revealing significant improvements in detailed pixel-level description and question-answering abilities.

Connected Papers

Building similarity graph...

Analyzing shared references across papers

Discussion

Cite this study

Lin et al. (Fri,) studied this question.

www.synapsesocial.com/papers/68e71cc2b6db6435876969a0 — DOI: https://doi.org/10.48550/arxiv.2403.20271

Also consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

Authors

Weifeng Lin

Xinyu Wei

Ruichuan An

Actions

References and Citations

Connected Papers

Building similarity graph...

Analyzing shared references across papers

Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want

Key Points

Abstract

Citation Network

Connected Papers

Discussion

Cite this study

Also consider

Authors

Actions

References and Citations

Citation Network

Connected Papers

Discussion