PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 17, 20240 citationsOpen Access

ManipVQA: Injecting Robotic Affordance and Physically Grounded Information into Multi-Modal Large Language Models

View Full Paper
SHSiyuan HuangIPIaroslav PonomarenkoZJZhengkai Jiang

Key Points

Key points are not available for this paper at this time.

Abstract

The integration of Multimodal Large Language Models (MLLMs) with robotic systems has significantly enhanced the ability of robots to interpret and act upon natural language instructions. Despite these advancements, conventional MLLMs are typically trained on generic image-text pairs, lacking essential robotics knowledge such as affordances and physical knowledge, which hampers their efficacy in manipulation tasks. To bridge this gap, we introduce ManipVQA, a novel framework designed to endow MLLMs with Manipulation-centric knowledge through a Visual Question-Answering format. This approach not only encompasses tool detection and affordance recognition but also extends to a comprehensive understanding of physical concepts. Our approach starts with collecting a varied set of images displaying interactive objects, which presents a broad range of challenges in tool object detection, affordance, and physical concept predictions. To seamlessly integrate this robotic-specific knowledge with the inherent vision-reasoning capabilities of MLLMs, we adopt a unified VQA format and devise a fine-tuning strategy that preserves the original vision-reasoning abilities while incorporating the new robotic insights. Empirical evaluations conducted in robotic simulators and across various vision task benchmarks demonstrate the robust performance of ManipVQA. Code and dataset will be made publicly available at https://github.com/SiyuanHuang95/ManipVQA.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Huang et al. (2024) studied this question.

synapsesocial.com/papers/68e73a8db6db6435876b496bhttps://doi.org/10.48550/arxiv.2403.11289
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets2025
  2. 2ManipLVM-R1: Reinforcement Learning for Reasoning in Embodied Manipulation with Large Vision-Language Models2025
  3. 3MOKA: Open-World Robotic Manipulation through Mark-Based Visual Prompting2024 · 5 citations
  4. 4Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces2025
  5. 5NaturalVLM: Leveraging Fine-grained Natural Language for Affordance-Guided Visual Manipulation2024 · 1 citations