PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 20, 20250 citationsOpen Access

MedGround-R1: Advancing Medical Image Grounding via Spatial-Semantic Rewarded Group Relative Policy Optimization

View Full Paper
HXHuihui XuYNYuanpeng NieHWHualiang Wang

Key Points

  • State-of-the-art performance in medical image grounding was achieved without requiring costly reasoning annotations.
  • Our method utilized spatial-semantic rewards to effectively balance spatial accuracy and semantic consistency.
  • Experiments on multiple datasets demonstrated the effectiveness of the proposed spatial-semantic rewarded optimization.
  • We integrated visual information from bounding boxes into the reasoning process to improve model performance.

Abstract

Medical Image Grounding (MIG), which involves localizing specific regions in medical images based on textual descriptions, requires models to not only perceive regions but also deduce spatial relationships of these regions. Existing Vision-Language Models (VLMs) for MIG often rely on Supervised Fine-Tuning (SFT) with large amounts of Chain-of-Thought (CoT) reasoning annotations, which are expensive and time-consuming to acquire. Recently, DeepSeek-R1 demonstrated that Large Language Models (LLMs) can acquire reasoning abilities through Group Relative Policy Optimization (GRPO) without requiring CoT annotations. In this paper, we adapt the GRPO reinforcement learning framework to VLMs for Medical Image Grounding. We propose the Spatial-Semantic Rewarded Group Relative Policy Optimization to train the model without CoT reasoning annotations. Specifically, we introduce Spatial-Semantic Rewards, which combine spatial accuracy reward and semantic consistency reward to provide nuanced feedback for both spatially positive and negative completions. Additionally, we propose to use the Chain-of-Box template, which integrates visual information of referring bounding boxes into the reasoning process, enabling the model to explicitly reason about spatial regions during intermediate steps. Experiments on three datasets MS-CXR, ChestX-ray8, and M3D-RefSeg demonstrate that our method achieves state-of-the-art performance in Medical Image Grounding. Ablation studies further validate the effectiveness of each component in our approach. Code, checkpoints, and datasets are available at https://github.com/bio-mlhui/MedGround-R1

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Xu et al. (2025) studied this question.

synapsesocial.com/papers/68f5fcdc8d54a28a75cf23d6https://doi.org/10.48550/arxiv.2507.02994
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Med-R1: Reinforcement Learning for Generalizable Medical Reasoning in Vision-Language Models2026 · 14 citations
  2. 2MedGR$^2$: Breaking the Data Barrier for Medical Reasoning via Generative Reward Learning2025
  3. 3MRG-R1: Reinforcement Learning for Clinically Aligned Medical Report Generation2025
  4. 4RadVLM-GRPO : enhancing chest X-ray report generation and visual grounding via reinforcement learning2026
  5. 5Seeing the Trees for the Forest: Rethinking Weakly-Supervised Medical Visual Grounding2025