PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
December 21, 20250 citationsOpen Access

ReinboT: Amplifying Robot Visual-Language Manipulation with Reinforcement Learning

View Full Paper
HZHongyin ZhangZZZifeng ZhuangHZHan Zhao

Key Points

  • This research aims to enhance robot decision-making tasks through reinforcement learning and visual-language manipulation.
  • Introduced ReinboT, an end-to-end Vision-Language-Action model
  • Integrated reinforcement learning to maximize cumulative rewards
  • Conducted extensive experiments on the CALVIN mixed-quality dataset
  • ReinboT shows state-of-the-art performance on the CALVIN dataset
  • Demonstrates superior few-shot learning capabilities
  • Exhibits strong out-of-distribution generalization in real-world tasks

Abstract

Vision-Language-Action (VLA) models have shown great potential in general robotic decision-making tasks via imitation learning. However, the variable quality of training data often constrains the performance of these models. On the other hand, offline Reinforcement Learning (RL) excels at learning robust policy models from mixed-quality data. In this paper, we introduce Reinforced robot GPT (ReinboT), a novel end-to-end VLA model that integrates the RL principle of maximizing cumulative reward. ReinboT achieves a deeper understanding of the data quality distribution by predicting dense returns that capture the nuances of manipulation tasks. The dense return prediction capability enables the robot to generate more robust decision-making actions, oriented towards maximizing future benefits. Extensive experiments show that ReinboT achieves state-of-the-art performance on the CALVIN mixed-quality dataset and exhibits superior few-shot learning and out-of-distribution generalization capabilities in real-world tasks.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhang et al. (2025) studied this question.

synapsesocial.com/papers/69473b64db9c958d0dfca888https://doi.org/10.48550/arxiv.2505.07395
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning2025
  2. 2A Vision-Language-Action-Critic Model for Robotic Real-World Reinforcement Learning2025
  3. 3Sample-Efficient Robot Skill Learning for Construction Tasks: Benchmarking Hierarchical Reinforcement Learning and Vision-Language-Action Model2026
  4. 4VLA-R1: Enhancing Reasoning in Vision-Language-Action Models2025
  5. 5Affordance-Guided Reinforcement Learning via Visual Prompting2024