What question did this study set out to answer?

The aim is to rigorously evaluate the capabilities of multitask robot manipulation policies known as large behavior models.

April 17, 2026

A careful examination of large behavior models for multitask dexterous manipulation

Key Points

The aim is to rigorously evaluate the capabilities of multitask robot manipulation policies known as large behavior models.
Extended the diffusion policy paradigm across simulated and real-world data.
Developed an evaluation pipeline to analyze model capabilities with statistical confidence.
Conducted blind, randomized trials comparing multitask and single-task baselines in controlled settings.
Multitask pretraining increased policy success and robustness.
Teaching complex tasks became faster using less data than single-task models.
Performance improved with greater pretraining scale and diversity.

Abstract

Robot manipulation has seen tremendous progress in recent years, with imitation learning policies enabling successful performance of dexterous and hard-to-model tasks. Concurrently, scaling data and model size has led to the development of capable language and vision foundation models, motivating large-scale efforts to create general-purpose robot foundation models. Although these models have garnered considerable enthusiasm and investment, meaningful evaluation of real-world performance remains a challenge, limiting the pace of development and inhibiting a nuanced understanding of current capabilities. Here, we rigorously evaluated multitask robot manipulation policies, referred to as large behavior models, by extending the diffusion policy paradigm across a corpus of simulated and real-world robot data. We proposed and validated an evaluation pipeline to rigorously analyze the capabilities of these models with statistical confidence. We compared against single-task baselines through blind, randomized trials in a controlled setting, using both simulation and real-world experiments. We found that multitask pretraining made the policies more successful and robust and enabled teaching complex new tasks more quickly, using a fraction of the data when compared with single-task baselines. Moreover, performance predictably increased as pretraining scale and diversity grows.

Connected Papers

Building similarity graph...

Analyzing shared references across papers

Discussion

Authors

Jose Barreiros

Andrew Beaulieu

Aditya Bhat

Journals

Science Robotics

Actions

Institutions

Massachusetts Institute of Technology

Cornell University

Toyota Research Institute

References and Citations

Connected Papers

Building similarity graph...

Analyzing shared references across papers

A careful examination of large behavior models for multitask dexterous manipulation

Key Points

Abstract

Citation Network

Connected Papers

Discussion

Authors

Journals

Actions

Institutions

References and Citations

Citation Network

Connected Papers

Discussion

Cite this study

Also consider