PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 16, 20251 citationsOpen Access

DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments

View Full Paper
YZYu-Xiang ZhengDFDayuan FuXHXiangkun Hu

Key Points

  • DeepResearcher shows up to 28.9 points improvement over traditional prompt engineering methods, enhancing research efficiency.
  • The innovative framework utilizes reinforcement learning in unstructured environments, representing a significant advancement in deep research agent development.
  • Extensive experiments validate the framework's performance, revealing the importance of real-world web training for achieving robust research capabilities.
  • Cognitive behaviors, such as information cross-validation and self-reflection, emerge as key aspects of enhanced deep research agents.

Abstract

Large Language Models (LLMs) equipped with web search capabilities have demonstrated impressive potential for deep research tasks. However, current approaches predominantly rely on either manually engineered prompts (prompt engineering-based) with brittle performance or reinforcement learning within controlled Retrieval-Augmented Generation (RAG) environments (RAG-based) that fail to capture the complexities of real-world interaction. In this paper, we introduce DeepResearcher, the first comprehensive framework for end-to-end training of LLM-based deep research agents through scaling reinforcement learning (RL) in real-world environments with authentic web search interactions. Unlike RAG-based approaches that assume all necessary information exists within a fixed corpus, our method trains agents to navigate the noisy, unstructured, and dynamic nature of the open web. We implement a specialized multi-agent architecture where browsing agents extract relevant information from various webpage structures and overcoming significant technical challenges. Extensive experiments on open-domain research tasks demonstrate that DeepResearcher achieves substantial improvements of up to 28.9 points over prompt engineering-based baselines and up to 7.2 points over RAG-based RL agents. Our qualitative analysis reveals emergent cognitive behaviors from end-to-end RL training, including the ability to formulate plans, cross-validate information from multiple sources, engage in self-reflection to redirect research, and maintain honesty when unable to find definitive answers. Our results highlight that end-to-end training in real-world web environments is not merely an implementation detail but a fundamental requirement for developing robust research capabilities aligned with real-world applications. We release DeepResearcher at https://github.com/GAIR-NLP/DeepResearcher.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zheng et al. (2025) studied this question.

synapsesocial.com/papers/68f147cc724575985c3fcfb0https://doi.org/10.48550/arxiv.2504.03160
Ask AI
Helpful
Bookmark
Share
View Full Paper