Research agents trained with reinforcement learning
A paper dated September 9, 2025 describes autonomous agents for deep research using web search, browsing, Python, and summaries of earlier results when context is limited.
Published on September 9, 2025, the SFR-DeepResearch paper presents autonomous agents trained with reinforcement learning for deep research tasks. They use web search, browsing, and Python, and can summarize earlier results when available context becomes limited.
The work describes an approach to tool use, memory for long tasks, and end-to-end reinforcement learning in research agents. To check its scope and details, consult the original paper and compare its claims with the text and results presented by the authors; the available summary does not specify metrics or quantitative findings.