Skip to content
Rota Nacional

Radar ·

RLVMR combines rewards with step supervision

In a publication dated August 15, 2025, Tencent researchers proposed RLVMR, a method combining outcome rewards with supervision of agent steps. The text reports 83.6% on the hardest unseen tasks in ALFWorld and ScienceWorld for a 7B agent.

On August 15, 2025, Tencent researchers presented RLVMR, a proposed agent-training method that combines rewards for final outcomes with supervision of steps such as planning, exploration, and reflection. The idea is to guide not only what an agent achieves, but also parts of the process leading to the result.

According to the published summary, a 7B agent reportedly scored 83.6% on the hardest unseen tasks in ALFWorld and ScienceWorld. The text suggests step rewards may help reduce redundant actions and support recovery from errors; however, it does not provide enough detail to assess the protocol or compare the result. Consult the original publication and check its metrics, test conditions, and task definitions.

Get new articles

Privacy, AI engineering and security in your inbox.

Rota Nacional

Bring privacy into your workflow.

30 days, no card, with a starting quota. After that, Pix credit from R$ 5,00.

Try free